15a3728ae9e53c66576ff2ac91363c6ea33e5516
132 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
15a3728ae9 |
feat: bake shellcheck and add a client-side pre-push lint gate
v1.8.14 made shell lint a RELEASE gate (scripts/lint-shell.sh, shared by lint.yml
and the new lint-gate job that resolve-versions depends on), and that script
correctly exits 2 when shellcheck is absent -- "a gate that cannot run must not
pass". Measured on v1.8.14 on 2026-09-09 by three routes (command -v, dpkg -l, a
filesystem search): shellcheck was NOT IN THE IMAGE AT ALL. So the gate could not
be run by a developer in any container, only in CI, and the loop stayed
write-shell -> push -> wait for CI -> discover. That is the loop the gate was
added to shorten, after v1.8.14's first attempt burned ~46 min on a tree whose
lint had already been red for 24 hours.
shellcheck 0.10.0-1 added to the Dockerfile.base apt block: ~39 MB installed
(Installed-Size 40112 KB), measured to pull ZERO additional packages under
--no-install-recommends because libc6/libffi8/libgmp10 are already present.
NOTE this forces one full base rebuild -- base-decide hashes Dockerfile.base +
rootfs/, so unlike a scripts/ change it cannot reuse the existing base- layer.
hooks/pre-push is opt-in per clone (git config core.hooksPath hooks), bypassable
with --no-verify, and execs scripts/lint-shell.sh rather than reimplementing it
-- one copy, because a duplicated check that drifts is the failure this repo
keeps paying for. Matches the idiom skillset/ and myconfigs/ already use.
WHY THIS REPO HAD NO HOOKS, since it was reported as drift and is not: a peer
asked tor-ms22 for core.hooksPath per clone on the premise that unset meant the
gates were unverified there. Measured: pi-devbox unset, skillset hooks, myconfigs
common/hooks, pi-toolkit unset -- but `git ls-files | grep -i hook` is EMPTY in
both pi-devbox and pi-toolkit, so there was nothing to point at on any machine
and unset was the only correct value. This closes the real half for pi-devbox;
pi-toolkit still ships none.
Verified, expected result written down before each check:
* refusal paths -- shellcheck absent => rc 2 with the remedy named; linter
missing => rc 2. Never waved through on the assumption CI will catch it.
* the hook is IN the scan set -- "Checking 14 shell file(s)" with it present,
13 with it moved aside, so the extensionless file is found by the shebang
half of the linter's two-signal union. This check exists because the first
attempt was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not reported,
which could equally have meant "not scanned" or "below -S error". It was the
latter. A count that moves is unambiguous; a clean run is not.
* it catches the REAL v1.8.14 defect -- planting `echo 'the fleet\'s thing'`
in hooks/pre-push yields SC1073/SC1072 at severity error, rc=1.
And the gate earned its keep inside this commit: the first version of the
smoke-test assertion carried a comment beginning "# shellcheck is a GATE
DEPENDENCY", and a comment whose first word is the tool's name is parsed as a
DIRECTIVE, not a comment. The new gate failed it with SC1073/SC1072 before the
push -- same family as the v1.8.14 apostrophe, a line that reads as prose to a
human and as syntax to the parser.
|
||
|
|
361babd4fd |
ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose own lint had been failing for 24 hours. shellcheck had already flagged the defect (SC2289, severity error) on the push that introduced it; the lint workflow went red at run 186 and nobody read it. lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged tree was already linted on main, and a tag-ref lint run sorts above the publish run, making a release look finished before anything ships. The missing invariant was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and only a job inside the publish workflow can enforce that. So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and call it from both places, then add a lint-gate job that resolve-versions depends on. resolve-versions is the graph root, so gating it gates everything. Cost is ~40 s at the front of a release; the alternative already cost fifty minutes. Extracted rather than copied on purpose. A second copy of a check is the drift this repo keeps paying for -- the same evening produced a skillset mirror that had sat 9579 B behind its upstream through two consecutive edits. The script adds one behaviour the inline version lacked: if shellcheck is not installed it exits 2 rather than silently finding nothing, inheriting the existing "a gate that cannot run must not pass" rule from hooks/pre-commit in the skillset repo. Without that, reordering the install step away would turn the gate into a green tick over zero checks. Verified locally with a stubbed shellcheck (the real binary is not in the devbox), five cases, each with its expectation stated first: absent shellcheck -> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half, naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own number: the inline version reported 12 files, the extracted one reports 13, the difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an assertion that resolve-versions needs lint-gate, and the repo's check-workflow-shell.sh guard still passes. |
||
|
|
70e675afee |
fix(smoke): keep prose out of the single-quoted exec_test body
The agent-browser execution guard added on 2026-09-07 carried its explanation
INSIDE the single-quoted script body, and the explanation contained an
apostrophe ("the fleet\'s only recurring amd64 runtime proof"). Inside '...'
bash treats a backslash as literal, so \' does not escape the quote -- it CLOSES
the string. The body truncated at that point and the remaining lines were parsed
by the calling shell.
Consequences, both measured rather than inferred:
- exec_test received 12 arguments instead of 2 (verified two-sided: the fixed
tree yields argc=2, HEAD yields argc=12).
- the leaked `v=$(agent-browser --version)` ran on the CI RUNNER instead of
inside the image. The runner has no agent-browser, so smoke and
smoke-studio both failed with "line 770: command not found" after
build-base had already spent ~46 minutes. Every downstream job was skipped.
- the truncated body still passed inside the container and printed its green
tick first, so the log shows a PASS immediately followed by the failure --
the tick was real, it just no longer covered the assertion.
The prose now sits above the exec_test call, where an apostrophe cannot
terminate anything, and a comment at that spot records why it must stay there.
Not a new failure class: shellcheck flagged it as SC2289 at severity error the
same day, so the lint job has been red since run 186 (2026-09-07 21:21) and was
not read. The gate did its job; nobody looked.
|
||
|
|
601fc98a49 |
docs(changelog): release v1.8.14
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 14s
Lint / actionlint (push) Failing after 24s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 50m32s
Publish Docker Image / smoke-studio (push) Failing after 5m13s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 7m40s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Converts the Unreleased section and records what this build carries beyond it: the mempalace-toolkit bump that makes closing replies reach the mailbox (deriveClosed, 21023e7 -> e45f6b4), and the L0-L4 subtask documentation landing via pi-toolkit adfb553 + pi-extensions c64c122. Notes the mechanism that makes the toolkit fix land at all — the resolved toolkit SHA is folded into the content-addressed base tag, so the toolkit moving forces a base rebuild rather than waiting for one — and the consequence for the amd64 item already in this section: v1.8.13's base was cached, so Dockerfile.base:607's agent-browser assertion never ran. This base is not cached, so the native-amd64 proof is finally collected instead of discarded. |
||
|
|
6bd8b79d3a |
test(smoke): assert agent-browser EXECUTES — it was the discarded amd64 proof
Second instance of the same bug class as the node line, in the same file, found the same way. The agent-browser guard captured the version inside an echo with 2>/dev/null: echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null|head -n1)]" >&2 so the exit code was discarded and a binary that could not execute at all still PASSED, printing version=[]. Verified two-sided: a stub exiting 127 passes the old form and is caught by the new one. Why this exit code matters more than most: smoke runs platforms: linux/amd64 on an x86 runner, i.e. NATIVE amd64, so this line is the fleet's only recurring amd64 runtime proof for agent-browser's linux-x64 ELF. NO DEVBOX CAN EVER SUPPLY THAT PROOF. Every machine in the pi fleet is an Apple Silicon Mac: mbp-m1-2020; tor-ms22 = Mac Studio Mac13,1 M1 Max (fleet-ops hosts/tor-ms22.md, verified 2026-08-17 with system_profiler); emb-7kj4vr4g = Apple Silicon, verified 4 routes 2026-09-07. The open "amd64 runtime proof still needed" ask sent to two devices was asking for the impossible, and emb's reply naming tor-ms22 as "the only remaining candidate" is wrong for the same reason. CI had the answer all along and was throwing it away. Dockerfile.base:607 DOES assert it (`agent-browser --version && \`), but only when the base rebuilds, and v1.8.13's base was cached — so smoke is where the recurring gate belongs. |
||
|
|
fabf1274aa |
docs(changelog): Unreleased section for the smoke node assertion + agent-browser correction
Summarises what changed since v1.8.13: the node-major assertion (a bump would have passed the suite silently), the two-sided verification of the derivation, and the v1.8.13 agent-browser 0.35.2 -> 0.36.0 correction. No image content changes; NODE_VERSION still 22. |
||
|
|
5972a2c535 |
test+docs: assert the node major in smoke, and correct v1.8.13's agent-browser version
Two findings from a delegated read-only audit of this repo, both verified from the filesystem before patching. 1. No test asserted the node major, so a node-24 bump would have passed the smoke suite SILENTLY. scripts/smoke-test.sh:94 was a bare `run "node" "node --version"` — exit-0 and non-empty output only, the printed version compared to nothing — while the line above it uses run_expect against $EXPECTED_PI_VERSION for pi. A reader skimming the suite would reasonably assume node regressions were covered. Worse, this is where the "node v22.23.2 verified" line in the v1.8.13 recreate notes came from: printed output, not an assertion. Now gated on EXPECTED_NODE_MAJOR, which CI derives from Dockerfile.base's ARG NODE_VERSION — the single source of truth (Dockerfile.base:557 is the ONLY hard pin in the repo; Dockerfile.variant has no node install at all). That also catches a stale cached layer whose node disagrees with the declared ARG. Unset => previous behaviour, so this is backward compatible. Verified two-sided rather than assumed: the sed derivation yields 22 (empty would have silently disabled the assertion, reintroducing the bug); grep -Fq "v22." matches v22.23.2; "v24." does NOT match, so a wrong major is caught; and "v2." does not prefix-collide. Workflow YAML re-parsed after editing (9 jobs). 2. The v1.8.13 entry claimed "the image's own 0.35.2" for agent-browser. The image ships 0.36.0: /usr/lib/node_modules/agent-browser/package.json says version 0.36.0, engines.node >=24.0.0, and no 0.35.2 exists anywhere in the image. The claim was also internally incoherent, contrasting 0.36.0 against a version that is not present. Corrected in place with a visible note, since the entry is already released. The reasoning survives untouched: the engines floor really is vestigial, because /usr/bin/agent-browser is a prebuilt aarch64 ELF invoked directly and never through node — which is why 0.36.0 runs fine on 22.23.2. |
||
|
|
aa0fbc5ec0 |
fix: correct the pi-studio claim — CI publishes v0.9.59, not the v0.9.60-rc.0 label
Lint / actionlint (push) Successful in 17s
Lint / hadolint (push) Successful in 14s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 4m44s
Publish Docker Image / smoke-studio (push) Successful in 5m8s
Publish Docker Image / build-variant (push) Successful in 15m52s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / build-variant-studio (push) Successful in 21m22s
Measured at the wrong layer during the v1.8.13 audit. I read `ARG
PI_STUDIO_REF=main` in Dockerfile.variant, concluded the release would adopt
main (= v0.9.60-rc.0), set PI_STUDIO_VERSION to that, and wrote a comment plus a
CHANGELOG entry describing deliberate RC adoption. A Dockerfile default cannot
answer "what will CI publish?" when CI overrides it, and it does: build-variant
passes PI_STUDIO_REF=studio_ref and PI_STUDIO_VERSION=studio_tag (lines 598-599
and 787-788), and resolve-versions picks the newest STABLE semver tag via
`^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases.
Caught by reading run 639's own resolve-versions output rather than the
Dockerfile: studio_tag=v0.9.59, studio_ref=9eed84f = refs/tags/v0.9.59^{}, while
main/v0.9.60-rc.0 is 658536f and never gets built. So published v1.8.13 studio
images carry pi-studio v0.9.59.
ARG restored to `none` rather than pinned to v0.9.59: the local-build default
should not hardcode a tag that goes stale as soon as main moves, which is how the
previous value came to lie. The comment now leads with the override so the next
reader starts at the layer that decides. Upstream's tag-over-main policy is
deliberate (Releases stopped at v0.5.55, main receives half-finished commits), so
adopting an RC from CI would mean changing that filter, not this ARG.
Consequence kept on purpose: the RC's opt-in Studio network binding is in NO
published v1.8.13 image, so it needs no audit this release.
Doc/label-only: base_tag hashes Dockerfile.base + rootfs/** + both entrypoints +
mempalace_toolkit_ref, none of which this touches, so the in-flight base build
(base-ad9faf00f2b2) stays valid and the tag run will reuse it. Verified with CI's
pinned linters: hadolint 2.14.0 exit 0 on both Dockerfiles, actionlint 1.7.7 exit
0, shellcheck 0.10.0 -S error exit 0.
|
||
|
|
f561acc89a |
skills: refresh vendored mempalace snapshot a12fe5e -> e9e09d9, re-pin the canary
Folded into v1.8.13 at zero marginal cost: the snapshot is hashed into
base_tag, but Dockerfile.base already changed this release, so the ~67 min base
rebuild was already being paid. vendor-mempalace-skill.sh --check reported exit
0 (stale-but-truthful) beforehand, so skipping was sanctioned -- this is the
deliberate call the release checklist asks for. Upstream content: the bare
project-name wing convention and the <harness>@<device> added_by rule, both
downstream of the attribution defect measured on this device 2026-09-06.
The canary re-pin matters more than the refresh. Its old pair ("Provenance is
stamped for you" present / "Attribute what you file yourself" absent) still
PASSED against the new snapshot, so leaving it would have yielded a canary
green on both old and new bytes -- blind to exactly the refresh it exists to
witness, the same false-green family as the pre-v1.8.5 canary. New pair chosen
by measuring direction against both files rather than reading the diff
("Diaries self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1), then tested two-sided: PASS on refreshed bytes, FAIL on the old
bytes recovered from git.
Gates after the change: smoke-test.sh parses, vendor --check exit 0,
check-base-hash exit 0.
|
||
|
|
0d984b1414 | changelog: cut v1.8.13 section | ||
|
|
adcf56f829 |
release: audited bumps (pi 0.85.1, mempalace 3.9.0, atelier v0.10.1) + two guards
Version audit for the next release. pi 0.84.4 -> 0.85.1, deliberately skipping 0.85.0 (it published internal experimental code and broke SDK imports, upstream #9132). mempalace 3.8.0 -> 3.9.0. pi-atelier v0.10.0 -> v0.10.1. PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0 so the floating main ref's RC status is visible at docker-inspect time instead of discovered later. PI_FORK_REF stays floating and adopts e69725c. The pi bump was verified by running it under a pty in five combinations rather than by reading the changelog, because this repo has already shipped a version pair no changelog flagged (atelier < 0.7.1 hangs pi >= 0.84). CPU delta 0.00-0.01s over 5s against a ~5s sustained-CPU hang signature, two-sided via the atelier sidebar painting identically to the 0.84.4 control. NODE_VERSION stays 22 on purpose: node 24 is technically safe (pi's five prebuilt addons are all NAPI, nothing declares a ceiling, agent-browser's engines.node >=24 is vestigial for the shipped aarch64 ELF), but this release already moves two minors and bakes an RC, and a node major would leave four suspects if the image misbehaves. Own release, smoke suite as the gate. Also corrects a stale claim at the mempalace ARG: synlig serves 3.8.0 server-side, not 3.7.1 (measured over ssh 2026-09-06). agent-browser volume shadowing: the image has shipped 0.35.2, but every session on mbp-m1-2020 ran 0.27.0 from a 2026-07-17 hand-install in ~/.pi/npm-global (a VOLUME, at PATH position 2 vs /usr/bin at 8). Third package hit by this hazard after pi and pi-atelier, so the guard is now generalised: entrypoint-user.sh retires the copy by moving it aside (reversible, only when the image ships its own), recreate-sanity-check.sh asserts resolution under /usr where the volume is real, smoke-test.sh carries the build-time half and says in the source why it is weak. The real damage was the stale BUNDLED SKILL (3 skillsets/17.6 KB vs 8/31.5 KB, ten subcommands undocumented to the agent) - a stale tool errors, a stale skill quietly teaches wrong commands. pi-fork capability floor (extensions: []): forks were measured across four dispatches ignoring their brief, answering in the user's voice, fabricating self-referential measurements, and once filing a diary entry as agent_name=pi. Cause is upstream by design - the child gets getHeader()+getBranch(), the whole active session branch, with the brief as the final user message. Not a model-capability problem: the same model as the fast profile obeyed the identical brief perfectly with a fresh session and no inherited context. extensions: [] runs children with --no-extensions, so the mempalace bridge is absent and palace writes are impossible by construction (verified by asking a child to enumerate its tools: read, bash, edit, write). Removes palace writes, not filesystem writes. |
||
|
|
c8622ece9d |
skills: correct the credential-incident-response §5 premise about chroma metadata
§5 said embedding_metadata.string_value holds "metadata fields only". False, measured directly: chroma also stores a copy of the document text there, under key chroma:document. Confirmed with a disposable sentinel drawer (pi@tor-ms22, 2026-08-30): one row in fts_content AND one row in embedding_metadata for the same drawer. This was a real mistake in shipped guidance, not a nitpick: this section's own scanning advice was written to guard against explaining a zero with a mechanism nobody verified from source, and the section itself did exactly that -- I downgraded a census to "a floor" on the strength of a metadata-blind claim I never checked against chroma's actual storage layout. The practical scan order is unchanged (fts_content is still the direct target, raw bytes are still the backstop); only the stated REASON for a metadata zero changes: it needs a different explanation now (key filter, query shape, escaping), not "structurally absent". §6's row-gone/bytes-gone claim is upgraded from asserted to measured, same sentinel: delete_by_source took both fts_content and embedding_metadata 1->0, raw bytes stayed 4->4 (freed pages persist until VACUUM). Also records the method that unblocked the measurement: not a better instrument, a disposable sentinel drawer instead of risking real fleet data. No image behaviour changes. |
||
|
|
05843ecfae |
changelog: reopen an Unreleased section after v1.8.12
v1.8.12's release retitled the previous Unreleased heading, leaving the file with no place to put the next change — so the next contributor either invents a heading or appends to a released section. The note under it points at the release checklist step that renames it, so the convention is discoverable from the file rather than only from AGENTS.md. Also the first push after moving CI off synlig: lint.yml should now run on runner-a1 (8 vCPU / 16 GB, Debian 13, upstream Docker CE) instead of the box that hosts the palace. |
||
|
|
a2846a5f7e |
release: adopt pi 0.84.4 + pi-atelier v0.10.0, and fix the doc claim the pi bump invalidates
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Successful in 59m39s
Publish Docker Image / smoke-studio (push) Successful in 5m22s
Publish Docker Image / smoke (push) Successful in 17m53s
Publish Docker Image / build-variant-studio (push) Successful in 17m14s
Publish Docker Image / build-variant (push) Successful in 18m22s
Publish Docker Image / promote-base-latest (push) Successful in 12s
Publish Docker Image / update-description (push) Successful in 20s
pi 0.84.3 -> 0.84.4 (no Breaking Changes / Removed heading in that section,
grepped). Adopted for three fixes that land on machinery this fleet runs:
#6879 (large tool results crossing the auto-compaction threshold were sent to
the provider before compacting), #8345 (a resumed session corrupted its next
appended entry when the JSONL lacked a trailing newline -- that file is the
memory feeder's input; measured 49/49 clean here beforehand), and #8537
(triggerTurn:false messages sent mid-run were inserted between a tool call and
its result). The mempalace mailbox is outside #8537's precondition: it delivers
at agent_settled with deliverAs:"steer" and no triggerTurn, and 0.84.4 leaves
the documented steer semantics unchanged.
pi-atelier v0.8.2 -> v0.10.0: two minor releases, both UI-only, no BREAKING
notice. v0.9.0 raises its minimum pi to 0.84.0 and, unlike the
0.7.1-under-pi-0.84 startup-hang precedent, encodes it in peerDependencies
(>=0.84.0). Satisfied by PI_VERSION=0.84.4. Both executable floors compare with
sort -V, so 0.10.0 >= 0.7.1 evaluates correctly.
docs/observational-memory.md: pi's own compaction.md gained one paragraph in
0.84.4 -- autoCompact is now also checked mid-run, after a tool batch's results
are appended. Our text said compaction is checked only when pi goes idle and so
"never interrupts a turn"; that was only ever true of the OM trigger. The
section now states both entry points into session_before_compact and the
diagram carries the second edge (mermaid checker re-run: 6 blocks, 44 labels,
0 soft-wrapped, no cut glyphs at 1280px and 800px).
README: the version-pin table had been wrong since v1.8.6 --
|
||
|
|
58c22afb04 |
skills: the fingerprint advice was missing its precondition, and the skill had no section on proving absence
Docs only; no image behaviour changes. WHY THIS AND NOT A PRIVATE NOTE. pi@emb-7kj4vr4g reported itself for printing sha256[:8] fingerprints of GIT_USER_EMAIL, GIT_USER_NAME and HOST_SSH_USER, and wrote a private rule forbidding it. It had not broken a rule. It followed §2 of this skill as written, and §2 is incomplete: it says a fingerprint lets you compare a credential "without ever materialising the secret" with no condition attached. When two agents independently make the same mistake, the artifact that taught them both is the bug. §2 NOW CARRIES THE PRECONDITION. A fingerprint is 32 bits over its INPUT SPACE, so publishing fp8(x) hands anyone a MEMBERSHIP ORACLE: they can test x == v for every candidate v they can generate. Safe for a 40-char random token; a wordlist for a hostname, username, e-mail, port, path, commit SHA or weak password. "High entropy" is the usual sufficient condition, NOT the test — a commit SHA is 160-bit and still fully enumerable from the repo. Operationally: if you can imagine writing the wordlist, you cannot publish the fingerprint. Also added: candidate fingerprints are working memory and never output (an extractor hashes hostnames and paths too, so the tempting "print what the scanner saw" debug step leaks low-entropy fingerprints wholesale), and a plain statement that a fingerprint register is a CONFIRMATION ORACLE for anyone already holding a candidate corpus — which is exactly how a retired token is identified in old transcripts, and works identically for someone else holding those same files. NEW §6, "Proving absence: instrument strength, and four ways a scan lies clean", placed next to §5 on purpose: §5 optimises against false POSITIVES, and every failure in §6 is a false NEGATIVE. Triage optimises precision, a gate optimises recall, and conflating them is what produced three clean reports over secrets that were really there. Contents: instrument ranking (exact-byte value search > class/structure pass > fingerprint census) with the instruction to state which one produced your zero; census vs class passes as different questions, both failure modes measured on this fleet; the tokenisation trap where quoting alone decided detectability; scan the index or pushed tree, never the working tree; git filters never run on symlinks while check-attr claims they do; two-sided self-tests that abort, incl. the fixture-interaction artifact; row-gone is not bytes-gone. Attribution kept per finding: the census/class split and the instrument ranking's provenance are pi@emb-7kj4vr4g's; exact-byte search over index blobs is pi@tor-ms22's. The credential sense of "census" originated in this skill, not with either agent. TRAP FOR THE NEXT EDITOR, also in the CHANGELOG: the frontmatter description is now 1022 of 1024 characters. Trim before adding, or the skill silently fails to load. Verified by parsing the frontmatter (1022 chars, name intact, every prior trigger phrase retained). Deployment: baked skill -> needs an image rebuild AND a container recreate to reach a running container. |
||
|
|
30094782df |
shell: source cli_utils' functions, and install the iproute2 that one of them needs
v1.8.11 linked cli_utils' bin/ COMMANDS onto PATH and stopped there. Nothing ever
sourced cli_utils.sh, so its 14 FUNCTIONS were missing from every interactive shell
whose $HOME has no zsh rc -- which is the normal case, not an edge case: the
container's interactive shell is bash and zsh is not installed in the image. A
symlink cannot carry a shell function and a function cannot be reached from a
non-interactive shell, so the two mechanisms are disjoint and both are required.
The image was already paying this layer's dependency cost (fzf, bat, fd, rg, jq are
baked partly FOR these functions) while delivering none of its benefit.
The two changes ship together because they are coupled: portcheck is one of the 14,
and it was a hard stub in every image up to v1.8.11 -- neither ss nor ip nor lsof
nor netstat was present, so it printed "portcheck requires at least one of: ss,
lsof, netstat" and exited. Wiring the functions in without iproute2 would have
shipped a visibly broken one.
MEASURED, not assumed:
- the loader is bash-safe despite the *.zsh filenames: `bash --noprofile --norc`
exits 0, defines all 14, and they run (pathls, mkcd, up, extract, agents-sync,
fhist verified). The tree's one zsh-only construct (print -z in fzf/fhist.zsh)
is already guarded by [[ -n $ZSH_VERSION ]] with a bash fallback.
- fresh-$HOME seeding resolves 14/14; CLI_UTILS_SOURCE=0 is honoured; an absent
checkout is a genuinely silent no-op (no output, no leaked _cu).
- interactive shell startup 12 ms -> 17 ms.
- iproute2 is ~5.5 MB (4.2 MB itself + 6 libs under --no-install-recommends;
libpam-cap is a Recommends and correctly dropped). ss lands at /usr/bin/ss,
ip at /usr/sbin/ip, both already on the developer PATH, and `portcheck --all`
then correctly identifies the socat listener on 8765.
- hadolint clean on both Dockerfiles; repo-wide shellcheck -S error and bash -n
clean. .bash_aliases is outside CI's discovery (no shebang, not *.sh), so it
was checked by hand with -s bash at error AND warning level.
Named explicitly per this repo's floating-ref rule: /workspace/cli_utils is a HOST
BIND MOUNT, not a pinned ref, so the image now executes unpinned content in every
interactive shell. Errors are left visible rather than sent to /dev/null so that a
future zsh-only file in that repo is diagnosable rather than mysterious, and
CLI_UTILS_SOURCE=0 is the documented escape hatch. It is deliberately independent
of CLI_UTILS_LINK=0: the two disable independent mechanisms.
Deployment: needs a rebuild AND a recreate. $HOME is the container's writable layer
rather than a named volume (verified -- ~/.bash_aliases carries the container start
mtime while ~/.bashrc carries the image's), so the skel file is re-seeded on every
recreate; a host-bind-mounted ~/.bash_aliases is still never overwritten.
|
||
|
|
d9a7fe101b |
changelog: an Unreleased section for a rule that was already there
Records the two skill commits ahead of tomorrow's build, and states the finding that shaped them: the "a negative result is usually your own filter" rule was already baked, already symlinked in at every container start, and already survived every recreate — then was violated five times by a session that had it available. The gap was activation, not persistence, which is why the cross-cutting form went into the always-appended AGENTS block instead of into a skill that only loads when a task description matches. Also notes what the entry's own subject implies for the reader: neither change reaches a running container until the image is rebuilt AND the container recreated, since ~/.agents/skills and the global AGENTS.md both live in the image rather than in a volume or a mount. |
||
|
|
b615571913 |
changelog: release v1.8.11 — the mail nobody could deliver, and the ship that skipped a file
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 10s
Publish Docker Image / build-base (push) Successful in 53m20s
Publish Docker Image / smoke (push) Successful in 4m48s
Publish Docker Image / smoke-studio (push) Successful in 4m58s
Publish Docker Image / build-variant-studio (push) Successful in 17m0s
Publish Docker Image / build-variant (push) Successful in 26m43s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 24s
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:
- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
(a real behaviour change to every remote-mode client) and RFC 003
§7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
the mempalace skill's from_agent identity rule, plus the vendored
fallback snapshot re-pinned to match (this pi-devbox commit).
Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
|
||
|
|
45850bc973 |
entrypoint: put back the shell state a recreate eats
cli_utils' install.sh reaches PATH by symlinking bin/ into ~/.local/bin. That is persistent on a host and ephemeral in a container, so the same installer produced opposite durability and every --force-recreate sent the human back to typing /workspace/cli_utils/bin/git-status-all. Re-link at start instead, and add a per-device boot hook so the next question of this shape needs no image change at all. Symlinks rather than a PATH edit in an rc file, deliberately: ~/.local/bin is already ahead of /usr/local/bin in ENV PATH, so links resolve in NON-interactive shells too (docker exec, agent tool shells, scripts). An rc-file PATH edit cannot reach those because ~/.bashrc returns early when not interactive — measured, that asymmetry is exactly why `command -v git-status-all` failed in one shell and worked in another on the same box. Shell FUNCTIONS remain the sourced file's job; a symlink cannot carry them. Guards, because ~/.local/bin is shared: a real file is never clobbered, a symlink pointing elsewhere is never stolen, ours are refreshed, and links into a cli_utils/bin whose target vanished are pruned — a dangling link on PATH reads as a broken container rather than a removed script. The hook (~/.config/devbox-shell/init.sh) adds no trust boundary: that dir is already sourced into every interactive shell by the baked bash_aliases, so it is already arbitrary code from the same owner. Only WHEN it runs is new. bash <file>, never sourced, exit status ignored, output to a log. Caught before commit, and the reason the loops use `if` bodies instead of `&&` chains: under `set -euo pipefail` a for-loop whose last command is a false test exits non-zero, and with no match the /workspace/*/cli_utils glob stays literal — so the first draft would have failed to START a container on every machine that does not have this repo, rather than merely skipping the links. Re-tested with set -e in place: no-cli_utils/empty-HOME no-op, guards, idempotence, CLI_UTILS_LINK=0, a hook that exits 7, and a hook that tries to mutate CLI_UTILS_BIN — all exit 0 with intact state. Not covered by CI: docker-publish.yml runs only on tags and lint.yml lints workflow run: steps, so neither executes this file. Validated by extracting both sections and running them against fixtures, then for real in a live v1.8.10 container. No smoke assertion added on purpose — the positive path needs a /workspace mount smoke does not have, and asserting it there would repeat the v1.8.0 mistake of a smoke check written against a stage that does not exist at run time. Moves the base hash (base-decide folds `cat entrypoint.sh entrypoint-user.sh`), so this rides along with the next tag's ~40-minute base rebuild rather than justifying a tag of its own. |
||
|
|
6891dc32b8 |
changelog: release v1.8.10 — deploy the scrubber, and name the check that must not be skipped
Publish Docker Image / build-variant-studio (push) Successful in 17m4s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 14s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-variant (push) Successful in 18m45s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / smoke (push) Successful in 4m55s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Audited every component against upstream rather than assuming. Only two moved since v1.8.9: - mempalace-toolkit 5b8d78f -> b2b50af (ours): feeder scrubber, the symlink fix that stops it refusing to stage, hlc owed-set join, queued-delivery note, explicit MEMPALACE_MAILBOX_NOTIFY protocol modes, RFC 003 + fleet-memory docs. - pi-studio v0.9.48 -> v0.9.52 (upstream, studio variant only): 22 commits, all additive — PDF previews, hideable header, contextual side questions. No removals or renames. Its pi floor is >=0.84.3 against our exact 0.84.3 pin: satisfied, zero headroom, named as a watch item because the next floor bump breaks the studio job only, after core has already published. Verified unchanged, with dates proving they predate the v1.8.9 bake: pi 0.84.3 (= npm latest), mempalace 3.8.0 (= PyPI latest), pi-atelier v0.8.2, playwright 1.62.1 (no drift this cycle despite floating on latest), pi-fork bf702b4, pi-obsmem ce9fc98, pi-toolkit 0e1369e, pi-extensions 2022887. No breaking changes anywhere, so everything can ship in one tag. The reason for tagging now is not features. The scrubber has been live on exactly ONE device since this morning; every other device has kept staging unscrubbed transcripts into a shared palace, and cleanup after the fact is manual redaction (done twice today, with freed-page residue left behind by choice). The entry leads with the acceptance check instead of burying it, because this release's worst failure is silent: the feeder is fail-closed, so a packaging or path mistake stops the fleet's memory feed and nothing complains — refusing to stage looks exactly like a quiet session. f0bffd1 exists because that very bug was real (BASH_SOURCE reports the symlink path, not the target). First client on the new image must confirm a [scrub] summary line appears, the drawer count moves, and exit 3 did not fire. Silence is the failure signal, not success. |
||
|
|
8a673ec143 |
docs: unclip the diagrams, and answer what compaction leaves behind
Two problems reported against docs/observational-memory.md, one cosmetic and one
substantive. Both turned out to be worth more than the fix.
Clipping. Several boxes lost their bottom line of text in the viewer, and neither
the source nor my own render showed it. Cause: Mermaid measures a node label with
its own font metrics, commits to a box size, then renders that label as real HTML
in a <foreignObject> — so any host stylesheet touching line-height or font-size
inflates the text past a box that is already fixed, and the overflow is clipped.
The error accumulates per line, which is why it always eats the last line of the
tallest labels.
Rejected the obvious fix after testing it rather than assuming it:
%%{init: {'flowchart': {'htmlLabels': false}}}%% *is* honoured (labels switch from
16 foreignObject to 7 tspan) and clips identically, because inflated font-size
inherits into SVG text too. The fix that works is a hard limit of two short lines
per node, with the detail moved into prose under each diagram — one- and two-line
boxes have the vertical slack to absorb inflation, three- and four-line boxes do
not. The nodes were carrying paragraph-sized text; the diagrams are better for
losing it.
Regression harness: render every block with line-height: 1.7 !important forced
onto the label HTML, screenshot, read it. That caught two survivors of the rewrite
that looked fine in the clean render — a long unbreakable /opt path wrapping to a
third line, and a cylinder shape whose curved bottom leaves less room than a
rectangle for the same two lines.
New §4, because the document explained that compaction folds the ledger and never
said what that leaves in the context. Asked directly: is the session back to
knowing nothing? No. A verbatim tail survives, sized by keepRecentTokens (20k) and
cut only at turn boundaries; the system prompt and AGENTS.md were never in the
compacted region because they are rebuilt from disk each request; nothing is
deleted from disk, since compaction appends a compaction entry rather than
rewriting lines; and recall keeps resolving ids whose sources left the context
because it reads the full branch via sessionManager.getBranch() and never consults
the context window. Repeated compaction renders from live records, not from the
previous summary's prose, so there is no generation-loss spiral.
Also corrects this repo's own "compaction calls no model" to the steady-state
claim it actually is: with an empty ledger the hook returns nothing and explicitly
declines ownership, and pi's native model-based summariser runs. Snippet quoted in
the doc.
And a bug shipped in the first version: §9 said the ledger entries are
custom_message and specifically not custom. Exactly backwards, so the one grep
that section existed to get right was the one it got wrong. Verified empirically
against the live session file — 11 om.observations.recorded and 6
om.reflections.recorded, all "type":"custom", next to "type":"custom_message"
entries whose customType is mempalace-mailbox and mempalace-wakeup. That is where
the confusion came from, and the distinction is load-bearing rather than
cosmetic: the mailbox uses the context-visible append API, om's ledger uses the
invisible one. Which makes it a feature the doc now advertises — the ledger costs
zero context until it is folded.
|
||
|
|
cdb6fc0950 |
changelog: name what the floating toolkit ref will pull into the next tag
MEMPALACE_TOOLKIT_REF=main floats and docker-publish.yml resolves it to a SHA at build time, so toolkit main moving 5b8d78f -> f0bffd1 (10 commits) ships in the next tagged image whether or not this repo has a commit. v1.8.9 adopted the rule after that shape bit twice (553d865, 5b8d78f): name the behaviour change BEFORE tagging. This is that rule obeyed rather than re-learned — the work was pushed to toolkit main earlier today and this entry was missing, which is exactly the gap that caused a cross-host misattribution in v1.8.7. Contents: the feeder-side secret scrubber (three tiers, T3 report-only after a measured 403 -> 29 false-positive calibration on 52 MB of real transcripts, fail closed); the symlink near-miss that fail-closed would have turned into a fleet-wide silent memory outage at bake time, caught before tagging; the mailbox work (hlc owed-set join, queued-until-next-turn note, explicit notify protocol modes, tmux path documented unverified); and the documentation set (RFC 003, fleet-memory.md, secret-hygiene.md). Also states what remains unscrubbed: the opencode bridge write path and the unbuilt server-side layer. No tag pushed — per the release protocol, no tag means no build. |
||
|
|
14371e2da6 |
docs: explain the memory that runs itself, and fix a claim the mailbox falsified
`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.
Scoped by who is authoritative, so there is one copy of each claim:
- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
section contrasting it with MemPalace, on the line "observational memory keeps
a session coherent, the palace keeps the fleet coherent".
Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.
Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.
Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since
|
||
|
|
aac4a1c323 |
release: v1.8.9 — the version flag that blamed the wrong component
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 4m51s
Publish Docker Image / smoke-studio (push) Successful in 4m59s
Publish Docker Image / build-variant-studio (push) Successful in 16m58s
Publish Docker Image / build-variant (push) Successful in 28m35s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 17s
Two versions, two flags. `--expected-version` has only ever asserted
`pi --version`, but AGENTS.md step 4 spelled it `X.Y.Z` inside a checklist where
every other X.Y.Z is the pi-devbox tag. Run as documented for v1.8.8 the final
runtime gate of the release printed
✗ pi version mismatch: expected 1.8.8, got 0.84.3
and exited 1 — a red accusing the image of being the wrong version. Not one
reader's slip: the v1.8.8 release-readiness handoff from pi@emb-7kj4vr4g
propagated the same wrong spelling twice while correctly calling step 4 "not
ceremonial", so two independent readers converged on it. README.md had it right
all along, which means the two documents disagreed.
- new --expected-image-version asserts the pi-devbox release tag, read from
release_tag in /etc/pi-devbox/build-manifest.json (no checkout, no network);
leading `v` optional on either side
- both flags detect being handed the other one's value, and the test is exact
rather than heuristic: the value is compared against the other quantity the
image itself reports, so it can only fire on a real mix-up
- neither flag is required now. With none, live `pi --version` is asserted
against the manifest's pi_version — not a tautology, since a stale pi in the
~/.pi/npm-global volume can shadow the baked one, exactly as a stale
npm:pi-atelier can in packages[]
- the header note replaced was stale and load-bearing: it claimed pi is resolved
from 'latest' and cannot be self-derived, while Dockerfile.variant pins
ARG PI_VERSION=0.84.3 and docker-publish.yml reads that ARG as its source of
truth. The same withdrawn claim also sat in cli_utils' pi-devbox-sanity --help
- argument parsing: a missing value, or a value that is another flag, is a usage
error instead of silently consuming the next argument; --help works
All fourteen flag combinations exercised by execution, including the two
manifest-absent branches and the shadowing branch a healthy container cannot
reach — mutation-tested with a doctored manifest so each failure branch was
observed firing rather than assumed present.
CHANGELOG also names what no commit here causes: mempalace-toolkit main moved
e70bef2 -> 5b8d78f, so this tag ships the auto-delivered logstream mailbox
because base_tag folds the resolved toolkit SHA. It would have landed either
way; going unnamed is the 553d865 shape that already caused one cross-host
misattribution. Component audit found nothing else to bump — pi, mempalace,
pi-atelier all equal their upstream latest, and every other floating ref
resolves to the commit already baked.
|
||
|
|
34cf1e3810 |
release: v1.8.8, and a notice that named the wrong remedy
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 19s
Publish Docker Image / build-base (push) Successful in 1h3m31s
Publish Docker Image / smoke (push) Successful in 4m47s
Publish Docker Image / smoke-studio (push) Successful in 8m8s
Publish Docker Image / build-variant-studio (push) Successful in 20m4s
Publish Docker Image / build-variant (push) Successful in 26m9s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 14s
Freezes the v1.8.8 section and clears the two non-code checklist items
pi@emb-7kj4vr4g handed over (evt_20260826T134919_a614ecfc2d4f), plus the two
carried nits from its round-2 verification (evt_20260826T133356_d56792791a49).
Every claim below was re-measured here rather than taken from the handoff.
THE STALENESS NOTICE ASSERTED A DIRECTION IT NEVER TESTED — Blocker 1's shape,
one layer down, in the message I added to replace the message that named the
wrong cause. The notice fired on "recorded != HEAD" and then announced HEAD as
the newer side without testing ancestry, so a clone that was merely BEHIND got
"has moved to 82a8d3c; the snapshot describes the older 5fd0d5c" when 82a8d3c is
5fd0d5c's ANCESTOR. Found by EMB against the real state of its own host, not a
fabrication. The verdict was never wrong (rc 0, nothing mis-verified) but the
remedy it implies is a ~67-minute base rebuild when the actual fix is `git pull`
— the only one of the two carried nits with a price tag, which is why it went
first. Now tests ancestry with the merge-base --is-ancestor primitive the refresh
path 60 lines below already used, and reports three verdicts: stale (refresh),
clone behind (pull, do NOT refresh), diverged (reconcile). All three verified by
execution; only the first was correct before. --help no longer errors on the one
script whose argument order was itself a landmine. The `-s "$VENDORED"` guard is
now commented as load-bearing: it makes the empty-stdin collision unreachable by
construction, which also means no test below exercises it any more, so deleting
it as "redundant with the probes" would silently restore the false OK.
CHANGELOG: retitled, and three stale spots fixed in what becomes the permanent
record. It cited the skill at 82a8d3c (twice superseded); it RE-ASSERTED the
retracted mailbox measurement as live evidence 200 lines after withdrawing it,
which is the exact non-contradiction failure this release exists to fix; and its
warning block still described the pre-e8ddeaf world ("still records c04cd15",
"now exits 1") while quoting as exemplary the very notice whose direction was
unverified. Per EMB's steer the conclusion was kept and only the evidence
replaced: the status filter does drop broadcast noise, it just never computed
owed-ness. Honest replacement, measured on both machines: raw filter returns 2
here and 1 there, EVERY ONE already answered, derivation returns 0 for both.
SNAPSHOT RESYNCED AGAIN, 5fd0d5c -> 6eb20af, because skillset 6eb20af adds the
limit of my own seq ordering test: seq is REPLICA-LOCAL, equal to origin_seq only
because one replica authors for all four machines, so use hlc once mesh_peers
reports a peer. Recorded as reasoning not measurement — a second replica cannot
be stood up here. The durable half is the asymmetry: seq skew makes an ANSWERED
item resurface (noise, visible, self-correcting) while created_at SUPPRESSES AN
UNANSWERED ask forever (silent, permanent), so the skill now says outright that
"fixing" a resurfacing item with a timestamp trades the safe failure for the
dangerous one. The resync was free: rootfs/ was already changing, so the base
rebuild was forced regardless — the ordering warning about accidental staleness
does not apply to a deliberate refresh before the tag.
--check is a clean OK at 6eb20af with no notice, canary re-verified bidirectionally
(present 3, withdrawn 0), baked snapshot 0644, tree hash recomputed at build time
and re-verified in-container. bash -n clean; shellcheck/hadolint/actionlint remain
absent locally, so CI is still the only evidence for those.
|
||
|
|
e8ddeaf89f |
skills: a gate that could pass without checking, and a mailbox that never empties
Fixes the three blockers and seven should-fixes from pi@emb-7kj4vr4g's review (logstream correlation skills-provenance-review, full text in drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce). Every finding was reproduced by execution here before being fixed; two were refined by that reproduction rather than taken as given. BLOCKER 1 — the provenance gate could print OK and exit 0 without verifying. `git show <ref>:<path> | sha256sum` hashes EMPTY STDIN when the ref does not resolve, so at_ref was never empty and the UNKNOWN branch was dead code. Measured: a bogus ref reported MISMATCH — accusing the snapshot of lying when the real cause was an incomplete clone, and the operator's natural remedy for MISMATCH is to re-run the refresh, which rewrites provenance to silence the complaint; and with a 0-byte snapshot against a 0-byte upstream file it printed "OK: exactly skillset@aaaaaaa" with exit 0 for a ref that does not exist. The script already had the sha_empty idiom and had applied it to blob_sha but not to at_ref. Existence is now PROVEN with git cat-file -e before anything is hashed, at two levels (ref resolves / path exists at it) because those deserve different messages. Same defect class as the canary it replaces: a check that can succeed without checking. A second, unflagged instance of the same pipeline shape in blob_sha was found and fixed too. Exit codes split, because the old contract failed the sanctioned case: 0 truthful (including stale, with a NOTICE), 1 a lying record only, 2 cannot determine. AGENTS.md step 2 promised "the message distinguishes the two" and was the thing this branch was breaking; rewritten to state all three. BLOCKER 3 — VENDORED.md contradicted itself in the release whose stated invariant is non-contradiction: its hand-maintained provenance line named skillset 670f7f1, seven commits behind the ARG and itself the commit that told agents to hand-stamp added_by — the withdrawn instruction this work exists to stop shipping — while its cp recipe contradicted the "not cp" rule 20 lines above. Line removed (nothing forced it to move when the ARGs did); 670f7f1 kept only as a labelled cautionary example. The pi-extensions half was verified redundant (CI require_sha resolves PI_EXTENSIONS_REF) before removal. SHOULD-FIXES: `<root> --check`, the spelling VENDORED.md documented, silently ran a REFRESH because only $1 was parsed (both tools now parse all args and reject unknown ones); refresh at a detached/older HEAD silently rewound ref and bytes (now refused unless the recorded ref is an ancestor, --force to override); upstream_dirty was computed and never used in check mode; --no-skills --json printed human text and broke jq; --help was a hardcoded sed range this branch had already made stale; the fingerprint hashed SKILL.md alone so a live skill differing only in a sibling file reported "identical", and pi-extensions already ships two files, so it is now a per-skill TREE hash with the manifest field renamed skillset_snapshot_tree_sha256; the --no-skills smoke assertion was negative-only and passed on a crashed binary. mktemp+mv left files 0600 — CI was unaffected since the index records 100644, so the blast radius was local builds only, narrower than the review inferred. Snapshot resynced c04cd15 -> 5fd0d5c so the no-clone fallback carries the CORRECTED coordination protocol rather than the withdrawn one; --check is now OK with no staleness notice, and the bidirectional canary re-verified against the new bytes. Local validation is bash -n only (shellcheck, hadolint and actionlint are all absent in this container) — CI remains the shellcheck gate. |
||
|
|
49a6534093 |
docs: write down the coordination channel the fleet already runs on
The logstream has carried cross-machine work since 2026-08-18 — patch handoff, review, a v1->v2 supersede — and nothing in this repo said it existed. That gap had a measurable cost this morning: another host addressed a retraction to pi@tor-ms22 by name and it was read only because the human said "read the logstream", while the agent was actively rebuilding the thing it warned about. Split by what each document is authoritative for, so there is one copy of each claim rather than three that drift: - README § Cross-machine agent coordination — what the CONTAINER needs. MEMPALACE_REMOTE_URL selects the shared palace; MEMPALACE_PI_DEVICE is what makes this machine reachable, because where every host is a thin client of one palace the stamped agent name is the only thing that distinguishes them. Stated as a rule with teeth: set both or neither, since a container missing the device var can read the log but is addressable by nobody. - AGENTS.md release checklist step 2 — the vendored-snapshot refresh, as a MECHANISM in the document a releasing agent actually reads, not a comment hoping to be noticed. It says the refresh costs a base rebuild, that skipping it is legitimate (every enrolled host reads its live clone), and that skipping it silently is not. - CHANGELOG — the three-way split itself, plus the measurement that shaped the ack contract: unfiltered, the mailbox returned 5 events, 4 of them finished broadcasts from eight days earlier; with status="open", exactly the 1 that needed an answer. Norms live in the skillset skill (82a8d3c, already live on every host that mounts the skillset — no rebuild) and mechanism in mempalace-toolkit's extensions/pi/README.md (e70bef2, which also documents the edge stamper that 553d8657 shipped undocumented). Deliberately NOT duplicated here. Consequence recorded rather than hidden: the skill edit lands in the skillset, so this repo's SKILLSET_SNAPSHOT_REF now honestly reports itself behind, and --check exits 1 with "has moved to 82a8d3c; the snapshot describes the older c04cd15". That message is also fixed in this commit — it previously blamed "the working tree" even when the tree was clean and only the ref had moved, which is the same defect class as a canary pinned to a phrase the release deleted: a message that names the wrong cause. Now distinguishes moved-HEAD from dirty-tree, verified against both plus the in-sync case. |
||
|
|
e070e0bcbf |
skills: record the vendored snapshot's provenance, and report which copy wins
Found while verifying v1.8.7 from inside a fresh container: the baked mempalace
snapshot is read by no host on this fleet. devbox-skill-reconcile repoints
~/.agents/skills/mempalace at the mounted live clone (the v1.8.5 fix working as
designed), and all four compose stacks mount a workspace containing the
skillset. So the phrase canary that blocked v1.8.7's first tag polices a file
nobody opens, while the drift that could actually mislead an agent — a git pull
nobody ran in /workspace/skillset — was invisible from inside the container and
is invisible to CI by construction.
Record provenance instead of policing it, and move the check to where the
skillset actually is:
- Dockerfile.variant: ARG SKILLSET_SNAPSHOT_REF (the claim) + a sha256 of the
shipped bytes measured in the manifest layer (the fact), as manifest siblings
rather than components{} members, plus an OCI label. An ARG default, not a
CI-resolved output: no credential for the private skillset, no change at any
of the four variant build call sites, and a local docker build records what CI
does. Variant-only, so no base rebuild — check-base-hash.sh scans
Dockerfile.base alone, verified by running it.
- pi-devbox-version: a skills: section naming baked vs live <repo> @ <sha> per
vendored skill, and for mempalace whether the live copy is identical to the
baked fingerprint, at the same commit with uncommitted edits, or divergent.
entrypoint-user.sh passes the new --no-skills, because the banner prints
before the links exist and long before the reconcile runs.
- scripts/vendor-mempalace-skill.sh: refresh the file and rewrite the ref
together (a cp without an ARG bump makes the manifest lie, which is worse than
anonymity); --check verifies the claim against a real clone.
- 5 new smoke assertions (78 -> 83), mutation-tested through the real sh -c
path: 6 fabricated manifests, where a well-formed hash of the wrong file
proves the two manifest assertions are not redundant; the all-baked reporting
test verified to FAIL against a live-skillset environment.
Reviewed mid-flight by pi@emb-7kj4vr4g over the logstream (correlation
skillset-vendor-drift), which retracted its own earlier recommendation of a
build-time byte-compare against skillset HEAD and supplied the better framing:
the invariant is NON-CONTRADICTION, not currency. Byte parity on a fallback
would have cost a resync commit plus a ~67-min base rebuild for each of the four
skillset commits pushed in one evening. Its warning also found a real bug here:
the script now CONSTRUCTS the snapshot from `git show HEAD:<path>` instead of
copying the working tree, because a clean `git diff` says nothing about an
untracked file — the one input the first draft would have recorded a false ref
for. Tested: untracked, unstaged and staged-but-uncommitted all refuse, atomically.
Also fixes three stale in-repo markers of the same class the canary belongs to
(true when written, silently false at release): two dangling "Unreleased"
pointers and a typst line still marked Unreleased five releases after v1.4.0.
|
||
|
|
f645e6654f |
smoke: fix the snapshot canary that blocked v1.8.7, and make it bidirectional
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 27s
Publish Docker Image / resolve-versions (push) Successful in 1m5s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 41m8s
Publish Docker Image / smoke (push) Successful in 4m49s
Publish Docker Image / smoke-studio (push) Successful in 18m28s
Publish Docker Image / build-variant (push) Successful in 15m46s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / update-description (push) Successful in 20s
Publish Docker Image / build-variant-studio (push) Successful in 16m52s
Run 589 built the base cleanly and then failed both smoke jobs 81-passed/1-failed on 'mempalace skill snapshot is current'. That canary greps a phrase from the vendored mempalace skill to detect a stale snapshot, and the phrase it pinned was 'Attribute what you file yourself' — the heading of the hand-stamping instruction that THIS release withdraws. So it fired correctly: the snapshot changed and the expectation did not. Every publish job was skipped, so nothing reached the registry and v1.8.7 was never consumed. Rather than bump the string: * the assertion is now BIDIRECTIONAL — the new phrase must be present AND the withdrawn one absent. A one-way canary only catches half the drift: it cannot notice a re-vendored stale snapshot that happens to contain the pinned phrase. Verified against v1.8.6's snapshot, which now correctly fails. * the comment records the structural limit rather than just the fix: a phrase canary can only ever detect 'older than what I remembered to pin', never 'older than skillset main'. Only a diff against the skillset repo can do that, which is now a Still-open item — it needs a CI clone credential for a private repo, i.e. a policy decision, not a code change. Changelog consolidated: the SSH sidecar multiplexing default moves from Unreleased into v1.8.7, since the retag will sit on a commit that contains it, and the v1.8.7 summary now records the failed first attempt rather than quietly presenting the second one as the whole story. |
||
|
|
657b1ad856 |
ssh sidecar: default to multiplexing, as a default and not an override
A target whose ~/.ssh/config entry never mentioned ControlMaster got no
multiplexing from the sidecar (only ControlPath was supplied), so every ssh call
opened a fresh TCP connection. On 2026-08-25 that produced ~12 connections to
one host in 15 min and a fail2ban block that looked like an outage — the tell
being that HTTPS to the same estate stayed healthy.
The correctness of this depends entirely on WHERE the block goes. ssh_config is
first-value-wins:
ControlPath before the Include -> override (the user's value points at
read-only ~/.ssh and cannot work in the container)
ControlMaster after the Include -> default (an explicit per-host
'ControlMaster no' must keep winning)
Force what is broken, default what is merely absent. The first draft put both in
the leading block and would have silently overridden an explicit 'no'.
Verified with ssh -G rather than from the man page, including the counterfactual:
under the shipped layout an explicit 'no' resolves to controlmaster false while a
silent host resolves to auto; under the rejected layout the 'no' host flips to
auto. So the test discriminates position, not presence. Plus a sandbox render of
the real script, bash -n, and shellcheck -S error (the v1.8.7 gate) clean.
Effect measured on 41 real host aliases: 22 silent entries gain auto+10m, 0
overridden. Note the fleet's one deliberate opt-out is written as absence plus a
comment ('# No ControlMaster — VPN means direct route'), which ssh cannot
distinguish from no opinion; that host now multiplexes, which its own comment
says is unnecessary rather than harmful.
Skill documents the sidecar-vs-~/.ssh trap (the failure misleads: read-only
ControlPath makes multiplexing look impossible rather than misconfigured) and
the stale-master recovery, ssh -O check / -O exit.
|
||
|
|
ebd0de0be2 |
changelog: v1.8.7 — device provenance reaches the fleet
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 42m4s
Publish Docker Image / smoke (push) Failing after 4m44s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 7m46s
Publish Docker Image / build-variant-studio (push) Has been skipped
The provenance fix's client half lives in mempalace-toolkit, which the image clones at build time, so it only reaches the fleet through a tag. Records both routes (extension via MEMPALACE_TOOLKIT_REF folded into base_tag; vendored skill via rootfs), the three design points (stamp in the client not the agent; diary marker in TEXT because metadata is invisible to readers; solitary devbox stamps nothing), and why the allowlist is per tool (3.8.0 hard-fails -32602 on undeclared args). Carries the CI-hardening work already sitting in Unreleased, and a Still open block for the three known bounds. |
||
|
|
9e744d701f |
lint: shellcheck the repo's own shell scripts, not just workflow run: steps
lint.yml has shellchecked every workflow `run:` step since the dash-vs-bash incidents, but nothing had ever pointed shellcheck at entrypoint.sh, scripts/*.sh or the extensionless tools under rootfs/usr/local/bin/. That gap is not hypothetical: the skillset repo's ci-release-watcher template shipped `echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)` for two months, where the heredoc IS python's stdin (no script arg) so the load hit EOF and the function silently returned nothing. shellcheck names exactly that at severity ERROR — SC2259, "This redirection overrides piped input" — and could have named it the whole time. New step in the existing actionlint job, so no second container pull: shellcheck -S error plus bash -n over every shell file, discovered as *.sh UNION a shebang scan (the glob alone misses pi-devbox-version, devbox-skill-reconcile, dot-watch and studio-expose; a shebang scan alone would miss a sourced fragment without one). Fails loudly on a zero-file match, because a green tick over an empty set is not a check. Severity chosen by measurement, not taste: -S error is 0 findings across all 11 shell files today, so the gate is green on arrival with no cleanup, while -S warning is NOT free (19x SC2088 tilde-in-quotes in recreate-sanity-check.sh plus assorted SC2016, all intentional) and would train everyone to ignore the job — the same reasoning as the SHELLCHECK_OPTS exclusions already on the actionlint step. |
||
|
|
2b8c3a4db4 |
ci: audit MEMPALACE_VERSION the way PI_VERSION is audited
Closes the item v1.8.6 (and v1.8.5 before it) listed as "Still open": the palace pin was a literal string in Dockerfile.base with zero references in docker-publish.yml, while PI_VERSION had a concreteness gate, a published-on-registry check and a never-silently-adopt drift warning. resolve-versions now applies all of those to MEMPALACE_VERSION, read from Dockerfile.base so a local `docker build` and CI install the same version by construction, plus one gate pi does not need: a YANKED release is refused, because an exact pin installs one silently under PEP 592 and would have shipped a withdrawn palace client to the whole fleet. smoke gains `installed mempalace matches CI's audited pin` via a new EXPECTED_MEMPALACE_VERSION threaded into both smoke jobs. It is not redundant with `manifest mempalace_version matches the installed core`: that compares two properties of one image and cannot notice that both are the wrong version. The case this covers is a variant built FROM a cached base carrying an older pin — internally consistent, silently stale. Mutation-tested by extracting the shipped block out of the YAML and stubbing curl: 9 cases covering every gate, then once end-to-end against live PyPI. That found a real defect in the first draft — the yank message inlined a jq program inside $(...) inside a double-quoted string, where the escaping broke the filter (jq compile error) while the surrounding `exit 1` still fired: a gate that looked correct and reported garbage. Note: correcting Dockerfile.base's now-false "known gap, carried forward" comment forces a base rebuild (~67 min) on the next tag. Leaving a comment asserting the audit does not exist was the worse option. |
||
|
|
cb7b8ad2ae |
smoke: assert manifest VALUES, not the presence of field names
Four build-provenance assertions grepped the manifest for a field name and
never looked at the value:
run_expect "manifest records pi_version" "cat …manifest.json" '"pi_version"'
which passes on {"pi_version": ""} and on {"pi_version": null}. The tell was
in its own passing output the whole time — `✅ manifest records pi_version (got
"pi_version")` echoes the key back as the thing it claims to have found. Found
while reading run 579's smoke log to confirm v1.8.6's new assertions had really
executed rather than merely gone green.
Now checked against values, and against ground truth where it exists:
- every required component key present, naming the one that vanished
- every component value a full 40-hex SHA (null allowed for pi-studio alone,
which is legitimately absent in the non-studio variant)
- pi_version equal to `pi --version`, mirroring the mempalace ground-truth check
- release_tag non-empty; source_revision 40-hex and build_date ISO-8601 *when
populated*, since both default empty on a plain local `docker build` and
demanding them would fail honest local smoke runs
- --json compared byte-for-byte with the file, which is assertable because that
mode is a verbatim cat; the old form grepped its output for "release_tag"
Key presence and value shape are deliberately SEPARATE assertions: a single
"all values are valid SHAs" loop passes vacuously on components:{}, because
jq's all() over an empty list is true. Combining them would reproduce the same
shape of hole as the three false greens already recorded in CHANGELOG.md.
Dropped `manifest has no unresolved ('unknown') components`: the 40-hex check
strictly subsumes it ("unknown" is not 40-hex, and only rev() emits it, feeding
components{} exclusively). Removed rather than kept, because a check that can
no longer fail independently is one more green tick that means nothing.
Mutation-tested twice rather than reasoned about: nine fabricated manifests
through the raw jq filters, then twelve through the shipped assertions using
the real `run` helper's `sh -c` quoting path — the quoting is load-bearing,
since a jq filter dying on a quoting error exits non-zero and looks exactly
like a caught defect. Measured on the same twelve defects: old caught 3, missed
9; new catches 12. Three legitimate variations stay green (empty
source_revision, empty build_date, null pi-studio).
Also corrects a factually wrong "Still open" bullet in the released v1.8.6
entry, which claimed pi-devbox-version's human output does not show
mempalace_version and that only --json surfaces it. Both halves are false: it
prints a `palace:` line with live-vs-baked drift annotation, verified against
fabricated manifests (match, skew, and pre-v1.8.6 absent-field cases). Left as
a struck-through correction rather than deleted, since v1.8.6 is published.
|
||
|
|
93f986e90e |
v1.8.6: adopt pi 0.84.3 + mempalace 3.8.0, close the v1.8.5 doc/observability gaps
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Lint / actionlint (push) Successful in 1m9s
Publish Docker Image / build-base (push) Successful in 41m23s
Publish Docker Image / smoke (push) Successful in 4m50s
Publish Docker Image / smoke-studio (push) Successful in 5m5s
Publish Docker Image / build-variant (push) Successful in 15m46s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 15s
Publish Docker Image / build-variant-studio (push) Successful in 19m59s
Three coupled pieces of work, all of which ride on the base rebuild that the
mempalace bump forces anyway.
DRIFT ADOPTED
- pi 0.84.2 -> 0.84.3. Its release notes carry a "Breaking Changes" line
(GoogleThinkingLevel -> GoogleApiThinkingLevel). Audited before adopting:
zero references across all four vendored companions (pi-fork,
pi-observational-memory, pi-atelier, pi-studio), so it is inert for us. The
reason to adopt is two skill-discovery fixes that land directly on v1.8.5's
vendored-skill work: nested Markdown skills inside grouping directories were
not discovered, and root README.md/AGENTS.md in skill dirs were reported as
broken skills.
- mempalace core 3.7.1 -> 3.8.0. Additive/reliability only. Its sync fix
(#2320/#2322) stops sync --apply deleting drawers whose source_file was
unreachable *at that moment* -- which does NOT relax the standing landmine
against sync on the shared palace, because that landmine is about paths
permanently absent from whichever host runs the sync. Different failure
shape; the caution stands.
DOCS -- three defects, one of them public
- DOCKER_HUB.md advertised "neovim (LazyVim defaults)". Nothing in the image
installs LazyVim; the only nvim config is a 19-line sysinit.vim. CI PATCHes
this file into the Docker Hub description on every release, so this was a
false claim published to the world. Removed.
- agent-browser + Playwright + Chromium is the single largest addition in the
image (~625 MB) and had zero mentions in README, DOCKER_HUB or THIRD_PARTY --
it was documented only to agents, in the AGENTS.md managed block. Now
documented to humans, including the Chromium licence dimension.
- typst and socat appeared in README prose but not in the "What's inside"
inventory. Added.
OBSERVABILITY -- the three gaps v1.8.5 listed as still open
- build-manifest.json now records mempalace core, read from the live binary
(ground truth, not the build ARG). Placed as a sibling of pi_version rather
than inside components{}, because pi-devbox-version renders that map through
[0:12] and would truncate a version string.
- smoke asserts the pi-observational-memory clone actually CONTAINS the ce9fc98
auth fix, pinned to src/runtime.ts. Deliberately not a repo-wide grep: two of
the three markers also live under tests/, so the repo-wide form stays green
with the fix site reverted. That is the third false-green of this exact family
in this repo (canary phrase in both snapshots; reconciler fixture using a
non-owned name; now this) -- pin containment checks to the fix site.
- smoke asserts the feeder's pi@<device> agent default behaviourally. The
earlier audit concluded this needed a --print-config added upstream; it does
not. AGENT is assigned before arg parsing, so `bash -x mempalace-pi-session
--help` observes the real resolution with no toolkit change. Two-sided:
device set => pi@<device>, unset => must not be pi@*.
- pi-devbox-version now prints a palace: line with the same live-vs-baked drift
detection pi already had. This matters more than it looks: mempalace is the
one component that is both client (here) and server (synlig), so skew between
them is a real failure mode. Degrades quietly on pre-v1.8.6 images.
Deferred deliberately: a native arm64 act_runner on tor-ms22 (the current
runner is on synlig, x86_64, so every arm64 layer ships QEMU-emulated).
Analysis and caveats filed to the palace rather than actioned here.
|
||
|
|
26f223568d |
.env.example: document MEMPALACE_PALACE_PATH and why nothing exports it
The only MemPalace variable the template never mentioned, and the one that moves the feeders' stage as a side effect: the palace root resolves as $MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json -> ~/.mempalace/palace, and the stage is derived from it (<palace-root>/pi-stage). The comment states that precedence, records why neither the image nor the entrypoint exports it (pinning the palace without carrying the stage along re-creates the split a shared root removed, v1.8.2), warns that a stage whose persistence differs from the palace makes a scoped `mempalace sync` prune conversation drawers whose dedup key is the staged path, and notes it is a path INSIDE the container unlike the host-side WORKSPACE_PATH/SSH_KEY_PATH above it. Found while auditing a live host whose .env sets it redundantly to the default value. |
||
|
|
01abda3456 |
v1.8.5: the skillset owns its skills, and a dangling link no longer kills boot
Publish Docker Image / base-decide (push) Successful in 18s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / smoke (push) Successful in 7m31s
Publish Docker Image / smoke-studio (push) Successful in 20m56s
Publish Docker Image / build-variant (push) Successful in 16m11s
Publish Docker Image / promote-base-latest (push) Successful in 15s
Publish Docker Image / update-description (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 17m32s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / resolve-versions (push) Successful in 15s
Lint / actionlint (push) Successful in 20s
No pin moved except mempalace-toolkit fd8b15f5 -> 0fe64c4 (feeder defaults --agent to pi@$MEMPALACE_PI_DEVICE, so hand-filed palace writes carry provenance; $USER remains the fallback when the var is unset). pi 0.84.2 and mempalace 3.7.1 are still upstream-latest, atelier v0.8.2 keeps the >=0.7.1 floor for pi >=0.84, and om master is still ce9fc982 -- nothing landed after the merge that fixed the eight-week silent-observation bug. Expect a full multi-arch base rebuild: entrypoint-user.sh, rootfs/** and the resolved toolkit SHA all feed the base hash. |
||
|
|
b5810654f6 |
skills: let the skillset own the skills it owns, and stop a dangling link from killing boot
Baked skill links won over the live skillset clone for all three vendored skills, so a pushed edit to skills/mempalace/SKILL.md was invisible in every container until the next image build -- measured on two hosts (live md5 129bcc4752 vs baked 5236024fef). Cause was ordering, not intent: the baked links are created early with a create-only-when-absent guard to close a smoke readiness race, and the skillset deploy runs last and treats them as foreign. The comment claimed the opposite of the behaviour. The fix is not "skillset always wins". Ownership is per-skill: pi-extensions is owned by its package repo and copied over the snapshot at build time, so the skillset's lagging duplicate must keep losing; pi-devbox-environment is authored here. Only mempalace is skillset-owned. devbox-skill-reconcile therefore runs after the deploy and repoints only the names in skills/skillset-owned.txt, replacing a link solely when it points into the baked tree, so a real directory or a link pointing elsewhere is never disturbed. Precedence is now user override -> live clone (owned names) -> baked snapshot, with the early links intact as the fallback so the readiness race stays closed. Reviewing that turned up a latent boot-abort in the pre-existing baked-link block: `[ ! -e "$link" ]` is TRUE for a dangling symlink, so once a link can point into /workspace/skillset, a vanished mount makes plain `ln -s` fail with "File exists" -- and under `set -euo pipefail` that aborts container start before `exec "$@"`. Reachable on `docker restart` or a host reboot, not on a recreate, since ~/.agents is not a volume on any host. Now `ln -sfn`, which heals the link back to the baked fallback. Smoke additions cover what let this ship: the stale-snapshot canary grepped a phrase present in BOTH the stale and fresh copies, so it passed throughout; it now pins the newest section. Link targets are asserted, not just `test -L`; the owned-list content is asserted both ways; and the reconciler's replace path -- which no CI container exercises, since none mounts a skillset -- is covered by fabricating one. A mutation test showed the obvious three assertions still pass with the "is this link ours?" guard deleted, so a discriminating case was added: an owned name whose link is a user override outside the baked tree. Also refreshes the mempalace snapshot to skillset 670f7f1 (without it the fix helps only hosts that mount skillset) and corrects README, which documented the old, wrong precedence in three places. Verified with 12 fixture cases plus 2 mutants: ownership respected against the real trees, user overrides preserved, relative/trailing-slash/CRLF/space/glob inputs handled, dangling link healed, read-only skills dir exits 0, idempotent. |
||
|
|
4f6f470518 |
changelog: file the vendored-skill shadowing bug for v1.8.5
~/.agents/skills is asymmetric: the three pi-devbox-specific skills resolve to the baked copies while all others resolve to the live skillset clone. Root cause is ordering in entrypoint-user.sh -- baked links are created early (line 65) with a create-only-when-absent guard, and the skillset deploy runs last (line 387) and leaves them alone as foreign links. The comment at line 61 states the intent as protecting the skillset skill from being clobbered, but the effect is the reverse. Cost measured on two hosts: a pushed edit to the mempalace skill (live md5 129bcc4752) was invisible to both containers, which kept loading the baked copy (md5 5236024fef). Editing those three skills appears to work and silently does nothing until a rebuild. Filed as a known issue with a proposed fix rather than fixed here: changing symlink precedence is image behaviour and wants its own review plus a smoke assertion, and the early-link ordering exists to close a readiness race that must not regress. |
||
|
|
fbc1f86612 |
docs: a negative result is usually your own filter (skill + AGENTS.md)
Three false negatives in one session, all self-inflicted, all convincing because the command "succeeded": a `| head -20` proved an SSH peer absent that sits at line 454 of a ~500-line config; `ssh mac 'docker ps'` proved the host had no Docker, when the non-interactive PATH simply lacks /usr/local/bin; and `grep 'ssh '` proved no ControlMaster was running, when those processes rename themselves to `ssh: <path> [mux]`. Same root cause each time, so it goes in the skill rather than in a commit message: a positive result carries its own evidence, absence has to be earned. The skill (rootfs/, symlinked into ~/.agents/skills) is BAKED, so this is an image change and is logged in CHANGELOG Unreleased accordingly. Its §3 also now records that a live ControlMaster socket makes later commands authenticate not at all -- after editing a peer's authorized_keys, "it still works" proves nothing; prove it with -o ControlPath=none, or the breakage waits for a future session that has no memory of the edit. AGENTS.md: corrected a stale CI claim while placing the pointer. It said a tag push produces two runs including lint; lint.yml has since been scoped to branches: ['**'], which excludes tag refs, and refs/tags/v1.8.4 duly produced run 571 (publish) and nothing else. Kept the head_sha + workflow path filter advice, which is cheap and guards against a future v*-triggered workflow. Added a short section on verifying this repo from inside a container, including that docker-compose.yml here is a TEMPLATE pinning :latest while a real host runs its own per-machine file -- recreating from the repo copy can silently move a host off :latest-studio. Placement note: AGENTS.md is only auto-read when the cwd is this repo, so the durable rule lives in the skill, which loads by description match in any pi-devbox session. |
||
|
|
2ebf00d6d4 |
v1.8.4: the om fix lands upstream, pi-atelier v0.8.2, todo edit
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 22s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Successful in 42m16s
Publish Docker Image / smoke-studio (push) Successful in 5m0s
Publish Docker Image / smoke (push) Successful in 19m0s
Publish Docker Image / build-variant-studio (push) Successful in 17m38s
Publish Docker Image / build-variant (push) Successful in 28m36s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 10s
Headline: pi-observational-memory 37986b6 -> ce9fc98. The ambient-credential gate fix (6f694e6 + 699ccc7) was merged upstream as PR #52 on 2026-08-22, closing issue #51, so this release picks it up through the ordinary PI_OBSMEM_REF=master path with nothing carried locally. Every image up to and including v1.8.3 silently recorded zero observations on a Bedrock host using ambient AWS credentials; from this one on, /opt is the fix and the settings.json packages[] workaround should be deleted (verify against /etc/pi-devbox/build-manifest.json first). npm still ships the broken 3.0.4, which does not matter here because the image clones the ref instead. Pin bump: PI_ATELIER_REF / PI_ATELIER_VERSION v0.8.1 -> v0.8.2, audited per the floor note above the ARG. The only version in between is 0.8.2 itself and both its entries are Workspace-Pulse-internal (inspection coalescing and serialization; fresh inspection guaranteed at Turn end, retired sessions can no longer publish stale results). Nothing touches pi's private TUI renderer, which is the coupling behind the 0.6.0/0.7.0-under-pi-0.84 startup hang, and pi is unchanged at 0.84.2 - so the bump stays outside that risk class. Also baked by this build, no pin needed: pi-extensions 98eb07b -> 2022887 (the todo edit action), pi-fork 4a09af4 -> f1ff808, pi-studio 0.9.44 -> v0.9.48, mempalace-toolkit b609cf5 -> fd8b15f (docs only - the pi-session false-success guard was already baked in v1.8.3, confirmed by ancestry), aws-cli 2.36.24 -> 2.36.29 and the other *_VERSION=latest tools. README: the pin table also said mempalace 3.6.0, stale since v1.8.3 bumped it to 3.7.1. Fixed in passing. Unchanged and verified current: pi 0.84.2 (npm latest, published 2026-08-14), MEMPALACE_VERSION 3.7.1 (PyPI latest), pi-toolkit 0e1369e. |
||
|
|
c3b6d36778 |
docs(CHANGELOG): Unreleased — the todo extension gains an edit action
The commit itself lives in pi-extensions (2022887), but /opt/pi-extensions is baked into this image, so "which todo behaviour does this image have" is an image question. PI_EXTENSIONS_REF=main means the next build picks it up with no pin to bump -- worth stating explicitly, since a reader who expects a version bump will otherwise go looking for one. Also recorded, because it confused a session today: pi-atelier's tool_result hook replaces todo output with "N/M done - see sidebar" whenever the sidebar panel is visible, so an agent sees the counter and not the item text. Upstream's list returns every item; the terseness is atelier's deliberate context saving, not a limitation of the tool. |
||
|
|
3a509077c2 |
v1.8.3: mempalace 3.7.1, refreshed skill snapshot, census on PATH
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 23s
Publish Docker Image / resolve-versions (push) Successful in 13s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Successful in 53m49s
Publish Docker Image / smoke (push) Successful in 7m22s
Publish Docker Image / smoke-studio (push) Successful in 18m6s
Publish Docker Image / build-variant (push) Successful in 19m6s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / build-variant-studio (push) Successful in 27m5s
mempalace 3.6.0 -> 3.7.1. Verified against the 3.7.1 source rather than its
changelog, because the risk lands on palaces users cannot reconstruct: legacy
drawers lack the new chunk_total marker and both decision sites trust them, so
no mass re-mine; NORMALIZE_VERSION is 2 in both; chromadb stays <2 so no
index-format migration; no auto-migration exists; logstream.sqlite3 is created
lazily. Downgrade remains possible (3.6.0 has zero references to chunk_total).
Two behaviour changes documented in the CHANGELOG: ALLOW_PEER_WRITER no longer
works on local/chroma palaces, and writer-lock setup failures fail closed.
Neither affects this image's MCP-server-plus-CLI-feeder pattern, which already
serialised on the same lock under 3.6.0 -- the upstream "process-lifetime
single-writer" entry describes tightened escape hatches, not a new lease.
The motivation is the shared central palace: 3.7.1 drops the stale chromadb
SharedSystemClient cache on reconnect (3.6.0 could let a stale in-memory HNSW
segment overwrite a peer's writes, "index count going backwards"), releases the
writer lease on SIGTERM/SIGHUP, and stops treating an interrupted mine as
complete. The fleet primary was upgraded to 3.7.1 and restarted before this tag,
because 3.7.1 refuses writes when the served library drifts and reconnect cannot
clear that. opencode-devbox still pins 3.6.0, so the lockstep is broken until it
cuts its own release.
Vendored mempalace skill snapshot refreshed to skillset 936fed8 (was 63f3bf5).
This closes a gap that had been invisible for two commits: ~/.agents/skills/
mempalace symlinks to the IMAGE-BAKED copy, entrypoint-user.sh creates that link
first, and the skillset deploy never clobbers an existing name -- so in a devbox
container the vendored snapshot always wins and editing skillset alone changes
nothing a container reads. Brings the multi-machine shared-palace section and
the hand-crafted-provenance guard.
pi-global-AGENTS.append.md already carried the three shared-palace damage rules
(
|
||
|
|
ffd54750b9 |
docs: CHANGELOG for v1.8.2 — the silent transcript-feed failure and its guard
Publish Docker Image / resolve-versions (push) Successful in 25s
Lint / actionlint (push) Successful in 30s
Publish Docker Image / base-decide (push) Successful in 8s
Lint / hadolint (push) Successful in 46s
Publish Docker Image / build-base (push) Successful in 41m26s
Publish Docker Image / smoke (push) Successful in 4m45s
Publish Docker Image / smoke-studio (push) Successful in 5m11s
Publish Docker Image / build-variant-studio (push) Successful in 17m44s
Publish Docker Image / build-variant (push) Successful in 23m41s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 10s
|
||
|
|
53b41cd76b |
smoke: assert the pi stage $HOME-relative, and add a smoke_only dispatch
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 13s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m20s
Publish Docker Image / smoke (push) Successful in 14m40s
Publish Docker Image / build-variant-studio (push) Successful in 21m51s
Publish Docker Image / build-variant (push) Successful in 15m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 17s
v1.8.0 never shipped: smoke (67/68) and smoke-studio (70/71) each failed the
same single assertion, so build-variant and everything downstream skipped and
latest stayed on v1.7.0.
The assertion was wrong, not the product. It grepped for a literal
stage=/home/developer/.mempalace/pi-stage/, but run() invokes
docker run --rm --entrypoint="" "$IMAGE" sh -c "$cmd"
and neither Dockerfile sets USER or ENV HOME — the published base image config
has no HOME at all; it is normally set by entrypoint-user.sh, which
--entrypoint="" skips on purpose. So the assertion executed as root with
HOME=/root, mempalace-pi-session correctly resolved
stage=/root/.mempalace/pi-stage/... (the stage is $HOME-relative by design), and
the literal grep could never match under any circumstances.
The tell was one line below in the log: the sibling assertion "pi stage follows
MEMPALACE_PALACE_PATH" PASSED, because it sets the variable explicitly and never
consults HOME. Default fails + explicit passes = wrong HOME, not broken staging.
Now asserts the invariant actually intended — the stage sits beside the resolved
palace, sharing its lifetime — which is user-independent:
case "$stage" in "stage=$HOME/.mempalace/pi-stage/"*) exit 0 ;; *) exit 1 ;; esac
$HOME is expanded by the container's own shell, so it holds as root, as
developer, or under any future user. Verified all four cases against the real
bin/mempalace-pi-session by extracting the committed assertion bodies and
running them under sh -c: virgin HOME -> exit 0; HOME=/home/developer -> exit 0;
MEMPALACE_PI_STAGE pinned to a .cache path -> exit 1 (the regression this
assertion exists to catch still fails it); developer-identity companion -> 0.
Added that companion assertion, "pi stage is palace-adjacent for the developer
user", which covers the deployment-specific path properly by SUPPLYING
HOME=/home/developer rather than assuming it.
Why this took a release to surface: docker-publish.yml triggers on push tags v*
only. The assertion was added on a push to main (
|
||
|
|
29b62093f0 |
v1.8.0: bump pi 0.84.1 → 0.84.2 and pi-atelier v0.8.0 → v0.8.1, audited
Publish Docker Image / resolve-versions (push) Successful in 12s
Lint / actionlint (push) Successful in 23s
Publish Docker Image / base-decide (push) Successful in 8s
Lint / hadolint (push) Successful in 1m20s
Publish Docker Image / build-base (push) Successful in 41m34s
Publish Docker Image / smoke (push) Failing after 4m41s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 8m22s
Publish Docker Image / build-variant-studio (push) Has been skipped
pi 0.84.2 closes the Amazon Bedrock tool-argument poison pill that v1.6.4 recorded as "Not fixed upstream". pi-ai 0.84.2 adds a recursive sanitizeBedrockDocument() and applies it at exactly the site that entry named (dist/api/bedrock-converse-stream.js, line 692 -> 704; upstream PR #7882): - toolUse: { ..., input: c.arguments }, + toolUse: { ..., input: sanitizeBedrockDocument(c.arguments) }, It strips object members whose key is the empty string, recursing through arrays and nested objects. It runs at request-build time, so it covers the live turn and a resume alike: a session already bricked by an empty-key tool argument now replays instead of dying on a Bedrock ValidationException. pi-session-repair is therefore no longer the recovery path on this image -- it stays useful for older images and for inspection, since the fix sanitises what is sent, not what was recorded. Bumping PI_VERSION is the ONLY way to get that fix: pi publishes an npm-shrinkwrap.json, so pi 0.84.1 pins pi-ai to exactly 0.84.1 even though its package.json range (^0.84.1) would admit 0.84.2. Transitive upstream fixes never leak into this image. pi-atelier v0.8.1 is the matching companion -- both sides changed fullscreen input handling within three days. Its only code change (src/split-pane.ts) stops atelier writing its own 1002h/1006h pair around a sidebar resize under pi's fullscreen renderer, which had been tearing down the mouse reporting pi itself enabled and leaving the wheel dead. Audited against the surfaces the pin policies name: - session .jsonl: identical migrateV1ToV2/migrateV2ToV3 ladder - node engine floor: unchanged >=22.19.0 (image ships 22.23.2) - atelier's three private couplings all intact in pi-tui 0.84.2 -- class TuiAltScreen extends TuiBase (detected by constructor NAME, so a rename would fail silently), inputListeners still `new Set()` at the same line 103, own render(width) descriptor still present - pi's mouse sequences byte-identical between 0.84.1 and 0.84.2 - atelier metadata unchanged: engines >=22.19.0, peerDeps >=0.80.7, zero runtime deps, so the "no npm install step" note holds Not proven by execution: CI smoke does not drive the TUI, and 0.84.2 adds a focused fullscreen search overlay that also participates in input handling. The pairing is reasoned from the diffs. Worth an alt+a plus a sidebar resize and wheel scroll in fullscreen on first use. |
||
|
|
cbd7cf5c67 |
entrypoint: announce the "remote palace, no inbox" skip instead of vanishing
When MEMPALACE_REMOTE_URL is set but MEMPALACE_PI_SSH_TARGET is not, the feeder has nowhere to ship staged transcripts, so skipping is correct. The problem was that the branch was a bare `:` AND the skip happens before the subshell that writes ~/.pi/agent/mempalace-catchup.log -- so a container in that state contributed nothing to the palace and left no artifact at all, not even an empty log, to explain why. It is indistinguishable from a healthy run that had nothing to file, which is the worst property a memory system can have: the failure looks exactly like success. Surfaced while flipping the first client onto the shared palace, where this is the single most likely way to end up quietly memory-less -- the palace is the only thing that survives a container recreate. The notice goes to both the container start output (docker logs) and the log path anyone debugging looks at first. It names both variables, says what still works (MCP tools read/write the shared palace; only this container's own conversations go nowhere), and points at MEMPALACE_FEED=0 for anyone who meant it -- "HTTPS first, mining later" is a documented interim state, so the notice has to be silenceable without being ignorable. Guarded against becoming a startup failure. mkdir -p in a branch that previously touched no filesystem is a new risk: an unwritable ~/.pi (root-owned volume, a classic Docker accident) fails under set -e and would abort the entire entrypoint. It now degrades to stdout-only. Verified all five paths by executing the extracted block under `set -euo pipefail`: the trap prints and writes the log; unwritable ~/.pi still exits 0 and still prints; MEMPALACE_FEED=0 stays completely silent (no message, no file); and both normal-remote and local mode still background the feeder with no notice. Two smoke assertions guard against a regression to the silent no-op. They test the entrypoint as shipped in the image rather than behaviour, because this branch only runs at container start and a `docker run` one-shot cannot reach it. entrypoint-user.sh is COPY'd in Dockerfile.base, so this rides the base rebuild the Unreleased feeder work already needs. |
||
|
|
7c00dd6001 |
mempalace: drop the stage ENV pin, fix the shared-server compose
The feeder now defaults to <palace-root>/pi-stage upstream, so pinning
MEMPALACE_PI_STAGE into ~/.pi here is unnecessary -- and was actively wrong. It
created a second convention that could still diverge from the palace: keep the
devbox-palace volume, drop devbox-pi-config, and a scoped `mempalace sync`
prunes every conversation drawer, because dedup keys on the staged path. Both
the ENV and the entrypoint export are gone; a comment explains why adding one
back re-introduces the split it was meant to fix.
docker-compose.mempalace.yml was broken on mempalace 3.6.0 in both directions:
- `--host 0.0.0.0` with no token in the environment makes the server refuse
to start, crash-looping under `restart: unless-stopped`.
- Supply a token and the healthcheck's unauthenticated `tools/list` POST 401s,
marking a perfectly healthy server unhealthy forever.
Now the token is required via ${MEMPALACE_REMOTE_TOKEN:?...} so it fails fast at
`docker compose up` with a readable message, and the healthcheck probes the
deliberately token-free /healthz. The "no authentication of its own" security
note has been stale since 3.6.0 and is replaced with the actual posture
(bearer token + Host pin + Origin allowlist), including why browser-shaped auth
must not be put in front of it.
Dockerfile.base: mempalace-pi-session symlinked onto PATH, with a build-time
`--help` check so a broken feeder fails the image build rather than the first
session.
smoke-test: assert the stage resolves beside the palace (default, and following
$MEMPALACE_PALACE_PATH) instead of asserting the removed ENV pin. The two
behavioural guards -- a synthetic session that must be captured, an abandoned
one that must not be -- are unchanged.
.env.example: recommend `mempalace serve` on the docker0 gateway rather than
`mempalace-mcp --transport http --host 0.0.0.0`, with the two binds to avoid.
|
||
|
|
43cd6e22f2 |
v1.7.0: bundle pi-atelier at a pinned tag; pin pi to an audited 0.84.1
Publish Docker Image / resolve-versions (push) Successful in 9s
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 13s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 41m22s
Publish Docker Image / smoke-studio (push) Successful in 5m16s
Publish Docker Image / smoke (push) Successful in 7m31s
Publish Docker Image / build-variant-studio (push) Successful in 18m31s
Publish Docker Image / build-variant (push) Successful in 27m18s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / update-description (push) Successful in 12s
Two changes that belong together, because the first is what makes the second dangerous to get wrong. pi-atelier (TUI sidebar + status rail) is now vendored to /opt/pi-atelier at PI_ATELIER_REF=v0.8.0 and registered by entrypoint-user.sh — the pi-fork / pi-observational-memory / pi-studio pattern, deliberately NOT `pi install npm:pi-atelier`, which writes into ~/.pi/npm-global on the config volume where it shadows the image and pins nothing. Unlike its siblings it gets no `npm install`: atelier declares zero runtime deps (peerDeps only, satisfied by the baked pi) and has no build step, so pi loads its TypeScript straight from the checkout via package.json `pi.extensions`. pi is no longer resolved to npm `latest` at build time. The pin lives in Dockerfile.variant and CI reads it from there, so a local `docker build` and a CI release ship the same versions by construction. The pin is a CHECKPOINT, NOT A FREEZE: bumping stays a one-line change; what stops is *unreviewed* adoption of whatever shipped that morning, in the same build that then gets tagged and published. CI fails when a pin is not concrete or not actually published on npm, and warns — never adopts — when npm latest moves ahead, naming what to re-check. Why this pairing needed care: pi-atelier 0.6.0/0.7.0 wrap pi's PRIVATE TUI renderer, and under pi 0.84 that wrapper recurses — pi hangs at startup burning CPU with no error. Upstream fixed the recursion in 0.7.1 and restored the non-overlapping split in 0.7.2; 0.8.0 is additive on top. atelier's own peerDependencies still say >=0.80.7, which does not express that floor, so nothing in npm metadata could have warned us. The floor is therefore encoded as an executable rule — pi >= 0.84 => pi-atelier >= 0.7.1 — asserted in both smoke-test.sh (build time) and recreate-sanity-check.sh (after a real recreate), verified against a 4x4 version matrix. Existing volumes needed migration, not just vendoring: a hand-installed `npm:pi-atelier` entry is counted as already-registered by the entrypoint guard, so every existing volume would have kept its unpinned npm copy — and a 0.6.x copy next to pi 0.84 is exactly the startup hang. The entrypoint now drops that one exact string (settings.json.bak.atelier.<ts> backup, distinct prefix so it cannot clobber the template merge's backup in the same second) and lets the pinned /opt copy register. Tested against a real settings.json: only that entry removed, other packages and all keys intact, idempotent, and unparseable JSON leaves the file untouched. DEVBOX_ATELIER=0 opts out entirely — in the entrypoint rather than via `pi uninstall`, because this component's failure mode is "pi will not start", which cannot be repaired from inside pi. 0.84.1 was audited for this release, not merely adopted: theme/TUI changes are additive, the session format is unchanged (CURRENT_SESSION_VERSION = 3 in both 0.83.0 and 0.84.1 with an identical migrateV1ToV2/migrateV2ToV3 ladder, so existing transcripts are neither migrated nor at risk and pi-session-repair stays valid), and the Node engine floor is unmoved at >=22.19.0. CI resolves the atelier tag to its PEELED commit SHA — atelier uses annotated tags, so the unpeeled ref is a tag object, not a commit; pi-studio's lightweight tags never exposed that distinction. Also: docs for overriding the read-only ~/.ssh/config from the container — container-only keys in ~/.ssh-local, hardened authorized_keys, the fact that `from=` must allow the HOST's addresses because container egress is NAT'd through it, and the macOS-only-keyword trap (`UseKeychain` is fatal to Linux OpenSSH and takes out dssh/pi --ssh while the host keeps working). Corrects two claims in "Naming LAN peers": ssh-lan.conf is not ProxyJump-only, and first-time creation does need one restart because the Include is emitted only when the file already exists at start. |
||
|
|
572430237f |
Bump mempalace pin 3.5.0 → 3.6.0 (lockstep with opencode-devbox v2.9.0)
3.6.0 (2026-07-17, PyPI latest) is additive/reliability only: secure
`mempalace serve` remote mode, optional Milvus backend, atomic KG
supersede(), conversation chronology, mining exclusions, plus recovery and
locking fixes.
Reviewed for MCP tool-schema changes before bumping — that being the exact
regression class this pin exists to catch, after an unpinned install once
swept in the broken 3.3.x/3.4.0 diary_write schema. There are none, and
nothing touches diary_write, so the perl workaround removed in v1.2.2 stays
removed.
Two fixes matter for how this image uses mempalace: read-only mode now covers
checkpoint + delete_by_source in _MUTATING_TOOLS (#1930), and agent
attribution is preserved in mempalace_checkpoint (#2023/#2034) — the latter
because the diary protocol relies on per-agent attribution.
Also adds a CHANGELOG Unreleased block that backfills the per-variant image
description labels (
|