Files
mempalace-toolkit/docs/phase-1-exposure-runbook.md
T
Joakim Persson 00a95d1a2f docs: why the tunnel and the feeder's SSH path are not redundant; newt done
Asked "why do we need Pangolin if you proposed rsync/ssh?", and the runbook did
not actually answer it -- it stated both were needed without saying why neither
substitutes. New section 1.3:

  - Pangolin/HTTPS carries the MCP tool surface (search, add_drawer,
    diary_write, kg_*) -- every live tool call, from any MCP client.
  - SSH/rsync carries transcript *files* only, because mempalace_mine expands
    its source path server-side, so the server can only mine its own disk.

HTTPS alone is a palace you can query but cannot feed; SSH alone is files with
no query API. The rsync is not a transport preference, it is a workaround for
where `mine` resolves paths.

Records honestly that `ssh -L 8765:172.17.0.1:8765 synlig` *would* replace the
tunnel for MCP, and why we don't: synlig dials out (reaching for a dial-out
tunnel is itself the evidence inbound was unavailable), MCP clients want a
durable URL rather than a per-session forward, and the forward must be up on
every device before every session.

And the design's weak point, stated instead of glossed: the rsync runs
client -> synlig, so mining needs synlig's SSH reachable *from the client*. Were
that true everywhere, no tunnel would be needed for MCP either. Honest
expectation after Phase 1 is therefore: query/write from anywhere, mine only
from devices that can reach synlig's SSH. Section 4 now carries the upstream ask
that would close the gap -- have the feeder send content over MCP (add_drawer /
diary_write, which it already calls) instead of asking the server to mine a path
it must first rsync there.

New section 3.7, a live trap for the imminent client flip:
MEMPALACE_REMOTE_URL on its own does not degrade to local feeding, it *stops*
feeding. auto mode switches to remote as soon as the URL is set (:286) and
remote mode then exits 1 without MEMPALACE_PI_SSH_TARGET (:298-300), before
anything is staged or filed -- so a cron feeder just starts failing, and the
loudest symptom is silence. Two safe orders given: set all three variables in
one edit, or set URL+token and pin --mode local until the SSH target exists.

newt is installed on synlig and connected to Pangolin (done 2026-08-12), marked
here and in the synlig runbook's item 2; the blocker is now item 3, the one
sudo. Added the follow-up that "connected to Pangolin" only proves newt reached
nyvaken -- reaching the *palace* is a separate claim that fails independently,
so probe 172.17.0.1:8765/healthz from inside newt's namespace.

All seven code citations verified against the source at commit time.
2026-08-13 00:15:57 +02:00

14 KiB

Phase 1 exposure — newt on synlig, DNS, and the client auth model

Companion to rfc-001-global-palace.md (design + decisions) and synlig-primary-runbook.md (what is already installed on the primary). This doc covers only the step the other two leave open: making the primary reachable — runbook §4 items 2 and 5.

Status 2026-08-12 — Pangolin updated on nyvaken (done, yours). newt not yet installed on synlig. Nothing exposed. No client .env flipped.

Read this before touching Pangolin: three of the four questions this step raises were already decided in RFC §6.2 on 2026-08-09, and re-deciding them differently is how the fleet ends up in two states.


1. The four questions, answered

Question Answer Where it was decided
Which port? 8765, path /mcp (liveness: /healthz) cli.py:2141 default; runbook §2.4
What does newt target? 172.17.0.1:8765 (docker0), never 127.0.0.1 RFC §6.2 Transport; runbook §2.4
Open, or authenticated? Authenticated. The primary is never an open public resource. RFC §6.2 Network posture
Per-device credentials? No — Phase 1 ships the single shared bearer token. Per-device tokens are Phase 4. RFC §6.2 Authentication

1.1 Why not per-device users at the proxy

The instinct — "create a Pangolin user per container, put the credentials in each .env, keep the usernames distinct" — is the right goal (revocation, attribution) reached through the wrong layer, twice over:

  1. mempalace validates exactly one token. hmac.compare_digest(provided, f"Bearer {srv.auth_token}") (mcp_server.py:5292-5295) — there is no user table and no second credential. Per-device HTTP identity is not a configuration you can express today; it is Phase 4 work (a server-side token → {device_id, scopes} registry). RFC §6.2 chose the shared token for Phase 1 deliberately: "iterate more feature rich but more complex solutions over time."

  2. Pangolin's HTTP auth is browser-shaped; the clients are not. SSO login, resource PIN and resource password all assume something that can follow a redirect, render a form and hold a session cookie. Every MemPalace client here is a headless JSON-RPC POST with an Authorization header — pi's extension, opencode's type:remote MCP entry, and mempalace-pi-session --mode remote's urllib.request.urlopen. Point those at a user-authenticated resource and they receive a login page where JSON should be. Enabling that protection breaks precisely the clients it is meant to protect.

So: Pangolin terminates TLS and nothing more (RFC §6.2 Transport, decided 2026-08-09). The bearer token is the authentication. This is not "unprotected" — an unauthenticated request to /mcp gets a 401 from mempalace itself, verified A4/A5 in runbook §2.4.

Consequence to accept consciously (RFC §6.2, §7.3.2): until Phase 4 the primary cannot tell devices apart. origin_device is client-asserted and advisory — nothing load-bearing may depend on it, and revoking one laptop means rotating the token everywhere.

1.2 The one place per-device identity does exist today

Remote mode is not only HTTP. mempalace_mine expands its source path in the server process, so a client's staged transcripts must physically exist on the primary. The feeder therefore ships them over SSH into a per-device inbox before asking the server to mine its own copy:

rsync -a --update -e "$ssh_cmd" "$STAGE/" "${SSH_TARGET%/}/$DEVICE/"   # bin/mempalace-pi-session:677-680

That SSH key is per-device identity, and it is individually revocable (one line out of authorized_keys) years before Phase 4 lands. It costs nothing extra, because the mining path needs SSH regardless.

Two implications people miss:

  • mempalace.jordbo.se alone does not enable mining. HTTPS covers the read/write tool surface (search, add_drawer, diary_write, kg_*) — genuinely useful on its own, and the reason to do this at all. But --mode remote also needs MEMPALACE_PI_SSH_TARGET reachable. Budget for both paths.
  • DEVICE defaults to $(hostname) (bin/mempalace-pi-session:163). In a container that is the container hostname: either random per recreate (inboxes proliferate; each recreate re-mines into a fresh empty inbox) or identical across sibling devboxes (two containers writing one inbox). Set MEMPALACE_PI_DEVICE explicitly per container. It is a label, not a secret, so put it somewhere reviewable — a committed compose file — where duplicates are visible. That, not username hygiene in .env, is the discipline this design actually asks of you.

1.3 "Then why Pangolin at all, if the feeder uses SSH?"

Because they are not alternatives — they carry different traffic, and neither substitutes for the other.

Pangolin/newt (HTTPS) SSH + rsync
Carries the MCP tool surface: search, add_drawer, diary_write, kg_* — every live tool call transcript files only, once per session or cron run
Used by the pi extension, opencode type:remote, any MCP client the feeder, internally (bin/mempalace-pi-session:677-680)
Needed because clients need one stable URL, reachable from wherever they are mempalace_mine expands its source path server-side, so the server can only mine files on its own disk

HTTPS alone is a palace you can query but cannot feed. SSH alone is files shipped with no live query API. The rsync is not a transport preference; it is a workaround for where mine resolves paths.

Could SSH replace Pangolin? Partly, and it is worth being honest about it: ssh -L 8765:172.17.0.1:8765 synlig yields a working local MCP endpoint with no public HTTPS at all. Three reasons this runbook does not do that:

  1. Direction. synlig dials out through newt. That we reached for a dial-out tunnel rather than a port-forward is itself the evidence that inbound was not available — a corporate host does not accept connections from a phone on a foreign network.
  2. MCP clients want a durable URL, not a per-session forwarded port. opencode type:remote takes a URL; a forward that drops takes the tools down mid-session.
  3. The forward must be up on every device before every session. Pangolin is up once.

The weak point, stated plainly. The rsync runs client → synlig, so it needs synlig's SSH reachable from the client. Were that already true everywhere, no tunnel would be needed for MCP either. So the honest expectation after Phase 1 is: query and write from anywhere, mine only from devices that can reach synlig's SSH (corporate network / VPN / LAN). See §4 for the change that would remove that limit.


2. The bind trap, in full

RFC §6.2 and runbook §2.4 already say do not bind loopback behind the tunnel, because enforce_host_pin = _http_is_loopback(host) (mcp_server.py:5367) makes a loopback bind reject the proxy's forwarded Host: with a 403 that reads exactly like a Pangolin misconfiguration.

Additional finding, 2026-08-12 — the same reflex also silently removes authentication. Token resolution in cmd_serve (cli.py:1447-1450) is:

loopback = _server_is_loopback(host)
if not token and not loopback and not args.allow_insecure:
    token, token_created = _load_or_create_server_token(palace_path)

Auto-minting is gated on the bind being non-loopback. A loopback bind therefore starts with no token at all — no error, no warning, --allow-insecure not required — because the server has concluded it is only reachable locally, while the tunnel is serving it to the internet. Bind loopback behind newt and you get a 403 wall and, the moment anything relaxes the Host pin, an unauthenticated palace.

Both failure modes have the same cure, already implemented in contrib/systemd/mempalace-serve.service: bind 172.17.0.1. Non-loopback, so the Host pin relaxes and the token is mandatory; docker0-only, so newt reaches it and the LAN does not.

Belt and braces: set MEMPALACE_MCP_HTTP_TOKEN explicitly in the unit rather than relying on auto-minting. Then no future bind change can quietly drop authentication.


3. Steps

Ordered so nothing is reachable before it is authenticated.

3.1 Start the primary (runbook §4.3 — one sudo, unit already staged)

sudo loginctl enable-linger ecsjper
cd ~/.config/systemd/user && mv mempalace-serve.service.staged mempalace-serve.service
systemctl --user daemon-reload && systemctl --user enable --now mempalace-serve

curl -s 172.17.0.1:8765/healthz     # expect ok
curl -s 127.0.0.1:8765/healthz      # expect 403 — correct, not a bug (§2)
ss -ltnp | grep 8765                # expect 172.17.0.1:8765 only

3.2 Collect the shared token

cat ~/.mempalace/server/f5d849287f6d73f0141b29d7/token

Directory name is sha256(realpath(palace))[:24] — it changes if the palace path ever moves. Store via the .env.age flow, 0600 (RFC §6.2).

3.3 newt on synlig — done 2026-08-12 (installed, connected to Pangolin)

synlig runs Docker (Gitea Actions runner + digikam) but no tunnel client — runbook §4.2. Pangolin on nyvaken cannot dial in; synlig must dial out. Add a newt container with the credentials Pangolin issues for a new site.

Because newt runs in Docker on this box, the docker0 bind is already correct for it: from inside the container the primary is 172.17.0.1:8765. Verify from inside newt's network namespace, not from the host, before touching DNS.

synlig has 7.8 GiB shared with a CI runner (runbook §1). newt is small, but do not colocate anything else here casually.

Confirm next, now that newt is up. "Connected to Pangolin" proves newt reached nyvaken — a different claim from newt reaching the palace, and the two fail independently:

# from inside newt's namespace, not from the host
docker exec <newt-container> wget -qO- http://172.17.0.1:8765/healthz    # expect ok

If that hangs or refuses while the Pangolin dashboard shows the site online, the tunnel is fine and the target is wrong — look at the resource's upstream address (§3.5), not at newt. Note this check needs §3.1 done first: if mempalace-serve is not running yet, it fails for that reason alone.

3.4 DNS at the web hotel

One CNAME: mempalacethe same target your existing Pangolin resources use (nyvaken's public hostname). RFC §6.2 costed this as "one DNS record per service on the web hotel is the whole setup cost."

⚠️ Not verified from here: nyvaken's public FQDN, and whether your web hotel permits a CNAME at that label (some require an A record, or forbid CNAME where other records exist). Confirm before assuming a 5-minute job.

3.5 Pangolin resource

  • Target: newt site → 172.17.0.1:8765, path /mcp (plus /healthz if you want the external probe).
  • Auth: none at the Pangolin layer (§1.1). TLS termination only.
  • Do not attach an Origin-injecting proxy or browser client: a present non-loopback Origin is a hard 403 with no override (runbook §2.4 B3).

3.6 Verify end-to-end before flipping any client

curl -s https://mempalace.jordbo.se/healthz                      # ok
curl -s -o /dev/null -w '%{http_code}\n' -X POST \
     https://mempalace.jordbo.se/mcp                             # 401 — the token is doing its job
curl -s -X POST https://mempalace.jordbo.se/mcp \
     -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
     -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -c 300   # 36 tools

The 401 check matters as much as the 200: it is the only evidence that the thing you just published to the internet is not open. Then, and only then, Phase 1 client flip — one machine first (RFC §8), and remember opencode containers need the §4.1 sidecar merge before their .env takes effect.

3.7 Flipping a client: set three variables, or none

MEMPALACE_REMOTE_URL on its own does not degrade to local feeding — it stops feeding. auto mode switches to remote the moment the URL is set, and remote mode then refuses to run without an SSH target:

auto) if [[ -n "$REMOTE_URL" ]]; then MODE="remote"; else MODE="local"; fi ;;   # :286
...
command -v rsync >/dev/null 2>&1 || { echo "error: rsync not found ..."; exit 3; }   # :297
if [[ -z "$SSH_TARGET" ]]; then
  echo "error: MEMPALACE_PI_SSH_TARGET unset (needed for --mode remote)" >&2; exit 1  # :298-300
fi

That exit happens before anything is staged or filed, and a cron-driven feeder will simply start failing — the loudest symptom is silence, which is the hardest kind to notice. Two safe orders:

  • Both paths at once: set MEMPALACE_REMOTE_URL, MEMPALACE_REMOTE_TOKEN and MEMPALACE_PI_SSH_TARGET (plus MEMPALACE_PI_DEVICE, §1.2) in the same edit.
  • HTTPS first, mining later: set the URL and token, and pin the feeder to --mode local until the SSH target exists. Tools then read/write the shared palace while transcripts keep landing in the local one.

Either way, run the feeder once by hand and read its exit code before trusting the timer. This is precisely the failure "one machine first" is meant to contain.


4. Still open

  • Feeding without SSH — the upstream ask that would close §1.3's gap. Have the feeder send content over MCP (add_drawer / diary_write, which it already calls) instead of asking the server to mine a path it must first rsync there. HTTPS would then be genuinely sufficient and mining would work from any network. Until then, mining is limited to devices that can reach synlig's SSH.
  • Per-device tokens — Phase 4. Until then origin_device is advisory (§1.1).
  • §7.6 diary dedup must be settled before the first §4.4 join; replay duplicates every entry.
  • §7.2: never run mempalace sync against the shared palace. Doubly true now that the pi/opencode feeders stage inside the palace root, which puts staged sources in scope for a sync of the palace dir.
  • nyvaken's public FQDN and the web hotel's CNAME rules — unverified (§3.4).