08e344b047
Both corrections come from the first real Phase 1 start on synlig (2026-08-12).
1. The Pangolin resource target was documented as bare `172.17.0.1:8765` with no
scheme, and the obvious guess from that is `https://` -- which cannot work.
`contrib/systemd/mempalace-serve.service` runs `serve --host 172.17.0.1
--port 8765` with no --tls-cert, so the primary speaks plaintext HTTP; TLS
terminates at Pangolin, which is the whole point of the RFC 6.2 decision.
Point a proxy at https:// and it attempts a TLS handshake against a plaintext
listener: 502 from outside, while `curl 172.17.0.1:8765/healthz` on the box
still says ok -- a confusing pair of symptoms. Now spelled `http://` with the
failure mode named, in the runbook and in the unit's comments.
2. `curl -s 127.0.0.1:8765/healthz` was documented as "expect 403". Wrong: the
real run returns empty. With a docker0-only bind nothing is listening on
loopback, so the connection is refused at TCP level before any header is sent
(%{http_code} -> 000, exit 7). The 403 is the *loopback-bind* case verified
2026-08-10 -- server on 127.0.0.1 answering a proxy-forwarded foreign Host.
Two distinct behaviours had been collapsed into one expectation in three
places (both runbooks and the unit).
Worth stating why the correction matters rather than just fixing it: refusal
is the *stronger* signal. A 403 proves only that a request was rejected; a
refused connection proves the loopback and LAN surface is not listening at
all. Someone who expected 403, saw silence, and "fixed" it by rebinding to
0.0.0.0 would have converted a correct configuration into an exposed one.
The docs now also say what to do if it hangs, or if ss shows 0.0.0.0:8765.
271 lines
15 KiB
Markdown
271 lines
15 KiB
Markdown
# Phase 1 exposure — newt on synlig, DNS, and the client auth model
|
|
|
|
Companion to [`rfc-001-global-palace.md`](./rfc-001-global-palace.md) (design + decisions) and
|
|
[`synlig-primary-runbook.md`](./synlig-primary-runbook.md) (what is already installed on the primary).
|
|
This doc covers only the step the other two leave open: **making the primary reachable** — runbook §4
|
|
items 2 and 5.
|
|
|
|
**Status 2026-08-12 — Pangolin updated on nyvaken (done, yours). newt not yet installed on synlig.
|
|
Nothing exposed. No client `.env` flipped.**
|
|
|
|
Read this before touching Pangolin: three of the four questions this step raises were **already decided**
|
|
in RFC §6.2 on 2026-08-09, and re-deciding them differently is how the fleet ends up in two states.
|
|
|
|
---
|
|
|
|
## 1. The four questions, answered
|
|
|
|
| Question | Answer | Where it was decided |
|
|
| --- | --- | --- |
|
|
| Which port? | **8765**, path **`/mcp`** (liveness: `/healthz`) | `cli.py:2141` default; runbook §2.4 |
|
|
| What does newt target? | **`172.17.0.1:8765`** (docker0), **never** `127.0.0.1` | RFC §6.2 Transport; runbook §2.4 |
|
|
| Open, or authenticated? | **Authenticated. The primary is never an open public resource.** | RFC §6.2 Network posture |
|
|
| Per-device credentials? | **No — Phase 1 ships the single shared bearer token.** Per-device tokens are Phase 4. | RFC §6.2 Authentication |
|
|
|
|
### 1.1 Why not per-device users at the proxy
|
|
|
|
The instinct — "create a Pangolin user per container, put the credentials in each `.env`, keep the
|
|
usernames distinct" — is the right *goal* (revocation, attribution) reached through the wrong *layer*,
|
|
twice over:
|
|
|
|
1. **mempalace validates exactly one token.** `hmac.compare_digest(provided, f"Bearer {srv.auth_token}")`
|
|
(`mcp_server.py:5292-5295`) — there is no user table and no second credential. Per-device HTTP identity
|
|
is not a configuration you can express today; it is Phase 4 work (a server-side
|
|
`token → {device_id, scopes}` registry). RFC §6.2 chose the shared token for Phase 1 deliberately:
|
|
*"iterate more feature rich but more complex solutions over time."*
|
|
|
|
2. **Pangolin's HTTP auth is browser-shaped; the clients are not.** SSO login, resource PIN and resource
|
|
password all assume something that can follow a redirect, render a form and hold a session cookie.
|
|
Every MemPalace client here is a headless JSON-RPC `POST` with an `Authorization` header — pi's
|
|
extension, opencode's `type:remote` MCP entry, and `mempalace-pi-session --mode remote`'s
|
|
`urllib.request.urlopen`. Point those at a user-authenticated resource and they receive a login page
|
|
where JSON should be. Enabling that protection breaks precisely the clients it is meant to protect.
|
|
|
|
So: **Pangolin terminates TLS and nothing more** (RFC §6.2 Transport, decided 2026-08-09). The bearer token
|
|
is the authentication. This is not "unprotected" — an unauthenticated request to `/mcp` gets a 401 from
|
|
mempalace itself, verified A4/A5 in runbook §2.4.
|
|
|
|
**Consequence to accept consciously** (RFC §6.2, §7.3.2): until Phase 4 the primary **cannot tell devices
|
|
apart**. `origin_device` is client-asserted and advisory — nothing load-bearing may depend on it, and
|
|
revoking one laptop means rotating the token everywhere.
|
|
|
|
### 1.2 The one place per-device identity *does* exist today
|
|
|
|
Remote mode is not only HTTP. `mempalace_mine` expands its source path in the **server** process, so a
|
|
client's staged transcripts must physically exist on the primary. The feeder therefore ships them over
|
|
SSH into a **per-device inbox** before asking the server to mine its own copy:
|
|
|
|
```sh
|
|
rsync -a --update -e "$ssh_cmd" "$STAGE/" "${SSH_TARGET%/}/$DEVICE/" # bin/mempalace-pi-session:677-680
|
|
```
|
|
|
|
That SSH key **is** per-device identity, and it is individually revocable (one line out of
|
|
`authorized_keys`) years before Phase 4 lands. It costs nothing extra, because the mining path needs SSH
|
|
regardless.
|
|
|
|
Two implications people miss:
|
|
|
|
- **`mempalace.jordbo.se` alone does not enable mining.** HTTPS covers the read/write tool surface
|
|
(`search`, `add_drawer`, `diary_write`, `kg_*`) — genuinely useful on its own, and the reason to do this
|
|
at all. But `--mode remote` also needs `MEMPALACE_PI_SSH_TARGET` reachable. Budget for both paths.
|
|
- **`DEVICE` defaults to `$(hostname)`** (`bin/mempalace-pi-session:163`). In a container that is the
|
|
container hostname: either random per recreate (inboxes proliferate; each recreate re-mines into a fresh
|
|
empty inbox) or identical across sibling devboxes (two containers writing one inbox). **Set
|
|
`MEMPALACE_PI_DEVICE` explicitly per container.** It is a label, not a secret, so put it somewhere
|
|
reviewable — a committed compose file — where duplicates are visible. That, not username hygiene in
|
|
`.env`, is the discipline this design actually asks of you.
|
|
|
|
### 1.3 "Then why Pangolin at all, if the feeder uses SSH?"
|
|
|
|
Because they are not alternatives — they carry different traffic, and neither substitutes for the other.
|
|
|
|
| | Pangolin/newt (HTTPS) | SSH + rsync |
|
|
| --- | --- | --- |
|
|
| Carries | the **MCP tool surface**: `search`, `add_drawer`, `diary_write`, `kg_*` — every live tool call | **transcript files only**, once per session or cron run |
|
|
| Used by | the pi extension, opencode `type:remote`, any MCP client | the feeder, internally (`bin/mempalace-pi-session:677-680`) |
|
|
| Needed because | clients need one stable URL, reachable from wherever they are | `mempalace_mine` expands its source path **server-side**, so the server can only mine files on its own disk |
|
|
|
|
HTTPS alone is a palace you can query but cannot feed. SSH alone is files shipped with no live query API.
|
|
The rsync is not a transport preference; it is a workaround for *where `mine` resolves paths*.
|
|
|
|
**Could SSH replace Pangolin?** Partly, and it is worth being honest about it:
|
|
`ssh -L 8765:172.17.0.1:8765 synlig` yields a working local MCP endpoint with no public HTTPS at all.
|
|
Three reasons this runbook does not do that:
|
|
|
|
1. **Direction.** synlig dials *out* through newt. That we reached for a dial-out tunnel rather than a
|
|
port-forward is itself the evidence that inbound was not available — a corporate host does not accept
|
|
connections from a phone on a foreign network.
|
|
2. **MCP clients want a durable URL**, not a per-session forwarded port. opencode `type:remote` takes a
|
|
URL; a forward that drops takes the tools down mid-session.
|
|
3. The forward must be up on **every device before every session**. Pangolin is up once.
|
|
|
|
**The weak point, stated plainly.** The rsync runs *client → synlig*, so it needs synlig's SSH reachable
|
|
**from the client**. Were that already true everywhere, no tunnel would be needed for MCP either. So the
|
|
honest expectation after Phase 1 is: **query and write from anywhere, mine only from devices that can
|
|
reach synlig's SSH** (corporate network / VPN / LAN). See §4 for the change that would remove that limit.
|
|
|
|
---
|
|
|
|
## 2. The bind trap, in full
|
|
|
|
RFC §6.2 and runbook §2.4 already say **do not bind loopback behind the tunnel**, because
|
|
`enforce_host_pin = _http_is_loopback(host)` (`mcp_server.py:5367`) makes a loopback bind reject the
|
|
proxy's forwarded `Host:` with a **403** that reads exactly like a Pangolin misconfiguration.
|
|
|
|
**Additional finding, 2026-08-12 — the same reflex also silently removes authentication.** Token
|
|
resolution in `cmd_serve` (`cli.py:1447-1450`) is:
|
|
|
|
```python
|
|
loopback = _server_is_loopback(host)
|
|
if not token and not loopback and not args.allow_insecure:
|
|
token, token_created = _load_or_create_server_token(palace_path)
|
|
```
|
|
|
|
Auto-minting is gated on the bind being **non-loopback**. A loopback bind therefore starts with **no token
|
|
at all** — no error, no warning, `--allow-insecure` not required — because the server has concluded it is
|
|
only reachable locally, while the tunnel is serving it to the internet. Bind loopback behind newt and you
|
|
get a 403 wall *and*, the moment anything relaxes the Host pin, an unauthenticated palace.
|
|
|
|
Both failure modes have the same cure, already implemented in
|
|
`contrib/systemd/mempalace-serve.service`: **bind `172.17.0.1`**. Non-loopback, so the Host pin relaxes and
|
|
the token is mandatory; docker0-only, so newt reaches it and the LAN does not.
|
|
|
|
> Belt and braces: set `MEMPALACE_MCP_HTTP_TOKEN` explicitly in the unit rather than relying on
|
|
> auto-minting. Then no future bind change can quietly drop authentication.
|
|
|
|
---
|
|
|
|
## 3. Steps
|
|
|
|
Ordered so nothing is reachable before it is authenticated.
|
|
|
|
### 3.1 Start the primary (runbook §4.3 — one `sudo`, unit already staged)
|
|
|
|
```sh
|
|
sudo loginctl enable-linger ecsjper
|
|
cd ~/.config/systemd/user && mv mempalace-serve.service.staged mempalace-serve.service
|
|
systemctl --user daemon-reload && systemctl --user enable --now mempalace-serve
|
|
|
|
curl -s 172.17.0.1:8765/healthz # expect ok
|
|
curl -s 127.0.0.1:8765/healthz # expect NOTHING — connection refused, exit 7 (see below)
|
|
ss -ltnp | grep 8765 # expect 172.17.0.1:8765 only
|
|
```
|
|
|
|
⚠ **The loopback probe returns empty, not 403** — corrected 2026-08-12 against the real run. Nothing is
|
|
listening on `127.0.0.1`, so the connection is refused at TCP level and `curl -s` prints nothing; check it
|
|
with `-w '%{http_code}'` → `000` and `$?` → `7`. The 403 belongs to a *different* configuration: server
|
|
bound **to loopback**, receiving a proxy-forwarded foreign `Host:` (§2, verified 2026-08-10). With a
|
|
docker0-only bind you cannot get 403 from loopback, because you never get far enough to send a header.
|
|
Refusal is the stronger signal of the two: it proves the loopback and LAN surface is not listening at all.
|
|
If it *hangs* instead, or `ss` shows `0.0.0.0:8765`, stop — that is not this configuration.
|
|
```sh
|
|
|
|
### 3.2 Collect the shared token
|
|
|
|
```sh
|
|
cat ~/.mempalace/server/f5d849287f6d73f0141b29d7/token
|
|
```
|
|
|
|
Directory name is `sha256(realpath(palace))[:24]` — it changes if the palace path ever moves. Store via the
|
|
`.env.age` flow, 0600 (RFC §6.2).
|
|
|
|
### 3.3 newt on synlig — ✅ done 2026-08-12 (installed, connected to Pangolin)
|
|
|
|
synlig runs Docker (Gitea Actions runner + digikam) but **no tunnel client** — runbook §4.2. Pangolin on
|
|
nyvaken cannot dial in; synlig must dial out. Add a `newt` container with the credentials Pangolin issues
|
|
for a new site.
|
|
|
|
Because newt runs in Docker on this box, the docker0 bind is already correct for it: from inside the
|
|
container the primary is `172.17.0.1:8765`. Verify from *inside* newt's network namespace, not from the
|
|
host, before touching DNS.
|
|
|
|
> synlig has 7.8 GiB shared with a CI runner (runbook §1). newt is small, but do not colocate anything
|
|
> else here casually.
|
|
|
|
**Confirm next, now that newt is up.** "Connected to Pangolin" proves newt reached *nyvaken* — a different
|
|
claim from newt reaching *the palace*, and the two fail independently:
|
|
|
|
```sh
|
|
# from inside newt's namespace, not from the host
|
|
docker exec <newt-container> wget -qO- http://172.17.0.1:8765/healthz # expect ok
|
|
```
|
|
|
|
If that hangs or refuses while the Pangolin dashboard shows the site online, the tunnel is fine and the
|
|
*target* is wrong — look at the resource's upstream address (§3.5), not at newt. Note this check needs
|
|
§3.1 done first: if `mempalace-serve` is not running yet, it fails for that reason alone.
|
|
|
|
### 3.4 DNS at the web hotel
|
|
|
|
One CNAME: `mempalace` → **the same target your existing Pangolin resources use** (nyvaken's public
|
|
hostname). RFC §6.2 costed this as *"one DNS record per service on the web hotel is the whole setup cost."*
|
|
|
|
⚠️ Not verified from here: nyvaken's public FQDN, and whether your web hotel permits a CNAME at that label
|
|
(some require an A record, or forbid CNAME where other records exist). Confirm before assuming a 5-minute job.
|
|
|
|
### 3.5 Pangolin resource
|
|
|
|
- Target: newt site → **`http://172.17.0.1:8765`** — path `/mcp` (plus `/healthz` for the external probe).
|
|
- ⚠️ **The scheme is `http`, not `https`.** The primary runs `serve --host 172.17.0.1 --port 8765` with no
|
|
cert (`contrib/systemd/mempalace-serve.service`): TLS terminates **at Pangolin**, which is the entire
|
|
point of the §6.2 decision. Point the resource at `https://172.17.0.1:8765` and Pangolin attempts a TLS
|
|
handshake against a plaintext listener — you get a 502/Bad Gateway from outside while the server itself
|
|
looks perfectly healthy on `curl 172.17.0.1:8765/healthz`. (Got this wrong on the first attempt
|
|
2026-08-12, because this line used to omit the scheme.)
|
|
- **Auth: none at the Pangolin layer** (§1.1). TLS termination only.
|
|
- Do **not** attach an `Origin`-injecting proxy or browser client: a *present* non-loopback `Origin` is a
|
|
hard 403 with no override (runbook §2.4 B3).
|
|
|
|
### 3.6 Verify end-to-end before flipping any client
|
|
|
|
```sh
|
|
curl -s https://mempalace.jordbo.se/healthz # ok
|
|
curl -s -o /dev/null -w '%{http_code}\n' -X POST \
|
|
https://mempalace.jordbo.se/mcp # 401 — the token is doing its job
|
|
curl -s -X POST https://mempalace.jordbo.se/mcp \
|
|
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
|
|
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -c 300 # 36 tools
|
|
```
|
|
|
|
The 401 check matters as much as the 200: it is the only evidence that the thing you just published to the
|
|
internet is not open. Then, and only then, Phase 1 client flip — **one machine first** (RFC §8), and
|
|
remember opencode containers need the §4.1 sidecar merge before their `.env` takes effect.
|
|
|
|
### 3.7 Flipping a client: set three variables, or none
|
|
|
|
⚠ **`MEMPALACE_REMOTE_URL` on its own does not degrade to local feeding — it stops feeding.** `auto` mode
|
|
switches to `remote` the moment the URL is set, and remote mode then refuses to run without an SSH target:
|
|
|
|
```sh
|
|
auto) if [[ -n "$REMOTE_URL" ]]; then MODE="remote"; else MODE="local"; fi ;; # :286
|
|
...
|
|
command -v rsync >/dev/null 2>&1 || { echo "error: rsync not found ..."; exit 3; } # :297
|
|
if [[ -z "$SSH_TARGET" ]]; then
|
|
echo "error: MEMPALACE_PI_SSH_TARGET unset (needed for --mode remote)" >&2; exit 1 # :298-300
|
|
fi
|
|
```
|
|
|
|
That exit happens **before anything is staged or filed**, and a cron-driven feeder will simply start
|
|
failing — the loudest symptom is silence, which is the hardest kind to notice. Two safe orders:
|
|
|
|
- **Both paths at once:** set `MEMPALACE_REMOTE_URL`, `MEMPALACE_REMOTE_TOKEN` **and**
|
|
`MEMPALACE_PI_SSH_TARGET` (plus `MEMPALACE_PI_DEVICE`, §1.2) in the same edit.
|
|
- **HTTPS first, mining later:** set the URL and token, and pin the feeder to `--mode local` until the SSH
|
|
target exists. Tools then read/write the shared palace while transcripts keep landing in the local one.
|
|
|
|
Either way, run the feeder once by hand and read its exit code before trusting the timer. This is
|
|
precisely the failure "one machine first" is meant to contain.
|
|
|
|
---
|
|
|
|
## 4. Still open
|
|
|
|
- **Feeding without SSH — the upstream ask that would close §1.3's gap.** Have the feeder send *content*
|
|
over MCP (`add_drawer` / `diary_write`, which it already calls) instead of asking the server to mine a
|
|
path it must first rsync there. HTTPS would then be genuinely sufficient and mining would work from any
|
|
network. Until then, mining is limited to devices that can reach synlig's SSH.
|
|
- **Per-device tokens** — Phase 4. Until then `origin_device` is advisory (§1.1).
|
|
- **§7.6 diary dedup** must be settled *before* the first §4.4 join; replay duplicates every entry.
|
|
- **§7.2**: never run `mempalace sync` against the shared palace. Doubly true now that the pi/opencode
|
|
feeders stage *inside* the palace root, which puts staged sources in scope for a sync of the palace dir.
|
|
- **nyvaken's public FQDN and the web hotel's CNAME rules** — unverified (§3.4).
|