ssh sidecar: default to multiplexing, as a default and not an override
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 16s

A target whose ~/.ssh/config entry never mentioned ControlMaster got no
multiplexing from the sidecar (only ControlPath was supplied), so every ssh call
opened a fresh TCP connection. On 2026-08-25 that produced ~12 connections to
one host in 15 min and a fail2ban block that looked like an outage — the tell
being that HTTPS to the same estate stayed healthy.

The correctness of this depends entirely on WHERE the block goes. ssh_config is
first-value-wins:

  ControlPath   before the Include -> override (the user's value points at
                read-only ~/.ssh and cannot work in the container)
  ControlMaster after  the Include -> default  (an explicit per-host
                'ControlMaster no' must keep winning)

Force what is broken, default what is merely absent. The first draft put both in
the leading block and would have silently overridden an explicit 'no'.

Verified with ssh -G rather than from the man page, including the counterfactual:
under the shipped layout an explicit 'no' resolves to controlmaster false while a
silent host resolves to auto; under the rejected layout the 'no' host flips to
auto. So the test discriminates position, not presence. Plus a sandbox render of
the real script, bash -n, and shellcheck -S error (the v1.8.7 gate) clean.

Effect measured on 41 real host aliases: 22 silent entries gain auto+10m, 0
overridden. Note the fleet's one deliberate opt-out is written as absence plus a
comment ('# No ControlMaster — VPN means direct route'), which ssh cannot
distinguish from no opinion; that host now multiplexes, which its own comment
says is unnecessary rather than harmful.

Skill documents the sidecar-vs-~/.ssh trap (the failure misleads: read-only
ControlPath makes multiplexing look impossible rather than misconfigured) and
the stale-master recovery, ssh -O check / -O exit.
This commit is contained in:
pi
2026-08-25 23:09:46 +02:00
parent ebd0de0be2
commit 657b1ad856
3 changed files with 130 additions and 0 deletions
@@ -185,6 +185,16 @@ entrypoint's `setup-lan-access.sh` writes a **writable SSH sidecar** at
- A `Host *` block redirecting `ControlPath` into the writable `~/.ssh-local/cm`
(because `~/.ssh` is typically bind-mounted **read-only**, so a master socket
can't be created under it), plus `Include ~/.ssh/config`.
- A **trailing** `Host *` block supplying `ControlMaster auto` + `ControlPersist
10m` as a *default*. Position is the design: `ControlPath` sits **before** the
`Include` (an override — the value in your own config points at read-only
`~/.ssh` and cannot work here), while `ControlMaster` sits **after** it (a
default — an explicit per-host `ControlMaster no`/`auto` in your own config
still wins, because ssh_config is first-value-wins). **Force what is broken,
default what is merely absent.** Without this, a target whose entry never
mentioned `ControlMaster` opens a fresh TCP connection per `ssh` call, and an
agent making a dozen calls in a few minutes can trip fail2ban or a CGNAT
flow-table cap on the far end.
- Aliases **`host` / `mac`** → `host.docker.internal` (user comes from
`HOST_SSH_USER`) — i.e. SSH back into the Docker host.
- On VM-backed hosts only: an **SSH-jump-via-host** block so the container can
@@ -199,6 +209,25 @@ ssh -F "$HOME/.ssh-local/config" mac 'hostname; whoami' # reach the host
ssh -F "$HOME/.ssh-local/config" <lan-peer> '…' # reach a LAN peer (if configured)
```
**Always go through the sidecar, never `-F ~/.ssh/config`.** This is the single
easiest way to break SSH from inside the container, and the failure actively
misleads: the read-only path makes the master socket uncreatable, so
multiplexing appears *impossible* rather than misconfigured. What follows is a
burst of fresh connections and, on a rate-limiting peer, a block that looks like
an outage. The tell that it is rate-limiting and not an outage: HTTPS to the same
estate keeps working while port 22 stops answering. (Recorded 2026-08-25 — an
agent hit exactly this, concluded "ControlMaster is impossible here", disabled
multiplexing, and filed that as a lesson. The sidecar had solved it since v1.4.)
If every `ssh` to one host suddenly hangs, suspect a **stale master** — socket
file present, daemon gone, typically after the host suspended or changed
network. Check and clear it:
```sh
ssh -F "$HOME/.ssh-local/config" -O check <host> # "Master running (pid=…)" or no master
ssh -F "$HOME/.ssh-local/config" -O exit <host> # tear down a stale one
```
Two related mechanisms (don't reinvent them):
- **ControlMaster multiplexing** is preconfigured (`/tmp/sshcm/`) to survive