CHAOS NIGHT round 4: the box repairs its own tunnel in 97 seconds
gates / gates (push) Successful in 22s

Killed cloudflared and deliberately never restarted it, because whether it
returns by itself IS the measurement. It returned in ~97 seconds (killed
~21:41:30Z, running again 21:43:07.478Z), and the box did it, not Docker:
RestartCount=0 proves the unless-stopped policy never acted, and the controller
log shows "[infra] deploying cloudflared -> /opt/docker/stacks/cloudflared".
That is the product's protected-infra recovery repairing one of its own
infrastructure stacks unasked.

health_critical (error) fired at 21:43 - "Rendszer allapot kritikus (volt: ok)"
- which is exactly what the ladder predicts for a missing protected container,
and it is true. Whether health_recovered closes the pair is checked at the end
of the round, not guessed at now.

A METHOD CORRECTION that retro-labels every front-door reading tonight: the 530
during the outage is a Cloudflare status, which exposed that `curl -sL` was
following traefik's 301 out to the public hostname and back down the tunnel. So
every "front door" reading so far measured the PUBLIC path, not the LAN. It does
not invalidate the readings - a 200 by that route proves more, not less - but it
invalidates the label, and with it any claim of the form "the app is fine, only
the tunnel is down". The two paths are now measured separately: during this
outage the public route gave 530 while traefik answered 301 locally throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 23:44:54 +02:00
parent 34d22a1a92
commit ec84eadc19
@@ -4,3 +4,53 @@
2026-09-16T21:41:29Z --- ACTION: offsite-run on mealie ---
ABORT: no session — mine, not the product's
2026-09-16T21:41:32Z --- ACCIDENT: tunnel-down-10min (injected after the action started) ---
## A METHOD CORRECTION that retro-labels every front-door reading tonight
During the tunnel outage the front doors returned **530**. That is a **Cloudflare** status, not one
traefik would produce — which exposed what my sweep was actually measuring.
`curl -sL -H "Host: x.enkicsifelhom.hu" http://192.168.0.116/` gets a **301** from traefik and then
**follows it** to `https://x.enkicsifelhom.hu/`, which resolves through public DNS to Cloudflare and
comes back down the tunnel. So every „front door" reading I have taken tonight — the 200s in Phase 0,
the 10-of-12 and 11-of-12 sweeps, the per-round `door()` calls — has been measuring the **PUBLIC
path**: box -> tunnel -> Cloudflare -> back. Not the LAN.
**What that does and does not invalidate.** It does NOT invalidate the readings: a 200 by that route
is a stronger statement than a LAN 200, because it proves traefik, the app, the tunnel and Cloudflare
all worked. What it invalidates is the LABEL „from DooPlex over the LAN", and with it any conclusion
of the form „the app is fine, only the tunnel is down" — that distinction was never being drawn.
**Corrected method, used from here on:** the two paths are measured separately —
* LAN: no redirect following, straight to the box
* public: the hostname pinned to the box, or the ordinary public route
so „the app serves" and „the tunnel serves" are two answers, not one.
This is the eighth instrument error of mine tonight and the most consequential, because it silently
changed what a green reading MEANT rather than producing an obviously wrong one.
## The tunnel came back BY ITSELF in ~97 seconds — and the CONTROLLER did it, not Docker
killed ~21:41:30Z (`docker kill cloudflared`, deliberately never restarted by hand)
cloudflared Exited (137) at 21:42:44Z, 25 of 26 containers running
back up **StartedAt 2026-09-16T21:43:07.478Z**, Status=running, 26 containers
elapsed **~97 seconds**
**Who restarted it is answerable, and it matters.** `RestartCount=**0**` — Docker's own
`unless-stopped` policy did NOT act, because that would have incremented the counter. The controller's
log says what did:
21:43:07 [INFO] [infra] deploying cloudflared → /opt/docker/stacks/cloudflared
i.e. the product's **protected-infra recovery** noticed a protected stack was gone and redeployed it.
That is a stronger result than „docker restarted it": the box repaired one of its own infrastructure
stacks without anyone asking.
## The alarm the ladder predicted, and it fired TRUE
Sep 16 21:43 **error health_critical** „Rendszer állapot kritikus (volt: ok)"
A missing protected container forces the health verdict to fail, and `NotifyHealthChange` is
edge-triggered on the rank increase — so exactly one alarm, at error severity, customer-facing and ON
by default. It fired within the outage rather than after it, and it is TRUE: the tunnel really was
down and the box really was not fully healthy.
**Front doors during the outage, measured on both paths (see the method correction above):**
public route (box -> tunnel -> Cloudflare) **530** at 21:42:44Z, **200** again at 21:43:15Z
LAN route straight to the box traefik answered **301** throughout
So the household's apps never stopped serving locally; what failed, and recovered, was the way in
from outside.