diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt index d782c257..304a7dea 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt @@ -4,3 +4,53 @@ 2026-09-16T21:41:29Z --- ACTION: offsite-run on mealie --- ABORT: no session — mine, not the product's 2026-09-16T21:41:32Z --- ACCIDENT: tunnel-down-10min (injected after the action started) --- + +## A METHOD CORRECTION that retro-labels every front-door reading tonight +During the tunnel outage the front doors returned **530**. That is a **Cloudflare** status, not one +traefik would produce — which exposed what my sweep was actually measuring. + +`curl -sL -H "Host: x.enkicsifelhom.hu" http://192.168.0.116/` gets a **301** from traefik and then +**follows it** to `https://x.enkicsifelhom.hu/`, which resolves through public DNS to Cloudflare and +comes back down the tunnel. So every „front door" reading I have taken tonight — the 200s in Phase 0, +the 10-of-12 and 11-of-12 sweeps, the per-round `door()` calls — has been measuring the **PUBLIC +path**: box -> tunnel -> Cloudflare -> back. Not the LAN. + +**What that does and does not invalidate.** It does NOT invalidate the readings: a 200 by that route +is a stronger statement than a LAN 200, because it proves traefik, the app, the tunnel and Cloudflare +all worked. What it invalidates is the LABEL „from DooPlex over the LAN", and with it any conclusion +of the form „the app is fine, only the tunnel is down" — that distinction was never being drawn. + +**Corrected method, used from here on:** the two paths are measured separately — + * LAN: no redirect following, straight to the box + * public: the hostname pinned to the box, or the ordinary public route +so „the app serves" and „the tunnel serves" are two answers, not one. + +This is the eighth instrument error of mine tonight and the most consequential, because it silently +changed what a green reading MEANT rather than producing an obviously wrong one. + +## The tunnel came back BY ITSELF in ~97 seconds — and the CONTROLLER did it, not Docker + killed ~21:41:30Z (`docker kill cloudflared`, deliberately never restarted by hand) + cloudflared Exited (137) at 21:42:44Z, 25 of 26 containers running + back up **StartedAt 2026-09-16T21:43:07.478Z**, Status=running, 26 containers + elapsed **~97 seconds** + +**Who restarted it is answerable, and it matters.** `RestartCount=**0**` — Docker's own +`unless-stopped` policy did NOT act, because that would have incremented the counter. The controller's +log says what did: + 21:43:07 [INFO] [infra] deploying cloudflared → /opt/docker/stacks/cloudflared +i.e. the product's **protected-infra recovery** noticed a protected stack was gone and redeployed it. +That is a stronger result than „docker restarted it": the box repaired one of its own infrastructure +stacks without anyone asking. + +## The alarm the ladder predicted, and it fired TRUE + Sep 16 21:43 **error health_critical** „Rendszer állapot kritikus (volt: ok)" +A missing protected container forces the health verdict to fail, and `NotifyHealthChange` is +edge-triggered on the rank increase — so exactly one alarm, at error severity, customer-facing and ON +by default. It fired within the outage rather than after it, and it is TRUE: the tunnel really was +down and the box really was not fully healthy. + +**Front doors during the outage, measured on both paths (see the method correction above):** + public route (box -> tunnel -> Cloudflare) **530** at 21:42:44Z, **200** again at 21:43:15Z + LAN route straight to the box traefik answered **301** throughout +So the household's apps never stopped serving locally; what failed, and recovered, was the way in +from outside.