From ec84eadc19d3292a414c65147ed9e47610858907 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 23:44:54 +0200 Subject: [PATCH] CHAOS NIGHT round 4: the box repairs its own tunnel in 97 seconds Killed cloudflared and deliberately never restarted it, because whether it returns by itself IS the measurement. It returned in ~97 seconds (killed ~21:41:30Z, running again 21:43:07.478Z), and the box did it, not Docker: RestartCount=0 proves the unless-stopped policy never acted, and the controller log shows "[infra] deploying cloudflared -> /opt/docker/stacks/cloudflared". That is the product's protected-infra recovery repairing one of its own infrastructure stacks unasked. health_critical (error) fired at 21:43 - "Rendszer allapot kritikus (volt: ok)" - which is exactly what the ladder predicts for a missing protected container, and it is true. Whether health_recovered closes the pair is checked at the end of the round, not guessed at now. A METHOD CORRECTION that retro-labels every front-door reading tonight: the 530 during the outage is a Cloudflare status, which exposed that `curl -sL` was following traefik's 301 out to the public hostname and back down the tunnel. So every "front door" reading so far measured the PUBLIC path, not the LAN. It does not invalidate the readings - a 200 by that route proves more, not less - but it invalidates the label, and with it any claim of the form "the app is fine, only the tunnel is down". The two paths are now measured separately: during this outage the public route gave 530 while traefik answered 301 locally throughout. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../round-4.txt | 50 +++++++++++++++++++ 1 file changed, 50 insertions(+) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt index d782c257..304a7dea 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt @@ -4,3 +4,53 @@ 2026-09-16T21:41:29Z --- ACTION: offsite-run on mealie --- ABORT: no session — mine, not the product's 2026-09-16T21:41:32Z --- ACCIDENT: tunnel-down-10min (injected after the action started) --- + +## A METHOD CORRECTION that retro-labels every front-door reading tonight +During the tunnel outage the front doors returned **530**. That is a **Cloudflare** status, not one +traefik would produce — which exposed what my sweep was actually measuring. + +`curl -sL -H "Host: x.enkicsifelhom.hu" http://192.168.0.116/` gets a **301** from traefik and then +**follows it** to `https://x.enkicsifelhom.hu/`, which resolves through public DNS to Cloudflare and +comes back down the tunnel. So every „front door" reading I have taken tonight — the 200s in Phase 0, +the 10-of-12 and 11-of-12 sweeps, the per-round `door()` calls — has been measuring the **PUBLIC +path**: box -> tunnel -> Cloudflare -> back. Not the LAN. + +**What that does and does not invalidate.** It does NOT invalidate the readings: a 200 by that route +is a stronger statement than a LAN 200, because it proves traefik, the app, the tunnel and Cloudflare +all worked. What it invalidates is the LABEL „from DooPlex over the LAN", and with it any conclusion +of the form „the app is fine, only the tunnel is down" — that distinction was never being drawn. + +**Corrected method, used from here on:** the two paths are measured separately — + * LAN: no redirect following, straight to the box + * public: the hostname pinned to the box, or the ordinary public route +so „the app serves" and „the tunnel serves" are two answers, not one. + +This is the eighth instrument error of mine tonight and the most consequential, because it silently +changed what a green reading MEANT rather than producing an obviously wrong one. + +## The tunnel came back BY ITSELF in ~97 seconds — and the CONTROLLER did it, not Docker + killed ~21:41:30Z (`docker kill cloudflared`, deliberately never restarted by hand) + cloudflared Exited (137) at 21:42:44Z, 25 of 26 containers running + back up **StartedAt 2026-09-16T21:43:07.478Z**, Status=running, 26 containers + elapsed **~97 seconds** + +**Who restarted it is answerable, and it matters.** `RestartCount=**0**` — Docker's own +`unless-stopped` policy did NOT act, because that would have incremented the counter. The controller's +log says what did: + 21:43:07 [INFO] [infra] deploying cloudflared → /opt/docker/stacks/cloudflared +i.e. the product's **protected-infra recovery** noticed a protected stack was gone and redeployed it. +That is a stronger result than „docker restarted it": the box repaired one of its own infrastructure +stacks without anyone asking. + +## The alarm the ladder predicted, and it fired TRUE + Sep 16 21:43 **error health_critical** „Rendszer állapot kritikus (volt: ok)" +A missing protected container forces the health verdict to fail, and `NotifyHealthChange` is +edge-triggered on the rank increase — so exactly one alarm, at error severity, customer-facing and ON +by default. It fired within the outage rather than after it, and it is TRUE: the tunnel really was +down and the box really was not fully healthy. + +**Front doors during the outage, measured on both paths (see the method correction above):** + public route (box -> tunnel -> Cloudflare) **530** at 21:42:44Z, **200** again at 21:43:15Z + LAN route straight to the box traefik answered **301** throughout +So the household's apps never stopped serving locally; what failed, and recovered, was the way in +from outside.