From 3129d4f6b9778a7b5f7f8cbebca7ff9cf8e85a87 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 00:23:38 +0200 Subject: [PATCH] CHAOS NIGHT round 7: the cut is invisible at home, ten minutes long from away Round 7 (use nextcloud, internet cut ten minutes). What the customer saw depends entirely on where they stood: at home nothing at all - traefik answered 301 throughout and every app kept serving; away from home, ten minutes of 502, then 200 again about a minute after the block lifted. The box kept all 26 containers running, needed no repair, and re-established the way in unaided in ~64 seconds. No alarm fired, and none should have: the ladder puts node_stale at 30 minutes and this was ten. The question the round was meant to answer is recorded as NOT EXERCISED rather than passed. Are alarms raised while the hub is unreachable retried and then silently dropped? The box reports every 15m0s (measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports, so nothing was attempted and nothing could be lost. Rounds 8 and 9 are also ten-minute cuts and one should contain a scheduled report naturally - they are not re-timed to force it. The fence held, checked against a baseline taken BEFORE the round: demo-hp is back to -P FORWARD ACCEPT, zero physdev rules, sysctl 0. That host also carries guests 9201 and 9202, so an abandoned rule would have been a fence breach rather than an untidy drill. Also labelled honestly: the runner's own 530 reading was taken three seconds after unblocking and measures nothing about recovery. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/DRILL-chaos-night-2026-09-17.md | 30 +++++++- .../round-7-accident.txt | 3 + .../round-7.txt | 75 +++++++++++++++++++ 3 files changed, 107 insertions(+), 1 deletion(-) diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 432bf3b7..048ef0d0 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -309,7 +309,35 @@ falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nes stands, but I first wrote „I stopped the backup", which was wrong: I stopped **one leg**, and the off-site leg started by itself seconds later and succeeded. Counted as an intervention either way. -### Rounds 7-12 +### Round 7 — `use` nextcloud / accident: **internet cut for ten minutes** + +*Drawn as `update`; ran as `use`, because the catalog bump was never pushed — the catalog gates +returned INCONCLUSIVE and undetermined is never a pass. Decided and recorded before the round, not +after it.* + +| the five things | | +|---|---| +| what the customer saw | **depends where they stood.** At home: nothing — traefik answered `301` throughout and every app kept serving. Away from home: ten minutes of nothing — the public route gave **502**, then **200** about a minute after the block lifted | +| what the box did by itself | kept all **26** containers running, needed no repair, and re-established the way in unaided: 530 three seconds after unblocking, **all four apps 200 by 22:21:52Z — ~64 s** | +| time to steady | the apps never left steady; only the path in broke and healed, in **~64 s** | +| alarm fired / true? | **none, and none should have** — the ladder puts `node_stale` at 30 minutes and this was ten. Nothing false was raised either | +| should have fired, did not | **none** | + +**Household loop: 10 operations, 0 failures — narrower than it looks.** It does not follow redirects, +so it measured the box (up throughout) and was blind to the public outage. + +**The question this round was meant to answer, and honestly did not.** Are alarms raised while the hub +is unreachable retried and then silently dropped? **Not exercised.** The box reports every **15m0s** +(measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports — the next was due +~22:23:43Z, after it lifted. Nothing was attempted, so nothing could be lost. Recorded as *not +exercised*, never as *passed*. + +**The fence held, and it was checked against a baseline taken beforehand.** The accident flips a +host-wide sysctl on demo-hp and inserts two `physdev` rules. Afterwards: `-P FORWARD ACCEPT`, **0** +physdev rules, sysctl **0** — identical to the pre-round reading. demo-hp also carries guests 9201 +and 9202, so an abandoned rule would have been a fence breach, not an untidy drill. + +### Rounds 8-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt index 75796e29..719a1fc3 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt @@ -1,3 +1,6 @@ 2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7 2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up 2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes +2026-09-16T22:20:48Z unblocked; host sysctl restored to 0 and both rules removed +-P FORWARD ACCEPT +2026-09-16T22:20:48Z accident internet-gone-10min complete diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt index fa358e71..954eeb0e 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt @@ -108,3 +108,78 @@ at ~25-minute spacing against a 15-minute report cycle, so one of them will almo scheduled report and will exercise the drop behaviour properly. **The rounds are NOT re-timed to make that happen** — re-timing a round to obtain a better result is choosing the night after the fact. If it happens naturally, it is measured; if it does not, that is recorded too. + 2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7 + 2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up + 2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes + 2026-09-16T22:20:48Z unblocked; host sysctl restored to 0 and both rules removed + -P FORWARD ACCEPT + 2026-09-16T22:20:48Z accident internet-gone-10min complete +2026-09-16T22:20:48Z --- AFTER: what the box did BY ITSELF --- +2026-09-16T22:20:50Z t+603s containers=26 (before 26) +2026-09-16T22:20:50Z STEADY after 603s +2026-09-16T22:20:51Z front doors: cloud=530 status=530 paste=530 wiki=530 +2026-09-16T22:20:51Z household lines this round: 10 failures: 0 +2026-09-16T22:20:51Z --- alarms --- + | Time | Severity | Type | Message | Source + | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller + | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller + | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller + | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller + | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller + | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller +2026-09-16T22:20:51Z ================ END ROUND 7 ================ + +## The block, and the fence check that proves the host was left as it was found + 22:10:48Z blocked (LAN allowed, everything else dropped) — ten minutes + 22:20:48Z unblocked; host sysctl restored to 0 and both rules removed + 22:20:48Z accident internet-gone-10min complete + +**Verified independently on demo-hp at 22:21:08Z, against the baseline taken BEFORE the round:** + iptables -S FORWARD -> `-P FORWARD ACCEPT` (baseline: identical) + physdev rules -> **0** (baseline: 0) + net.bridge.bridge-nf-call-iptables = **0** (baseline: 0) + +Byte for byte the state I recorded before touching anything. That matters more than tidiness: demo-hp +is a Tier-0 host that also carries guests **9201 and 9202**, so an abandoned FORWARD rule would be a +fence breach rather than an untidy drill. The unconditional cleanup trap added before this round was +not needed in the end — the happy path ran — but it was the right insurance for a ten-minute sleep on +a shared host. + +## A reading that measures nothing, and is labelled as such +The runner's own front-door sample at **22:20:51Z** shows `cloud/status/paste/wiki = 530` — taken +**three seconds** after the block lifted, before Cloudflare could re-establish the tunnel. It is not a +recovery measurement and is not treated as one; the recovery time comes from a later reading. + +## ROUND 7 — the five things +**action:** `use` nextcloud (drawn as `update`; ran as `use` because the catalog bump was never +pushed — the catalog gates returned INCONCLUSIVE and undetermined is never a pass) +**accident:** the box's traffic blocked off-LAN for ten minutes (22:10:48Z → 22:20:48Z) + +1. **What the customer saw.** Depends entirely on where they were standing. **At home: nothing** — + traefik answered `301` on the box throughout and every app kept serving. **Away from home: ten + minutes of nothing** — the public route returned **502** while the tunnel was alive but could reach + nothing, then **200** again about a minute after the block lifted. +2. **What the box did by itself.** Kept all **26** containers running and needed no repair. When the + block lifted it re-established the way in **unaided**: 22:20:51Z still 530 (three seconds after + unblocking), **22:21:52Z all four apps 200** — so **~64 seconds** from network restored to front + doors serving. +3. **Time to steady.** The apps never left steady (26 before, 26 after). The only thing that broke and + healed was the path in: **~64 s**. +4. **Alarm fired / true?** **None fired, and none should have.** The ladder puts `node_stale` at 30 + minutes; a ten-minute outage is far inside that. Nothing false was raised either. +5. **Should have fired and did not.** **None.** + +**Household loop: 10 operations, 0 failures — and again narrower than it looks.** It does not follow +redirects, so it measured the box (up throughout) and was blind to the public outage. Its zero is not +evidence the household was unaffected. + +## The question this round was supposed to answer, and honestly did not +Are alarms raised while the hub is unreachable retried and then **silently dropped**? **Not +exercised.** The box reports every **15m0s** (measured: pushes at 21:53:49Z and 22:08:43Z), and the +block fell entirely between two reports — the next was due ~22:23:43Z, after it lifted. Nothing was +attempted, so nothing could be lost. Recorded as *not exercised*, never as *passed*. + +**Rounds 8 and 9 are also ten-minute cuts** and, at ~25-minute spacing against a 15-minute cycle, one +of them should contain a scheduled report naturally. They are not re-timed to force it.