diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 432bf3b7..048ef0d0 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -309,7 +309,35 @@ falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nes stands, but I first wrote „I stopped the backup", which was wrong: I stopped **one leg**, and the off-site leg started by itself seconds later and succeeded. Counted as an intervention either way. -### Rounds 7-12 +### Round 7 — `use` nextcloud / accident: **internet cut for ten minutes** + +*Drawn as `update`; ran as `use`, because the catalog bump was never pushed — the catalog gates +returned INCONCLUSIVE and undetermined is never a pass. Decided and recorded before the round, not +after it.* + +| the five things | | +|---|---| +| what the customer saw | **depends where they stood.** At home: nothing — traefik answered `301` throughout and every app kept serving. Away from home: ten minutes of nothing — the public route gave **502**, then **200** about a minute after the block lifted | +| what the box did by itself | kept all **26** containers running, needed no repair, and re-established the way in unaided: 530 three seconds after unblocking, **all four apps 200 by 22:21:52Z — ~64 s** | +| time to steady | the apps never left steady; only the path in broke and healed, in **~64 s** | +| alarm fired / true? | **none, and none should have** — the ladder puts `node_stale` at 30 minutes and this was ten. Nothing false was raised either | +| should have fired, did not | **none** | + +**Household loop: 10 operations, 0 failures — narrower than it looks.** It does not follow redirects, +so it measured the box (up throughout) and was blind to the public outage. + +**The question this round was meant to answer, and honestly did not.** Are alarms raised while the hub +is unreachable retried and then silently dropped? **Not exercised.** The box reports every **15m0s** +(measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports — the next was due +~22:23:43Z, after it lifted. Nothing was attempted, so nothing could be lost. Recorded as *not +exercised*, never as *passed*. + +**The fence held, and it was checked against a baseline taken beforehand.** The accident flips a +host-wide sysctl on demo-hp and inserts two `physdev` rules. Afterwards: `-P FORWARD ACCEPT`, **0** +physdev rules, sysctl **0** — identical to the pre-round reading. demo-hp also carries guests 9201 +and 9202, so an abandoned rule would have been a fence breach, not an untidy drill. + +### Rounds 8-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt index 75796e29..719a1fc3 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt @@ -1,3 +1,6 @@ 2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7 2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up 2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes +2026-09-16T22:20:48Z unblocked; host sysctl restored to 0 and both rules removed +-P FORWARD ACCEPT +2026-09-16T22:20:48Z accident internet-gone-10min complete diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt index fa358e71..954eeb0e 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt @@ -108,3 +108,78 @@ at ~25-minute spacing against a 15-minute report cycle, so one of them will almo scheduled report and will exercise the drop behaviour properly. **The rounds are NOT re-timed to make that happen** — re-timing a round to obtain a better result is choosing the night after the fact. If it happens naturally, it is measured; if it does not, that is recorded too. + 2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7 + 2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up + 2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes + 2026-09-16T22:20:48Z unblocked; host sysctl restored to 0 and both rules removed + -P FORWARD ACCEPT + 2026-09-16T22:20:48Z accident internet-gone-10min complete +2026-09-16T22:20:48Z --- AFTER: what the box did BY ITSELF --- +2026-09-16T22:20:50Z t+603s containers=26 (before 26) +2026-09-16T22:20:50Z STEADY after 603s +2026-09-16T22:20:51Z front doors: cloud=530 status=530 paste=530 wiki=530 +2026-09-16T22:20:51Z household lines this round: 10 failures: 0 +2026-09-16T22:20:51Z --- alarms --- + | Time | Severity | Type | Message | Source + | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller + | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller + | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller + | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller + | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller + | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller +2026-09-16T22:20:51Z ================ END ROUND 7 ================ + +## The block, and the fence check that proves the host was left as it was found + 22:10:48Z blocked (LAN allowed, everything else dropped) — ten minutes + 22:20:48Z unblocked; host sysctl restored to 0 and both rules removed + 22:20:48Z accident internet-gone-10min complete + +**Verified independently on demo-hp at 22:21:08Z, against the baseline taken BEFORE the round:** + iptables -S FORWARD -> `-P FORWARD ACCEPT` (baseline: identical) + physdev rules -> **0** (baseline: 0) + net.bridge.bridge-nf-call-iptables = **0** (baseline: 0) + +Byte for byte the state I recorded before touching anything. That matters more than tidiness: demo-hp +is a Tier-0 host that also carries guests **9201 and 9202**, so an abandoned FORWARD rule would be a +fence breach rather than an untidy drill. The unconditional cleanup trap added before this round was +not needed in the end — the happy path ran — but it was the right insurance for a ten-minute sleep on +a shared host. + +## A reading that measures nothing, and is labelled as such +The runner's own front-door sample at **22:20:51Z** shows `cloud/status/paste/wiki = 530` — taken +**three seconds** after the block lifted, before Cloudflare could re-establish the tunnel. It is not a +recovery measurement and is not treated as one; the recovery time comes from a later reading. + +## ROUND 7 — the five things +**action:** `use` nextcloud (drawn as `update`; ran as `use` because the catalog bump was never +pushed — the catalog gates returned INCONCLUSIVE and undetermined is never a pass) +**accident:** the box's traffic blocked off-LAN for ten minutes (22:10:48Z → 22:20:48Z) + +1. **What the customer saw.** Depends entirely on where they were standing. **At home: nothing** — + traefik answered `301` on the box throughout and every app kept serving. **Away from home: ten + minutes of nothing** — the public route returned **502** while the tunnel was alive but could reach + nothing, then **200** again about a minute after the block lifted. +2. **What the box did by itself.** Kept all **26** containers running and needed no repair. When the + block lifted it re-established the way in **unaided**: 22:20:51Z still 530 (three seconds after + unblocking), **22:21:52Z all four apps 200** — so **~64 seconds** from network restored to front + doors serving. +3. **Time to steady.** The apps never left steady (26 before, 26 after). The only thing that broke and + healed was the path in: **~64 s**. +4. **Alarm fired / true?** **None fired, and none should have.** The ladder puts `node_stale` at 30 + minutes; a ten-minute outage is far inside that. Nothing false was raised either. +5. **Should have fired and did not.** **None.** + +**Household loop: 10 operations, 0 failures — and again narrower than it looks.** It does not follow +redirects, so it measured the box (up throughout) and was blind to the public outage. Its zero is not +evidence the household was unaffected. + +## The question this round was supposed to answer, and honestly did not +Are alarms raised while the hub is unreachable retried and then **silently dropped**? **Not +exercised.** The box reports every **15m0s** (measured: pushes at 21:53:49Z and 22:08:43Z), and the +block fell entirely between two reports — the next was due ~22:23:43Z, after it lifted. Nothing was +attempted, so nothing could be lost. Recorded as *not exercised*, never as *passed*. + +**Rounds 8 and 9 are also ten-minute cuts** and, at ~25-minute spacing against a 15-minute cycle, one +of them should contain a scheduled report naturally. They are not re-timed to force it.