CHAOS NIGHT round 7: the cut is invisible at home, ten minutes long from away
gates / gates (push) Successful in 21s

Round 7 (use nextcloud, internet cut ten minutes). What the customer saw depends
entirely on where they stood: at home nothing at all - traefik answered 301
throughout and every app kept serving; away from home, ten minutes of 502, then
200 again about a minute after the block lifted. The box kept all 26 containers
running, needed no repair, and re-established the way in unaided in ~64 seconds.

No alarm fired, and none should have: the ladder puts node_stale at 30 minutes
and this was ten.

The question the round was meant to answer is recorded as NOT EXERCISED rather
than passed. Are alarms raised while the hub is unreachable retried and then
silently dropped? The box reports every 15m0s (measured: 21:53:49Z, 22:08:43Z)
and the cut fell entirely between two reports, so nothing was attempted and
nothing could be lost. Rounds 8 and 9 are also ten-minute cuts and one should
contain a scheduled report naturally - they are not re-timed to force it.

The fence held, checked against a baseline taken BEFORE the round: demo-hp is
back to -P FORWARD ACCEPT, zero physdev rules, sysctl 0. That host also carries
guests 9201 and 9202, so an abandoned rule would have been a fence breach rather
than an untidy drill.

Also labelled honestly: the runner's own 530 reading was taken three seconds
after unblocking and measures nothing about recovery.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 00:23:38 +02:00
parent e61aac1d8f
commit 3129d4f6b9
3 changed files with 107 additions and 1 deletions
@@ -309,7 +309,35 @@ falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nes
stands, but I first wrote „I stopped the backup", which was wrong: I stopped **one leg**, and the
off-site leg started by itself seconds later and succeeded. Counted as an intervention either way.
### Rounds 7-12
### Round 7 — `use` nextcloud / accident: **internet cut for ten minutes**
*Drawn as `update`; ran as `use`, because the catalog bump was never pushed — the catalog gates
returned INCONCLUSIVE and undetermined is never a pass. Decided and recorded before the round, not
after it.*
| the five things | |
|---|---|
| what the customer saw | **depends where they stood.** At home: nothing — traefik answered `301` throughout and every app kept serving. Away from home: ten minutes of nothing — the public route gave **502**, then **200** about a minute after the block lifted |
| what the box did by itself | kept all **26** containers running, needed no repair, and re-established the way in unaided: 530 three seconds after unblocking, **all four apps 200 by 22:21:52Z — ~64 s** |
| time to steady | the apps never left steady; only the path in broke and healed, in **~64 s** |
| alarm fired / true? | **none, and none should have** — the ladder puts `node_stale` at 30 minutes and this was ten. Nothing false was raised either |
| should have fired, did not | **none** |
**Household loop: 10 operations, 0 failures — narrower than it looks.** It does not follow redirects,
so it measured the box (up throughout) and was blind to the public outage.
**The question this round was meant to answer, and honestly did not.** Are alarms raised while the hub
is unreachable retried and then silently dropped? **Not exercised.** The box reports every **15m0s**
(measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports — the next was due
~22:23:43Z, after it lifted. Nothing was attempted, so nothing could be lost. Recorded as *not
exercised*, never as *passed*.
**The fence held, and it was checked against a baseline taken beforehand.** The accident flips a
host-wide sysctl on demo-hp and inserts two `physdev` rules. Afterwards: `-P FORWARD ACCEPT`, **0**
physdev rules, sysctl **0** — identical to the pre-round reading. demo-hp also carries guests 9201
and 9202, so an abandoned rule would have been a fence breach, not an untidy drill.
### Rounds 8-12
PENDING
@@ -1,3 +1,6 @@
2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7
2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up
2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes
2026-09-16T22:20:48Z unblocked; host sysctl restored to 0 and both rules removed
-P FORWARD ACCEPT
2026-09-16T22:20:48Z accident internet-gone-10min complete
@@ -108,3 +108,78 @@ at ~25-minute spacing against a 15-minute report cycle, so one of them will almo
scheduled report and will exercise the drop behaviour properly. **The rounds are NOT re-timed to make
that happen** — re-timing a round to obtain a better result is choosing the night after the fact. If
it happens naturally, it is measured; if it does not, that is recorded too.
2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7
2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up
2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes
2026-09-16T22:20:48Z unblocked; host sysctl restored to 0 and both rules removed
-P FORWARD ACCEPT
2026-09-16T22:20:48Z accident internet-gone-10min complete
2026-09-16T22:20:48Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T22:20:50Z t+603s containers=26 (before 26)
2026-09-16T22:20:50Z STEADY after 603s
2026-09-16T22:20:51Z front doors: cloud=530 status=530 paste=530 wiki=530
2026-09-16T22:20:51Z household lines this round: 10 failures: 0
2026-09-16T22:20:51Z --- alarms ---
| Time | Severity | Type | Message | Source
| Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
| Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
2026-09-16T22:20:51Z ================ END ROUND 7 ================
## The block, and the fence check that proves the host was left as it was found
22:10:48Z blocked (LAN allowed, everything else dropped) — ten minutes
22:20:48Z unblocked; host sysctl restored to 0 and both rules removed
22:20:48Z accident internet-gone-10min complete
**Verified independently on demo-hp at 22:21:08Z, against the baseline taken BEFORE the round:**
iptables -S FORWARD -> `-P FORWARD ACCEPT` (baseline: identical)
physdev rules -> **0** (baseline: 0)
net.bridge.bridge-nf-call-iptables = **0** (baseline: 0)
Byte for byte the state I recorded before touching anything. That matters more than tidiness: demo-hp
is a Tier-0 host that also carries guests **9201 and 9202**, so an abandoned FORWARD rule would be a
fence breach rather than an untidy drill. The unconditional cleanup trap added before this round was
not needed in the end — the happy path ran — but it was the right insurance for a ten-minute sleep on
a shared host.
## A reading that measures nothing, and is labelled as such
The runner's own front-door sample at **22:20:51Z** shows `cloud/status/paste/wiki = 530` — taken
**three seconds** after the block lifted, before Cloudflare could re-establish the tunnel. It is not a
recovery measurement and is not treated as one; the recovery time comes from a later reading.
## ROUND 7 — the five things
**action:** `use` nextcloud (drawn as `update`; ran as `use` because the catalog bump was never
pushed — the catalog gates returned INCONCLUSIVE and undetermined is never a pass)
**accident:** the box's traffic blocked off-LAN for ten minutes (22:10:48Z → 22:20:48Z)
1. **What the customer saw.** Depends entirely on where they were standing. **At home: nothing** —
traefik answered `301` on the box throughout and every app kept serving. **Away from home: ten
minutes of nothing** — the public route returned **502** while the tunnel was alive but could reach
nothing, then **200** again about a minute after the block lifted.
2. **What the box did by itself.** Kept all **26** containers running and needed no repair. When the
block lifted it re-established the way in **unaided**: 22:20:51Z still 530 (three seconds after
unblocking), **22:21:52Z all four apps 200** — so **~64 seconds** from network restored to front
doors serving.
3. **Time to steady.** The apps never left steady (26 before, 26 after). The only thing that broke and
healed was the path in: **~64 s**.
4. **Alarm fired / true?** **None fired, and none should have.** The ladder puts `node_stale` at 30
minutes; a ten-minute outage is far inside that. Nothing false was raised either.
5. **Should have fired and did not.** **None.**
**Household loop: 10 operations, 0 failures — and again narrower than it looks.** It does not follow
redirects, so it measured the box (up throughout) and was blind to the public outage. Its zero is not
evidence the household was unaffected.
## The question this round was supposed to answer, and honestly did not
Are alarms raised while the hub is unreachable retried and then **silently dropped**? **Not
exercised.** The box reports every **15m0s** (measured: pushes at 21:53:49Z and 22:08:43Z), and the
block fell entirely between two reports — the next was due ~22:23:43Z, after it lifted. Nothing was
attempted, so nothing could be lost. Recorded as *not exercised*, never as *passed*.
**Rounds 8 and 9 are also ten-minute cuts** and, at ~25-minute spacing against a 15-minute cycle, one
of them should contain a scheduled report naturally. They are not re-timed to force it.