chaos night round 8: the accident did not do what its name said
gates / gates (push) Successful in 20s

Round 8 (backup-app nextcloud + internet cut) passed on the product side:
26 containers throughout, LAN doors served the whole ten minutes, public
path restored unaided in <=43 s, no alarm fired and none should have.

The finding is against my own instrument. The hub report due at 22:38:43Z
fell inside the cut and SUCCEEDED, because hub.felhom.eu resolves to a LAN
address (192.168.0.192, measured from guest and host) and the injector
allowed the whole LAN. So rounds 7 and 8 never tested hub unreachability,
and the dropped-event behaviour is still unmeasured.

Two fixes, both to the harness, neither to the product:
  * inject.sh now blocks the hub address from the VM's side (the host tap
    rule). The hub itself is untouched - the fence is kept. It refuses to
    inject at all if it cannot resolve the hub.
  * run_round.sh no longer reads the front doors 3 s after an unblock. Every
    door reading now states its own timestamp and a second reading is taken
    60 s later. This was the sixth mistimed reading of the night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 00:52:40 +02:00
parent e45fb5e37f
commit 418f3a2c20
4 changed files with 143 additions and 6 deletions
@@ -337,6 +337,48 @@ host-wide sysctl on demo-hp and inserts two `physdev` rules. Afterwards: `-P FOR
physdev rules, sysctl **0** — identical to the pre-round reading. demo-hp also carries guests 9201
and 9202, so an abandoned rule would have been a fence breach, not an untidy drill.
### Round 8 — `backup-app` nextcloud / accident: **internet cut for ten minutes**
**22:35:48Z–22:47:59Z.** The app-data backup ran first and finished in 1 m 55 s
(`db_dump` count 5, `success:true`, 22:37:48Z). The cut began six seconds later, so the two barely
overlapped — and would not have interacted in any case: `backup-app` is the **local** app-data tier
and needs no internet. The off-site tier is a different action.
| the five things | |
|---|---|
| what the customer saw | **Nothing at home.** Every front door kept serving on the LAN for the whole ten minutes. From outside the house the sites were unreachable — the public path was down. 26 apps up before, 26 after, never fewer. |
| what the box did by itself | Kept every container running, kept backing up, kept reporting to the hub, and rebuilt the public path unaided when the link returned. No restart, no intervention, noaction from me. |
| time to steady | **≤43 s** after the link returned (public doors 200 again at 22:48:38Z). Containers never left steady at all. |
| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold and this was ten. Nothing false was raised. |
| should have fired, did not | **none** |
**And the finding of the round is against my own instrument, not the box.**
The hub report due at **22:38:43Z fell inside the cut** — and it **succeeded**:
„Hub report pushed successfully (15526 bytes)". It succeeded because `hub.felhom.eu` resolves to
**192.168.0.192**, a LAN address (measured from both the guest and the host), and my injector blocks
everything **except** the LAN. So the accident named „internet gone" only ever removed the **public**
path. The box never lost the hub, in round 7 or in round 8.
Two consequences, both stated plainly:
1. **The dropped-event question is still unmeasured** after two rounds that appeared to measure it.
Events pushed while the hub is unreachable are retried three times and then dropped permanently,
with no queue — that behaviour has still never been seen live.
2. **My own memory file carries this exact warning** („hub.felhom.eu resolves to the LAN here; an
internet-cut drill must block it too") and I did not apply it. A warning that is written down and
not read is worth nothing, which is the same class of failure as an unread alarm.
The injector is corrected for round 9's drawn internet cut so that the hub address is blocked too —
**from the VM's side, at the host's tap rule.** The hub itself is never touched; the fence is kept.
This is a repair to a broken instrument, not a re-draw: the drawn action, app and accident for every
remaining round are unchanged.
**A sixth mistimed reading, and the fix is in the runner now.** The round's own AFTER step read the
public doors **three seconds** after the unblock and reported 530 on all four. That reading could
never have been fair. The re-measurement 43 s later read 200 on all four. The runner measured the
recovery before the recovery could begin — so the runner is the thing that gets fixed, not the note.
### Rounds 8-12
PENDING
@@ -45,23 +45,28 @@ case "$A" in
internet-gone-10min)
TAP=$(H "ls /sys/class/net | grep -E \"^tap${VM}i0$\"")
[ -n "$TAP" ] || { say "NO TAP FOUND for VM $VM — accident NOT injected, and that is recorded as such"; exit 1; }
say "blocking the box's traffic off-LAN at the HOST, on $TAP; the LAN stays up"
HUBIP=$(H "getent hosts hub.felhom.eu | awk '{print \$1}' | head -1" | tr -d ' \r')
[ -n "$HUBIP" ] || { say "HUB ADDRESS NOT RESOLVABLE - accident NOT injected. An instrument that cannot find its target must not pretend it blocked it."; exit 1; }
say "blocking the box traffic off-LAN at the HOST, on $TAP; the LAN stays up EXCEPT the hub ($HUBIP)"
H "sysctl -w net.bridge.bridge-nf-call-iptables=1 >/dev/null
iptables -I FORWARD 1 -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT
iptables -I FORWARD 2 -m physdev --physdev-in $TAP -j DROP"
iptables -I FORWARD 2 -m physdev --physdev-in $TAP -j DROP
iptables -I FORWARD 1 -m physdev --physdev-in $TAP -d $HUBIP -j DROP"
# UNCONDITIONAL cleanup: if this script is killed during the sleep, or the SSH drops, the host
# must NOT be left with the rules in place and the sysctl flipped. demo-hp is a Tier-0 host that
# 9201 and 9202 also live on, and an abandoned FORWARD rule is a fence breach, not a measurement.
netcleanup(){
H "iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP 2>/dev/null
H "iptables -D FORWARD -m physdev --physdev-in $TAP -d $HUBIP -j DROP 2>/dev/null
iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP 2>/dev/null
iptables -D FORWARD -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT 2>/dev/null
sysctl -w net.bridge.bridge-nf-call-iptables=0 >/dev/null 2>&1" >/dev/null 2>&1
say "net cleanup ran (rules removed, sysctl restored)"
}
trap 'netcleanup' EXIT INT TERM
say "blocked (LAN allowed, everything else dropped) — 10 minutes"
say "blocked (LAN allowed EXCEPT the hub at $HUBIP, everything else dropped) - 10 minutes"
sleep 600
H "iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP
H "iptables -D FORWARD -m physdev --physdev-in $TAP -d $HUBIP -j DROP
iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP
iptables -D FORWARD -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT
sysctl -w net.bridge.bridge-nf-call-iptables=0 >/dev/null"
trap - EXIT INT TERM
@@ -0,0 +1,83 @@
-rwxrwxr-x 1 kisfenyo kisfenyo 5650 Sep 16 23:42 /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh
-rwxrwxr-x 1 kisfenyo kisfenyo 5712 Sep 16 23:52 /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh
waiter armed; will launch round 8 at 2026-09-16T22:35:48Z
=== round 8 launched 2026-09-16T22:35:48Z (due 22:35:48Z) ===
2026-09-16T22:35:48Z ================ ROUND 8 : backup-app nextcloud, while: internet-gone-10min ================
2026-09-16T22:35:50Z --- BEFORE --- containers=26 cloud=200 status=200 paste=200
2026-09-16T22:35:50Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T22:35:50Z --- ACTION: backup-app on nextcloud ---
POST /api/backup/run -> 200
{"ok":true,"message":"Mentés elindítva"}
{"ok":true,"data":{"enabled":true,"running":true}}
{"ok":true,"data":{"enabled":true,"running":true}}
{"ok":true,"data":{"enabled":true,"running":true}}
{"ok":true,"data":{"enabled":true,"running":true}}
{"ok":true,"data":{"enabled":true,"running":true}}
{"ok":true,"data":{"db_dump":{"count":5,"duration":"1m55.578787777s","last_run":"2026-09-16T22:37:48.900952183Z","success":true},"enabled":true,"running":false}}
2026-09-16T22:37:54Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---
2026-09-16T22:37:54Z ACCIDENT=internet-gone-10min round=8
2026-09-16T22:37:55Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up
2026-09-16T22:37:55Z blocked (LAN allowed, everything else dropped) — 10 minutes
2026-09-16T22:47:55Z unblocked; host sysctl restored to 0 and both rules removed
-P FORWARD ACCEPT
2026-09-16T22:47:55Z accident internet-gone-10min complete
2026-09-16T22:47:55Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T22:47:57Z t+603s containers=26 (before 26)
2026-09-16T22:47:57Z STEADY after 603s
2026-09-16T22:47:58Z front doors: cloud=530 status=530 paste=530 wiki=530
2026-09-16T22:47:58Z household lines this round: 12 failures: 0
2026-09-16T22:47:58Z --- alarms ---
| Time | Severity | Type | Message | Source
| Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
| Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
2026-09-16T22:47:59Z ================ END ROUND 8 ================
[exited with code 0]
## POST-ROUND CORRECTIONS, written immediately (2026-09-16T22:48Z)
### 1. The "530 on every door" line above is a MISTIMED READING, not a fault.
It was taken 3 seconds after the unblock (22:47:58Z; unblock 22:47:55Z). Re-measured at
22:48:38Z, 43 s after the unblock:
cloud LAN=301 public=200
status LAN=301 public=200
paste LAN=301 public=200
wiki LAN=301 public=200
So the public path came back by itself in <=43 s. This is the SIXTH mistimed reading tonight and
the runner's own AFTER step is the one still doing it -- it measures the public door 3 s after the
network returns, which can never be a fair reading. Fix applied below.
### 2. THE ACCIDENT DID NOT DO WHAT ITS NAME SAYS. This is my fault, and it is the bigger finding.
The hub report due at 22:38:43Z fell INSIDE the ten-minute cut (22:37:55 -> 22:47:55).
It SUCCEEDED: "[report] Hub report pushed successfully (15526 bytes)".
It succeeded because hub.felhom.eu resolves to a LAN address in this lab, and my injector blocks
everything EXCEPT the LAN ("LAN allowed, everything else dropped"). So:
* the box never lost contact with the hub in round 7 OR round 8;
* "internet gone" as injected means only "the PUBLIC path is gone";
* the dropped-event question (events pushed while the hub is unreachable are retried 3x then
dropped permanently, no queue) is STILL unmeasured after two rounds that looked like they
measured it.
My own memory file carries this exact warning -- "hub.felhom.eu resolves to the LAN here; an
internet-cut drill must block it too" -- and I did not apply it. Recorded as an instrument fault,
not a product defect. Nothing the box did was wrong.
### 3. The action and the accident barely overlapped, and would not have interacted anyway.
The backup finished at 22:37:48Z (db_dump count=5, 1m55s, success=true). The cut began 22:37:54Z,
six seconds LATER. So "backup-app during an internet cut" was not really exercised. It would not
have mattered: backup-app is the LOCAL app-data tier (DB dump + volumes) and needs no internet.
The off-site tier is a different action (offsite-run), drawn in rounds 1 and 4.
### 4. What the box actually did, which is all true and all good.
containers 26 -> 26, never dropped, at no point during the ten minutes
household loop: 12 lines this round, 0 failures
front doors on the LAN: served throughout
public path: lost during the cut, restored by itself in <=43 s, unaided
alarms: NONE fired, and none should have (node_stale threshold is 30 min; and the box was
never actually stale, because it was reporting to the hub the whole time)
@@ -83,7 +83,14 @@ for i in $(seq 1 40); do
[ "$BEFORE" -gt 0 ] && [ "$N2" -ge "$BEFORE" ] && { say " STEADY after $(( $(date +%s) - T0 ))s"; break; }
sleep 15
done
say " front doors: $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)"
# A door reading taken seconds after a network accident ends measures the accident, not the
# recovery. Round 8 read 530 on every door 3 s after the unblock, and 200 on every door 43 s
# later. The reading now STATES ITS OWN PRECONDITION, and a second one is always taken.
say " front doors, FIRST reading at $(date -u +%FT%TZ) - TOO EARLY to trust if the accident just ended:"
say " $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)"
sleep 60
say " front doors, SECOND reading at $(date -u +%FT%TZ), 60 s later - THIS is the one to trust:"
say " $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)"
HL1=$(BOX 'wc -l < /root/household.log' | tr -d ' \r')
say " household lines this round: $(( ${HL1:-0} - ${HL0:-0} )) failures: $(BOX "tail -n +$(( ${HL0:-0} + 1 )) /root/household.log | grep -cE 'FAILED|UNREACHABLE'" | tr -d ' \r')"
say "--- alarms ---"; bash $E/events.sh 8 | tee -a $OUT