diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 048ef0d0..8715dd96 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -337,6 +337,48 @@ host-wide sysctl on demo-hp and inserts two `physdev` rules. Afterwards: `-P FOR physdev rules, sysctl **0** — identical to the pre-round reading. demo-hp also carries guests 9201 and 9202, so an abandoned rule would have been a fence breach, not an untidy drill. +### Round 8 — `backup-app` nextcloud / accident: **internet cut for ten minutes** + +**22:35:48Z–22:47:59Z.** The app-data backup ran first and finished in 1 m 55 s +(`db_dump` count 5, `success:true`, 22:37:48Z). The cut began six seconds later, so the two barely +overlapped — and would not have interacted in any case: `backup-app` is the **local** app-data tier +and needs no internet. The off-site tier is a different action. + +| the five things | | +|---|---| +| what the customer saw | **Nothing at home.** Every front door kept serving on the LAN for the whole ten minutes. From outside the house the sites were unreachable — the public path was down. 26 apps up before, 26 after, never fewer. | +| what the box did by itself | Kept every container running, kept backing up, kept reporting to the hub, and rebuilt the public path unaided when the link returned. No restart, no intervention, noaction from me. | +| time to steady | **≤43 s** after the link returned (public doors 200 again at 22:48:38Z). Containers never left steady at all. | +| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold and this was ten. Nothing false was raised. | +| should have fired, did not | **none** | + +**And the finding of the round is against my own instrument, not the box.** + +The hub report due at **22:38:43Z fell inside the cut** — and it **succeeded**: +„Hub report pushed successfully (15526 bytes)". It succeeded because `hub.felhom.eu` resolves to +**192.168.0.192**, a LAN address (measured from both the guest and the host), and my injector blocks +everything **except** the LAN. So the accident named „internet gone" only ever removed the **public** +path. The box never lost the hub, in round 7 or in round 8. + +Two consequences, both stated plainly: + +1. **The dropped-event question is still unmeasured** after two rounds that appeared to measure it. + Events pushed while the hub is unreachable are retried three times and then dropped permanently, + with no queue — that behaviour has still never been seen live. +2. **My own memory file carries this exact warning** („hub.felhom.eu resolves to the LAN here; an + internet-cut drill must block it too") and I did not apply it. A warning that is written down and + not read is worth nothing, which is the same class of failure as an unread alarm. + +The injector is corrected for round 9's drawn internet cut so that the hub address is blocked too — +**from the VM's side, at the host's tap rule.** The hub itself is never touched; the fence is kept. +This is a repair to a broken instrument, not a re-draw: the drawn action, app and accident for every +remaining round are unchanged. + +**A sixth mistimed reading, and the fix is in the runner now.** The round's own AFTER step read the +public doors **three seconds** after the unblock and reported 530 on all four. That reading could +never have been fair. The re-measurement 43 s later read 200 on all four. The runner measured the +recovery before the recovery could begin — so the runner is the thing that gets fixed, not the note. + ### Rounds 8-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh b/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh index 5f29dfe4..29f99327 100755 --- a/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh +++ b/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh @@ -45,23 +45,28 @@ case "$A" in internet-gone-10min) TAP=$(H "ls /sys/class/net | grep -E \"^tap${VM}i0$\"") [ -n "$TAP" ] || { say "NO TAP FOUND for VM $VM — accident NOT injected, and that is recorded as such"; exit 1; } - say "blocking the box's traffic off-LAN at the HOST, on $TAP; the LAN stays up" + HUBIP=$(H "getent hosts hub.felhom.eu | awk '{print \$1}' | head -1" | tr -d ' \r') + [ -n "$HUBIP" ] || { say "HUB ADDRESS NOT RESOLVABLE - accident NOT injected. An instrument that cannot find its target must not pretend it blocked it."; exit 1; } + say "blocking the box traffic off-LAN at the HOST, on $TAP; the LAN stays up EXCEPT the hub ($HUBIP)" H "sysctl -w net.bridge.bridge-nf-call-iptables=1 >/dev/null iptables -I FORWARD 1 -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT - iptables -I FORWARD 2 -m physdev --physdev-in $TAP -j DROP" + iptables -I FORWARD 2 -m physdev --physdev-in $TAP -j DROP + iptables -I FORWARD 1 -m physdev --physdev-in $TAP -d $HUBIP -j DROP" # UNCONDITIONAL cleanup: if this script is killed during the sleep, or the SSH drops, the host # must NOT be left with the rules in place and the sysctl flipped. demo-hp is a Tier-0 host that # 9201 and 9202 also live on, and an abandoned FORWARD rule is a fence breach, not a measurement. netcleanup(){ - H "iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP 2>/dev/null + H "iptables -D FORWARD -m physdev --physdev-in $TAP -d $HUBIP -j DROP 2>/dev/null + iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP 2>/dev/null iptables -D FORWARD -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT 2>/dev/null sysctl -w net.bridge.bridge-nf-call-iptables=0 >/dev/null 2>&1" >/dev/null 2>&1 say "net cleanup ran (rules removed, sysctl restored)" } trap 'netcleanup' EXIT INT TERM - say "blocked (LAN allowed, everything else dropped) — 10 minutes" + say "blocked (LAN allowed EXCEPT the hub at $HUBIP, everything else dropped) - 10 minutes" sleep 600 - H "iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP + H "iptables -D FORWARD -m physdev --physdev-in $TAP -d $HUBIP -j DROP + iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP iptables -D FORWARD -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT sysctl -w net.bridge.bridge-nf-call-iptables=0 >/dev/null" trap - EXIT INT TERM diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-8.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-8.txt new file mode 100644 index 00000000..5d52b0d4 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-8.txt @@ -0,0 +1,83 @@ + -rwxrwxr-x 1 kisfenyo kisfenyo 5650 Sep 16 23:42 /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh + -rwxrwxr-x 1 kisfenyo kisfenyo 5712 Sep 16 23:52 /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh + waiter armed; will launch round 8 at 2026-09-16T22:35:48Z +=== round 8 launched 2026-09-16T22:35:48Z (due 22:35:48Z) === +2026-09-16T22:35:48Z ================ ROUND 8 : backup-app nextcloud, while: internet-gone-10min ================ +2026-09-16T22:35:50Z --- BEFORE --- containers=26 cloud=200 status=200 paste=200 +2026-09-16T22:35:50Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) +2026-09-16T22:35:50Z --- ACTION: backup-app on nextcloud --- + POST /api/backup/run -> 200 +{"ok":true,"message":"Mentés elindítva"} + + {"ok":true,"data":{"enabled":true,"running":true}} + {"ok":true,"data":{"enabled":true,"running":true}} + {"ok":true,"data":{"enabled":true,"running":true}} + {"ok":true,"data":{"enabled":true,"running":true}} + {"ok":true,"data":{"enabled":true,"running":true}} + {"ok":true,"data":{"db_dump":{"count":5,"duration":"1m55.578787777s","last_run":"2026-09-16T22:37:48.900952183Z","success":true},"enabled":true,"running":false}} +2026-09-16T22:37:54Z --- ACCIDENT: internet-gone-10min (injected after the action started) --- + 2026-09-16T22:37:54Z ACCIDENT=internet-gone-10min round=8 + 2026-09-16T22:37:55Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up + 2026-09-16T22:37:55Z blocked (LAN allowed, everything else dropped) — 10 minutes + 2026-09-16T22:47:55Z unblocked; host sysctl restored to 0 and both rules removed + -P FORWARD ACCEPT + 2026-09-16T22:47:55Z accident internet-gone-10min complete +2026-09-16T22:47:55Z --- AFTER: what the box did BY ITSELF --- +2026-09-16T22:47:57Z t+603s containers=26 (before 26) +2026-09-16T22:47:57Z STEADY after 603s +2026-09-16T22:47:58Z front doors: cloud=530 status=530 paste=530 wiki=530 +2026-09-16T22:47:58Z household lines this round: 12 failures: 0 +2026-09-16T22:47:58Z --- alarms --- + | Time | Severity | Type | Message | Source + | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller + | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller + | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller + | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller + | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller + | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller +2026-09-16T22:47:59Z ================ END ROUND 8 ================ + +[exited with code 0] + +## POST-ROUND CORRECTIONS, written immediately (2026-09-16T22:48Z) + +### 1. The "530 on every door" line above is a MISTIMED READING, not a fault. +It was taken 3 seconds after the unblock (22:47:58Z; unblock 22:47:55Z). Re-measured at +22:48:38Z, 43 s after the unblock: + cloud LAN=301 public=200 + status LAN=301 public=200 + paste LAN=301 public=200 + wiki LAN=301 public=200 +So the public path came back by itself in <=43 s. This is the SIXTH mistimed reading tonight and +the runner's own AFTER step is the one still doing it -- it measures the public door 3 s after the +network returns, which can never be a fair reading. Fix applied below. + +### 2. THE ACCIDENT DID NOT DO WHAT ITS NAME SAYS. This is my fault, and it is the bigger finding. +The hub report due at 22:38:43Z fell INSIDE the ten-minute cut (22:37:55 -> 22:47:55). +It SUCCEEDED: "[report] Hub report pushed successfully (15526 bytes)". +It succeeded because hub.felhom.eu resolves to a LAN address in this lab, and my injector blocks +everything EXCEPT the LAN ("LAN allowed, everything else dropped"). So: + * the box never lost contact with the hub in round 7 OR round 8; + * "internet gone" as injected means only "the PUBLIC path is gone"; + * the dropped-event question (events pushed while the hub is unreachable are retried 3x then + dropped permanently, no queue) is STILL unmeasured after two rounds that looked like they + measured it. +My own memory file carries this exact warning -- "hub.felhom.eu resolves to the LAN here; an +internet-cut drill must block it too" -- and I did not apply it. Recorded as an instrument fault, +not a product defect. Nothing the box did was wrong. + +### 3. The action and the accident barely overlapped, and would not have interacted anyway. +The backup finished at 22:37:48Z (db_dump count=5, 1m55s, success=true). The cut began 22:37:54Z, +six seconds LATER. So "backup-app during an internet cut" was not really exercised. It would not +have mattered: backup-app is the LOCAL app-data tier (DB dump + volumes) and needs no internet. +The off-site tier is a different action (offsite-run), drawn in rounds 1 and 4. + +### 4. What the box actually did, which is all true and all good. + containers 26 -> 26, never dropped, at no point during the ten minutes + household loop: 12 lines this round, 0 failures + front doors on the LAN: served throughout + public path: lost during the cut, restored by itself in <=43 s, unaided + alarms: NONE fired, and none should have (node_stale threshold is 30 min; and the box was + never actually stale, because it was reporting to the hub the whole time) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh b/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh index bf4620bb..7aa2df86 100755 --- a/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh +++ b/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh @@ -83,7 +83,14 @@ for i in $(seq 1 40); do [ "$BEFORE" -gt 0 ] && [ "$N2" -ge "$BEFORE" ] && { say " STEADY after $(( $(date +%s) - T0 ))s"; break; } sleep 15 done -say " front doors: $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)" +# A door reading taken seconds after a network accident ends measures the accident, not the +# recovery. Round 8 read 530 on every door 3 s after the unblock, and 200 on every door 43 s +# later. The reading now STATES ITS OWN PRECONDITION, and a second one is always taken. +say " front doors, FIRST reading at $(date -u +%FT%TZ) - TOO EARLY to trust if the accident just ended:" +say " $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)" +sleep 60 +say " front doors, SECOND reading at $(date -u +%FT%TZ), 60 s later - THIS is the one to trust:" +say " $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)" HL1=$(BOX 'wc -l < /root/household.log' | tr -d ' \r') say " household lines this round: $(( ${HL1:-0} - ${HL0:-0} )) failures: $(BOX "tail -n +$(( ${HL0:-0} + 1 )) /root/household.log | grep -cE 'FAILED|UNREACHABLE'" | tr -d ' \r')" say "--- alarms ---"; bash $E/events.sh 8 | tee -a $OUT