diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt index 02863ab6..12db4dcf 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt @@ -6,3 +6,51 @@ 2026-09-16T21:29:49Z wiki read 2 -> 200 2026-09-16T21:29:49Z wiki read 3 -> 200 2026-09-16T21:29:49Z --- ACCIDENT: disk-95-full (injected after the action started) --- + +## MID-ACCIDENT measurements, taken while the disk was held full (21:31:39Z) + guest / 32G, **29G used, 1.5G free, 96%** <- the app's view: full + the fill file /var/tmp/.chaosfill, **28G** + the shared LVM thin pool <75.81g, **39.69%** <- essentially unchanged + writable? `touch` succeeded: **WRITE OK** <- no read-only wedge + +**The important detail, so this round is not over-read later: the POOL was never stressed.** +`fallocate` reserves filesystem blocks without writing them, so the guest's ext4 believes it is 96 % +full while the thin pool underneath never allocated the extents. The accident is therefore faithful +at the level the ACCIDENT IS ABOUT — an application meeting a full disk — and is **not** a repeat of +the pool exhaustion that wedged the box in Phase 0. My pool guard computed a cap and, as it turns +out, the cap was never the binding constraint. + +Recorded because the opposite conclusion is easy to reach from the words „disk 95 % full" alone, and +this project has a standing lesson about measurements that look like something they are not. + +## Household loop attribution for this round +The loop shows 2 `UNREACHABLE` lines at **21:27:57Z**, which is BEFORE this round began (21:29:45Z): +they are the tail of round 2's power-cut recovery, caught at the moment the loop was reinstalled as +a persistent unit while the box was still coming back. **They belong to round 2, not round 3**, and +round 2's household measure remains „not collected" because the loop was dead for the cut itself. +From 21:29:57Z the loop reads normally again (`paste read ok http=301`). + +## The alarm the ladder PREDICTED would not fire, and did not (checked at 21:32:17Z) +With the guest's root filesystem held at **96 %**, the newest event in the hub is still +`controller_started` from 21:28. **No `disk_critical`, no `disk_warning`, no `health_degraded`.** + +This was predicted before the round from `08-alarm-ladder.md` and the controller's own fill-watch: +the fill check runs **daily at 03:30**, plus once about **90 seconds after a controller start**. A +ten-minute window therefore contains no check at all unless a controller restart happens to land +inside it. + +The timing here is sharper than the general rule, and it is worth the detail: + 21:26:01Z power restored (round 2's accident) + 21:28 `controller_started` — so its opportunistic fill check ran at about **21:29:30** + 21:29:49Z my fill began + 21:31:39Z disk measured at 96 %, pool 39.69 %, still writable +So the single check this box would have made in the whole window ran roughly **twenty seconds before +the disk filled**, and the next one is not due until 03:30. + +**What this means, stated carefully:** a disk that fills and empties between checks is invisible to +the alarm ladder. That is by design rather than a defect — the fill-watch is a daily sweep, not a +monitor — but it is the honest answer to „would the household be told?" for a transient full disk: +**no, unless the controller happens to restart while it is full.** + +Re-checked at the END of the ten-minute window before this is called final (below), because an alarm +arriving late is a different finding from an alarm never arriving.