From 7221367f389f30af40314f137fa1f7390f0ebe0c Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 23:33:08 +0200 Subject: [PATCH] CHAOS NIGHT round 3: the disk fills, and nothing is told about it Measured while the guest's root filesystem was held at 96%: - the app's view is a genuinely full disk (29G used, 1.5G free, 28G fill file) - the shared thin pool stayed at 39.69% - fallocate reserves blocks without writing them, so this round is NOT a repeat of the pool exhaustion that wedged the box in Phase 0. Recorded explicitly, because "disk 95% full" invites exactly that wrong reading. - the filesystem stayed writable (a real touch, not the mount flags) No disk alarm fired, and that was PREDICTED from the ladder before the round: the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller start. The timing is sharper still - the controller restarted at 21:28 after round 2's power cut, so its single opportunistic check ran about twenty seconds BEFORE the disk filled. A disk that fills and empties between checks is invisible; that is by design, but it is the honest answer to "would the household be told?" - no. Household lines are attributed to the right round: the two UNREACHABLE entries at 21:27:57Z are round 2's recovery tail, not round 3's accident. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../round-3.txt | 48 +++++++++++++++++++ 1 file changed, 48 insertions(+) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt index 02863ab6..12db4dcf 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt @@ -6,3 +6,51 @@ 2026-09-16T21:29:49Z wiki read 2 -> 200 2026-09-16T21:29:49Z wiki read 3 -> 200 2026-09-16T21:29:49Z --- ACCIDENT: disk-95-full (injected after the action started) --- + +## MID-ACCIDENT measurements, taken while the disk was held full (21:31:39Z) + guest / 32G, **29G used, 1.5G free, 96%** <- the app's view: full + the fill file /var/tmp/.chaosfill, **28G** + the shared LVM thin pool <75.81g, **39.69%** <- essentially unchanged + writable? `touch` succeeded: **WRITE OK** <- no read-only wedge + +**The important detail, so this round is not over-read later: the POOL was never stressed.** +`fallocate` reserves filesystem blocks without writing them, so the guest's ext4 believes it is 96 % +full while the thin pool underneath never allocated the extents. The accident is therefore faithful +at the level the ACCIDENT IS ABOUT — an application meeting a full disk — and is **not** a repeat of +the pool exhaustion that wedged the box in Phase 0. My pool guard computed a cap and, as it turns +out, the cap was never the binding constraint. + +Recorded because the opposite conclusion is easy to reach from the words „disk 95 % full" alone, and +this project has a standing lesson about measurements that look like something they are not. + +## Household loop attribution for this round +The loop shows 2 `UNREACHABLE` lines at **21:27:57Z**, which is BEFORE this round began (21:29:45Z): +they are the tail of round 2's power-cut recovery, caught at the moment the loop was reinstalled as +a persistent unit while the box was still coming back. **They belong to round 2, not round 3**, and +round 2's household measure remains „not collected" because the loop was dead for the cut itself. +From 21:29:57Z the loop reads normally again (`paste read ok http=301`). + +## The alarm the ladder PREDICTED would not fire, and did not (checked at 21:32:17Z) +With the guest's root filesystem held at **96 %**, the newest event in the hub is still +`controller_started` from 21:28. **No `disk_critical`, no `disk_warning`, no `health_degraded`.** + +This was predicted before the round from `08-alarm-ladder.md` and the controller's own fill-watch: +the fill check runs **daily at 03:30**, plus once about **90 seconds after a controller start**. A +ten-minute window therefore contains no check at all unless a controller restart happens to land +inside it. + +The timing here is sharper than the general rule, and it is worth the detail: + 21:26:01Z power restored (round 2's accident) + 21:28 `controller_started` — so its opportunistic fill check ran at about **21:29:30** + 21:29:49Z my fill began + 21:31:39Z disk measured at 96 %, pool 39.69 %, still writable +So the single check this box would have made in the whole window ran roughly **twenty seconds before +the disk filled**, and the next one is not due until 03:30. + +**What this means, stated carefully:** a disk that fills and empties between checks is invisible to +the alarm ladder. That is by design rather than a defect — the fill-watch is a daily sweep, not a +monitor — but it is the honest answer to „would the household be told?" for a transient full disk: +**no, unless the controller happens to restart while it is full.** + +Re-checked at the END of the ten-minute window before this is called final (below), because an alarm +arriving late is a different finding from an alarm never arriving.