CHAOS NIGHT: round 3 written up, and the household loop's blind spot stated
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
Round 3 in the findings document: the disk sat at 96% for ten minutes, twelve apps kept serving, the household loop logged 12 operations with zero failures, and NOTHING was ever raised. The silence is the finding, and it was predicted from the ladder before the round: the fill-watch is a daily sweep plus one check ~90s after a controller start, and that single check ran about twenty seconds before the disk filled. Recorded with it: I twice labelled a mid-window reading "end of window", estimating the clock instead of reading it. The readings were unchanged but the label was wrong, and "nothing yet" is not "nothing ever". And a limit of my own instrument, stated before its numbers get quoted: the household loop does not follow redirects, so it measures "is the app serving on the box" and never "can the household reach it from outside". It logged zero failures straight through round 4's tunnel outage while the public route was returning 530. So "0 household failures in round 4" must not be read as "the household was unaffected" - someone away from home would have met 530 for about ninety seconds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -229,7 +229,29 @@ and named the exact container** — worth recording against this project's stand
|
||||
signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in
|
||||
the alarm feed, correctly labelled, the whole time.
|
||||
|
||||
### Rounds 3-12
|
||||
### Round 3 — `use` bookstack / accident: **system disk filled to 96 % for ten minutes**
|
||||
|
||||
| the five things | |
|
||||
|---|---|
|
||||
| what the customer saw | **nothing.** wiki, status and paste all answered before, during and after. No banner, no warning, no mail — the household was never told the disk was full |
|
||||
| what the box did by itself | kept all twelve apps running on a 96 %-full root filesystem and released the space cleanly when the fill was removed (29 G used → 944 M used). The shared thin pool never moved (**39.69 %**) and the filesystem stayed writable |
|
||||
| time to steady | the box never left steady — **26 containers before, 26 after**, none restarted |
|
||||
| alarm fired / true? | **none fired**, checked twice independently after the fill was released |
|
||||
| should have fired, did not | `disk_critical` is defined at ≥95 % used and the disk sat at **96 % for ten minutes**. **But this is the ladder working as designed, not a miss:** the fill-watch is a daily sweep (03:30) plus one check ~90 s after a controller start. Predicted before the round, confirmed after |
|
||||
|
||||
**Household loop: 12 operations, 0 failures.** The household kept using its apps normally throughout.
|
||||
|
||||
**The finding is the silence.** The honest answer to „would the household be told their disk is
|
||||
full?" is **no** — unless the controller happens to restart while it is full. Here the timing was
|
||||
almost comic: the controller restarted at 21:28 after round 2's power cut, so its one opportunistic
|
||||
check ran about twenty seconds *before* the disk filled, and the next is not due until 03:30.
|
||||
|
||||
**A correction, recorded where it happened:** I twice labelled a mid-window reading „end of window",
|
||||
estimating the clock instead of reading it. The readings were unchanged, but „nothing yet, five
|
||||
minutes in" and „nothing in the whole window" are different findings. From here the end-of-window
|
||||
check is taken when the round's own runner reports completion.
|
||||
|
||||
### Rounds 4-12
|
||||
|
||||
PENDING
|
||||
|
||||
|
||||
@@ -54,3 +54,21 @@ down and the box really was not fully healthy.
|
||||
LAN route straight to the box traefik answered **301** throughout
|
||||
So the household's apps never stopped serving locally; what failed, and recovered, was the way in
|
||||
from outside.
|
||||
|
||||
## A LIMIT of the household loop, stated before its numbers are read
|
||||
Through the tunnel outage the loop logged:
|
||||
21:41:57Z paperless read ok http=301 21:41:57Z paperless dash ok http=301
|
||||
21:43:57Z paste read ok http=301 21:43:57Z paste dash ok http=301
|
||||
i.e. **no failures at all**, while the public way in was returning 530.
|
||||
|
||||
That is correct behaviour for the loop and a real blind spot at the same time. The loop does **not**
|
||||
follow redirects: it takes traefik's 301 on the box as a success, so it measures *„is the app serving
|
||||
on the box?"* and never *„can the household reach it from outside?"*. During round 4 the first answer
|
||||
stayed yes throughout, which is why it saw nothing.
|
||||
|
||||
**Therefore: „0 household failures in round 4" must NOT be read as „the household was unaffected".**
|
||||
A household member away from home would have met 530 for about ninety seconds. The loop's number is
|
||||
true and answers a narrower question than it appears to.
|
||||
|
||||
This is the same distinction as the front-door method correction above, and it is recorded in both
|
||||
places because the number and the label live in different files.
|
||||
|
||||
Reference in New Issue
Block a user