CHAOS NIGHT round 3: the disk fills for ten minutes and nobody is told
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
Round 3 (use bookstack, system disk held at 96% for ten minutes): - the household saw nothing wrong: wiki, status and paste all answered 200 before, during and after, and the background loop logged 12 operations with ZERO failures on a 96%-full disk - the box kept all 26 containers running and released the space cleanly (29G used -> 944M used) with the thin pool untouched at 39.69% throughout - NO alarm fired at any point, checked twice independently after the fill was released That silence is the finding, and it was predicted from the ladder before the round rather than discovered after: disk_critical is defined at >=95% used, but the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller start. The controller happened to restart at 21:28, so its single opportunistic check ran about twenty seconds BEFORE the disk filled. A disk that fills and empties between sweeps is invisible - by design, but the honest answer to "would the household be told?" is no. Also fixed and explained: my injector printed "unexpected EOF" while the accident itself completed. bash -n passes, so it was not local syntax - G() flattens its argument through `pct exec`, so a nested bash -c '...' has its quoting re-parsed remotely. Both instances were in the disk branch only; the accidents still to come use plain commands. The experiment was verified on the box, not from the script's own account. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -92,13 +92,16 @@ case "$A" in
|
||||
say " read-only; writing the full 95 % would re-wedge the box rather than measure it."
|
||||
WANT=$CAP
|
||||
fi
|
||||
G "bash -c 'df -h / | tail -1'" | tee -a "$LOG"
|
||||
G "df -h /" | tail -1 | tee -a "$LOG"
|
||||
G "fallocate -l ${WANT}M /var/tmp/.chaosfill" >/dev/null 2>&1
|
||||
G "df -h / | tail -1" | tee -a "$LOG"
|
||||
say " pool after the fill: $(BOX "lvs --noheadings -o data_percent pve/data" | tr -d ' \r')"
|
||||
say "full — holding 10 minutes"
|
||||
sleep 600
|
||||
G "bash -c 'rm -f /var/tmp/.chaosfill; df -h / | tail -1'" | tee -a "$LOG"
|
||||
# two plain calls: G() flattens its argument through `pct exec`, so a nested bash -c '...' has
|
||||
# its quoting re-parsed by the remote shell and dies with "unexpected EOF". Measured in round 3.
|
||||
G "rm -f /var/tmp/.chaosfill"
|
||||
G "df -h /" | tail -1 | tee -a "$LOG"
|
||||
say "freed"
|
||||
;;
|
||||
*) say "UNKNOWN ACCIDENT $A"; exit 2 ;;
|
||||
|
||||
@@ -5,3 +5,5 @@
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 29G 1.5G 96% /
|
||||
2026-09-16T21:29:55Z pool after the fill: 39.69
|
||||
2026-09-16T21:29:55Z full — holding 10 minutes
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 944M 29G 4% /
|
||||
2026-09-16T21:39:56Z freed
|
||||
|
||||
@@ -87,3 +87,74 @@ it released the fill; I was guessing.
|
||||
|
||||
This is the seventh item on tonight's list of my own instrument errors, and it has the same shape as
|
||||
the rest — **a confident label on a measurement whose precondition was not met.**
|
||||
2026-09-16T21:29:49Z ACCIDENT=disk-95-full round=3
|
||||
2026-09-16T21:29:49Z filling the customer guest's SYSTEM disk to ~95 % — with a POOL GUARD
|
||||
2026-09-16T21:29:51Z pool now: 39% used; guest / has 29352 MB free
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 944M 29G 4% /
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 29G 1.5G 96% /
|
||||
2026-09-16T21:29:55Z pool after the fill: 39.69
|
||||
2026-09-16T21:29:55Z full — holding 10 minutes
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 944M 29G 4% /
|
||||
2026-09-16T21:39:56Z freed
|
||||
/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh: line 106: unexpected EOF while looking for matching `''
|
||||
2026-09-16T21:39:56Z --- AFTER: what the box did BY ITSELF ---
|
||||
2026-09-16T21:39:58Z t+609s containers=26 (before 26)
|
||||
2026-09-16T21:39:58Z STEADY after 609s
|
||||
2026-09-16T21:39:59Z front doors: wiki=200 status=200 paste=200 wiki=200
|
||||
2026-09-16T21:39:59Z household lines this round: 12 failures: 0
|
||||
2026-09-16T21:39:59Z --- alarms ---
|
||||
| Time | Severity | Type | Message | Source
|
||||
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
|
||||
| Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
|
||||
| Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
|
||||
| Sep 16 21:19 | info | app_deploy_started | Alkalmazás telepítése elindult: Immich | controller
|
||||
| Sep 16 21:19 | info | app_removed | Alkalmazás eltávolítva: immich | controller
|
||||
2026-09-16T21:40:00Z ================ END ROUND 3 ================
|
||||
|
||||
## ROUND 3 — the five things
|
||||
**action:** `use` bookstack (three reads through its own front door)
|
||||
**accident:** the guest's system disk filled to ~96 % and held for ten minutes
|
||||
|
||||
1. **What the customer saw.** Nothing. `wiki` answered 200 before, during and after; `status` and
|
||||
`paste` likewise. No banner, no warning, no mail — the household was never told the disk was full.
|
||||
2. **What the box did by itself.** Kept all twelve apps running on a 96 %-full root filesystem and
|
||||
released the space cleanly when the fill was removed:
|
||||
21:29:55Z 32G, **29G used, 1.5G free, 96%** — held
|
||||
21:39:56Z 32G, **944M used, 29G free, 4%** — freed
|
||||
The shared thin pool never moved (**39.69 %** throughout), and the filesystem stayed writable.
|
||||
3. **Time to steady.** The box never left steady — 26 containers before, **26 after**, none restarted.
|
||||
(The runner's „STEADY after 609s" is simply the elapsed window, not a recovery time.)
|
||||
4. **Alarm fired / true?** **None fired.** Checked twice independently — the runner's own dump at
|
||||
21:39:59Z and a separate query at 21:40:13Z, after the fill was released. Newest event in both is
|
||||
still `controller_started` from 21:28.
|
||||
5. **Should have fired and did not.** `disk_critical` is defined for ≥95 % used, and the disk sat at
|
||||
**96 % for ten minutes**. It did not fire — **but that is the ladder working as designed, not a
|
||||
miss**: the fill-watch is a daily sweep (03:30) plus one check ~90 s after a controller start, so a
|
||||
transient full disk between sweeps is invisible. Predicted from the ladder before the round, and
|
||||
confirmed. The honest answer to „would the household be told?" is **no**, unless the controller
|
||||
happens to restart while the disk is full.
|
||||
|
||||
**Household loop:** **12 operations in the window, 0 failures.** The household kept reading its apps
|
||||
normally throughout — which is the other half of the finding: nothing broke, and nothing was said.
|
||||
|
||||
**Known pre-existing condition, unchanged:** immich still down (`app_oom`, its Postgres killed by the
|
||||
memory limit). Not caused by this round.
|
||||
|
||||
## An error in MY injector, diagnosed rather than waved away
|
||||
The run printed:
|
||||
inject.sh: line 106: unexpected EOF while looking for matching `''
|
||||
yet `bash -n inject.sh` passes cleanly. The fault is not local syntax: my `G()` helper runs
|
||||
`pct exec 9201 -- $*`, which FLATTENS its argument, so the nested `bash -c 'rm -f …; df -h …'` had its
|
||||
quoting re-parsed by the remote shell and fell apart there.
|
||||
|
||||
**The accident itself completed regardless**, and that was verified on the box rather than taken from
|
||||
the script's own account: the fill file was gone, the filesystem back to 4 %, the pool unmoved at
|
||||
39.69 %, 26 containers running and a real write succeeding. So the error was in the reporting layer,
|
||||
not in the experiment — the same distinction I have had to draw repeatedly tonight.
|
||||
|
||||
Fixed by splitting it into two plain `G` calls with no nested quotes. Round 4's accident
|
||||
(`docker kill cloudflared`) and round 5's (`systemctl restart docker`) contain no nested quoting and
|
||||
were unaffected, which is why round 4 was allowed to start before this was tidied.
|
||||
|
||||
@@ -0,0 +1,3 @@
|
||||
2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4
|
||||
2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest
|
||||
2026-09-16T21:41:33Z tunnel killed — 10 minutes
|
||||
@@ -0,0 +1,6 @@
|
||||
2026-09-16T21:41:25Z ================ ROUND 4 : offsite-run mealie, while: tunnel-down-10min ================
|
||||
2026-09-16T21:41:28Z --- BEFORE --- containers=26 recipes=200 status=200 paste=200
|
||||
2026-09-16T21:41:28Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T21:41:29Z --- ACTION: offsite-run on mealie ---
|
||||
ABORT: no session — mine, not the product's
|
||||
2026-09-16T21:41:32Z --- ACCIDENT: tunnel-down-10min (injected after the action started) ---
|
||||
Reference in New Issue
Block a user