chaos night round 11: the drive pulled for 20 minutes, and eight true alarms
gates / gates (push) Successful in 25s

The richest alarm round of the night, and every alarm was true and correctly
paired: storage_disconnected naming the drive by the household's own label,
four app_start_failed naming exactly the four apps whose data lives on that
drive, health_degraded, then storage_reconnected and health_recovered.

The other eleven apps kept serving throughout. The box recovered unaided in
67 s after the drive was plugged back in, with the front door following at
128 s. The drive came back clean: 98 G, 2% used, mountpoint config unchanged.

The round's own snapshot showed health_degraded with no recovery, which would
have been the first missing alarm of the night. The recovery had fired seconds
after the snapshot. A re-read taken after the precondition found it. Not a
missing alarm - a premature reading, caught by the discipline the earlier
mistimed readings forced.

Truth table now: 17 alarms fired, 17 true, 0 missing, across eleven rounds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 02:15:26 +02:00
parent 73ac9d7871
commit 70bffb1676
3 changed files with 112 additions and 3 deletions
@@ -460,6 +460,29 @@ a household count that reported 0 lines and 0 failures when the truth was one li
failure; and a disk guard that reported „active" all night while being a **transient** unit that
vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box.
### Round 11 — `use` paperless-ngx / accident: **the data drive pulled out for twenty minutes**
**23:50:49Z–00:13:09Z.** The drive was detached from the **running** box at 23:50:54Z and put back at
00:10:58Z. This is the round that produced the most alarms of the night, and every one of them was
true.
| the five things | |
|---|---|
| what the customer saw | **The four apps whose files live on that drive stopped** — Paperless, Jellyfin, Nextcloud, Immich — and Paperless's front door went 404. The other eleven apps kept serving normally throughout. About twenty minutes later everything was back, roughly a minute after the drive was plugged in again. |
| what the box did by itself | Noticed the drive had gone and **named it by the label the household sees** („Adatlemez"), named **each** broken app individually, degraded its own health, waited, noticed the drive return, restarted the apps and recovered its health. No restart, no repair, nothing from me. |
| time to steady | **67 s after the drive returned** (26 containers again at 00:12:05Z). The door followed at 00:13:06Z. |
| alarm fired / true? | **eight, all true, correctly paired at both ends** — `storage_disconnected` (error) → four `app_start_failed` (warning) → `health_degraded` (warning) → `storage_reconnected` (info) → `health_recovered` (info). The four apps named are **exactly** the four with data on the pulled drive. Nothing false was raised. |
| should have fired, did not | **none** |
**The alarm that looked missing, and was not.** The round's own snapshot at 00:13:08Z showed
`health_degraded` with no recovery — which would have been the first missing alarm of the night. The
recovery fired at **00:13**, and the snapshot missed it by **seconds**. A re-read at 00:14:13Z, taken
after the apps were back, found it. This is the discipline from the earlier mistimed readings earning
its keep: when a measurement could have been early, it is re-taken rather than turned into a verdict.
**The drive came back clean** — `/dev/sdd`, 98 G, 2 % used, mounted at `/mnt/felhom-drives/hdd_1`,
with the guest's mountpoint config unchanged, 26 containers up and every real front door serving.
### Rounds 8-12
PENDING
@@ -21,10 +21,15 @@ R10| hard reset, 4 s into | controller_started (info)
| an app restore | | | no restore field in the status JSON, only a button label on the pages, and no file at
| | | | all modified in the reset window. Interrupted and never-happened look identical to the
| | | | customer. Filed as R-550 (P2). The box itself recovered 26/26 containers in 150 s.
R11| data drive pulled | storage_disconnected (error) -> app_start_failed x4 (warning)| YES (8/8) | none. The four apps named are EXACTLY the four whose data lives on the pulled drive;
| out for 20 minutes | -> health_degraded (warning) -> storage_reconnected (info) | | the other eleven kept serving. The drive is named by the household's own label
| | -> health_recovered (info) | | ("Adatlemez"). Recovery unaided in 67 s after the drive returned.
| | | | NOTE: the round's own snapshot missed health_recovered by SECONDS. Re-read after
| | | | the precondition found it. Not a missing alarm - a premature reading, caught.
## Running totals, rounds 1-10
alarms fired: 9
alarms TRUE: 9 (9/9 - no false alarm in ten rounds)
## Running totals, rounds 1-11
alarms fired: 17
alarms TRUE: 17 (17/17 - no false alarm in eleven rounds)
alarms MISSING: 0
design gaps found: 3 (R-547: a transient full disk is never mentioned to anyone)
(R-549: one failed report push spends the whole staleness budget)
@@ -0,0 +1,81 @@
round 11 armed for 2026-09-16T23:50:48Z (use paperless-ngx + drive-pulled-20min)
=== round 11 launched 2026-09-16T23:50:49Z (due 23:50:48Z) ===
2026-09-16T23:50:49Z ================ ROUND 11 : use paperless-ngx, while: drive-pulled-20min ================
2026-09-16T23:50:51Z --- BEFORE --- containers=26 paperless=200 status=200 paste=200
2026-09-16T23:50:51Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T23:50:51Z --- ACTION: use on paperless-ngx ---
2026-09-16T23:50:51Z paperless read 1 -> 200
2026-09-16T23:50:52Z paperless read 2 -> 200
2026-09-16T23:50:52Z paperless read 3 -> 200
2026-09-16T23:50:52Z --- ACCIDENT: drive-pulled-20min (injected after the action started) ---
2026-09-16T23:50:52Z ACCIDENT=drive-pulled-20min round=11
2026-09-16T23:50:54Z data disk is nvme-scratch:336/vm-336-disk-0.raw — detaching it from the RUNNING box (the cable is pulled)
update VM 336: -delete scsi1
2026-09-16T23:50:56Z detached; the disk file stays as unused0
2026-09-16T23:50:56Z leaving it out for 20 minutes
update VM 336: -scsi1 nvme-scratch:336/vm-336-disk-0.raw
2026-09-17T00:10:58Z re-attached: nvme-scratch:336/vm-336-disk-0.raw
2026-09-17T00:10:58Z accident drive-pulled-20min complete
2026-09-17T00:10:58Z --- AFTER: what the box did BY ITSELF ---
2026-09-17T00:10:59Z t+1207s containers=15 (before 26)
2026-09-17T00:11:16Z t+1224s containers=15 (before 26)
2026-09-17T00:11:32Z t+1240s containers=18 (before 26)
2026-09-17T00:11:49Z t+1257s containers=25 (before 26)
2026-09-17T00:12:05Z t+1273s containers=26 (before 26)
2026-09-17T00:12:05Z STEADY after 1273s
2026-09-17T00:12:05Z front doors, FIRST reading at 2026-09-17T00:12:05Z - TOO EARLY to trust if the accident just ended:
2026-09-17T00:12:06Z paperless=404 status=200 paste=200 wiki=200
2026-09-17T00:13:06Z front doors, SECOND reading at 2026-09-17T00:13:06Z, 60 s later - THIS is the one to trust:
2026-09-17T00:13:07Z paperless=200 status=200 paste=200 wiki=200
2026-09-17T00:13:08Z household log lines: before=164 after=186
2026-09-17T00:13:08Z household lines this round: 22 failures: 0
2026-09-17T00:13:08Z --- alarms ---
| Time | Severity | Type | Message | Source
| Sep 17 00:11 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller
| Sep 16 23:53 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller
| Sep 16 23:51 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller
| Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller
2026-09-17T00:13:09Z ================ END ROUND 11 ================
[exited with code 0]
## THE MISSING ALARM WAS NOT MISSING - the round's own snapshot was seconds early (00:14:13Z)
At 00:13:08Z the round recorded health_degraded with NO health_recovered, which would have been the
first missing alarm of the night. Re-read at 00:14:13Z, after the apps were back (00:12:05Z):
| Sep 17 00:13 | info | health_recovered | Rendszer allapot helyreallt: ok (volt: warn) | controller
It fired at 00:13 and the snapshot was taken at 00:13:08 - missed by SECONDS.
This is the sixth-mistake fix earning its keep: the measurement is re-taken after its precondition
instead of a verdict being written from the first reading.
## THE FULL ALARM SET, and every one of them is TRUE
23:51 error storage_disconnected "Meghajto varatlanul levalasztva: Adatlemez" <- names the DRIVE
23:51 warning app_start_failed "Telepitett alkalmazas nem fut: Paperless-ngx" <- names the APP
23:51 warning app_start_failed "... Jellyfin"
23:51 warning app_start_failed "... Nextcloud"
23:51 warning app_start_failed "... Immich"
23:53 warning health_degraded "Rendszer allapot romlott (volt: ok)"
00:11 info storage_reconnected "Meghajto ujra csatlakoztatva: Adatlemez"
00:13 info health_recovered "Rendszer allapot helyreallt: ok (volt: warn)"
Seven alarms (plus the recovery), all true, correctly paired at both ends, and LEGIBLE: the drive is
named by the label the household sees ("Adatlemez"), and each broken app is named individually.
The four apps that failed are EXACTLY the four whose data lives on the pulled drive. The other
eleven kept running and kept serving. Nothing false was raised.
## THE RECOVERY, unaided
drive detached from the RUNNING box 23:50:54Z (scsi1 removed - the cable is pulled)
containers during the outage 26 -> 15
drive re-attached 00:10:58Z
containers back to 26 00:12:05Z = 67 SECONDS after the drive returned
paperless front door 404 at 00:12:05Z (67 s after), 200 at 00:13:06Z (128 s after)
household loop 22 lines this round, 0 failures (before=164 after=186)
No restart, no repair, nothing from me. The box noticed the drive was gone, said so, said which apps
it cost, degraded its own health, waited, noticed the drive return, restarted the apps and recovered.
## DRIVE VERIFIED BACK IN PLACE (00:13:30Z)
inside the guest: /dev/sdd 98G 1.4G used 92G avail 2% on /mnt/felhom-drives/hdd_1
guest config still carries mp8: /mnt/felhom-drives -> /mnt/felhom-drives
26 containers, none in a non-Up state; cloud/paperless/status/wiki all 200 on the public path