chaos night round 11: the drive pulled for 20 minutes, and eight true alarms
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
The richest alarm round of the night, and every alarm was true and correctly paired: storage_disconnected naming the drive by the household's own label, four app_start_failed naming exactly the four apps whose data lives on that drive, health_degraded, then storage_reconnected and health_recovered. The other eleven apps kept serving throughout. The box recovered unaided in 67 s after the drive was plugged back in, with the front door following at 128 s. The drive came back clean: 98 G, 2% used, mountpoint config unchanged. The round's own snapshot showed health_degraded with no recovery, which would have been the first missing alarm of the night. The recovery had fired seconds after the snapshot. A re-read taken after the precondition found it. Not a missing alarm - a premature reading, caught by the discipline the earlier mistimed readings forced. Truth table now: 17 alarms fired, 17 true, 0 missing, across eleven rounds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -460,6 +460,29 @@ a household count that reported 0 lines and 0 failures when the truth was one li
|
||||
failure; and a disk guard that reported „active" all night while being a **transient** unit that
|
||||
vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box.
|
||||
|
||||
### Round 11 — `use` paperless-ngx / accident: **the data drive pulled out for twenty minutes**
|
||||
|
||||
**23:50:49Z–00:13:09Z.** The drive was detached from the **running** box at 23:50:54Z and put back at
|
||||
00:10:58Z. This is the round that produced the most alarms of the night, and every one of them was
|
||||
true.
|
||||
|
||||
| the five things | |
|
||||
|---|---|
|
||||
| what the customer saw | **The four apps whose files live on that drive stopped** — Paperless, Jellyfin, Nextcloud, Immich — and Paperless's front door went 404. The other eleven apps kept serving normally throughout. About twenty minutes later everything was back, roughly a minute after the drive was plugged in again. |
|
||||
| what the box did by itself | Noticed the drive had gone and **named it by the label the household sees** („Adatlemez"), named **each** broken app individually, degraded its own health, waited, noticed the drive return, restarted the apps and recovered its health. No restart, no repair, nothing from me. |
|
||||
| time to steady | **67 s after the drive returned** (26 containers again at 00:12:05Z). The door followed at 00:13:06Z. |
|
||||
| alarm fired / true? | **eight, all true, correctly paired at both ends** — `storage_disconnected` (error) → four `app_start_failed` (warning) → `health_degraded` (warning) → `storage_reconnected` (info) → `health_recovered` (info). The four apps named are **exactly** the four with data on the pulled drive. Nothing false was raised. |
|
||||
| should have fired, did not | **none** |
|
||||
|
||||
**The alarm that looked missing, and was not.** The round's own snapshot at 00:13:08Z showed
|
||||
`health_degraded` with no recovery — which would have been the first missing alarm of the night. The
|
||||
recovery fired at **00:13**, and the snapshot missed it by **seconds**. A re-read at 00:14:13Z, taken
|
||||
after the apps were back, found it. This is the discipline from the earlier mistimed readings earning
|
||||
its keep: when a measurement could have been early, it is re-taken rather than turned into a verdict.
|
||||
|
||||
**The drive came back clean** — `/dev/sdd`, 98 G, 2 % used, mounted at `/mnt/felhom-drives/hdd_1`,
|
||||
with the guest's mountpoint config unchanged, 26 containers up and every real front door serving.
|
||||
|
||||
### Rounds 8-12
|
||||
|
||||
PENDING
|
||||
|
||||
@@ -21,10 +21,15 @@ R10| hard reset, 4 s into | controller_started (info)
|
||||
| an app restore | | | no restore field in the status JSON, only a button label on the pages, and no file at
|
||||
| | | | all modified in the reset window. Interrupted and never-happened look identical to the
|
||||
| | | | customer. Filed as R-550 (P2). The box itself recovered 26/26 containers in 150 s.
|
||||
R11| data drive pulled | storage_disconnected (error) -> app_start_failed x4 (warning)| YES (8/8) | none. The four apps named are EXACTLY the four whose data lives on the pulled drive;
|
||||
| out for 20 minutes | -> health_degraded (warning) -> storage_reconnected (info) | | the other eleven kept serving. The drive is named by the household's own label
|
||||
| | -> health_recovered (info) | | ("Adatlemez"). Recovery unaided in 67 s after the drive returned.
|
||||
| | | | NOTE: the round's own snapshot missed health_recovered by SECONDS. Re-read after
|
||||
| | | | the precondition found it. Not a missing alarm - a premature reading, caught.
|
||||
|
||||
## Running totals, rounds 1-10
|
||||
alarms fired: 9
|
||||
alarms TRUE: 9 (9/9 - no false alarm in ten rounds)
|
||||
## Running totals, rounds 1-11
|
||||
alarms fired: 17
|
||||
alarms TRUE: 17 (17/17 - no false alarm in eleven rounds)
|
||||
alarms MISSING: 0
|
||||
design gaps found: 3 (R-547: a transient full disk is never mentioned to anyone)
|
||||
(R-549: one failed report push spends the whole staleness budget)
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
round 11 armed for 2026-09-16T23:50:48Z (use paperless-ngx + drive-pulled-20min)
|
||||
=== round 11 launched 2026-09-16T23:50:49Z (due 23:50:48Z) ===
|
||||
2026-09-16T23:50:49Z ================ ROUND 11 : use paperless-ngx, while: drive-pulled-20min ================
|
||||
2026-09-16T23:50:51Z --- BEFORE --- containers=26 paperless=200 status=200 paste=200
|
||||
2026-09-16T23:50:51Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T23:50:51Z --- ACTION: use on paperless-ngx ---
|
||||
2026-09-16T23:50:51Z paperless read 1 -> 200
|
||||
2026-09-16T23:50:52Z paperless read 2 -> 200
|
||||
2026-09-16T23:50:52Z paperless read 3 -> 200
|
||||
2026-09-16T23:50:52Z --- ACCIDENT: drive-pulled-20min (injected after the action started) ---
|
||||
2026-09-16T23:50:52Z ACCIDENT=drive-pulled-20min round=11
|
||||
2026-09-16T23:50:54Z data disk is nvme-scratch:336/vm-336-disk-0.raw — detaching it from the RUNNING box (the cable is pulled)
|
||||
update VM 336: -delete scsi1
|
||||
2026-09-16T23:50:56Z detached; the disk file stays as unused0
|
||||
2026-09-16T23:50:56Z leaving it out for 20 minutes
|
||||
update VM 336: -scsi1 nvme-scratch:336/vm-336-disk-0.raw
|
||||
2026-09-17T00:10:58Z re-attached: nvme-scratch:336/vm-336-disk-0.raw
|
||||
2026-09-17T00:10:58Z accident drive-pulled-20min complete
|
||||
2026-09-17T00:10:58Z --- AFTER: what the box did BY ITSELF ---
|
||||
2026-09-17T00:10:59Z t+1207s containers=15 (before 26)
|
||||
2026-09-17T00:11:16Z t+1224s containers=15 (before 26)
|
||||
2026-09-17T00:11:32Z t+1240s containers=18 (before 26)
|
||||
2026-09-17T00:11:49Z t+1257s containers=25 (before 26)
|
||||
2026-09-17T00:12:05Z t+1273s containers=26 (before 26)
|
||||
2026-09-17T00:12:05Z STEADY after 1273s
|
||||
2026-09-17T00:12:05Z front doors, FIRST reading at 2026-09-17T00:12:05Z - TOO EARLY to trust if the accident just ended:
|
||||
2026-09-17T00:12:06Z paperless=404 status=200 paste=200 wiki=200
|
||||
2026-09-17T00:13:06Z front doors, SECOND reading at 2026-09-17T00:13:06Z, 60 s later - THIS is the one to trust:
|
||||
2026-09-17T00:13:07Z paperless=200 status=200 paste=200 wiki=200
|
||||
2026-09-17T00:13:08Z household log lines: before=164 after=186
|
||||
2026-09-17T00:13:08Z household lines this round: 22 failures: 0
|
||||
2026-09-17T00:13:08Z --- alarms ---
|
||||
| Time | Severity | Type | Message | Source
|
||||
| Sep 17 00:11 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller
|
||||
| Sep 16 23:53 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller
|
||||
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
|
||||
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller
|
||||
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller
|
||||
| Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller
|
||||
| Sep 16 23:51 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller
|
||||
| Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
2026-09-17T00:13:09Z ================ END ROUND 11 ================
|
||||
|
||||
[exited with code 0]
|
||||
|
||||
## THE MISSING ALARM WAS NOT MISSING - the round's own snapshot was seconds early (00:14:13Z)
|
||||
At 00:13:08Z the round recorded health_degraded with NO health_recovered, which would have been the
|
||||
first missing alarm of the night. Re-read at 00:14:13Z, after the apps were back (00:12:05Z):
|
||||
| Sep 17 00:13 | info | health_recovered | Rendszer allapot helyreallt: ok (volt: warn) | controller
|
||||
It fired at 00:13 and the snapshot was taken at 00:13:08 - missed by SECONDS.
|
||||
This is the sixth-mistake fix earning its keep: the measurement is re-taken after its precondition
|
||||
instead of a verdict being written from the first reading.
|
||||
|
||||
## THE FULL ALARM SET, and every one of them is TRUE
|
||||
23:51 error storage_disconnected "Meghajto varatlanul levalasztva: Adatlemez" <- names the DRIVE
|
||||
23:51 warning app_start_failed "Telepitett alkalmazas nem fut: Paperless-ngx" <- names the APP
|
||||
23:51 warning app_start_failed "... Jellyfin"
|
||||
23:51 warning app_start_failed "... Nextcloud"
|
||||
23:51 warning app_start_failed "... Immich"
|
||||
23:53 warning health_degraded "Rendszer allapot romlott (volt: ok)"
|
||||
00:11 info storage_reconnected "Meghajto ujra csatlakoztatva: Adatlemez"
|
||||
00:13 info health_recovered "Rendszer allapot helyreallt: ok (volt: warn)"
|
||||
Seven alarms (plus the recovery), all true, correctly paired at both ends, and LEGIBLE: the drive is
|
||||
named by the label the household sees ("Adatlemez"), and each broken app is named individually.
|
||||
The four apps that failed are EXACTLY the four whose data lives on the pulled drive. The other
|
||||
eleven kept running and kept serving. Nothing false was raised.
|
||||
|
||||
## THE RECOVERY, unaided
|
||||
drive detached from the RUNNING box 23:50:54Z (scsi1 removed - the cable is pulled)
|
||||
containers during the outage 26 -> 15
|
||||
drive re-attached 00:10:58Z
|
||||
containers back to 26 00:12:05Z = 67 SECONDS after the drive returned
|
||||
paperless front door 404 at 00:12:05Z (67 s after), 200 at 00:13:06Z (128 s after)
|
||||
household loop 22 lines this round, 0 failures (before=164 after=186)
|
||||
No restart, no repair, nothing from me. The box noticed the drive was gone, said so, said which apps
|
||||
it cost, degraded its own health, waited, noticed the drive return, restarted the apps and recovered.
|
||||
|
||||
## DRIVE VERIFIED BACK IN PLACE (00:13:30Z)
|
||||
inside the guest: /dev/sdd 98G 1.4G used 92G avail 2% on /mnt/felhom-drives/hdd_1
|
||||
guest config still carries mp8: /mnt/felhom-drives -> /mnt/felhom-drives
|
||||
26 containers, none in a non-Up state; cloud/paperless/status/wiki all 200 on the public path
|
||||
Reference in New Issue
Block a user