From 70bffb16769b96f9f5ebb413f59d385704599ab0 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 02:15:26 +0200 Subject: [PATCH] chaos night round 11: the drive pulled for 20 minutes, and eight true alarms The richest alarm round of the night, and every alarm was true and correctly paired: storage_disconnected naming the drive by the household's own label, four app_start_failed naming exactly the four apps whose data lives on that drive, health_degraded, then storage_reconnected and health_recovered. The other eleven apps kept serving throughout. The box recovered unaided in 67 s after the drive was plugged back in, with the front door following at 128 s. The drive came back clean: 98 G, 2% used, mountpoint config unchanged. The round's own snapshot showed health_degraded with no recovery, which would have been the first missing alarm of the night. The recovery had fired seconds after the snapshot. A re-read taken after the precondition found it. Not a missing alarm - a premature reading, caught by the discipline the earlier mistimed readings forced. Truth table now: 17 alarms fired, 17 true, 0 missing, across eleven rounds. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/DRILL-chaos-night-2026-09-17.md | 23 ++++++ .../alarm-truth-table.txt | 11 ++- .../round-11.txt | 81 +++++++++++++++++++ 3 files changed, 112 insertions(+), 3 deletions(-) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-11.txt diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 7d259c53..1d3f8bdd 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -460,6 +460,29 @@ a household count that reported 0 lines and 0 failures when the truth was one li failure; and a disk guard that reported „active" all night while being a **transient** unit that vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box. +### Round 11 — `use` paperless-ngx / accident: **the data drive pulled out for twenty minutes** + +**23:50:49Z–00:13:09Z.** The drive was detached from the **running** box at 23:50:54Z and put back at +00:10:58Z. This is the round that produced the most alarms of the night, and every one of them was +true. + +| the five things | | +|---|---| +| what the customer saw | **The four apps whose files live on that drive stopped** — Paperless, Jellyfin, Nextcloud, Immich — and Paperless's front door went 404. The other eleven apps kept serving normally throughout. About twenty minutes later everything was back, roughly a minute after the drive was plugged in again. | +| what the box did by itself | Noticed the drive had gone and **named it by the label the household sees** („Adatlemez"), named **each** broken app individually, degraded its own health, waited, noticed the drive return, restarted the apps and recovered its health. No restart, no repair, nothing from me. | +| time to steady | **67 s after the drive returned** (26 containers again at 00:12:05Z). The door followed at 00:13:06Z. | +| alarm fired / true? | **eight, all true, correctly paired at both ends** — `storage_disconnected` (error) → four `app_start_failed` (warning) → `health_degraded` (warning) → `storage_reconnected` (info) → `health_recovered` (info). The four apps named are **exactly** the four with data on the pulled drive. Nothing false was raised. | +| should have fired, did not | **none** | + +**The alarm that looked missing, and was not.** The round's own snapshot at 00:13:08Z showed +`health_degraded` with no recovery — which would have been the first missing alarm of the night. The +recovery fired at **00:13**, and the snapshot missed it by **seconds**. A re-read at 00:14:13Z, taken +after the apps were back, found it. This is the discipline from the earlier mistimed readings earning +its keep: when a measurement could have been early, it is re-taken rather than turned into a verdict. + +**The drive came back clean** — `/dev/sdd`, 98 G, 2 % used, mounted at `/mnt/felhom-drives/hdd_1`, +with the guest's mountpoint config unchanged, 26 containers up and every real front door serving. + ### Rounds 8-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/alarm-truth-table.txt b/documentation/audits/evidence-chaos-night-2026-09-17/alarm-truth-table.txt index b6fea207..bcdf4c0c 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/alarm-truth-table.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/alarm-truth-table.txt @@ -21,10 +21,15 @@ R10| hard reset, 4 s into | controller_started (info) | an app restore | | | no restore field in the status JSON, only a button label on the pages, and no file at | | | | all modified in the reset window. Interrupted and never-happened look identical to the | | | | customer. Filed as R-550 (P2). The box itself recovered 26/26 containers in 150 s. +R11| data drive pulled | storage_disconnected (error) -> app_start_failed x4 (warning)| YES (8/8) | none. The four apps named are EXACTLY the four whose data lives on the pulled drive; + | out for 20 minutes | -> health_degraded (warning) -> storage_reconnected (info) | | the other eleven kept serving. The drive is named by the household's own label + | | -> health_recovered (info) | | ("Adatlemez"). Recovery unaided in 67 s after the drive returned. + | | | | NOTE: the round's own snapshot missed health_recovered by SECONDS. Re-read after + | | | | the precondition found it. Not a missing alarm - a premature reading, caught. -## Running totals, rounds 1-10 - alarms fired: 9 - alarms TRUE: 9 (9/9 - no false alarm in ten rounds) +## Running totals, rounds 1-11 + alarms fired: 17 + alarms TRUE: 17 (17/17 - no false alarm in eleven rounds) alarms MISSING: 0 design gaps found: 3 (R-547: a transient full disk is never mentioned to anyone) (R-549: one failed report push spends the whole staleness budget) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-11.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-11.txt new file mode 100644 index 00000000..bfb1a82e --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-11.txt @@ -0,0 +1,81 @@ + round 11 armed for 2026-09-16T23:50:48Z (use paperless-ngx + drive-pulled-20min) +=== round 11 launched 2026-09-16T23:50:49Z (due 23:50:48Z) === +2026-09-16T23:50:49Z ================ ROUND 11 : use paperless-ngx, while: drive-pulled-20min ================ +2026-09-16T23:50:51Z --- BEFORE --- containers=26 paperless=200 status=200 paste=200 +2026-09-16T23:50:51Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) +2026-09-16T23:50:51Z --- ACTION: use on paperless-ngx --- +2026-09-16T23:50:51Z paperless read 1 -> 200 +2026-09-16T23:50:52Z paperless read 2 -> 200 +2026-09-16T23:50:52Z paperless read 3 -> 200 +2026-09-16T23:50:52Z --- ACCIDENT: drive-pulled-20min (injected after the action started) --- + 2026-09-16T23:50:52Z ACCIDENT=drive-pulled-20min round=11 + 2026-09-16T23:50:54Z data disk is nvme-scratch:336/vm-336-disk-0.raw — detaching it from the RUNNING box (the cable is pulled) + update VM 336: -delete scsi1 + 2026-09-16T23:50:56Z detached; the disk file stays as unused0 + 2026-09-16T23:50:56Z leaving it out for 20 minutes + update VM 336: -scsi1 nvme-scratch:336/vm-336-disk-0.raw + 2026-09-17T00:10:58Z re-attached: nvme-scratch:336/vm-336-disk-0.raw + 2026-09-17T00:10:58Z accident drive-pulled-20min complete +2026-09-17T00:10:58Z --- AFTER: what the box did BY ITSELF --- +2026-09-17T00:10:59Z t+1207s containers=15 (before 26) +2026-09-17T00:11:16Z t+1224s containers=15 (before 26) +2026-09-17T00:11:32Z t+1240s containers=18 (before 26) +2026-09-17T00:11:49Z t+1257s containers=25 (before 26) +2026-09-17T00:12:05Z t+1273s containers=26 (before 26) +2026-09-17T00:12:05Z STEADY after 1273s +2026-09-17T00:12:05Z front doors, FIRST reading at 2026-09-17T00:12:05Z - TOO EARLY to trust if the accident just ended: +2026-09-17T00:12:06Z paperless=404 status=200 paste=200 wiki=200 +2026-09-17T00:13:06Z front doors, SECOND reading at 2026-09-17T00:13:06Z, 60 s later - THIS is the one to trust: +2026-09-17T00:13:07Z paperless=200 status=200 paste=200 wiki=200 +2026-09-17T00:13:08Z household log lines: before=164 after=186 +2026-09-17T00:13:08Z household lines this round: 22 failures: 0 +2026-09-17T00:13:08Z --- alarms --- + | Time | Severity | Type | Message | Source + | Sep 17 00:11 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller + | Sep 16 23:53 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller + | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller + | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller + | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller + | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller + | Sep 16 23:51 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller + | Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller +2026-09-17T00:13:09Z ================ END ROUND 11 ================ + +[exited with code 0] + +## THE MISSING ALARM WAS NOT MISSING - the round's own snapshot was seconds early (00:14:13Z) +At 00:13:08Z the round recorded health_degraded with NO health_recovered, which would have been the +first missing alarm of the night. Re-read at 00:14:13Z, after the apps were back (00:12:05Z): + | Sep 17 00:13 | info | health_recovered | Rendszer allapot helyreallt: ok (volt: warn) | controller +It fired at 00:13 and the snapshot was taken at 00:13:08 - missed by SECONDS. +This is the sixth-mistake fix earning its keep: the measurement is re-taken after its precondition +instead of a verdict being written from the first reading. + +## THE FULL ALARM SET, and every one of them is TRUE + 23:51 error storage_disconnected "Meghajto varatlanul levalasztva: Adatlemez" <- names the DRIVE + 23:51 warning app_start_failed "Telepitett alkalmazas nem fut: Paperless-ngx" <- names the APP + 23:51 warning app_start_failed "... Jellyfin" + 23:51 warning app_start_failed "... Nextcloud" + 23:51 warning app_start_failed "... Immich" + 23:53 warning health_degraded "Rendszer allapot romlott (volt: ok)" + 00:11 info storage_reconnected "Meghajto ujra csatlakoztatva: Adatlemez" + 00:13 info health_recovered "Rendszer allapot helyreallt: ok (volt: warn)" +Seven alarms (plus the recovery), all true, correctly paired at both ends, and LEGIBLE: the drive is +named by the label the household sees ("Adatlemez"), and each broken app is named individually. +The four apps that failed are EXACTLY the four whose data lives on the pulled drive. The other +eleven kept running and kept serving. Nothing false was raised. + +## THE RECOVERY, unaided + drive detached from the RUNNING box 23:50:54Z (scsi1 removed - the cable is pulled) + containers during the outage 26 -> 15 + drive re-attached 00:10:58Z + containers back to 26 00:12:05Z = 67 SECONDS after the drive returned + paperless front door 404 at 00:12:05Z (67 s after), 200 at 00:13:06Z (128 s after) + household loop 22 lines this round, 0 failures (before=164 after=186) +No restart, no repair, nothing from me. The box noticed the drive was gone, said so, said which apps +it cost, degraded its own health, waited, noticed the drive return, restarted the apps and recovered. + +## DRIVE VERIFIED BACK IN PLACE (00:13:30Z) + inside the guest: /dev/sdd 98G 1.4G used 92G avail 2% on /mnt/felhom-drives/hdd_1 + guest config still carries mp8: /mnt/felhom-drives -> /mnt/felhom-drives + 26 containers, none in a non-Up state; cloud/paperless/status/wiki all 200 on the public path