  round 11 armed for 2026-09-16T23:50:48Z (use paperless-ngx + drive-pulled-20min)
=== round 11 launched 2026-09-16T23:50:49Z (due 23:50:48Z) ===
2026-09-16T23:50:49Z ================ ROUND 11 : use paperless-ngx, while: drive-pulled-20min ================
2026-09-16T23:50:51Z --- BEFORE --- containers=26  paperless=200  status=200  paste=200
2026-09-16T23:50:51Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T23:50:51Z --- ACTION: use on paperless-ngx ---
2026-09-16T23:50:51Z     paperless read 1 -> 200
2026-09-16T23:50:52Z     paperless read 2 -> 200
2026-09-16T23:50:52Z     paperless read 3 -> 200
2026-09-16T23:50:52Z --- ACCIDENT: drive-pulled-20min (injected after the action started) ---
    2026-09-16T23:50:52Z ACCIDENT=drive-pulled-20min round=11
    2026-09-16T23:50:54Z data disk is nvme-scratch:336/vm-336-disk-0.raw — detaching it from the RUNNING box (the cable is pulled)
    update VM 336: -delete scsi1
    2026-09-16T23:50:56Z detached; the disk file stays as unused0
    2026-09-16T23:50:56Z leaving it out for 20 minutes
    update VM 336: -scsi1 nvme-scratch:336/vm-336-disk-0.raw
    2026-09-17T00:10:58Z re-attached: nvme-scratch:336/vm-336-disk-0.raw
    2026-09-17T00:10:58Z accident drive-pulled-20min complete
2026-09-17T00:10:58Z --- AFTER: what the box did BY ITSELF ---
2026-09-17T00:10:59Z     t+1207s containers=15 (before 26)
2026-09-17T00:11:16Z     t+1224s containers=15 (before 26)
2026-09-17T00:11:32Z     t+1240s containers=18 (before 26)
2026-09-17T00:11:49Z     t+1257s containers=25 (before 26)
2026-09-17T00:12:05Z     t+1273s containers=26 (before 26)
2026-09-17T00:12:05Z     STEADY after 1273s
2026-09-17T00:12:05Z     front doors, FIRST reading at 2026-09-17T00:12:05Z - TOO EARLY to trust if the accident just ended:
2026-09-17T00:12:06Z       paperless=404  status=200  paste=200  wiki=200
2026-09-17T00:13:06Z     front doors, SECOND reading at 2026-09-17T00:13:06Z, 60 s later - THIS is the one to trust:
2026-09-17T00:13:07Z       paperless=200  status=200  paste=200  wiki=200
2026-09-17T00:13:08Z     household log lines: before=164 after=186
2026-09-17T00:13:08Z     household lines this round: 22  failures: 0
2026-09-17T00:13:08Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 17 00:11 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller
  | Sep 16 23:53 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller
  | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
  | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller
  | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller
  | Sep 16 23:51 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller
  | Sep 16 23:51 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller
  | Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller
2026-09-17T00:13:09Z ================ END ROUND 11 ================

[exited with code 0]

## THE MISSING ALARM WAS NOT MISSING - the round's own snapshot was seconds early (00:14:13Z)
At 00:13:08Z the round recorded health_degraded with NO health_recovered, which would have been the
first missing alarm of the night. Re-read at 00:14:13Z, after the apps were back (00:12:05Z):
  | Sep 17 00:13 | info | health_recovered | Rendszer allapot helyreallt: ok (volt: warn) | controller
It fired at 00:13 and the snapshot was taken at 00:13:08 - missed by SECONDS.
This is the sixth-mistake fix earning its keep: the measurement is re-taken after its precondition
instead of a verdict being written from the first reading.

## THE FULL ALARM SET, and every one of them is TRUE
  23:51 error    storage_disconnected  "Meghajto varatlanul levalasztva: Adatlemez"   <- names the DRIVE
  23:51 warning  app_start_failed      "Telepitett alkalmazas nem fut: Paperless-ngx" <- names the APP
  23:51 warning  app_start_failed      "... Jellyfin"
  23:51 warning  app_start_failed      "... Nextcloud"
  23:51 warning  app_start_failed      "... Immich"
  23:53 warning  health_degraded       "Rendszer allapot romlott (volt: ok)"
  00:11 info     storage_reconnected   "Meghajto ujra csatlakoztatva: Adatlemez"
  00:13 info     health_recovered      "Rendszer allapot helyreallt: ok (volt: warn)"
Seven alarms (plus the recovery), all true, correctly paired at both ends, and LEGIBLE: the drive is
named by the label the household sees ("Adatlemez"), and each broken app is named individually.
The four apps that failed are EXACTLY the four whose data lives on the pulled drive. The other
eleven kept running and kept serving. Nothing false was raised.

## THE RECOVERY, unaided
  drive detached from the RUNNING box   23:50:54Z   (scsi1 removed - the cable is pulled)
  containers during the outage          26 -> 15
  drive re-attached                     00:10:58Z
  containers back to 26                 00:12:05Z   = 67 SECONDS after the drive returned
  paperless front door                  404 at 00:12:05Z (67 s after), 200 at 00:13:06Z (128 s after)
  household loop                        22 lines this round, 0 failures (before=164 after=186)
No restart, no repair, nothing from me. The box noticed the drive was gone, said so, said which apps
it cost, degraded its own health, waited, noticed the drive return, restarted the apps and recovered.

## DRIVE VERIFIED BACK IN PLACE (00:13:30Z)
  inside the guest: /dev/sdd  98G  1.4G used  92G avail  2%  on /mnt/felhom-drives/hdd_1
  guest config still carries mp8: /mnt/felhom-drives -> /mnt/felhom-drives
  26 containers, none in a non-Up state; cloud/paperless/status/wiki all 200 on the public path
