From ac6600be3bedb3040722115421c976f427726db1 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 14 Sep 2026 23:05:22 +0200 Subject: [PATCH] BIGNIGHT F7 disk full: box holds, English banner, operator alarm silenced by cooldown; R-516/R-521 amended --- .../alarm-truth-table.md | 2 + .../evidence-bignight-2026-09-14/journal.md | 16 + .../phase5/F7-disk-full.txt | 27 ++ .../phase5/F7/backup-on-full-disk.txt | 18 + .../phase5/F7/controller.log | 437 ++++++++++++++++++ .../phase5/F7/recovery.txt | 14 + .../phase5/F7/screens.txt | 9 + .../phase5/F8-internet-gone.txt | 11 + documentation/backlog/OPEN-ITEMS.md | 4 +- 9 files changed, 536 insertions(+), 2 deletions(-) create mode 100644 documentation/audits/evidence-bignight-2026-09-14/phase5/F7-disk-full.txt create mode 100644 documentation/audits/evidence-bignight-2026-09-14/phase5/F7/backup-on-full-disk.txt create mode 100644 documentation/audits/evidence-bignight-2026-09-14/phase5/F7/controller.log create mode 100644 documentation/audits/evidence-bignight-2026-09-14/phase5/F7/recovery.txt create mode 100644 documentation/audits/evidence-bignight-2026-09-14/phase5/F7/screens.txt create mode 100644 documentation/audits/evidence-bignight-2026-09-14/phase5/F8-internet-gone.txt diff --git a/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md b/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md index b742b1f6..110d1372 100644 --- a/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md +++ b/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md @@ -27,6 +27,7 @@ Source: hub log (`alarms.sh`, CEST times), the operator mailbox and the customer | 22:38:33 | `storage_disconnected (error)` | F6 (second unplug) | **suppressed — cooldown** | true, **not delivered** (R-521) | | 22:38:45 | `app_start_failed (warning)` × 4 | F6 consequence | **suppressed — cooldown** × 4 | true, not delivered | | 22:39:45 | `health_degraded (warning)` „Rendszer állapot romlott (volt: ok)" | F6 | operator | true | +| 22:54:46 | `health_degraded (warning)` | F7 system disk 95 % | **suppressed — cooldown** | true, not delivered; no disk-specific event exists (R-521) | Hub WARN lines without an event or mail: `host tester-1-a61396 backup FAILED: target=felhom-pbs … does not exist` at 21:12:43 and 21:27:43 (the agent's own PBS attempts). @@ -43,3 +44,4 @@ Hub WARN lines without an event or mail: `host tester-1-a61396 backup FAILED: ta | F4 drive unplugged | drive lost | `storage_disconnected (error)` (+4 redundant) | correct | | F5 drive back after 30 min | drive back | `storage_reconnected (info)` | correct; customer banner stale 2–10 min | | F6 drive lost during backup | drive lost; backup incomplete | `storage_disconnected` logged, **mail suppressed by cooldown**; the run reported `success:true` | **MISSED** on both counts (R-521, R-519) | +| F7 system disk 95 % | disk nearly full | `health_degraded` only, mail suppressed; customer banner in English | **MISSED** for the operator (R-521); customer told, in English (R-516) | diff --git a/documentation/audits/evidence-bignight-2026-09-14/journal.md b/documentation/audits/evidence-bignight-2026-09-14/journal.md index bab0b617..0b13da2b 100644 --- a/documentation/audits/evidence-bignight-2026-09-14/journal.md +++ b/documentation/audits/evidence-bignight-2026-09-14/journal.md @@ -570,3 +570,19 @@ storage page (control „Tárhely" present) — the stale banner cleared within | drive back | all four drive apps running again at **+122 s** after re-attach | | the next run's honesty | `POST /api/backup/run` 20:47:24Z → 2 m 11 s, `success:true`; **all four drive apps dumped** (immich, jellyfin, nextcloud, paperless volume dumps 20:48–20:49Z); points nextcloud 20:48:59Z, immich 20:48:13Z, jellyfin 20:48:26Z, paperless 20:49:15Z; immich's torn `.tmp` gone. **Honest.** One leftover: F3's `pre-restore-20260914T195203Z-nextcloud-mariadb.sql.tmp` still in nextcloud's unit | | alarms | `storage_disconnected (error)` + 4 `app_start_failed` — **all five operator mails suppressed by cooldown** (a second, separate drive loss 40 min after the first, a reconnect in between) → **added to R-521**; `health_degraded (warning)` mailed — true | + +### F7 — the system disk fills up (`phase5/F7-disk-full.txt`, `phase5/F7/`) + +| | | +|---|---| +| how | `fallocate -l 41 591 249 306` in the guest at `/mnt/sys_drive/felhom-data/userdata/import/F7-nagy-fajl.bin` at **20:50:52Z** (sys_drive 69 G: 36 % → **95 %**, 3.5 G left). The box's thin pool did not grow (fallocate on ext4: `vm-9201-disk-1` data 36.94 %) | +| what the customer saw | dashboard from +3 s: „Rendszer (/) · **61.8 GB / 68.7 GB (90%)** · Kritikusan kevés hely" — **90 %, while `df` says 95 %** (the tile ignores the filesystem's reserved blocks; and the tile says „/" for the data volume `/mnt/sys_drive`). From ≈ +5 min a banner on every page: **„SSD disk usage high: 90%" — English** · „Rendszermonitor →". `/backups`: no warning. Deploy page (homebox): memory line only, **nothing about the disk** | +| does the box stay reachable | yes — dashboard and BookStack 200 throughout | +| what refuses first | nothing refused: a 4 MiB BookStack attachment still uploaded (200); a backup was started on the full disk (below) | +| alarms | hub 22:54:46 CEST `health_degraded (warning)` — **operator mail suppressed by cooldown** (the F6 `health_degraded` 15 min earlier). **No disk-specific event** reached the hub | +| a backup on the full disk | `POST /api/backup/run` 20:58:5xZ → finished 21:00:57Z, `success:true`; sys_drive stayed 95 % (units are replaced in place; the 3.5 G left sufficed); no `no space`/`ENOSPC` line; controller `[monitor] Disk (SSD) threshold breached: 90% (limit: 80%)`; apps running after (vaultwarden briefly `starting` from its own volume dump) | +| recovery | file removed 21:01:19Z → sys_drive 36 %; the English banner gone at +212 s (21:04:53Z); hub 23:04:45 CEST `health_recovered (info) — Rendszer állapot helyreállt: ok (volt: warn)` | +| alarms | see above: `health_degraded` (mail suppressed by cooldown) → `health_recovered`; **no disk-specific alarm** → R-521 amended; the English banner → R-516 amended | + +Verdict **PASS for the box** (reachable, nothing refused or broke at 95 %, the backup still completed); **the operator +was not told**. diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F7-disk-full.txt b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7-disk-full.txt new file mode 100644 index 00000000..3a80e342 --- /dev/null +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7-disk-full.txt @@ -0,0 +1,27 @@ +--- before +/dev/mapper/pve-vm--9201--disk--1 73793978368 24738168832 45280948224 36% /mnt/sys_drive +/dev/mapper/pve-vm--9201--disk--0 33501757440 989749248 30777245696 4% / +size=73793978368 avail=45280948224 -> fallocate 41591249306 B to leave 5% +F7 fill at 20:50:52 +/dev/mapper/pve-vm--9201--disk--1 69G 62G 3.5G 95% /mnt/sys_drive +20:50:55 +3s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:51:25 +33s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:51:55 +63s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:52:25 +93s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:52:56 +124s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:53:26 +154s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:53:57 +185s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:54:27 +215s wiki=200 dashboard: ↗ | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:54:57 +245s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:55:27 +275s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:55:58 +306s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:56:28 +336s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:56:59 +367s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:57:29 +397s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:57:59 +427s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +20:58:29 +457s wiki=200 dashboard: ↗ | SSD disk usage high: 90% | Rendszermonitor → | || Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés he +--- a family member writes: nextcloud is on the drive; bookstack attachment on the system disk +bookstack 4 MiB attachment on a 95% disk -> 200 {"name":"F7-teszt.bin","extension":"bin","uploaded_to":2,"created_by":1,"updated_by":1,"order":2,"updated_at":"2026-09-1 +--- backup now on a full disk +{"ok":true,"message":"Mentés elindítva"} + diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/backup-on-full-disk.txt b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/backup-on-full-disk.txt new file mode 100644 index 00000000..6bb63727 --- /dev/null +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/backup-on-full-disk.txt @@ -0,0 +1,18 @@ +20:59:18 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +20:59:29 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +20:59:41 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +20:59:53 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +21:00:05 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +21:00:17 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +21:00:29 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +21:00:40 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +21:00:52 backup(running,success,last)=True True 2026-09-14T20:49:36 sys_drive(used,avail,%)=62G 3.5G 95% +21:01:04 backup(running,success,last)=False True 2026-09-14T21:00:57 sys_drive(used,avail,%)=62G 3.5G 95% +--- errors during the full-disk backup +2026/09/14 20:59:45 [WARN] [monitor] Disk (SSD) threshold breached: 90% (limit: 80%) +--- states +bookstack running +docmost running +vaultwarden starting +mealie running +adventurelog running diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/controller.log b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/controller.log new file mode 100644 index 00000000..22daeaba --- /dev/null +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/controller.log @@ -0,0 +1,437 @@ +2026/09/14 20:50:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 20:50:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:50:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 20:50:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:50:45 [INFO] [scheduler] Job stack-scan completed (took 120ms) +2026/09/14 20:50:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:50:45 [INFO] [scheduler] Job agent-channel-health completed (took 179ms) +2026/09/14 20:50:49 [INFO] [web] Login from 172.18.0.10:55646 +2026/09/14 20:50:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:51:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:51:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:51:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:51:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:51:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:51:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:51:45 [INFO] [scheduler] Job agent-channel-health completed (took 181ms) +2026/09/14 20:51:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 20:52:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 20:52:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:45 [INFO] [scheduler] Job stack-scan completed (took 85ms) +2026/09/14 20:52:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:52:45 [INFO] [scheduler] Job agent-channel-health completed (took 199ms) +2026/09/14 20:52:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:52:55 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:53:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:53:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:53:15 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:53:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:53:25 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 20:53:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:53:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:53:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:53:45 [INFO] [scheduler] Job agent-channel-health completed (took 179ms) +2026/09/14 20:53:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:54:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:54:05 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:54:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:54:15 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 20:54:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:54:25 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:54:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:54:44 [INFO] [sync] Starting catalog sync +2026/09/14 20:54:44 [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main) +2026/09/14 20:54:44 [INFO] [sync] Catalog sync complete +2026/09/14 20:54:44 [INFO] [sync] Periodic sync: Sablonok naprakészek — nincs változás +2026/09/14 20:54:45 [INFO] [scheduler] Running job: backup-cache +2026/09/14 20:54:45 [INFO] [scheduler] Running job: hub-report +2026/09/14 20:54:45 [INFO] [report] Building system report +2026/09/14 20:54:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 20:54:45 [INFO] [scheduler] Running job: system-health +2026/09/14 20:54:45 [INFO] [scheduler] Running job: offsite-credential-retry +2026/09/14 20:54:45 [INFO] [scheduler] Job offsite-credential-retry completed (took 0s) +2026/09/14 20:54:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:54:45 [WARN] [monitor] Disk (SSD) threshold breached: 90% (limit: 80%) +2026/09/14 20:54:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 20:54:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:54:45 [INFO] [scheduler] Job stack-scan completed (took 131ms) +2026/09/14 20:54:45 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:54:45 [INFO] [backup] Found 6 DB dump files across drives +2026/09/14 20:54:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:54:45 [INFO] [scheduler] Running job: disk-health-check +2026/09/14 20:54:45 [WARN] [monitor] Disk (SSD) threshold breached: 90% (limit: 80%) +2026/09/14 20:54:45 [INFO] [monitor] Health check: status=warn +2026/09/14 20:54:45 [INFO] [scheduler] Job system-health completed (took 337ms) +2026/09/14 20:54:45 [INFO] [scheduler] Job agent-channel-health completed (took 197ms) +2026/09/14 20:54:45 [INFO] [backup] Discovered 6 databases +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bookstack/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/gokapi/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/immich/docker-compose.yml +2026/09/14 20:54:45 [INFO] [monitor] Health check: status=warn +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/mealie/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 2 HDD mounts for /opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/14 20:54:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/14 20:54:45 [INFO] [backup] Discovered app data: 12 apps +2026/09/14 20:54:46 [INFO] [backup] Backup status cache refreshed +2026/09/14 20:54:46 [INFO] [scheduler] Job backup-cache completed (took 482ms) +2026/09/14 20:54:46 [INFO] Event pushed: health_degraded (warning) — Rendszer állapot romlott (volt: ok) +2026/09/14 20:54:46 [INFO] [web] disk-health check complete: 0 disk(s) evaluated, 0 alert(s) +2026/09/14 20:54:46 [INFO] [scheduler] Job disk-health-check completed (took 347ms) +2026/09/14 20:54:46 [INFO] [report] Hub report pushed successfully (16304 bytes) +2026/09/14 20:54:46 [INFO] [settings] Settings saved +2026/09/14 20:54:46 [INFO] [scheduler] Job hub-report completed (took 1.021s) +2026/09/14 20:54:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:55:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:55:05 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 20:55:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:55:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:55:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:55:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:55:45 [INFO] [deadapp] check alive: 120 scans since boot, 12 deployed app(s) evaluated, 0 currently down +2026/09/14 20:55:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:55:45 [INFO] [scheduler] Job agent-channel-health completed (took 204ms) +2026/09/14 20:55:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:56:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:56:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:56:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:56:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:56:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 20:56:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:56:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 20:56:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:56:45 [INFO] [scheduler] Job stack-scan completed (took 94ms) +2026/09/14 20:56:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:56:45 [INFO] [scheduler] Job agent-channel-health completed (took 187ms) +2026/09/14 20:56:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:57:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:57:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:57:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:57:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:57:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:57:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:57:45 [INFO] [scheduler] Job agent-channel-health completed (took 171ms) +2026/09/14 20:57:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:57:55 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:58:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:58:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:58:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:58:25 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:58:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:58:35 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 20:58:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 20:58:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 20:58:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:58:45 [INFO] [scheduler] Job stack-scan completed (took 61ms) +2026/09/14 20:58:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:58:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:58:45 [INFO] [scheduler] Job agent-channel-health completed (took 194ms) +2026/09/14 20:58:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:00 [INFO] [api] Manual app-data backup (DB dump) triggered +2026/09/14 20:59:00 [INFO] [backup] Starting database dump run +2026/09/14 20:59:01 [INFO] [backup] Discovered 6 databases +2026/09/14 20:59:01 [INFO] [backup] Discovered 6 database(s): paperless-postgres(postgres), nextcloud-db(mariadb), immich-postgres(postgres), docmost-postgres(postgres), bookstack-db(mariadb), adventurelog-postgres(postgres) +2026/09/14 20:59:01 [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (320.3 KB, 311ms, 72 tables) +2026/09/14 20:59:01 [INFO] [settings] Settings saved +2026/09/14 20:59:02 [INFO] [backup] DB dump: nextcloud-db → nextcloud-mariadb.sql (1.4 MB, 640ms, 131 tables) +2026/09/14 20:59:02 [INFO] [settings] Settings saved +2026/09/14 20:59:03 [INFO] [backup] DB dump: immich-postgres → immich-postgres.sql (51.0 MB, 807ms, 66 tables) +2026/09/14 20:59:03 [INFO] [settings] Settings saved +2026/09/14 20:59:03 [INFO] [backup] DB dump: docmost-postgres → docmost-postgres.sql (139.9 KB, 347ms, 42 tables) +2026/09/14 20:59:03 [INFO] [settings] Settings saved +2026/09/14 20:59:03 [INFO] [backup] DB dump: bookstack-db → bookstack-mariadb.sql (66.6 KB, 315ms, 41 tables) +2026/09/14 20:59:03 [INFO] [settings] Settings saved +2026/09/14 20:59:04 [INFO] [backup] DB dump: adventurelog-postgres → adventurelog-postgres.sql (153.0 KB, 702ms, 47 tables) +2026/09/14 20:59:04 [INFO] [settings] Settings saved +2026/09/14 20:59:04 [INFO] [backup] Stopping adventurelog for safe volume dump +2026/09/14 20:59:04 [INFO] [stacks] Stopping stack: adventurelog +2026/09/14 20:59:05 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 20:59:07 [INFO] [stacks] Stack adventurelog stopped successfully (took 3.0s) +2026/09/14 20:59:07 [INFO] [stacks] Status refresh: 25 containers across 57 stacks +2026/09/14 20:59:08 [INFO] [backup] Volume dump: adventurelog/adventurelog_adventurelog_media → 2.0 KB +2026/09/14 20:59:08 [INFO] [backup] Volume dump: adventurelog/adventurelog_adventurelog_postgres_data → 111.4 MB +2026/09/14 20:59:08 [INFO] [backup] Restarting adventurelog after volume dump +2026/09/14 20:59:08 [INFO] [stacks] Starting stack: adventurelog +2026/09/14 20:59:15 [INFO] [stacks] Stack adventurelog started successfully (took 6.1s) +2026/09/14 20:59:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:15 [INFO] [backup] Stopping bookstack for safe volume dump +2026/09/14 20:59:15 [INFO] [stacks] Stopping stack: bookstack +2026/09/14 20:59:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:15 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:59:18 [INFO] [stacks] Stack adventurelog post-start status: +2026/09/14 20:59:18 [INFO] [stacks] adventurelog ghcr.io/seanmorley15/adventurelog-backend:v0.12.1 running Up 3 seconds (health: starting) +2026/09/14 20:59:18 [INFO] [stacks] adventurelog-frontend ghcr.io/seanmorley15/adventurelog-frontend:v0.12.1 running Up 9 seconds (healthy) +2026/09/14 20:59:18 [INFO] [stacks] adventurelog-postgres postgis/postgis:16-3.5-alpine running Up 9 seconds (healthy) +2026/09/14 20:59:19 [INFO] [stacks] Stack bookstack stopped successfully (took 4.4s) +2026/09/14 20:59:19 [INFO] [stacks] Status refresh: 26 containers across 57 stacks +2026/09/14 20:59:20 [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_db_data → 153.4 MB +2026/09/14 20:59:20 [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_config → 5.1 MB +2026/09/14 20:59:20 [INFO] [backup] Restarting bookstack after volume dump +2026/09/14 20:59:20 [INFO] [stacks] Starting stack: bookstack +2026/09/14 20:59:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:25 [INFO] Health probes: 3 ok (of 3 probed) +2026/09/14 20:59:26 [INFO] [stacks] Stack bookstack started successfully (took 6.2s) +2026/09/14 20:59:27 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:27 [INFO] [backup] Stopping docmost for safe volume dump +2026/09/14 20:59:27 [INFO] [stacks] Stopping stack: docmost +2026/09/14 20:59:27 [INFO] [stacks] Stack docmost stopped successfully (took 0.7s) +2026/09/14 20:59:27 [INFO] [stacks] Status refresh: 25 containers across 57 stacks +2026/09/14 20:59:28 [INFO] [backup] Volume dump: docmost/docmost_docmost_storage → 1.5 KB +2026/09/14 20:59:28 [INFO] [backup] Volume dump: docmost/docmost_docmost_postgres_data → 65.5 MB +2026/09/14 20:59:29 [INFO] [backup] Volume dump: docmost/docmost_docmost_redis_data → 673.0 KB +2026/09/14 20:59:29 [INFO] [backup] Restarting docmost after volume dump +2026/09/14 20:59:29 [INFO] [stacks] Starting stack: docmost +2026/09/14 20:59:30 [INFO] [stacks] Stack bookstack post-start status: +2026/09/14 20:59:30 [INFO] [stacks] bookstack lscr.io/linuxserver/bookstack:26.05.2 running Up 3 seconds (health: starting) +2026/09/14 20:59:30 [INFO] [stacks] bookstack-db mariadb:12.3 running Up 9 seconds (healthy) +2026/09/14 20:59:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:35 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:59:40 [INFO] [stacks] Stack docmost started successfully (took 11.3s) +2026/09/14 20:59:40 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:40 [INFO] [backup] Stopping gokapi for safe volume dump +2026/09/14 20:59:40 [INFO] [stacks] Stopping stack: gokapi +2026/09/14 20:59:41 [INFO] [stacks] Stack gokapi stopped successfully (took 0.3s) +2026/09/14 20:59:41 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 20:59:41 [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_config → 3.5 KB +2026/09/14 20:59:41 [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_data → 51.4 MB +2026/09/14 20:59:41 [INFO] [backup] Restarting gokapi after volume dump +2026/09/14 20:59:41 [INFO] [stacks] Starting stack: gokapi +2026/09/14 20:59:42 [INFO] [stacks] Stack gokapi started successfully (took 0.4s) +2026/09/14 20:59:42 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:42 [INFO] [backup] Stopping immich for safe volume dump +2026/09/14 20:59:42 [INFO] [stacks] Stopping stack: immich +2026/09/14 20:59:43 [INFO] [stacks] Stack immich stopped successfully (took 0.8s) +2026/09/14 20:59:43 [INFO] [stacks] Status refresh: 24 containers across 57 stacks +2026/09/14 20:59:43 [INFO] [stacks] Stack docmost post-start status: +2026/09/14 20:59:43 [INFO] [stacks] docmost docmost/docmost:0.95.0 running Up 3 seconds (health: starting) +2026/09/14 20:59:43 [INFO] [stacks] docmost-postgres postgres:16-alpine running Up 14 seconds (healthy) +2026/09/14 20:59:43 [INFO] [stacks] docmost-redis redis:7-alpine running Up 14 seconds (healthy) +2026/09/14 20:59:44 [INFO] [backup] Volume dump: immich/immich_immich_ml_cache → 785.5 MB +2026/09/14 20:59:45 [INFO] [scheduler] Running job: offsite-credential-retry +2026/09/14 20:59:45 [INFO] [scheduler] Running job: system-health +2026/09/14 20:59:45 [INFO] [scheduler] Running job: backup-cache +2026/09/14 20:59:45 [INFO] [scheduler] Job offsite-credential-retry completed (took 0s) +2026/09/14 20:59:45 [INFO] [stacks] Status refresh: 24 containers across 57 stacks +2026/09/14 20:59:45 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:59:45 [INFO] [backup] Found 6 DB dump files across drives +2026/09/14 20:59:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 20:59:45 [INFO] [backup] Volume dump: immich/immich_immich_postgres_data → 327.8 MB +2026/09/14 20:59:45 [WARN] [monitor] Disk (SSD) threshold breached: 90% (limit: 80%) +2026/09/14 20:59:45 [INFO] [stacks] Stack gokapi post-start status: +2026/09/14 20:59:45 [INFO] [stacks] gokapi f0rc3/gokapi:v1.9.6 running Up 3 seconds (health: starting) +2026/09/14 20:59:45 [INFO] [scheduler] Job agent-channel-health completed (took 213ms) +2026/09/14 20:59:45 [INFO] [backup] Discovered 5 databases +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bookstack/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/gokapi/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/immich/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/mealie/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 2 HDD mounts for /opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/14 20:59:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/14 20:59:45 [INFO] [backup] Discovered app data: 12 apps +2026/09/14 20:59:46 [INFO] [backup] Backup status cache refreshed +2026/09/14 20:59:46 [INFO] [scheduler] Job backup-cache completed (took 477ms) +2026/09/14 20:59:46 [INFO] [monitor] Health check: status=warn +2026/09/14 20:59:46 [INFO] [scheduler] Job system-health completed (took 490ms) +2026/09/14 20:59:46 [INFO] [backup] Volume dump: immich/immich_immich_redis_data → 3.7 MB +2026/09/14 20:59:46 [INFO] [backup] Restarting immich after volume dump +2026/09/14 20:59:46 [INFO] [stacks] Starting stack: immich +2026/09/14 20:59:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:55 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 20:59:57 [INFO] [stacks] Stack immich started successfully (took 11.3s) +2026/09/14 20:59:57 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:57 [INFO] [backup] Stopping jellyfin for safe volume dump +2026/09/14 20:59:57 [INFO] [stacks] Stopping stack: jellyfin +2026/09/14 20:59:58 [INFO] [stacks] Stack jellyfin stopped successfully (took 0.5s) +2026/09/14 20:59:58 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 20:59:58 [INFO] [backup] Volume dump: jellyfin/jellyfin_jellyfin_cache → 3.0 KB +2026/09/14 20:59:59 [INFO] [backup] Volume dump: jellyfin/jellyfin_jellyfin_config → 705.0 KB +2026/09/14 20:59:59 [INFO] [backup] Restarting jellyfin after volume dump +2026/09/14 20:59:59 [INFO] [stacks] Starting stack: jellyfin +2026/09/14 20:59:59 [INFO] [stacks] Stack jellyfin started successfully (took 0.4s) +2026/09/14 20:59:59 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 20:59:59 [INFO] [backup] Stopping mealie for safe volume dump +2026/09/14 20:59:59 [INFO] [stacks] Stopping stack: mealie +2026/09/14 21:00:01 [INFO] [stacks] Stack immich post-start status: +2026/09/14 21:00:01 [INFO] [stacks] immich-machine-learning ghcr.io/immich-app/immich-machine-learning:v3.0.3 running Up 14 seconds (healthy) +2026/09/14 21:00:01 [INFO] [stacks] immich-postgres ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 running Up 14 seconds (healthy) +2026/09/14 21:00:01 [INFO] [stacks] immich-redis redis:7-alpine running Up 14 seconds (healthy) +2026/09/14 21:00:01 [INFO] [stacks] immich-server ghcr.io/immich-app/immich-server:v3.0.3 running Up 3 seconds (health: starting) +2026/09/14 21:00:01 [INFO] [stacks] Stack mealie stopped successfully (took 2.1s) +2026/09/14 21:00:01 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 21:00:02 [INFO] [backup] Volume dump: mealie/mealie_mealie_data → 1.5 MB +2026/09/14 21:00:02 [INFO] [backup] Restarting mealie after volume dump +2026/09/14 21:00:02 [INFO] [stacks] Starting stack: mealie +2026/09/14 21:00:02 [INFO] [stacks] Stack jellyfin post-start status: +2026/09/14 21:00:02 [INFO] [stacks] jellyfin jellyfin/jellyfin:10.11.11 running Up 3 seconds (health: starting) +2026/09/14 21:00:02 [INFO] [stacks] Stack mealie started successfully (took 0.6s) +2026/09/14 21:00:03 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:03 [INFO] [backup] Stopping nextcloud for safe volume dump +2026/09/14 21:00:03 [INFO] [stacks] Stopping stack: nextcloud +2026/09/14 21:00:05 [INFO] [stacks] Status refresh: 26 containers across 57 stacks +2026/09/14 21:00:05 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 21:00:05 [INFO] [stacks] Stack nextcloud stopped successfully (took 2.5s) +2026/09/14 21:00:05 [INFO] [stacks] Status refresh: 25 containers across 57 stacks +2026/09/14 21:00:06 [INFO] [stacks] Stack mealie post-start status: +2026/09/14 21:00:06 [INFO] [stacks] mealie ghcr.io/mealie-recipes/mealie:v3.20.1 running Up 3 seconds (health: starting) +2026/09/14 21:00:06 [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_db_data → 174.9 MB +2026/09/14 21:00:15 [INFO] [stacks] Status refresh: 25 containers across 57 stacks +2026/09/14 21:00:15 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 21:00:21 [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_html → 745.1 MB +2026/09/14 21:00:22 [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_redis_data → 5.2 MB +2026/09/14 21:00:22 [INFO] [backup] Restarting nextcloud after volume dump +2026/09/14 21:00:22 [INFO] [stacks] Starting stack: nextcloud +2026/09/14 21:00:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:28 [INFO] [stacks] Stack nextcloud started successfully (took 6.3s) +2026/09/14 21:00:28 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:28 [INFO] [backup] Stopping paperless-ngx for safe volume dump +2026/09/14 21:00:28 [INFO] [stacks] Stopping stack: paperless-ngx +2026/09/14 21:00:31 [INFO] [stacks] Stack nextcloud post-start status: +2026/09/14 21:00:31 [INFO] [stacks] nextcloud nextcloud:34.0.1-apache running Up 3 seconds (health: starting) +2026/09/14 21:00:31 [INFO] [stacks] nextcloud-db mariadb:11.6 running Up 9 seconds (healthy) +2026/09/14 21:00:31 [INFO] [stacks] nextcloud-redis redis:7-alpine running Up 9 seconds (healthy) +2026/09/14 21:00:35 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 21:00:35 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 21:00:35 [INFO] [stacks] Stack paperless-ngx stopped successfully (took 6.9s) +2026/09/14 21:00:35 [INFO] [stacks] Status refresh: 25 containers across 57 stacks +2026/09/14 21:00:36 [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 67.8 MB +2026/09/14 21:00:36 [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 231.5 KB +2026/09/14 21:00:37 [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 274.5 KB +2026/09/14 21:00:37 [INFO] [backup] Restarting paperless-ngx after volume dump +2026/09/14 21:00:37 [INFO] [stacks] Starting stack: paperless-ngx +2026/09/14 21:00:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 21:00:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 21:00:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:45 [INFO] [scheduler] Job stack-scan completed (took 101ms) +2026/09/14 21:00:45 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 21:00:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 21:00:45 [INFO] [scheduler] Job agent-channel-health completed (took 180ms) +2026/09/14 21:00:48 [INFO] [stacks] Stack paperless-ngx started successfully (took 11.2s) +2026/09/14 21:00:48 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:48 [INFO] [backup] Stopping privatebin for safe volume dump +2026/09/14 21:00:48 [INFO] [stacks] Stopping stack: privatebin +2026/09/14 21:00:48 [INFO] [stacks] Stack privatebin stopped successfully (took 0.3s) +2026/09/14 21:00:48 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 21:00:49 [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 17.0 KB +2026/09/14 21:00:49 [INFO] [backup] Restarting privatebin after volume dump +2026/09/14 21:00:49 [INFO] [stacks] Starting stack: privatebin +2026/09/14 21:00:49 [INFO] [stacks] Stack privatebin started successfully (took 0.4s) +2026/09/14 21:00:49 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:49 [INFO] [backup] Stopping uptime-kuma for safe volume dump +2026/09/14 21:00:49 [INFO] [stacks] Stopping stack: uptime-kuma +2026/09/14 21:00:51 [INFO] [stacks] Stack paperless-ngx post-start status: +2026/09/14 21:00:51 [INFO] [stacks] paperless-postgres postgres:16-alpine running Up 14 seconds (healthy) +2026/09/14 21:00:51 [INFO] [stacks] paperless-redis redis:7-alpine running Up 14 seconds (healthy) +2026/09/14 21:00:51 [INFO] [stacks] paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 running Up 3 seconds (health: starting) +2026/09/14 21:00:53 [INFO] [stacks] Stack privatebin post-start status: +2026/09/14 21:00:53 [INFO] [stacks] privatebin privatebin/pdo:2.0.6 running Up 3 seconds (health: starting) +2026/09/14 21:00:54 [INFO] [stacks] Stack uptime-kuma stopped successfully (took 4.5s) +2026/09/14 21:00:54 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 21:00:54 [INFO] [backup] Volume dump: uptime-kuma/uptime-kuma_uptime_kuma_data → 428.5 KB +2026/09/14 21:00:54 [INFO] [backup] Restarting uptime-kuma after volume dump +2026/09/14 21:00:54 [INFO] [stacks] Starting stack: uptime-kuma +2026/09/14 21:00:55 [INFO] [stacks] Stack uptime-kuma started successfully (took 0.5s) +2026/09/14 21:00:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:55 [INFO] [backup] Stopping vaultwarden for safe volume dump +2026/09/14 21:00:55 [INFO] [stacks] Stopping stack: vaultwarden +2026/09/14 21:00:56 [INFO] [stacks] Stack vaultwarden stopped successfully (took 0.4s) +2026/09/14 21:00:56 [INFO] [stacks] Status refresh: 27 containers across 57 stacks +2026/09/14 21:00:56 [INFO] [backup] Volume dump: vaultwarden/vaultwarden_vaultwarden_data → 433.0 KB +2026/09/14 21:00:56 [INFO] [backup] Restarting vaultwarden after volume dump +2026/09/14 21:00:56 [INFO] [stacks] Starting stack: vaultwarden +2026/09/14 21:00:57 [INFO] [stacks] Stack vaultwarden started successfully (took 0.7s) +2026/09/14 21:00:57 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:00:57 [INFO] [backup] App-data backup completed: 6 databases (53.1 MB total), 12 volume dump(s) (1m56.653s) +2026/09/14 21:00:58 [INFO] [stacks] Stack uptime-kuma post-start status: +2026/09/14 21:00:58 [INFO] [stacks] uptime-kuma louislam/uptime-kuma:2.4.0 running Up 3 seconds (health: starting) +2026/09/14 21:01:00 [INFO] [stacks] Stack vaultwarden post-start status: +2026/09/14 21:01:00 [INFO] [stacks] vaultwarden vaultwarden/server:1.36.0-alpine running Up 3 seconds (health: starting) +2026/09/14 21:01:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:01:05 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 21:01:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:01:15 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 21:01:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:01:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:01:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:01:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 21:01:45 [INFO] [scheduler] Job agent-channel-health completed (took 186ms) +2026/09/14 21:01:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:02:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:02:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:02:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:02:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:02:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 21:02:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:02:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 21:02:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:02:45 [INFO] [scheduler] Job stack-scan completed (took 98ms) +2026/09/14 21:02:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 21:02:45 [INFO] [scheduler] Job agent-channel-health completed (took 180ms) +2026/09/14 21:02:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:03:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:03:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:03:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:03:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:03:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:03:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 21:03:45 [INFO] [scheduler] Job agent-channel-health completed (took 183ms) +2026/09/14 21:03:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:04:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:04:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:04:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:04:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:04:35 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 21:04:45 [INFO] [scheduler] Running job: backup-cache +2026/09/14 21:04:45 [INFO] [scheduler] Running job: offsite-credential-retry +2026/09/14 21:04:45 [INFO] [scheduler] Job offsite-credential-retry completed (took 0s) +2026/09/14 21:04:45 [INFO] [scheduler] Running job: stack-scan +2026/09/14 21:04:45 [INFO] [scheduler] Running job: system-health +2026/09/14 21:04:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:04:45 [INFO] [stacks] ScanStacks complete: 57 stacks found (12 deployed, 45 available) +2026/09/14 21:04:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:04:45 [INFO] [scheduler] Job stack-scan completed (took 112ms) +2026/09/14 21:04:45 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 21:04:45 [INFO] [backup] Found 6 DB dump files across drives +2026/09/14 21:04:45 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 21:04:45 [INFO] [monitor] Health check: status=ok +2026/09/14 21:04:45 [INFO] [scheduler] Job system-health completed (took 245ms) +2026/09/14 21:04:45 [INFO] [backup] Discovered 6 databases +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bookstack/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/gokapi/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/immich/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/mealie/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 2 HDD mounts for /opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/14 21:04:45 [INFO] Event pushed: health_recovered (info) — Rendszer állapot helyreállt: ok (volt: warn) +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/14 21:04:45 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/14 21:04:45 [INFO] [backup] Discovered app data: 12 apps +2026/09/14 21:04:45 [INFO] [backup] Backup status cache refreshed +2026/09/14 21:04:45 [INFO] [scheduler] Job backup-cache completed (took 316ms) +2026/09/14 21:04:45 [INFO] [scheduler] Job agent-channel-health completed (took 169ms) +2026/09/14 21:04:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:05:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 21:05:05 [INFO] Health probes: 2 ok (of 2 probed) +2026/09/14 21:05:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/recovery.txt b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/recovery.txt new file mode 100644 index 00000000..a10858e7 --- /dev/null +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/recovery.txt @@ -0,0 +1,14 @@ +F7 remove the file at 21:01:19 +/dev/mapper/pve-vm--9201--disk--1 69G 24G 43G 36% /mnt/sys_drive +21:01:21 +0s banner-SSD 1 kritikus 0 +21:01:51 +30s banner-SSD 1 kritikus 0 +21:02:22 +61s banner-SSD 1 kritikus 0 +21:02:52 +91s banner-SSD 1 kritikus 0 +21:03:22 +121s banner-SSD 1 kritikus 0 +21:03:52 +151s banner-SSD 1 kritikus 0 +21:04:23 +182s banner-SSD 1 kritikus 0 +21:04:53 +212s banner-SSD 0 kritikus 0 +--- hub +2026/09/14 22:54:46 [INFO] Event from tester-1: health_degraded (warning) — Rendszer állapot romlott (volt: ok) +2026/09/14 22:54:46 [INFO] Operator email suppressed for tester-1/health_degraded — cooldown (key=tester-1:health_degraded) +2026/09/14 23:04:45 [INFO] Event from tester-1: health_recovered (info) — Rendszer állapot helyreállt: ok (volt: warn) diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/screens.txt b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/screens.txt new file mode 100644 index 00000000..1fc77090 --- /dev/null +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F7/screens.txt @@ -0,0 +1,9 @@ +=== /dashboard 20:51:42 + kritikus ctx: lt | enkicsifelhom.hu | 16 | Futó alkalmazás | 0 | Leállítva | 57 | Összes alkalmazás | Memória | 3.8 GB / 12 GB (33%) | CPU | 5% | Load: 1.16 / 1.74 / 1.26 | Rendszer (/) | 61.8 GB / 68.7 GB (90%) | Kritikusan kevés hely | Adatlemez | 3.47 GB / 97.9 GB (4%) | Lemezek állapota | QEMU QEMU HARDDISK | Nincs adat | 0 °C | QEMU QEMU HARDDISK | Nincs adat | 0 °C | Biztonsági mentés | Utolsó mentés: | 2026-09-14 20:49 | Adatbázisok: | 6 mentve | Telepített alkal + deploy memory/disk ctx: +=== /backups 20:51:42 + kritikus ctx: absent + deploy memory/disk ctx: +=== /stacks/homebox/deploy 20:51:42 + kritikus ctx: absent + deploy memory/disk ctx: Memória | 3932 MB / 11444 MB (34%) | Jelenlegi használat (3882 MB) | Homebox (+50 MB) | Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. | Normál használat mellett ez nem okoz problémát. | Hol lesznek az adatok: | ennél az alkalmazásnál nincs külön adatmeghajtó-választás, diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F8-internet-gone.txt b/documentation/audits/evidence-bignight-2026-09-14/phase5/F8-internet-gone.txt new file mode 100644 index 00000000..17a7e787 --- /dev/null +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F8-internet-gone.txt @@ -0,0 +1,11 @@ +Error: syntax error, unexpected fwd, expecting string or last +add chain bridge bignight fwd { type filter hook forward priority 0; policy accept; } + ^^^ +Error: syntax error, unexpected policy +add chain bridge bignight fwd { type filter hook forward priority 0; policy accept; } + ^^^^^^ +Error: syntax error, unexpected '}' +add chain bridge bignight fwd { type filter hook forward priority 0; policy accept; } + ^ +F8 internet cut at 21:05:19 +21:05:20 +1s LAN-dashboard=200 guest->internet=302 public-name=502 diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 0225b3f4..94887002 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -719,12 +719,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-513** | **[P1-HIGH — SECURITY] Every box's file manager (FileBrowser, `files.`, a launcher tile) accepts the login `admin` / `admin`, and on demo-hp that login page is on the public internet.** MEASURED 2026-09-14 (BIGNIGHT): `POST /api/auth/login?username=admin` with `X-Password: admin` → **200 + a session token** on VM 333 (fresh ISO 1.27.1 install, controller 0.242.0) and on demo-hp guests **9201 and 9202** (loopback, `Host: files.enkisfelhom.hu`); negative control `admin` / wrong → **401** on all three. On VM 333 that token lists both sources — „Adatlemez" (the data drive's `userdata`: documents, media, photos…) and „Beolvasás" (with `paperless`) — `GET /api/users?id=self` 200. **Public exposure, measured by GET only:** `https://files.enkisfelhom.hu/` from DooPlex through Cloudflare → 200, FileBrowser Quantum, `passwordAvailable:true, noAuth:false` (`phase3/filebrowser-public-reachability-demo-hp.txt`); no login was attempted over the internet. The generated `config.yaml` sets no admin credential (FileBrowser's own default applies); no screen shows the customer any FileBrowser login. Geo-restriction narrows who can reach it, it does not authenticate. **Not changed tonight** (9201's standing state is fenced; no product code). **Fix shape:** the controller sets a generated admin password (or proxy auth behind the dashboard session) at stack creation and on every existing box, and shows it where the customer finds app credentials. | **READY — rank P1-HIGH; owner: CC (controller) · operator (rotate on live boxes first)** | | **R-514** | **[P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, catalog paperless-ngx 2.20.15, `paperless-webserver` limit 805 306 368 B): 20 small 3-page PDFs posted through `/api/documents/post_document/` at 18:29:57Z. At 18:31:22Z the VM kernel logged `Memory cgroup out of memory: Killed process … (gs)` ×2, `([celeryd: celer)`, `([celery beat] -)` — constraint MEMCG of that container; `docker inspect` → `oomkilled=true restarts=0`. Paperless's own task list: **11 FAILURE (`WorkerLostError`), 1 STARTED, 8 PENDING**, unchanged 13 minutes later; **0 documents**. The controller shows the app running and healthy; no event, no alarm (`phase3/paperless-tasks.txt`, `paperless-oom-check.txt`). A household scanning a drawer of bills sees nothing arrive and no reason. **Fix shape:** raise the cap or set `PAPERLESS_TASK_WORKERS=1` / `PAPERLESS_THREADS_PER_WORKER=1` in the template, and let the dead-app/health check see an OOM-killed worker. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | | **R-515** | **[P3-LOW] The Paperless-ngx app page tells the customer to log in with `admin / admin`, and that login does not exist: the deploy form generates the admin password.** MEASURED 2026-09-14 (BIGNIGHT, VM 333): `/apps/paperless-ngx` „Első lépések — Jelentkezz be: admin / admin" and „Alapértelmezett belépés admin / admin"; the catalog template generates `PAPERLESS_ADMIN_PASSWORD` (`password:16`) and the token call with the generated value succeeded (`phase3/seed-paperless.txt`). A household following the page is refused at its first login. Same card also sends documents to „FileBrowser … import/paperless" — the FileBrowser login is R-513. **Fix:** the card points at „Beállítások → Automatikusan generált értékek" as gokapi's does. | **READY — rank P3-LOW; owner: CC (catalog)** | -| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. | **READY — rank P3-LOW; owner: CC (controller + catalog)** | +| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. | **READY — rank P3-LOW; owner: CC (controller + catalog)** | | **R-517** | **[P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, controller 0.242.0, customer `tester-1` with the DR tier ticked but never provisioned — R-511): „Mentés most" at 19:03:23Z; the local tier succeeded (8 877 619 753 B, 362 s); the `felhom-pbs` tier then failed — agent: `could not activate storage 'felhom-pbs': storage 'felhom-pbs' does not exist`; `pvesm status` lists only `local` and `local-lvm`. At 19:15:36Z `/backups` read: „✗ · **Utolsó teljes mentés 2026-09-14 21:09 (5 perce) · 0 B · Biztonsági szerver – külön hardver (PBS) · Naprakész**" and „✓ **Távoli rendszermentés — külön hardveren (PBS)**" (`phase4/backup-pages-after.txt`). The successful 8.9 GB local backup is no longer shown; a 0-byte failed attempt is labelled up to date; the remote tier is ticked as present. The CLAUDE.md rule „presence is not success" in page form: an attempt's timestamp stands in for a result. The hub did raise a true `whole_guest_backup_failed (error)` to the operator; the customer's page says the opposite. **Fix shape:** the tile shows the newest SUCCESSFUL backup per tier, a failed tier as failed, and a tier whose storage does not exist as „nincs beállítva". **Also measured after F2's reboot (19:45:41Z):** the whole-system tile read „– · Utolsó teljes mentés – · Méret / cél · **Naprakész**" — no backup listed at all, still labelled up to date (the 21:03 CEST local success no longer shown). | **READY — rank P1-HIGH; owner: CC (controller)** | | **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** | -| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** | +| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |