diff --git a/documentation/audits/BIGNIGHT-household-month-2026-09-14.md b/documentation/audits/BIGNIGHT-household-month-2026-09-14.md index 107ae384..0a1cde73 100644 --- a/documentation/audits/BIGNIGHT-household-month-2026-09-14.md +++ b/documentation/audits/BIGNIGHT-household-month-2026-09-14.md @@ -72,3 +72,8 @@ reached the box 14 m 25 s after the push, „Frissítés elérhető — ma", **D | F6 | drive unplugged during a backup | nothing about skipped apps | skipped 4 volume dumps but `success:true`; torn `.tmp` not promoted; next run honest and complete | 122 s after re-attach | `storage_disconnected` + 4, **all mails suppressed by cooldown** | the second drive loss (R-521); the incomplete run (R-519) | | F7 | system disk to 95 % | dashboard „90 % · Kritikusan kevés hely"; banner **English** „SSD disk usage high: 90%" | stayed reachable; uploads and a backup still succeeded; warnings cleared 3½ min after cleanup | — | `health_degraded` / true, **mail suppressed by cooldown**; no disk event | operator told nothing (R-521) | | F8 | internet gone 17½ min (venue: hub on the LAN, so the hub's ingress was blocked too) | LAN dashboard 200 throughout, **no offline notice, tunnel tile „Fut"** (R-522) | report push failed and backed off; on return pushed at once, tunnel re-registered +9 s (all 4 +58 s) | ≈ 1 min | `node_stale` + mail at 31 min since last report, `node_recovered` + mail / true | customer notice (R-522) | +| F9 | `docker kill` the controller 4 s into a deploy | dashboard **502 for 33 min**, no page at all | **nothing** restarted it (`unless-stopped` ignores a kill; bootstrap is one-shot; agent silent); the deployed app itself came up healthy | — until power-cycle | `node_stale` at 30 min, **mail suppressed by cooldown** | everything: **R-523 (P1)** | +| F10–F12 | photo folder deleted · forgotten credentials · two quick reboots | **not run** | — | — | — | **the brief's stop rule was met at F9** | + +**Stop rule, applied as written (22:08Z):** F9 left the box in a state the product did not leave on its own and no +customer screen can reach. Filed P1 (R-523), no further faults injected, the box recovered by a power-cycle. diff --git a/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md b/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md index 97e2752b..62638b6c 100644 --- a/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md +++ b/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md @@ -30,6 +30,8 @@ Source: hub log (`alarms.sh`, CEST times), the operator mailbox and the customer | 22:54:46 | `health_degraded (warning)` | F7 system disk 95 % | **suppressed — cooldown** | true, not delivered; no disk-specific event exists (R-521) | | 23:25:43 | `node_stale` (staleness ok → stale) | F8 internet cut (31 min after the last report) | operator | true | | 23:26:43 | `node_recovered` | F8 internet back | operator | true | +| 23:34:35 | `app_deployed (info)` Homebox | F9 deploy (sent before the kill) | — | true | +| 00:04:43 (09-15) | `node_stale` | F9 controller dead 30 min | **suppressed — cooldown** (F8's key) | true, **not delivered** (R-523, R-521) | Hub WARN lines without an event or mail: `host tester-1-a61396 backup FAILED: target=felhom-pbs … does not exist` at 21:12:43 and 21:27:43 (the agent's own PBS attempts). @@ -48,3 +50,7 @@ Hub WARN lines without an event or mail: `host tester-1-a61396 backup FAILED: ta | F6 drive lost during backup | drive lost; backup incomplete | `storage_disconnected` logged, **mail suppressed by cooldown**; the run reported `success:true` | **MISSED** on both counts (R-521, R-519) | | F7 system disk 95 % | disk nearly full | `health_degraded` only, mail suppressed; customer banner in English | **MISSED** for the operator (R-521); customer told, in English (R-516) | | F8 internet gone 17½ min | box offline | `node_stale` + mail, then `node_recovered` + mail | correct for the operator; the customer's dashboard says „Fut" (R-522) | +| F9 controller killed mid-deploy | the dashboard is down / the controller is not running | nothing for 30 min, then `node_stale` with its mail **suppressed** | **MISSED** — the stop-rule finding (R-523) | +| F10 child deletes the photo folder | — | **not run: stop rule met at F9** | — | +| F11 forgotten credentials | — | **not run: stop rule met at F9** | — | +| F12 two reboots in two minutes | — | **not run: stop rule met at F9** | — | diff --git a/documentation/audits/evidence-bignight-2026-09-14/journal.md b/documentation/audits/evidence-bignight-2026-09-14/journal.md index 48c74d3b..2ff3e5a8 100644 --- a/documentation/audits/evidence-bignight-2026-09-14/journal.md +++ b/documentation/audits/evidence-bignight-2026-09-14/journal.md @@ -643,3 +643,10 @@ itself, through the product's restore endpoint, and reads the data back — reco customer cannot recover from any screen (there is no dashboard). Filed **R-523 (P1)** before acting. Per the brief: **no further faults are injected — F10, F11 and F12 are not run**; the box is recovered by the household's only lever, a power-cycle, and the night moves to Phase 6. + +**F9 recovery — the household's only lever, a power-cycle** (`phase5/F9/recovery-power-cycle.txt`, +`controller-after-reboot.log`): `qm reset 333` 22:08:17Z → the in-guest bootstrap started the controller at boot +(22:10:23Z, `restarts=0`); dashboard 200 at **+131 s**; all 12 apps `running` at **+229 s** (paperless again started by +the boot reconciler); **homebox `running`, `deploying false`, `deployed true`, „Fut · Naprakész" — the interrupted deploy +is not stuck**. Hub: `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z — **mail suppressed by cooldown**. +So a reboot heals it; nothing short of a reboot does, and no screen tells a household to reboot. diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/after-33-min-dead.txt b/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/after-33-min-dead.txt index c0eb63cf..88e6931c 100644 --- a/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/after-33-min-dead.txt +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/after-33-min-dead.txt @@ -2,3 +2,8 @@ 2026/09/14 23:34:35 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Homebox 2026/09/15 00:04:43 [INFO] Staleness: tester-1 ok → stale (node_stale) 2026/09/15 00:04:43 [INFO] Operator email suppressed for tester-1/node_stale — cooldown (key=tester-1:node_stale) +2026/09/15 00:04:43 [INFO] Staleness: tester-1 ok → stale (node_stale) +2026/09/15 00:04:43 [INFO] Operator email suppressed for tester-1/node_stale — cooldown (key=tester-1:node_stale) +2026/09/15 00:10:32 [INFO] Event from tester-1: controller_started (info) — Controller elindult (0.242.0) +2026/09/15 00:10:43 [INFO] Staleness: tester-1 stale → ok (node_recovered) +2026/09/15 00:10:43 [INFO] Operator email suppressed for tester-1/node_recovered — cooldown (key=tester-1:node_recovered) diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/controller-after-reboot.log b/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/controller-after-reboot.log new file mode 100644 index 00000000..ff473a77 --- /dev/null +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/controller-after-reboot.log @@ -0,0 +1,200 @@ +2026/09/14 22:10:24 [INFO] local-api: channel up (agent 169.254.253.1:8443) — guest 9201, 3 mount(s) visible +2026/09/14 22:10:24 [INFO] local-api: mount mp8 → /mnt/felhom-drives (storage=/mnt/felhom-drives, class=, backup=false) +2026/09/14 22:10:24 [INFO] local-api: mount mp9 → /etc/felhom-bootstrap (storage=/var/lib/felhom-agent/guests/9201/bootstrap, class=, backup=false) +2026/09/14 22:10:24 [INFO] local-api: mount mp0 → /var/lib/felhom (storage=local-lvm, class=, backup=true) +2026/09/14 22:10:24 [INFO] felhom-controller 0.242.0 starting (customer: tester-1, domain: enkicsifelhom.hu) +2026/09/14 22:10:24 [INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json +2026/09/14 22:10:24 [INFO] Encryption key loaded from /opt/docker/felhom-controller/data/encryption.key +2026/09/14 22:10:25 [INFO] [stacks] Using compose command: docker compose +2026/09/14 22:10:25 [INFO] [stacks] ScanStacks complete: 57 stacks found (13 deployed, 44 available) +2026/09/14 22:10:25 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:25 [INFO] [stacks] InjectMissingFields: processed 13 stacks +2026/09/14 22:10:25 [INFO] [stacks] Encryption migration: no stacks needed migration +2026/09/14 22:10:25 [INFO] [stacks] desired-state backfill: 0 app(s) recorded as running, 0 left unrecorded (state ambiguous — legacy boot behaviour retained) +2026/09/14 22:10:25 [INFO] [stacks] installed-images backfill: 0 app(s) recorded, 13 already had a record, 0 left unrecorded (could not be observed completely — unknown, which renders nothing) +2026/09/14 22:10:25 [INFO] [stacks] pin adoption: 0 pinned, 13 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers) +2026/09/14 22:10:25 [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s) +2026/09/14 22:10:25 [INFO] [quiesce] loop started (poll 5m0s, max-quiesce 30m0s) +2026/09/14 22:10:25 [INFO] [sync] Starting catalog sync +2026/09/14 22:10:25 [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main) +2026/09/14 22:10:25 [INFO] Metrics store opened at /opt/docker/felhom-controller/data/metrics.db +2026/09/14 22:10:25 [INFO] Metrics collector started (60s interval) +2026/09/14 22:10:25 [INFO] Notifier enabled (hub: https://hub.felhom.eu) +2026/09/14 22:10:25 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: status-refresh (every 10s) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: health-probes (every 10s) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: system-health (every 5m0s) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: ring-spill (every 30s) +2026/09/14 22:10:25 [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-15 02:30 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s) +2026/09/14 22:10:25 [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-15 03:30 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-15 04:15 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-15 05:10 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-15 06:00 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-15 05:30 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-15 04:00 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-15 03:30 CEST +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: hub-report (every 15m0s) +2026/09/14 22:10:25 [INFO] Hub reporting enabled (every 15m0s to https://hub.felhom.eu) +2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s) +2026/09/14 22:10:25 [INFO] ========== Startup Self-Test ========== +2026/09/14 22:10:26 [INFO] [PASS] Docker socket: reachable (v29.8.0) +2026/09/14 22:10:26 [INFO] [PASS] Stacks directory: /opt/docker/stacks +2026/09/14 22:10:26 [INFO] [PASS] Data directory: /opt/docker/felhom-controller/data (writable) +2026/09/14 22:10:26 [INFO] [PASS] System data path: /mnt/sys_drive +2026/09/14 22:10:26 [INFO] [PASS] Storage paths: 1 connected, 0 disconnected +2026/09/14 22:10:26 [INFO] [PASS] Git catalog: 53 app definitions found +2026/09/14 22:10:26 [INFO] [infra] connected felhom-controller to traefik-public +2026/09/14 22:10:27 [INFO] [PASS] Hub connectivity: https://hub.felhom.eu reachable (HTTP 200) +2026/09/14 22:10:27 [INFO] [PASS] Metrics DB: 0.5 MB +2026/09/14 22:10:27 [INFO] ======================================== +2026/09/14 22:10:27 [INFO] Self-test complete: 8 passed, 0 warnings, 0 failed +2026/09/14 22:10:27 [INFO] [scheduler] Starting scheduler with 18 jobs +2026/09/14 22:10:27 [INFO] [scheduler] Registered periodic job: geo-verify (every 6h0m0s) +2026/09/14 22:10:27 [INFO] Geo-restriction support enabled (CF API token configured) +2026/09/14 22:10:27 [INFO] [report] hub wait channel active (hold ≤240s) +2026/09/14 22:10:27 [INFO] [web] Auth: using password from settings.json +2026/09/14 22:10:27 [INFO] [scheduler] Registered periodic job: disk-health-check (every 1h0m0s) +2026/09/14 22:10:27 [INFO] [scheduler] Registered periodic job: agent-channel-health (every 1m0s) +2026/09/14 22:10:27 [INFO] Web UI listening on :8080 +2026/09/14 22:10:27 [INFO] [backup] Found 6 DB dump files across drives +2026/09/14 22:10:27 [INFO] [gate] boot 1789423711-10047: waiting (≤2m0s) for live drive bind(s) [/mnt/felhom-drives/hdd_1] before recreating drive-backed apps +2026/09/14 22:10:27 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app immich (state=starting) onto /mnt/felhom-drives/hdd_1 +2026/09/14 22:10:27 [INFO] [stacks] Stopping stack: immich +2026/09/14 22:10:28 [INFO] [sync] Catalog sync complete +2026/09/14 22:10:28 [INFO] [sync] Initial sync: Sablonok naprakészek — nincs változás +2026/09/14 22:10:29 [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s)) +2026/09/14 22:10:29 [INFO] [monitor] Health check: status=ok +2026/09/14 22:10:29 [INFO] [backup] Discovered 6 databases +2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bookstack/docker-compose.yml +2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/gokapi/docker-compose.yml +2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/homebox/docker-compose.yml +2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/immich/docker-compose.yml +2026/09/14 22:10:29 [INFO] [web] Login from 172.18.0.16:51988 +2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/mealie/docker-compose.yml +2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 2 HDD mounts for /opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml +2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/14 22:10:30 [INFO] [backup] Discovered app data: 13 apps +2026/09/14 22:10:30 [INFO] [backup] Recovery unit captured for homebox → /mnt/sys_drive/felhom-data/backups/primary/homebox (images=1, secrets-referenced=1, data_keys=1, portable-carried=1/1, withheld=0) +2026/09/14 22:10:30 [INFO] [backup] Backup status cache refreshed +2026/09/14 22:10:30 [INFO] [stacks] Stack immich stopped successfully (took 2.5s) +2026/09/14 22:10:30 [INFO] [stacks] Status refresh: 25 containers across 57 stacks +2026/09/14 22:10:30 [INFO] [stacks] Starting stack: immich +2026/09/14 22:10:30 [INFO] [stacks] Status refresh: 25 containers across 57 stacks +2026/09/14 22:10:32 [INFO] [report] Building system report +2026/09/14 22:10:32 [INFO] Event pushed: controller_started (info) — Controller elindult (0.242.0) +2026/09/14 22:10:33 [INFO] [monitor] Health check: status=ok +2026/09/14 22:10:34 [INFO] [report] Hub report pushed successfully (11575 bytes) +2026/09/14 22:10:34 [INFO] [settings] Settings saved +2026/09/14 22:10:34 [INFO] Startup hub report sent +2026/09/14 22:10:35 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:37 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:37 [WARN] Health probe jellyfin: API GET :8096/health → Get "http://jellyfin:8096/health": dial tcp: lookup jellyfin on 192.168.0.250:53: no such host +2026/09/14 22:10:37 [WARN] Health probe nextcloud: API GET :80/status.php → Get "http://nextcloud:80/status.php": dial tcp: lookup nextcloud on 192.168.0.250:53: no such host +2026/09/14 22:10:37 [WARN] Health probe immich: API GET :2283/api/server/ping → Get "http://immich-postgres:2283/api/server/ping": dial tcp: lookup immich-postgres on 192.168.0.250:53: no such host +2026/09/14 22:10:37 [WARN] Health probe homebox: API GET :7745/api/v1/status → Get "http://homebox:7745/api/v1/status": dial tcp: lookup homebox on 192.168.0.250:53: no such host +2026/09/14 22:10:37 [WARN] Health probe gokapi: HTTP GET :53842/ → Get "http://gokapi:53842/": dial tcp: lookup gokapi on 192.168.0.250:53: no such host +2026/09/14 22:10:37 [WARN] Health probe vaultwarden: API GET :80/alive → Get "http://vaultwarden:80/alive": dial tcp: lookup vaultwarden on 192.168.0.250:53: no such host +2026/09/14 22:10:37 [WARN] Health probes: 1 ok, 6 unhealthy (of 7 probed) +2026/09/14 22:10:40 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:43 [INFO] [stacks] Stack immich started successfully (took 13.5s) +2026/09/14 22:10:46 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:46 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:46 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app jellyfin (state=starting) onto /mnt/felhom-drives/hdd_1 +2026/09/14 22:10:46 [INFO] [stacks] Stopping stack: jellyfin +2026/09/14 22:10:47 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:50 [INFO] [stacks] Stack immich post-start status: +2026/09/14 22:10:50 [INFO] [stacks] immich-machine-learning ghcr.io/immich-app/immich-machine-learning:v3.0.3 running Up 18 seconds (health: starting) +2026/09/14 22:10:50 [INFO] [stacks] immich-postgres ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 running Up 18 seconds (healthy) +2026/09/14 22:10:50 [INFO] [stacks] immich-redis redis:7-alpine running Up 18 seconds (healthy) +2026/09/14 22:10:50 [INFO] [stacks] immich-server ghcr.io/immich-app/immich-server:v3.0.3 running Up 7 seconds (health: starting) +2026/09/14 22:10:50 [WARN] Health probe jellyfin: API GET :8096/health → expected status 200, got 503 +2026/09/14 22:10:50 [WARN] Health probes: 5 ok, 1 unhealthy (of 6 probed) +2026/09/14 22:10:51 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:54 [INFO] [stacks] Stack jellyfin stopped successfully (took 7.9s) +2026/09/14 22:10:54 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 22:10:54 [INFO] [stacks] Starting stack: jellyfin +2026/09/14 22:10:55 [INFO] [stacks] Stack jellyfin started successfully (took 0.6s) +2026/09/14 22:10:55 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:55 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app nextcloud (state=starting) onto /mnt/felhom-drives/hdd_1 +2026/09/14 22:10:55 [INFO] [stacks] Stopping stack: nextcloud +2026/09/14 22:10:56 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:10:57 [INFO] [stacks] Status refresh: 28 containers across 57 stacks +2026/09/14 22:10:57 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 22:10:57 [INFO] [stacks] Stack nextcloud stopped successfully (took 2.2s) +2026/09/14 22:10:57 [INFO] [stacks] Status refresh: 26 containers across 57 stacks +2026/09/14 22:10:57 [INFO] [stacks] Starting stack: nextcloud +2026/09/14 22:10:58 [INFO] [stacks] Stack jellyfin post-start status: +2026/09/14 22:10:58 [INFO] [stacks] jellyfin jellyfin/jellyfin:10.11.11 running Up 4 seconds (health: starting) +2026/09/14 22:10:59 [INFO] [selfupdate] Current version 0.242.0 is up to date +2026/09/14 22:11:01 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:06 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:07 [WARN] Health probe jellyfin: API GET :8096/health → Get "http://jellyfin:8096/health": dial tcp 172.18.0.4:8096: connect: connection refused +2026/09/14 22:11:07 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:07 [WARN] Health probes: 0 ok, 1 unhealthy (of 1 probed) +2026/09/14 22:11:09 [INFO] [stacks] Stack nextcloud started successfully (took 12.0s) +2026/09/14 22:11:10 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:10 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app paperless-ngx (state=starting) onto /mnt/felhom-drives/hdd_1 +2026/09/14 22:11:10 [INFO] [stacks] Stopping stack: paperless-ngx +2026/09/14 22:11:12 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:13 [INFO] [stacks] Stack nextcloud post-start status: +2026/09/14 22:11:13 [INFO] [stacks] nextcloud nextcloud:34.0.1-apache running Up 4 seconds (health: starting) +2026/09/14 22:11:13 [INFO] [stacks] nextcloud-db mariadb:11.6 running Up 15 seconds (healthy) +2026/09/14 22:11:13 [INFO] [stacks] nextcloud-redis redis:7-alpine running Up 15 seconds (healthy) +2026/09/14 22:11:17 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:17 [INFO] Health probes: 1 ok (of 1 probed) +2026/09/14 22:11:17 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:21 [INFO] [stacks] Stack paperless-ngx stopped successfully (took 11.3s) +2026/09/14 22:11:21 [INFO] [stacks] Status refresh: 26 containers across 57 stacks +2026/09/14 22:11:21 [INFO] [stacks] Starting stack: paperless-ngx +2026/09/14 22:11:22 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:22 [INFO] [bootrecon] boot window: budget 50s exhausted while the fleet was still changing — sweeping anyway +2026/09/14 22:11:22 [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [paperless-ngx] — up to 2 attempt(s) +2026/09/14 22:11:22 [INFO] [stacks] Starting stack: paperless-ngx +2026/09/14 22:11:27 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:27 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 22:11:27 [INFO] Health probes: 5 ok (of 5 probed) +2026/09/14 22:11:27 [INFO] [scheduler] Job agent-channel-health completed (took 185ms) +2026/09/14 22:11:32 [INFO] [stacks] Stack paperless-ngx started successfully (took 9.9s) +2026/09/14 22:11:32 [INFO] [stacks] Stack paperless-ngx started successfully (took 11.4s) +2026/09/14 22:11:32 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:32 [INFO] [bootrecon] Boot reconciliation attempt 1/2: started "paperless-ngx" (took 10.1s) +2026/09/14 22:11:32 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:32 [INFO] [bootrecon] Boot reconciliation complete: 1 app(s) recovered in 1 attempt(s): [paperless-ngx] +2026/09/14 22:11:33 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:33 [INFO] [gate] boot 1789423711-10047: re-syncing FileBrowser mounts against the live binds +2026/09/14 22:11:33 [INFO] [settings] Settings saved +2026/09/14 22:11:33 [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s)) +2026/09/14 22:11:35 [INFO] [stacks] Stack paperless-ngx post-start status: +2026/09/14 22:11:35 [INFO] [stacks] paperless-postgres postgres:16-alpine running Up 14 seconds (healthy) +2026/09/14 22:11:35 [INFO] [stacks] paperless-redis redis:7-alpine running Up 14 seconds (healthy) +2026/09/14 22:11:35 [INFO] [stacks] paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 running Up 3 seconds (health: starting) +2026/09/14 22:11:36 [INFO] [stacks] Stack paperless-ngx post-start status: +2026/09/14 22:11:36 [INFO] [stacks] paperless-postgres postgres:16-alpine running Up 14 seconds (healthy) +2026/09/14 22:11:36 [INFO] [stacks] paperless-redis redis:7-alpine running Up 14 seconds (healthy) +2026/09/14 22:11:36 [INFO] [stacks] paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 running Up 3 seconds (health: starting) +2026/09/14 22:11:37 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:47 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:11:55 [INFO] [fillwatch] checked 3 filesystem(s), 0 unreadable/skipped, 0 notification(s); bands: all ok +2026/09/14 22:11:57 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:12:07 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:12:07 [INFO] [web] Login from 172.18.0.16:51988 +2026/09/14 22:12:17 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:12:27 [INFO] [scheduler] Running job: stack-scan +2026/09/14 22:12:27 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:12:27 [INFO] [stacks] ScanStacks complete: 57 stacks found (13 deployed, 44 available) +2026/09/14 22:12:27 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:12:27 [INFO] [scheduler] Job stack-scan completed (took 84ms) +2026/09/14 22:12:27 [INFO] [scheduler] Running job: agent-channel-health +2026/09/14 22:12:27 [INFO] [scheduler] Job agent-channel-health completed (took 187ms) +2026/09/14 22:12:37 [INFO] [stacks] Status refresh: 29 containers across 57 stacks +2026/09/14 22:12:47 [INFO] [stacks] Status refresh: 29 containers across 57 stacks diff --git a/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/recovery-power-cycle.txt b/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/recovery-power-cycle.txt index 91d5fd45..7b155b7d 100644 --- a/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/recovery-power-cycle.txt +++ b/documentation/audits/evidence-bignight-2026-09-14/phase5/F9/recovery-power-cycle.txt @@ -1,2 +1,26 @@ F9 power-cycle (qm reset) at 22:08:17 22:08:23 +5s health=000 +22:08:43 +25s health=000 +22:09:01 +43s health=000 +22:09:19 +61s health=000 +22:09:37 +79s health=000 +22:09:56 +98s health=000 +22:10:14 +116s health=000 +22:10:29 +131s health=200 +F9 + 131s dashboard /api/health 200 +F9 + 132s privatebin running +F9 + 132s gokapi running +F9 + 138s nextcloud running +F9 + 138s vaultwarden running +F9 + 138s jellyfin running +F9 + 149s bookstack running +F9 + 156s uptime-kuma running +F9 + 178s mealie running +F9 + 188s docmost running +F9 + 194s immich running +F9 + 194s adventurelog running +F9 + 229s paperless-ngx running +F9 + 229s ALL 12 running +--- homebox after reboot +state running deploying False deployed True {'homebox': 'ghcr.io/sysadminsmedia/homebox:0.26.2'} +Fut Naprakész diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 3827a740..1a6873f2 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -726,7 +726,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** | | **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** | | **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** | -| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle: see the F9 journal (measured after filing). `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** | +| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle (`qm reset` 22:08:17Z): the bootstrap started the controller at boot (22:10:23Z), dashboard 200 at +131 s, all 12 apps running at +229 s, the deployed homebox `running · deploying false · deployed true` — „Fut · Naprakész", **not stuck**. Hub `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z with its mail **also suppressed by cooldown**. `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** |