BIGNIGHT F9 recovery: power-cycle heals, deploy not stuck, recovered mail suppressed; R-523 amended; table + audit
gates / gates (push) Successful in 20s

This commit is contained in:
2026-09-15 00:13:11 +02:00
parent adc38d1f72
commit c664fa0326
7 changed files with 248 additions and 1 deletions
@@ -72,3 +72,8 @@ reached the box 14 m 25 s after the push, „Frissítés elérhető — ma", **D
| F6 | drive unplugged during a backup | nothing about skipped apps | skipped 4 volume dumps but `success:true`; torn `.tmp` not promoted; next run honest and complete | 122 s after re-attach | `storage_disconnected` + 4, **all mails suppressed by cooldown** | the second drive loss (R-521); the incomplete run (R-519) |
| F7 | system disk to 95 % | dashboard „90 % · Kritikusan kevés hely"; banner **English** „SSD disk usage high: 90%" | stayed reachable; uploads and a backup still succeeded; warnings cleared 3½ min after cleanup | — | `health_degraded` / true, **mail suppressed by cooldown**; no disk event | operator told nothing (R-521) |
| F8 | internet gone 17½ min (venue: hub on the LAN, so the hub's ingress was blocked too) | LAN dashboard 200 throughout, **no offline notice, tunnel tile „Fut"** (R-522) | report push failed and backed off; on return pushed at once, tunnel re-registered +9 s (all 4 +58 s) | ≈ 1 min | `node_stale` + mail at 31 min since last report, `node_recovered` + mail / true | customer notice (R-522) |
| F9 | `docker kill` the controller 4 s into a deploy | dashboard **502 for 33 min**, no page at all | **nothing** restarted it (`unless-stopped` ignores a kill; bootstrap is one-shot; agent silent); the deployed app itself came up healthy | — until power-cycle | `node_stale` at 30 min, **mail suppressed by cooldown** | everything: **R-523 (P1)** |
| F10–F12 | photo folder deleted · forgotten credentials · two quick reboots | **not run** | — | — | — | **the brief's stop rule was met at F9** |
**Stop rule, applied as written (22:08Z):** F9 left the box in a state the product did not leave on its own and no
customer screen can reach. Filed P1 (R-523), no further faults injected, the box recovered by a power-cycle.
@@ -30,6 +30,8 @@ Source: hub log (`alarms.sh`, CEST times), the operator mailbox and the customer
| 22:54:46 | `health_degraded (warning)` | F7 system disk 95 % | **suppressed — cooldown** | true, not delivered; no disk-specific event exists (R-521) |
| 23:25:43 | `node_stale` (staleness ok → stale) | F8 internet cut (31 min after the last report) | operator | true |
| 23:26:43 | `node_recovered` | F8 internet back | operator | true |
| 23:34:35 | `app_deployed (info)` Homebox | F9 deploy (sent before the kill) | — | true |
| 00:04:43 (09-15) | `node_stale` | F9 controller dead 30 min | **suppressed — cooldown** (F8's key) | true, **not delivered** (R-523, R-521) |
Hub WARN lines without an event or mail: `host tester-1-a61396 backup FAILED: target=felhom-pbs … does not exist` at
21:12:43 and 21:27:43 (the agent's own PBS attempts).
@@ -48,3 +50,7 @@ Hub WARN lines without an event or mail: `host tester-1-a61396 backup FAILED: ta
| F6 drive lost during backup | drive lost; backup incomplete | `storage_disconnected` logged, **mail suppressed by cooldown**; the run reported `success:true` | **MISSED** on both counts (R-521, R-519) |
| F7 system disk 95 % | disk nearly full | `health_degraded` only, mail suppressed; customer banner in English | **MISSED** for the operator (R-521); customer told, in English (R-516) |
| F8 internet gone 17½ min | box offline | `node_stale` + mail, then `node_recovered` + mail | correct for the operator; the customer's dashboard says „Fut" (R-522) |
| F9 controller killed mid-deploy | the dashboard is down / the controller is not running | nothing for 30 min, then `node_stale` with its mail **suppressed** | **MISSED** — the stop-rule finding (R-523) |
| F10 child deletes the photo folder | — | **not run: stop rule met at F9** | — |
| F11 forgotten credentials | — | **not run: stop rule met at F9** | — |
| F12 two reboots in two minutes | — | **not run: stop rule met at F9** | — |
@@ -643,3 +643,10 @@ itself, through the product's restore endpoint, and reads the data back — reco
customer cannot recover from any screen (there is no dashboard). Filed **R-523 (P1)** before acting. Per the brief:
**no further faults are injected — F10, F11 and F12 are not run**; the box is recovered by the household's only lever,
a power-cycle, and the night moves to Phase 6.
**F9 recovery — the household's only lever, a power-cycle** (`phase5/F9/recovery-power-cycle.txt`,
`controller-after-reboot.log`): `qm reset 333` 22:08:17Z → the in-guest bootstrap started the controller at boot
(22:10:23Z, `restarts=0`); dashboard 200 at **+131 s**; all 12 apps `running` at **+229 s** (paperless again started by
the boot reconciler); **homebox `running`, `deploying false`, `deployed true`, „Fut · Naprakész" — the interrupted deploy
is not stuck**. Hub: `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z — **mail suppressed by cooldown**.
So a reboot heals it; nothing short of a reboot does, and no screen tells a household to reboot.
@@ -2,3 +2,8 @@
2026/09/14 23:34:35 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Homebox
2026/09/15 00:04:43 [INFO] Staleness: tester-1 ok → stale (node_stale)
2026/09/15 00:04:43 [INFO] Operator email suppressed for tester-1/node_stale — cooldown (key=tester-1:node_stale)
2026/09/15 00:04:43 [INFO] Staleness: tester-1 ok → stale (node_stale)
2026/09/15 00:04:43 [INFO] Operator email suppressed for tester-1/node_stale — cooldown (key=tester-1:node_stale)
2026/09/15 00:10:32 [INFO] Event from tester-1: controller_started (info) — Controller elindult (0.242.0)
2026/09/15 00:10:43 [INFO] Staleness: tester-1 stale → ok (node_recovered)
2026/09/15 00:10:43 [INFO] Operator email suppressed for tester-1/node_recovered — cooldown (key=tester-1:node_recovered)
@@ -0,0 +1,200 @@
2026/09/14 22:10:24 [INFO] local-api: channel up (agent 169.254.253.1:8443) — guest 9201, 3 mount(s) visible
2026/09/14 22:10:24 [INFO] local-api: mount mp8 → /mnt/felhom-drives (storage=/mnt/felhom-drives, class=, backup=false)
2026/09/14 22:10:24 [INFO] local-api: mount mp9 → /etc/felhom-bootstrap (storage=/var/lib/felhom-agent/guests/9201/bootstrap, class=, backup=false)
2026/09/14 22:10:24 [INFO] local-api: mount mp0 → /var/lib/felhom (storage=local-lvm, class=, backup=true)
2026/09/14 22:10:24 [INFO] felhom-controller 0.242.0 starting (customer: tester-1, domain: enkicsifelhom.hu)
2026/09/14 22:10:24 [INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json
2026/09/14 22:10:24 [INFO] Encryption key loaded from /opt/docker/felhom-controller/data/encryption.key
2026/09/14 22:10:25 [INFO] [stacks] Using compose command: docker compose
2026/09/14 22:10:25 [INFO] [stacks] ScanStacks complete: 57 stacks found (13 deployed, 44 available)
2026/09/14 22:10:25 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:25 [INFO] [stacks] InjectMissingFields: processed 13 stacks
2026/09/14 22:10:25 [INFO] [stacks] Encryption migration: no stacks needed migration
2026/09/14 22:10:25 [INFO] [stacks] desired-state backfill: 0 app(s) recorded as running, 0 left unrecorded (state ambiguous — legacy boot behaviour retained)
2026/09/14 22:10:25 [INFO] [stacks] installed-images backfill: 0 app(s) recorded, 13 already had a record, 0 left unrecorded (could not be observed completely — unknown, which renders nothing)
2026/09/14 22:10:25 [INFO] [stacks] pin adoption: 0 pinned, 13 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers)
2026/09/14 22:10:25 [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
2026/09/14 22:10:25 [INFO] [quiesce] loop started (poll 5m0s, max-quiesce 30m0s)
2026/09/14 22:10:25 [INFO] [sync] Starting catalog sync
2026/09/14 22:10:25 [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main)
2026/09/14 22:10:25 [INFO] Metrics store opened at /opt/docker/felhom-controller/data/metrics.db
2026/09/14 22:10:25 [INFO] Metrics collector started (60s interval)
2026/09/14 22:10:25 [INFO] Notifier enabled (hub: https://hub.felhom.eu)
2026/09/14 22:10:25 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: status-refresh (every 10s)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: health-probes (every 10s)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: system-health (every 5m0s)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: ring-spill (every 30s)
2026/09/14 22:10:25 [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-15 02:30 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s)
2026/09/14 22:10:25 [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-15 03:30 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-15 04:15 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-15 05:10 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-15 06:00 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-15 05:30 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-15 04:00 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-15 03:30 CEST
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: hub-report (every 15m0s)
2026/09/14 22:10:25 [INFO] Hub reporting enabled (every 15m0s to https://hub.felhom.eu)
2026/09/14 22:10:25 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/14 22:10:25 [INFO] ========== Startup Self-Test ==========
2026/09/14 22:10:26 [INFO] [PASS] Docker socket: reachable (v29.8.0)
2026/09/14 22:10:26 [INFO] [PASS] Stacks directory: /opt/docker/stacks
2026/09/14 22:10:26 [INFO] [PASS] Data directory: /opt/docker/felhom-controller/data (writable)
2026/09/14 22:10:26 [INFO] [PASS] System data path: /mnt/sys_drive
2026/09/14 22:10:26 [INFO] [PASS] Storage paths: 1 connected, 0 disconnected
2026/09/14 22:10:26 [INFO] [PASS] Git catalog: 53 app definitions found
2026/09/14 22:10:26 [INFO] [infra] connected felhom-controller to traefik-public
2026/09/14 22:10:27 [INFO] [PASS] Hub connectivity: https://hub.felhom.eu reachable (HTTP 200)
2026/09/14 22:10:27 [INFO] [PASS] Metrics DB: 0.5 MB
2026/09/14 22:10:27 [INFO] ========================================
2026/09/14 22:10:27 [INFO] Self-test complete: 8 passed, 0 warnings, 0 failed
2026/09/14 22:10:27 [INFO] [scheduler] Starting scheduler with 18 jobs
2026/09/14 22:10:27 [INFO] [scheduler] Registered periodic job: geo-verify (every 6h0m0s)
2026/09/14 22:10:27 [INFO] Geo-restriction support enabled (CF API token configured)
2026/09/14 22:10:27 [INFO] [report] hub wait channel active (hold ≤240s)
2026/09/14 22:10:27 [INFO] [web] Auth: using password from settings.json
2026/09/14 22:10:27 [INFO] [scheduler] Registered periodic job: disk-health-check (every 1h0m0s)
2026/09/14 22:10:27 [INFO] [scheduler] Registered periodic job: agent-channel-health (every 1m0s)
2026/09/14 22:10:27 [INFO] Web UI listening on :8080
2026/09/14 22:10:27 [INFO] [backup] Found 6 DB dump files across drives
2026/09/14 22:10:27 [INFO] [gate] boot 1789423711-10047: waiting (≤2m0s) for live drive bind(s) [/mnt/felhom-drives/hdd_1] before recreating drive-backed apps
2026/09/14 22:10:27 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app immich (state=starting) onto /mnt/felhom-drives/hdd_1
2026/09/14 22:10:27 [INFO] [stacks] Stopping stack: immich
2026/09/14 22:10:28 [INFO] [sync] Catalog sync complete
2026/09/14 22:10:28 [INFO] [sync] Initial sync: Sablonok naprakészek — nincs változás
2026/09/14 22:10:29 [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s))
2026/09/14 22:10:29 [INFO] [monitor] Health check: status=ok
2026/09/14 22:10:29 [INFO] [backup] Discovered 6 databases
2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/adventurelog/docker-compose.yml
2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bookstack/docker-compose.yml
2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/gokapi/docker-compose.yml
2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/homebox/docker-compose.yml
2026/09/14 22:10:29 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/immich/docker-compose.yml
2026/09/14 22:10:29 [INFO] [web] Login from 172.18.0.16:51988
2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/jellyfin/docker-compose.yml
2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/mealie/docker-compose.yml
2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 2 HDD mounts for /opt/docker/stacks/paperless-ngx/docker-compose.yml
2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml
2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml
2026/09/14 22:10:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/vaultwarden/docker-compose.yml
2026/09/14 22:10:30 [INFO] [backup] Discovered app data: 13 apps
2026/09/14 22:10:30 [INFO] [backup] Recovery unit captured for homebox → /mnt/sys_drive/felhom-data/backups/primary/homebox (images=1, secrets-referenced=1, data_keys=1, portable-carried=1/1, withheld=0)
2026/09/14 22:10:30 [INFO] [backup] Backup status cache refreshed
2026/09/14 22:10:30 [INFO] [stacks] Stack immich stopped successfully (took 2.5s)
2026/09/14 22:10:30 [INFO] [stacks] Status refresh: 25 containers across 57 stacks
2026/09/14 22:10:30 [INFO] [stacks] Starting stack: immich
2026/09/14 22:10:30 [INFO] [stacks] Status refresh: 25 containers across 57 stacks
2026/09/14 22:10:32 [INFO] [report] Building system report
2026/09/14 22:10:32 [INFO] Event pushed: controller_started (info) — Controller elindult (0.242.0)
2026/09/14 22:10:33 [INFO] [monitor] Health check: status=ok
2026/09/14 22:10:34 [INFO] [report] Hub report pushed successfully (11575 bytes)
2026/09/14 22:10:34 [INFO] [settings] Settings saved
2026/09/14 22:10:34 [INFO] Startup hub report sent
2026/09/14 22:10:35 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:37 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:37 [WARN] Health probe jellyfin: API GET :8096/health → Get "http://jellyfin:8096/health": dial tcp: lookup jellyfin on 192.168.0.250:53: no such host
2026/09/14 22:10:37 [WARN] Health probe nextcloud: API GET :80/status.php → Get "http://nextcloud:80/status.php": dial tcp: lookup nextcloud on 192.168.0.250:53: no such host
2026/09/14 22:10:37 [WARN] Health probe immich: API GET :2283/api/server/ping → Get "http://immich-postgres:2283/api/server/ping": dial tcp: lookup immich-postgres on 192.168.0.250:53: no such host
2026/09/14 22:10:37 [WARN] Health probe homebox: API GET :7745/api/v1/status → Get "http://homebox:7745/api/v1/status": dial tcp: lookup homebox on 192.168.0.250:53: no such host
2026/09/14 22:10:37 [WARN] Health probe gokapi: HTTP GET :53842/ → Get "http://gokapi:53842/": dial tcp: lookup gokapi on 192.168.0.250:53: no such host
2026/09/14 22:10:37 [WARN] Health probe vaultwarden: API GET :80/alive → Get "http://vaultwarden:80/alive": dial tcp: lookup vaultwarden on 192.168.0.250:53: no such host
2026/09/14 22:10:37 [WARN] Health probes: 1 ok, 6 unhealthy (of 7 probed)
2026/09/14 22:10:40 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:43 [INFO] [stacks] Stack immich started successfully (took 13.5s)
2026/09/14 22:10:46 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:46 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:46 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app jellyfin (state=starting) onto /mnt/felhom-drives/hdd_1
2026/09/14 22:10:46 [INFO] [stacks] Stopping stack: jellyfin
2026/09/14 22:10:47 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:50 [INFO] [stacks] Stack immich post-start status:
2026/09/14 22:10:50 [INFO] [stacks] immich-machine-learning ghcr.io/immich-app/immich-machine-learning:v3.0.3 running Up 18 seconds (health: starting)
2026/09/14 22:10:50 [INFO] [stacks] immich-postgres ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 running Up 18 seconds (healthy)
2026/09/14 22:10:50 [INFO] [stacks] immich-redis redis:7-alpine running Up 18 seconds (healthy)
2026/09/14 22:10:50 [INFO] [stacks] immich-server ghcr.io/immich-app/immich-server:v3.0.3 running Up 7 seconds (health: starting)
2026/09/14 22:10:50 [WARN] Health probe jellyfin: API GET :8096/health → expected status 200, got 503
2026/09/14 22:10:50 [WARN] Health probes: 5 ok, 1 unhealthy (of 6 probed)
2026/09/14 22:10:51 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:54 [INFO] [stacks] Stack jellyfin stopped successfully (took 7.9s)
2026/09/14 22:10:54 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
2026/09/14 22:10:54 [INFO] [stacks] Starting stack: jellyfin
2026/09/14 22:10:55 [INFO] [stacks] Stack jellyfin started successfully (took 0.6s)
2026/09/14 22:10:55 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:55 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app nextcloud (state=starting) onto /mnt/felhom-drives/hdd_1
2026/09/14 22:10:55 [INFO] [stacks] Stopping stack: nextcloud
2026/09/14 22:10:56 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:10:57 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
2026/09/14 22:10:57 [INFO] Health probes: 1 ok (of 1 probed)
2026/09/14 22:10:57 [INFO] [stacks] Stack nextcloud stopped successfully (took 2.2s)
2026/09/14 22:10:57 [INFO] [stacks] Status refresh: 26 containers across 57 stacks
2026/09/14 22:10:57 [INFO] [stacks] Starting stack: nextcloud
2026/09/14 22:10:58 [INFO] [stacks] Stack jellyfin post-start status:
2026/09/14 22:10:58 [INFO] [stacks] jellyfin jellyfin/jellyfin:10.11.11 running Up 4 seconds (health: starting)
2026/09/14 22:10:59 [INFO] [selfupdate] Current version 0.242.0 is up to date
2026/09/14 22:11:01 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:06 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:07 [WARN] Health probe jellyfin: API GET :8096/health → Get "http://jellyfin:8096/health": dial tcp 172.18.0.4:8096: connect: connection refused
2026/09/14 22:11:07 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:07 [WARN] Health probes: 0 ok, 1 unhealthy (of 1 probed)
2026/09/14 22:11:09 [INFO] [stacks] Stack nextcloud started successfully (took 12.0s)
2026/09/14 22:11:10 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:10 [INFO] [gate] boot 1789423711-10047: live bind confirmed — recreating drive-backed app paperless-ngx (state=starting) onto /mnt/felhom-drives/hdd_1
2026/09/14 22:11:10 [INFO] [stacks] Stopping stack: paperless-ngx
2026/09/14 22:11:12 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:13 [INFO] [stacks] Stack nextcloud post-start status:
2026/09/14 22:11:13 [INFO] [stacks] nextcloud nextcloud:34.0.1-apache running Up 4 seconds (health: starting)
2026/09/14 22:11:13 [INFO] [stacks] nextcloud-db mariadb:11.6 running Up 15 seconds (healthy)
2026/09/14 22:11:13 [INFO] [stacks] nextcloud-redis redis:7-alpine running Up 15 seconds (healthy)
2026/09/14 22:11:17 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:17 [INFO] Health probes: 1 ok (of 1 probed)
2026/09/14 22:11:17 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:21 [INFO] [stacks] Stack paperless-ngx stopped successfully (took 11.3s)
2026/09/14 22:11:21 [INFO] [stacks] Status refresh: 26 containers across 57 stacks
2026/09/14 22:11:21 [INFO] [stacks] Starting stack: paperless-ngx
2026/09/14 22:11:22 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:22 [INFO] [bootrecon] boot window: budget 50s exhausted while the fleet was still changing — sweeping anyway
2026/09/14 22:11:22 [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [paperless-ngx] — up to 2 attempt(s)
2026/09/14 22:11:22 [INFO] [stacks] Starting stack: paperless-ngx
2026/09/14 22:11:27 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:27 [INFO] [scheduler] Running job: agent-channel-health
2026/09/14 22:11:27 [INFO] Health probes: 5 ok (of 5 probed)
2026/09/14 22:11:27 [INFO] [scheduler] Job agent-channel-health completed (took 185ms)
2026/09/14 22:11:32 [INFO] [stacks] Stack paperless-ngx started successfully (took 9.9s)
2026/09/14 22:11:32 [INFO] [stacks] Stack paperless-ngx started successfully (took 11.4s)
2026/09/14 22:11:32 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:32 [INFO] [bootrecon] Boot reconciliation attempt 1/2: started "paperless-ngx" (took 10.1s)
2026/09/14 22:11:32 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:32 [INFO] [bootrecon] Boot reconciliation complete: 1 app(s) recovered in 1 attempt(s): [paperless-ngx]
2026/09/14 22:11:33 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:33 [INFO] [gate] boot 1789423711-10047: re-syncing FileBrowser mounts against the live binds
2026/09/14 22:11:33 [INFO] [settings] Settings saved
2026/09/14 22:11:33 [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s))
2026/09/14 22:11:35 [INFO] [stacks] Stack paperless-ngx post-start status:
2026/09/14 22:11:35 [INFO] [stacks] paperless-postgres postgres:16-alpine running Up 14 seconds (healthy)
2026/09/14 22:11:35 [INFO] [stacks] paperless-redis redis:7-alpine running Up 14 seconds (healthy)
2026/09/14 22:11:35 [INFO] [stacks] paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 running Up 3 seconds (health: starting)
2026/09/14 22:11:36 [INFO] [stacks] Stack paperless-ngx post-start status:
2026/09/14 22:11:36 [INFO] [stacks] paperless-postgres postgres:16-alpine running Up 14 seconds (healthy)
2026/09/14 22:11:36 [INFO] [stacks] paperless-redis redis:7-alpine running Up 14 seconds (healthy)
2026/09/14 22:11:36 [INFO] [stacks] paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 running Up 3 seconds (health: starting)
2026/09/14 22:11:37 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:47 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:11:55 [INFO] [fillwatch] checked 3 filesystem(s), 0 unreadable/skipped, 0 notification(s); bands: all ok
2026/09/14 22:11:57 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:12:07 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:12:07 [INFO] [web] Login from 172.18.0.16:51988
2026/09/14 22:12:17 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:12:27 [INFO] [scheduler] Running job: stack-scan
2026/09/14 22:12:27 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:12:27 [INFO] [stacks] ScanStacks complete: 57 stacks found (13 deployed, 44 available)
2026/09/14 22:12:27 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:12:27 [INFO] [scheduler] Job stack-scan completed (took 84ms)
2026/09/14 22:12:27 [INFO] [scheduler] Running job: agent-channel-health
2026/09/14 22:12:27 [INFO] [scheduler] Job agent-channel-health completed (took 187ms)
2026/09/14 22:12:37 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
2026/09/14 22:12:47 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
@@ -1,2 +1,26 @@
F9 power-cycle (qm reset) at 22:08:17
22:08:23 +5s health=000
22:08:43 +25s health=000
22:09:01 +43s health=000
22:09:19 +61s health=000
22:09:37 +79s health=000
22:09:56 +98s health=000
22:10:14 +116s health=000
22:10:29 +131s health=200
F9 + 131s dashboard /api/health 200
F9 + 132s privatebin running
F9 + 132s gokapi running
F9 + 138s nextcloud running
F9 + 138s vaultwarden running
F9 + 138s jellyfin running
F9 + 149s bookstack running
F9 + 156s uptime-kuma running
F9 + 178s mealie running
F9 + 188s docmost running
F9 + 194s immich running
F9 + 194s adventurelog running
F9 + 229s paperless-ngx running
F9 + 229s ALL 12 running
--- homebox after reboot
state running deploying False deployed True {'homebox': 'ghcr.io/sysadminsmedia/homebox:0.26.2'}
Fut Naprakész
+1 -1
View File
@@ -726,7 +726,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** |
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle: see the F9 journal (measured after filing). `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** |
| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle (`qm reset` 22:08:17Z): the bootstrap started the controller at boot (22:10:23Z), dashboard 200 at +131 s, all 12 apps running at +229 s, the deployed homebox `running · deploying false · deployed true` — „Fut · Naprakész", **not stuck**. Hub `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z with its mail **also suppressed by cooldown**. `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.