BIGNIGHT F9: controller killed stays dead 33 min, no alarm delivered — R-523 (P1), stop rule met; F10-F12 not run
gates / gates (push) Successful in 19s

This commit is contained in:
2026-09-15 00:08:27 +02:00
parent ce011b823d
commit adc38d1f72
10 changed files with 133 additions and 0 deletions
+1
View File
@@ -726,6 +726,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** |
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle: see the F9 journal (measured after filing). `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.