BIGNIGHT phase 6: apps healthy, R-524 (downgrade offered as update), backup pages re-read
gates / gates (push) Successful in 19s
gates / gates (push) Successful in 19s
This commit is contained in:
@@ -727,6 +727,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
|
||||
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle (`qm reset` 22:08:17Z): the bootstrap started the controller at boot (22:10:23Z), dashboard 200 at +131 s, all 12 apps running at +229 s, the deployed homebox `running · deploying false · deployed true` — „Fut · Naprakész", **not stuck**. Hub `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z with its mail **also suppressed by cooldown**. `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** |
|
||||
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user