BIGNIGHT F9: controller killed stays dead 33 min, no alarm delivered — R-523 (P1), stop rule met; F10-F12 not run
gates / gates (push) Successful in 19s
gates / gates (push) Successful in 19s
This commit is contained in:
@@ -627,3 +627,19 @@ itself, through the product's restore endpoint, and reads the data back — reco
|
||||
| hub shows DOWN, and when | never DOWN (threshold 1 h). Status „warn · Last report N min ago" climbing 14 → 31 min; **`node_stale` 21:25:43Z** (31 min after the last report, 17 min into the cut) + operator mail; **`node_recovered` 21:26:43Z** + mail |
|
||||
| alarms true? | `node_stale` true; `node_recovered` true — a cut shorter than ≈ 13 min after a report would raise nothing, by design |
|
||||
| should have fired, did not | a customer-side notice that the box is offline (R-522) |
|
||||
|
||||
### F9 — the controller dies in the middle of a deploy (`phase5/F9-controller-killed-mid-deploy.txt`, `phase5/F9/`)
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| how | deploy of a throwaway app (homebox) `POST` 21:34:35Z → „Telepítés elindítva" (with a memory warning: „Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normál használat mellett ez nem okoz problémát.") → **`docker kill felhom-controller` at 21:34:39Z** (4 s in) |
|
||||
| what the box did by itself | **nothing brought the controller back**: `Exited (137)`, `restart=unless-stopped` — Docker does not restart a container stopped by `kill`; the in-guest `felhom-controller-bootstrap` unit is a one-shot („active (exited)" since boot) and does not watch it. 5 min later still dead, dashboard **502** on the LAN |
|
||||
| the deploy | the app itself came up: container `homebox Up 5 minutes (healthy)`; `app.yaml` and `applied-compose.yml` written at 21:34 |
|
||||
| what the customer saw | the dashboard answers 502 — no page at all |
|
||||
| 23 min later (21:57:39Z) | still `Exited (137) 22 minutes ago`, LAN health **502**; hub shows only `app_deployed (info) — Homebox` (the controller's last report reached the hub at **21:34:38Z**, 3 s before the kill, so the 30-minute staleness is due ≈ 22:04:38Z). The host agent's journal meanwhile: drive reconcile and stale-lock scans every 20 s, **nothing about the controller** |
|
||||
| 33 min later (22:08:04Z) | still `Exited (137) 33 minutes ago`, LAN 502. Hub **22:04:43Z `node_stale` — operator mail suppressed by cooldown** (key `tester-1:node_stale`, set by F8 at 21:25:43Z) |
|
||||
|
||||
**STOP RULE MET (22:08Z).** F9 left the box in a state the product did not recover from on its own (33 min) and that a
|
||||
customer cannot recover from any screen (there is no dashboard). Filed **R-523 (P1)** before acting. Per the brief:
|
||||
**no further faults are injected — F10, F11 and F12 are not run**; the box is recovered by the household's only lever,
|
||||
a power-cycle, and the night moves to Phase 6.
|
||||
|
||||
@@ -11,3 +11,6 @@
|
||||
21:29:23 guest->1.1.1.1=301 guest->hub=302 cloudflared(last120s registered=0 errors=8)
|
||||
21:31:14 guest->1.1.1.1=301 guest->hub=302 cloudflared(last120s registered=0 errors=8)
|
||||
21:33:05 guest->1.1.1.1=301 guest->hub=302 cloudflared(last120s registered=0 errors=8)
|
||||
21:34:57 guest->1.1.1.1=301 guest->hub=302 cloudflared(last120s registered=0 errors=10)
|
||||
21:36:49 guest->1.1.1.1=301 guest->hub=302 cloudflared(last120s registered=0 errors=8)
|
||||
21:38:40 guest->1.1.1.1=301 guest->hub=302 cloudflared(last120s registered=0 errors=8)
|
||||
|
||||
+12
@@ -54,3 +54,15 @@ LAN dashboard banners: ↗ |
|
||||
hub: Last report: 7 min ago · Controller 0.242.0 | Auto-refresh | (paused) | Tester 1 | ok | Controller | 0.242.0 |
|
||||
LAN dashboard banners: ↗ |
|
||||
tunnel tile: Cloudflare Tunnel | Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. | Fut | Védett | Doc
|
||||
=== 21:34:32
|
||||
hub: Last report: 8 min ago · Controller 0.242.0 | Auto-refresh | (paused) | Tester 1 | ok | Controller | 0.242.0 |
|
||||
LAN dashboard banners: ↗ |
|
||||
tunnel tile: Cloudflare Tunnel | Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. | Fut | Védett | Doc
|
||||
=== 21:36:23
|
||||
hub: Last report: 2 min ago · Controller 0.242.0 | Auto-refresh | (paused) | Tester 1 | ok | Controller | 0.242.0 |
|
||||
LAN dashboard banners: (none)
|
||||
tunnel tile:
|
||||
=== 21:38:13
|
||||
hub: Last report: 4 min ago · Controller 0.242.0 | Auto-refresh | (paused) | Tester 1 | ok | Controller | 0.242.0 |
|
||||
LAN dashboard banners: (none)
|
||||
tunnel tile:
|
||||
|
||||
+45
@@ -0,0 +1,45 @@
|
||||
F9 docker kill felhom-controller at 21:34:39 (deploy started: POST 21:34:35)
|
||||
felhom-controller
|
||||
felhom-controller Exited (137) 3 seconds ago
|
||||
21:34:45 +1s controller=[Exited (137) 4 seconds ago] health=502
|
||||
21:35:02 +18s controller=[Exited (137) 20 seconds ago] health=502
|
||||
21:35:18 +34s controller=[Exited (137) 37 seconds ago] health=502
|
||||
21:35:35 +51s controller=[Exited (137) 53 seconds ago] health=502
|
||||
21:35:51 +67s controller=[Exited (137) About a minute ago] health=502
|
||||
21:36:08 +84s controller=[Exited (137) About a minute ago] health=502
|
||||
21:36:25 +101s controller=[Exited (137) About a minute ago] health=502
|
||||
21:36:41 +117s controller=[Exited (137) About a minute ago] health=502
|
||||
21:36:57 +133s controller=[Exited (137) 2 minutes ago] health=502
|
||||
21:37:14 +150s controller=[Exited (137) 2 minutes ago] health=502
|
||||
21:37:30 +166s controller=[Exited (137) 2 minutes ago] health=502
|
||||
21:37:47 +183s controller=[Exited (137) 3 minutes ago] health=502
|
||||
21:38:03 +199s controller=[Exited (137) 3 minutes ago] health=502
|
||||
21:38:20 +216s controller=[Exited (137) 3 minutes ago] health=502
|
||||
21:38:36 +232s controller=[Exited (137) 3 minutes ago] health=502
|
||||
21:38:53 +249s controller=[Exited (137) 4 minutes ago] health=502
|
||||
21:39:10 +266s controller=[Exited (137) 4 minutes ago] health=502
|
||||
21:39:26 +282s controller=[Exited (137) 4 minutes ago] health=502
|
||||
21:39:43 +299s controller=[Exited (137) 5 minutes ago] health=502
|
||||
21:39:59 +315s controller=[Exited (137) 5 minutes ago] health=502
|
||||
21:40:16 +332s controller=[Exited (137) 5 minutes ago] health=502
|
||||
21:40:32 +348s controller=[Exited (137) 5 minutes ago] health=502
|
||||
21:40:49 +365s controller=[Exited (137) 6 minutes ago] health=502
|
||||
21:41:05 +381s controller=[Exited (137) 6 minutes ago] health=502
|
||||
21:41:22 +398s controller=[Exited (137) 6 minutes ago] health=502
|
||||
21:41:38 +414s controller=[Exited (137) 6 minutes ago] health=502
|
||||
21:41:55 +431s controller=[Exited (137) 7 minutes ago] health=502
|
||||
21:42:11 +447s controller=[Exited (137) 7 minutes ago] health=502
|
||||
21:42:28 +464s controller=[Exited (137) 7 minutes ago] health=502
|
||||
21:42:44 +480s controller=[Exited (137) 8 minutes ago] health=502
|
||||
21:43:01 +497s controller=[Exited (137) 8 minutes ago] health=502
|
||||
21:43:17 +513s controller=[Exited (137) 8 minutes ago] health=502
|
||||
21:43:34 +530s controller=[Exited (137) 8 minutes ago] health=502
|
||||
21:43:50 +546s controller=[Exited (137) 9 minutes ago] health=502
|
||||
21:44:07 +563s controller=[Exited (137) 9 minutes ago] health=502
|
||||
21:44:23 +579s controller=[Exited (137) 9 minutes ago] health=502
|
||||
21:44:40 +596s controller=[Exited (137) 9 minutes ago] health=502
|
||||
21:44:57 +613s controller=[Exited (137) 10 minutes ago] health=502
|
||||
21:45:13 +629s controller=[Exited (137) 10 minutes ago] health=502
|
||||
21:45:30 +646s controller=[Exited (137) 10 minutes ago] health=502
|
||||
--- app after restart
|
||||
Bad Gateway
|
||||
@@ -0,0 +1,4 @@
|
||||
=== 21:57:39 controller still dead? Exited (137) 22 minutes ago
|
||||
LAN health: 502
|
||||
=== hub alarms 21:34Z..
|
||||
2026/09/14 23:34:35 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Homebox
|
||||
@@ -0,0 +1,4 @@
|
||||
=== 22:08:04 controller: Exited (137) 33 minutes ago · LAN health 502
|
||||
2026/09/14 23:34:35 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Homebox
|
||||
2026/09/15 00:04:43 [INFO] Staleness: tester-1 ok → stale (node_stale)
|
||||
2026/09/15 00:04:43 [INFO] Operator email suppressed for tester-1/node_stale — cooldown (key=tester-1:node_stale)
|
||||
@@ -0,0 +1,29 @@
|
||||
2026/09/14 21:33:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:33:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:33:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:33:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:33:45 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:33:45 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/09/14 21:33:45 [INFO] [scheduler] Job agent-channel-health completed (took 189ms)
|
||||
2026/09/14 21:33:55 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:34:05 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:34:15 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:34:25 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:34:34 [INFO] [web] Login from 172.18.0.10:37154
|
||||
2026/09/14 21:34:35 [INFO] [api] Deploy requested for stack: homebox
|
||||
2026/09/14 21:34:35 [INFO] [stacks] Memory check: total=11828MB, reserved=384MB, usable=11444MB, committed_used=4126MB, new_req=50MB, remaining=7268MB
|
||||
2026/09/14 21:34:35 [INFO] [stacks] Deploying stack homebox with 3 env vars: [DOMAIN, HBOX_AUTH_API_KEY_PEPPER, SUBDOMAIN]
|
||||
2026/09/14 21:34:35 [INFO] [stacks] SaveAppConfig: saved config for homebox
|
||||
2026/09/14 21:34:35 [INFO] Event pushed: app_deployed (info) — Alkalmazás telepítve: Homebox
|
||||
2026/09/14 21:34:35 [INFO] [stacks] Status refresh: 28 containers across 57 stacks
|
||||
2026/09/14 21:34:37 [INFO] [report] Building system report
|
||||
2026/09/14 21:34:37 [INFO] [monitor] Health check: status=ok
|
||||
2026/09/14 21:34:38 [INFO] [report] Hub report pushed successfully (13817 bytes)
|
||||
2026/09/14 21:34:38 [INFO] [settings] Settings saved
|
||||
2026/09/14 21:34:40 [INFO] [stacks] Stack homebox deployed successfully (took 5.7s)
|
||||
2026/09/14 21:34:40 [INFO] [stacks] SaveAppConfig: saved config for homebox
|
||||
2026/09/14 21:34:41 [INFO] [stacks] installed-images homebox: recorded 1 service(s) (homebox=ghcr.io/sysadminsmedia/homebox:0.26.2 (sha256:b1ad7e3c63f7…))
|
||||
2026/09/14 21:34:41 [INFO] [stacks] SaveAppConfig: saved config for homebox
|
||||
2026/09/14 21:34:41 [INFO] [stacks] SaveAppConfig: saved config for homebox
|
||||
2026/09/14 21:34:41 [INFO] [stacks] pin homebox: homebox=ghcr.io/sysadminsmedia/homebox:0.26.2
|
||||
2026/09/14 21:34:41 [INFO] [stacks] Status refresh: 29 containers across 57 stacks
|
||||
+17
@@ -0,0 +1,17 @@
|
||||
● felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)
|
||||
Loaded: loaded (/etc/systemd/system/felhom-controller-bootstrap.service; enabled; preset: enabled)
|
||||
Active: active (exited) since Mon 2026-09-14 19:54:43 UTC; 1h 45min ago
|
||||
Invocation: acc822de0aa047f7a7e65fb6cef0f587
|
||||
TriggeredBy: ● felhom-controller-bootstrap.path
|
||||
Process: 3204 ExecStart=/usr/local/sbin/felhom-controller-bootstrap.sh (code=exited, status=0/SUCCESS)
|
||||
Main PID: 3204 (code=exited, status=0/SUCCESS)
|
||||
Mem peak: 48.6M
|
||||
restart=unless-stopped exit=137 oom=false finished=2026-09-14T21:34:41.542160118Z
|
||||
homebox Up 5 minutes (healthy)
|
||||
total 24
|
||||
drwxr-xr-x 2 root root 4096 Sep 14 21:34 .
|
||||
drwxr-xr-x 59 root root 4096 Sep 14 18:50 ..
|
||||
-rw-r--r-- 1 root root 2601 Sep 14 17:58 .felhom.yml
|
||||
-rw------- 1 root root 723 Sep 14 21:34 app.yaml
|
||||
-rw-r--r-- 1 root root 1636 Sep 14 21:34 applied-compose.yml
|
||||
-rw-r--r-- 1 root root 1636 Sep 14 17:58 docker-compose.yml
|
||||
@@ -0,0 +1,2 @@
|
||||
F9 power-cycle (qm reset) at 22:08:17
|
||||
22:08:23 +5s health=000
|
||||
Reference in New Issue
Block a user