chaos-night fixes: R-539 closed PROVEN-LIVE, the morning note, the report
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
R-539 closed: five real controller kills on demo-hp 9201 with the production 24 h window raised controller_slow_crashloop, and exactly one operator mail arrived (09:29:40Z). The fast brake never armed. Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed). unproven.py: 35 of 55 not walked, no number moved. The report names the brief's wrong claims first and one recommendation not followed: the controller floor was not raised - validated on one guest, a gap of my own found during validation, and a floor above the golden reaches Peti's box too. The operator's call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,33 @@
|
||||
# Part C LIVE - the slow crash loop with the PRODUCTION 24 h window (no test-only interval)
|
||||
# 9201 on demo-hp, agent 0.132.0. Restart #1 already persisted from proof A (08:52:29Z).
|
||||
# Four more kills, 8 min apart: any 15-minute window holds at most two restarts, so the fast brake never arms.
|
||||
2026-09-17T08:57:38Z persisted before: {"restarts":["2026-09-17T08:52:29.969961515Z"],"slow_crashloop_since":"0001-01-01T00:00:00Z"}
|
||||
2026-09-17T08:57:39Z KILL for restart #2: felhom-controller
|
||||
2026-09-17T08:58:36Z controller back: running after 56 s (not by hand)
|
||||
Sep 17 10:58:32 demo-hp felhom-agent[1195783]: time=2026-09-17T10:58:32.336+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=2
|
||||
2026-09-17T09:05:40Z KILL for restart #3: felhom-controller
|
||||
2026-09-17T09:06:37Z controller back: running after 55 s (not by hand)
|
||||
Sep 17 11:06:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:06:32.247+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=3
|
||||
2026-09-17T09:13:42Z KILL for restart #4: felhom-controller
|
||||
2026-09-17T09:14:33Z controller back: running after 50 s (not by hand)
|
||||
Sep 17 11:14:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:14:32.222+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=4
|
||||
2026-09-17T09:21:43Z KILL for restart #5: felhom-controller
|
||||
2026-09-17T09:22:35Z controller back: running after 50 s (not by hand)
|
||||
Sep 17 11:22:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:22:32.184+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=5
|
||||
Sep 17 11:22:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:22:32.184+02:00 level=WARN msg="controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop" vmid=9201 restarts_24h=5 window=24h0m0s threshold=5
|
||||
2026-09-17T09:22:35Z persisted after: {"restarts":["2026-09-17T08:52:29.969961515Z","2026-09-17T08:58:30.062356117Z","2026-09-17T09:06:29.97157901Z","2026-09-17T09:14:29.970146017Z","2026-09-17T09:22:29.963142311Z"],"slow_crashloop_since":"2026-09-17T09:22:29.963142311Z"}
|
||||
2026-09-17T09:22:35Z waiting for the hub to mint controller_slow_crashloop (the agent's heartbeat carries it) ...
|
||||
2026-09-17T09:29:50Z hub customer page mentions of controller_slow_crashloop: 3
|
||||
warning controller_slow_crashloop Host demo-hp-bb76ea guest 9201: the controller keeps dying — the agent restarted it 5 times in 24 hours (each one too far apart for the 15-minute brake). It is still being restarted; this is the warning that it will not stay up. Last reason: controller container exited on 2 consecutive sweeps hub
|
||||
hub log: 2026/09/17 11:29:40 [INFO] Controller supervisor: controller_slow_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps)
|
||||
hub log: 2026/09/17 11:29:40 [INFO] Operator email sent for demo-hp/controller_slow_crashloop
|
||||
|
||||
## DELIVERED, not just "sent" - the mailbox itself (read 2026-09-17T09:31Z)
|
||||
exactly ONE mail for the five restarts:
|
||||
2026-09-17T09:29:40Z monitoring@felhom.eu -> admin@felhom.eu
|
||||
"[Felhom] ⚠️ demo-hp: controller_slow_crashloop"
|
||||
"Customer: demo-hp Event: controller_slow_crashloop Severity: warning ... Host demo-hp-bb76ea guest 9201:
|
||||
the controller keeps dying — the agent restarted it 5 times in 24 hours ..."
|
||||
## After: 9201 controller 0.246.0 Up (healthy), 24 containers running.
|
||||
## Standing effect on the demo box, stated: slow_crashloop stays raised on 9201 until 09:22Z tomorrow and
|
||||
## the counter holds 5 restarts; it cannot re-mail inside that window (moves at most once per 24 h).
|
||||
Reference in New Issue
Block a user