REPORT + STATUS: P1 fixes — supervisor proven incl. crash-loop resume and hub events; publish; rows
gates / gates (push) Successful in 18s

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-15 11:45:21 +02:00
parent 0a6cf60bf8
commit 351296114c
5 changed files with 99 additions and 38 deletions
@@ -0,0 +1,7 @@
## 2026-09-15T09:36:30Z dashboard health: 200
Sep 15 11:34:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:34:52.719+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
Sep 15 11:35:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:35:22.727+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
Sep 15 11:35:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:35:52.706+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
Sep 15 11:36:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:36:22.781+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exited unit=felhom-controller-bootstrap.service
Sep 15 11:36:24 demo-hp felhom-agent[1526161]: time=2026-09-15T11:36:24.022+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps"
status=running started=2026-09-15T09:36:23.860962402Z
@@ -2,3 +2,6 @@
2026/09/15 11:14:37 [INFO] Controller supervisor: controller_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps)
2026/09/15 11:14:37 [INFO] Controller supervisor: controller_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps)
2026/09/15 11:14:38 [INFO] Operator email sent for demo-hp/controller_crashloop
## 2026-09-15T09:44:52Z after the 09:36:24Z resume restart:
2026/09/15 11:44:37 [INFO] Controller supervisor: controller_restarted_by_agent demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps)
## NOTE: the crash-loop line appears twice above only because the watch script printed it twice; the single hub pod logged it ONCE (checked per pod).
@@ -9,3 +9,10 @@ Sep 15 11:09:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:09:52.697+02:0
Sep 15 11:10:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:10:22.739+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
## teardown: remove http=502 at 2026-09-15T09:10:40Z
## CORRECTION 2026-09-15T09:11:57Z: the 'health 200 at 09:09:08Z — 248 s after the kill' line is WRONG. docker inspect: felhom-controller exited 09:05:02Z (exit 137, the kill) and was NOT restarted. What answered 200 once is unexplained. What happened instead, and it is the guard working as designed: the agent had restarted this controller at 08:54:24Z, 08:56:53Z and 08:58:24Z (my idle, unpark and failed-swap tests; the swap's own rollback at 09:01:57Z was correctly NOT counted), so on the second not-running sweep after this kill it logged 'CRASH-LOOP … restarts_in_window=3' at 09:05:52Z and paused restarts for 30 min. Mid-deploy restart timing is therefore NOT measured here; the resume after the pause is captured separately.
## homebox after resume: {'state': 'not_deployed', 'deployed': False, 'deploying': False}
homebox not deployed — nothing to remove
## after cleanup: 0
## after cleanup: /opt/docker/stacks/homebox/app.yaml
## homebox app.yaml: 2026-09-15 09:04:55.994631593 +0000 451
## homebox app.yaml: deployed deployed_at env locked_fields desired_state
## homebox app.yaml: moved-aside