## F9'' - three more controller kills, 20 minutes apart, at IDLE (no deploy running)
##   Question to MEASURE (not to change): does a restart after 20 minutes of healthy uptime
##   count against the 3-restarts-per-15-minutes budget, or does the window start fresh?
##   Baseline: the last restart before this run was F9' at 12:32:16 CEST (1 of 3 at that time).
## kill 1 at 2026-09-16T11:12:41Z
felhom-controller
   dashboard 200 again after 61 s
Sep 16 13:12:37 tester1 felhom-agent[1128]: time=2026-09-16T13:12:37.326+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=40 guests_evaluated=1 controllers_not_running=0
Sep 16 13:13:07 tester1 felhom-agent[1128]: time=2026-09-16T13:13:07.323+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
Sep 16 13:13:37 tester1 felhom-agent[1128]: time=2026-09-16T13:13:37.415+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
Sep 16 13:13:38 tester1 felhom-agent[1128]: time=2026-09-16T13:13:38.655+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
   RESTARTED lines in the whole journal so far: 2
   waiting 20 minutes before the next kill
## kill 2 at 2026-09-16T11:34:05Z
felhom-controller
   dashboard 200 again after 41 s
Sep 16 13:32:37 tester1 felhom-agent[1128]: time=2026-09-16T13:32:37.415+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=80 guests_evaluated=1 controllers_not_running=0
Sep 16 13:34:07 tester1 felhom-agent[1128]: time=2026-09-16T13:34:07.341+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
Sep 16 13:34:37 tester1 felhom-agent[1128]: time=2026-09-16T13:34:37.379+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
Sep 16 13:34:38 tester1 felhom-agent[1128]: time=2026-09-16T13:34:38.709+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
   RESTARTED lines in the whole journal so far: 3
   waiting 20 minutes before the next kill
## kill 3 at 2026-09-16T11:55:09Z
felhom-controller
   dashboard 200 again after 61 s
Sep 16 13:52:37 tester1 felhom-agent[1128]: time=2026-09-16T13:52:37.361+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=120 guests_evaluated=1 controllers_not_running=0
Sep 16 13:55:37 tester1 felhom-agent[1128]: time=2026-09-16T13:55:37.379+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
Sep 16 13:56:07 tester1 felhom-agent[1128]: time=2026-09-16T13:56:07.300+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
Sep 16 13:56:08 tester1 felhom-agent[1128]: time=2026-09-16T13:56:08.536+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
   RESTARTED lines in the whole journal so far: 4
## F9'' finished at 2026-09-16T11:56:32Z
## RESULT - F9'' (three kills at IDLE, 20 minutes apart)
##  kill 1 11:12:41Z -> dashboard 200 again after 61 s  (restart #2 in the journal)
##  kill 2 11:34:05Z -> dashboard 200 again after 41 s  (restart #3)
##  kill 3 11:55:09Z -> dashboard 200 again after 61 s  (restart #4)
##  Every kill was seen on one sweep, confirmed on the next, and restarted - 30 to 90 s of dashboard
##  downtime each time, and the apps themselves never stopped (they do not depend on the controller).
##
##  THE ANSWER TO THE BUDGET QUESTION, measured and NOT changed:
##   Restarts 20 minutes apart NEVER accumulate. The budget is 3 restarts inside a 15-minute window, so
##   each of these four restarts started a fresh window and the 30-minute pause was never armed.
##   Total this session: 4 restarts, 0 pauses, 0 crash-loop events.
##  CONSEQUENCE, for the operator to rule on (a design question, not a defect):
##   A box whose controller dies every 20 minutes is restarted forever, quietly. The only signal is the
##   info-level `controller_restarted_by_agent` event, which mails nobody. The brake catches a FAST loop
##   (3 in 15 min) and is blind to a SLOW one. Options: (a) leave it - the box self-heals and the record
##   is on the timeline; (b) add a second, longer counter (for example 5 restarts in 6 hours) that raises
##   a warning-severity event; (c) raise the severity of the Nth restart in a day. This session measured
##   only; it changed nothing.
## HONEST GAP in this run, caused by MY teardown timing, not by the product:
##  The hub minted `controller_restarted_by_agent` for the kill-1 and kill-2 restarts (13:22:04 and
##  13:37:04 CEST in its log). The kill-3 restart happened at 13:56:08 and I destroyed the VM at
##  ~13:57:30 - about 84 s later, inside the box's own report cycle. So the third restart never reached
##  the hub. This is expected: the supervisor's record RIDES the host report, and a box that disappears
##  before its next report takes the unsent record with it. It is recorded here so nobody later reads
##  "2 events for 3 restarts" as a defect. A real box does not vanish 84 s after a restart.
