drill 0.243.0: F9'' answers the supervisor budget question; machine torn down
gates / gates (push) Successful in 20s

Three more kills 20 minutes apart: recovered in 61 s / 41 s / 61 s, and none of
them accumulated, because the window is 15 minutes. Four restarts, zero pauses.
So the brake catches a FAST crash loop and is blind to a SLOW one — a controller
dying every 20 minutes is restarted forever, and the only trace is an info event
that mails nobody. Measured, not changed: the options are written into R-531 for
the operator to rule on.

Machine layer torn down: VM 334 purged with its disks, demo-hp's own containers
9201 and 9202 untouched. Evidence copied off the box first, token-leak control 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 13:59:27 +02:00
parent db58af80a2
commit 0b63574293
8 changed files with 2308 additions and 1 deletions
@@ -11,3 +11,46 @@ Sep 16 13:13:37 tester1 felhom-agent[1128]: time=2026-09-16T13:13:37.415+02:00 l
Sep 16 13:13:38 tester1 felhom-agent[1128]: time=2026-09-16T13:13:38.655+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
RESTARTED lines in the whole journal so far: 2
waiting 20 minutes before the next kill
## kill 2 at 2026-09-16T11:34:05Z
felhom-controller
dashboard 200 again after 41 s
Sep 16 13:32:37 tester1 felhom-agent[1128]: time=2026-09-16T13:32:37.415+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=80 guests_evaluated=1 controllers_not_running=0
Sep 16 13:34:07 tester1 felhom-agent[1128]: time=2026-09-16T13:34:07.341+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
Sep 16 13:34:37 tester1 felhom-agent[1128]: time=2026-09-16T13:34:37.379+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
Sep 16 13:34:38 tester1 felhom-agent[1128]: time=2026-09-16T13:34:38.709+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
RESTARTED lines in the whole journal so far: 3
waiting 20 minutes before the next kill
## kill 3 at 2026-09-16T11:55:09Z
felhom-controller
dashboard 200 again after 61 s
Sep 16 13:52:37 tester1 felhom-agent[1128]: time=2026-09-16T13:52:37.361+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=120 guests_evaluated=1 controllers_not_running=0
Sep 16 13:55:37 tester1 felhom-agent[1128]: time=2026-09-16T13:55:37.379+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
Sep 16 13:56:07 tester1 felhom-agent[1128]: time=2026-09-16T13:56:07.300+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
Sep 16 13:56:08 tester1 felhom-agent[1128]: time=2026-09-16T13:56:08.536+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
RESTARTED lines in the whole journal so far: 4
## F9'' finished at 2026-09-16T11:56:32Z
## RESULT - F9'' (three kills at IDLE, 20 minutes apart)
## kill 1 11:12:41Z -> dashboard 200 again after 61 s (restart #2 in the journal)
## kill 2 11:34:05Z -> dashboard 200 again after 41 s (restart #3)
## kill 3 11:55:09Z -> dashboard 200 again after 61 s (restart #4)
## Every kill was seen on one sweep, confirmed on the next, and restarted - 30 to 90 s of dashboard
## downtime each time, and the apps themselves never stopped (they do not depend on the controller).
##
## THE ANSWER TO THE BUDGET QUESTION, measured and NOT changed:
## Restarts 20 minutes apart NEVER accumulate. The budget is 3 restarts inside a 15-minute window, so
## each of these four restarts started a fresh window and the 30-minute pause was never armed.
## Total this session: 4 restarts, 0 pauses, 0 crash-loop events.
## CONSEQUENCE, for the operator to rule on (a design question, not a defect):
## A box whose controller dies every 20 minutes is restarted forever, quietly. The only signal is the
## info-level `controller_restarted_by_agent` event, which mails nobody. The brake catches a FAST loop
## (3 in 15 min) and is blind to a SLOW one. Options: (a) leave it - the box self-heals and the record
## is on the timeline; (b) add a second, longer counter (for example 5 restarts in 6 hours) that raises
## a warning-severity event; (c) raise the severity of the Nth restart in a day. This session measured
## only; it changed nothing.
## HONEST GAP in this run, caused by MY teardown timing, not by the product:
## The hub minted `controller_restarted_by_agent` for the kill-1 and kill-2 restarts (13:22:04 and
## 13:37:04 CEST in its log). The kill-3 restart happened at 13:56:08 and I destroyed the VM at
## ~13:57:30 - about 84 s later, inside the box's own report cycle. So the third restart never reached
## the hub. This is expected: the supervisor's record RIDES the host report, and a box that disappears
## before its next report takes the unsent record with it. It is recorded here so nobody later reads
## "2 events for 3 restarts" as a defect. A real box does not vanish 84 s after a restart.
@@ -0,0 +1,6 @@
## 2026-09-16T11:58:30Z LAYER 3 - delete the HOST record, KEEP the customer
delete forms on the host page:
chosen action:
form fields: NONE
## waiting for the host record to fall stale (delete is refused while ONLINE, by design)
2026-09-16T11:59:21Z status=ok deletable=False
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,37 @@
arch: amd64
cores: 3
features: nesting=1,keyctl=1
hookscript: local:snippets/felhom-guest-hook.sh
hostname: tester-1
memory: 4096
mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G
mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives
mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1
net0: name=eth0,bridge=vmbr0,hwaddr=BC:24:11:9A:B0:DC,ip=dhcp,type=veth
net1: name=eth1,bridge=vmbr9,hwaddr=BC:24:11:21:9E:D9,ip=169.254.253.2/30,type=veth
onboot: 1
ostype: debian
rootfs: local-lvm:vm-9201-disk-0,size=32G
swap: 512
unprivileged: 1
---
---
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
local dir active 14145416 7776604 5628460 54.98%
local-lvm lvmthin active 12378112 10723158 1654953 86.63%
---
---
VMID Status Lock Name
9201 running tester-1
felhom-controller Up About a minute (healthy)
nextcloud Up 50 minutes (healthy)
nextcloud-redis Up 50 minutes (healthy)
nextcloud-db Up 50 minutes (healthy)
vaultwarden Up 54 minutes (healthy)
privatebin Up 54 minutes (healthy)
paperless-webserver Up 54 minutes (healthy)
paperless-postgres Up 54 minutes (healthy)
paperless-redis Up 54 minutes (healthy)
filebrowser Up About an hour (healthy)
cloudflared Up About an hour
traefik Up About an hour
@@ -0,0 +1,91 @@
2026/09/16 11:56:08 [INFO] local-api: channel up (agent 169.254.253.1:8443) — guest 9201, 3 mount(s) visible
2026/09/16 11:56:08 [INFO] local-api: mount mp8 → /mnt/felhom-drives (storage=/mnt/felhom-drives, class=, backup=false)
2026/09/16 11:56:08 [INFO] local-api: mount mp9 → /etc/felhom-bootstrap (storage=/var/lib/felhom-agent/guests/9201/bootstrap, class=, backup=false)
2026/09/16 11:56:08 [INFO] local-api: mount mp0 → /var/lib/felhom (storage=local-lvm, class=, backup=true)
2026/09/16 11:56:08 [INFO] felhom-controller 0.243.0 starting (customer: tester-1, domain: enkicsifelhom.hu)
2026/09/16 11:56:08 [INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json
2026/09/16 11:56:08 [INFO] Encryption key loaded from /opt/docker/felhom-controller/data/encryption.key
2026/09/16 11:56:08 [INFO] [stacks] Using compose command: docker compose
2026/09/16 11:56:08 [INFO] [stacks] ScanStacks complete: 56 stacks found (4 deployed, 52 available)
2026/09/16 11:56:08 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:08 [INFO] [stacks] InjectMissingFields: processed 4 stacks
2026/09/16 11:56:08 [INFO] [stacks] Encryption migration: no stacks needed migration
2026/09/16 11:56:08 [INFO] [stacks] desired-state backfill: 0 app(s) recorded as running, 0 left unrecorded (state ambiguous — legacy boot behaviour retained)
2026/09/16 11:56:08 [INFO] [quiesce] loop started (poll 5m0s, max-quiesce 30m0s)
2026/09/16 11:56:08 [INFO] [stacks] installed-images backfill: 0 app(s) recorded, 4 already had a record, 0 left unrecorded (could not be observed completely — unknown, which renders nothing)
2026/09/16 11:56:08 [INFO] [stacks] pin adoption: 0 pinned, 4 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers)
2026/09/16 11:56:08 [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
2026/09/16 11:56:08 [INFO] [sync] Starting catalog sync
2026/09/16 11:56:08 [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main)
2026/09/16 11:56:08 [INFO] Metrics store opened at /opt/docker/felhom-controller/data/metrics.db
2026/09/16 11:56:08 [INFO] Metrics collector started (60s interval)
2026/09/16 11:56:08 [INFO] Notifier enabled (hub: https://hub.felhom.eu)
2026/09/16 11:56:08 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: status-refresh (every 10s)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: health-probes (every 10s)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: system-health (every 5m0s)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: ring-spill (every 30s)
2026/09/16 11:56:08 [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-17 02:30 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s)
2026/09/16 11:56:08 [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-17 03:30 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-17 04:15 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-17 05:10 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-17 06:00 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-17 05:30 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-17 04:00 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-17 03:30 CEST
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: hub-report (every 15m0s)
2026/09/16 11:56:08 [INFO] Hub reporting enabled (every 15m0s to https://hub.felhom.eu)
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/16 11:56:08 [INFO] ========== Startup Self-Test ==========
2026/09/16 11:56:08 [INFO] [PASS] Docker socket: reachable (v29.8.1)
2026/09/16 11:56:08 [INFO] [PASS] Stacks directory: /opt/docker/stacks
2026/09/16 11:56:08 [INFO] [PASS] Data directory: /opt/docker/felhom-controller/data (writable)
2026/09/16 11:56:08 [INFO] [PASS] System data path: /mnt/sys_drive
2026/09/16 11:56:08 [INFO] [PASS] Storage paths: 1 connected, 0 disconnected
2026/09/16 11:56:08 [INFO] [PASS] Git catalog: 53 app definitions found
2026/09/16 11:56:09 [INFO] [PASS] Hub connectivity: https://hub.felhom.eu reachable (HTTP 200)
2026/09/16 11:56:09 [INFO] [PASS] Metrics DB: 0.0 MB
2026/09/16 11:56:09 [INFO] ========================================
2026/09/16 11:56:09 [INFO] Self-test complete: 8 passed, 0 warnings, 0 failed
2026/09/16 11:56:09 [INFO] [scheduler] Starting scheduler with 18 jobs
2026/09/16 11:56:09 [INFO] [scheduler] Registered periodic job: geo-verify (every 6h0m0s)
2026/09/16 11:56:09 [INFO] Geo-restriction support enabled (CF API token configured)
2026/09/16 11:56:09 [INFO] [report] hub wait channel active (hold ≤240s)
2026/09/16 11:56:09 [INFO] [backup] Found 2 DB dump files across drives
2026/09/16 11:56:09 [INFO] [web] Auth: using password from settings.json
2026/09/16 11:56:09 [INFO] [scheduler] Registered periodic job: disk-health-check (every 1h0m0s)
2026/09/16 11:56:09 [INFO] [scheduler] Registered periodic job: agent-channel-health (every 1m0s)
2026/09/16 11:56:09 [INFO] Web UI listening on :8080
2026/09/16 11:56:09 [INFO] [infra] connected felhom-controller to traefik-public
2026/09/16 11:56:09 [INFO] [sync] Catalog sync complete
2026/09/16 11:56:09 [INFO] [sync] Initial sync: Sablonok naprakészek — nincs változás
2026/09/16 11:56:09 [INFO] [backup] Discovered 2 databases
2026/09/16 11:56:09 [INFO] [backup] Backup status cache refreshed
2026/09/16 11:56:09 [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s))
2026/09/16 11:56:09 [INFO] [monitor] Health check: status=ok
2026/09/16 11:56:13 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:14 [INFO] [report] Building system report
2026/09/16 11:56:14 [INFO] Event pushed: controller_started (info) — Controller elindult (0.243.0)
2026/09/16 11:56:14 [INFO] [monitor] Health check: status=ok
2026/09/16 11:56:14 [INFO] [report] Hub report pushed successfully (4703 bytes)
2026/09/16 11:56:14 [INFO] [settings] Settings saved
2026/09/16 11:56:14 [INFO] Startup hub report sent
2026/09/16 11:56:18 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:19 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:19 [INFO] Health probes: 3 ok (of 3 probed)
2026/09/16 11:56:23 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:23 [INFO] [bootrecon] boot window: fleet settled after 10s (3 identical samples 5s apart) — sweeping
2026/09/16 11:56:23 [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
2026/09/16 11:56:29 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:39 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:39 [INFO] [selfupdate] Current version 0.243.0 is up to date
2026/09/16 11:56:49 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:56:59 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:57:09 [INFO] [scheduler] Running job: agent-channel-health
2026/09/16 11:57:09 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
2026/09/16 11:57:09 [INFO] [scheduler] Job agent-channel-health completed (took 192ms)
2026/09/16 11:57:19 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
@@ -0,0 +1,84 @@
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
purging VM 334 from related configurations..
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
---
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40453376 27577580 10788680 68.17%
local-lvm lvmthin active 56487936 25272702 31215233 44.74%
nvme-scratch dir active 983379700 14205108 919147980 1.44%
---
ls: cannot access '/mnt/hdd_1/images/334': No such file or directory
## 2026-09-16T11:57:58Z LAYER 1 verification - is the machine really gone?
--- qm list:
--- any 334 disk left on hdd_1:
ls: cannot access '/mnt/hdd_1/images/334': No such file or directory
--- storage:
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40453376 27577584 10788676 68.17%
local-lvm lvmthin active 56487936 25272702 31215233 44.74%
nvme-scratch dir active 983379700 14205108 919147980 1.44%
## 2026-09-16T11:58:24Z FENCE CHECK - demo-hp's own guests and apps must be untouched
--- pct list (demo-hp own containers):
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
--- where 334 disks lived, before:
--- the before-file record of VM 334 disks:
334 tester1-drill-0243 running 8192 32.00 1581189
@@ -0,0 +1,46 @@
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
VMID NAME STATUS MEM(MB) BOOTDISK(GB) PID
334 tester1-drill-0243 running 8192 32.00 1581189
---
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40453376 27577572 10788688 68.17%
local-lvm lvmthin active 56487936 25272702 31215233 44.74%
nvme-scratch dir active 983379700 39080732 894272356 3.97%