drill 0.243.0: F9'' answers the supervisor budget question; machine torn down
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
Three more kills 20 minutes apart: recovered in 61 s / 41 s / 61 s, and none of them accumulated, because the window is 15 minutes. Four restarts, zero pauses. So the brake catches a FAST crash loop and is blind to a SLOW one — a controller dying every 20 minutes is restarted forever, and the only trace is an info event that mails nobody. Measured, not changed: the options are written into R-531 for the operator to rule on. Machine layer torn down: VM 334 purged with its disks, demo-hp's own containers 9201 and 9202 untouched. Evidence copied off the box first, token-leak control 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -11,3 +11,46 @@ Sep 16 13:13:37 tester1 felhom-agent[1128]: time=2026-09-16T13:13:37.415+02:00 l
|
||||
Sep 16 13:13:38 tester1 felhom-agent[1128]: time=2026-09-16T13:13:38.655+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
|
||||
RESTARTED lines in the whole journal so far: 2
|
||||
waiting 20 minutes before the next kill
|
||||
## kill 2 at 2026-09-16T11:34:05Z
|
||||
felhom-controller
|
||||
dashboard 200 again after 41 s
|
||||
Sep 16 13:32:37 tester1 felhom-agent[1128]: time=2026-09-16T13:32:37.415+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=80 guests_evaluated=1 controllers_not_running=0
|
||||
Sep 16 13:34:07 tester1 felhom-agent[1128]: time=2026-09-16T13:34:07.341+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
|
||||
Sep 16 13:34:37 tester1 felhom-agent[1128]: time=2026-09-16T13:34:37.379+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
|
||||
Sep 16 13:34:38 tester1 felhom-agent[1128]: time=2026-09-16T13:34:38.709+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
|
||||
RESTARTED lines in the whole journal so far: 3
|
||||
waiting 20 minutes before the next kill
|
||||
## kill 3 at 2026-09-16T11:55:09Z
|
||||
felhom-controller
|
||||
dashboard 200 again after 61 s
|
||||
Sep 16 13:52:37 tester1 felhom-agent[1128]: time=2026-09-16T13:52:37.361+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=120 guests_evaluated=1 controllers_not_running=0
|
||||
Sep 16 13:55:37 tester1 felhom-agent[1128]: time=2026-09-16T13:55:37.379+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
|
||||
Sep 16 13:56:07 tester1 felhom-agent[1128]: time=2026-09-16T13:56:07.300+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
|
||||
Sep 16 13:56:08 tester1 felhom-agent[1128]: time=2026-09-16T13:56:08.536+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
|
||||
RESTARTED lines in the whole journal so far: 4
|
||||
## F9'' finished at 2026-09-16T11:56:32Z
|
||||
## RESULT - F9'' (three kills at IDLE, 20 minutes apart)
|
||||
## kill 1 11:12:41Z -> dashboard 200 again after 61 s (restart #2 in the journal)
|
||||
## kill 2 11:34:05Z -> dashboard 200 again after 41 s (restart #3)
|
||||
## kill 3 11:55:09Z -> dashboard 200 again after 61 s (restart #4)
|
||||
## Every kill was seen on one sweep, confirmed on the next, and restarted - 30 to 90 s of dashboard
|
||||
## downtime each time, and the apps themselves never stopped (they do not depend on the controller).
|
||||
##
|
||||
## THE ANSWER TO THE BUDGET QUESTION, measured and NOT changed:
|
||||
## Restarts 20 minutes apart NEVER accumulate. The budget is 3 restarts inside a 15-minute window, so
|
||||
## each of these four restarts started a fresh window and the 30-minute pause was never armed.
|
||||
## Total this session: 4 restarts, 0 pauses, 0 crash-loop events.
|
||||
## CONSEQUENCE, for the operator to rule on (a design question, not a defect):
|
||||
## A box whose controller dies every 20 minutes is restarted forever, quietly. The only signal is the
|
||||
## info-level `controller_restarted_by_agent` event, which mails nobody. The brake catches a FAST loop
|
||||
## (3 in 15 min) and is blind to a SLOW one. Options: (a) leave it - the box self-heals and the record
|
||||
## is on the timeline; (b) add a second, longer counter (for example 5 restarts in 6 hours) that raises
|
||||
## a warning-severity event; (c) raise the severity of the Nth restart in a day. This session measured
|
||||
## only; it changed nothing.
|
||||
## HONEST GAP in this run, caused by MY teardown timing, not by the product:
|
||||
## The hub minted `controller_restarted_by_agent` for the kill-1 and kill-2 restarts (13:22:04 and
|
||||
## 13:37:04 CEST in its log). The kill-3 restart happened at 13:56:08 and I destroyed the VM at
|
||||
## ~13:57:30 - about 84 s later, inside the box's own report cycle. So the third restart never reached
|
||||
## the hub. This is expected: the supervisor's record RIDES the host report, and a box that disappears
|
||||
## before its next report takes the unsent record with it. It is recorded here so nobody later reads
|
||||
## "2 events for 3 restarts" as a defect. A real box does not vanish 84 s after a restart.
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
## 2026-09-16T11:58:30Z LAYER 3 - delete the HOST record, KEEP the customer
|
||||
delete forms on the host page:
|
||||
chosen action:
|
||||
form fields: NONE
|
||||
## waiting for the host record to fall stale (delete is refused while ONLINE, by design)
|
||||
2026-09-16T11:59:21Z status=ok deletable=False
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,37 @@
|
||||
arch: amd64
|
||||
cores: 3
|
||||
features: nesting=1,keyctl=1
|
||||
hookscript: local:snippets/felhom-guest-hook.sh
|
||||
hostname: tester-1
|
||||
memory: 4096
|
||||
mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G
|
||||
mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives
|
||||
mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1
|
||||
net0: name=eth0,bridge=vmbr0,hwaddr=BC:24:11:9A:B0:DC,ip=dhcp,type=veth
|
||||
net1: name=eth1,bridge=vmbr9,hwaddr=BC:24:11:21:9E:D9,ip=169.254.253.2/30,type=veth
|
||||
onboot: 1
|
||||
ostype: debian
|
||||
rootfs: local-lvm:vm-9201-disk-0,size=32G
|
||||
swap: 512
|
||||
unprivileged: 1
|
||||
---
|
||||
---
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
local dir active 14145416 7776604 5628460 54.98%
|
||||
local-lvm lvmthin active 12378112 10723158 1654953 86.63%
|
||||
---
|
||||
---
|
||||
VMID Status Lock Name
|
||||
9201 running tester-1
|
||||
felhom-controller Up About a minute (healthy)
|
||||
nextcloud Up 50 minutes (healthy)
|
||||
nextcloud-redis Up 50 minutes (healthy)
|
||||
nextcloud-db Up 50 minutes (healthy)
|
||||
vaultwarden Up 54 minutes (healthy)
|
||||
privatebin Up 54 minutes (healthy)
|
||||
paperless-webserver Up 54 minutes (healthy)
|
||||
paperless-postgres Up 54 minutes (healthy)
|
||||
paperless-redis Up 54 minutes (healthy)
|
||||
filebrowser Up About an hour (healthy)
|
||||
cloudflared Up About an hour
|
||||
traefik Up About an hour
|
||||
@@ -0,0 +1,91 @@
|
||||
2026/09/16 11:56:08 [INFO] local-api: channel up (agent 169.254.253.1:8443) — guest 9201, 3 mount(s) visible
|
||||
2026/09/16 11:56:08 [INFO] local-api: mount mp8 → /mnt/felhom-drives (storage=/mnt/felhom-drives, class=, backup=false)
|
||||
2026/09/16 11:56:08 [INFO] local-api: mount mp9 → /etc/felhom-bootstrap (storage=/var/lib/felhom-agent/guests/9201/bootstrap, class=, backup=false)
|
||||
2026/09/16 11:56:08 [INFO] local-api: mount mp0 → /var/lib/felhom (storage=local-lvm, class=, backup=true)
|
||||
2026/09/16 11:56:08 [INFO] felhom-controller 0.243.0 starting (customer: tester-1, domain: enkicsifelhom.hu)
|
||||
2026/09/16 11:56:08 [INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json
|
||||
2026/09/16 11:56:08 [INFO] Encryption key loaded from /opt/docker/felhom-controller/data/encryption.key
|
||||
2026/09/16 11:56:08 [INFO] [stacks] Using compose command: docker compose
|
||||
2026/09/16 11:56:08 [INFO] [stacks] ScanStacks complete: 56 stacks found (4 deployed, 52 available)
|
||||
2026/09/16 11:56:08 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:08 [INFO] [stacks] InjectMissingFields: processed 4 stacks
|
||||
2026/09/16 11:56:08 [INFO] [stacks] Encryption migration: no stacks needed migration
|
||||
2026/09/16 11:56:08 [INFO] [stacks] desired-state backfill: 0 app(s) recorded as running, 0 left unrecorded (state ambiguous — legacy boot behaviour retained)
|
||||
2026/09/16 11:56:08 [INFO] [quiesce] loop started (poll 5m0s, max-quiesce 30m0s)
|
||||
2026/09/16 11:56:08 [INFO] [stacks] installed-images backfill: 0 app(s) recorded, 4 already had a record, 0 left unrecorded (could not be observed completely — unknown, which renders nothing)
|
||||
2026/09/16 11:56:08 [INFO] [stacks] pin adoption: 0 pinned, 4 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers)
|
||||
2026/09/16 11:56:08 [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
|
||||
2026/09/16 11:56:08 [INFO] [sync] Starting catalog sync
|
||||
2026/09/16 11:56:08 [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main)
|
||||
2026/09/16 11:56:08 [INFO] Metrics store opened at /opt/docker/felhom-controller/data/metrics.db
|
||||
2026/09/16 11:56:08 [INFO] Metrics collector started (60s interval)
|
||||
2026/09/16 11:56:08 [INFO] Notifier enabled (hub: https://hub.felhom.eu)
|
||||
2026/09/16 11:56:08 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: status-refresh (every 10s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: health-probes (every 10s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: system-health (every 5m0s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: ring-spill (every 30s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-17 02:30 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-17 03:30 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-17 04:15 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-17 05:10 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-17 06:00 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-17 05:30 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-17 04:00 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-17 03:30 CEST
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: hub-report (every 15m0s)
|
||||
2026/09/16 11:56:08 [INFO] Hub reporting enabled (every 15m0s to https://hub.felhom.eu)
|
||||
2026/09/16 11:56:08 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
|
||||
2026/09/16 11:56:08 [INFO] ========== Startup Self-Test ==========
|
||||
2026/09/16 11:56:08 [INFO] [PASS] Docker socket: reachable (v29.8.1)
|
||||
2026/09/16 11:56:08 [INFO] [PASS] Stacks directory: /opt/docker/stacks
|
||||
2026/09/16 11:56:08 [INFO] [PASS] Data directory: /opt/docker/felhom-controller/data (writable)
|
||||
2026/09/16 11:56:08 [INFO] [PASS] System data path: /mnt/sys_drive
|
||||
2026/09/16 11:56:08 [INFO] [PASS] Storage paths: 1 connected, 0 disconnected
|
||||
2026/09/16 11:56:08 [INFO] [PASS] Git catalog: 53 app definitions found
|
||||
2026/09/16 11:56:09 [INFO] [PASS] Hub connectivity: https://hub.felhom.eu reachable (HTTP 200)
|
||||
2026/09/16 11:56:09 [INFO] [PASS] Metrics DB: 0.0 MB
|
||||
2026/09/16 11:56:09 [INFO] ========================================
|
||||
2026/09/16 11:56:09 [INFO] Self-test complete: 8 passed, 0 warnings, 0 failed
|
||||
2026/09/16 11:56:09 [INFO] [scheduler] Starting scheduler with 18 jobs
|
||||
2026/09/16 11:56:09 [INFO] [scheduler] Registered periodic job: geo-verify (every 6h0m0s)
|
||||
2026/09/16 11:56:09 [INFO] Geo-restriction support enabled (CF API token configured)
|
||||
2026/09/16 11:56:09 [INFO] [report] hub wait channel active (hold ≤240s)
|
||||
2026/09/16 11:56:09 [INFO] [backup] Found 2 DB dump files across drives
|
||||
2026/09/16 11:56:09 [INFO] [web] Auth: using password from settings.json
|
||||
2026/09/16 11:56:09 [INFO] [scheduler] Registered periodic job: disk-health-check (every 1h0m0s)
|
||||
2026/09/16 11:56:09 [INFO] [scheduler] Registered periodic job: agent-channel-health (every 1m0s)
|
||||
2026/09/16 11:56:09 [INFO] Web UI listening on :8080
|
||||
2026/09/16 11:56:09 [INFO] [infra] connected felhom-controller to traefik-public
|
||||
2026/09/16 11:56:09 [INFO] [sync] Catalog sync complete
|
||||
2026/09/16 11:56:09 [INFO] [sync] Initial sync: Sablonok naprakészek — nincs változás
|
||||
2026/09/16 11:56:09 [INFO] [backup] Discovered 2 databases
|
||||
2026/09/16 11:56:09 [INFO] [backup] Backup status cache refreshed
|
||||
2026/09/16 11:56:09 [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s))
|
||||
2026/09/16 11:56:09 [INFO] [monitor] Health check: status=ok
|
||||
2026/09/16 11:56:13 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:14 [INFO] [report] Building system report
|
||||
2026/09/16 11:56:14 [INFO] Event pushed: controller_started (info) — Controller elindult (0.243.0)
|
||||
2026/09/16 11:56:14 [INFO] [monitor] Health check: status=ok
|
||||
2026/09/16 11:56:14 [INFO] [report] Hub report pushed successfully (4703 bytes)
|
||||
2026/09/16 11:56:14 [INFO] [settings] Settings saved
|
||||
2026/09/16 11:56:14 [INFO] Startup hub report sent
|
||||
2026/09/16 11:56:18 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:19 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:19 [INFO] Health probes: 3 ok (of 3 probed)
|
||||
2026/09/16 11:56:23 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:23 [INFO] [bootrecon] boot window: fleet settled after 10s (3 identical samples 5s apart) — sweeping
|
||||
2026/09/16 11:56:23 [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
|
||||
2026/09/16 11:56:29 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:39 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:39 [INFO] [selfupdate] Current version 0.243.0 is up to date
|
||||
2026/09/16 11:56:49 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:56:59 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:57:09 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/09/16 11:57:09 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
2026/09/16 11:57:09 [INFO] [scheduler] Job agent-channel-health completed (took 192ms)
|
||||
2026/09/16 11:57:19 [INFO] [stacks] Status refresh: 11 containers across 56 stacks
|
||||
@@ -0,0 +1,84 @@
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
purging VM 334 from related configurations..
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
---
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
felhom-pbs pbs active 0 0 0 0.00%
|
||||
local dir active 40453376 27577580 10788680 68.17%
|
||||
local-lvm lvmthin active 56487936 25272702 31215233 44.74%
|
||||
nvme-scratch dir active 983379700 14205108 919147980 1.44%
|
||||
---
|
||||
ls: cannot access '/mnt/hdd_1/images/334': No such file or directory
|
||||
## 2026-09-16T11:57:58Z LAYER 1 verification - is the machine really gone?
|
||||
--- qm list:
|
||||
--- any 334 disk left on hdd_1:
|
||||
ls: cannot access '/mnt/hdd_1/images/334': No such file or directory
|
||||
--- storage:
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
felhom-pbs pbs active 0 0 0 0.00%
|
||||
local dir active 40453376 27577584 10788676 68.17%
|
||||
local-lvm lvmthin active 56487936 25272702 31215233 44.74%
|
||||
nvme-scratch dir active 983379700 14205108 919147980 1.44%
|
||||
## 2026-09-16T11:58:24Z FENCE CHECK - demo-hp's own guests and apps must be untouched
|
||||
--- pct list (demo-hp own containers):
|
||||
VMID Status Lock Name
|
||||
9201 running demo-hp
|
||||
9202 running demo-hp-scratch
|
||||
--- where 334 disks lived, before:
|
||||
--- the before-file record of VM 334 disks:
|
||||
334 tester1-drill-0243 running 8192 32.00 1581189
|
||||
@@ -0,0 +1,46 @@
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
VMID NAME STATUS MEM(MB) BOOTDISK(GB) PID
|
||||
334 tester1-drill-0243 running 8192 32.00 1581189
|
||||
---
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
felhom-pbs pbs active 0 0 0 0.00%
|
||||
local dir active 40453376 27577572 10788688 68.17%
|
||||
local-lvm lvmthin active 56487936 25272702 31215233 44.74%
|
||||
nvme-scratch dir active 983379700 39080732 894272356 3.97%
|
||||
@@ -728,7 +728,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** |
|
||||
| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** |
|
||||
| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** |
|
||||
| **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
|
||||
Reference in New Issue
Block a user