Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
15 KiB
REPORT — burn-down round 2: the operator's answer recorded, R-887 re-diagnosed, small rows fixed WITH releases — 2026-10-05 (late night)
| Part | Result |
|---|---|
| A — rulings, then R-887 | done — rulings commit 301fe45 (count after: 249); R-887 re-diagnosed from the logs (the restart idea refuted; the mechanism then SEEN in Gitea's own log), dated check 2026-10-12 |
| B — R-124, the small rows, the 23 unchecked | done — R-124 fixed (agent v0.147.0 + runbook); 50 more rows fixed and closed with tests and red-proofs; the 23 checked from source (1 duplicate closed, facts added to 8 rows, the rest left as they need a live box or a decision) |
| B.4 — releases, delivered the normal way | done — hub v0.137.0 deployed; agent v0.147.0 released + signed jobs (binary and bundle) to demo-hp, demo-felhom, Tester 1; controller v0.297.0 + golden 0.297.0 baked, vouched, floors raised, all three boxes on 0.297.0; catalog pushed |
| C — numbers and record | done — STATUS shows 199 and asks nothing about the closed list |
| Rows before | Rows after | Opened | Closed |
|---|---|---|---|
| 292 | 199 | 1 (R-888) | 94 (43 accepted by the operator + 51 fixed/merged) |
Counted by register_shape_gate.py's method. Target ≤ 220: met.
Baselines (re-verified at the start)
felhom.eu e8c56c440a (hub v0.136.0) · controller 114ff2761a (v0.296.0) · agent d83316326e (v0.146.1) · catalog
29ac711d26 · golden 0.296.0 · register 292. The agent clone had a stray scripts/__pycache__/ from round 1 — removed.
Part A — the rulings commit and R-887
301fe45: 43 rows closed as „accepted by the operator, 2026-10-05", each with its one-line reason from the list; R-124 and R-698 kept (R-698 owner → operator); R-831/R-870 carry the not-rotated rulings; R-887 records the screenshot (one runner, ID 2, online). STATUS: the list and the rotate/runners requests removed. Count after: 249. CI run 1363 success.- R-887, from the logs: the runner's last restart was 13:24:42Z; the lost attempts started 15:05–15:46Z — not a
restart. Four lost attempts (not two): each without a runner
taskline, each failed at a :38-second mark 10–13 min after assignment. Gitea's log for that hour had rotated. Then it happened again at 17:15Z with the log intact:slow POST …/RunnerService/FetchTask for 10.42.0.42, elapsed 3192ms→context canceled→ 17:28:39clear_tasks.go … stopTasks() … task 1371— the runner abandoned its fetch after Gitea assigned the task; Gitea's zombie stop failed it. Load at that minute: an outside crawler on public commit pages, and this session's CI waiter (15-page job listings at 13–31 s each). The waiter now makes ONEruns?head_sha=call a minute. A lost run re-runs withPOST …/actions/runs/<id>/rerun(used twice: controller run 1357 → success; catalog run 1368 → success). Dated check 2026-10-12 in DUE-CHECKS. The fix on DooPlex (runner fetch timeout, crawler) is the operator's.
Part B — fixes by repo
agent v0.147.0 (f1b9b41, CI 1365; tag v0.147.0; binary sha256 642c4d19…, bundle 326527d0…, verified by
download; CHANGELOG 208fac8, CI 1367): R-124, R-118, R-269, R-317 — red-proofs audits/burndown2-2026-10-05/r124-red-proof.txt,
agent-red-proofs.txt. Delivery: vouched (agent 0.147.0, golden 0.296.0 first), signed agent_update ×3, then
agent_config_update ×3 (felhom-op-1, ttl 45 m); hub System page: demo-hp, demo-felhom, Tester 1 — agent 0.147.0,
root files 0.147.0 (delivery/). Tester 2 offline — nothing sent.
hub v0.137.0 (557629d, CI 1369; manifest 81d04a6; CI 1370): R-277, R-581, R-600, R-544, R-855, R-134, R-92,
R-292, R-599, R-725, R-728, R-208 (hub half) — red-proofs felhom-eu-red-proofs.txt. Deployed: ArgoCD Synced/Healthy
at d75ad0f, image felhom-hub:0.137.0, felhom-hub 0.137.0 starting, healthz 200; R-855's new line seen live
(„after 2 healthy ring-0 night(s)"). (The build ran while a helper was still appending to an audit text file outside
hub/ — the image is the committed hub/ tree; said here because the clean-tree gate is literal.)
felhom.eu gates/tools/docs (same commits): R-819 (stands gate), R-857, R-555, R-364 (hu_grep.py + REUSE.md),
R-587, R-571, R-129 (demo-hp authenticates with DooPlex's own key — corrected everywhere it said „no key"), R-124 runbook.
New script tests pass under a BusyBox + bash + python3 + git PATH (the CI runner's tools): 19/19.
controller v0.297.0 (1453cfc; CI run 1371 FAILED — the new gofmt gate was INCONCLUSIVE on the Go-less runner;
fixed in 6f1ba1f, CI 1372 success): R-591, R-568, R-567, R-363, R-547, R-10, R-552, R-251, R-104, R-619, R-362, R-675,
R-256, R-257, R-240, R-365, R-425, R-565, R-564, R-603, R-454, R-208, R-457 (swept, nothing left) + two twins found and
fixed on the way (the top-bar countdown at 0 days; nine more shared references in deepCopyStack). Red-proofs
controller-red-proofs.txt (two first attempts that did not convict are marked, with valid re-runs). MinAgent 0.131.0.
Image felhom-controller:0.297.0. Golden 0.297.0 baked per RUNBOOK §4.0–4.1 (documentation/tests/golden-0.297.0-2026-10-05/:
all pass markers, round trip sha 8cebc42e…, token leak 0 with a working control, teardown to virgin). Vouched
(agent 0.147.0, golden 0.297.0, min_agent 0.131.0) and floors 0.297.0 for demo-hp, demo-felhom, tester-1.
Delivered: demo-hp and demo-felhom felhom-controller:0.297.0 … (healthy); Tester 1 reports Controller 0.297.0
(„Controller frissítve: 0.296.0 → 0.297.0").
catalog (4828dc7; CI run 1368 lost by R-887, re-run success): R-593, R-760, R-594, R-605, R-781, R-806 (scheme half;
row narrowed), plus a stale runner test (expected 11 gates, 12 exist) and a test that never ran (outside its class) —
fixed, not filed.
Not done, and why: R-469 and R-605's exit-code line in the catalog's CLAUDE.md — the permission check refused
the instruction-file edit; the operator is asked (rule 5). R-126 needs an operator choice. R-325 needs a same-step
felhom.eu gate change (left). R-377 (CONTEXT headings) not attempted. Installer rows (R-179, R-180, R-275, R-276, R-306,
R-130, R-310, R-881), R-136 (logs every operator out), R-502 (Docker in CI), R-798 (a live app definition) and the
larger controller rows (R-492, R-569, R-575, R-615, R-616, R-498, R-718) were left on purpose.
Opened: R-888 — two report fields the hub never reads (a decision). Seen, not a row: Tester 1's crash guard reads
TRIPPED since 07:57Z — the morning's two deliberate test crashes; it re-arms by itself after 24 h (runbooks/crash-guard.md).
The 23 rows the first burn-down could not check
Checked from source by a read-only agent (audits/burndown2-2026-10-05/unchecked-results.jsonl). Closed: R-350 (duplicate
of R-132, facts merged). Facts added to the open rows R-607, R-883, R-886, R-884, R-756, R-91, R-338, R-488. The three
„not worth it" ones are on STATUS for the operator. The rest need a live box reading (the settle command is in the table).
| Row | Group | Evidence / how to settle (abridged) |
|---|---|---|
| R-76 | UNCHECKABLE-FROM-SOURCE | Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 FileBrowserImage = "gtstef/filebrowser:1.5.6-stable". The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --gre |
| R-91 | UNCHECKABLE-FROM-SOURCE | Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of delet |
| R-209a | UNCHECKABLE-FROM-SOURCE | Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test. — settle: uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h / |
| R-337 | NOT-WORTH-IT | The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 b, err : |
| R-375 | NOT-WORTH-IT | The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is pvesm status showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 if st == nil // st.Type == "pbs" // st.Avail <= 0 { return true, "" } (space preflight sk |
| R-488 | STILL-TRUE-SMALL | The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded interval := 5 * time.Second and time.Sleep(3 * time.Second) // initial settling time, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go: |
| R-504 | UNCHECKABLE-FROM-SOURCE | Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html. — settle: curl -sI https://iso.felhom.eu/ / head -1 |
| R-644 | UNCHECKABLE-FROM-SOURCE | Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs --deployment-password; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches |
| R-814 | UNCHECKABLE-FROM-SOURCE | Hetzner account state; nothing in source records a deletion. — settle: Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status |
| R-815 | UNCHECKABLE-FROM-SOURCE | PBS server-side state on ep0; no GC completion record in the docs (grep). — settle: ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 / grep -i garbage' |
| R-884 | UNCHECKABLE-FROM-SOURCE | Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml prom/prometheus:v3.14.0 -> v3.15.0 (monitoring.yaml:419), and the monitoring Application has no automated syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an un |
| R-132 | UNCHECKABLE-FROM-SOURCE | Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key operator_password_hash, setSetting :2200-2207 writes updated_at = datetime('now'). No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31 |
| R-298 | UNCHECKABLE-FROM-SOURCE | Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 if(d.role==='user-data'){ else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/ |
| R-338 | UNCHECKABLE-FROM-SOURCE | nodes.md:86-88 still claims demo-hp is on the R-50 island (local_api on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent con |
| R-350 | DUPLICATE | of R-132 — Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. |
| R-542 | NOT-WORTH-IT | Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 initialize = append(initialize, c) // every unclaimed disk can be initialized; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkf |
| R-607 | STILL-TRUE-SMALL | Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY if len(newApps) > 0 // len(updated) > 0, and :257-258 says 'nincs változás' when both are empty. updated counts stack-dir copies only (copyTemplates, :447 updated = append(updated, appName) a |
| R-683 | UNCHECKABLE-FROM-SOURCE | Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power- |
| R-756 | UNCHECKABLE-FROM-SOURCE | Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH |
| R-862 | UNCHECKABLE-FROM-SOURCE | Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83). — settle: Ask the operator wheth |
| R-882 | UNCHECKABLE-FROM-SOURCE | Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it). — settle: sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p Act |
| R-883 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/rec |
| R-886 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root nobody image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts |
Teardown
Machines: drill VM — build guest destroyed, token/scripts/log shredded, powered off, disk back on virgin. Boxes: only the
normal deliveries above. Hub: only the deploy, the vouch and the floors. Scratch secrets (hub password file, hub key file,
signed envelopes) are shredded at the end of the session.