Files
felhom.eu/REPORT-burndown2-2026-10-05.md
T

15 KiB
Raw Blame History

REPORT — burn-down round 2: the operator's answer recorded, R-887 re-diagnosed, small rows fixed WITH releases — 2026-10-05 (late night)

Part Result
A — rulings, then R-887 done — rulings commit 301fe45 (count after: 249); R-887 re-diagnosed from the logs (the restart idea refuted; the mechanism then SEEN in Gitea's own log), dated check 2026-10-12
B — R-124, the small rows, the 23 unchecked done — R-124 fixed (agent v0.147.0 + runbook); 50 more rows fixed and closed with tests and red-proofs; the 23 checked from source (1 duplicate closed, facts added to 8 rows, the rest left as they need a live box or a decision)
B.4 — releases, delivered the normal way done — hub v0.137.0 deployed; agent v0.147.0 released + signed jobs (binary and bundle) to demo-hp, demo-felhom, Tester 1; controller v0.297.0 + golden 0.297.0 baked, vouched, floors raised, all three boxes on 0.297.0; catalog pushed
C — numbers and record done — STATUS shows 199 and asks nothing about the closed list
Rows before Rows after Opened Closed
292 199 1 (R-888) 94 (43 accepted by the operator + 51 fixed/merged)

Counted by register_shape_gate.py's method. Target ≤ 220: met.

Baselines (re-verified at the start)

felhom.eu e8c56c440a (hub v0.136.0) · controller 114ff2761a (v0.296.0) · agent d83316326e (v0.146.1) · catalog 29ac711d26 · golden 0.296.0 · register 292. The agent clone had a stray scripts/__pycache__/ from round 1 — removed.

Part A — the rulings commit and R-887

  • 301fe45: 43 rows closed as „accepted by the operator, 2026-10-05", each with its one-line reason from the list; R-124 and R-698 kept (R-698 owner → operator); R-831/R-870 carry the not-rotated rulings; R-887 records the screenshot (one runner, ID 2, online). STATUS: the list and the rotate/runners requests removed. Count after: 249. CI run 1363 success.
  • R-887, from the logs: the runner's last restart was 13:24:42Z; the lost attempts started 15:05–15:46Z — not a restart. Four lost attempts (not two): each without a runner task line, each failed at a :38-second mark 10–13 min after assignment. Gitea's log for that hour had rotated. Then it happened again at 17:15Z with the log intact: slow POST …/RunnerService/FetchTask for 10.42.0.42, elapsed 3192ms → context canceled → 17:28:39 clear_tasks.go … stopTasks() … task 1371 — the runner abandoned its fetch after Gitea assigned the task; Gitea's zombie stop failed it. Load at that minute: an outside crawler on public commit pages, and this session's CI waiter (15-page job listings at 13–31 s each). The waiter now makes ONE runs?head_sha= call a minute. A lost run re-runs with POST …/actions/runs/<id>/rerun (used twice: controller run 1357 → success; catalog run 1368 → success). Dated check 2026-10-12 in DUE-CHECKS. The fix on DooPlex (runner fetch timeout, crawler) is the operator's.

Part B — fixes by repo

agent v0.147.0 (f1b9b41, CI 1365; tag v0.147.0; binary sha256 642c4d19…, bundle 326527d0…, verified by download; CHANGELOG 208fac8, CI 1367): R-124, R-118, R-269, R-317 — red-proofs audits/burndown2-2026-10-05/r124-red-proof.txt, agent-red-proofs.txt. Delivery: vouched (agent 0.147.0, golden 0.296.0 first), signed agent_update ×3, then agent_config_update ×3 (felhom-op-1, ttl 45 m); hub System page: demo-hp, demo-felhom, Tester 1 — agent 0.147.0, root files 0.147.0 (delivery/). Tester 2 offline — nothing sent.

hub v0.137.0 (557629d, CI 1369; manifest 81d04a6; CI 1370): R-277, R-581, R-600, R-544, R-855, R-134, R-92, R-292, R-599, R-725, R-728, R-208 (hub half) — red-proofs felhom-eu-red-proofs.txt. Deployed: ArgoCD Synced/Healthy at d75ad0f, image felhom-hub:0.137.0, felhom-hub 0.137.0 starting, healthz 200; R-855's new line seen live („after 2 healthy ring-0 night(s)"). (The build ran while a helper was still appending to an audit text file outside hub/ — the image is the committed hub/ tree; said here because the clean-tree gate is literal.)

felhom.eu gates/tools/docs (same commits): R-819 (stands gate), R-857, R-555, R-364 (hu_grep.py + REUSE.md), R-587, R-571, R-129 (demo-hp authenticates with DooPlex's own key — corrected everywhere it said „no key"), R-124 runbook. New script tests pass under a BusyBox + bash + python3 + git PATH (the CI runner's tools): 19/19.

controller v0.297.0 (1453cfc; CI run 1371 FAILED — the new gofmt gate was INCONCLUSIVE on the Go-less runner; fixed in 6f1ba1f, CI 1372 success): R-591, R-568, R-567, R-363, R-547, R-10, R-552, R-251, R-104, R-619, R-362, R-675, R-256, R-257, R-240, R-365, R-425, R-565, R-564, R-603, R-454, R-208, R-457 (swept, nothing left) + two twins found and fixed on the way (the top-bar countdown at 0 days; nine more shared references in deepCopyStack). Red-proofs controller-red-proofs.txt (two first attempts that did not convict are marked, with valid re-runs). MinAgent 0.131.0. Image felhom-controller:0.297.0. Golden 0.297.0 baked per RUNBOOK §4.0–4.1 (documentation/tests/golden-0.297.0-2026-10-05/: all pass markers, round trip sha 8cebc42e…, token leak 0 with a working control, teardown to virgin). Vouched (agent 0.147.0, golden 0.297.0, min_agent 0.131.0) and floors 0.297.0 for demo-hp, demo-felhom, tester-1. Delivered: demo-hp and demo-felhom felhom-controller:0.297.0 … (healthy); Tester 1 reports Controller 0.297.0 („Controller frissítve: 0.296.0 → 0.297.0").

catalog (4828dc7; CI run 1368 lost by R-887, re-run success): R-593, R-760, R-594, R-605, R-781, R-806 (scheme half; row narrowed), plus a stale runner test (expected 11 gates, 12 exist) and a test that never ran (outside its class) — fixed, not filed.

Not done, and why: R-469 and R-605's exit-code line in the catalog's CLAUDE.md — the permission check refused the instruction-file edit; the operator is asked (rule 5). R-126 needs an operator choice. R-325 needs a same-step felhom.eu gate change (left). R-377 (CONTEXT headings) not attempted. Installer rows (R-179, R-180, R-275, R-276, R-306, R-130, R-310, R-881), R-136 (logs every operator out), R-502 (Docker in CI), R-798 (a live app definition) and the larger controller rows (R-492, R-569, R-575, R-615, R-616, R-498, R-718) were left on purpose.

Opened: R-888 — two report fields the hub never reads (a decision). Seen, not a row: Tester 1's crash guard reads TRIPPED since 07:57Z — the morning's two deliberate test crashes; it re-arms by itself after 24 h (runbooks/crash-guard.md).

The 23 rows the first burn-down could not check

Checked from source by a read-only agent (audits/burndown2-2026-10-05/unchecked-results.jsonl). Closed: R-350 (duplicate of R-132, facts merged). Facts added to the open rows R-607, R-883, R-886, R-884, R-756, R-91, R-338, R-488. The three „not worth it" ones are on STATUS for the operator. The rest need a live box reading (the settle command is in the table).

Row Group Evidence / how to settle (abridged)
R-76 UNCHECKABLE-FROM-SOURCE Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 FileBrowserImage = "gtstef/filebrowser:1.5.6-stable". The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --gre
R-91 UNCHECKABLE-FROM-SOURCE Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of delet
R-209a UNCHECKABLE-FROM-SOURCE Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test. — settle: uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h /
R-337 NOT-WORTH-IT The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 b, err :
R-375 NOT-WORTH-IT The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is pvesm status showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 if st == nil // st.Type == "pbs" // st.Avail <= 0 { return true, "" } (space preflight sk
R-488 STILL-TRUE-SMALL The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded interval := 5 * time.Second and time.Sleep(3 * time.Second) // initial settling time, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:
R-504 UNCHECKABLE-FROM-SOURCE Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html. — settle: curl -sI https://iso.felhom.eu/ / head -1
R-644 UNCHECKABLE-FROM-SOURCE Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs --deployment-password; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches
R-814 UNCHECKABLE-FROM-SOURCE Hetzner account state; nothing in source records a deletion. — settle: Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status
R-815 UNCHECKABLE-FROM-SOURCE PBS server-side state on ep0; no GC completion record in the docs (grep). — settle: ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 / grep -i garbage'
R-884 UNCHECKABLE-FROM-SOURCE Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml prom/prometheus:v3.14.0 -> v3.15.0 (monitoring.yaml:419), and the monitoring Application has no automated syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an un
R-132 UNCHECKABLE-FROM-SOURCE Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key operator_password_hash, setSetting :2200-2207 writes updated_at = datetime('now'). No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31
R-298 UNCHECKABLE-FROM-SOURCE Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 if(d.role==='user-data'){ else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/
R-338 UNCHECKABLE-FROM-SOURCE nodes.md:86-88 still claims demo-hp is on the R-50 island (local_api on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent con
R-350 DUPLICATE of R-132 — Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence.
R-542 NOT-WORTH-IT Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 initialize = append(initialize, c) // every unclaimed disk can be initialized; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkf
R-607 STILL-TRUE-SMALL Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY if len(newApps) > 0 // len(updated) > 0, and :257-258 says 'nincs változás' when both are empty. updated counts stack-dir copies only (copyTemplates, :447 updated = append(updated, appName) a
R-683 UNCHECKABLE-FROM-SOURCE Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power-
R-756 UNCHECKABLE-FROM-SOURCE Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH
R-862 UNCHECKABLE-FROM-SOURCE Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83). — settle: Ask the operator wheth
R-882 UNCHECKABLE-FROM-SOURCE Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it). — settle: sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p Act
R-883 STILL-TRUE-SMALL homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/rec
R-886 STILL-TRUE-SMALL homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root nobody image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts

Teardown

Machines: drill VM — build guest destroyed, token/scripts/log shredded, powered off, disk back on virgin. Boxes: only the normal deliveries above. Hub: only the deploy, the vouch and the floors. Scratch secrets (hub password file, hub key file, signed envelopes) are shredded at the end of the session.