burn-down round 2: controller v0.297.0 rows closed (23), golden 0.297.0 evidence, delivery evidence, 23-row unchecked table, STATUS/CONTEXT/REPORT (292 -> 199; 1 opened, 94 closed)
gates / gates (push) Successful in 1m59s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 20:41:50 +02:00
parent d75ad0fdf3
commit 30650cad6e
13 changed files with 640 additions and 37 deletions
+10
View File
@@ -16,6 +16,16 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-05 (late night) — burn-down round 2 (releases).** Register 292 → 199 (1 opened: R-888; 94 closed: 43
> accepted by the operator 18:23, 51 fixed). Releases: hub v0.137.0 (`557629d`, deployed), agent v0.147.0 (tag, sha
> `642c4d19…`, bundle `326527d0…`, signed jobs to 3 boxes), controller v0.297.0 (`1453cfc` + `6f1ba1f`), golden 0.297.0
> (`8cebc42e…`, vouched with agent 0.147.0 / min_agent 0.131.0; floors 0.297.0 for demo-hp, demo-felhom, tester-1),
> catalog `4828dc7`. R-124: recipe root namespace = `""` (+ runbook). R-887 mechanism from Gitea's log: a FetchTask the
> runner abandons after assignment → zombie stop after ~10 min; load = an outside crawler + the session's own 15-page CI
> polling (now one `runs?head_sha=` call per minute). New gates: `stands` (felhom.eu), `gofmt` (controller; NOT
> CHECKED out loud on the Go-less runner). R-469 not done: the permission check refused the catalog CLAUDE.md edit.
> Report: `REPORT-burndown2-2026-10-05.md`.
> **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 292 (1 opened — R-887 CI runner fault — 45
> closed): 24 fixed by later work + 2 duplicates (each re-checked; `audits/burndown-2026-10-05/partA-table.md` holds all
> 317 P3/P4 verdicts), 19 small fixes with tests/red-proofs (catalog `29ac711`, agent `d833163`, controller `114ff27`,
+115
View File
@@ -0,0 +1,115 @@
# REPORT — burn-down round 2: the operator's answer recorded, R-887 re-diagnosed, small rows fixed WITH releases — 2026-10-05 (late night)
| Part | Result |
|---|---|
| **A** — rulings, then R-887 | **done** — rulings commit `301fe45` (count after: **249**); R-887 re-diagnosed from the logs (the restart idea refuted; the mechanism then SEEN in Gitea's own log), dated check 2026-10-12 |
| **B** — R-124, the small rows, the 23 unchecked | **done** — R-124 fixed (agent v0.147.0 + runbook); 50 more rows fixed and closed with tests and red-proofs; the 23 checked from source (1 duplicate closed, facts added to 8 rows, the rest left as they need a live box or a decision) |
| **B.4** — releases, delivered the normal way | **done** — hub v0.137.0 deployed; agent v0.147.0 released + signed jobs (binary and bundle) to demo-hp, demo-felhom, Tester 1; controller v0.297.0 + golden 0.297.0 baked, vouched, floors raised, all three boxes on 0.297.0; catalog pushed |
| **C** — numbers and record | **done** — STATUS shows 199 and asks nothing about the closed list |
| Rows before | Rows after | Opened | Closed |
|---|---|---|---|
| **292** | **199** | **1** (R-888) | **94** (43 accepted by the operator + 51 fixed/merged) |
Counted by `register_shape_gate.py`'s method. Target ≤ 220: met.
## Baselines (re-verified at the start)
felhom.eu `e8c56c440a` (hub v0.136.0) · controller `114ff2761a` (v0.296.0) · agent `d83316326e` (v0.146.1) · catalog
`29ac711d26` · golden 0.296.0 · register 292. The agent clone had a stray `scripts/__pycache__/` from round 1 — removed.
## Part A — the rulings commit and R-887
- `301fe45`: 43 rows closed as „accepted by the operator, 2026-10-05", each with its one-line reason from the list;
R-124 and R-698 kept (R-698 owner → operator); R-831/R-870 carry the not-rotated rulings; R-887 records the screenshot
(one runner, ID 2, online). STATUS: the list and the rotate/runners requests removed. **Count after: 249.** CI run
1363 success.
- **R-887, from the logs:** the runner's last restart was 13:24:42Z; the lost attempts started 15:05–15:46Z — **not a
restart**. Four lost attempts (not two): each without a runner `task` line, each failed at a :38-second mark 10–13 min
after assignment. Gitea's log for that hour had rotated. **Then it happened again at 17:15Z with the log intact:**
`slow POST …/RunnerService/FetchTask for 10.42.0.42, elapsed 3192ms` → `context canceled` → 17:28:39
`clear_tasks.go … stopTasks() … task 1371` — the runner abandoned its fetch after Gitea assigned the task; Gitea's
zombie stop failed it. Load at that minute: an outside crawler on public commit pages, and this session's CI waiter
(15-page job listings at 13–31 s each). The waiter now makes ONE `runs?head_sha=` call a minute. A lost run re-runs
with `POST …/actions/runs/<id>/rerun` (used twice: controller run 1357 → success; catalog run 1368 → success).
**Dated check 2026-10-12** in DUE-CHECKS. The fix on DooPlex (runner fetch timeout, crawler) is the operator's.
## Part B — fixes by repo
**agent v0.147.0** (`f1b9b41`, CI 1365; tag `v0.147.0`; binary sha256 `642c4d19…`, bundle `326527d0…`, verified by
download; CHANGELOG `208fac8`, CI 1367): R-124, R-118, R-269, R-317 — red-proofs `audits/burndown2-2026-10-05/r124-red-proof.txt`,
`agent-red-proofs.txt`. **Delivery:** vouched (agent 0.147.0, golden 0.296.0 first), signed `agent_update` ×3, then
`agent_config_update` ×3 (felhom-op-1, ttl 45 m); hub System page: demo-hp, demo-felhom, Tester 1 — agent 0.147.0,
root files 0.147.0 (`delivery/`). Tester 2 offline — nothing sent.
**hub v0.137.0** (`557629d`, CI 1369; manifest `81d04a6`; CI 1370): R-277, R-581, R-600, R-544, R-855, R-134, R-92,
R-292, R-599, R-725, R-728, R-208 (hub half) — red-proofs `felhom-eu-red-proofs.txt`. **Deployed:** ArgoCD Synced/Healthy
at `d75ad0f`, image `felhom-hub:0.137.0`, `felhom-hub 0.137.0 starting`, healthz 200; R-855's new line seen live
(„after 2 healthy ring-0 night(s)"). (The build ran while a helper was still appending to an audit text file outside
`hub/` — the image is the committed `hub/` tree; said here because the clean-tree gate is literal.)
**felhom.eu gates/tools/docs** (same commits): R-819 (`stands` gate), R-857, R-555, R-364 (`hu_grep.py` + REUSE.md),
R-587, R-571, R-129 (demo-hp authenticates with DooPlex's own key — corrected everywhere it said „no key"), R-124 runbook.
New script tests pass under a BusyBox + bash + python3 + git PATH (the CI runner's tools): 19/19.
**controller v0.297.0** (`1453cfc`; CI run 1371 **FAILED** — the new gofmt gate was INCONCLUSIVE on the Go-less runner;
fixed in `6f1ba1f`, CI 1372 success): R-591, R-568, R-567, R-363, R-547, R-10, R-552, R-251, R-104, R-619, R-362, R-675,
R-256, R-257, R-240, R-365, R-425, R-565, R-564, R-603, R-454, R-208, R-457 (swept, nothing left) + two twins found and
fixed on the way (the top-bar countdown at 0 days; nine more shared references in `deepCopyStack`). Red-proofs
`controller-red-proofs.txt` (two first attempts that did not convict are marked, with valid re-runs). **MinAgent 0.131.0.**
**Image** `felhom-controller:0.297.0`. **Golden 0.297.0** baked per RUNBOOK §4.0–4.1 (`documentation/tests/golden-0.297.0-2026-10-05/`:
all pass markers, round trip sha `8cebc42e…`, token leak 0 with a working control, teardown to `virgin`). **Vouched**
(agent 0.147.0, golden 0.297.0, min_agent 0.131.0) and **floors** 0.297.0 for demo-hp, demo-felhom, tester-1.
**Delivered:** demo-hp and demo-felhom `felhom-controller:0.297.0 … (healthy)`; Tester 1 reports Controller 0.297.0
(„Controller frissítve: 0.296.0 → 0.297.0").
**catalog** (`4828dc7`; CI run 1368 lost by R-887, re-run success): R-593, R-760, R-594, R-605, R-781, R-806 (scheme half;
row narrowed), plus a stale runner test (expected 11 gates, 12 exist) and a test that never ran (outside its class) —
fixed, not filed.
**Not done, and why:** R-469 and R-605's exit-code line in the catalog's `CLAUDE.md` — **the permission check refused
the instruction-file edit**; the operator is asked (rule 5). R-126 needs an operator choice. R-325 needs a same-step
felhom.eu gate change (left). R-377 (CONTEXT headings) not attempted. Installer rows (R-179, R-180, R-275, R-276, R-306,
R-130, R-310, R-881), R-136 (logs every operator out), R-502 (Docker in CI), R-798 (a live app definition) and the
larger controller rows (R-492, R-569, R-575, R-615, R-616, R-498, R-718) were left on purpose.
**Opened:** R-888 — two report fields the hub never reads (a decision). **Seen, not a row:** Tester 1's crash guard reads
TRIPPED since 07:57Z — the morning's two deliberate test crashes; it re-arms by itself after 24 h (`runbooks/crash-guard.md`).
## The 23 rows the first burn-down could not check
Checked from source by a read-only agent (`audits/burndown2-2026-10-05/unchecked-results.jsonl`). Closed: R-350 (duplicate
of R-132, facts merged). Facts added to the open rows R-607, R-883, R-886, R-884, R-756, R-91, R-338, R-488. The three
„not worth it" ones are on STATUS for the operator. The rest need a live box reading (the settle command is in the table).
| Row | Group | Evidence / how to settle (abridged) |
|---|---|---|
| R-76 | UNCHECKABLE-FROM-SOURCE | Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 FileBrowserImage = "gtstef/filebrowser:1.5.6-stable". The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --gre |
| R-91 | UNCHECKABLE-FROM-SOURCE | Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of delet |
| R-209a | UNCHECKABLE-FROM-SOURCE | Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test. — settle: uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h / |
| R-337 | NOT-WORTH-IT | The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 b, err : |
| R-375 | NOT-WORTH-IT | The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is pvesm status showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 if st == nil // st.Type == "pbs" // st.Avail <= 0 { return true, "" } (space preflight sk |
| R-488 | STILL-TRUE-SMALL | The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded interval := 5 * time.Second and time.Sleep(3 * time.Second) // initial settling time, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go: |
| R-504 | UNCHECKABLE-FROM-SOURCE | Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html. — settle: curl -sI https://iso.felhom.eu/ / head -1 |
| R-644 | UNCHECKABLE-FROM-SOURCE | Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs --deployment-password; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches |
| R-814 | UNCHECKABLE-FROM-SOURCE | Hetzner account state; nothing in source records a deletion. — settle: Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status |
| R-815 | UNCHECKABLE-FROM-SOURCE | PBS server-side state on ep0; no GC completion record in the docs (grep). — settle: ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 / grep -i garbage' |
| R-884 | UNCHECKABLE-FROM-SOURCE | Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml prom/prometheus:v3.14.0 -> v3.15.0 (monitoring.yaml:419), and the monitoring Application has no automated syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an un |
| R-132 | UNCHECKABLE-FROM-SOURCE | Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key operator_password_hash, setSetting :2200-2207 writes updated_at = datetime('now'). No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31 |
| R-298 | UNCHECKABLE-FROM-SOURCE | Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 if(d.role==='user-data'){ else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/ |
| R-338 | UNCHECKABLE-FROM-SOURCE | nodes.md:86-88 still claims demo-hp is on the R-50 island (local_api on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent con |
| R-350 | DUPLICATE | of R-132 — Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. |
| R-542 | NOT-WORTH-IT | Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 initialize = append(initialize, c) // every unclaimed disk can be initialized; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkf |
| R-607 | STILL-TRUE-SMALL | Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY if len(newApps) > 0 // len(updated) > 0, and :257-258 says 'nincs változás' when both are empty. updated counts stack-dir copies only (copyTemplates, :447 updated = append(updated, appName) a |
| R-683 | UNCHECKABLE-FROM-SOURCE | Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power- |
| R-756 | UNCHECKABLE-FROM-SOURCE | Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH |
| R-862 | UNCHECKABLE-FROM-SOURCE | Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83). — settle: Ask the operator wheth |
| R-882 | UNCHECKABLE-FROM-SOURCE | Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it). — settle: sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p Act |
| R-883 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/rec |
| R-886 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root nobody image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts |
## Teardown
Machines: drill VM — build guest destroyed, token/scripts/log shredded, powered off, disk back on `virgin`. Boxes: only the
normal deliveries above. Hub: only the deploy, the vouch and the floors. Scratch secrets (hub password file, hub key file,
signed envelopes) are shredded at the end of the session.
+28 -1
View File
@@ -3,7 +3,34 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
sent to it.**
**Updated 2026-10-05 (late night, burn-down round 2 — in progress): 249 open rows after your answer (was 292).**
**Updated 2026-10-05 (late night, burn-down round 2): every box of ours healthy. The open-items list is at 199 (was 292
at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10-05.md`.**
## Tonight, last (2026-10-05): the list at 199
**What happened:**
- **Your answer is recorded** (below): 43 rows closed as accepted.
- **51 more rows fixed and closed**, with one release per repository, delivered the normal way: hub **0.137.0**
(live), agent **0.147.0** (on demo-hp, demo-felhom and Tester 1, root files too), controller **0.297.0** (on all
three), new-install image **0.297.0** (baked, checked, approved), app catalog updated.
- **R-124 is fixed** (your ruling): the recovery recipe now writes the backup-server namespace the way the server reads it.
- **The CI fault (R-887) is understood:** when Gitea is busy, the CI runner's request for work can time out after Gitea
already gave it the job; the job is then never run and is failed 10–13 minutes later. Gitea was busy because of an
outside web crawler and because of MY CI checks, which asked too much — mine now ask once a minute in one small request.
- **Two of my own CI misses tonight:** a new check needed Go, which the CI machine does not have (fixed); a red run was
the fault above (re-run passed).
**The numbers:** 292 before → **199 after**; 1 opened; 94 closed.
**Needs you (none urgent; if you do nothing, each stays open as it is):**
1. **R-469** — a one-paragraph rewording in the app catalog's instruction file; my permission check refused editing
instruction files. Say "go" and it is done in a minute.
2. **R-126** — should a network share be offered as an export destination? Two options in the row; pick one.
3. **R-888** — two facts the boxes report that the hub never shows (installed-app list, retired drives). Needed or not?
4. **R-887** — CI: raise the runner's fetch timeout and/or slow the crawler on Gitea's public pages. If nothing: now and
then a CI run fails without running; it can be re-run.
5. Three more rows look „not worth doing" (R-337, R-375, R-542 — the check found each is by design or harmless). Close
them as accepted? If you say nothing they stay.
## Your answer to the burn-down list (2026-10-05 18:23), recorded
@@ -497,3 +497,9 @@ rc=0
--- PASS: TestR365_BannerSaysDueAtZeroDays (0.12s)
PASS
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.132s
### CI fix (lead, 2026-10-05 ~18:20Z): the new gofmt gate was INCONCLUSIVE on the CI runner (no Go) — CI run 1371 FAILED
Fix: in CI (GITEA_ACTIONS/GITHUB_ACTIONS=true) with no gofmt reachable the gate prints "NOT CHECKED in CI" and exits 0;
elsewhere a missing gofmt stays INCONCLUSIVE. Decoys gofmt/ci-without-go and gofmt/dev-without-go (empty PATH).
Simulated CI (env -i, empty PATH, GITEA_ACTIONS=true): "gofmt gate NOT CHECKED in CI …" rc=0.
Red-proof (CI branch disabled): FAIL: gofmt/ci-without-go: want rc=0 and 'NOT CHECKED in CI', got rc=2.
@@ -0,0 +1,4 @@
== controller delivery 2026-10-05T18:28:59Z
demo-hp gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 32 seconds (healthy)
demo-felhom gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 35 seconds (healthy)
tester-1 (hub customer page, version strings seen): 6 0.297.0 4 0.296.0 4 0.295.0
@@ -0,0 +1,11 @@
== hub deploy 2026-10-05T18:11:56Z
image=gitea.dooplex.hu/admin/felhom-hub:0.137.0
sync=Synced health=Healthy op=Succeeded rev=d75ad0fdf3daf3f3b2a1690746d9a6a70ee4104c
2026/10/05 20:11:05 [INFO] felhom-hub 0.137.0 starting
healthz 200
2026/10/05 20:11:06 [INFO] osupdates: the Docker engine set is approved only by the operator, after 2 healthy ring-0 night(s)
== System page 2026-10-05T18:12:05Z: root-files / agent cells
73: 0 armed unknown 0.142.0 → 0.147.0 (since 2026-10-05)
89: 0 armed 0.147.0 0.147.0
105: 2 armed 0.147.0 0.147.0
121: 2 TRIPPED 2026-10-05T07:57:17Z 0.147.0 0.147.0
@@ -0,0 +1,11 @@
== vouch 2026-10-05T18:27:46Z: agent 0.147.0, golden 0.297.0, min_agent 0.131.0
HTTP/1.1 303 See Other
Location: /configuration?flash=artifacts_set
== floors 2026-10-05T18:28:16Z: POST /customers/<id>/floor min_controller_version=0.297.0 min_agent=0.131.0
demo-hp: Location: /customers/demo-hp?flash=floor_set
demo-felhom: Location: /customers/demo-felhom?flash=floor_set
tester-1: Location: /customers/tester-1?flash=floor_set
2026/10/05 20:28:16 [INFO] Artifact manifest set: agent=0.147.0 golden=0.297.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d"
2026/10/05 20:28:16 [INFO] Customer demo-hp controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
2026/10/05 20:28:17 [INFO] Customer demo-felhom controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
2026/10/05 20:28:17 [INFO] Customer tester-1 controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
@@ -0,0 +1,23 @@
{"id": "R-76", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 `FileBrowserImage = \"gtstef/filebrowser:1.5.6-stable\"`. The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --grep 'R-76|setgid' in controller: only 2026-06 commits). Whether 1.5.6 still drops setgid is runtime behaviour of the image.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- find /mnt/felhom-drives -path '*/userdata/*' -mindepth 3 -maxdepth 5 -type d ! -perm -2000 -printf '%m %u:%g %p\\n'\" | head (any folder a customer made in FileBrowser showing 755 without setgid = still true on 1.5.6)", "minutes_spent": 3}
{"id": "R-91", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word.", "dup_of": null, "unique_facts": "Stale doc to fix in the same commit as the delete: documentation/runbooks/offsite-endpoint.md:24 still says datastore `felhom-offsite` is at `/srv/pbs-felhom`; the real path since 2026-07-27 is `/mnt/pbs-datastore` (RUNBOOK-ep0-datastore-volume-2026-07-27.md:8). CONTEXT.md:3666 is the other line to change.", "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'du -sh /srv/pbs-felhom 2>&1; proxmox-backup-manager datastore list; df -h /'", "minutes_spent": 4}
{"id": "R-209a", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h /", "minutes_spent": 1}
{"id": "R-337", "sev": "P4", "category": "Monitoring & notifications", "group": "NOT-WORTH-IT", "evidence": "The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 `b, err := tier.Service.BackupWithSnapshotHook(...)` then :901 `s.store.RecordBackup(b)` (grep RecordBackup: no other caller). So it is NOT a collection cadence; the status appears when the agent's own job finishes (WaitTask + archive resolve, internal/backup/runner.go:205-232), and a backup the agent did not run itself never appears. The 4-min demo-hp skew is the gap between PBS writing the manifest and the job returning -- unmeasured.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: measure why demo-hp's job returned ~4 min after the manifest landed; cost: a constructed live repro on demo-hp with task-log timing; if never: during an incident the box's backup status can trail the PBS manifest by minutes after an out-of-schedule run, and a run started outside the agent never shows; pick: close with the source fact above written into the row (status = agent job end, not polling), reopen only if a lag is seen on a scheduled run.", "settle_cmd": null, "minutes_spent": 7}
{"id": "R-375", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is `pvesm status` showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 `if st == nil || st.Type == \"pbs\" || st.Avail <= 0 { return true, \"\" }` (space preflight skips PBS and fails open on 0), and the same report says the hub's gauge reads the real 3.7/97.9 GB.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: confirm on ep0 that the namespace-scoped token lacks Datastore.Audit; cost: a read-only ep0 session; if never: the PVE UI on a box shows 0/0/0 for felhom-pbs, which no Felhom code reads (runner.go:265); pick: close as cosmetic with this pointer.", "settle_cmd": "(if ever wanted) ssh root@ep0 'proxmox-backup-manager acl list' ; on a box: pvesm status --storage felhom-pbs", "minutes_spent": 6}
{"id": "R-488", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded `interval := 5 * time.Second` and `time.Sleep(3 * time.Second) // initial settling time`, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:90, restore_unit.go:471. No commit since 2026-09-13 touching internal/backup mentions R-488/test speed. Runtime (333 s) was NOT re-measured here (read-only).", "dup_of": null, "unique_facts": null, "small_fix": "controller: add Manager fields healthSettle/healthInterval (defaults 3s/5s, set in the constructor) used by waitForHealthy; a test helper (the existing Manager fixture constructor) sets them to 0/10ms. Test: TestWaitForHealthy_DefaultsAreProduction asserts a fresh Manager has 3s/5s (pins production), and measure `go test ./internal/backup` wall time before/after in the commit message (expect the 89 >=1 s tests to drop). Run with the R-650 docker-free seams.", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}
{"id": "R-504", "sev": "P4", "category": "Install & onboarding", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "If 404 is confirmed it is cosmetic (row's own re-rank); the operator could close it as accepted rather than add a Cloudflare rule.", "settle_cmd": "curl -sI https://iso.felhom.eu/ | head -1", "minutes_spent": 2}
{"id": "R-644", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs `--deployment-password`; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches the crash text. Not proven.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- docker ps -a --filter name=gokapi --format '{{.Names}} {{.Status}}'; pct exec 9202 -- grep -s -c '\\\"Password\\\":\\\"\\\"' /opt/docker/stacks/gokapi/config/config.json\"", "minutes_spent": 3}
{"id": "R-814", "sev": "P4", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Hetzner account state; nothing in source records a deletion.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status", "minutes_spent": 1}
{"id": "R-815", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "PBS server-side state on ep0; no GC completion record in the docs (grep).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 | grep -i garbage'", "minutes_spent": 2}
{"id": "R-884", "sev": "P4", "category": "Monitoring & notifications", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live.", "dup_of": null, "unique_facts": "Cause candidate: Renovate 53c6e99 prometheus v3.15.0 merged 2026-10-03, never synced because monitoring has no auto-sync in git.", "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n mon-system get deploy prometheus -o jsonpath='{.spec.template.spec.containers[0].image}' (v3.14.0 => the diff is the Renovate bump)", "minutes_spent": 5}
{"id": "R-132", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key `operator_password_hash`, setSetting :2200-2207 writes `updated_at = datetime('now')`. No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "On a copy of the hub DB (hub.db + -wal + -shm): sqlite3 hub.db \"SELECT updated_at FROM hub_settings WHERE key='operator_password_hash'\" -- rotated only if later than 2026-09-18 (the last exposure, R-580 folded here); no row = still the ConfigMap seed", "minutes_spent": 4}
{"id": "R-298", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 `if(d.role==='user-data'){` else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/role.go:172-186 -- a dir storage (felhom-backup) is user-data unless its backing device is on the system disk; disks.go:1236-1240 deliberately does not reclassify the backup-target drive. role.go unchanged since 2026-08-09. demo-hp was reprovisioned before 2026-09-21, so the 2026-08-10 topology may no longer hold.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp 'grep -A2 felhom-backup /etc/pve/storage.cfg; lsblk -no PKNAME $(findmnt -no SOURCE /) ; lsblk -no PKNAME $(findmnt -no SOURCE /mnt/nvme-1tb)' (same parent disk => role=system => the page still locks it => R-298 true; different => user-data => not reproducible on this box)", "minutes_spent": 8}
{"id": "R-338", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"grep -E 'listen_addr|island_' /etc/felhom-agent/agent.json; pct config 9201 | grep ^net; ip -br link show master vmbr9\"", "minutes_spent": 6}
{"id": "R-350", "sev": "P3", "category": "Security & access", "group": "DUPLICATE", "evidence": "Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence.", "dup_of": "R-132", "unique_facts": "(1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked.", "small_fix": null, "not_worth": null, "settle_cmd": null, "minutes_spent": 3}
{"id": "R-542", "sev": "P3", "category": "Storage & devices", "group": "NOT-WORTH-IT", "evidence": "Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 `initialize = append(initialize, c) // every unclaimed disk can be initialized`; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkfs wrapper explicitly allows it: configs/felhom-mkfs-guarded.sh:59-60 'Mounts under /mnt/felhom-drives are our own drives (the agent detaches before a re-init) -> allowed'. Controller passes initialize through untouched (agent_disk_handlers.go:158-162). `already_mounted` is only set for controller-contributed stores (agentapi/client.go:447-451, omitempty) -- so 'null' is expected for agent candidates.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: drop felhom-mounted disks from `initialize`; cost: reverses the R-220/re-init design (wrapper comment) and needs a decision on how a registered drive is re-initialized; if never: the raw endpoint lists a registered drive under initialize while the page (customer view) filters it -- only a session reading the raw endpoint is misled; pick: close as by-design, add one comment line at disks.go:438 saying a registered felhom drive is listed here deliberately.", "settle_cmd": null, "minutes_spent": 9}
{"id": "R-607", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-SMALL", "evidence": "Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY `if len(newApps) > 0 || len(updated) > 0`, and :257-258 says 'nincs változás' when both are empty. `updated` counts stack-dir copies only (copyTemplates, :447 `updated = append(updated, appName)` after a hash mismatch); for a deployed+pinned app whose catalog moved, renderSource (:457-469 table) copies the STORED definition, so the hash matches and nothing is 'updated' although the git cache moved. CatalogImages is read from the catalog cache only inside ScanStacks (stacks/manager.go:666-672, assigned :690). So (a) the message measures the stack dir, not the catalog; (b) CatalogImages refreshes only on a ScanStacks, which this sync does not trigger. The 29 s nextcloud case (needed several rounds) is not explained by this.", "dup_of": null, "unique_facts": null, "small_fix": "controller sync.go: record the catalog git HEAD before and after gitCloneOrPull; if it moved, call s.rescanFn() even when newApps/updated are empty, and say 'Katalógus frissítve — az alkalmazások nem változtak' (and EN) instead of 'nincs változás'. Test (red first): a Syncer with a fake pull that moves HEAD and a frozen pinned app (renderSource returns the stored definition) asserts rescanFn was called once and the message is not the no-change one; and a no-move pull asserts rescanFn NOT called.", "not_worth": null, "settle_cmd": null, "minutes_spent": 9}
{"id": "R-683", "sev": "P3", "category": "App updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power-cut-mid-update drill: night-2026-10-04/MORNING-NOTE.md:64 'A1 power cut mid-update (demo-hp): NOT RUN'; DRILL-night-2026-09-25.md:102 was a cut during romm's verifying, not checked for the backup choice.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Not a read-only command: the next power-cut-during-update drill with the controller log saved at arm time, then grep \"phase backing-up\\|Tier-2\" in the saved pre-cut log", "minutes_spent": 7}
{"id": "R-756", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 `if !m.DriveLive(hddPath)` -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 `return m.isMountPoint(hddPath)` -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names `.../scratch_hdd/userdata/calibre-web`, so on 9202 HDD_PATH is a per-app subfolder, which can never be a mount point -> the 409 is certain for that value. Open: whether that HDD_PATH was written by the product (handlers.go:550 prefill from place.Drive) or by the test venue.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- grep -H HDD_PATH /opt/docker/stacks/calibre-web/app.yaml /opt/docker/stacks/grimmory/app.yaml; pct exec 9202 -- findmnt -no TARGET,SOURCE /mnt/felhom-drives/scratch_hdd\" (HDD_PATH = per-app subfolder => product/venue wrote a non-root HDD_PATH; HDD_PATH = drive root and not mounted => scratch drive not registered)", "minutes_spent": 10}
{"id": "R-862", "sev": "P3", "category": "Box system & updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Ask the operator whether the three bootstrap commands ran; read-only proof: the hub's operator view of Tester 2 (OS/config-bundle sha vs the vouched sha), or whether os_config_bundle_behind has fired for it", "minutes_spent": 2}
{"id": "R-882", "sev": "P3", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p ActiveEnterTimestamp (instance-manager older than a k3s/containerd/iscsid restart => the stale-PID state can recur)", "minutes_spent": 2}
{"id": "R-883", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/recipe-importer:latest (pull Always :27); adventurelog-system/adventurelog.yaml:100 and :256 adventurelog-backend/frontend:latest (pull Always :101,:257); jarrs-system/jarr-dev.yaml:311,:345,:630 gitea.dooplex.hu/admin/jarr:latest (pull Always). Zipline fixed (90f60e4, 4c8ec7a). The repo's own rule homelab-manifests/CLAUDE.md:120 'Image tags always pinned'. Helm values files (external-dns, pihole, plex, authentik, cnpg) not checked for tag fields beyond a `tag: latest|dev|empty` grep (no hits).", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests only: for each line above, read the running digest/version (`sudo kubectl get pod -n <ns> -o jsonpath='{..imageID}'`), pin that exact tag (or @sha256 for the self-built gitea.dooplex.hu jarr/recipe-importer images, which have no version tags), add the version-checker match-regex annotation per CLAUDE.md:120, and set imagePullPolicy IfNotPresent. Test: `grep -rnE 'image:.*(:latest|:dev)\\s*$' --include=*.yaml .` returns nothing; after ArgoCD sync each pod's imageID equals the pre-change one (no upgrade). Operator-owned DooPlex change: needs the operator's word.", "not_worth": null, "settle_cmd": null, "minutes_spent": 6}
{"id": "R-886", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager real storage').", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests: add pod `securityContext: {fsGroup: 65534, fsGroupChangePolicy: OnRootMismatch}` (65534 = nobody, the image's user -- confirm with `sudo kubectl -n mon-system exec deploy/alertmanager -- id`) to the alertmanager Deployment. Test (consequence, per CLAUDE.md): create a silence via amtool/API, delete the pod, after it returns the silence is still listed AND the log has no 'Running maintenance failed ... permission denied' within 15 min (positive observable: a 'maintenance done' line).", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}
+23
View File
@@ -103,6 +103,29 @@ The full text of every row below: `git show e8c56c44:documentation/backlog/OPEN-
| **R-269** | **A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): token store reloads on growth before answering; `TestTokenStore_RotatedOutTokenRejectedFirst`; red-proof agent-red-proofs.txt |
| **R-317** | **The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): dnsmasq install probed by its service unit; `TestEnsureDnsmasq_*`; red-proof agent-red-proofs.txt |
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-132 (its unique facts moved there) | Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. |
| **R-591** | **[P3-LOW] `Stack.Copy()` is a deep copy with one shallow field, and the field is new.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `deepCopyStack` copies every reference in Meta (i18n overlay and 11 more found by `TestDeepCopyStackMetaSharesNoReference`); `TestDeepCopyStackI18nIsNotShared` |
| **R-568** | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: disk-health rows sorted by durable id; `TestDiskHealthRows_OrderIsStable` |
| **R-567** | **[P3-LOW] The two drive wizard pages (/storage/init, /storage/attach) do not highlight the Tárhely menu group — the sidebar reads as if the household left the storage section.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: storage wizard pages open the Storage nav group; `TestStorageWizardPages_OpenTheStorageNavGroup`; two parity fixtures re-captured (nav only) |
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: fill-watch also every 10 min (`sched.Every`), daily + start-up kept; `TestFillWatchRunsOnAnInterval` |
| **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: same change as R-363 (the interval watch); `TestFillWatchRunsOnAnInterval` |
| **R-10** | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state `idea`.** Moved verbatim; nothing added (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: dump directory fsynced after the rename; `TestDumpOneTo_SyncsTheDumpDirectoryAfterRename` |
| **R-552** | **[P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: remove clears the interrupted-restore notice (`ClearInterruptedRestore`, wired in the remove path); `TestR552_RemoveClearsTheInterruptedRestoreNotice` |
| **R-251** | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: off-site marker tag not listed as an app; `TestR251_MarkerTagIsNotAnApp` |
| **R-104** | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a lock surviving the self-heal is classed `locked` with its own cause line; `TestR104_SurvivingLockIsNamed` |
| **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: password deploy fields served as required by the API (fresh metadata per call); `TestR619_PasswordFieldIsServedAsRequired` |
| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a restore onto a detached drive names the drive; `TestR362_DetachedDriveIsNamed` |
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: files-restore refusal names the second drive's whole copy (and is in both languages now); `TestR675_RefusalNamesTheWholeCopy` |
| **R-256** | **C2 — „A mentéskezelő nem elérhető." names no route at all.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `flash.offbox.mgr_unavailable/_unreachable` name a route (hu+en); `TestR256_R257_OffboxRefusalsNameARoute` |
| **R-257** | **C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `flash.offbox.not_orphaned` reworded (hu+en); `TestR256_R257_OffboxRefusalsNameARoute` |
| **R-240** | **A backup that covered nothing calls itself „Sikeres".** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a zero-selection run no longer says „Sikeres"; `TestR240_ZeroSelectionRunDoesNotSaySuccess` |
| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a 0-day deletion countdown says it is due — page AND top bar (`TestR365_OverdueCountdownIsNotFutureTense`, `TestR365_BannerSaysDueAtZeroDays`) |
| **R-425** | **`offbox_rename_gate.py` scans a fixed three-entry `FILES` list.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: offbox rename gate finds files by pattern and judges the bundle; 3 decoys |
| **R-565** | **[P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: English-page test knows ASCII Hungarian words; a real `mp` leak became a key; `TestI18nEnglishPages` |
| **R-564** | **[P3-LOW] The retrieval-promise gate's Hungarian stems cannot see a SPLIT verb — „csak akkor állíthatók vissza", „hozod vissza" — so those Hungarian sentences were never scanned; the English translation exposed them.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: retrieval-promise gate knows split-verb Hungarian; 7 occurrences registered; 2 decoys |
| **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: Go-side check for HTML-escapable values; `TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping` |
| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `scripts/gofmt_gate.py` (NOT CHECKED out loud on the Go-less CI runner, INCONCLUSIVE elsewhere; 3 decoys); 12 files formatted |
| **R-208** | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: controller and hub Dockerfiles: `ARG VERSION…` just above `go build`; `TestR208_DockerfileVersionArgsSitBelowModuleDownload`, felhom.eu `scripts/test_dockerfile_arg_order.py` (hub v0.137.0 deployed) |
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** (P4) | CLOSED 2026-10-05 — CHECKED, NOTHING LEFT (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: swept: none of the six candidate files has a date literal feeding an assertion against the real clock (identity/format/ordering checks only) — nothing to change; the faked-future-date CI idea is a separate, larger job |
---
File diff suppressed because one or more lines are too long
@@ -0,0 +1,5 @@
== round trip 2026-10-05T18:26:49Z: anonymous GET .../generic/felhom-golden/0.297.0/golden.tar.zst
HTTP 200
bytes 648208028
sha256 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
printed 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
@@ -0,0 +1,52 @@
# Golden 0.297.0 — bake + publish + vouch, 2026-10-05 (night, burn-down round 2)
Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–5, in the drill VM on DooPlex.
| | Previous (`../golden-0.296.0-2026-10-05/`) | This bake |
|---|---|---|
| `build-golden.sh` | sha256 `645b3b659cba…` | same file, unchanged (agent repo `configs/build-golden.sh`) |
| Controller | `felhom-controller:0.296.0` | **`felhom-controller:0.297.0`** (MinAgent 0.131.0, unchanged) |
| Docker engine | the approved set `os-docker-20261004-142842` | same pinned set (same `GOLDEN_DOCKER_PKGS` as the 0.296.0 bake) |
| Guest packages | template | template — `GOLDEN_GUEST_PKGS` EMPTY |
## Launch
- No qemu running before; drill VM reverted to `virgin`, cold-booted per §4.0; `pveversion` = `pve-manager/9.2.2`.
- `pveam update` → `update successful`; `pveam available` listed `debian-13-standard_13.6-1_amd64.tar.zst` (downloaded).
- `/root/bake-run.sh` reads the token from the file; transient unit `golden-bake`. Token copied file → file (`scp`);
`systemctl show golden-bake -p Environment -p ExecStart | grep -c -F <token>` = **0**.
## Pass markers (from `bake.log`, this folder)
```
docker OK (overlay2; data-root /var/lib/docker)
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.297.0
GOLDEN_SHA256=8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
```
No `excluding` and no `FATAL` in the log (grep count 0).
## Round trip — `02-round-trip.txt`
Anonymous GET of `…/generic/felhom-golden/0.297.0/golden.tar.zst`: HTTP 200, 648 208 028 bytes, sha256 equals the printed one.
## Secrets
Saved-log leak grep for the literal token: **0**; positive control (a throwaway copy with the token appended): **1**,
copy shredded.
## Vouch (step 5)
`POST /configuration/artifacts` (Basic + `X-Felhom-Operator`, hub v0.137.0): agent **0.147.0**, golden **0.297.0**,
`min_agent` **0.131.0** → `303 flash=artifacts_set`; hub log `Artifact manifest set: agent=0.147.0 golden=0.297.0
min_agent="0.131.0" … bundle_sha="326527d0…"`. Per-customer floors 0.297.0 (declared MinAgent 0.131.0) for demo-hp,
demo-felhom, tester-1 (`../../audits/burndown2-2026-10-05/delivery/vouch-golden-floors.txt`). The global floor unchanged.
## Teardown
`pct destroy 9100 --purge` (rc 0); `shred -u` of the token, runner script, bake script and log in the VM (log copied off
first); `poweroff`; qemu gone (`ps -eo comm | grep -c qemu-system-x86` = 0); `qemu-img snapshot -a virgin`. Host:
nothing provisioned.
@@ -0,0 +1,339 @@
[golden] build-golden.sh v3.2.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.297.0
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
Logical volume "vm-9100-disk-0" created.
Logical volume pve/vm-9100-disk-0 changed.
Creating filesystem with 8388608 4k blocks and 2097152 inodes
Filesystem UUID: 3f3e66a1-e05c-42f3-91af-3ff9fc5c8f65
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624
Logical volume "vm-9100-disk-1" created.
Logical volume pve/vm-9100-disk-1 changed.
Creating filesystem with 6291456 4k blocks and 1572864 inodes
Filesystem UUID: 9327fbdd-d6c5-4299-a41f-3cb56b1238b3
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
Total bytes read: 553512960 (528MiB, 115MiB/s)
Detected container architecture: amd64
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
done: SHA256:6hCAi5WjsL3daO1FsrQR2bko1QAWr49AoJUmsNXkGSY root@felhom-golden
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
done: SHA256:STGSy09tGfb2oiy9XwC4UHzFcSawljx6IhtUka8tiJI root@felhom-golden
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:rh5fSFDYSXH/CINtl163usIqQy+m082joU7xzA5LENU root@felhom-golden
[golden] starting + installing Docker (official repo, trixie channel) …
[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docker-ce=5:29.8.2-1~debian.13~trixie docker-ce-cli=5:29.8.2-1~debian.13~trixie docker-ce-rootless-extras=5:29.8.2-1~debian.13~trixie docker-compose-plugin=5.6.0-1~debian.13~trixie
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
installed: containerd.io 2.3.6-1~debian.13~trixie
installed: docker-buildx-plugin 0.37.1-1~debian.13~trixie
installed: docker-ce 5:29.8.2-1~debian.13~trixie
installed: docker-ce-cli 5:29.8.2-1~debian.13~trixie
installed: docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie
installed: docker-compose-plugin 5.6.0-1~debian.13~trixie
[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a
[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
Unable to find image 'hello-world:latest' locally
latest: Pulling from library/hello-world
4f55086f7dd0: Pulling fs layer
4f55086f7dd0: Download complete
4f55086f7dd0: Pull complete
Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8
Status: Downloaded newer image for hello-world:latest
docker OK (overlay2; data-root /var/lib/docker)
live-restore: on
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.297.0 (no registry cred at deploy) …
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
Configure a credential helper to remove this warning. See
https://docs.docker.com/go/credential-store/
0.297.0: Pulling from admin/felhom-controller
774043ccc8cc: Pulling fs layer
ab6b448d4be9: Pulling fs layer
23a5bfa58353: Pulling fs layer
862a57157567: Pulling fs layer
6db4169d1fd9: Pulling fs layer
167f80584563: Pulling fs layer
862a57157567: Waiting
6db4169d1fd9: Waiting
167f80584563: Waiting
774043ccc8cc: Verifying Checksum
774043ccc8cc: Download complete
862a57157567: Verifying Checksum
862a57157567: Download complete
23a5bfa58353: Verifying Checksum
23a5bfa58353: Download complete
167f80584563: Verifying Checksum
167f80584563: Download complete
6db4169d1fd9: Verifying Checksum
6db4169d1fd9: Download complete
ab6b448d4be9: Verifying Checksum
ab6b448d4be9: Download complete
774043ccc8cc: Pull complete
ab6b448d4be9: Pull complete
23a5bfa58353: Pull complete
862a57157567: Pull complete
6db4169d1fd9: Pull complete
167f80584563: Pull complete
Digest: sha256:23e4e0ffd9c9e2df28196e55bce8b89d4dfc1f9d4f8e52ca3ae2872e365073a0
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.297.0
gitea.dooplex.hu/admin/felhom-controller:0.297.0
[golden] asking the controller which infra images it manages …
[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
v3.7.13: Pulling from library/traefik
e2de96513ba9: Pulling fs layer
b686a4f73445: Pulling fs layer
78cb21c375ca: Pulling fs layer
acb2f33459b1: Pulling fs layer
acb2f33459b1: Waiting
e2de96513ba9: Verifying Checksum
e2de96513ba9: Download complete
b686a4f73445: Download complete
acb2f33459b1: Verifying Checksum
acb2f33459b1: Download complete
e2de96513ba9: Pull complete
78cb21c375ca: Verifying Checksum
78cb21c375ca: Download complete
b686a4f73445: Pull complete
78cb21c375ca: Pull complete
acb2f33459b1: Pull complete
Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0
Status: Downloaded newer image for traefik:v3.7.13
docker.io/library/traefik:v3.7.13
2026.9.3: Pulling from cloudflare/cloudflared
2cc7ee286bf3: Pulling fs layer
c172f21841df: Pulling fs layer
218cf840d0d9: Pulling fs layer
f6069939f718: Pulling fs layer
d6b1b89eccac: Pulling fs layer
2780920e5dbf: Pulling fs layer
7c12895b777b: Pulling fs layer
3214acf345c0: Pulling fs layer
52630fc75a18: Pulling fs layer
dd64bf2dd177: Pulling fs layer
b839dfae01f6: Pulling fs layer
ebddc55facdc: Pulling fs layer
c4bc6f35ff5e: Pulling fs layer
b96fe2995f90: Pulling fs layer
58c0c263dc73: Pulling fs layer
bd8962e29291: Pulling fs layer
cac2ae0193cb: Pulling fs layer
f0383d5ebc47: Pulling fs layer
3214acf345c0: Waiting
52630fc75a18: Waiting
dd64bf2dd177: Waiting
b839dfae01f6: Waiting
ebddc55facdc: Waiting
c4bc6f35ff5e: Waiting
b96fe2995f90: Waiting
58c0c263dc73: Waiting
bd8962e29291: Waiting
cac2ae0193cb: Waiting
f0383d5ebc47: Waiting
f6069939f718: Waiting
d6b1b89eccac: Waiting
2780920e5dbf: Waiting
7c12895b777b: Waiting
2cc7ee286bf3: Download complete
218cf840d0d9: Verifying Checksum
218cf840d0d9: Download complete
c172f21841df: Verifying Checksum
c172f21841df: Download complete
f6069939f718: Verifying Checksum
f6069939f718: Download complete
2cc7ee286bf3: Pull complete
d6b1b89eccac: Download complete
2780920e5dbf: Verifying Checksum
2780920e5dbf: Download complete
7c12895b777b: Verifying Checksum
7c12895b777b: Download complete
3214acf345c0: Verifying Checksum
3214acf345c0: Download complete
52630fc75a18: Verifying Checksum
52630fc75a18: Download complete
dd64bf2dd177: Download complete
c172f21841df: Pull complete
b839dfae01f6: Verifying Checksum
b839dfae01f6: Download complete
ebddc55facdc: Verifying Checksum
ebddc55facdc: Download complete
c4bc6f35ff5e: Verifying Checksum
c4bc6f35ff5e: Download complete
58c0c263dc73: Verifying Checksum
58c0c263dc73: Download complete
bd8962e29291: Verifying Checksum
bd8962e29291: Download complete
b96fe2995f90: Verifying Checksum
b96fe2995f90: Download complete
cac2ae0193cb: Verifying Checksum
cac2ae0193cb: Download complete
218cf840d0d9: Pull complete
f0383d5ebc47: Verifying Checksum
f0383d5ebc47: Download complete
f6069939f718: Pull complete
d6b1b89eccac: Pull complete
2780920e5dbf: Pull complete
7c12895b777b: Pull complete
3214acf345c0: Pull complete
52630fc75a18: Pull complete
dd64bf2dd177: Pull complete
b839dfae01f6: Pull complete
ebddc55facdc: Pull complete
c4bc6f35ff5e: Pull complete
b96fe2995f90: Pull complete
58c0c263dc73: Pull complete
bd8962e29291: Pull complete
cac2ae0193cb: Pull complete
f0383d5ebc47: Pull complete
Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c
Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3
docker.io/cloudflare/cloudflared:2026.9.3
1.5.6-stable: Pulling from gtstef/filebrowser
55afa1ecc21d: Pulling fs layer
8ed8f35f8d4f: Pulling fs layer
989b226a579c: Pulling fs layer
660aeead31d5: Pulling fs layer
4f4fb700ef54: Pulling fs layer
adce24567e4c: Pulling fs layer
f17ea56b313b: Pulling fs layer
6b6f3b3efe88: Pulling fs layer
4ed1ca4f3fce: Pulling fs layer
e6fc9c6a5757: Pulling fs layer
d47782d1182a: Pulling fs layer
660aeead31d5: Waiting
6b6f3b3efe88: Waiting
4ed1ca4f3fce: Waiting
e6fc9c6a5757: Waiting
d47782d1182a: Waiting
4f4fb700ef54: Waiting
adce24567e4c: Waiting
f17ea56b313b: Waiting
55afa1ecc21d: Verifying Checksum
55afa1ecc21d: Download complete
660aeead31d5: Verifying Checksum
660aeead31d5: Download complete
8ed8f35f8d4f: Verifying Checksum
8ed8f35f8d4f: Download complete
4f4fb700ef54: Verifying Checksum
4f4fb700ef54: Download complete
55afa1ecc21d: Pull complete
989b226a579c: Verifying Checksum
989b226a579c: Download complete
f17ea56b313b: Verifying Checksum
f17ea56b313b: Download complete
6b6f3b3efe88: Verifying Checksum
6b6f3b3efe88: Download complete
4ed1ca4f3fce: Verifying Checksum
4ed1ca4f3fce: Download complete
adce24567e4c: Verifying Checksum
adce24567e4c: Download complete
d47782d1182a: Verifying Checksum
d47782d1182a: Download complete
8ed8f35f8d4f: Pull complete
e6fc9c6a5757: Verifying Checksum
e6fc9c6a5757: Download complete
989b226a579c: Pull complete
660aeead31d5: Pull complete
4f4fb700ef54: Pull complete
adce24567e4c: Pull complete
f17ea56b313b: Pull complete
6b6f3b3efe88: Pull complete
4ed1ca4f3fce: Pull complete
e6fc9c6a5757: Pull complete
d47782d1182a: Pull complete
Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8
Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable
docker.io/gtstef/filebrowser:1.5.6-stable
1.1.0: Pulling from admin/felhom-samba
897d797d2723: Pulling fs layer
3051591aa250: Pulling fs layer
ce57a3f93416: Pulling fs layer
fb94eeec2fe1: Pulling fs layer
fb94eeec2fe1: Waiting
ce57a3f93416: Verifying Checksum
ce57a3f93416: Download complete
fb94eeec2fe1: Verifying Checksum
fb94eeec2fe1: Download complete
897d797d2723: Verifying Checksum
897d797d2723: Download complete
897d797d2723: Pull complete
3051591aa250: Verifying Checksum
3051591aa250: Download complete
3051591aa250: Pull complete
ce57a3f93416: Pull complete
fb94eeec2fe1: Pull complete
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
gitea.dooplex.hu/admin/felhom-samba:1.1.0
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
[golden] identity-clean + minimize …
[golden] stop + archive …
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
INFO: archive file size: 618MB
INFO: Finished Backup of VM 9100 (00:00:29)
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_05-20_24_09.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
[golden] publishing golden (648208028 bytes, sha256 8cebc42e15b091f0…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.297.0/golden.tar.zst
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.297.0
GOLDEN_SHA256=8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.297.0 / 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)