burn-down round 2: controller v0.297.0 rows closed (23), golden 0.297.0 evidence, delivery evidence, 23-row unchecked table, STATUS/CONTEXT/REPORT (292 -> 199; 1 opened, 94 closed)
gates / gates (push) Successful in 1m59s
gates / gates (push) Successful in 1m59s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+10
@@ -16,6 +16,16 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **2026-10-05 (late night) — burn-down round 2 (releases).** Register 292 → 199 (1 opened: R-888; 94 closed: 43
|
||||
> accepted by the operator 18:23, 51 fixed). Releases: hub v0.137.0 (`557629d`, deployed), agent v0.147.0 (tag, sha
|
||||
> `642c4d19…`, bundle `326527d0…`, signed jobs to 3 boxes), controller v0.297.0 (`1453cfc` + `6f1ba1f`), golden 0.297.0
|
||||
> (`8cebc42e…`, vouched with agent 0.147.0 / min_agent 0.131.0; floors 0.297.0 for demo-hp, demo-felhom, tester-1),
|
||||
> catalog `4828dc7`. R-124: recipe root namespace = `""` (+ runbook). R-887 mechanism from Gitea's log: a FetchTask the
|
||||
> runner abandons after assignment → zombie stop after ~10 min; load = an outside crawler + the session's own 15-page CI
|
||||
> polling (now one `runs?head_sha=` call per minute). New gates: `stands` (felhom.eu), `gofmt` (controller; NOT
|
||||
> CHECKED out loud on the Go-less runner). R-469 not done: the permission check refused the catalog CLAUDE.md edit.
|
||||
> Report: `REPORT-burndown2-2026-10-05.md`.
|
||||
|
||||
> **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 292 (1 opened — R-887 CI runner fault — 45
|
||||
> closed): 24 fixed by later work + 2 duplicates (each re-checked; `audits/burndown-2026-10-05/partA-table.md` holds all
|
||||
> 317 P3/P4 verdicts), 19 small fixes with tests/red-proofs (catalog `29ac711`, agent `d833163`, controller `114ff27`,
|
||||
|
||||
@@ -0,0 +1,115 @@
|
||||
# REPORT — burn-down round 2: the operator's answer recorded, R-887 re-diagnosed, small rows fixed WITH releases — 2026-10-05 (late night)
|
||||
|
||||
| Part | Result |
|
||||
|---|---|
|
||||
| **A** — rulings, then R-887 | **done** — rulings commit `301fe45` (count after: **249**); R-887 re-diagnosed from the logs (the restart idea refuted; the mechanism then SEEN in Gitea's own log), dated check 2026-10-12 |
|
||||
| **B** — R-124, the small rows, the 23 unchecked | **done** — R-124 fixed (agent v0.147.0 + runbook); 50 more rows fixed and closed with tests and red-proofs; the 23 checked from source (1 duplicate closed, facts added to 8 rows, the rest left as they need a live box or a decision) |
|
||||
| **B.4** — releases, delivered the normal way | **done** — hub v0.137.0 deployed; agent v0.147.0 released + signed jobs (binary and bundle) to demo-hp, demo-felhom, Tester 1; controller v0.297.0 + golden 0.297.0 baked, vouched, floors raised, all three boxes on 0.297.0; catalog pushed |
|
||||
| **C** — numbers and record | **done** — STATUS shows 199 and asks nothing about the closed list |
|
||||
|
||||
| Rows before | Rows after | Opened | Closed |
|
||||
|---|---|---|---|
|
||||
| **292** | **199** | **1** (R-888) | **94** (43 accepted by the operator + 51 fixed/merged) |
|
||||
|
||||
Counted by `register_shape_gate.py`'s method. Target ≤ 220: met.
|
||||
|
||||
## Baselines (re-verified at the start)
|
||||
|
||||
felhom.eu `e8c56c440a` (hub v0.136.0) · controller `114ff2761a` (v0.296.0) · agent `d83316326e` (v0.146.1) · catalog
|
||||
`29ac711d26` · golden 0.296.0 · register 292. The agent clone had a stray `scripts/__pycache__/` from round 1 — removed.
|
||||
|
||||
## Part A — the rulings commit and R-887
|
||||
|
||||
- `301fe45`: 43 rows closed as „accepted by the operator, 2026-10-05", each with its one-line reason from the list;
|
||||
R-124 and R-698 kept (R-698 owner → operator); R-831/R-870 carry the not-rotated rulings; R-887 records the screenshot
|
||||
(one runner, ID 2, online). STATUS: the list and the rotate/runners requests removed. **Count after: 249.** CI run
|
||||
1363 success.
|
||||
- **R-887, from the logs:** the runner's last restart was 13:24:42Z; the lost attempts started 15:05–15:46Z — **not a
|
||||
restart**. Four lost attempts (not two): each without a runner `task` line, each failed at a :38-second mark 10–13 min
|
||||
after assignment. Gitea's log for that hour had rotated. **Then it happened again at 17:15Z with the log intact:**
|
||||
`slow POST …/RunnerService/FetchTask for 10.42.0.42, elapsed 3192ms` → `context canceled` → 17:28:39
|
||||
`clear_tasks.go … stopTasks() … task 1371` — the runner abandoned its fetch after Gitea assigned the task; Gitea's
|
||||
zombie stop failed it. Load at that minute: an outside crawler on public commit pages, and this session's CI waiter
|
||||
(15-page job listings at 13–31 s each). The waiter now makes ONE `runs?head_sha=` call a minute. A lost run re-runs
|
||||
with `POST …/actions/runs/<id>/rerun` (used twice: controller run 1357 → success; catalog run 1368 → success).
|
||||
**Dated check 2026-10-12** in DUE-CHECKS. The fix on DooPlex (runner fetch timeout, crawler) is the operator's.
|
||||
|
||||
## Part B — fixes by repo
|
||||
|
||||
**agent v0.147.0** (`f1b9b41`, CI 1365; tag `v0.147.0`; binary sha256 `642c4d19…`, bundle `326527d0…`, verified by
|
||||
download; CHANGELOG `208fac8`, CI 1367): R-124, R-118, R-269, R-317 — red-proofs `audits/burndown2-2026-10-05/r124-red-proof.txt`,
|
||||
`agent-red-proofs.txt`. **Delivery:** vouched (agent 0.147.0, golden 0.296.0 first), signed `agent_update` ×3, then
|
||||
`agent_config_update` ×3 (felhom-op-1, ttl 45 m); hub System page: demo-hp, demo-felhom, Tester 1 — agent 0.147.0,
|
||||
root files 0.147.0 (`delivery/`). Tester 2 offline — nothing sent.
|
||||
|
||||
**hub v0.137.0** (`557629d`, CI 1369; manifest `81d04a6`; CI 1370): R-277, R-581, R-600, R-544, R-855, R-134, R-92,
|
||||
R-292, R-599, R-725, R-728, R-208 (hub half) — red-proofs `felhom-eu-red-proofs.txt`. **Deployed:** ArgoCD Synced/Healthy
|
||||
at `d75ad0f`, image `felhom-hub:0.137.0`, `felhom-hub 0.137.0 starting`, healthz 200; R-855's new line seen live
|
||||
(„after 2 healthy ring-0 night(s)"). (The build ran while a helper was still appending to an audit text file outside
|
||||
`hub/` — the image is the committed `hub/` tree; said here because the clean-tree gate is literal.)
|
||||
|
||||
**felhom.eu gates/tools/docs** (same commits): R-819 (`stands` gate), R-857, R-555, R-364 (`hu_grep.py` + REUSE.md),
|
||||
R-587, R-571, R-129 (demo-hp authenticates with DooPlex's own key — corrected everywhere it said „no key"), R-124 runbook.
|
||||
New script tests pass under a BusyBox + bash + python3 + git PATH (the CI runner's tools): 19/19.
|
||||
|
||||
**controller v0.297.0** (`1453cfc`; CI run 1371 **FAILED** — the new gofmt gate was INCONCLUSIVE on the Go-less runner;
|
||||
fixed in `6f1ba1f`, CI 1372 success): R-591, R-568, R-567, R-363, R-547, R-10, R-552, R-251, R-104, R-619, R-362, R-675,
|
||||
R-256, R-257, R-240, R-365, R-425, R-565, R-564, R-603, R-454, R-208, R-457 (swept, nothing left) + two twins found and
|
||||
fixed on the way (the top-bar countdown at 0 days; nine more shared references in `deepCopyStack`). Red-proofs
|
||||
`controller-red-proofs.txt` (two first attempts that did not convict are marked, with valid re-runs). **MinAgent 0.131.0.**
|
||||
**Image** `felhom-controller:0.297.0`. **Golden 0.297.0** baked per RUNBOOK §4.0–4.1 (`documentation/tests/golden-0.297.0-2026-10-05/`:
|
||||
all pass markers, round trip sha `8cebc42e…`, token leak 0 with a working control, teardown to `virgin`). **Vouched**
|
||||
(agent 0.147.0, golden 0.297.0, min_agent 0.131.0) and **floors** 0.297.0 for demo-hp, demo-felhom, tester-1.
|
||||
**Delivered:** demo-hp and demo-felhom `felhom-controller:0.297.0 … (healthy)`; Tester 1 reports Controller 0.297.0
|
||||
(„Controller frissítve: 0.296.0 → 0.297.0").
|
||||
|
||||
**catalog** (`4828dc7`; CI run 1368 lost by R-887, re-run success): R-593, R-760, R-594, R-605, R-781, R-806 (scheme half;
|
||||
row narrowed), plus a stale runner test (expected 11 gates, 12 exist) and a test that never ran (outside its class) —
|
||||
fixed, not filed.
|
||||
|
||||
**Not done, and why:** R-469 and R-605's exit-code line in the catalog's `CLAUDE.md` — **the permission check refused
|
||||
the instruction-file edit**; the operator is asked (rule 5). R-126 needs an operator choice. R-325 needs a same-step
|
||||
felhom.eu gate change (left). R-377 (CONTEXT headings) not attempted. Installer rows (R-179, R-180, R-275, R-276, R-306,
|
||||
R-130, R-310, R-881), R-136 (logs every operator out), R-502 (Docker in CI), R-798 (a live app definition) and the
|
||||
larger controller rows (R-492, R-569, R-575, R-615, R-616, R-498, R-718) were left on purpose.
|
||||
|
||||
**Opened:** R-888 — two report fields the hub never reads (a decision). **Seen, not a row:** Tester 1's crash guard reads
|
||||
TRIPPED since 07:57Z — the morning's two deliberate test crashes; it re-arms by itself after 24 h (`runbooks/crash-guard.md`).
|
||||
|
||||
## The 23 rows the first burn-down could not check
|
||||
|
||||
Checked from source by a read-only agent (`audits/burndown2-2026-10-05/unchecked-results.jsonl`). Closed: R-350 (duplicate
|
||||
of R-132, facts merged). Facts added to the open rows R-607, R-883, R-886, R-884, R-756, R-91, R-338, R-488. The three
|
||||
„not worth it" ones are on STATUS for the operator. The rest need a live box reading (the settle command is in the table).
|
||||
|
||||
| Row | Group | Evidence / how to settle (abridged) |
|
||||
|---|---|---|
|
||||
| R-76 | UNCHECKABLE-FROM-SOURCE | Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 FileBrowserImage = "gtstef/filebrowser:1.5.6-stable". The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --gre |
|
||||
| R-91 | UNCHECKABLE-FROM-SOURCE | Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of delet |
|
||||
| R-209a | UNCHECKABLE-FROM-SOURCE | Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test. — settle: uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h / |
|
||||
| R-337 | NOT-WORTH-IT | The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 b, err : |
|
||||
| R-375 | NOT-WORTH-IT | The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is pvesm status showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 if st == nil // st.Type == "pbs" // st.Avail <= 0 { return true, "" } (space preflight sk |
|
||||
| R-488 | STILL-TRUE-SMALL | The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded interval := 5 * time.Second and time.Sleep(3 * time.Second) // initial settling time, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go: |
|
||||
| R-504 | UNCHECKABLE-FROM-SOURCE | Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html. — settle: curl -sI https://iso.felhom.eu/ / head -1 |
|
||||
| R-644 | UNCHECKABLE-FROM-SOURCE | Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs --deployment-password; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches |
|
||||
| R-814 | UNCHECKABLE-FROM-SOURCE | Hetzner account state; nothing in source records a deletion. — settle: Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status |
|
||||
| R-815 | UNCHECKABLE-FROM-SOURCE | PBS server-side state on ep0; no GC completion record in the docs (grep). — settle: ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 / grep -i garbage' |
|
||||
| R-884 | UNCHECKABLE-FROM-SOURCE | Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml prom/prometheus:v3.14.0 -> v3.15.0 (monitoring.yaml:419), and the monitoring Application has no automated syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an un |
|
||||
| R-132 | UNCHECKABLE-FROM-SOURCE | Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key operator_password_hash, setSetting :2200-2207 writes updated_at = datetime('now'). No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31 |
|
||||
| R-298 | UNCHECKABLE-FROM-SOURCE | Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 if(d.role==='user-data'){ else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/ |
|
||||
| R-338 | UNCHECKABLE-FROM-SOURCE | nodes.md:86-88 still claims demo-hp is on the R-50 island (local_api on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent con |
|
||||
| R-350 | DUPLICATE | of R-132 — Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. |
|
||||
| R-542 | NOT-WORTH-IT | Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 initialize = append(initialize, c) // every unclaimed disk can be initialized; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkf |
|
||||
| R-607 | STILL-TRUE-SMALL | Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY if len(newApps) > 0 // len(updated) > 0, and :257-258 says 'nincs változás' when both are empty. updated counts stack-dir copies only (copyTemplates, :447 updated = append(updated, appName) a |
|
||||
| R-683 | UNCHECKABLE-FROM-SOURCE | Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power- |
|
||||
| R-756 | UNCHECKABLE-FROM-SOURCE | Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH |
|
||||
| R-862 | UNCHECKABLE-FROM-SOURCE | Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83). — settle: Ask the operator wheth |
|
||||
| R-882 | UNCHECKABLE-FROM-SOURCE | Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it). — settle: sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p Act |
|
||||
| R-883 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/rec |
|
||||
| R-886 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root nobody image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts |
|
||||
|
||||
## Teardown
|
||||
|
||||
Machines: drill VM — build guest destroyed, token/scripts/log shredded, powered off, disk back on `virgin`. Boxes: only the
|
||||
normal deliveries above. Hub: only the deploy, the vouch and the floors. Scratch secrets (hub password file, hub key file,
|
||||
signed envelopes) are shredded at the end of the session.
|
||||
@@ -3,7 +3,34 @@
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
|
||||
sent to it.**
|
||||
|
||||
**Updated 2026-10-05 (late night, burn-down round 2 — in progress): 249 open rows after your answer (was 292).**
|
||||
**Updated 2026-10-05 (late night, burn-down round 2): every box of ours healthy. The open-items list is at 199 (was 292
|
||||
at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10-05.md`.**
|
||||
|
||||
## Tonight, last (2026-10-05): the list at 199
|
||||
|
||||
**What happened:**
|
||||
- **Your answer is recorded** (below): 43 rows closed as accepted.
|
||||
- **51 more rows fixed and closed**, with one release per repository, delivered the normal way: hub **0.137.0**
|
||||
(live), agent **0.147.0** (on demo-hp, demo-felhom and Tester 1, root files too), controller **0.297.0** (on all
|
||||
three), new-install image **0.297.0** (baked, checked, approved), app catalog updated.
|
||||
- **R-124 is fixed** (your ruling): the recovery recipe now writes the backup-server namespace the way the server reads it.
|
||||
- **The CI fault (R-887) is understood:** when Gitea is busy, the CI runner's request for work can time out after Gitea
|
||||
already gave it the job; the job is then never run and is failed 10–13 minutes later. Gitea was busy because of an
|
||||
outside web crawler and because of MY CI checks, which asked too much — mine now ask once a minute in one small request.
|
||||
- **Two of my own CI misses tonight:** a new check needed Go, which the CI machine does not have (fixed); a red run was
|
||||
the fault above (re-run passed).
|
||||
|
||||
**The numbers:** 292 before → **199 after**; 1 opened; 94 closed.
|
||||
|
||||
**Needs you (none urgent; if you do nothing, each stays open as it is):**
|
||||
1. **R-469** — a one-paragraph rewording in the app catalog's instruction file; my permission check refused editing
|
||||
instruction files. Say "go" and it is done in a minute.
|
||||
2. **R-126** — should a network share be offered as an export destination? Two options in the row; pick one.
|
||||
3. **R-888** — two facts the boxes report that the hub never shows (installed-app list, retired drives). Needed or not?
|
||||
4. **R-887** — CI: raise the runner's fetch timeout and/or slow the crawler on Gitea's public pages. If nothing: now and
|
||||
then a CI run fails without running; it can be re-run.
|
||||
5. Three more rows look „not worth doing" (R-337, R-375, R-542 — the check found each is by design or harmless). Close
|
||||
them as accepted? If you say nothing they stay.
|
||||
|
||||
## Your answer to the burn-down list (2026-10-05 18:23), recorded
|
||||
|
||||
|
||||
@@ -497,3 +497,9 @@ rc=0
|
||||
--- PASS: TestR365_BannerSaysDueAtZeroDays (0.12s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.132s
|
||||
|
||||
### CI fix (lead, 2026-10-05 ~18:20Z): the new gofmt gate was INCONCLUSIVE on the CI runner (no Go) — CI run 1371 FAILED
|
||||
Fix: in CI (GITEA_ACTIONS/GITHUB_ACTIONS=true) with no gofmt reachable the gate prints "NOT CHECKED in CI" and exits 0;
|
||||
elsewhere a missing gofmt stays INCONCLUSIVE. Decoys gofmt/ci-without-go and gofmt/dev-without-go (empty PATH).
|
||||
Simulated CI (env -i, empty PATH, GITEA_ACTIONS=true): "gofmt gate NOT CHECKED in CI …" rc=0.
|
||||
Red-proof (CI branch disabled): FAIL: gofmt/ci-without-go: want rc=0 and 'NOT CHECKED in CI', got rc=2.
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
== controller delivery 2026-10-05T18:28:59Z
|
||||
demo-hp gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 32 seconds (healthy)
|
||||
demo-felhom gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 35 seconds (healthy)
|
||||
tester-1 (hub customer page, version strings seen): 6 0.297.0 4 0.296.0 4 0.295.0
|
||||
@@ -0,0 +1,11 @@
|
||||
== hub deploy 2026-10-05T18:11:56Z
|
||||
image=gitea.dooplex.hu/admin/felhom-hub:0.137.0
|
||||
sync=Synced health=Healthy op=Succeeded rev=d75ad0fdf3daf3f3b2a1690746d9a6a70ee4104c
|
||||
2026/10/05 20:11:05 [INFO] felhom-hub 0.137.0 starting
|
||||
healthz 200
|
||||
2026/10/05 20:11:06 [INFO] osupdates: the Docker engine set is approved only by the operator, after 2 healthy ring-0 night(s)
|
||||
== System page 2026-10-05T18:12:05Z: root-files / agent cells
|
||||
73: 0 armed unknown 0.142.0 → 0.147.0 (since 2026-10-05)
|
||||
89: 0 armed 0.147.0 0.147.0
|
||||
105: 2 armed 0.147.0 0.147.0
|
||||
121: 2 TRIPPED 2026-10-05T07:57:17Z 0.147.0 0.147.0
|
||||
@@ -0,0 +1,11 @@
|
||||
== vouch 2026-10-05T18:27:46Z: agent 0.147.0, golden 0.297.0, min_agent 0.131.0
|
||||
HTTP/1.1 303 See Other
|
||||
Location: /configuration?flash=artifacts_set
|
||||
== floors 2026-10-05T18:28:16Z: POST /customers/<id>/floor min_controller_version=0.297.0 min_agent=0.131.0
|
||||
demo-hp: Location: /customers/demo-hp?flash=floor_set
|
||||
demo-felhom: Location: /customers/demo-felhom?flash=floor_set
|
||||
tester-1: Location: /customers/tester-1?flash=floor_set
|
||||
2026/10/05 20:28:16 [INFO] Artifact manifest set: agent=0.147.0 golden=0.297.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d"
|
||||
2026/10/05 20:28:16 [INFO] Customer demo-hp controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
|
||||
2026/10/05 20:28:17 [INFO] Customer demo-felhom controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
|
||||
2026/10/05 20:28:17 [INFO] Customer tester-1 controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
|
||||
@@ -0,0 +1,23 @@
|
||||
{"id": "R-76", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 `FileBrowserImage = \"gtstef/filebrowser:1.5.6-stable\"`. The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --grep 'R-76|setgid' in controller: only 2026-06 commits). Whether 1.5.6 still drops setgid is runtime behaviour of the image.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- find /mnt/felhom-drives -path '*/userdata/*' -mindepth 3 -maxdepth 5 -type d ! -perm -2000 -printf '%m %u:%g %p\\n'\" | head (any folder a customer made in FileBrowser showing 755 without setgid = still true on 1.5.6)", "minutes_spent": 3}
|
||||
{"id": "R-91", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word.", "dup_of": null, "unique_facts": "Stale doc to fix in the same commit as the delete: documentation/runbooks/offsite-endpoint.md:24 still says datastore `felhom-offsite` is at `/srv/pbs-felhom`; the real path since 2026-07-27 is `/mnt/pbs-datastore` (RUNBOOK-ep0-datastore-volume-2026-07-27.md:8). CONTEXT.md:3666 is the other line to change.", "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'du -sh /srv/pbs-felhom 2>&1; proxmox-backup-manager datastore list; df -h /'", "minutes_spent": 4}
|
||||
{"id": "R-209a", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h /", "minutes_spent": 1}
|
||||
{"id": "R-337", "sev": "P4", "category": "Monitoring & notifications", "group": "NOT-WORTH-IT", "evidence": "The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 `b, err := tier.Service.BackupWithSnapshotHook(...)` then :901 `s.store.RecordBackup(b)` (grep RecordBackup: no other caller). So it is NOT a collection cadence; the status appears when the agent's own job finishes (WaitTask + archive resolve, internal/backup/runner.go:205-232), and a backup the agent did not run itself never appears. The 4-min demo-hp skew is the gap between PBS writing the manifest and the job returning -- unmeasured.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: measure why demo-hp's job returned ~4 min after the manifest landed; cost: a constructed live repro on demo-hp with task-log timing; if never: during an incident the box's backup status can trail the PBS manifest by minutes after an out-of-schedule run, and a run started outside the agent never shows; pick: close with the source fact above written into the row (status = agent job end, not polling), reopen only if a lag is seen on a scheduled run.", "settle_cmd": null, "minutes_spent": 7}
|
||||
{"id": "R-375", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is `pvesm status` showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 `if st == nil || st.Type == \"pbs\" || st.Avail <= 0 { return true, \"\" }` (space preflight skips PBS and fails open on 0), and the same report says the hub's gauge reads the real 3.7/97.9 GB.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: confirm on ep0 that the namespace-scoped token lacks Datastore.Audit; cost: a read-only ep0 session; if never: the PVE UI on a box shows 0/0/0 for felhom-pbs, which no Felhom code reads (runner.go:265); pick: close as cosmetic with this pointer.", "settle_cmd": "(if ever wanted) ssh root@ep0 'proxmox-backup-manager acl list' ; on a box: pvesm status --storage felhom-pbs", "minutes_spent": 6}
|
||||
{"id": "R-488", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded `interval := 5 * time.Second` and `time.Sleep(3 * time.Second) // initial settling time`, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:90, restore_unit.go:471. No commit since 2026-09-13 touching internal/backup mentions R-488/test speed. Runtime (333 s) was NOT re-measured here (read-only).", "dup_of": null, "unique_facts": null, "small_fix": "controller: add Manager fields healthSettle/healthInterval (defaults 3s/5s, set in the constructor) used by waitForHealthy; a test helper (the existing Manager fixture constructor) sets them to 0/10ms. Test: TestWaitForHealthy_DefaultsAreProduction asserts a fresh Manager has 3s/5s (pins production), and measure `go test ./internal/backup` wall time before/after in the commit message (expect the 89 >=1 s tests to drop). Run with the R-650 docker-free seams.", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}
|
||||
{"id": "R-504", "sev": "P4", "category": "Install & onboarding", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "If 404 is confirmed it is cosmetic (row's own re-rank); the operator could close it as accepted rather than add a Cloudflare rule.", "settle_cmd": "curl -sI https://iso.felhom.eu/ | head -1", "minutes_spent": 2}
|
||||
{"id": "R-644", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs `--deployment-password`; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches the crash text. Not proven.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- docker ps -a --filter name=gokapi --format '{{.Names}} {{.Status}}'; pct exec 9202 -- grep -s -c '\\\"Password\\\":\\\"\\\"' /opt/docker/stacks/gokapi/config/config.json\"", "minutes_spent": 3}
|
||||
{"id": "R-814", "sev": "P4", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Hetzner account state; nothing in source records a deletion.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status", "minutes_spent": 1}
|
||||
{"id": "R-815", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "PBS server-side state on ep0; no GC completion record in the docs (grep).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 | grep -i garbage'", "minutes_spent": 2}
|
||||
{"id": "R-884", "sev": "P4", "category": "Monitoring & notifications", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live.", "dup_of": null, "unique_facts": "Cause candidate: Renovate 53c6e99 prometheus v3.15.0 merged 2026-10-03, never synced because monitoring has no auto-sync in git.", "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n mon-system get deploy prometheus -o jsonpath='{.spec.template.spec.containers[0].image}' (v3.14.0 => the diff is the Renovate bump)", "minutes_spent": 5}
|
||||
{"id": "R-132", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key `operator_password_hash`, setSetting :2200-2207 writes `updated_at = datetime('now')`. No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "On a copy of the hub DB (hub.db + -wal + -shm): sqlite3 hub.db \"SELECT updated_at FROM hub_settings WHERE key='operator_password_hash'\" -- rotated only if later than 2026-09-18 (the last exposure, R-580 folded here); no row = still the ConfigMap seed", "minutes_spent": 4}
|
||||
{"id": "R-298", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 `if(d.role==='user-data'){` else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/role.go:172-186 -- a dir storage (felhom-backup) is user-data unless its backing device is on the system disk; disks.go:1236-1240 deliberately does not reclassify the backup-target drive. role.go unchanged since 2026-08-09. demo-hp was reprovisioned before 2026-09-21, so the 2026-08-10 topology may no longer hold.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp 'grep -A2 felhom-backup /etc/pve/storage.cfg; lsblk -no PKNAME $(findmnt -no SOURCE /) ; lsblk -no PKNAME $(findmnt -no SOURCE /mnt/nvme-1tb)' (same parent disk => role=system => the page still locks it => R-298 true; different => user-data => not reproducible on this box)", "minutes_spent": 8}
|
||||
{"id": "R-338", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"grep -E 'listen_addr|island_' /etc/felhom-agent/agent.json; pct config 9201 | grep ^net; ip -br link show master vmbr9\"", "minutes_spent": 6}
|
||||
{"id": "R-350", "sev": "P3", "category": "Security & access", "group": "DUPLICATE", "evidence": "Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence.", "dup_of": "R-132", "unique_facts": "(1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked.", "small_fix": null, "not_worth": null, "settle_cmd": null, "minutes_spent": 3}
|
||||
{"id": "R-542", "sev": "P3", "category": "Storage & devices", "group": "NOT-WORTH-IT", "evidence": "Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 `initialize = append(initialize, c) // every unclaimed disk can be initialized`; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkfs wrapper explicitly allows it: configs/felhom-mkfs-guarded.sh:59-60 'Mounts under /mnt/felhom-drives are our own drives (the agent detaches before a re-init) -> allowed'. Controller passes initialize through untouched (agent_disk_handlers.go:158-162). `already_mounted` is only set for controller-contributed stores (agentapi/client.go:447-451, omitempty) -- so 'null' is expected for agent candidates.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: drop felhom-mounted disks from `initialize`; cost: reverses the R-220/re-init design (wrapper comment) and needs a decision on how a registered drive is re-initialized; if never: the raw endpoint lists a registered drive under initialize while the page (customer view) filters it -- only a session reading the raw endpoint is misled; pick: close as by-design, add one comment line at disks.go:438 saying a registered felhom drive is listed here deliberately.", "settle_cmd": null, "minutes_spent": 9}
|
||||
{"id": "R-607", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-SMALL", "evidence": "Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY `if len(newApps) > 0 || len(updated) > 0`, and :257-258 says 'nincs változás' when both are empty. `updated` counts stack-dir copies only (copyTemplates, :447 `updated = append(updated, appName)` after a hash mismatch); for a deployed+pinned app whose catalog moved, renderSource (:457-469 table) copies the STORED definition, so the hash matches and nothing is 'updated' although the git cache moved. CatalogImages is read from the catalog cache only inside ScanStacks (stacks/manager.go:666-672, assigned :690). So (a) the message measures the stack dir, not the catalog; (b) CatalogImages refreshes only on a ScanStacks, which this sync does not trigger. The 29 s nextcloud case (needed several rounds) is not explained by this.", "dup_of": null, "unique_facts": null, "small_fix": "controller sync.go: record the catalog git HEAD before and after gitCloneOrPull; if it moved, call s.rescanFn() even when newApps/updated are empty, and say 'Katalógus frissítve — az alkalmazások nem változtak' (and EN) instead of 'nincs változás'. Test (red first): a Syncer with a fake pull that moves HEAD and a frozen pinned app (renderSource returns the stored definition) asserts rescanFn was called once and the message is not the no-change one; and a no-move pull asserts rescanFn NOT called.", "not_worth": null, "settle_cmd": null, "minutes_spent": 9}
|
||||
{"id": "R-683", "sev": "P3", "category": "App updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power-cut-mid-update drill: night-2026-10-04/MORNING-NOTE.md:64 'A1 power cut mid-update (demo-hp): NOT RUN'; DRILL-night-2026-09-25.md:102 was a cut during romm's verifying, not checked for the backup choice.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Not a read-only command: the next power-cut-during-update drill with the controller log saved at arm time, then grep \"phase backing-up\\|Tier-2\" in the saved pre-cut log", "minutes_spent": 7}
|
||||
{"id": "R-756", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 `if !m.DriveLive(hddPath)` -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 `return m.isMountPoint(hddPath)` -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names `.../scratch_hdd/userdata/calibre-web`, so on 9202 HDD_PATH is a per-app subfolder, which can never be a mount point -> the 409 is certain for that value. Open: whether that HDD_PATH was written by the product (handlers.go:550 prefill from place.Drive) or by the test venue.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- grep -H HDD_PATH /opt/docker/stacks/calibre-web/app.yaml /opt/docker/stacks/grimmory/app.yaml; pct exec 9202 -- findmnt -no TARGET,SOURCE /mnt/felhom-drives/scratch_hdd\" (HDD_PATH = per-app subfolder => product/venue wrote a non-root HDD_PATH; HDD_PATH = drive root and not mounted => scratch drive not registered)", "minutes_spent": 10}
|
||||
{"id": "R-862", "sev": "P3", "category": "Box system & updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Ask the operator whether the three bootstrap commands ran; read-only proof: the hub's operator view of Tester 2 (OS/config-bundle sha vs the vouched sha), or whether os_config_bundle_behind has fired for it", "minutes_spent": 2}
|
||||
{"id": "R-882", "sev": "P3", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p ActiveEnterTimestamp (instance-manager older than a k3s/containerd/iscsid restart => the stale-PID state can recur)", "minutes_spent": 2}
|
||||
{"id": "R-883", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/recipe-importer:latest (pull Always :27); adventurelog-system/adventurelog.yaml:100 and :256 adventurelog-backend/frontend:latest (pull Always :101,:257); jarrs-system/jarr-dev.yaml:311,:345,:630 gitea.dooplex.hu/admin/jarr:latest (pull Always). Zipline fixed (90f60e4, 4c8ec7a). The repo's own rule homelab-manifests/CLAUDE.md:120 'Image tags always pinned'. Helm values files (external-dns, pihole, plex, authentik, cnpg) not checked for tag fields beyond a `tag: latest|dev|empty` grep (no hits).", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests only: for each line above, read the running digest/version (`sudo kubectl get pod -n <ns> -o jsonpath='{..imageID}'`), pin that exact tag (or @sha256 for the self-built gitea.dooplex.hu jarr/recipe-importer images, which have no version tags), add the version-checker match-regex annotation per CLAUDE.md:120, and set imagePullPolicy IfNotPresent. Test: `grep -rnE 'image:.*(:latest|:dev)\\s*$' --include=*.yaml .` returns nothing; after ArgoCD sync each pod's imageID equals the pre-change one (no upgrade). Operator-owned DooPlex change: needs the operator's word.", "not_worth": null, "settle_cmd": null, "minutes_spent": 6}
|
||||
{"id": "R-886", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager real storage').", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests: add pod `securityContext: {fsGroup: 65534, fsGroupChangePolicy: OnRootMismatch}` (65534 = nobody, the image's user -- confirm with `sudo kubectl -n mon-system exec deploy/alertmanager -- id`) to the alertmanager Deployment. Test (consequence, per CLAUDE.md): create a silence via amtool/API, delete the pod, after it returns the silence is still listed AND the log has no 'Running maintenance failed ... permission denied' within 15 min (positive observable: a 'maintenance done' line).", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}
|
||||
@@ -103,6 +103,29 @@ The full text of every row below: `git show e8c56c44:documentation/backlog/OPEN-
|
||||
| **R-269** | **A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): token store reloads on growth before answering; `TestTokenStore_RotatedOutTokenRejectedFirst`; red-proof agent-red-proofs.txt |
|
||||
| **R-317** | **The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): dnsmasq install probed by its service unit; `TestEnsureDnsmasq_*`; red-proof agent-red-proofs.txt |
|
||||
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-132 (its unique facts moved there) | Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. |
|
||||
| **R-591** | **[P3-LOW] `Stack.Copy()` is a deep copy with one shallow field, and the field is new.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `deepCopyStack` copies every reference in Meta (i18n overlay and 11 more found by `TestDeepCopyStackMetaSharesNoReference`); `TestDeepCopyStackI18nIsNotShared` |
|
||||
| **R-568** | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: disk-health rows sorted by durable id; `TestDiskHealthRows_OrderIsStable` |
|
||||
| **R-567** | **[P3-LOW] The two drive wizard pages (/storage/init, /storage/attach) do not highlight the Tárhely menu group — the sidebar reads as if the household left the storage section.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: storage wizard pages open the Storage nav group; `TestStorageWizardPages_OpenTheStorageNavGroup`; two parity fixtures re-captured (nav only) |
|
||||
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: fill-watch also every 10 min (`sched.Every`), daily + start-up kept; `TestFillWatchRunsOnAnInterval` |
|
||||
| **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: same change as R-363 (the interval watch); `TestFillWatchRunsOnAnInterval` |
|
||||
| **R-10** | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state `idea`.** Moved verbatim; nothing added (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: dump directory fsynced after the rename; `TestDumpOneTo_SyncsTheDumpDirectoryAfterRename` |
|
||||
| **R-552** | **[P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: remove clears the interrupted-restore notice (`ClearInterruptedRestore`, wired in the remove path); `TestR552_RemoveClearsTheInterruptedRestoreNotice` |
|
||||
| **R-251** | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: off-site marker tag not listed as an app; `TestR251_MarkerTagIsNotAnApp` |
|
||||
| **R-104** | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a lock surviving the self-heal is classed `locked` with its own cause line; `TestR104_SurvivingLockIsNamed` |
|
||||
| **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: password deploy fields served as required by the API (fresh metadata per call); `TestR619_PasswordFieldIsServedAsRequired` |
|
||||
| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a restore onto a detached drive names the drive; `TestR362_DetachedDriveIsNamed` |
|
||||
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: files-restore refusal names the second drive's whole copy (and is in both languages now); `TestR675_RefusalNamesTheWholeCopy` |
|
||||
| **R-256** | **C2 — „A mentéskezelő nem elérhető." names no route at all.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `flash.offbox.mgr_unavailable/_unreachable` name a route (hu+en); `TestR256_R257_OffboxRefusalsNameARoute` |
|
||||
| **R-257** | **C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `flash.offbox.not_orphaned` reworded (hu+en); `TestR256_R257_OffboxRefusalsNameARoute` |
|
||||
| **R-240** | **A backup that covered nothing calls itself „Sikeres".** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a zero-selection run no longer says „Sikeres"; `TestR240_ZeroSelectionRunDoesNotSaySuccess` |
|
||||
| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a 0-day deletion countdown says it is due — page AND top bar (`TestR365_OverdueCountdownIsNotFutureTense`, `TestR365_BannerSaysDueAtZeroDays`) |
|
||||
| **R-425** | **`offbox_rename_gate.py` scans a fixed three-entry `FILES` list.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: offbox rename gate finds files by pattern and judges the bundle; 3 decoys |
|
||||
| **R-565** | **[P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: English-page test knows ASCII Hungarian words; a real `mp` leak became a key; `TestI18nEnglishPages` |
|
||||
| **R-564** | **[P3-LOW] The retrieval-promise gate's Hungarian stems cannot see a SPLIT verb — „csak akkor állíthatók vissza", „hozod vissza" — so those Hungarian sentences were never scanned; the English translation exposed them.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: retrieval-promise gate knows split-verb Hungarian; 7 occurrences registered; 2 decoys |
|
||||
| **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: Go-side check for HTML-escapable values; `TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping` |
|
||||
| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `scripts/gofmt_gate.py` (NOT CHECKED out loud on the Go-less CI runner, INCONCLUSIVE elsewhere; 3 decoys); 12 files formatted |
|
||||
| **R-208** | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: controller and hub Dockerfiles: `ARG VERSION…` just above `go build`; `TestR208_DockerfileVersionArgsSitBelowModuleDownload`, felhom.eu `scripts/test_dockerfile_arg_order.py` (hub v0.137.0 deployed) |
|
||||
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** (P4) | CLOSED 2026-10-05 — CHECKED, NOTHING LEFT (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: swept: none of the six candidate files has a date literal feeding an assertion against the real clock (identity/format/ordering checks only) — nothing to change; the faked-future-date CI idea is a separate, larger job |
|
||||
|
||||
---
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,5 @@
|
||||
== round trip 2026-10-05T18:26:49Z: anonymous GET .../generic/felhom-golden/0.297.0/golden.tar.zst
|
||||
HTTP 200
|
||||
bytes 648208028
|
||||
sha256 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
|
||||
printed 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
|
||||
@@ -0,0 +1,52 @@
|
||||
# Golden 0.297.0 — bake + publish + vouch, 2026-10-05 (night, burn-down round 2)
|
||||
|
||||
Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–5, in the drill VM on DooPlex.
|
||||
|
||||
| | Previous (`../golden-0.296.0-2026-10-05/`) | This bake |
|
||||
|---|---|---|
|
||||
| `build-golden.sh` | sha256 `645b3b659cba…` | same file, unchanged (agent repo `configs/build-golden.sh`) |
|
||||
| Controller | `felhom-controller:0.296.0` | **`felhom-controller:0.297.0`** (MinAgent 0.131.0, unchanged) |
|
||||
| Docker engine | the approved set `os-docker-20261004-142842` | same pinned set (same `GOLDEN_DOCKER_PKGS` as the 0.296.0 bake) |
|
||||
| Guest packages | template | template — `GOLDEN_GUEST_PKGS` EMPTY |
|
||||
|
||||
## Launch
|
||||
|
||||
- No qemu running before; drill VM reverted to `virgin`, cold-booted per §4.0; `pveversion` = `pve-manager/9.2.2`.
|
||||
- `pveam update` → `update successful`; `pveam available` listed `debian-13-standard_13.6-1_amd64.tar.zst` (downloaded).
|
||||
- `/root/bake-run.sh` reads the token from the file; transient unit `golden-bake`. Token copied file → file (`scp`);
|
||||
`systemctl show golden-bake -p Environment -p ExecStart | grep -c -F <token>` = **0**.
|
||||
|
||||
## Pass markers (from `bake.log`, this folder)
|
||||
|
||||
```
|
||||
docker OK (overlay2; data-root /var/lib/docker)
|
||||
INFO: including mount point rootfs ('/') in backup
|
||||
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||
[golden] upload OK (HTTP 201)
|
||||
GOLDEN_VERSION=0.297.0
|
||||
GOLDEN_SHA256=8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
|
||||
```
|
||||
|
||||
No `excluding` and no `FATAL` in the log (grep count 0).
|
||||
|
||||
## Round trip — `02-round-trip.txt`
|
||||
|
||||
Anonymous GET of `…/generic/felhom-golden/0.297.0/golden.tar.zst`: HTTP 200, 648 208 028 bytes, sha256 equals the printed one.
|
||||
|
||||
## Secrets
|
||||
|
||||
Saved-log leak grep for the literal token: **0**; positive control (a throwaway copy with the token appended): **1**,
|
||||
copy shredded.
|
||||
|
||||
## Vouch (step 5)
|
||||
|
||||
`POST /configuration/artifacts` (Basic + `X-Felhom-Operator`, hub v0.137.0): agent **0.147.0**, golden **0.297.0**,
|
||||
`min_agent` **0.131.0** → `303 flash=artifacts_set`; hub log `Artifact manifest set: agent=0.147.0 golden=0.297.0
|
||||
min_agent="0.131.0" … bundle_sha="326527d0…"`. Per-customer floors 0.297.0 (declared MinAgent 0.131.0) for demo-hp,
|
||||
demo-felhom, tester-1 (`../../audits/burndown2-2026-10-05/delivery/vouch-golden-floors.txt`). The global floor unchanged.
|
||||
|
||||
## Teardown
|
||||
|
||||
`pct destroy 9100 --purge` (rc 0); `shred -u` of the token, runner script, bake script and log in the VM (log copied off
|
||||
first); `poweroff`; qemu gone (`ps -eo comm | grep -c qemu-system-x86` = 0); `qemu-img snapshot -a virgin`. Host:
|
||||
nothing provisioned.
|
||||
@@ -0,0 +1,339 @@
|
||||
[golden] build-golden.sh v3.2.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.297.0
|
||||
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
|
||||
Logical volume "vm-9100-disk-0" created.
|
||||
Logical volume pve/vm-9100-disk-0 changed.
|
||||
Creating filesystem with 8388608 4k blocks and 2097152 inodes
|
||||
Filesystem UUID: 3f3e66a1-e05c-42f3-91af-3ff9fc5c8f65
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
4096000, 7962624
|
||||
Logical volume "vm-9100-disk-1" created.
|
||||
Logical volume pve/vm-9100-disk-1 changed.
|
||||
Creating filesystem with 6291456 4k blocks and 1572864 inodes
|
||||
Filesystem UUID: 9327fbdd-d6c5-4299-a41f-3cb56b1238b3
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
|
||||
Total bytes read: 553512960 (528MiB, 115MiB/s)
|
||||
Detected container architecture: amd64
|
||||
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
|
||||
done: SHA256:6hCAi5WjsL3daO1FsrQR2bko1QAWr49AoJUmsNXkGSY root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
|
||||
done: SHA256:STGSy09tGfb2oiy9XwC4UHzFcSawljx6IhtUka8tiJI root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
|
||||
done: SHA256:rh5fSFDYSXH/CINtl163usIqQy+m082joU7xzA5LENU root@felhom-golden
|
||||
[golden] starting + installing Docker (official repo, trixie channel) …
|
||||
[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docker-ce=5:29.8.2-1~debian.13~trixie docker-ce-cli=5:29.8.2-1~debian.13~trixie docker-ce-rootless-extras=5:29.8.2-1~debian.13~trixie docker-compose-plugin=5.6.0-1~debian.13~trixie
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
installed: containerd.io 2.3.6-1~debian.13~trixie
|
||||
installed: docker-buildx-plugin 0.37.1-1~debian.13~trixie
|
||||
installed: docker-ce 5:29.8.2-1~debian.13~trixie
|
||||
installed: docker-ce-cli 5:29.8.2-1~debian.13~trixie
|
||||
installed: docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie
|
||||
installed: docker-compose-plugin 5.6.0-1~debian.13~trixie
|
||||
[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a
|
||||
[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49
|
||||
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
|
||||
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
|
||||
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
|
||||
Unable to find image 'hello-world:latest' locally
|
||||
latest: Pulling from library/hello-world
|
||||
4f55086f7dd0: Pulling fs layer
|
||||
4f55086f7dd0: Download complete
|
||||
4f55086f7dd0: Pull complete
|
||||
Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8
|
||||
Status: Downloaded newer image for hello-world:latest
|
||||
docker OK (overlay2; data-root /var/lib/docker)
|
||||
live-restore: on
|
||||
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
|
||||
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
|
||||
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
|
||||
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.297.0 (no registry cred at deploy) …
|
||||
|
||||
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
|
||||
Configure a credential helper to remove this warning. See
|
||||
https://docs.docker.com/go/credential-store/
|
||||
|
||||
0.297.0: Pulling from admin/felhom-controller
|
||||
774043ccc8cc: Pulling fs layer
|
||||
ab6b448d4be9: Pulling fs layer
|
||||
23a5bfa58353: Pulling fs layer
|
||||
862a57157567: Pulling fs layer
|
||||
6db4169d1fd9: Pulling fs layer
|
||||
167f80584563: Pulling fs layer
|
||||
862a57157567: Waiting
|
||||
6db4169d1fd9: Waiting
|
||||
167f80584563: Waiting
|
||||
774043ccc8cc: Verifying Checksum
|
||||
774043ccc8cc: Download complete
|
||||
862a57157567: Verifying Checksum
|
||||
862a57157567: Download complete
|
||||
23a5bfa58353: Verifying Checksum
|
||||
23a5bfa58353: Download complete
|
||||
167f80584563: Verifying Checksum
|
||||
167f80584563: Download complete
|
||||
6db4169d1fd9: Verifying Checksum
|
||||
6db4169d1fd9: Download complete
|
||||
ab6b448d4be9: Verifying Checksum
|
||||
ab6b448d4be9: Download complete
|
||||
774043ccc8cc: Pull complete
|
||||
ab6b448d4be9: Pull complete
|
||||
23a5bfa58353: Pull complete
|
||||
862a57157567: Pull complete
|
||||
6db4169d1fd9: Pull complete
|
||||
167f80584563: Pull complete
|
||||
Digest: sha256:23e4e0ffd9c9e2df28196e55bce8b89d4dfc1f9d4f8e52ca3ae2872e365073a0
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.297.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.297.0
|
||||
[golden] asking the controller which infra images it manages …
|
||||
[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
|
||||
v3.7.13: Pulling from library/traefik
|
||||
e2de96513ba9: Pulling fs layer
|
||||
b686a4f73445: Pulling fs layer
|
||||
78cb21c375ca: Pulling fs layer
|
||||
acb2f33459b1: Pulling fs layer
|
||||
acb2f33459b1: Waiting
|
||||
e2de96513ba9: Verifying Checksum
|
||||
e2de96513ba9: Download complete
|
||||
b686a4f73445: Download complete
|
||||
acb2f33459b1: Verifying Checksum
|
||||
acb2f33459b1: Download complete
|
||||
e2de96513ba9: Pull complete
|
||||
78cb21c375ca: Verifying Checksum
|
||||
78cb21c375ca: Download complete
|
||||
b686a4f73445: Pull complete
|
||||
78cb21c375ca: Pull complete
|
||||
acb2f33459b1: Pull complete
|
||||
Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0
|
||||
Status: Downloaded newer image for traefik:v3.7.13
|
||||
docker.io/library/traefik:v3.7.13
|
||||
2026.9.3: Pulling from cloudflare/cloudflared
|
||||
2cc7ee286bf3: Pulling fs layer
|
||||
c172f21841df: Pulling fs layer
|
||||
218cf840d0d9: Pulling fs layer
|
||||
f6069939f718: Pulling fs layer
|
||||
d6b1b89eccac: Pulling fs layer
|
||||
2780920e5dbf: Pulling fs layer
|
||||
7c12895b777b: Pulling fs layer
|
||||
3214acf345c0: Pulling fs layer
|
||||
52630fc75a18: Pulling fs layer
|
||||
dd64bf2dd177: Pulling fs layer
|
||||
b839dfae01f6: Pulling fs layer
|
||||
ebddc55facdc: Pulling fs layer
|
||||
c4bc6f35ff5e: Pulling fs layer
|
||||
b96fe2995f90: Pulling fs layer
|
||||
58c0c263dc73: Pulling fs layer
|
||||
bd8962e29291: Pulling fs layer
|
||||
cac2ae0193cb: Pulling fs layer
|
||||
f0383d5ebc47: Pulling fs layer
|
||||
3214acf345c0: Waiting
|
||||
52630fc75a18: Waiting
|
||||
dd64bf2dd177: Waiting
|
||||
b839dfae01f6: Waiting
|
||||
ebddc55facdc: Waiting
|
||||
c4bc6f35ff5e: Waiting
|
||||
b96fe2995f90: Waiting
|
||||
58c0c263dc73: Waiting
|
||||
bd8962e29291: Waiting
|
||||
cac2ae0193cb: Waiting
|
||||
f0383d5ebc47: Waiting
|
||||
f6069939f718: Waiting
|
||||
d6b1b89eccac: Waiting
|
||||
2780920e5dbf: Waiting
|
||||
7c12895b777b: Waiting
|
||||
2cc7ee286bf3: Download complete
|
||||
218cf840d0d9: Verifying Checksum
|
||||
218cf840d0d9: Download complete
|
||||
c172f21841df: Verifying Checksum
|
||||
c172f21841df: Download complete
|
||||
f6069939f718: Verifying Checksum
|
||||
f6069939f718: Download complete
|
||||
2cc7ee286bf3: Pull complete
|
||||
d6b1b89eccac: Download complete
|
||||
2780920e5dbf: Verifying Checksum
|
||||
2780920e5dbf: Download complete
|
||||
7c12895b777b: Verifying Checksum
|
||||
7c12895b777b: Download complete
|
||||
3214acf345c0: Verifying Checksum
|
||||
3214acf345c0: Download complete
|
||||
52630fc75a18: Verifying Checksum
|
||||
52630fc75a18: Download complete
|
||||
dd64bf2dd177: Download complete
|
||||
c172f21841df: Pull complete
|
||||
b839dfae01f6: Verifying Checksum
|
||||
b839dfae01f6: Download complete
|
||||
ebddc55facdc: Verifying Checksum
|
||||
ebddc55facdc: Download complete
|
||||
c4bc6f35ff5e: Verifying Checksum
|
||||
c4bc6f35ff5e: Download complete
|
||||
58c0c263dc73: Verifying Checksum
|
||||
58c0c263dc73: Download complete
|
||||
bd8962e29291: Verifying Checksum
|
||||
bd8962e29291: Download complete
|
||||
b96fe2995f90: Verifying Checksum
|
||||
b96fe2995f90: Download complete
|
||||
cac2ae0193cb: Verifying Checksum
|
||||
cac2ae0193cb: Download complete
|
||||
218cf840d0d9: Pull complete
|
||||
f0383d5ebc47: Verifying Checksum
|
||||
f0383d5ebc47: Download complete
|
||||
f6069939f718: Pull complete
|
||||
d6b1b89eccac: Pull complete
|
||||
2780920e5dbf: Pull complete
|
||||
7c12895b777b: Pull complete
|
||||
3214acf345c0: Pull complete
|
||||
52630fc75a18: Pull complete
|
||||
dd64bf2dd177: Pull complete
|
||||
b839dfae01f6: Pull complete
|
||||
ebddc55facdc: Pull complete
|
||||
c4bc6f35ff5e: Pull complete
|
||||
b96fe2995f90: Pull complete
|
||||
58c0c263dc73: Pull complete
|
||||
bd8962e29291: Pull complete
|
||||
cac2ae0193cb: Pull complete
|
||||
f0383d5ebc47: Pull complete
|
||||
Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c
|
||||
Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3
|
||||
docker.io/cloudflare/cloudflared:2026.9.3
|
||||
1.5.6-stable: Pulling from gtstef/filebrowser
|
||||
55afa1ecc21d: Pulling fs layer
|
||||
8ed8f35f8d4f: Pulling fs layer
|
||||
989b226a579c: Pulling fs layer
|
||||
660aeead31d5: Pulling fs layer
|
||||
4f4fb700ef54: Pulling fs layer
|
||||
adce24567e4c: Pulling fs layer
|
||||
f17ea56b313b: Pulling fs layer
|
||||
6b6f3b3efe88: Pulling fs layer
|
||||
4ed1ca4f3fce: Pulling fs layer
|
||||
e6fc9c6a5757: Pulling fs layer
|
||||
d47782d1182a: Pulling fs layer
|
||||
660aeead31d5: Waiting
|
||||
6b6f3b3efe88: Waiting
|
||||
4ed1ca4f3fce: Waiting
|
||||
e6fc9c6a5757: Waiting
|
||||
d47782d1182a: Waiting
|
||||
4f4fb700ef54: Waiting
|
||||
adce24567e4c: Waiting
|
||||
f17ea56b313b: Waiting
|
||||
55afa1ecc21d: Verifying Checksum
|
||||
55afa1ecc21d: Download complete
|
||||
660aeead31d5: Verifying Checksum
|
||||
660aeead31d5: Download complete
|
||||
8ed8f35f8d4f: Verifying Checksum
|
||||
8ed8f35f8d4f: Download complete
|
||||
4f4fb700ef54: Verifying Checksum
|
||||
4f4fb700ef54: Download complete
|
||||
55afa1ecc21d: Pull complete
|
||||
989b226a579c: Verifying Checksum
|
||||
989b226a579c: Download complete
|
||||
f17ea56b313b: Verifying Checksum
|
||||
f17ea56b313b: Download complete
|
||||
6b6f3b3efe88: Verifying Checksum
|
||||
6b6f3b3efe88: Download complete
|
||||
4ed1ca4f3fce: Verifying Checksum
|
||||
4ed1ca4f3fce: Download complete
|
||||
adce24567e4c: Verifying Checksum
|
||||
adce24567e4c: Download complete
|
||||
d47782d1182a: Verifying Checksum
|
||||
d47782d1182a: Download complete
|
||||
8ed8f35f8d4f: Pull complete
|
||||
e6fc9c6a5757: Verifying Checksum
|
||||
e6fc9c6a5757: Download complete
|
||||
989b226a579c: Pull complete
|
||||
660aeead31d5: Pull complete
|
||||
4f4fb700ef54: Pull complete
|
||||
adce24567e4c: Pull complete
|
||||
f17ea56b313b: Pull complete
|
||||
6b6f3b3efe88: Pull complete
|
||||
4ed1ca4f3fce: Pull complete
|
||||
e6fc9c6a5757: Pull complete
|
||||
d47782d1182a: Pull complete
|
||||
Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8
|
||||
Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable
|
||||
docker.io/gtstef/filebrowser:1.5.6-stable
|
||||
1.1.0: Pulling from admin/felhom-samba
|
||||
897d797d2723: Pulling fs layer
|
||||
3051591aa250: Pulling fs layer
|
||||
ce57a3f93416: Pulling fs layer
|
||||
fb94eeec2fe1: Pulling fs layer
|
||||
fb94eeec2fe1: Waiting
|
||||
ce57a3f93416: Verifying Checksum
|
||||
ce57a3f93416: Download complete
|
||||
fb94eeec2fe1: Verifying Checksum
|
||||
fb94eeec2fe1: Download complete
|
||||
897d797d2723: Verifying Checksum
|
||||
897d797d2723: Download complete
|
||||
897d797d2723: Pull complete
|
||||
3051591aa250: Verifying Checksum
|
||||
3051591aa250: Download complete
|
||||
3051591aa250: Pull complete
|
||||
ce57a3f93416: Pull complete
|
||||
fb94eeec2fe1: Pull complete
|
||||
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
|
||||
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
|
||||
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
|
||||
[golden] identity-clean + minimize …
|
||||
[golden] stop + archive …
|
||||
INFO: including mount point rootfs ('/') in backup
|
||||
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||
INFO: archive file size: 618MB
|
||||
INFO: Finished Backup of VM 9100 (00:00:29)
|
||||
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_05-20_24_09.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
|
||||
[golden] publishing golden (648208028 bytes, sha256 8cebc42e15b091f0…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.297.0/golden.tar.zst
|
||||
[golden] pre-delete existing: HTTP 404 (404/204 expected)
|
||||
[golden] upload OK (HTTP 201)
|
||||
GOLDEN_VERSION=0.297.0
|
||||
GOLDEN_SHA256=8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
|
||||
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.297.0 / 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad
|
||||
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
|
||||
Reference in New Issue
Block a user