burn-down round 2: controller v0.297.0 rows closed (23), golden 0.297.0 evidence, delivery evidence, 23-row unchecked table, STATUS/CONTEXT/REPORT (292 -> 199; 1 opened, 94 closed)
gates / gates (push) Successful in 1m59s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 20:41:50 +02:00
parent d75ad0fdf3
commit 30650cad6e
13 changed files with 640 additions and 37 deletions
@@ -497,3 +497,9 @@ rc=0
--- PASS: TestR365_BannerSaysDueAtZeroDays (0.12s)
PASS
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.132s
### CI fix (lead, 2026-10-05 ~18:20Z): the new gofmt gate was INCONCLUSIVE on the CI runner (no Go) — CI run 1371 FAILED
Fix: in CI (GITEA_ACTIONS/GITHUB_ACTIONS=true) with no gofmt reachable the gate prints "NOT CHECKED in CI" and exits 0;
elsewhere a missing gofmt stays INCONCLUSIVE. Decoys gofmt/ci-without-go and gofmt/dev-without-go (empty PATH).
Simulated CI (env -i, empty PATH, GITEA_ACTIONS=true): "gofmt gate NOT CHECKED in CI …" rc=0.
Red-proof (CI branch disabled): FAIL: gofmt/ci-without-go: want rc=0 and 'NOT CHECKED in CI', got rc=2.
@@ -0,0 +1,4 @@
== controller delivery 2026-10-05T18:28:59Z
demo-hp gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 32 seconds (healthy)
demo-felhom gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 35 seconds (healthy)
tester-1 (hub customer page, version strings seen): 6 0.297.0 4 0.296.0 4 0.295.0
@@ -0,0 +1,11 @@
== hub deploy 2026-10-05T18:11:56Z
image=gitea.dooplex.hu/admin/felhom-hub:0.137.0
sync=Synced health=Healthy op=Succeeded rev=d75ad0fdf3daf3f3b2a1690746d9a6a70ee4104c
2026/10/05 20:11:05 [INFO] felhom-hub 0.137.0 starting
healthz 200
2026/10/05 20:11:06 [INFO] osupdates: the Docker engine set is approved only by the operator, after 2 healthy ring-0 night(s)
== System page 2026-10-05T18:12:05Z: root-files / agent cells
73: 0 armed unknown 0.142.0 → 0.147.0 (since 2026-10-05)
89: 0 armed 0.147.0 0.147.0
105: 2 armed 0.147.0 0.147.0
121: 2 TRIPPED 2026-10-05T07:57:17Z 0.147.0 0.147.0
@@ -0,0 +1,11 @@
== vouch 2026-10-05T18:27:46Z: agent 0.147.0, golden 0.297.0, min_agent 0.131.0
HTTP/1.1 303 See Other
Location: /configuration?flash=artifacts_set
== floors 2026-10-05T18:28:16Z: POST /customers/<id>/floor min_controller_version=0.297.0 min_agent=0.131.0
demo-hp: Location: /customers/demo-hp?flash=floor_set
demo-felhom: Location: /customers/demo-felhom?flash=floor_set
tester-1: Location: /customers/tester-1?flash=floor_set
2026/10/05 20:28:16 [INFO] Artifact manifest set: agent=0.147.0 golden=0.297.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d"
2026/10/05 20:28:16 [INFO] Customer demo-hp controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
2026/10/05 20:28:17 [INFO] Customer demo-felhom controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
2026/10/05 20:28:17 [INFO] Customer tester-1 controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
@@ -0,0 +1,23 @@
{"id": "R-76", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 `FileBrowserImage = \"gtstef/filebrowser:1.5.6-stable\"`. The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --grep 'R-76|setgid' in controller: only 2026-06 commits). Whether 1.5.6 still drops setgid is runtime behaviour of the image.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- find /mnt/felhom-drives -path '*/userdata/*' -mindepth 3 -maxdepth 5 -type d ! -perm -2000 -printf '%m %u:%g %p\\n'\" | head (any folder a customer made in FileBrowser showing 755 without setgid = still true on 1.5.6)", "minutes_spent": 3}
{"id": "R-91", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word.", "dup_of": null, "unique_facts": "Stale doc to fix in the same commit as the delete: documentation/runbooks/offsite-endpoint.md:24 still says datastore `felhom-offsite` is at `/srv/pbs-felhom`; the real path since 2026-07-27 is `/mnt/pbs-datastore` (RUNBOOK-ep0-datastore-volume-2026-07-27.md:8). CONTEXT.md:3666 is the other line to change.", "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'du -sh /srv/pbs-felhom 2>&1; proxmox-backup-manager datastore list; df -h /'", "minutes_spent": 4}
{"id": "R-209a", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h /", "minutes_spent": 1}
{"id": "R-337", "sev": "P4", "category": "Monitoring & notifications", "group": "NOT-WORTH-IT", "evidence": "The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 `b, err := tier.Service.BackupWithSnapshotHook(...)` then :901 `s.store.RecordBackup(b)` (grep RecordBackup: no other caller). So it is NOT a collection cadence; the status appears when the agent's own job finishes (WaitTask + archive resolve, internal/backup/runner.go:205-232), and a backup the agent did not run itself never appears. The 4-min demo-hp skew is the gap between PBS writing the manifest and the job returning -- unmeasured.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: measure why demo-hp's job returned ~4 min after the manifest landed; cost: a constructed live repro on demo-hp with task-log timing; if never: during an incident the box's backup status can trail the PBS manifest by minutes after an out-of-schedule run, and a run started outside the agent never shows; pick: close with the source fact above written into the row (status = agent job end, not polling), reopen only if a lag is seen on a scheduled run.", "settle_cmd": null, "minutes_spent": 7}
{"id": "R-375", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is `pvesm status` showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 `if st == nil || st.Type == \"pbs\" || st.Avail <= 0 { return true, \"\" }` (space preflight skips PBS and fails open on 0), and the same report says the hub's gauge reads the real 3.7/97.9 GB.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: confirm on ep0 that the namespace-scoped token lacks Datastore.Audit; cost: a read-only ep0 session; if never: the PVE UI on a box shows 0/0/0 for felhom-pbs, which no Felhom code reads (runner.go:265); pick: close as cosmetic with this pointer.", "settle_cmd": "(if ever wanted) ssh root@ep0 'proxmox-backup-manager acl list' ; on a box: pvesm status --storage felhom-pbs", "minutes_spent": 6}
{"id": "R-488", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded `interval := 5 * time.Second` and `time.Sleep(3 * time.Second) // initial settling time`, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:90, restore_unit.go:471. No commit since 2026-09-13 touching internal/backup mentions R-488/test speed. Runtime (333 s) was NOT re-measured here (read-only).", "dup_of": null, "unique_facts": null, "small_fix": "controller: add Manager fields healthSettle/healthInterval (defaults 3s/5s, set in the constructor) used by waitForHealthy; a test helper (the existing Manager fixture constructor) sets them to 0/10ms. Test: TestWaitForHealthy_DefaultsAreProduction asserts a fresh Manager has 3s/5s (pins production), and measure `go test ./internal/backup` wall time before/after in the commit message (expect the 89 >=1 s tests to drop). Run with the R-650 docker-free seams.", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}
{"id": "R-504", "sev": "P4", "category": "Install & onboarding", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "If 404 is confirmed it is cosmetic (row's own re-rank); the operator could close it as accepted rather than add a Cloudflare rule.", "settle_cmd": "curl -sI https://iso.felhom.eu/ | head -1", "minutes_spent": 2}
{"id": "R-644", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs `--deployment-password`; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches the crash text. Not proven.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- docker ps -a --filter name=gokapi --format '{{.Names}} {{.Status}}'; pct exec 9202 -- grep -s -c '\\\"Password\\\":\\\"\\\"' /opt/docker/stacks/gokapi/config/config.json\"", "minutes_spent": 3}
{"id": "R-814", "sev": "P4", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Hetzner account state; nothing in source records a deletion.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status", "minutes_spent": 1}
{"id": "R-815", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "PBS server-side state on ep0; no GC completion record in the docs (grep).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 | grep -i garbage'", "minutes_spent": 2}
{"id": "R-884", "sev": "P4", "category": "Monitoring & notifications", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live.", "dup_of": null, "unique_facts": "Cause candidate: Renovate 53c6e99 prometheus v3.15.0 merged 2026-10-03, never synced because monitoring has no auto-sync in git.", "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n mon-system get deploy prometheus -o jsonpath='{.spec.template.spec.containers[0].image}' (v3.14.0 => the diff is the Renovate bump)", "minutes_spent": 5}
{"id": "R-132", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key `operator_password_hash`, setSetting :2200-2207 writes `updated_at = datetime('now')`. No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "On a copy of the hub DB (hub.db + -wal + -shm): sqlite3 hub.db \"SELECT updated_at FROM hub_settings WHERE key='operator_password_hash'\" -- rotated only if later than 2026-09-18 (the last exposure, R-580 folded here); no row = still the ConfigMap seed", "minutes_spent": 4}
{"id": "R-298", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 `if(d.role==='user-data'){` else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/role.go:172-186 -- a dir storage (felhom-backup) is user-data unless its backing device is on the system disk; disks.go:1236-1240 deliberately does not reclassify the backup-target drive. role.go unchanged since 2026-08-09. demo-hp was reprovisioned before 2026-09-21, so the 2026-08-10 topology may no longer hold.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp 'grep -A2 felhom-backup /etc/pve/storage.cfg; lsblk -no PKNAME $(findmnt -no SOURCE /) ; lsblk -no PKNAME $(findmnt -no SOURCE /mnt/nvme-1tb)' (same parent disk => role=system => the page still locks it => R-298 true; different => user-data => not reproducible on this box)", "minutes_spent": 8}
{"id": "R-338", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"grep -E 'listen_addr|island_' /etc/felhom-agent/agent.json; pct config 9201 | grep ^net; ip -br link show master vmbr9\"", "minutes_spent": 6}
{"id": "R-350", "sev": "P3", "category": "Security & access", "group": "DUPLICATE", "evidence": "Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence.", "dup_of": "R-132", "unique_facts": "(1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked.", "small_fix": null, "not_worth": null, "settle_cmd": null, "minutes_spent": 3}
{"id": "R-542", "sev": "P3", "category": "Storage & devices", "group": "NOT-WORTH-IT", "evidence": "Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 `initialize = append(initialize, c) // every unclaimed disk can be initialized`; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkfs wrapper explicitly allows it: configs/felhom-mkfs-guarded.sh:59-60 'Mounts under /mnt/felhom-drives are our own drives (the agent detaches before a re-init) -> allowed'. Controller passes initialize through untouched (agent_disk_handlers.go:158-162). `already_mounted` is only set for controller-contributed stores (agentapi/client.go:447-451, omitempty) -- so 'null' is expected for agent candidates.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: drop felhom-mounted disks from `initialize`; cost: reverses the R-220/re-init design (wrapper comment) and needs a decision on how a registered drive is re-initialized; if never: the raw endpoint lists a registered drive under initialize while the page (customer view) filters it -- only a session reading the raw endpoint is misled; pick: close as by-design, add one comment line at disks.go:438 saying a registered felhom drive is listed here deliberately.", "settle_cmd": null, "minutes_spent": 9}
{"id": "R-607", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-SMALL", "evidence": "Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY `if len(newApps) > 0 || len(updated) > 0`, and :257-258 says 'nincs változás' when both are empty. `updated` counts stack-dir copies only (copyTemplates, :447 `updated = append(updated, appName)` after a hash mismatch); for a deployed+pinned app whose catalog moved, renderSource (:457-469 table) copies the STORED definition, so the hash matches and nothing is 'updated' although the git cache moved. CatalogImages is read from the catalog cache only inside ScanStacks (stacks/manager.go:666-672, assigned :690). So (a) the message measures the stack dir, not the catalog; (b) CatalogImages refreshes only on a ScanStacks, which this sync does not trigger. The 29 s nextcloud case (needed several rounds) is not explained by this.", "dup_of": null, "unique_facts": null, "small_fix": "controller sync.go: record the catalog git HEAD before and after gitCloneOrPull; if it moved, call s.rescanFn() even when newApps/updated are empty, and say 'Katalógus frissítve — az alkalmazások nem változtak' (and EN) instead of 'nincs változás'. Test (red first): a Syncer with a fake pull that moves HEAD and a frozen pinned app (renderSource returns the stored definition) asserts rescanFn was called once and the message is not the no-change one; and a no-move pull asserts rescanFn NOT called.", "not_worth": null, "settle_cmd": null, "minutes_spent": 9}
{"id": "R-683", "sev": "P3", "category": "App updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power-cut-mid-update drill: night-2026-10-04/MORNING-NOTE.md:64 'A1 power cut mid-update (demo-hp): NOT RUN'; DRILL-night-2026-09-25.md:102 was a cut during romm's verifying, not checked for the backup choice.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Not a read-only command: the next power-cut-during-update drill with the controller log saved at arm time, then grep \"phase backing-up\\|Tier-2\" in the saved pre-cut log", "minutes_spent": 7}
{"id": "R-756", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 `if !m.DriveLive(hddPath)` -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 `return m.isMountPoint(hddPath)` -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names `.../scratch_hdd/userdata/calibre-web`, so on 9202 HDD_PATH is a per-app subfolder, which can never be a mount point -> the 409 is certain for that value. Open: whether that HDD_PATH was written by the product (handlers.go:550 prefill from place.Drive) or by the test venue.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- grep -H HDD_PATH /opt/docker/stacks/calibre-web/app.yaml /opt/docker/stacks/grimmory/app.yaml; pct exec 9202 -- findmnt -no TARGET,SOURCE /mnt/felhom-drives/scratch_hdd\" (HDD_PATH = per-app subfolder => product/venue wrote a non-root HDD_PATH; HDD_PATH = drive root and not mounted => scratch drive not registered)", "minutes_spent": 10}
{"id": "R-862", "sev": "P3", "category": "Box system & updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Ask the operator whether the three bootstrap commands ran; read-only proof: the hub's operator view of Tester 2 (OS/config-bundle sha vs the vouched sha), or whether os_config_bundle_behind has fired for it", "minutes_spent": 2}
{"id": "R-882", "sev": "P3", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p ActiveEnterTimestamp (instance-manager older than a k3s/containerd/iscsid restart => the stale-PID state can recur)", "minutes_spent": 2}
{"id": "R-883", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/recipe-importer:latest (pull Always :27); adventurelog-system/adventurelog.yaml:100 and :256 adventurelog-backend/frontend:latest (pull Always :101,:257); jarrs-system/jarr-dev.yaml:311,:345,:630 gitea.dooplex.hu/admin/jarr:latest (pull Always). Zipline fixed (90f60e4, 4c8ec7a). The repo's own rule homelab-manifests/CLAUDE.md:120 'Image tags always pinned'. Helm values files (external-dns, pihole, plex, authentik, cnpg) not checked for tag fields beyond a `tag: latest|dev|empty` grep (no hits).", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests only: for each line above, read the running digest/version (`sudo kubectl get pod -n <ns> -o jsonpath='{..imageID}'`), pin that exact tag (or @sha256 for the self-built gitea.dooplex.hu jarr/recipe-importer images, which have no version tags), add the version-checker match-regex annotation per CLAUDE.md:120, and set imagePullPolicy IfNotPresent. Test: `grep -rnE 'image:.*(:latest|:dev)\\s*$' --include=*.yaml .` returns nothing; after ArgoCD sync each pod's imageID equals the pre-change one (no upgrade). Operator-owned DooPlex change: needs the operator's word.", "not_worth": null, "settle_cmd": null, "minutes_spent": 6}
{"id": "R-886", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager real storage').", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests: add pod `securityContext: {fsGroup: 65534, fsGroupChangePolicy: OnRootMismatch}` (65534 = nobody, the image's user -- confirm with `sudo kubectl -n mon-system exec deploy/alertmanager -- id`) to the alertmanager Deployment. Test (consequence, per CLAUDE.md): create a silence via amtool/API, delete the pod, after it returns the silence is still listed AND the log has no 'Running maintenance failed ... permission denied' within 15 min (positive observable: a 'maintenance done' line).", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}