diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index bc9cd0e5..45e0c094 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -1,5 +1,14 @@ # 08 — The app-down alarm ladder +> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as +> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the +> 2026-08-22 pass): +> +> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet. +> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation. +> +> **An unmarked statement means "not yet classified", never "observed"** (R-376). + **Written 2026-08-23, with controller v0.222.0 (R-384).** **The absence is the finding.** Until this file existed, no document owned the question *"when does a diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index b62b5e19..8e5dc7ee 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -1,5 +1,14 @@ # 09 — How an app update works, and what it is becoming +> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as +> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the +> 2026-08-22 pass): +> +> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet. +> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation. +> +> **An unmarked statement means "not yet classified", never "observed"** (R-376). + > **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.** > Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen > deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who @@ -620,6 +629,11 @@ R-636's louder repeated alarm. controller images are deleted by the same in-use rule as decision 53 — *operator ruling 2026-10-01 (R-745, option 3A).* **Why:** about 50 old controller versions sat on each demo box (R-745); a release is ~400 MB unpacked and several ship a day. Registry tags are never deleted by this. + *Clarified 2026-10-05 (R-817), the ruling unchanged:* a swap records the image RUNNING when it starts + (`felhom-agent internal/localapi/controllerswap.go:236-240`, `st.Previous`) and a failed swap writes exactly that + image back (`:289`). After a good swap that image is "the one before" the running one — so „the one before it" here + and R-745's „rolls back to the RUNNING image" name the same image, seen before and after the swap. The controller + never hands the agent an older roll-back target. ### 2026-10-01 — decided by CC unattended, operator may reverse diff --git a/documentation/architecture/11-os-updates.md b/documentation/architecture/11-os-updates.md index 3802b72f..2180ce56 100644 --- a/documentation/architecture/11-os-updates.md +++ b/documentation/architecture/11-os-updates.md @@ -1,5 +1,14 @@ # 11 — Operating-system updates: the host, the guest and the Docker engine +> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as +> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the +> 2026-08-22 pass): +> +> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet. +> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation. +> +> **An unmarked statement means "not yet classified", never "observed"** (R-376). + > | | | > |---|---| > | **Status** | **NOT RATIFIED — a PROPOSAL with operator rulings, corrected by the 2026-10-04 spike (§7.1, C1–C12); §8 step 2 BUILT 2026-10-04 (§8.1).** Ratification is Viktor's review, not an editor's. | diff --git a/documentation/audits/burndown-2026-10-05/partA-results.jsonl b/documentation/audits/burndown-2026-10-05/partA-results.jsonl new file mode 100644 index 00000000..4ba3b7b2 --- /dev/null +++ b/documentation/audits/burndown-2026-10-05/partA-results.jsonl @@ -0,0 +1,317 @@ +{"id": "R-10", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/appbackup/dbdump.go:364 `if err := tmpFile.Sync(); err != nil {` then :390 `if err := os.Rename(tmpPath, finalPath); err != nil {` with no directory Sync after; the twin at controller/internal/backup/backup.go:948 `_ = dir.Sync()` does sync the dir. Origin: audits/CAMPAIGN-6E-2026-07-15.md:129.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/appbackup/dbdump.go"], "change": "After the os.Rename in DumpOne, open filepath.Dir(finalPath) and call a best-effort dir.Sync() (log at DEBUG on error), mirroring atomicPromoteTar in backup.go:948.", "test": "Unit test that DumpOne still produces the final file and leaves no .tmp; fsync itself is not observable in a unit test, so add a small syncDir seam and assert it is called with the dump directory (red-proof by removing the call).", "minutes": 30}, "not_worth": null, "minutes_spent": 4} +{"id": "R-25", "sev": "P4", "category": "Storage & devices", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/web/storage_handlers.go:153 `uuid := resolveEnrollUUID(ctx, agent, device)` still resolves by device PATH after format, then AssignDisk(uuid) at the next step; FormatResult (controller/internal/agentapi/client.go:384-395) carries DurableID only for the confirmation path, not the new fs UUID. Binding resolve+assign to the format's durable-id needs the agent to return the new fs identity (two repos).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6} +{"id": "R-76", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKED", "evidence": "Behaviour is FileBrowser-image runtime behaviour (mode/setgid of UI-created folders), only observable on a live box. The image has changed since the finding: controller/internal/infra/infra.go:27 `FileBrowserImage = \"gtstef/filebrowser:1.5.6-stable\"` (finding was on 1.3.3). The comment at infra.go:207-208 still asserts `umask 002 so folders the customer creates here come out group-writable (2775 with the parent's setgid)`, which the row says was false on 1.3.3 — needs a re-measure on 1.5.6 before deciding. Note: the register row itself is truncated mid-sentence (\"does not say t\").", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5} +{"id": "R-89", "sev": "P4", "category": "Business & legal", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No retention policy object in hub: `grep -rln -i 'retentionpolicy|retention_policy' felhom.eu/hub` returns nothing (felhom.eu@53d8131b). Commercial per-customer policy = money/product decision + new reconciler.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2} +{"id": "R-91", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKED", "evidence": "Whether /srv/pbs-felhom still exists on ep0 is live-only (ep0 is protected; not touched). Source-side: CONTEXT.md:3656 still reads \"`/srv/pbs-felhom` is 13 G of dead weight on `/` awaiting R-91's go-ahead\". Extra fact found: documentation/runbooks/offsite-endpoint.md:24 still says the datastore `felhom-offsite` is at `/srv/pbs-felhom` and :119 `proxmox-backup-manager datastore create felhom-offsite /srv/pbs-felhom`, contradicting RUNBOOK-ep0-datastore-volume-2026-07-27.md:8 (moved to /mnt/pbs-datastore). The row's CONTEXT.md:1018 citation is stale (now :3656).", "dup_of": null, "unique_facts": "offsite-endpoint.md:24 and :119 still name /srv/pbs-felhom as the live datastore path (stale since the 2026-07-27 move to /mnt/pbs-datastore) — a doc fix independent of the deletion; CONTEXT citation moved from :1018 to :3656.", "small_fix": null, "not_worth": null, "minutes_spent": 5} +{"id": "R-92", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/web/pbsdr_box.go:57 and :64 `view.UsedStr = fmtBytesGB(snap.UsedBytes)`; hub/internal/web/offsite_box.go:54 `return fmt.Sprintf(\"%.1f GB\", float64(b)/float64(int64(1)<<30))` — still 0.1 GB granular. Note the row's own trigger (\"when retention becomes customer-visible\") has not fired.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/web/pbsdr_box.go", "hub/internal/web/offsite_box.go", "hub/internal/web/templates/offsite.html"], "change": "Add an exact-bytes value to the PBS DR view (e.g. UsedBytesExact rendered as a title= tooltip or a MB-precision string below 10 GB) without changing fmtBytesGB for other callers.", "test": "Table test on the view builder: two snapshots 50 MB apart render different exact strings; render test of offsite.html shows the exact value.", "minutes": 30}, "not_worth": null, "minutes_spent": 4} +{"id": "R-93", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Premise gone per the row itself (R-461, CLOSED-ITEMS.md:636): drill-r50 VM no longer exists on either demo box. target-selection.md:111 keeps the fence with the note \"the VM does not exist anywhere, so the fence currently protects nothing\". Only remaining references are comments/tests (hub/internal/monitor/deadline_anchor_test.go:16, deadline_tiers.go:59).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A row about choosing between two fixtures, neither of which exists any more.", "cost": "Building a new synthetic drift fixture is a design task (M), not a fix of this row.", "if_never": "Nothing breaks; there is no drift fixture either way. If one is wanted, it is a new row.", "pick": "close-as-accepted (operator word needed: close, or reopen as 'build a drift fixture'); also drop the dead drill-r50 fence in target-selection.md:111 at close"}, "minutes_spent": 4} +{"id": "R-99", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No phantom-snapshot cleanup in felhom.eu/hub or felhom-agent (grep -i phantom finds only agent runner/test detection code; no removal path). Deletion on a customer datastore is a separate operator ruling per the row — customer data.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3} +{"id": "R-104", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/backup/offbox.go:193-222 ClassifyOffsiteFailure has cases NoUnits/NoRepo/Transport and `default: return OffsiteFailUnknown` — no lock case, although offbox.go:826 already defines `var offboxLockRe = regexp.MustCompile(`repository is already locked`)`. The self-heal half is built (offbox.go:839-858 unlock --remove-all + retry once), as the row's 2026-08-22 note says.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/backup/offbox.go", "controller/internal/i18n/locales/hu.json", "controller/internal/i18n/locales/en.json", "controller/internal/backup/*_test.go"], "change": "Add an OffsiteFailLocked class matched by offboxLockRe in ClassifyOffsiteFailure (before transport) and a cause line in OffsiteFailureMessage telling the operator the repository is locked by an interrupted run and how it clears.", "test": "Table test: a restic 'repository is already locked' error classifies as Locked (red-proof: fails today as Unknown); i18n parity gate for the new key.", "minutes": 50}, "not_worth": null, "minutes_spent": 5} +{"id": "R-124", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "Still true: felhom-agent@e06ed97 internal/hub/dr_recipe.go:61 `const PBSRootNamespace = \"root\"` and :276 `c.Namespace = PBSRootNamespace`; agent internal/pbs/report.go:24 `ns = \"root\"`. The comment at dr_recipe.go:57-58 already documents that PBS spells it \"\".", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The disaster-recovery recipe writes the PBS root namespace as the word 'root', but PBS itself uses an empty name, so a pasted '--ns root' fails.", "cost": "Changing it alters a wire field read by the hub (cross-repo wire contract + recipe producers), for a case no customer has: every box writes a per-customer namespace.", "if_never": "An operator restoring a box with NO namespace line would get one failed command and have to drop --ns; no data risk. The constant's comment already warns.", "pick": "close-as-accepted"}, "minutes_spent": 4} +{"id": "R-129", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "Docs still say no key: felhom.eu@53d8131b documentation/operations/nodes.md:110 `### Access — there is no baked SSH key` and :112 \"no operator public key is on this box\"; target-selection.md:111 still flags R-129 unresolved; MEMORY.md:34 says `ssh demo-hp`, NO KEY→G1. Whether the key works today is a live fact (not checked — no ssh in this pass).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["documentation/operations/nodes.md", "documentation/runbooks/target-selection.md"], "change": "After one read-only `ssh -o BatchMode=yes demo-hp true` (and reading root's authorized_keys comment to name the key), rewrite nodes.md 'Access' section to the measured truth and drop the R-129 caveat in target-selection.md:111-112 (also update the memory index line).", "test": "Positive control: the BatchMode ssh succeeds/fails as the doc now states; repo_gates.py doc gates pass.", "minutes": 30}, "not_worth": null, "minutes_spent": 4} +{"id": "R-134", "sev": "P4", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/cloudflare/unblock.go:117 `for _, name := range []string{domain, parentDomain(domain)} {` and :136-141 parentDomain strips exactly one label (`strings.SplitN(domain, \".\", 2)`); controller strips progressively (controller/internal/cloudflare/zone.go per row).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/cloudflare/unblock.go", "hub/internal/cloudflare/unblock_test.go (new)"], "change": "Extract a pure zoneCandidates(domain) []string that yields the name and every parent down to two labels, and loop resolveZone over it (same order: most specific first).", "test": "Table test on zoneCandidates: 'a.b.felhom.eu' yields [a.b.felhom.eu b.felhom.eu felhom.eu] (red-proof: one-label version yields only two); no HTTP needed because apiBase is a const.", "minutes": 40}, "not_worth": null, "minutes_spent": 4} +{"id": "R-161", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Automatic half exists: app-catalog-felhom.eu@917a779 .gitea/workflows/gates.yml:40 `run: cd ws/app-catalog-felhom.eu && python3 scripts/catalog_gates.py --fast`; .githooks/pre-push:84 runs the same. Only the runtime volume-persistence gate stays a manual periodic run, by operator ruling 2026-08-02.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The runtime check that app data lands on a volume is run by hand, not on every push.", "cost": "Automating it means CI pulling and starting ~53+ app images per push — slow, and the row itself says such CI gets disabled.", "if_never": "A template that writes data outside its volume can ship until the next periodic run catches it; the static gates and pre-push still run.", "pick": "close-as-accepted (residual is a deliberate ruling; owner operator)"}, "minutes_spent": 3} +{"id": "R-162", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Gate is app-catalog-felhom.eu scripts/check-volume-persistence.py (`docker diff` at :41, :72); behaviour on a non-overlay driver is not reachable from source and the row says it fails closed. Status WATCHING, no defect.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "If Docker ever ran on a storage driver where `docker diff` does not work, the persistence gate would refuse to report and blame the prober instead of the driver.", "cost": "A driver probe + reworded message in the catalog script, ~1 h, for a driver nobody runs.", "if_never": "Nothing, unless a non-overlay driver ships; even then the gate fails closed (no false green).", "pick": "close-as-accepted"}, "minutes_spent": 3} +{"id": "R-164", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Predicate still absent: felhom-controller controller/internal/appbackup/dbdump.go:544 still only WARNs `its accounts table has NO rows`; restore still replays dump + tar (internal/backup/restore_unit.go:114-118 hasReplayableDump). Blocked on a design (live-vs-dump per-table counts).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3} +{"id": "R-169", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Working-style ruling owed by operator; nothing in source to fix. Current nets per row: pre-push hooks + Gitea runner alarm (e.g. app-catalog .gitea/workflows/gates.yml:40).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "CI only reports after a push lands, because every repo pushes straight to main with no pull request.", "cost": "Making CI blocking needs branch protection plus a PR workflow for every change — a slower way of working for a one-operator project.", "if_never": "A `--no-verify` push can land broken code until the operator reads the CI alarm e-mail.", "pick": "close-as-accepted (row itself says decide only if the window ever costs something)"}, "minutes_spent": 2} +{"id": "R-177", "sev": "P4", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller@7690c27 controller/cmd/controller/main.go:1546 `sched.Daily(\"fill-watch\", \"03:30\", func(ctx context.Context) error { return fillWatcher.Check() })`; internal/scheduler/scheduler.go:269 has GetJobs but grep finds no RunNow/Trigger method and no run-job route in internal/web. Needs a new operator-gated trigger endpoint (auth surface) — a new mechanism, solve together with R-279.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4} +{"id": "R-184", "sev": "P4", "category": "Box system & updates", "group": "FIXED-BY-LATER-WORK", "evidence": "Fixed by felhom.eu b55fc17d \"hub v0.102.0 — refuse to vouch a version that cannot be installed (R-273)\" — exactly shape (b), validate at vouch time in the hub. felhom.eu/hub/internal/web/configs.go:1358 `res := s.gitea.PackageDownloadable(ctx, t.pkg, t.version, t.file)` and :1365 `s.logger.Printf(\"[WARN] artifact vouch REFUSED: %s package %s is NOT downloadable (R-287)\", ...)`; tag leg at :1343 TagServesFile; unreachable registry also refuses.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4} +{"id": "R-194", "sev": "P4", "category": "Box system & updates", "group": "NOT-WORTH-IT", "evidence": "PVE behaviour, not our code; row states the self-repair already tolerates it (fires on the next probe after the cache clears). No source change to check.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Proxmox caches permissions, so a removed storage grant can still read as present for seconds to minutes; our self-repair notices only after the cache expires.", "cost": "Adding a second signal (storage content listing) to the agent's grant probe is a new mechanism, needs live measurement on a box.", "if_never": "A lost grant is noticed up to ~16 min late; the repair still happens on its own.", "pick": "close-as-accepted"}, "minutes_spent": 2} +{"id": "R-206", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "homelab-manifests@87dfc29 (/home/kisfenyo/git/homelab-manifests): no daemon.json template in homelab-ansible (grep finds only a comment at roles/node_housekeeping/templates/node-housekeeping.sh.j2:17 and homelab-ansible/CLAUDE.md:54). Part (b) was superseded by fc9fbb8 (\"correct the expired Docker rationale\"): the script now says at :13-20 do NOT add docker calls, the GC policy in daemon.json is the control point. DooPlex work — not unprompted.", "dup_of": null, "unique_facts": "Part (b) (prune in the role) is superseded by fc9fbb8 — the role now deliberately forbids docker calls and names daemon.json GC policy as the control; remaining scope is (a) template daemon.json in Ansible + (c) restart-and-verify.", "small_fix": null, "not_worth": null, "minutes_spent": 4} +{"id": "R-207", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "Fixed by homelab-manifests fc9fbb8 \"node_housekeeping: guard DRY_RUN, correct the expired Docker rationale, pin container log rotation\". /home/kisfenyo/git/homelab-manifests/homelab-ansible/roles/node_housekeeping/templates/node-housekeeping.sh.j2:137 `if [[ \"${DRY_RUN}\" == \"1\" ]]; then` inside write_metrics, :138 logs \"file left untouched\".", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3} +{"id": "R-208", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/Dockerfile:12 `ARG VERSION=dev` and :13 `ARG GIT_COMMIT=unknown` sit above :19 `RUN go mod download || true`; felhom.eu@53d8131b hub/Dockerfile:3 `ARG VERSION=dev`, :4 `ARG BUILD_TIME=unknown` above :9 `RUN go mod download || true`.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller (and the identical one-line move in felhom.eu/hub — two repos, each trivial)", "files": ["controller/Dockerfile", "felhom.eu: hub/Dockerfile"], "change": "Move the ARG VERSION/GIT_COMMIT (controller) and ARG VERSION/BUILD_TIME (hub) declarations down to just above the final `go build` RUN.", "test": "Build twice with different --build-arg VERSION on a clean tree; the second build must show `RUN go mod download` CACHED; `--version` of the built binary still shows the passed version.", "minutes": 30}, "not_worth": null, "minutes_spent": 3} +{"id": "R-209a", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKED", "evidence": "Live-only: whether DooPlex has rebooted and /var/log/felhom-store-postboot-check.log says PASS. Not read (DooPlex is Tier 2, operator ruled no reboot; this pass touches no machine). No source claim to check.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2} +{"id": "R-210", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Operator ruling owed; row records CC's view 'not worth doing for the space' (~27 GB reclaimable vs 199 GB free). Workspace CLAUDE.md also forbids `docker image prune -a` on DooPlex.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "193 old controller/hub images exist only on DooPlex and cannot be re-pulled; the question is whether to delete them.", "cost": "An operator decision plus a careful targeted delete on the production host; returns ~27 GB.", "if_never": "~27 GB stays used on a disk with ~199 GB free; old images remain as clutter (and as the only copies of very old builds).", "pick": "close-as-accepted"}, "minutes_spent": 2} +{"id": "R-213", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Row is a not-started design (live-vs-backup comparison, then put-back flow), operator-owned; nothing in source to verify against.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1} +{"id": "R-230", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Owed rulings, not code: (a) bulk-correction ruling on MEMORY.md staleness (MEMORY.md index still carries version literals, e.g. 'ctrl 0.224.0', 'hub 0.109.0'); (c) spec-as-failing-test pilot not started. (b) closed. Operator decision required.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2} +{"id": "R-246", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/store/store.go:3248 `func (s *Store) MarkEscrowStale(hostID string) error {` still has no production caller (grep: only definition + comments at offsite.go:208,216); stale_at still read (store.go:3182 clears it). Ruling owed by operator: evidential setter or retire the column (folds R-248).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3} +{"id": "R-256", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/i18n/locales/hu.json:1406 `\"flash.offbox.mgr_unavailable\": \"A mentéskezelő nem elérhető.\",` used at controller/internal/web/offbox_handlers.go:54 and :197; sibling :1407 mgr_unreachable used at offbox_handlers.go:582; en.json:1415-1416 same shape.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/i18n/locales/hu.json", "controller/internal/i18n/locales/en.json"], "change": "Rewrite flash.offbox.mgr_unavailable / mgr_unreachable in both languages to say the backup service is not running yet and give a route (try again in a few minutes; if it persists, contact support).", "test": "i18n parity/accent gates (controller_gates.py) pass; a handler test with nil backupMgr asserts the redirect carries the key (exists or add one). Owner is operator (copy) — needs a nod on the wording.", "minutes": 20}, "not_worth": null, "minutes_spent": 4} +{"id": "R-261", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/store/selfbind.go:111 `func (s *Store) CountSelfBindTokens(customerID string) (int, error) {`; only callers are hub/internal/web/customer_delete_test.go:511 and selfbind_automint_test.go:29 — no production caller.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/store/selfbind.go"], "change": "Reword the doc comment (selfbind.go:106-110) to say it is a test accessor and name the two tests that pin the auto-mint invariant (selfbind_automint_test.go, customer_delete_test.go) — or, if the operator prefers, add one post-mint production check that logs [WARN] when count != 1.", "test": "Comment-only option: existing tests stay green; production-check option: unit test that a pre-seeded extra token produces the WARN (red-proof).", "minutes": 20}, "not_worth": null, "minutes_spent": 3} +{"id": "R-263", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/settings/settings.go:1655 `// from every other. This is the ONLY writer of StoragePath.BackupTarget — registration must never set` while :1699 `s.StoragePaths[i].BackupTarget = false` (ClearBackupTarget) also writes it; :1684 is SetBackupTarget's write.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/settings/settings.go", "controller/internal/settings/backup_target_role_test.go"], "change": "Change the comment to 'the only writer that GRANTS the role' and add a source-scanning test that finds every `.BackupTarget =` assignment in non-test settings code and fails if any other than SetBackupTarget can assign a non-false value.", "test": "The new test passes today; red-proof by planting a temporary `BackupTarget = true` in another function and seeing it fail.", "minutes": 45}, "not_worth": null, "minutes_spent": 3} +{"id": "R-264", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu@53d8131b scripts/wire_contract_gate.py still allowlists the six with _R264: :242 selfupdate_pending, :246 selfupdate_pending_version, :255 restore_tests.mount_parity, :258 restore_tests.mount_inventory, :281 backup.last_db_dump, :282 backup.last_integrity_check. Each reader is a design per the row.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3} +{"id": "R-266", "sev": "P4", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/report/builder.go:94 `{Mount: \"/\", Label: \"SSD\", TotalGB: sysInfo.DiskTotalGB, UsedGB: sysInfo.DiskUsedGB, Percent: sysInfo.DiskPercent},` — no disk_known on the storage entry; hub has no disk_known (grep empty). Two-repo wire change gated by wire_contract_gate.py.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3} +{"id": "R-279", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No operator/hub path to start an off-site run: grep for offbox run triggers in felhom.eu/hub/internal finds nothing; the only run entry is the customer dashboard handler (felhom-controller controller/internal/web/offbox_handlers.go:270 `if !s.backupMgr.OffboxRunnable() {`). Needs a new operator-authenticated trigger — sibling of R-177, not a duplicate (different job).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3} +{"id": "R-284", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "Not a defect: felhom-controller@7690c27 controller/internal/web/templates/deploy.html:622 `