# CLOSED-ITEMS — finished work, compressed > **What this is.** Every register row that reached a terminal state, compressed to its title, the > version it shipped in, its evidence paths, and any sentence that states a RULE rather than a > narrative. **Nothing was deleted:** each entry names the commit that holds its full original text, > and `git show :documentation/backlog/OPEN-ITEMS.md` returns it verbatim. > > **Why it exists (operator ruling, 2026-08-22).** `OPEN-ITEMS.md` had grown to 672 KB across 286 > entries, over half of it finished work, with one single entry at 16 KB. A file that cannot be read > is a file that cannot be checked — and this project has already paid for that twice: a record > nobody could find because it sat inside an entry about something else, and a finding rediscovered > because nobody could see it. The register now holds **open work only**, so its size tracks the work > rather than the project's age. > > **A sibling rather than the bottom of the register**, deliberately: appending to the same file keeps > the byte count and the scroll, which is the thing being fixed. > > **Load-bearing reasoning was NOT compressed away.** Where a closed row states a rule, a fence or a > deliberate refusal, that sentence is carried here verbatim under **Reasoning kept**. Rules that > outlive their work item also live in their proper homes — `workspace-CLAUDE.md` standing rules, > `felhom.eu/CLAUDE.md`, `CONTEXT.md`, and the architecture folder — and this file is not their > primary record. > > **This file is not the register.** Nothing here is open. `OPEN-ITEMS.md` remains the single source > of truth for open work; `scripts/one_register_gate.py` enforces that against `ROADMAP.md`. --- ## 2026-10-05 (late night) — burn-down round 2: the operator's answer (43 accepted), small rows fixed with releases The full text of every row below: `git show e8c56c44:documentation/backlog/OPEN-ITEMS.md` (the commit before each closing commit; the burn-down evidence table is `audits/burndown-2026-10-05/`). | Row | What | Closed | Evidence | |---|---|---|---| | **R-93** | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's p (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A choice between two test fixtures that no longer exist. Fix would cost: a new fixture (a design job). If never: nothing breaks. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-161** | **The volume-persistence gate is enforced by CONVENTION, not automatically.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The check that app data lands where backups see it runs by hand, not on every push. Fix would cost: CI starting ~53 apps per push. If never: a bad template can ship until the next periodic check. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-162** | **`docker diff` is the gate's only witness, and its failure mode is quiet.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: If Docker ran on an unusual storage driver, one catalog check would blame the wrong thing. Fix would cost: ~1 h for a driver nobody runs. If never: nothing; the check still fails safely. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-169** | **CI can only report, because there is no gate in the road.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: CI reports after a push lands instead of blocking it (no pull requests here). Fix would cost: a pull-request workflow for every change. If never: a forbidden --no-verify push could land broken code until the CI mail. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A removed storage permission can still look present for minutes (Proxmox cache). Fix would cost: a new probe in the agent. If never: the self-repair notices up to ~16 min late, then repairs. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-210** | **Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by `docker pull`** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: 193 old controller/hub images sit only on DooPlex (~27 GB). Fix would cost: a careful delete on DooPlex. If never: 27 GB stays used; ~199 GB is free. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-284** | **„A kiválasztott tárhely majdnem megtelt." on a store that is 93 % FREE — an apparent inverted threshold.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: An 'almost full' warning reported on an empty disk was a misreading; the page never shows it there. Fix would cost: nothing to fix. If never: nothing. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-346** | **`ActiveEnterTimestamp` answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A warning about a mistake nobody has made (the audit found 0 cases). Fix would cost: nothing left. If never: nothing today. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-367** | **The database dumps already written under the wrong name are stranded, and nothing will ever collect them.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: One old 312 KB database dump sits on the demo-hp test box under an old folder name. Fix would cost: a hand delete on a demo box. If never: a small file stays on a test box. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-371** | **The off-site tier is the only backup tier that announces nothing on success.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The weekly off-site backup sends no 'done' message (failures and staleness already alarm). Fix would cost: a new event in two repos. If never: success stays silent, as now. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-372** | **A Tier-2 copy that has NEVER been produced because its source path is missing is not surfaced prominently to the operator.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: An idea to show 'second copy never made' apart from 'second copy failed' to the operator. Fix would cost: a design and a new state end to end. If never: the existing loud warning stays. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-374** | **Three C1 refusal cases were judged borderline, left unfiled, and never named — so nobody can re-open the judgement.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A July audit says three borderline cases were left out but never named them. Fix would cost: hours to guess which three. If never: nobody can re-judge them; later sweeps exist. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-393** | **A decision-log skill for unattended runs was considered and deliberately deferred.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A proposed tool to log every small decision an unattended run makes. Fix would cost: a small project. If never: the bigger decisions are already recorded under the rules file. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-420** | **`controller_gates.py` could not express a NON-BLOCKING gate before 2026-09-01** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The felhom.eu gate runner cannot mark a gate as advisory only. Fix would cost: add it when one is needed. If never: nothing today. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-424** | **`one_register_gate.py`: a real defect parked under the roadmap state `idea` is invisible to it.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The roadmap gate cannot tell a real defect filed as an idea from an idea. Fix would cost: no mechanical fix exists. If never: a person must read the roadmap. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The hub's memory suggestion for an app can use samples from a short test install. Fix would cost: a hub change (~1-2 h). If never: a misleading suggestion for up to 7 days after a test install. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: BookStack's uploaded files cannot be checked automatically after an upgrade. Fix would cost: an upstream change or a browser step. If never: upgrades stay half-checked automatically. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-503** | **[P3-LOW] SPIKE (not built): an install-time disk rule for the public ISO — "exactly one internal disk → install; otherwise stop in Hungarian".** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Letting the installer pick the disk when there is only one (you ruled no). Fix would cost: reversing your rulings. If never: a person keeps choosing the disk. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A catalog 'locked after install' flag does nothing visible (all settings are read-only anyway). Fix would cost: delete it everywhere, or build an edit page. If never: nothing a household meets. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Vaultwarden shows a sign-up form although sign-up is off; the server refuses it. Fix would cost: an upstream fix. If never: a stranger sees a form that fails. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-551** | **[P3-LOW] No Tier-0 box can put the escrow ceremony in the state R-546 fixes — paused AND connected to its agent — so the readiness branches are proven only by tests.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The escrow 'waiting for the agent' screens are tested but never seen on a real box. Fix would cost: a token setup on the scratch box. If never: small chance the live page differs; the state lasts ~17 min. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-610** | **[P2-MEDIUM] The power cut AFTER the new version has started was never measured — R-520 closed on the safe half only.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A power cut inside a sub-second startup phase was never measured (same code as the measured phase). Fix would cost: a new fault injector and a drill. If never: nothing new would be learned. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Old opengist /login bookmarks answer 404 after an upstream move. Fix would cost: one help-text line. If never: a household re-bookmarks once. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Three live proofs a scratch box cannot give, and one log line 20 minutes off. Fix would cost: special test venues. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Deleting a customer does not remove their Cloudflare tunnel and DNS; the dialog says so. Fix would cost: a new Cloudflare step in the delete. If never: you remove them by hand, guided by the dialog. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-768** | **[P3-LOW] Grimoire is not built: upstream rules out public exposure and publishes no image for its current line.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Grimoire is not offered (upstream rules out public use); the row only watches upstream. Fix would cost: a re-read each catalog campaign. If never: nothing; Karakeep covers bookmarks. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-793** | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Four apps contain paid-edition code that is off; the row says never turn it on. Fix would cost: nothing (the rule lives in the licence audit). If never: nothing unless someone enables it. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-796** | **[P3-LOW] MeTube's "send to MeTube" helpers (browser extensions, bookmarklets, phone apps) cannot work behind the family gate.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: MeTube's browser 'send' helpers cannot pass the family gate. Fix would cost: a new token design. If never: households paste links in the page. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-797** | **[P3-LOW] `check-family-gate.py` rule 3 (a family_gate template needs a baked golden ≥ 0.287.0) is checked only where the felhom.eu sibling exists — CI's single clone cannot.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Catalog CI cannot run one gate rule; it says 'not checked' and the push hook checks it. Fix would cost: a CI checkout of a second repo. If never: only a forbidden bypass could skip it. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-804** | **[P3-LOW] plant-it's image is gone from Docker Hub: `msdeluise/plant-it:0.10.0` answers "pull access denied … repository does not exist".** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: plant-it's image no longer exists; the template is already marked not installable. Fix would cost: hide the template (small). If never: nothing. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-190** | **A storage ACL that demonstrably WORKED in the morning was gone by mid-morning, and nothing recorded its removal.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A storage permission vanished once in August; the agent now restores it and mails you. Fix would cost: hours of live experiments. If never: one mail per recurrence; it self-repairs. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-412** | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A rare race can push one hollow off-site copy of one app; the next night repairs it. Fix would cost: a change at the backup/push boundary plus a drill. If never: rarely, one app's off-site copy is hollow for a day; the second-drive copy stays good. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A pinned (frozen) app can get a newer health check, which can only cause a false alarm. Fix would cost: a fiddly change on the sync path. If never: maybe one false alarm one day. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-584** | **[P2-MED] Credential-bearing probe scripts were left in a live guest's `/tmp` for hours, across three releases.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Helper scripts with the shared demo password were left in a demo box's /tmp. Fix would cost: a wrapper tool. If never: demo-password litter on throwaway boxes. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-586** | **[P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The ISO bootstrap harness runs at each ISO release, not on every push. Fix would cost: Docker-capable CI. If never: a break is caught at the next ISO release. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-738** | **[P2-MEDIUM] Every wger update that brings database migrations leaves wger broken, and the guarded Update reports `done`.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The update's health check sees only the front page, so data broken behind it passes. Fix would cost: a per-app data check in every template. If never: catalog tests keep catching it before release. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-778** | **[P3-LOW] A box that rolls back to a controller ≤ 0.285 after v0.286 keeps the new traefik (an old controller never rewrites a running traefik) — and the old `clientIP` believes the LEFTMOST X-Forwarded-For, which a stranger then writes: the dashboard's login counter becomes dodgeable until the box moves forward again.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A box that rolls back below controller 0.286 trusts a forged address in the login counter. Fix would cost: a patch release or a rollback rule. If never: a short window until the next update. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-783** | **[P3-LOW] SparkyFitness: three wrong sign-ins by anyone shut EVERY visitor out of sign-in for ~10 s — a stranger retrying every 10 s keeps the household out.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Three wrong SparkyFitness logins block sign-in for everyone for ~10 s. Fix would cost: an upstream setting. If never: a persistent stranger can annoy the household. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-853** | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: After a boot, the box's versions reach the hub up to ~15 min late. Fix would cost: a new agent mode and release. If never: one report late; nothing lost. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A drive move keeps the app's records: fixed in controller v0.276.0 with tests; never seen live on a two-drive box. Fix would cost: a live two-drive move. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A fresh install drops an old update hold: fixed in v0.278.0 with tests; not seen live. Fix would cost: a live install with a leftover hold. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-706** | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Removing an app with its backups also deletes its off-site test copy: fixed in v0.279.0 with tests; not seen live. Fix would cost: a live remove with an off-site copy. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: No 'box recovered' alarm in a new box's first hour: fixed in hub v0.126.0 with tests; not seen at a real first install. Fix would cost: watch the next real install. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | --- ## 2026-10-05 (night) — the burn-down: stale rows closed with evidence, small rows fixed (rule: fix small, do not file) The full text of every row below: `git show ab2b3049:documentation/backlog/OPEN-ITEMS.md` (the commit before each closing commit; the burn-down evidence table is `audits/burndown-2026-10-05/`). | Row | What | Closed | Evidence | |---|---|---|---| | **R-885** | **The Python tests under `felhom.eu/scripts/` did not run in CI (P4).** New gate `script-tests` (`scripts/script_tests_gate.py`) walks `scripts/` and runs every `test_*.py` on every push (13 suites, ~20 s); a suite's verdict is its exit code; finding none fails; nested runs step aside. The hub-DB script tests use a Python-sqlite3 stand-in when the CI runner has no `sqlite3` CLI (said out loud). Decoys (5) declared in `test_gate_decoys.py`; red-proofs R885-a/b convict. | CLOSED 2026-10-05 — FIXED (gate) | `scripts/script_tests_gate.py`; `scripts/test_script_tests_gate.py`; `audits/burndown-2026-10-05/r885-red-proof.txt` | | **R-184** | **Nothing prevents the hub from vouching an agent version that was never released.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Fixed by felhom.eu b55fc17d "hub v0.102.0 — refuse to vouch a version that cannot be installed (R-273)" — exactly shape (b), validate at vouch time in the hub. felhom.eu/hub/internal/web/configs.go:1358 `res := s.gitea.PackageDownloadable(ctx, t.pkg, t.version, t.file)` and :1365 `s.logger.Printf("[WARN] artifact vouch REFUSED: %s package %s is NOT downloadable (R-287)", ...)`; tag leg at :1343 TagServesFile; unreachable registry also refuses. | | **R-207** | **`DRY_RUN=1` on `node-housekeeping.sh` is NOT non-mutating — it destroys the metric history it is supposed to let you inspect** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Fixed by homelab-manifests fc9fbb8 "node_housekeeping: guard DRY_RUN, correct the expired Docker rationale, pin container log rotation". /home/kisfenyo/git/homelab-manifests/homelab-ansible/roles/node_housekeeping/templates/node-housekeeping.sh.j2:137 `if [[ "${DRY_RUN}" == "1" ]]; then` inside write_metrics, :138 logs "file left untouched". | | **R-287** | **`felhom-agent` CI is red for a TRUE reason, and the diagnosis it was filed under is wrong in every particular.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Deleter established 2026-08-10 (R-267 newest-10 prune, recorded in the row itself); CI fixed by felhom-agent 53d047a "Two guards, one number: bound the published check to the retention it must live with" (R-291). felhom-agent@e06ed97 scripts/check-published-versions.py:101 `RETENTION_FILE = os.path.join(os.path.dirname(os.path.abspath(__file__)), "retention-policy.json")`, :213 `keep = retention_kept()`; scripts/retention-policy.json:37 `"generic_versions_kept": 10,`. Follow-up row R-291 is open (OPEN-ITEMS.md:439). | | **R-289** | **R-182's register row describes a defect the code no longer has — an OPEN row that is a false alarm.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | R-182 was closed by felhom.eu ef6ac6fe (2026-08-22, register compression): documentation/backlog/CLOSED-ITEMS.md:474 `/ **R-182** / ... / **CLOSED — SHIPPED** (controller v0.194.0 + hub v0.90.0/.1, 2026-08-03) /`. The residue (digest never seen delivering) was since observed: documentation/audits/DRILL-chaos-night-2026-09-17.md:181 `backup_run_failures` „1 of 12 apps failed to back up in this nightly run: nextcloud" listed as an alarm that fired and was true. | | **R-373** | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Premise (20G/50G two-volume mismatch, 'nothing sets SysDataGrowGB') was retired by agent v0.120.0 one-data-volume work, commit cd6e267 'v0.120.0 — one data volume (R-165...)'. felhom-agent/internal/reconcile/bringup.go:191 '// SysDataGrowGB is a COMPATIBILITY INPUT since agent v0.120.0 (R-165). There is no longer a second' and :437 'growGB := spec.DataVolGrowGB + spec.SysDataGrowGB'; installer passes it (felhom-host-install.sh:3140 '-sysdata-grow "$SYSDATA_GROW"') and records the sizing at :2041-2058. Note: cd6e267 predates the row's filing date; the row quoted an older audit. | | **R-390** | **The golden-bake runbook omits `pveam update`, and the failure it produces names the wrong cause.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Commit 2344589a ('... runbook pveam note'); felhom.eu/documentation/runbooks/RUNBOOK-manual-build.md:154 '2. Run **`pveam update` first** — the `virgin` snapshot's template INDEX is stale too, and a stale index fails as a bogus'. | | **R-427** | **`closed_register_gate.py` checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in `OPEN-ITEMS.md` — is unchecked, and there are TWELVE.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Commit 71b8c8c6 (Backlog triage Part B: '... closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS'); felhom.eu/scripts/closed_register_gate.py:10 'RULE 3 — (2026-10-03) no row in `OPEN-ITEMS.md` may carry a CLOSED-family word (CLOSED, SHIPPED,'. Of the 12 named rows, R-385/387/341/378/405/88a/88b/123 are now only in CLOSED-ITEMS.md; R-190 and R-352 remain open (partly-closed, as the row predicted). | | **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | felhom.eu 71b8c8c6 'Backlog triage Part B: 125 finished rows + 20 id-less rows moved to CLOSED-ITEMS ... closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS (decoys, seen red) ... register 444 -> 325'. felhom.eu/scripts/closed_register_gate.py:10 'RULE 3 — (2026-10-03) no row in `OPEN-ITEMS.md` may carry a CLOSED-family word'. Gate run today: 'closed-register gate OK — no open work filed as closed, no id in both registers.' (456 closed / 336 open, 0 convicted). | | **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Lesson homed and harness uses the correct probe. app-catalog-felhom.eu b7ef0c4 'upgrade-test.py: record the engine's own view of its datadir'; app-catalog-felhom.eu/scripts/upgrade-test.py:278 '"mariadb-upgrade --check-if-upgrade-is-needed --user=root "'. felhom.eu d6837d98 (SPIKE R-459); felhom.eu/documentation/architecture/09-update-architecture.md:1647 '1. **Ask the engine, not the log.** MariaDB's entrypoint prints `MariaDB upgrade not required` on an' (cites R-464). | | **R-501** | **[P3-LOW] The documented "confirm your CI run" recipe reads only the LAST page of the jobs list, and that list is not in id order — so it can report a run as missing that exists and passed.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | felhom.eu a4993272 'CLAUDE.md: the CI-check recipe was wrong in two ways, both measured today'. felhom.eu/CLAUDE.md:176 'rows — a run can sit several pages earlier. **Scan every page** and match on `head_sha`; with a'; recipe at CLAUDE.md:166-168 loops every page. | | **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | felhom.eu e02bc038 'hub v0.119.0 — ... R-596/R-598 closed' added the finding; felhom.eu/documentation/architecture/10-localisation.md:809-812 'the `felhom_lang` cookie and got the **Hungarian** page for `en`. ... The cookie is the right instrument for the anonymous claim page and the **wrong**' and :509 '`langFor`'s order is fixed: `?lang=` → **the household's setting when a session exists**'. Only the optional pointer from the workspace live-validation rules is absent (grep ?lang=/felhom_lang in CLAUDE.md files and .claude/rules: none) — a 5-minute add if wanted. | | **R-617** | **[P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through `repos/migrate` — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it".** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | felhom.eu@462ab4a5 (2026-09-22) documentation/architecture/09-update-architecture.md:1969 "`POST /api/v1/repos/migrate` is the route that works (the project's Gitea tokens carry" - continues at :1970 "`write:repository` but not `write:user`, so `POST /user/repos` answers 403"; recipe at :1979. The one-line note the row asked for exists (in the architecture doc rather than operations/). The optional operator-scoped token is a separate wish, not the defect. | | **R-705** | **[P3-LOW] There is no way to run the night's chain now — only its pieces.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | The remaining half (manual whole-guest backup) EXISTS and predates the row: felhom-controller bbed5af (v0.47.0, 2026-06-12) 'backups page — whole-guest backup visibility + manual trigger'. Live source @7690c27: controller/internal/web/backup_handlers.go:324 `case r.URL.Path == "/api/guest-backup/trigger" && r.Method == http.MethodPost:`; :340 `if err := s.backupTrigger.TriggerNow(); err != nil {`; quiesce.go:427 "manual backup requested — quiescing now" (bypasses due-ness, all tiers); wired cmd/controller/main.go:2099 and the page button backups.html:229. Agent side: felhom-agent internal/localapi/server.go:514 `mux.HandleFunc("POST /backup", ...)`. The controller half was built v0.279.0 (night-chain, handler_debug.go:79). The row's 'no manual trigger' claim was not true at writing; it is not chained into night-chain, which the row did not require. | | **R-766** | **[P3-LOW] A new app's logo and screenshots reach the boxes only with the next HUB release — the website alone is not enough.** (P4) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Hub releases after the assets push (felhom.eu 40f07429, 2026-10-01): hub v0.131.0 (2026-10-04) .. v0.136.0 (d4be9f6f, 2026-10-05). The build copies website assets: felhom.eu scripts/build-hub.sh:98 `cp "${WEBSITE_ASSETS_DIR}"/*-logo.svg "${BUILD_DIR}/assets/" 2>/dev/null // true`, and the hub build workspace /mnt/5_hdd/felhom.eu/build/felhom-hub/workspace/assets/ holds radicale-logo.svg + 3 screenshots (also karakeep, dawarich) dated Oct 5 14:28; hub/Dockerfile:27 `COPY assets/ /usr/share/felhom/assets-seed/`. Not checked: what a live box shows (no machine access). Checked 2026-10-05: the live hub image (v0.136.0) `/usr/share/felhom/assets-seed/` holds radicale-, karakeep- and dawarich-logo.svg + screenshots. | | **R-50b** | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | Claim 'fetched via fetch_raw from raw/branch/main — no tag, no pin' no longer true: bee68484 (installer v1.23.0, R-110/R-183) pinned fetch_raw to the vouched agent tag — felhom.eu scripts/felhom-host-install.sh:533 '"$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/tag/v$ART_AGENT_VER/$path" \' (leg b). Leg (c)-like signed delivery: felhom-agent c9fa2e7 (R-840 config bundle) and configs/test_felhom_config_bundle.py:264 covers /usr/local/sbin/felhom-pbs-apply. Residual worth one line if kept: the 0440 sudoers drift visibility note. | | **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | 3d7a2761 (hub v0.135.0, R-530/R-604 'boxes left behind listed and alarmed'): felhom.eu hub/internal/osupdates/service.go:79 'EventAgentBehind = "agent_behind" // warning, operator' with :180 'AgentBehindAfter: a box runs an agent older than the vouched one this long → an operator alarm' (7 d window, the staleness window the row asked for). | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | The remaining half (customer-facing recovery-code form: yell → R form → preview) shipped as the recovery screen: felhom-controller 636c51e 'R-193: the recovery screen — unlocking, and only unlocking (v0.200.0)'; controller/internal/web/templates/recovery.html:82 '
', routed at internal/web/server.go:602. Plumbing half was 1b1366b (v0.196.0). The 64-hex inject-password route (server.go:783) stays a deliberate DR fallback with no form. | | **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | The row's only remainder was 'the other ten PostgreSQL apps need two-venue proof'. R-463 CLOSED 2026-09-30 by felhom.eu 25cb3eb9 ('The last six PostgreSQL apps decided'): 8 of 11 moved by the box's own conversion, 3 (zipline, adventurelog, immich) stay by decision 42; CLOSED-ITEMS.md:223. Source proof: app-catalog templates/docmost/docker-compose.yml:62 'image: postgres:18-alpine' (also rallly:67, outline:64, paperless-ngx:101). The per-app engine gate stays as the permanent rule (catalog CLAUDE.md:115-122). | | **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | The residual (a unit-restore-recreated volume has no compose label, so the remove answered []) was fixed in felhom-controller 206b035 (v0.268.0, R-658). controller/internal/stacks/delete.go:1089-1090: '// appVolumeSet is every volume the removal accounts for: the ones carrying the project label AND the // ones the app's definition declares that Docker holds by name (R-658, v0.268.0).' delete.go:703 'resp.VolumesRemoved = removedVolumes(volsBefore, m.appVolumeSet(name, stackDir))'. | | **R-573** | **[P3-LOW] The agent-channel and endpoint-drift banners reach the dashboard as finished Hungarian, so they stay Hungarian on an English page.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | felhom-controller 7c4a33b (v0.258.0, 'the last four Hungarian things an English household met ... R-573 the two channel banners'). controller/internal/web/alerts.go:113 'func (am *AlertManager) SetAgentChannelAlert(down bool, msgKey, msg string) {' and :136 'func (am *AlertManager) SetEndpointDriftAlert(drift bool, msgKey, msg string) {'. Both set MessageKey with msg only as a fail-open fallback. | | **R-622** | **[P2-MEDIUM] `adventurelog v0.13.0` migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | adventurelog v0.13.0 was diagnosed, fixed and promoted with a two-venue test record in app-catalog-felhom.eu 06ea7da (2026-09-27, 'adventurelog: v0.12.1 -> v0.13.0 with its health and world-data fixes in the same commit (R-655, 09 decision 41)'; bench healthy in 217 s, box 9202 through the guarded Update in 204 s). templates/adventurelog/docker-compose.yml:13 ' image: ghcr.io/seanmorley15/adventurelog-backend:v0.13.0'. The proven step is recorded at templates/adventurelog/.felhom.yml:132 (update_ladder entry from v0.12.1). | | **R-635** | **[P1-HIGH] `romm 5.3.0` does not fit the memory the template gives it, and the guarded Update called that a success — the app has been OOM-crash-looping on demo-hp for six hours at ~500% CPU.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | The open remainder (app_oom fires once per container run, no escalation) was built in felhom-controller 0054d4b (v0.265.0, 'OOM storm alarm', R-636). controller/internal/notify/notifier.go:746 '\t\tn.emit("app_oom_storm", "error",' fires once per run when 20 or more kills land in 30 min (:754-764, oomStormKills=20, oomStormWindowMin=30; pinned by TestR636_*). The 79 % headroom and the method lesson are carried by R-462 (per the row). | | **R-235** | **The appliance console keeps telling an already-paired box to go and pair itself.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | felhom.eu c033b3b6 'ISO 1.28.0 source: the console stops showing the pairing code once bound (R-535)'. scripts/iso/felhom-bootstrap.sh:538: `print_bound_banner # R-535: replace the pairing code on the console with the truth`. Same defect already CLOSED twice in CLOSED-ITEMS.md as R-535 (line 232) and R-214 (line 200, 'proven on a fresh install'). | | **R-282** | **One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing.** (P3) | CLOSED 2026-10-05 — FIXED BY LATER WORK (burn-down Part A) | felhom.eu 4d6ec7c 'hub v0.104.0: ... the hub half of the naming (R-295)' + controller v0.211.0 (R-295 CLOSED, CLOSED-ITEMS.md:207) + R-323 hub v0.105.0. hub/internal/notify/templates.go:204: '// R-295, HUB HALF (2026-08-13). ONE NAME PER SECRET, and it is „Beállító kód".'; hub/internal/claim/engine.go:51 `EmailReenroll EmailKind = "reenroll"` (mail names the setup page a rebuilt box shows). | | **R-376** | **The placement decision that cost four mis-filed defect reports was never recorded as a decision anywhere, and the architecture folder's own marker convention lives in one document of eight.** (P4) | CLOSED 2026-10-05 — DONE | The legend reached the three documents written after the 2026-08-22 pass (`08`, `09`, `11`); all eleven numbered architecture documents now carry it (`grep -c "not yet classified"` = 1 each; `07` is the original home). The hot/bulk decision was given its home on 2026-08-22 (`01` [DESIGN] + CONTEXT). Marking every statement stays a practice of each session that touches a document, not a defect. | | **R-817** | **`09` decision 56 and R-745 disagree about what the controller self-update rolls back to.** (P4) | CLOSED 2026-10-05 — CLARIFIED (no contradiction) | `felhom-agent internal/localapi/controllerswap.go:236-240` records the image running when a swap starts as `Previous`; `:289` writes it back on failure. After a good swap that is „the one before" the running image — decision 56 and R-745 name the same image. A dated clarification sits under decision 56 in `09`; the ruling text is unchanged. | | **R-818** | **Two changelogs cite register ids for other findings.** (P4) | CLOSED 2026-10-05 — CORRECTED | Dated correction notes under hub v0.109.0 (`hub/CHANGELOG.md`) and controller v0.224.0 + v0.225.0 (`felhom-controller/CHANGELOG.md`): those two findings never had register rows of their own — the triage's „the real ids are in CLOSED-ITEMS" was itself wrong (no closed row names hub v0.109.0 or controller v0.224.0/v0.225.0). Nothing renumbered. | | **R-755** | **[P3-LOW] wger runs Django's DEVELOPMENT server in production: `manage.py runserver`, because the template does not set `WGER_USE_GUNICORN=True`.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-762 (its unique fact moved there) | Still true: templates/wger/docker-compose.yml has no WGER_USE_GUNICORN (grep empty). R-762 (open, read) states 'Owner decides together with R-755 (same server question)' and its fix names 'the gunicorn switch of R-755'. | | **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-440 (its unique fact moved there) | felhom-controller/controller/internal/stacks/updateorder.go:96: `if len(s.CatalogDigests) == 0 // s.CatalogTestedAt.IsZero() { return false }` — blind only for apps with no ladder entry, i.e. the same 15 templates R-440 lists (app-catalog has no update_ladder for them). Both rows close by the same act: each app's first proven ladder step (R-462). | | **R-799** | **[P3-LOW] The MeTube fixture's `POST /add` leaves out `download_type`, which upstream's validator lists as required.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | app-catalog `29ac711`: `MeTube.add_body()` sends `download_type: video`; `scripts/test_upgrade_fixtures_metube.py` (red-proof: the field removed → FAIL). Not exercised on a box (the next MeTube step will). | | **R-761** | **[P3-LOW] The canonical example template tells a new app's author the logo is `-logo.webp`; the controller loads `-logo.svg`, then `.png`.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | app-catalog `29ac711`: the canonical template comment, `REUSE.md` and `NEW-APP-CHECKLIST.md` name `{slug}-logo.svg` then `.png` (controller `config.go` AppLogoURL/AppLogoPNGURL). Comment-only. | | **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | app-catalog `29ac711`: `CLAUDE.md` states its `REPORT.md` carries no observations section by convention (the row's second option); the shared gate was not copied. | | **R-291** | **CI's installability assertion is now BOUNDED by a retention number, and the narrowing is recorded here so it can be widened deliberately rather than discovered.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-agent `d833163`: `scripts/retention-policy.json` names the source of its 10 — the R-267 newest-10 prune, established 2026-08-10 (R-287) — and drops the non-existent `registry-retention.md` reader; `check-published-versions.py` still reads 10 (checked). The min_agent-floor bound stays recorded in the file as the better bound. | | **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-agent `d833163`: `internal/backup/store.go` says a restart blanks the reported backup list until the next run; only the hub's verdict (7-day look-back) is unaffected. Comment-only. | | **R-263** | **C7 — „This is the ONLY writer of `StoragePath.BackupTarget`" is false, and nothing pins it.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-controller `114ff27`: comment „the only writer that GRANTS"; `internal/settings/r263_backup_target_writers_test.go` scans every non-test file under internal/ and cmd/ (assignments and composite-literal keys). Red-proofs: ClearBackupTarget writing true; a `BackupTarget: true` literal in internal/web — both convict (`audits/burndown-2026-10-05/r263-red-proof.txt`). | | **R-368** | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-controller `114ff27`: the `IsDefault` comment says the deploy FORM pre-selects it and the deploy API applies no default (the row's second option; behaviour unchanged on purpose). | | **R-418** | **`repo_gates.py`'s docstring listed ELEVEN gates while THIRTEEN were registered** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `repo_gates.py` docstring lists all 17 gates; `scripts/test_repo_gates_docstring.py` asserts list == GATES in order (red-proof: one line removed → FAIL), run on every push by the script-tests gate. | | **R-345** | **`hub/Makefile` tags and pushes `:latest`, which the project's own rules forbid in two places.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `hub/Makefile` AND the real release script `scripts/build-hub.sh` (3 lines; the build dir links to it) no longer tag or push `felhom-hub:latest` — nothing pulls it (grep of all repos + homelab-manifests). `scripts/test_no_latest_push.py` (walk of hub/ + scripts/; red-proofs: the old Makefile and the old build-hub.sh each FAIL). Whether a stale `:latest` sits on the registry was not checked. | | **R-416** | **`closed_register_gate.py` still has no within-register duplicate-id rule.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `closed_register_gate.py` RULE 4 refuses an id twice in CLOSED-ITEMS.md (OPEN duplicates were already `register_shape_gate.py` RULE 3); 0 duplicates existed, so it registered green. Decoy `closed-register/duplicate-closed-id` (red-proof: RULE 4 off → LIVE HOLE). | | **R-261** | **C6 — `CountSelfBindTokens` exists so that callers can assert an invariant, and no production caller asserts it.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `hub/internal/store/selfbind.go` names `CountSelfBindTokens` a test accessor and the two tests that pin the auto-mint invariant. Comment-only; no hub release needed. | | **R-262** | **C7 — a comment claims a cross-repo contract is mirrored „field-for-field" and „the key-set tests guard drift"; it is two fields short, AND THE FIXTURE THE TEST READS OMITS THE SAME TWO FIELDS.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: the hub comment says hostRestoreTest is a deliberate SUBSET; `hub/internal/api/r262_restoretest_subset_test.go` pins the hub fields and the known-unmodelled agent fields and cross-checks the agent source beside it — which found a THIRD unmodelled field, `skipped` (agent v0.133.0, R-672; by design a skipped test reads as failed with its reason). Red-proofs: drop `skipped` from the list / stop decoding source_tier → FAIL. | | **R-286** | **A control drawn from the same channel as the measurement cannot detect a defect in that channel — and this one passed while the measurement was wrong.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: workspace standing rule 3 (both `CLAUDE.md` copies, identical) adds: a control must come from a DIFFERENT channel than the measurement; a hub-state check copies `hub.db-wal` or asks the running pod. | | **R-588** | **[P3-LOW] ISO release records live in two different places, so "was the gate run for this image?" cannot be answered by looking.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `runbooks/iso-release-gate.md` names `documentation/tests/iso-release--/` as the one home; `tests/iso-release-1.28.0-2026-09-16/README.md` points at the 1.28.0 record inside the 2026-09-16 audit. | | **R-423** | **`site_gates.py` checks a hardcoded `PAGES` list of seven files; a new page is not scanned at all.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `site_gates.py` walks website/ and fails on any *.html not in PAGES (all 9 listed today); decoy `site/unlisted-page` in `test_gate_decoys.py`; the `site` exemption is gone from `decoy_coverage_gate.py` (33 covered, 18 exempt). Red-proof: the walk switched off → LIVE HOLE (`audits/burndown-2026-10-05/r423-red-proof.txt`). | --- ## 2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; `09` rulings 125–127) The full text of every row below: `git show cea8502f:documentation/backlog/OPEN-ITEMS.md`. | Row | What | Closed | Evidence | |---|---|---|---| | **R-519** | **After a backup torn by a power cut, a restore point carried the new database dump's time over the previous run's files, and no screen said the run was interrupted (P2).** Fixed in controller v0.296.0 (page notice + synthesised status; dating by the oldest part since v0.275.0). **Live on 9202 (operator ruling 126):** a complete run, then a run cut by `docker restart felhom-controller` 0.8 s after bookstack's config dump and before its database volume; afterwards both /backups and /backups/apps carry `data-interrupted-run` („A legutóbbi mentés (2026-10-05 16:05) megszakadt …", negative control 0), the restore point reads 14:04:44Z = the run-1 database volume (its oldest part; config 14:05:58, SQL 14:05:54), the controller restarted bookstack itself; the next complete run cleared the notice and replaced the torn `.tar.tmp`. The throwaway bookstack was removed through the product (0 containers, volumes, folders, backups). | CLOSED 2026-10-05 — FIXED controller v0.296.0, proven live | `audits/hub-db-offsite-2026-10-05/partD/r519/`; `internal/backup/run_record_test.go` | --- ## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124) The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session). | Row | What | Closed | Evidence | |---|---|---|---| | **R-135** | **A cookie-less POST skipped the hub's CSRF gate, so a browser with cached Basic credentials could be made to POST cross-site.** hub v0.135.0: without a session a state change needs Basic credentials AND the header `X-Felhom-Operator` (decision 120); the gate sits before the route switch. 39 paths through RequireAuth→ServeHTTP; red-proof: the old shape lets all 39 through. Live: Basic + no header → 403 (also with `Origin: evil`, also on an unknown path); with the header → passes; header without credentials → 401. **Reasoning kept: a browser cannot add a custom header cross-site without a CORS preflight, which the hub never answers.** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partA/`; `web/r135_csrf_test.go` | | **R-133** | **Every box's break-glass console password was plaintext in hub.db.** hub v0.135.0: sealed with the off-site seal and key (decision 121); legacy rows sealed at start-up — live: 4 rows sealed, 0 left plain; the demo-hp reveal still returned a password that minted a PVE ticket (HTTP 200; a wrong one 401); a wrong key → 500, nothing in the body or the log, no event. **Reasoning kept: the running hub still holds the key — this closes the database-copy route only; a database backup without `OFFSITE_SECRET_KEY` cannot open the console passwords (R-173).** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partB/`; `store/r133_recovery_seal_test.go`, `web/r133_reveal_wrongkey_test.go` | | **R-604** | **A per-customer controller floor silently kept a box out of every global raise (demo-hp missed four).** hub v0.135.0: a global raise logs one line per customer whose own LOWER floor wins and sends ONE operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set; the System page's "Version floors" table lists every per-customer floor with its age and which ones the global cannot move. 2 red-proofs. Live: the table shows the three per-customer floors (age "unknown" — set before v0.135.0). The mail was not exercised live (it needs a global raise below an override). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `web/r604_floor_held_back_test.go` | | **R-530** | **Nothing listed which boxes still run an old agent (agents update only by a per-box signed job).** hub v0.135.0: the System page's Agent cell (box → vouched, how far, since when; red after the wait) and `agent_behind` after 7 days (decision 119). Live: Tester 2 reads `0.142.0 → 0.146.1`. Signing stays per box (the 2026-09-16 ruling: CC may sign until the first paying customer). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `osupdates/r530_agent_alarm_test.go` | | **R-508** | **Customer tester-1 had no registered e-mail, and the page did not say so.** The address has been set since 2026-09-14 (the connect mails reach it — Gmail-read 2026-10-05); hub v0.135.0 adds the page warning: a configured customer with no box and no e-mail shows a red line (three branches tested, red-proof). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partG/`; `web/r508_no_email_banner_test.go` | | **R-509** | **A box installed for an existing customer never got the connect e-mail.** Fixed in hub v0.114.0; the owed real-mail proof: three mails from the automatic "host delete" trigger, each within 1 s of the hub's own send line (2026-09-16 12:22:59 and 18:17:46, 2026-09-30 07:23:03 UTC), read through the Gmail connector (metadata only). The "e-mail set" trigger shares the send core and is unit-proven. | CLOSED 2026-10-05 — VERIFIED | `audits/hub-safety-2026-10-05/partG/r509-real-mails.txt` | | **R-880** | **An installed `felhom-os-apply` refuses a bundle naming a path it does not know (R16), so a release whose bundle ADDS a path cannot reach any box on an older bundle** (found 2026-10-05 before delivering v0.146.1, which adds four). Fixed by a step: `felhom-agent/scripts/build-step-bundle.py` — the box's current bundle with ONLY `felhom-os-apply` replaced (same paths), published as `0.146.1-step1`; then the release's bundle. Tests `StepBundle` (the R16 refusal reproduced; the step accepted; exactly one file changed). Live: demo-hp, demo-felhom and Tester 1 each took step1 (`written=1 same=20`) then 0.146.1 (`written=3 same=22`), self-check ok. **Reasoning kept: every future bundle that adds a path needs this step (decision 124); the step package stays published while any box may still be on the old bundle (Tester 2).** | CLOSED 2026-10-05 — FIXED (tooling, agent e4b5cf9) | `audits/hub-safety-2026-10-05/part{F,H}/`; memory `bundle-adding-a-path-needs-step-bundle` | --- ## 2026-10-05 (afternoon) — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0; rulings 109–111, CC decisions 112–118) | Row | What | Closed | Evidence | |---|---|---|---| | **R-871** | **No architecture covered a box that is not always on, and a missed night was never made up.** Decision 109 (option A) built: controller v0.295.0 `internal/nightchain` — a ledger of when each backup leg ran to its end; on a start or a host resume ONE catch-up 15 min later, backup legs only, never the update leg; a late daily timer after a suspend is skipped; the whole-guest backup and the catch-up wait for each other; decision 110's banner. Design `07` §6.1.1. Live: 9202 (dump made 15 min after the start; after a crash mid-wait, all three legs at the next start), demo-felhom (dump 15 min after the start; the household's timeline line reached the hub); banner served on 9202 and closed by its real route. 13 red-proofs. **Reasoning kept: the ledger records that a leg RAN, and is never read as evidence that a backup EXISTS.** Full text: `git show 1b0678fa:documentation/backlog/OPEN-ITEMS.md`. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partA/`, `partB/` | | **R-873** | **A household whose box is off every night was mailed "cannot be reached" every night.** hub v0.134.0: at most once per 7 days to the household (persisted), the operator every edge, the recovery mail stays paired (decision 116). Proven by test through the real dispatcher (red-proof); no live occurrence in the session (Tester 2 stayed off). | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r873-red-proof.txt` | | **R-874** | **A restore-test never ran on a box with short power-on sessions.** agent v0.145.0: first due-check 30 min after start (decision 117). Live on demo-felhom: start 07:38:46 UTC → `restore-test first evaluation after start (R-874)` at 08:08:46 → a due tier restored and passed in 29 s. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r874-*` | | **R-875** | **A kept report's reason said "the agent stopped mid-pass" for a hub-away pass.** agent v0.145.0: "sent late — kept on the box until the hub could take it". Test + red-proof. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r875-red-proof.txt` | | **R-876** | **After a power cut mid-update every later pass failed until a person ran `dpkg --configure -a`.** agent v0.145.0: dpkg's state = `--audit` AND the update journal in one call; repair on either; belt: repair + retry once when apt says "interrupted" (decision 118). Live (operator's go): crash at 07:56:03 UTC mid-unpack → back by itself → next pass `REPAIR configured=0 journal=1` → `DONE rc=0 upgraded=12`, no person, no mail; package list identical. **Reasoning kept: a check that reads one of two places dpkg keeps its state is a check that misses the other.** | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partD/` | | **R-877** | **The Tester 1 VM on demo-hp had no start-on-boot: the morning's demo-hp crash (06:14 UTC) left it off for 1 h 17 min, unnoticed** (the night-fixes report called every box healthy). Found 07:31 UTC; `qm set 341 --onboot 1`, started; the afternoon crash then brought it back by itself. Filed and closed in the same commit. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt` | ## 2026-10-05 (day) — the night's fixes: off-site clean-up guard, first-install image race, R8 download, a killed pass's report (controller v0.294.0, agent v0.144.0 + v0.144.1, golden 0.294.0; rulings 100–103, CC decisions 104–108) | Row | What | Closed | Evidence | |---|---|---|---| | **R-867** | **The off-site clean-up guard refused honest 7-day retention and mailed an error every window.** Controller v0.294.0: the guard's "young" line is keep-daily CALENDAR days, built from the same constants as the policy (decision 104); every other refusal kept. Tests run restic 0.14.0's policy itself (proven identical to the binary over 92 snapshots), 5 red-proofs. Live by hand: demo-felhom window 5 16 → 14, demo-hp window 6 145 → 127 — exactly the predicted snapshots; hub rows `pruned`, no event, no mail, key files clean. **Reasoning kept: a line that can sit inside the keep window is a line that refuses the honest case — derive it from the policy, never pick an age.** Full text: `git show 7221ee5c:documentation/backlog/OPEN-ITEMS.md`. | CLOSED 2026-10-05 — FIXED | `audits/night-fixes-2026-10-05/partA/` | | **R-95** | **The box could delete its own off-site history.** Append-only key since 2026-10-03 (decisions 68–69); the last open item — a clean-up window that actually removes snapshots — was observed 2026-10-05 on both demo boxes (R-867's live windows, opened by the operator's one-shot grant; the dated check's four conditions all hold). The unattended weekly window is the same code path and is due ~2026-10-11/12; it is not separately re-checked (the DUE-CHECKS entry is removed with this row). Residual: R-822 (an add-only attacker steering older keeps). Full text: `git show 7221ee5c:documentation/backlog/OPEN-ITEMS.md`. | CLOSED 2026-10-05 — FIXED | `audits/night-fixes-2026-10-05/partA/`; `audits/offsite-lock-build-2026-10-03/` | | **R-863** | **A new box's first app install could fail: the one-time image clean-up deleted the image compose had just pulled.** Controller v0.294.0: every compose command that can pull holds a shared lock (`dockerexec.BeginImageWork`); a clean-up pass takes it exclusively without waiting and otherwise does not run; the one-time pass is retried every 2 min and writes its marker only after a pass that ran (decision 105). Live on 9202: the clean-up fired at minute 3 (05:31:36 UTC) inside BookStack's 35.5 s install, skipped itself, BookStack installed first time; the retry at 05:33:36 ran and kept everything; teardown through the product. The update path was already guarded (`Updating`); restore/undo were exposed in principle and now hold the same lock. | CLOSED 2026-10-05 — FIXED | `audits/night-fixes-2026-10-05/partB/` | | **R-864** | **A failed compose logged the head of stderr (pull progress) and cut the reason.** Controller v0.294.0 `tailStr` (rune-safe), pinned with the night's real first line. | CLOSED 2026-10-05 — FIXED | `audits/night-fixes-2026-10-05/partB/red-proofs.txt` | | **R-869** | **The move-aside log line printed an empty destination.** Controller v0.294.0: assigned before the log line; test + red-proof. | CLOSED 2026-10-05 — FIXED | `audits/night-fixes-2026-10-05/partD/r869-red-proof.txt` | | **R-865** | **R8 measured every download as 0 B (`--print-uris` with `-s` prints no URIs).** Agent v0.144.0: no `-s`; the fake answers like real apt (verbatim 9202 output), 2 tests, red-proof. Live: the installed wrapper's `download_bytes` on demo-hp read **12 802 456 B** for 13 pending upgrades (0 before), nothing installed. A live R8 REFUSAL line was not produced: it needs < 500 MB free on demo-hp's guest, i.e. 28.5 GB written into a thin pool with 18.7 GB free — it would have stopped every guest. | CLOSED 2026-10-05 — FIXED | `audits/night-fixes-2026-10-05/partC/` | | **R-866** | **The debug OS pass could not run with the hub away.** Agent v0.144.0: the daemon saves the hub's block; the selftest falls back to it and says `block=SAVED(