diff --git a/STATUS.md b/STATUS.md index 81d40a8d..449e0c0c 100644 --- a/STATUS.md +++ b/STATUS.md @@ -3,8 +3,16 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was sent to it.** -**Updated 2026-10-05 (night, the burn-down session): every box of ours healthy (nothing was changed on any box). The -open-items list went from 336 to 292. Report: `REPORT-burndown-2026-10-05.md`.** +**Updated 2026-10-05 (late night, burn-down round 2 — in progress): 249 open rows after your answer (was 292).** + +## Your answer to the burn-down list (2026-10-05 18:23), recorded + +- **43 rows closed as accepted by you**, each with its reason from the list. +- **R-124 stays open and is fixed in this session** (a recovery step that fails during a real recovery). +- **R-698 stays open as a known risk** (a backup keeps an app's image name, not the image). It is yours; nobody works on it now. +- **R-831 and R-870** (the printed tokens) stay as they are, by your earlier rulings; the rows keep their steps. +- **R-887** (the CI jobs that never ran): your screenshot shows ONE runner, online. My „old copy of the runner" guess was + wrong; the row says so. ## Tonight, later (2026-10-05, night): the list got shorter @@ -19,76 +27,13 @@ open-items list went from 336 to 292. Report: `REPORT-burndown-2026-10-05.md`.** it; it is gone, and a test keeps it gone. - **You got 5 „CI failed" mails tonight — my fault, fixed.** The new test gate ran one of today's test suites on the CI machine for the first time; it uses a different `date` tool, and 9 tests failed. Fixed; CI is green again (job 1360). -- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes, with no log and no mail. One re-run passed; the other was again never picked up. It looks like a leftover runner registration, but I cannot see the runner list. If you do nothing: now and then a CI run reads „failed" without testing anything. Fix: Site Administration → Runners, remove any offline duplicate of `felhom-gates-runner`. +- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes, + with no log and no mail. (The „old runner copy" guess was wrong — see above.) - **A new rule, so the list stops growing:** a small problem found during work (about 30 minutes) is fixed in that session and never added to the list. Every report now states four numbers: rows before, after, opened, closed. **The numbers:** 336 before → **292 after**; 1 opened (the CI fault below); 45 closed. -**Needs you:** -1. **The list below: 45 rows I would close as „accepted".** Answer per row, or „all as picked". If you do nothing, - they stay open and the list stays at 292. -2. **Two leaked tokens to rotate** (R-831, R-870, at the end of the list). If you do nothing, whoever has those - transcripts keeps that access. -3. Earlier tonight's items are unchanged (pin the other „latest" apps on DooPlex; Alertmanager's file permissions). -4. **Look at the CI runners** (Gitea → Site Administration → Runners; R-887): remove any offline duplicate of - `felhom-gates-runner`. If you do nothing, some CI runs fail without running. - -## Your list: rows I would close as 'accepted' (burn-down 2026-10-05) - answer per row, or 'all as picked' - -Each line: what it is. What fixing costs. What happens if we never fix it. **My pick.** I closed none of these. - -- **R-93** - A choice between two test fixtures that no longer exist. Fix: a new fixture (a design job). If never: nothing breaks. **Pick: close.** -- **R-124** - The disaster-recovery recipe writes the backup-server namespace as the word 'root'; the server wants an empty name, so one pasted command fails. Fix: a change across hub and agent. If never: an operator doing a manual restore drops one option once; no data risk. **Pick: close.** -- **R-161** - The check that app data lands where backups see it runs by hand, not on every push. Fix: CI starting ~53 apps per push. If never: a bad template can ship until the next periodic check. **Pick: close.** -- **R-162** - If Docker ran on an unusual storage driver, one catalog check would blame the wrong thing. Fix: ~1 h for a driver nobody runs. If never: nothing; the check still fails safely. **Pick: close.** -- **R-169** - CI reports after a push lands instead of blocking it (no pull requests here). Fix: a pull-request workflow for every change. If never: a forbidden --no-verify push could land broken code until the CI mail. **Pick: close.** -- **R-194** - A removed storage permission can still look present for minutes (Proxmox cache). Fix: a new probe in the agent. If never: the self-repair notices up to ~16 min late, then repairs. **Pick: close.** -- **R-210** - 193 old controller/hub images sit only on DooPlex (~27 GB). Fix: a careful delete on DooPlex. If never: 27 GB stays used; ~199 GB is free. **Pick: close.** -- **R-284** - An 'almost full' warning reported on an empty disk was a misreading; the page never shows it there. Fix: nothing to fix. If never: nothing. **Pick: close.** -- **R-346** - A warning about a mistake nobody has made (the audit found 0 cases). Fix: nothing left. If never: nothing today. **Pick: close.** -- **R-367** - One old 312 KB database dump sits on the demo-hp test box under an old folder name. Fix: a hand delete on a demo box. If never: a small file stays on a test box. **Pick: close.** -- **R-371** - The weekly off-site backup sends no 'done' message (failures and staleness already alarm). Fix: a new event in two repos. If never: success stays silent, as now. **Pick: close.** -- **R-372** - An idea to show 'second copy never made' apart from 'second copy failed' to the operator. Fix: a design and a new state end to end. If never: the existing loud warning stays. **Pick: close.** -- **R-374** - A July audit says three borderline cases were left out but never named them. Fix: hours to guess which three. If never: nobody can re-judge them; later sweeps exist. **Pick: close.** -- **R-393** - A proposed tool to log every small decision an unattended run makes. Fix: a small project. If never: the bigger decisions are already recorded under the rules file. **Pick: close.** -- **R-420** - The felhom.eu gate runner cannot mark a gate as advisory only. Fix: add it when one is needed. If never: nothing today. **Pick: close.** -- **R-424** - The roadmap gate cannot tell a real defect filed as an idea from an idea. Fix: no mechanical fix exists. If never: a person must read the roadmap. **Pick: close.** -- **R-445** - The hub's memory suggestion for an app can use samples from a short test install. Fix: a hub change (~1-2 h). If never: a misleading suggestion for up to 7 days after a test install. **Pick: close.** -- **R-460** - BookStack's uploaded files cannot be checked automatically after an upgrade. Fix: an upstream change or a browser step. If never: upgrades stay half-checked automatically. **Pick: close.** -- **R-503** - Letting the installer pick the disk when there is only one (you ruled no). Fix: reversing your rulings. If never: a person keeps choosing the disk. **Pick: close.** -- **R-527** - A catalog 'locked after install' flag does nothing visible (all settings are read-only anyway). Fix: delete it everywhere, or build an edit page. If never: nothing a household meets. **Pick: close.** -- **R-532** - Vaultwarden shows a sign-up form although sign-up is off; the server refuses it. Fix: an upstream fix. If never: a stranger sees a form that fails. **Pick: close.** -- **R-551** - The escrow 'waiting for the agent' screens are tested but never seen on a real box. Fix: a token setup on the scratch box. If never: small chance the live page differs; the state lasts ~17 min. **Pick: close.** -- **R-610** - A power cut inside a sub-second startup phase was never measured (same code as the measured phase). Fix: a new fault injector and a drill. If never: nothing new would be learned. **Pick: close.** -- **R-654** - Old opengist /login bookmarks answer 404 after an upstream move. Fix: one help-text line. If never: a household re-bookmarks once. **Pick: close.** -- **R-687** - Three live proofs a scratch box cannot give, and one log line 20 minutes off. Fix: special test venues. If never: the tests stay the proof. **Pick: close.** -- **R-688** - Deleting a customer does not remove their Cloudflare tunnel and DNS; the dialog says so. Fix: a new Cloudflare step in the delete. If never: you remove them by hand, guided by the dialog. **Pick: close.** -- **R-768** - Grimoire is not offered (upstream rules out public use); the row only watches upstream. Fix: a re-read each catalog campaign. If never: nothing; Karakeep covers bookmarks. **Pick: close.** -- **R-793** - Four apps contain paid-edition code that is off; the row says never turn it on. Fix: nothing (the rule lives in the licence audit). If never: nothing unless someone enables it. **Pick: close.** -- **R-796** - MeTube's browser 'send' helpers cannot pass the family gate. Fix: a new token design. If never: households paste links in the page. **Pick: close.** -- **R-797** - Catalog CI cannot run one gate rule; it says 'not checked' and the push hook checks it. Fix: a CI checkout of a second repo. If never: only a forbidden bypass could skip it. **Pick: close.** -- **R-804** - plant-it's image no longer exists; the template is already marked not installable. Fix: hide the template (small). If never: nothing. **Pick: close.** -- **R-190** - A storage permission vanished once in August; the agent now restores it and mails you. Fix: hours of live experiments. If never: one mail per recurrence; it self-repairs. **Pick: close.** -- **R-412** - A rare race can push one hollow off-site copy of one app; the next night repairs it. Fix: a change at the backup/push boundary plus a drill. If never: rarely, one app's off-site copy is hollow for a day; the second-drive copy stays good. **Pick: close.** -- **R-458** - A pinned (frozen) app can get a newer health check, which can only cause a false alarm. Fix: a fiddly change on the sync path. If never: maybe one false alarm one day. **Pick: close.** -- **R-584** - Helper scripts with the shared demo password were left in a demo box's /tmp. Fix: a wrapper tool. If never: demo-password litter on throwaway boxes. **Pick: close.** -- **R-586** - The ISO bootstrap harness runs at each ISO release, not on every push. Fix: Docker-capable CI. If never: a break is caught at the next ISO release. **Pick: close.** -- **R-698** - A backup records an app's image name, not the image; restoring a version deleted upstream fails. Fix: a mirror or much more backup space. If never: that restore fails; the household uses another copy or version. **Pick: close.** -- **R-738** - The update's health check sees only the front page, so data broken behind it passes. Fix: a per-app data check in every template. If never: catalog tests keep catching it before release. **Pick: close.** -- **R-778** - A box that rolls back below controller 0.286 trusts a forged address in the login counter. Fix: a patch release or a rollback rule. If never: a short window until the next update. **Pick: close.** -- **R-783** - Three wrong SparkyFitness logins block sign-in for everyone for ~10 s. Fix: an upstream setting. If never: a persistent stranger can annoy the household. **Pick: close.** -- **R-853** - After a boot, the box's versions reach the hub up to ~15 min late. Fix: a new agent mode and release. If never: one report late; nothing lost. **Pick: close.** -- **R-700** - A drive move keeps the app's records: fixed in controller v0.276.0 with tests; never seen live on a two-drive box. Fix: a live two-drive move. If never: the tests stay the proof. **Pick: close.** -- **R-704** - A fresh install drops an old update hold: fixed in v0.278.0 with tests; not seen live. Fix: a live install with a leftover hold. If never: the tests stay the proof. **Pick: close.** -- **R-706** - Removing an app with its backups also deletes its off-site test copy: fixed in v0.279.0 with tests; not seen live. Fix: a live remove with an off-site copy. If never: the tests stay the proof. **Pick: close.** -- **R-723** - No 'box recovered' alarm in a new box's first hour: fixed in hub v0.126.0 with tests; not seen at a real first install. Fix: watch the next real install. If never: the tests stay the proof. **Pick: close.** - -**Not 'won't do' - these need YOU (leaked tokens; my pick is KEEP until you rotate):** -- **R-831** - the Hetzner storage API token was printed into one session transcript. Rotate it (~10 min). If never: whoever gets that transcript can manage Storage Box sub-accounts. -- **R-870** - Tester 1's two Cloudflare tokens were printed into one transcript. Rotate them (~15 min), or close when Tester 1 is retired. If never: someone could change that test zone's DNS. - - ## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box **Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 9bb91607..485ba3b6 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,58 @@ --- +## 2026-10-05 (late night) — burn-down round 2: the operator's answer (43 accepted), small rows fixed with releases + +The full text of every row below: `git show e8c56c44:documentation/backlog/OPEN-ITEMS.md` (the commit before each closing commit; the burn-down evidence table is `audits/burndown-2026-10-05/`). + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-93** | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's p (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A choice between two test fixtures that no longer exist. Fix would cost: a new fixture (a design job). If never: nothing breaks. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-161** | **The volume-persistence gate is enforced by CONVENTION, not automatically.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The check that app data lands where backups see it runs by hand, not on every push. Fix would cost: CI starting ~53 apps per push. If never: a bad template can ship until the next periodic check. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-162** | **`docker diff` is the gate's only witness, and its failure mode is quiet.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: If Docker ran on an unusual storage driver, one catalog check would blame the wrong thing. Fix would cost: ~1 h for a driver nobody runs. If never: nothing; the check still fails safely. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-169** | **CI can only report, because there is no gate in the road.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: CI reports after a push lands instead of blocking it (no pull requests here). Fix would cost: a pull-request workflow for every change. If never: a forbidden --no-verify push could land broken code until the CI mail. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A removed storage permission can still look present for minutes (Proxmox cache). Fix would cost: a new probe in the agent. If never: the self-repair notices up to ~16 min late, then repairs. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-210** | **Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by `docker pull`** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: 193 old controller/hub images sit only on DooPlex (~27 GB). Fix would cost: a careful delete on DooPlex. If never: 27 GB stays used; ~199 GB is free. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-284** | **„A kiválasztott tárhely majdnem megtelt." on a store that is 93 % FREE — an apparent inverted threshold.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: An 'almost full' warning reported on an empty disk was a misreading; the page never shows it there. Fix would cost: nothing to fix. If never: nothing. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-346** | **`ActiveEnterTimestamp` answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A warning about a mistake nobody has made (the audit found 0 cases). Fix would cost: nothing left. If never: nothing today. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-367** | **The database dumps already written under the wrong name are stranded, and nothing will ever collect them.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: One old 312 KB database dump sits on the demo-hp test box under an old folder name. Fix would cost: a hand delete on a demo box. If never: a small file stays on a test box. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-371** | **The off-site tier is the only backup tier that announces nothing on success.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The weekly off-site backup sends no 'done' message (failures and staleness already alarm). Fix would cost: a new event in two repos. If never: success stays silent, as now. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-372** | **A Tier-2 copy that has NEVER been produced because its source path is missing is not surfaced prominently to the operator.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: An idea to show 'second copy never made' apart from 'second copy failed' to the operator. Fix would cost: a design and a new state end to end. If never: the existing loud warning stays. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-374** | **Three C1 refusal cases were judged borderline, left unfiled, and never named — so nobody can re-open the judgement.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A July audit says three borderline cases were left out but never named them. Fix would cost: hours to guess which three. If never: nobody can re-judge them; later sweeps exist. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-393** | **A decision-log skill for unattended runs was considered and deliberately deferred.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A proposed tool to log every small decision an unattended run makes. Fix would cost: a small project. If never: the bigger decisions are already recorded under the rules file. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-420** | **`controller_gates.py` could not express a NON-BLOCKING gate before 2026-09-01** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The felhom.eu gate runner cannot mark a gate as advisory only. Fix would cost: add it when one is needed. If never: nothing today. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-424** | **`one_register_gate.py`: a real defect parked under the roadmap state `idea` is invisible to it.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The roadmap gate cannot tell a real defect filed as an idea from an idea. Fix would cost: no mechanical fix exists. If never: a person must read the roadmap. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The hub's memory suggestion for an app can use samples from a short test install. Fix would cost: a hub change (~1-2 h). If never: a misleading suggestion for up to 7 days after a test install. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: BookStack's uploaded files cannot be checked automatically after an upgrade. Fix would cost: an upstream change or a browser step. If never: upgrades stay half-checked automatically. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-503** | **[P3-LOW] SPIKE (not built): an install-time disk rule for the public ISO — "exactly one internal disk → install; otherwise stop in Hungarian".** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Letting the installer pick the disk when there is only one (you ruled no). Fix would cost: reversing your rulings. If never: a person keeps choosing the disk. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A catalog 'locked after install' flag does nothing visible (all settings are read-only anyway). Fix would cost: delete it everywhere, or build an edit page. If never: nothing a household meets. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Vaultwarden shows a sign-up form although sign-up is off; the server refuses it. Fix would cost: an upstream fix. If never: a stranger sees a form that fails. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-551** | **[P3-LOW] No Tier-0 box can put the escrow ceremony in the state R-546 fixes — paused AND connected to its agent — so the readiness branches are proven only by tests.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The escrow 'waiting for the agent' screens are tested but never seen on a real box. Fix would cost: a token setup on the scratch box. If never: small chance the live page differs; the state lasts ~17 min. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-610** | **[P2-MEDIUM] The power cut AFTER the new version has started was never measured — R-520 closed on the safe half only.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A power cut inside a sub-second startup phase was never measured (same code as the measured phase). Fix would cost: a new fault injector and a drill. If never: nothing new would be learned. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Old opengist /login bookmarks answer 404 after an upstream move. Fix would cost: one help-text line. If never: a household re-bookmarks once. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Three live proofs a scratch box cannot give, and one log line 20 minutes off. Fix would cost: special test venues. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Deleting a customer does not remove their Cloudflare tunnel and DNS; the dialog says so. Fix would cost: a new Cloudflare step in the delete. If never: you remove them by hand, guided by the dialog. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-768** | **[P3-LOW] Grimoire is not built: upstream rules out public exposure and publishes no image for its current line.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Grimoire is not offered (upstream rules out public use); the row only watches upstream. Fix would cost: a re-read each catalog campaign. If never: nothing; Karakeep covers bookmarks. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-793** | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Four apps contain paid-edition code that is off; the row says never turn it on. Fix would cost: nothing (the rule lives in the licence audit). If never: nothing unless someone enables it. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-796** | **[P3-LOW] MeTube's "send to MeTube" helpers (browser extensions, bookmarklets, phone apps) cannot work behind the family gate.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: MeTube's browser 'send' helpers cannot pass the family gate. Fix would cost: a new token design. If never: households paste links in the page. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-797** | **[P3-LOW] `check-family-gate.py` rule 3 (a family_gate template needs a baked golden ≥ 0.287.0) is checked only where the felhom.eu sibling exists — CI's single clone cannot.** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Catalog CI cannot run one gate rule; it says 'not checked' and the push hook checks it. Fix would cost: a CI checkout of a second repo. If never: only a forbidden bypass could skip it. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-804** | **[P3-LOW] plant-it's image is gone from Docker Hub: `msdeluise/plant-it:0.10.0` answers "pull access denied … repository does not exist".** (P4) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: plant-it's image no longer exists; the template is already marked not installable. Fix would cost: hide the template (small). If never: nothing. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-190** | **A storage ACL that demonstrably WORKED in the morning was gone by mid-morning, and nothing recorded its removal.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A storage permission vanished once in August; the agent now restores it and mails you. Fix would cost: hours of live experiments. If never: one mail per recurrence; it self-repairs. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-412** | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A rare race can push one hollow off-site copy of one app; the next night repairs it. Fix would cost: a change at the backup/push boundary plus a drill. If never: rarely, one app's off-site copy is hollow for a day; the second-drive copy stays good. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A pinned (frozen) app can get a newer health check, which can only cause a false alarm. Fix would cost: a fiddly change on the sync path. If never: maybe one false alarm one day. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-584** | **[P2-MED] Credential-bearing probe scripts were left in a live guest's `/tmp` for hours, across three releases.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Helper scripts with the shared demo password were left in a demo box's /tmp. Fix would cost: a wrapper tool. If never: demo-password litter on throwaway boxes. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-586** | **[P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The ISO bootstrap harness runs at each ISO release, not on every push. Fix would cost: Docker-capable CI. If never: a break is caught at the next ISO release. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-738** | **[P2-MEDIUM] Every wger update that brings database migrations leaves wger broken, and the guarded Update reports `done`.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: The update's health check sees only the front page, so data broken behind it passes. Fix would cost: a per-app data check in every template. If never: catalog tests keep catching it before release. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-778** | **[P3-LOW] A box that rolls back to a controller ≤ 0.285 after v0.286 keeps the new traefik (an old controller never rewrites a running traefik) — and the old `clientIP` believes the LEFTMOST X-Forwarded-For, which a stranger then writes: the dashboard's login counter becomes dodgeable until the box moves forward again.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A box that rolls back below controller 0.286 trusts a forged address in the login counter. Fix would cost: a patch release or a rollback rule. If never: a short window until the next update. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-783** | **[P3-LOW] SparkyFitness: three wrong sign-ins by anyone shut EVERY visitor out of sign-in for ~10 s — a stranger retrying every 10 s keeps the household out.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Three wrong SparkyFitness logins block sign-in for everyone for ~10 s. Fix would cost: an upstream setting. If never: a persistent stranger can annoy the household. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-853** | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: After a boot, the box's versions reach the hub up to ~15 min late. Fix would cost: a new agent mode and release. If never: one report late; nothing lost. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A drive move keeps the app's records: fixed in controller v0.276.0 with tests; never seen live on a two-drive box. Fix would cost: a live two-drive move. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A fresh install drops an old update hold: fixed in v0.278.0 with tests; not seen live. Fix would cost: a live install with a leftover hold. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-706** | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Removing an app with its backups also deletes its off-site test copy: fixed in v0.279.0 with tests; not seen live. Fix would cost: a live remove with an off-site copy. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: No 'box recovered' alarm in a new box's first hour: fixed in hub v0.126.0 with tests; not seen at a real first install. Fix would cost: watch the next real install. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | + +--- + ## 2026-10-05 (night) — the burn-down: stale rows closed with evidence, small rows fixed (rule: fix small, do not file) The full text of every row below: `git show ab2b3049:documentation/backlog/OPEN-ITEMS.md` (the commit before each closing commit; the burn-down evidence table is `audits/burndown-2026-10-05/`). diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 2a3dea9e..15587392 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -114,7 +114,7 @@ match what the reader sees is how an instrument stops being believed (R-421). It stopping line that lies. -## Install & onboarding — 15 rows (P3 9, P4 6) +## Install & onboarding — 14 rows (P3 9, P4 5) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -129,12 +129,11 @@ stopping line that lies. | **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-310** | Install & onboarding | P4 | **Two small edges on the installer, neither costing more than a moment.** (1) The R-297 operator-named refusal states the vouched version twice in consecutive sentences (*"…but the vouched golden is 0.213.0. The vouched golden is 0.213.0."*). (2) `--uninstall` reads its typed vmid confirmation from `/dev/tty` and `--force` deliberately does **not** bypass it, so teardown cannot be scripted without a pty — correct for an irreversible destroy, but undocumented; it surfaces as `line 891: /dev/tty: No such device or address` and an rc=1 that looks like a failure rather than a refusal to proceed unattended | **READY (S) — NEW 2026-08-12, RANK 4** | R-297 | Drop the duplicated sentence; add one runbook line naming the pty requirement | CC | | **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC | -| **R-503** | Install & onboarding | P4 | **[P3-LOW] SPIKE (not built): an install-time disk rule for the public ISO — "exactly one internal disk → install; otherwise stop in Hungarian".** Offered to the operator 2026-09-14 and **not chosen**: the 2026-07-31 ruling (a person chooses the disk) stands. Recorded so a reversal starts from measurements, not from the offer. **What must be measured first:** (1) whether the Proxmox auto-installer's HTTP answer mode can serve a per-machine answer from posted system info without network being a precondition a volunteer can miss; (2) whether USB transport is reliably visible in sysfs (`/sys/block/*/device` path, `removable`) where udev properties were measured blind (SPIKE-universal-iso-1 §3.2); (3) whether any refusal can be shown in Hungarian without modifying the Proxmox installer squashfs. **Reverses two rulings if built — needs an operator word.** | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (ruling), CC (spike)** **Re-ranked 2026-10-03: P3→P4: an unchosen idea that would reverse two rulings.** | — | — | operator | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | | **R-725** | Install & onboarding | P4 | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). **FIXED 2026-09-30:** the recovery wizard speaks „te" (controller v0.283.0; formal ceiling 18 → 17); the bind page says „a Felhom üzemeltetőjétől kaptál" (hub v0.126.0). **NARROWED — remaining:** the console's stray „V" (the installer/agent's banner, not these repos' text); the gate's English JSON to a phone app (the app shows its own error; left, deliberately); and the expired bind page still says „kérj újat az ügyfélszolgálattól" ABOVE the new „Új linket kérek" button (hub copy, next hub release). | **NARROWED — three small copy items; owner: CC** **Re-ranked 2026-10-03: P3→P4: three small copy slips left; nothing blocks the household.** | — | — | CC | | **R-881** | Install & onboarding | P4 | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (`scripts/felhom-host-install.sh` ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. | READY | — | Add the file to the uninstall list and fix the comment at the next installer tag | CC | -## Apps & catalog — 33 rows (P3 13, P4 20) +## Apps & catalog — 25 rows (P3 12, P4 13) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -146,53 +145,40 @@ stopping line that lies. | **R-613** | Apps & catalog | P3 | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. **— UPDATE NIGHT 2026-09-21:** the update night could not seed `uptime-kuma` for the same reason and left it out rather than faking it. **— FIXED 2026-09-23 night for uptime-kuma (catalog `a5a729a`):** `UPTIME_KUMA_DB_TYPE=sqlite` — the database choice is made, the wizard never appears, the real server starts (`/api/entry-page` → `entryPage`, `/metrics` 401), the database lands in `/app/data` (the backed-up volume). Red-proofed on 9202 through the product: before, the box read `running` over `setup-database`; after, the real server. (A first "after" run used a stale template — R-607's lag — and is kept.) **Still open:** the sweep for other templates whose probe passes on a setup wizard. | **READY — P3, narrowed to the sweep; owner: CC (catalog)** **Re-ranked 2026-10-03: P2->P3: uptime-kuma fixed; only a sweep of other templates remains.** | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-682** | Apps & catalog | P3 | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** | — | — | CC | -| **R-704** | Apps & catalog | P3 | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. **-- 2026-09-28 later: SECOND and worse instance, then FIXED in controller v0.278.0.** demo-hp's fresh nextcloud (installed 10:13) carried an UPDATE hold from a nextcloud of 2026-09-13 (set before v0.242.0 made removals clear update holds; nothing ever swept it). At the manual off-site run (15:17) the backup leg logged `Skipping volume dump for nextcloud — the app is HELD stopped`, captured no unit, and pushed a snapshot that `carried NO database dump and NO volume tar` — a freshly installed app silently NOT backed up. **Fix:** a removal also clears the crash-loop stop (`settings.ClearUpdateHold`), and a new install (plain or "use my kept data") drops a leftover update/crash-loop hold of an app that is not installed (`Router.dropLeftoverHold`); restore holds (R-379) untouched. Red-proofed RP4–RP6 (`audits/kept-offsite-2026-09-28/redproofs/`). Floor 0.278.0. **STILL OPEN: live proof of the install-time drop** (a box with a leftover hold on an uninstalled app). | **WATCHING — P2; owner: CC (install-time drop, live)** | — | — | CC | | **R-757** | Apps & catalog | P3 | **[P3-LOW] A template that gains a generated `secret` field makes the box INVENT that value for apps already installed — for calibre-web a login name the app never got.** MEASURED 2026-10-01 on demo-hp: 9 minutes after catalog `e9f50b5` (decision 61) synced, the controller logged `InjectMissingFields … injected missing fields: ADMIN_USER` (deploy.go:1337) and the app page's reveal returned a 10-character name that was NOT calibre-web's login (the app had kept its own name; `after_install` runs only after a fresh install). The page did not list the field, but the reveal answers it, and the password field's text now says "the user name above". demo-hp was fixed by renaming the app's user to the box's recorded name (credentials file updated). Any other installed calibre-web gets the same made-up name at its next sync while its login stays `admin` (no other box has one today: N100 none, Tester-2 not registered). **Needs:** InjectMissingFields must not invent a value an app has to have been GIVEN (a field consumed only by `after_install`), or such a field needs an "installed apps: ask" path. `audits/calibre-name-and-prune-2026-10-01/A/A2…, A3…` | **OPEN — rank P3-LOW; owner: CC (controller)** | — | — | CC | | **R-758** | Apps & catalog | P3 | **[P3-LOW] Eight templates declare a `mem_limit` that is not the sum of their compose limits, and no gate checks it.** FOUND 2026-10-01 by `scripts/onboarding_gaps.py` (the new-app checklist's gap page, row 5.3): adventurelog 384M vs 896M, bookstack 512M vs 768M, calcom 768M vs 1792M, claper 384M vs 640M, kimai 384M vs 640M, nextcloud 1024M vs 1664M, outline 768M vs 1152M, zipline 512M vs 768M — every one UNDER the sum, so the deploy screen and the box's capacity figure (`09` decision 22) read less memory than the app may take. REUSE.md §2 says `mem_limit` = the sum. The compose limits are what Docker enforces, so no app is starved by this; the figure the household and the capacity check read is wrong. **Needs:** the eight figures corrected (a template change: needs its own session, no image move), and a static gate (`--fast`) with a decoy, so a new app cannot repeat it. `app-catalog-felhom.eu/onboarding/EXISTING-APPS-GAPS.md` | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** | — | — | CC | | **R-774** | Apps & catalog | P3 | **[P3-LOW] Two things the new apps' pages do not show yet: Karakeep's mail-ON path is unproven, and its official phone app reports crashes to its makers.** READ/MEASURED 2026-10-01: Karakeep's `smtp_mapping` (plaintext :2526, as Cal.com) was proven only with mail OFF (a fresh install boots) — 9202 has no hub, so the relay cannot be exercised there; the official mobile app ships Sentry crash reporting with a hard-coded DSN (`apps/mobile/app/_layout.tsx`, FIT.md). **Needs:** one password-reset mail from Karakeep on a hub-enabled box (demo-hp 9201, a throwaway install); and a sentence on the page about the phone app (copy, freeze). `audits/new-apps-2026-10-01/bench/karakeep-mail-off-boot.txt` | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | -| **R-284** | Apps & catalog | P4 | **„A kiválasztott tárhely majdnem megtelt." on a store that is 93 % FREE — an apparent inverted threshold.** Calibre-Web's deploy page rendered `