burn-down: R-423 site page walk (exemption dropped); STATUS Part C list (45 rows for the operator); REPORT; register 336 -> 291 (0 opened, 45 closed)
gates / gates (push) Failing after 13m1s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 17:15:02 +02:00
parent b26d292a64
commit 8012ce4640
10 changed files with 221 additions and 8 deletions
+80 -3
View File
@@ -3,9 +3,86 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
sent to it.**
**Updated 2026-10-05 (evening, the hub-database session): every box of ours healthy. The hub database now leaves
DooPlex every night, locked, to ep0, is test-restored every Sunday, and an alarm mails you if either stops. Report:
`REPORT-hub-db-offsite-2026-10-05.md`.**
**Updated 2026-10-05 (night, the burn-down session): every box of ours healthy (nothing was changed on any box). The
open-items list went from 336 to 291. Report: `REPORT-burndown-2026-10-05.md`.**
## Tonight, later (2026-10-05, night): the list got shorter
**Decisions:** none of mine.
**What happened:**
- **Every lower-priority row was checked against today's code** (317 rows). 24 described problems a later change had
already fixed; 2 were duplicates. Those are closed, each with the change that fixed it.
- **19 small rows were fixed and closed** in four repositories — wrong comments and documents, missing tests, gates that
checked less than they claimed. No release was needed: nothing that runs on a box or on the hub changed.
- **One real fix found on the way:** the hub's build script pushed a `latest` image tag on every release. Nothing used
it; it is gone, and a test keeps it gone.
- **A new rule, so the list stops growing:** a small problem found during work (about 30 minutes) is fixed in that
session and never added to the list. Every report now states four numbers: rows before, after, opened, closed.
**The numbers:** 336 before → **291 after**; 0 opened; 45 closed.
**Needs you:**
1. **The list below: 45 rows I would close as „accepted".** Answer per row, or „all as picked". If you do nothing,
they stay open and the list stays at 291.
2. **Two leaked tokens to rotate** (R-831, R-870, at the end of the list). If you do nothing, whoever has those
transcripts keeps that access.
3. Earlier tonight's items are unchanged (pin the other „latest" apps on DooPlex; Alertmanager's file permissions).
## Your list: rows I would close as 'accepted' (burn-down 2026-10-05) - answer per row, or 'all as picked'
Each line: what it is. What fixing costs. What happens if we never fix it. **My pick.** I closed none of these.
- **R-93** - A choice between two test fixtures that no longer exist. Fix: a new fixture (a design job). If never: nothing breaks. **Pick: close.**
- **R-124** - The disaster-recovery recipe writes the backup-server namespace as the word 'root'; the server wants an empty name, so one pasted command fails. Fix: a change across hub and agent. If never: an operator doing a manual restore drops one option once; no data risk. **Pick: close.**
- **R-161** - The check that app data lands where backups see it runs by hand, not on every push. Fix: CI starting ~53 apps per push. If never: a bad template can ship until the next periodic check. **Pick: close.**
- **R-162** - If Docker ran on an unusual storage driver, one catalog check would blame the wrong thing. Fix: ~1 h for a driver nobody runs. If never: nothing; the check still fails safely. **Pick: close.**
- **R-169** - CI reports after a push lands instead of blocking it (no pull requests here). Fix: a pull-request workflow for every change. If never: a forbidden --no-verify push could land broken code until the CI mail. **Pick: close.**
- **R-194** - A removed storage permission can still look present for minutes (Proxmox cache). Fix: a new probe in the agent. If never: the self-repair notices up to ~16 min late, then repairs. **Pick: close.**
- **R-210** - 193 old controller/hub images sit only on DooPlex (~27 GB). Fix: a careful delete on DooPlex. If never: 27 GB stays used; ~199 GB is free. **Pick: close.**
- **R-284** - An 'almost full' warning reported on an empty disk was a misreading; the page never shows it there. Fix: nothing to fix. If never: nothing. **Pick: close.**
- **R-346** - A warning about a mistake nobody has made (the audit found 0 cases). Fix: nothing left. If never: nothing today. **Pick: close.**
- **R-367** - One old 312 KB database dump sits on the demo-hp test box under an old folder name. Fix: a hand delete on a demo box. If never: a small file stays on a test box. **Pick: close.**
- **R-371** - The weekly off-site backup sends no 'done' message (failures and staleness already alarm). Fix: a new event in two repos. If never: success stays silent, as now. **Pick: close.**
- **R-372** - An idea to show 'second copy never made' apart from 'second copy failed' to the operator. Fix: a design and a new state end to end. If never: the existing loud warning stays. **Pick: close.**
- **R-374** - A July audit says three borderline cases were left out but never named them. Fix: hours to guess which three. If never: nobody can re-judge them; later sweeps exist. **Pick: close.**
- **R-393** - A proposed tool to log every small decision an unattended run makes. Fix: a small project. If never: the bigger decisions are already recorded under the rules file. **Pick: close.**
- **R-420** - The felhom.eu gate runner cannot mark a gate as advisory only. Fix: add it when one is needed. If never: nothing today. **Pick: close.**
- **R-424** - The roadmap gate cannot tell a real defect filed as an idea from an idea. Fix: no mechanical fix exists. If never: a person must read the roadmap. **Pick: close.**
- **R-445** - The hub's memory suggestion for an app can use samples from a short test install. Fix: a hub change (~1-2 h). If never: a misleading suggestion for up to 7 days after a test install. **Pick: close.**
- **R-460** - BookStack's uploaded files cannot be checked automatically after an upgrade. Fix: an upstream change or a browser step. If never: upgrades stay half-checked automatically. **Pick: close.**
- **R-503** - Letting the installer pick the disk when there is only one (you ruled no). Fix: reversing your rulings. If never: a person keeps choosing the disk. **Pick: close.**
- **R-527** - A catalog 'locked after install' flag does nothing visible (all settings are read-only anyway). Fix: delete it everywhere, or build an edit page. If never: nothing a household meets. **Pick: close.**
- **R-532** - Vaultwarden shows a sign-up form although sign-up is off; the server refuses it. Fix: an upstream fix. If never: a stranger sees a form that fails. **Pick: close.**
- **R-551** - The escrow 'waiting for the agent' screens are tested but never seen on a real box. Fix: a token setup on the scratch box. If never: small chance the live page differs; the state lasts ~17 min. **Pick: close.**
- **R-610** - A power cut inside a sub-second startup phase was never measured (same code as the measured phase). Fix: a new fault injector and a drill. If never: nothing new would be learned. **Pick: close.**
- **R-654** - Old opengist /login bookmarks answer 404 after an upstream move. Fix: one help-text line. If never: a household re-bookmarks once. **Pick: close.**
- **R-687** - Three live proofs a scratch box cannot give, and one log line 20 minutes off. Fix: special test venues. If never: the tests stay the proof. **Pick: close.**
- **R-688** - Deleting a customer does not remove their Cloudflare tunnel and DNS; the dialog says so. Fix: a new Cloudflare step in the delete. If never: you remove them by hand, guided by the dialog. **Pick: close.**
- **R-768** - Grimoire is not offered (upstream rules out public use); the row only watches upstream. Fix: a re-read each catalog campaign. If never: nothing; Karakeep covers bookmarks. **Pick: close.**
- **R-793** - Four apps contain paid-edition code that is off; the row says never turn it on. Fix: nothing (the rule lives in the licence audit). If never: nothing unless someone enables it. **Pick: close.**
- **R-796** - MeTube's browser 'send' helpers cannot pass the family gate. Fix: a new token design. If never: households paste links in the page. **Pick: close.**
- **R-797** - Catalog CI cannot run one gate rule; it says 'not checked' and the push hook checks it. Fix: a CI checkout of a second repo. If never: only a forbidden bypass could skip it. **Pick: close.**
- **R-804** - plant-it's image no longer exists; the template is already marked not installable. Fix: hide the template (small). If never: nothing. **Pick: close.**
- **R-190** - A storage permission vanished once in August; the agent now restores it and mails you. Fix: hours of live experiments. If never: one mail per recurrence; it self-repairs. **Pick: close.**
- **R-412** - A rare race can push one hollow off-site copy of one app; the next night repairs it. Fix: a change at the backup/push boundary plus a drill. If never: rarely, one app's off-site copy is hollow for a day; the second-drive copy stays good. **Pick: close.**
- **R-458** - A pinned (frozen) app can get a newer health check, which can only cause a false alarm. Fix: a fiddly change on the sync path. If never: maybe one false alarm one day. **Pick: close.**
- **R-584** - Helper scripts with the shared demo password were left in a demo box's /tmp. Fix: a wrapper tool. If never: demo-password litter on throwaway boxes. **Pick: close.**
- **R-586** - The ISO bootstrap harness runs at each ISO release, not on every push. Fix: Docker-capable CI. If never: a break is caught at the next ISO release. **Pick: close.**
- **R-698** - A backup records an app's image name, not the image; restoring a version deleted upstream fails. Fix: a mirror or much more backup space. If never: that restore fails; the household uses another copy or version. **Pick: close.**
- **R-738** - The update's health check sees only the front page, so data broken behind it passes. Fix: a per-app data check in every template. If never: catalog tests keep catching it before release. **Pick: close.**
- **R-778** - A box that rolls back below controller 0.286 trusts a forged address in the login counter. Fix: a patch release or a rollback rule. If never: a short window until the next update. **Pick: close.**
- **R-783** - Three wrong SparkyFitness logins block sign-in for everyone for ~10 s. Fix: an upstream setting. If never: a persistent stranger can annoy the household. **Pick: close.**
- **R-853** - After a boot, the box's versions reach the hub up to ~15 min late. Fix: a new agent mode and release. If never: one report late; nothing lost. **Pick: close.**
- **R-700** - A drive move keeps the app's records: fixed in controller v0.276.0 with tests; never seen live on a two-drive box. Fix: a live two-drive move. If never: the tests stay the proof. **Pick: close.**
- **R-704** - A fresh install drops an old update hold: fixed in v0.278.0 with tests; not seen live. Fix: a live install with a leftover hold. If never: the tests stay the proof. **Pick: close.**
- **R-706** - Removing an app with its backups also deletes its off-site test copy: fixed in v0.279.0 with tests; not seen live. Fix: a live remove with an off-site copy. If never: the tests stay the proof. **Pick: close.**
- **R-723** - No 'box recovered' alarm in a new box's first hour: fixed in hub v0.126.0 with tests; not seen at a real first install. Fix: watch the next real install. If never: the tests stay the proof. **Pick: close.**
**Not 'won't do' - these need YOU (leaked tokens; my pick is KEEP until you rotate):**
- **R-831** - the Hetzner storage API token was printed into one session transcript. Rotate it (~10 min). If never: whoever gets that transcript can manage Storage Box sub-accounts.
- **R-870** - Tester 1's two Cloudflare tokens were printed into one transcript. Rotate them (~15 min), or close when Tester 1 is retired. If never: someone could change that test zone's DNS.
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box