operator rulings 2026-10-05 18:23: 43 burn-down rows closed as accepted (292 -> 249); R-124 kept (fix now), R-698 kept (operator's known risk); R-831/R-870 not rotated; R-887 one runner online — the stale-registration guess was wrong
gates / gates (push) Successful in 1m36s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 18:39:30 +02:00
parent e8c56c440a
commit 301fe4532a
3 changed files with 80 additions and 126 deletions
+12 -67
View File
@@ -3,8 +3,16 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
sent to it.**
**Updated 2026-10-05 (night, the burn-down session): every box of ours healthy (nothing was changed on any box). The
open-items list went from 336 to 292. Report: `REPORT-burndown-2026-10-05.md`.**
**Updated 2026-10-05 (late night, burn-down round 2 — in progress): 249 open rows after your answer (was 292).**
## Your answer to the burn-down list (2026-10-05 18:23), recorded
- **43 rows closed as accepted by you**, each with its reason from the list.
- **R-124 stays open and is fixed in this session** (a recovery step that fails during a real recovery).
- **R-698 stays open as a known risk** (a backup keeps an app's image name, not the image). It is yours; nobody works on it now.
- **R-831 and R-870** (the printed tokens) stay as they are, by your earlier rulings; the rows keep their steps.
- **R-887** (the CI jobs that never ran): your screenshot shows ONE runner, online. My „old copy of the runner" guess was
wrong; the row says so.
## Tonight, later (2026-10-05, night): the list got shorter
@@ -19,76 +27,13 @@ open-items list went from 336 to 292. Report: `REPORT-burndown-2026-10-05.md`.**
it; it is gone, and a test keeps it gone.
- **You got 5 „CI failed" mails tonight — my fault, fixed.** The new test gate ran one of today's test suites on the CI
machine for the first time; it uses a different `date` tool, and 9 tests failed. Fixed; CI is green again (job 1360).
- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes, with no log and no mail. One re-run passed; the other was again never picked up. It looks like a leftover runner registration, but I cannot see the runner list. If you do nothing: now and then a CI run reads „failed" without testing anything. Fix: Site Administration → Runners, remove any offline duplicate of `felhom-gates-runner`.
- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes,
with no log and no mail. (The „old runner copy" guess was wrong — see above.)
- **A new rule, so the list stops growing:** a small problem found during work (about 30 minutes) is fixed in that
session and never added to the list. Every report now states four numbers: rows before, after, opened, closed.
**The numbers:** 336 before → **292 after**; 1 opened (the CI fault below); 45 closed.
**Needs you:**
1. **The list below: 45 rows I would close as „accepted".** Answer per row, or „all as picked". If you do nothing,
they stay open and the list stays at 292.
2. **Two leaked tokens to rotate** (R-831, R-870, at the end of the list). If you do nothing, whoever has those
transcripts keeps that access.
3. Earlier tonight's items are unchanged (pin the other „latest" apps on DooPlex; Alertmanager's file permissions).
4. **Look at the CI runners** (Gitea → Site Administration → Runners; R-887): remove any offline duplicate of
`felhom-gates-runner`. If you do nothing, some CI runs fail without running.
## Your list: rows I would close as 'accepted' (burn-down 2026-10-05) - answer per row, or 'all as picked'
Each line: what it is. What fixing costs. What happens if we never fix it. **My pick.** I closed none of these.
- **R-93** - A choice between two test fixtures that no longer exist. Fix: a new fixture (a design job). If never: nothing breaks. **Pick: close.**
- **R-124** - The disaster-recovery recipe writes the backup-server namespace as the word 'root'; the server wants an empty name, so one pasted command fails. Fix: a change across hub and agent. If never: an operator doing a manual restore drops one option once; no data risk. **Pick: close.**
- **R-161** - The check that app data lands where backups see it runs by hand, not on every push. Fix: CI starting ~53 apps per push. If never: a bad template can ship until the next periodic check. **Pick: close.**
- **R-162** - If Docker ran on an unusual storage driver, one catalog check would blame the wrong thing. Fix: ~1 h for a driver nobody runs. If never: nothing; the check still fails safely. **Pick: close.**
- **R-169** - CI reports after a push lands instead of blocking it (no pull requests here). Fix: a pull-request workflow for every change. If never: a forbidden --no-verify push could land broken code until the CI mail. **Pick: close.**
- **R-194** - A removed storage permission can still look present for minutes (Proxmox cache). Fix: a new probe in the agent. If never: the self-repair notices up to ~16 min late, then repairs. **Pick: close.**
- **R-210** - 193 old controller/hub images sit only on DooPlex (~27 GB). Fix: a careful delete on DooPlex. If never: 27 GB stays used; ~199 GB is free. **Pick: close.**
- **R-284** - An 'almost full' warning reported on an empty disk was a misreading; the page never shows it there. Fix: nothing to fix. If never: nothing. **Pick: close.**
- **R-346** - A warning about a mistake nobody has made (the audit found 0 cases). Fix: nothing left. If never: nothing today. **Pick: close.**
- **R-367** - One old 312 KB database dump sits on the demo-hp test box under an old folder name. Fix: a hand delete on a demo box. If never: a small file stays on a test box. **Pick: close.**
- **R-371** - The weekly off-site backup sends no 'done' message (failures and staleness already alarm). Fix: a new event in two repos. If never: success stays silent, as now. **Pick: close.**
- **R-372** - An idea to show 'second copy never made' apart from 'second copy failed' to the operator. Fix: a design and a new state end to end. If never: the existing loud warning stays. **Pick: close.**
- **R-374** - A July audit says three borderline cases were left out but never named them. Fix: hours to guess which three. If never: nobody can re-judge them; later sweeps exist. **Pick: close.**
- **R-393** - A proposed tool to log every small decision an unattended run makes. Fix: a small project. If never: the bigger decisions are already recorded under the rules file. **Pick: close.**
- **R-420** - The felhom.eu gate runner cannot mark a gate as advisory only. Fix: add it when one is needed. If never: nothing today. **Pick: close.**
- **R-424** - The roadmap gate cannot tell a real defect filed as an idea from an idea. Fix: no mechanical fix exists. If never: a person must read the roadmap. **Pick: close.**
- **R-445** - The hub's memory suggestion for an app can use samples from a short test install. Fix: a hub change (~1-2 h). If never: a misleading suggestion for up to 7 days after a test install. **Pick: close.**
- **R-460** - BookStack's uploaded files cannot be checked automatically after an upgrade. Fix: an upstream change or a browser step. If never: upgrades stay half-checked automatically. **Pick: close.**
- **R-503** - Letting the installer pick the disk when there is only one (you ruled no). Fix: reversing your rulings. If never: a person keeps choosing the disk. **Pick: close.**
- **R-527** - A catalog 'locked after install' flag does nothing visible (all settings are read-only anyway). Fix: delete it everywhere, or build an edit page. If never: nothing a household meets. **Pick: close.**
- **R-532** - Vaultwarden shows a sign-up form although sign-up is off; the server refuses it. Fix: an upstream fix. If never: a stranger sees a form that fails. **Pick: close.**
- **R-551** - The escrow 'waiting for the agent' screens are tested but never seen on a real box. Fix: a token setup on the scratch box. If never: small chance the live page differs; the state lasts ~17 min. **Pick: close.**
- **R-610** - A power cut inside a sub-second startup phase was never measured (same code as the measured phase). Fix: a new fault injector and a drill. If never: nothing new would be learned. **Pick: close.**
- **R-654** - Old opengist /login bookmarks answer 404 after an upstream move. Fix: one help-text line. If never: a household re-bookmarks once. **Pick: close.**
- **R-687** - Three live proofs a scratch box cannot give, and one log line 20 minutes off. Fix: special test venues. If never: the tests stay the proof. **Pick: close.**
- **R-688** - Deleting a customer does not remove their Cloudflare tunnel and DNS; the dialog says so. Fix: a new Cloudflare step in the delete. If never: you remove them by hand, guided by the dialog. **Pick: close.**
- **R-768** - Grimoire is not offered (upstream rules out public use); the row only watches upstream. Fix: a re-read each catalog campaign. If never: nothing; Karakeep covers bookmarks. **Pick: close.**
- **R-793** - Four apps contain paid-edition code that is off; the row says never turn it on. Fix: nothing (the rule lives in the licence audit). If never: nothing unless someone enables it. **Pick: close.**
- **R-796** - MeTube's browser 'send' helpers cannot pass the family gate. Fix: a new token design. If never: households paste links in the page. **Pick: close.**
- **R-797** - Catalog CI cannot run one gate rule; it says 'not checked' and the push hook checks it. Fix: a CI checkout of a second repo. If never: only a forbidden bypass could skip it. **Pick: close.**
- **R-804** - plant-it's image no longer exists; the template is already marked not installable. Fix: hide the template (small). If never: nothing. **Pick: close.**
- **R-190** - A storage permission vanished once in August; the agent now restores it and mails you. Fix: hours of live experiments. If never: one mail per recurrence; it self-repairs. **Pick: close.**
- **R-412** - A rare race can push one hollow off-site copy of one app; the next night repairs it. Fix: a change at the backup/push boundary plus a drill. If never: rarely, one app's off-site copy is hollow for a day; the second-drive copy stays good. **Pick: close.**
- **R-458** - A pinned (frozen) app can get a newer health check, which can only cause a false alarm. Fix: a fiddly change on the sync path. If never: maybe one false alarm one day. **Pick: close.**
- **R-584** - Helper scripts with the shared demo password were left in a demo box's /tmp. Fix: a wrapper tool. If never: demo-password litter on throwaway boxes. **Pick: close.**
- **R-586** - The ISO bootstrap harness runs at each ISO release, not on every push. Fix: Docker-capable CI. If never: a break is caught at the next ISO release. **Pick: close.**
- **R-698** - A backup records an app's image name, not the image; restoring a version deleted upstream fails. Fix: a mirror or much more backup space. If never: that restore fails; the household uses another copy or version. **Pick: close.**
- **R-738** - The update's health check sees only the front page, so data broken behind it passes. Fix: a per-app data check in every template. If never: catalog tests keep catching it before release. **Pick: close.**
- **R-778** - A box that rolls back below controller 0.286 trusts a forged address in the login counter. Fix: a patch release or a rollback rule. If never: a short window until the next update. **Pick: close.**
- **R-783** - Three wrong SparkyFitness logins block sign-in for everyone for ~10 s. Fix: an upstream setting. If never: a persistent stranger can annoy the household. **Pick: close.**
- **R-853** - After a boot, the box's versions reach the hub up to ~15 min late. Fix: a new agent mode and release. If never: one report late; nothing lost. **Pick: close.**
- **R-700** - A drive move keeps the app's records: fixed in controller v0.276.0 with tests; never seen live on a two-drive box. Fix: a live two-drive move. If never: the tests stay the proof. **Pick: close.**
- **R-704** - A fresh install drops an old update hold: fixed in v0.278.0 with tests; not seen live. Fix: a live install with a leftover hold. If never: the tests stay the proof. **Pick: close.**
- **R-706** - Removing an app with its backups also deletes its off-site test copy: fixed in v0.279.0 with tests; not seen live. Fix: a live remove with an off-site copy. If never: the tests stay the proof. **Pick: close.**
- **R-723** - No 'box recovered' alarm in a new box's first hour: fixed in hub v0.126.0 with tests; not seen at a real first install. Fix: watch the next real install. If never: the tests stay the proof. **Pick: close.**
**Not 'won't do' - these need YOU (leaked tokens; my pick is KEEP until you rotate):**
- **R-831** - the Hetzner storage API token was printed into one session transcript. Rotate it (~10 min). If never: whoever gets that transcript can manage Storage Box sub-accounts.
- **R-870** - Tester 1's two Cloudflare tokens were printed into one transcript. Rotate them (~15 min), or close when Tester 1 is retired. If never: someone could change that test zone's DNS.
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box
**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three