R-547 (P3): a disk that fills and empties between sweeps is never mentioned to anyone. The guest's root filesystem sat at 96% for ten minutes and no alarm of any kind fired - checked twice, once by the round's runner and once independently after the fill was released. disk_critical is defined at >=95% used, but the fill-watch is a DAILY sweep plus one check ~90s after a controller start, so a ten-minute window contains no check. The timing was almost comic: the controller restarted at 21:28 after the previous round's power cut, so its single opportunistic check ran about twenty seconds before the disk filled. This is the ladder working as designed, not a missed alarm - it is filed because the honest answer to "would the household be told?" is no, and that is written down nowhere. R-548 (P3): the whole-guest backup's LOCAL tier cannot fit on a small-system-disk box and retries on that tier for ever. A ~29GB source into a 14GB pve-root, measured falling at ~16MB/s - under four minutes to a full / on the nested PVE. The product's behaviour is correct throughout: it failed the tier, named it, scheduled a retry, its status surface agreed, and the off-site tier then succeeded from the same snapshot in ~8.5 minutes taking no local disk. What is filed is the loop: on a box this shape the local tier can never succeed. Honest caveat recorded in the row - the 32GB system disk is this drill's own fixture choice - but nothing checks whether the local target could hold the source before starting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
documentation/backlog/
OPEN-ITEMS.md is the register of open work and the file to read first — it holds only what is
open, one row per item, every row with a state and an owner. ROADMAP.md is the full history and
reasoning behind the R-n IDs, including shipped and killed items; an ID is minted there, and a new
instance of an existing item attaches to that ID rather than getting its own.
The rest of this folder: verified-LIVE findings with implementable fix plans that are not yet
implemented. Preserved here
(instead of on git branches) per the trunk-based, no-branches rule — the fix itself is implemented later
directly on main, during a normal/supervised session.
-
FIX-M18-NOTES.md — dump re-validation runs every 5 min (perf). FIXED in controller v0.62.0 @
f8afe5c(2026-06-14). (was on the deletedfelhom-controllerbranchfix/m18-dump-validation-cache.) -
FIX-M19-NOTES.md —
deriveStackNamemisattribution edge (low-incidence correctness). FIXED in controller v0.62.0 @6bab68b(2026-06-14). (was on the deleted branchfix/m19-stackname-crossref.) -
FOLLOWUP-golden-default-controller-tag.md — the golden bakes a stale controller (
:0.43.0when queued; had rotted again to:0.85.1by resolution). FIXED in felhom-agent @ceca355(2026-07-03):build-golden.shv2.0.0 makes the controller tag a MANDATORY argument (a required arg cannot rot) and golden 0.98.3 was baked + clean-room-validated (bake → first-boot-current → self-manage → app deploy, on the drill VM — no supervised touch of live guests needed) + published + vouched. Evidence:../audits/DRILL-golden-098-2026-07-03.md.
Related: the live-drive fixspec (../audits/live-drive-fixspec-2026-06-14.md) carries the deferred
supervised items F9 (HDD provisioning/guest-attach), F20-BUG2 (durable_id scheme), F20-BUG3 (async
mkfs) — to be implemented in the agent/golden supervised session.