ROADMAP: - R-39 gains the full live diagnosis and REFUTES the brief's hypothesis. The generation IS bumped (SetHostDesired bumps unconditionally, 2->3) and applyPBSDR is exonerated, so no hub fix was shipped. The real mechanism is a signal mismatch: the hub's re-consume signal is a generation bump + poke, while the agent re-applies on a change of the DESCRIPTOR CONTENT HASH (manager.go ~L235). An ep0 re-issue re-keys the secret of an EXISTING token, so token_id/fingerprint are unchanged, the descriptor is byte-identical, the hash never moves, and the fresh secret is never consumed -> 401 forever. Proof: consumed-failed.json carries the same hash a4e5424... as the marker written two minutes before the re-issue. Records the second defect found while healing (wrapper reconcile passing --server, fixed in agent v0.90.1), marks the box HEALED with evidence (pvesm active, token 200, a real 9.7 GB encrypted backup listed PBS-side), and leaves the fleet fix explicitly pending its own spec. - R-33 collapses to SHIPPED (scripts v1.21.0), incl. why TimeoutStartSec=infinity is the load-bearing half. - Pre-invite checklist: golden target moves 0.145.x -> 0.146.0 and notes it is now MORE stale, since v0.146.0 is live on the demo box while the golden still bakes 0.143.0. REPORT overwritten with the train: R-39 diagnosis verbatim + heal evidence, the two ISO shas with the byte-identical-payload verification, the nav polish and why the screenshot leg could not be done (the demo controller password is customer-owned since the claim flow, so the build-server credentials are stale), Phase 4 skipped cleanly, and Phase 5 deferred rather than half-run. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
documentation/backlog/
Verified-LIVE findings with implementable fix plans that are not yet implemented. Preserved here
(instead of on git branches) per the trunk-based, no-branches rule — the fix itself is implemented later
directly on main, during a normal/supervised session.
-
FIX-M18-NOTES.md — dump re-validation runs every 5 min (perf). FIXED in controller v0.62.0 @
f8afe5c(2026-06-14). (was on the deletedfelhom-controllerbranchfix/m18-dump-validation-cache.) -
FIX-M19-NOTES.md —
deriveStackNamemisattribution edge (low-incidence correctness). FIXED in controller v0.62.0 @6bab68b(2026-06-14). (was on the deleted branchfix/m19-stackname-crossref.) -
FOLLOWUP-golden-default-controller-tag.md — the golden bakes a stale controller (
:0.43.0when queued; had rotted again to:0.85.1by resolution). FIXED in felhom-agent @ceca355(2026-07-03):build-golden.shv2.0.0 makes the controller tag a MANDATORY argument (a required arg cannot rot) and golden 0.98.3 was baked + clean-room-validated (bake → first-boot-current → self-manage → app deploy, on the drill VM — no supervised touch of live guests needed) + published + vouched. Evidence:../audits/DRILL-golden-098-2026-07-03.md.
Related: the live-drive fixspec (../audits/live-drive-fixspec-2026-06-14.md) carries the deferred
supervised items F9 (HDD provisioning/guest-attach), F20-BUG2 (durable_id scheme), F20-BUG3 (async
mkfs) — to be implemented in the agent/golden supervised session.