Night 2026-10-06: R-894 fixed on agent main (07 §6.1), R-892 still blocked (no agent key in this shell); night log started
gates / gates (push) Successful in 2m32s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-06 20:34:53 +02:00
parent 9246f62c3b
commit 8e2dc2049f
8 changed files with 73 additions and 2 deletions
@@ -374,6 +374,17 @@ never moved by an update or an undo, and the unit restore accepts it.
> **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.
**[FACT, agent main 2026-10-06 night, ships as v0.150.0 — R-894] When is a whole-guest tier due, and what if its
storage cannot be read?** The agent answers the controller's `GET /backup/due?target=…` from the tier's STORAGE first
(the newest archive there — the ground truth, R-84), then the newest success this process recorded. Until now a restart
emptied that record (memory only, R-348), so a restart followed by an unreachable storage read DUE as if no copy
existed — measured on demo-hp 2026-10-05 (the 7-day off-site tier, last copy 4 days old, read due; vzdump then failed).
**Now the newest success per tier is also kept on disk** (`backup-success-state.json` in the agent's state dir) and is
read **only when the storage cannot be read**: a fresh saved copy → not due; a saved copy older than the cadence → DUE
(the deliberate rule stays — an unreadable storage never suppresses a backup that is due); no saved copy → due, age
unknown, as before. A storage that answers always wins, so a pruned archive still makes the tier due. Tests:
`TestBackupDue_R894_*` (felhom-agent `internal/localapi`).
**[FACT, 2026-10-04 — agent v0.140.0, `11-os-updates.md` §8.1] The OS leg closes the night.** After the
whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent
waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a