v0.295.0: a box that was off at its backup time catches up once (R-871, decision 109); the missed-backup banner (decision 110); a late daily timer after a host suspend is skipped
gates / gates (push) Successful in 31s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 09:27:16 +02:00
parent e730a629fd
commit 635c33d381
24 changed files with 2074 additions and 18 deletions
+29
View File
@@ -1,3 +1,32 @@
## v0.295.0 — a box that was off at its backup time catches up once; the household sees why in a banner (R-871, `09` decisions 109–110) (2026-10-05)
**MinAgent: 0.131.0** (unchanged). New household strings: the timeline line `event.backup_catchup_done` and the banner
(`banner.missed_backup.*`, `layout.missed_backup_*`), Hungarian and English.
- **The catch-up (decision 109).** New `internal/nightchain`: a persisted night ledger records when each backup leg
(database dump, second copy, off-site copy) last ran to its end. On a controller start and on a host resume, a leg
that missed its last scheduled time is made up ONCE, 15 minutes later, in the night's order, through the same
wrapped leg bodies the scheduled jobs run (`withLeg`, `catchUpLegs`) — never the app-update leg, never a Docker
step. Several missed nights = one catch-up; a daytime restart after a normal night = none; a power cut mid-chain =
only the legs that did not end; a leg due within 30 min is left to its normal run; the ledger is seeded at the first
start of this release (no catch-up at the upgrade). One lock for every leg. The whole-guest backup defers while a
catch-up runs (`quiesce.SetCatchUpFn`) and the catch-up waits for a quiesce. Measured before the fix on 9202: off
across a 09:05 window, on at 09:08 → `db-dump scheduled for 2026-10-06 09:05`, nothing ran.
- **A host suspend.** A daily job whose timer fires more than 60 min after its wall-clock time is SKIPPED
(`scheduler.DailyLateLimit`) — Go timers run on CLOCK_MONOTONIC, which stops while suspended, so the app-update leg
would otherwise start at noon. A resume watch (wall vs monotonic clock, every minute) triggers the catch-up.
- **The banner (decision 110, the operator's idea).** When the last database dump (and the off-site copy, when
configured) is over 26 h old, every page of a logged-in household says when the last backup was, "the box was off
at backup time (02:30)" when the metrics record shows no sample then, and suggests the latest hour the box is
usually on (5 of the last 7 days, for 3 hours). A button opens the backup-time setting; it never changes the time.
Closing it (`POST /backups/missed-banner/dismiss`) lasts until the next missed backup time (durable, in the ledger);
a successful night removes it. New `MetricsStore.SampleCount` / `HourSampleCounts`.
- The off-site leg's "no target" log line no longer says "the update leg runs now" (the catch-up runs that body too).
- Tests: `internal/nightchain` (`TestCatchUp_*`, `TestResumeWatch_*`, `TestBanner_*`), `TestDaily_LateFireIsSkipped`,
`TestR871_ScheduledCycleWaitsForCatchUp`, `TestR871_CatchUpWiring`, `TestR871_CatchUpRunsNoUpdateLeg`,
`TestR871_Banner*` (real pages + the real dismiss route), parity fixture `launcher_missed_backup`. Red-proofs:
`felhom.eu/documentation/audits/catchup-2026-10-05/part{A,B}/red-proofs.txt`.
## v0.294.0 — the off-site clean-up deletes honest old copies; a first install survives the image clean-up; failed compose logs its reason; the move-aside line names its destination (R-867, R-863, R-864, R-869) (2026-10-05)
**MinAgent: 0.131.0** (unchanged). No new household string.