The ticker survives as the EVALUATION interval only. A tier is DUE when its
newest archive that has settled for `settle` (default 24h) has not been proven:
daily tier -> proved daily on yesterday's archive, weekly tier -> weekly on its
own, newborn -> UNKNOWN.
The trap avoided: the literal reading ("newest archive is >= 24h old") is NEVER
true on a daily tier, so it silently switches restore-testing off where it
matters most. Red-proved at 0 runs over 5 simulated days.
- state records WHICH archive was proven; legacy files keep their time and yield
no proven archive (each tier due once after the upgrade, deliberately)
- two knobs replace one: restore_test_eval_interval_seconds (6h, measured) and
restore_test_settle_seconds (24h). The old cadence key keeps its DISABLE
meaning verbatim and now seeds the settle lag, with a start-up WARN.
- due-check runs BEFORE the heavy-op gate (a frequent poll must not make a
starting backup record a failure, F-A1)
- candidate picker skips implausible archives (a phantom would be due forever)
- new read-only --selftest=restore-test-due prints the verdict + its cost
The scheduler could only ever see cfg.Backup.BackupTarget(), so the offsite
tier's archives were never candidates — which is why demo-hp's DR tier reported
'applied' with zero snapshots for five days and nobody noticed.
Selection: oldest-first (operator ruling, Option 1). Never-proven sorts first,
which is where the offsite tier starts. Ties break on target id so ordering is
deterministic rather than following Go's randomised map order. Rotation credit
only on SUCCESS — a permanently failing tier must keep sorting first, not look
freshly proven and stop being retried.
- backup.RestoreTestState: persisted last-success per tier (atomic tmp+rename).
This genuinely needs persistence unlike R-84: R-84 had ground truth to consult
(the archive is still on the storage), whereas a restore-test destroys its
scratch and leaves no artifact. Corrupt/missing file -> 'nothing proven'.
- backup.InFlight: host-wide one-heavy-op gate shared with the local-API backup
path. A LINK concern, not a lock one — an offsite restore pulls multi-GB over
the same tunnel a backup pushes one, and at ~33 MB/min both drift toward
timeout, which is how a healthy tier gets recorded as failed. Callers DEFER,
never cancel.
- PickRestoreCandidateOn: newest archive on a named tier; '' is not an error, or
every fresh box looks broken for its first week.
- An empty tier is skipped and the next tried; it cannot starve, since it is
still least-recently-proven once it has an archive.
- POST /backup joins the gate (409 naming the holder).
Red-proofs A/E/F observed with the documented text. Full suite green (29
packages, rc=0).