# REPORT — R-86: a restore-test proves each BACKUP, not the clock **Date:** 2026-08-03 · **Repo:** `felhom-agent` **v0.120.0 → v0.121.0 → v0.121.1** (`4618169`, `4d82591`, `53d0c6b`) · released, published, verified by independent download, deployed to demo-felhom and **proven live**. Sibling half: `felhom.eu` hub **v0.91.0 → v0.91.1** — the two ship together. --- ## 1. Baselines, re-read on arrival | Repo | `main` @ commit | Version | Matched §1? | |---|---|---|---| | `felhom-agent` | `1b14cfd0b48b` | `v0.120.0` | **yes** | | `felhom.eu` | `e34b614e5b65` | CHANGELOG `v0.90.0`, deployed image `0.90.1` | **yes — the discrepancy was real and is fixed** (entry backfilled) | Highest register ID in use was **R-184**; `R-185`–`R-187` established free by grep across all four repos and `documentation/`. ## 2. The rule, in one sentence — and the trap it avoids > Let **A** be the newest archive on a tier that has settled for at least the settle lag (24 h). > The tier is **DUE** when A exists and **A has not already been proven**. The literal reading of R-86's own wording — *"due when the newest archive is ≥ 24 h old"* — is **never true on a daily tier**, because a new archive lands each day and resets the newest-archive age to zero long before it reaches the lag. It would have switched restore-testing **off** for the tier that matters most, silently, while looking like the row was implemented. **Evidence that a daily tier does become due**, at three levels: 1. **Unit, time-driven** — `TestDue_DailyTierIsProvedDailyOnItsOwnArchive`: five simulated days, one archive a day, evaluated hourly (120 evaluations) → **exactly 5 runs**, and run *i* tests day *i−1*'s archive, never the still-settling one. 2. **The red-proof of the naive rule** — implemented and observed failing at **0 runs over 5 days** (§6), which is the trap made visible rather than argued about. 3. **Live** — the offsite tier on demo-felhom became due on its own archive and ran (§7). ## 3. What changed | Piece | File | Change | |---|---|---| | the due-check | `internal/backup/restoretest_due.go` (new) | `EvaluateDue` / `evaluateTier` — per-tier verdict + the reason, ordered oldest-proven first | | the trigger | `internal/backup/schedule.go` | the ticker is now the **evaluation interval**; `pickForThisRun` answers *"is anything due?"*, and "nothing" is a normal answer | | the state | `internal/backup/restoretest_state.go` | records **which archive** was proven, with migration | | the picker | `internal/backup/runner.go` | `PickSettledRestoreCandidateOn(ctx, target, notAfter)`; `PickRestoreCandidateOn` is a one-line call into it | | the knobs | `internal/config/config.go` | `restore_test_eval_interval_seconds` + `restore_test_settle_seconds`; the old key deprecated, not repurposed | | the wiring | `cmd/felhom-agent/main.go` | settle-aware picker + `Settle`; deprecation WARN; new `--selftest=restore-test-due` | | the observable (**v0.121.1**) | `internal/backup/schedule.go`, `restoretest_due.go` | a not-due evaluation logs one **INFO** line naming every tier's verdict — see §9 | ### The state records the archive (§8.2) A timestamp cannot answer *"have we proven **this** archive"* — it is the same class as the workspace rule that a timestamp records an *attempt*, not a *result*: here it records a result, but not **which** result. `RestoreTestState` now holds `{archive, proven_at}` per tier. **Migration:** a pre-R-86 file (`{"target": ""}`) keeps its **time** — rotation ordering survives a deploy, which is why the file exists at all — and yields **no proven archive**, so each tier is due exactly once after the upgrade. One extra test per tier, once, is the safe direction; reading a legacy time as proof of whatever archive is current would invent a guarantee. ### The config (§8.3) — and a correction to the spec The spec said *"`0` must keep meaning disabled"*. **In the code as it stands, `0` means *use the default* and NEGATIVE means disabled** (`RestoreTestCadence`, pre-existing). Making `0` disable would have switched restore-testing off on every box that leaves the key unset — the worst possible reading — so the actual semantics were preserved and this is flagged rather than silently followed. - `restore_test_eval_interval_seconds` — how often due-ness is **asked**. Default **6 h**. - `restore_test_settle_seconds` — how long an archive must sit. Default **24 h**. - `restore_test_cadence_seconds` — **deprecated**. Negative still **disables**, verbatim. A positive value seeds the **settle lag** (the quantity a person setting it was expressing: how long may pass between a backup and confidence that it restores), and the daemon logs one start-up WARN naming both replacements. It is deliberately **not** carried into the evaluation interval: a box that set 72 h to spare a weak endpoint would otherwise get a 72-hour-latency due-check, whereas what it wanted — fewer heavy restores — is what per-archive due-ness already gives it. ## 4. Part 1.4 — the measurement, and the interval chosen from it Measured on demo-felhom, 2026-08-03, via `--selftest=restore-test-due` and by timing the underlying API call directly (3 runs each): | tier | what it is | one due-check | |---|---|---| | `felhom-backup` (dir) | local, on-box | **18 ms** (18.7 / 18.3 / 18.5) | | `felhom-pbs` | offsite, **WAN to ep0** | **392 ms** (375 / 378 / 424) | | both together | one full evaluation | **430 ms** | **Cost does not set the interval** — even at one evaluation a minute the offsite leg would be ~0.7 % of the link's time. What sets it is the other bound, and it is not in the brief: **under a per-archive due-check a FAILING tier stays due, so the evaluation interval is also its RETRY interval — and a retry is a multi-GB restore.** Every few minutes would be an incident of its own; the old timer retried a broken tier once a day. **6 h chosen from both ends:** at most four heavy retries a day in the worst case, and at most 6 h of latency between an archive settling and its proof — negligible against a 24 h settle lag, so a daily tier is still proved daily. No second rate limiter was added (§8.4): the pacing remains one test per archive generation. ## 5. Two hazards the new frequency created, and their fixes Both are consequences of evaluating often rather than daily, and neither is in the brief: 1. **The due-check now runs BEFORE the heavy-operation gate is taken.** Holding that gate for a read that answers "nothing to do" would open a window at *every* evaluation in which a starting backup cannot acquire — and a backup that cannot acquire does not merely wait, it **records a failure and pages the operator** (F-A1). Nothing heavy starts before the gate; due-ness does not expire while we check. 2. **The candidate picker skips implausible archives.** Under per-archive due-ness an incomplete 1-byte phantom (F-CRIT-2's artefact, which server-side prune does **not** collect) would be picked forever, fail forever, never earn proof, and leave the tier due at *every* evaluation — turning the evaluation interval into the retry rate for a multi-GB restore. `archivePlausiblyComplete` (the canonical helper, with its warn-once companion) is applied in the shared scan, so both callers agree. **This is a behaviour change to `PickRestoreCandidateOn`** and is recorded as such. ## 6. Tests and red-proofs Green gate, both repos: `go build ./... && go vet ./... && go test ./...` — agent **29 packages ok, rc=0**; hub **rc=0**. The test run and the commit were always separate commands. | # | Test | Asserts | Mutation | Observed | |---|---|---|---|---| | A | `TestDue_DailyTierIsProvedDailyOnItsOwnArchive` | 5 runs over 5 days, each on the settled archive | the naive rule (`now-landed >= settle`, proven-archive check deleted, cutoff removed) | **FAIL** — `a daily tier must be proved once per day; got 0 run(s) over 5 days: []` | | B | `TestDue_WeeklyTierIsProvedOncePerArchive` | 84 evaluations over 3 weeks → exactly 3 runs, one per archive | — | pass | | C | `TestDue_RestartRunsNothing` | two restarts + 4 evaluations → **0 runs** | `ProvenArchive` reverted to per-tier time | **FAIL** — `2 restart(s) produced 4 run(s)` | | D | `TestDue_NewSettledArchiveMakesAProvedTierDueAgain` | a newly settled archive re-arms the tier, and the NEW archive is tested | — | pass | | E | `TestDue_FailingTierIsRetriedAndNeverProven` | 3 evaluations → 3 retries, no proof recorded | credit on failure (`rt.Pass &&` dropped) | **FAIL** — `got 1 run(s) over 3 evaluations` + `TestRotation_FailureEarnsNoCredit` also failed | | F | `TestDue_TwoDueTiersRunOneAtATime` | one evaluation → one run; the other is deferred and runs next | — | pass | | F′ | `TestDue_DeferredBehindABackupStaysDue` | the gate holds; a deferred tier stays DUE | — | pass | | H | `TestDue_NewbornTierIsNotDueAndNotAnError` | no archive → not due, no error, **and a reason** | — | pass | | — | `TestDue_UnsettledArchiveIsNotACandidate` | a 2 h-old archive is not a candidate under a 24 h lag | — | pass | | — | `TestDue_LookupFailureIsUnknownNotNotDue` | a tier that cannot be listed is UNKNOWN, the error travels, the other tier still runs | — | pass | | — | `TestRestoreTestState_LegacyFileMigratesToNothingProven` | legacy time kept, no archive claimed | — | pass | | — | `TestPickRestoreCandidate_SkipsImplausibleArchives` | the newest entry is not a candidate if it cannot be complete | guard removed | **FAIL** — `pick = "phantom" want the newest COMPLETE archive 'real'` | | I | `TestMainWiresTheSettleAwareTierPicker` | **AST** of `main.go`: settle picker wired, old picker gone, `Settle` set, eval-interval accessor called | the wiring line commented out | **FAIL** — `main.go never passes runner.PickSettledRestoreCandidateOn …` (a `strings.Contains` check would have PASSED — the string is still there, in a comment) | | I′ | `TestMainStillWiresTheHeavyOperationGateAndPerRunSpec` | R-85's gate + per-run spec survive | — | pass | Hub-side (Scenario G) is in `felhom.eu/REPORT.md`, including **a hollow test caught by its own red-proof**: the first weekly fixture had no jitter, sat on exactly 168 h, and PASSED under the flat-window mutation. ### Tests deliberately changed, and why `TestRotation_BothTiersExercisedAcrossCadences` asserted *4 ticks → 4 runs*. That was a faithful statement of the defect — every tick produced a heavy restore-test, because the ticker **was** the trigger. It is now `TestRotation_BothTiersExercisedOncePerArchive`: **2 runs across 4 evaluations**, one per tier, one per archive. Strictly stronger — it pins both the coverage R-85 won and the pacing R-86 adds. The old assertion is quoted in the test's comment so the change is legible. ## 7. The live run — triggered by due-ness, on real hardware Deployed to **demo-felhom** (Tier 0). The deployed binary is the **published artifact downloaded from Gitea**, not a local rebuild — see R-186. ### 7.1 The due verdict, per tier, before anything ran ``` eval_interval=6h0m0s settle=24h0m0s tier=felhom-backup due=false archive="" landed=- proven="" reason: no settled archive yet — nothing to prove (newborn or still settling) tier=felhom-pbs due=true archive="felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" landed=2026-07-28T04:49:43Z proven="" reason: newest settled archive (landed 2026-07-28T04:49:43Z) has not been proven; nothing proven yet ``` `felhom-backup` reads "no settled archive" for a reason that is **not** the one it appears to be — see **R-185**: the agent cannot list that storage at all. ### 7.2 A real run, started by the due-check Only the **evaluation interval** was shortened for the validation (a systemd drop-in, since removed): the due rule, the settle lag and the restore-test itself were untouched. ``` 15:14:38 backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers) target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z landed=2026-07-28T04:49:43Z reason="newest settled archive … has not been proven; nothing proven on this tier yet" 15:14:39 restore-test: full-fidelity restore params derived from the archive config scratch=990000 … proxmox-backup-client restore --crypt-mode=encrypt … (felhom-agent@pve!agent) 15:25:08 audit: gate decision class=guest_destroy guest=990000 source=one_shot_job allowed=true 15:25:14 restore-test: scratch guest torn down vmid=990000 15:25:14 backup: scheduled restore-test PASSED archive=felhom-pbs:… duration_s=635.1 ``` A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored into a scratch guest, booted, verified and destroyed — **635 s**, unattended, and started by *"this archive has not been proven"* rather than by a timer. ### 7.3 The state now names that archive, and a second evaluation runs nothing ```json { "felhom-pbs": { "archive": "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z", "proven_at": "2026-08-03T13:25:14Z" } } ``` ``` tier=felhom-pbs due=false proven="felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" reason: newest settled archive (landed 2026-07-28T04:49:43Z) is already proven ``` ### 7.4 Teardown — all three layers | Layer | Before | After | |---|---|---| | scratch guest 990000 | `stopped lock=create` during the run | **absent** from `pct list` | | its volumes | 5 thin LVs (32 G + 200 G + 50 G + 2×1 G) | **0** matches in `lvs` | | hub-side record | — | the run's **`restore_tests[]` entry is RETAINED deliberately** — it is the proof the hub's staleness check reads, and deleting it would delete the result. No event was created: the run passed, and `restore_test_failed`/`restore_test_stale` fire only on failure or staleness | `pvesm status` before and after: `local-lvm` 1.95 % used before the run, and the thin volumes are gone after it — the restore reclaimed to the same shape it started in. ### 7.5 The restart, which is the defect a person would actually notice ## 8. Release and deployment | Step | Evidence | |---|---| | Released via `scripts/release-agent.sh 0.121.0` | tag `v0.121.0` at `4d82591`, package published | | Verified by **independent download** | sha256 `b2128f3cd4539225a2842f541f56ffaf5390b1d97f3f3a80076ec5f53dbc7d7a`, 14 081 336 B, round-trip GET matched | | Gate re-run after release | `2 released version(s) to verify: 0.120.0, 0.121.0` → both **installable** | | Deployed | `felhom-agent --version` → **0.121.0**, `systemctl is-active` → **active**, prior binary kept as `.bak-0.120.0` | | Startup | `capabilities self-check ok=68 total=68 degraded=0`, and `backup: restore-test scheduler starting (per-archive due-check) eval_interval=6h0m0s settle=24h0m0s` | | **Vouching** | **NOT done — deliberately the operator's act.** Hub UI → Configs → Day-0 artifacts: agent `0.121.0`, sha above | ## 9. Findings filed (none fixed blind) - **R-185 — the agent cannot see demo-felhom's host backup tier at all.** The PVE token has no ACL on `/storage/felhom-backup`, so the content listing returns `{"data":[]}` where root sees three archives (6.1–6.3 GB, 08-01/02/03). Verified three ways, including `local` — which *has* a grant — returning its archives through the same token. **Pre-existing and independent of R-86** (R-85's rotation had the same blindness). The part worth fixing is the **silence**: a permission-blinded tier is today indistinguishable from a newborn one, and the agent already records the backups it wrote to that target, so the contradiction is detectable. - **R-186 — a released binary's sha cannot be reproduced from its tag.** `release-agent.sh` builds before tagging, so Go stamps a pseudo-version into the published bytes: published `b2128f3c…` (14 081 336 B) vs rebuild-at-tag `8302e396…` (14 077 240 B), identical source and toolchain. The build order is deliberate, so the fix is not to swap the steps blind. **Mitigated here** by deploying the published artifact. - **R-187 — R-115's one-command release had never run its publish leg** (`CLOSED`, fixed in the same session): `publish-agent.sh` has been mode `0644` since 2026-06-28 because every earlier caller used `bash …`, and `release-agent.sh` called it directly → `Permission denied` on the first real release. Fixed both ways: the mode bit restored **and** the call made mode-independent. The tag the failed run created was withdrawn (nothing had been published under it — verified 404) and recreated on the fix commit, so one version name still means one binary. ## 10. CI | Repo | Run | Commit | Result | |---|---|---|---| | `felhom-agent` | **#15** (id 83) | `4d82591` | **success** | | `felhom-agent` | #13 (id 81) | `4618169` | **failure — explained, and it is CI doing its job** | | `felhom.eu` | **#48** (id 86) | `ff2655c` | **success** | Run #13 fired on the **tag push** from the *failed* first release: `v0.121.0` existed as a tag while nothing was published, and `check-published-versions.py` correctly refused — *"every released agent version must be INSTALLABLE"*. That is precisely the state R-115's gate exists to catch, caught within minutes and self-resolved by the corrected release. Confirmed locally afterwards: both 0.120.0 and 0.121.0 verify. `--no-verify` was **not** used anywhere. ## 11. Observations — noticed, recorded, not acted on - **`felhom.eu/CONTEXT.md` has duplicate standing-ruling IDs** — three `S-14`s and two `S-15`s already in the file before this session. New rulings were numbered **S-17/S-18** rather than adding to the collision; the existing duplicates are untouched. - **`agent_gates.py --fast` skips the published-versions gate**, so the pre-push hook cannot catch an unpublished release — only CI can. That is the intended split (no network in a hook), and it is why run #13 mattered. - **The hub sweeps every 60 s and re-reads 14 days of host-reports per customer** for this check. Not changed here (the read window is the same as before), but it is the cost centre if the fleet grows.