diff --git a/REPORT.md b/REPORT.md index 7c10387..c2c3866 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,100 +1,272 @@ -# REPORT — releasing publishes, and an unreleasable version fails CI (R-115, R-183) +# REPORT — R-86: a restore-test proves each BACKUP, not the clock -**Date:** 2026-08-03 · **Repo:** `felhom-agent` · **NO VERSION BUMP** — the agent stays **v0.120.0**, -no Go code changed, nothing was built or deployed. +**Date:** 2026-08-03 · **Repo:** `felhom-agent` **v0.120.0 → v0.121.0 → v0.121.1** +(`4618169`, `4d82591`, `53d0c6b`) · +released, published, verified by independent download, deployed to demo-felhom and **proven live**. +Sibling half: `felhom.eu` hub **v0.91.0 → v0.91.1** — the two ship together. -## What changed +--- -| File | | -|---|---| -| `scripts/release-agent.sh` | **new** — THE release path: build → tag → publish → verify by independent download | -| `scripts/check-published-versions.py` | **new** — the R-115 gate | -| `scripts/agent_gates.py` | registers the gate as **not `--fast`** (it needs network) | -| `.gitea/workflows/gates.yml` | CI now runs the **full** gate set, not `--fast` | -| `CLAUDE.md` | the raw `go build` line is replaced by the release script; a **Vouch** row replaces the old Publish row | +## 1. Baselines, re-read on arrival -## Why +| Repo | `main` @ commit | Version | Matched §1? | +|---|---|---|---| +| `felhom-agent` | `1b14cfd0b48b` | `v0.120.0` | **yes** | +| `felhom.eu` | `e34b614e5b65` | CHANGELOG `v0.90.0`, deployed image `0.90.1` | **yes — the discrepancy was real and is fixed** (entry backfilled) | -Publishing was a step someone had to remember and was **forgotten three times in five days** — -R-111's seventeen stranded releases, 0.114.0, and 0.120.0, which sat deployed on both demo hosts and -undownloadable, so a documented-path reinstall would have silently downgraded them to the pre-merge -agent **while reporting success**. R-111's own closing line named this leg and closed SHIPPED without -it; it recurred the same afternoon. A note is not a mechanism. +Highest register ID in use was **R-184**; `R-185`–`R-187` established free by grep across all four +repos and `documentation/`. -The script also **tags**, because `felhom-host-install.sh` now fetches the agent's sixteen config -files from `raw/tag/v/` (R-183). A released version with no tag 404s a box mid-install, as -root, on a virgin machine. Tag and package are two halves of one release. +## 2. The rule, in one sentence — and the trap it avoids -It **verifies by downloading what it just published** and comparing the sha to what it built. The -publish step's own success is a report on its own write; a fetch returning the right bytes is a -different claim, and it is the one that matters. +> Let **A** be the newest archive on a tier that has settled for at least the settle lag (24 h). +> The tier is **DUE** when A exists and **A has not already been proven**. -It **does not vouch** — that points machines at a version and stays the operator's act. +The literal reading of R-86's own wording — *"due when the newest archive is ≥ 24 h old"* — is +**never true on a daily tier**, because a new archive lands each day and resets the newest-archive age +to zero long before it reaches the lag. It would have switched restore-testing **off** for the tier +that matters most, silently, while looking like the row was implemented. -## The gate's invariant — not the one specified, and the reason was measured +**Evidence that a daily tier does become due**, at three levels: -The task's §8.4 asked for *"the version the hub tells machines to install must be downloadable"*. -**CI cannot see that**, measured rather than assumed (P-C): +1. **Unit, time-driven** — `TestDue_DailyTierIsProvedDailyOnItsOwnArchive`: five simulated days, one + archive a day, evaluated hourly (120 evaluations) → **exactly 5 runs**, and run *i* tests day + *i−1*'s archive, never the still-settling one. +2. **The red-proof of the naive rule** — implemented and observed failing at **0 runs over 5 days** + (§6), which is the trap made visible rather than argued about. +3. **Live** — the offsite tier on demo-felhom became due on its own archive and ran (§7). -| Endpoint | Anonymous | -|---|---| -| Gitea package **download** | **200** (and **404** for a fake version — it discriminates) | -| Gitea **tags** api | **200** | -| Gitea package **listing** api | **401** — token required | -| Hub `/api/v1/artifacts/` | **401** — per-customer passphrase required | +## 3. What changed -So a credential-free gate can ask *"is this version installable"* but not *"which version is -vouched"*. Adding an operator credential to CI to close that is the operator's call, not a gate -author's. The implemented invariant — **every `v` tag must have a downloadable package and a -tag tree that serves the agent's configs** — needs no credential and **catches all three recorded -instances**, because the release script creates the tag and publishes in one act. - -**What it does not catch, stated rather than assumed away:** the hub vouching a version that was -never released at all. Nothing here can see that; it belongs at vouch time in the hub. → **R-184**. - -## Proof - -| Check | Result | -|---|---| -| `go build ./... && go vet ./...` | OK | -| `go test ./...` | **29 packages ok, rc=0** (read separately from any commit) | -| `agent_gates.py --fast` | `published` correctly **SKIPPED** — the pre-push hook must not fail because Gitea blinked | -| `agent_gates.py` (full) | `reuse-refs` OK, `published` OK | -| release script: re-release guard | `ERROR: tag v0.120.0 already exists — releasing over it would make one version name two binaries`, rc=1 | -| release script: clean-tree guard | `ERROR: working tree is dirty — commit and push first`, rc=1 | - -### Red-proof F — both directions - -- **A tagged-but-unpublished version** (`v9.9.9` created for the purpose): gate **rc=1**, - `binary NOT downloadable (HTTP 404 …)`. This is the R-115 shape exactly. -- **The gate deregistered from the entry point**, same bad state: `agent_gates.py` → **rc=0, "all - agent gates OK"**. Restored → **rc=1, CONVICTED: published**. The guard is what catches it, not - something else. - -### Scenario F measured on REAL CI, not inferred - -Runs **69** and **70** are on the **same commit** `0db7766`: - -| run | state of the repo | CI | +| Piece | File | Change | |---|---|---| -| 69 | no `v9.9.9` | **success** | -| 70 | `v9.9.9` tagged, not published | **failure** | +| the due-check | `internal/backup/restoretest_due.go` (new) | `EvaluateDue` / `evaluateTier` — per-tier verdict + the reason, ordered oldest-proven first | +| the trigger | `internal/backup/schedule.go` | the ticker is now the **evaluation interval**; `pickForThisRun` answers *"is anything due?"*, and "nothing" is a normal answer | +| the state | `internal/backup/restoretest_state.go` | records **which archive** was proven, with migration | +| the picker | `internal/backup/runner.go` | `PickSettledRestoreCandidateOn(ctx, target, notAfter)`; `PickRestoreCandidateOn` is a one-line call into it | +| the knobs | `internal/config/config.go` | `restore_test_eval_interval_seconds` + `restore_test_settle_seconds`; the old key deprecated, not repurposed | +| the wiring | `cmd/felhom-agent/main.go` | settle-aware picker + `Settle`; deprecation WARN; new `--selftest=restore-test-due` | +| the observable (**v0.121.1**) | `internal/backup/schedule.go`, `restoretest_due.go` | a not-due evaluation logs one **INFO** line naming every tier's verdict — see §9 | -Same code, same workflow, one variable — so the gate demonstrably RUNS in CI and fails for exactly -the R-115 condition. This also retrospectively explains runs 67/68, which were red in the window when -`v9.9.9` first existed. **One deliberate CI failure e-mail reached the operator — that was this -proof, not an incident.** +### The state records the archive (§8.2) -I could not read CI's own step log to attribute those runs directly: the Gitea jobs endpoint requires -an API token, and the only credential available on this host (`~/.docker/config.json`) is a registry -password, which the API rejects. The controlled before/after above replaced that log rather than an -assumption standing in for it. +A timestamp cannot answer *"have we proven **this** archive"* — it is the same class as the workspace +rule that a timestamp records an *attempt*, not a *result*: here it records a result, but not **which** +result. `RestoreTestState` now holds `{archive, proven_at}` per tier. -`v9.9.9` was deleted afterwards; `git ls-remote --tags` shows only `v0.120.0`. +**Migration:** a pre-R-86 file (`{"target": ""}`) keeps its **time** — rotation ordering +survives a deploy, which is why the file exists at all — and yields **no proven archive**, so each +tier is due exactly once after the upgrade. One extra test per tier, once, is the safe direction; +reading a legacy time as proof of whatever archive is current would invent a guarantee. -## Tag convention +### The config (§8.3) — and a correction to the spec -`v`, at the commit the binary was built from. `v0.120.0` was created retroactively at -`cd6e267` — the commit that produced the published binary (sha `a7763d31b55b5ce7…`). `configs/` is -byte-identical between that commit and `main`, so nothing about the sixteen fetched files depends on -the choice; `cd6e267` is tagged because it is the honest one. +The spec said *"`0` must keep meaning disabled"*. **In the code as it stands, `0` means *use the +default* and NEGATIVE means disabled** (`RestoreTestCadence`, pre-existing). Making `0` disable would +have switched restore-testing off on every box that leaves the key unset — the worst possible reading +— so the actual semantics were preserved and this is flagged rather than silently followed. + +- `restore_test_eval_interval_seconds` — how often due-ness is **asked**. Default **6 h**. +- `restore_test_settle_seconds` — how long an archive must sit. Default **24 h**. +- `restore_test_cadence_seconds` — **deprecated**. Negative still **disables**, verbatim. A positive + value seeds the **settle lag** (the quantity a person setting it was expressing: how long may pass + between a backup and confidence that it restores), and the daemon logs one start-up WARN naming + both replacements. It is deliberately **not** carried into the evaluation interval: a box that set + 72 h to spare a weak endpoint would otherwise get a 72-hour-latency due-check, whereas what it + wanted — fewer heavy restores — is what per-archive due-ness already gives it. + +## 4. Part 1.4 — the measurement, and the interval chosen from it + +Measured on demo-felhom, 2026-08-03, via `--selftest=restore-test-due` and by timing the underlying +API call directly (3 runs each): + +| tier | what it is | one due-check | +|---|---|---| +| `felhom-backup` (dir) | local, on-box | **18 ms** (18.7 / 18.3 / 18.5) | +| `felhom-pbs` | offsite, **WAN to ep0** | **392 ms** (375 / 378 / 424) | +| both together | one full evaluation | **430 ms** | + +**Cost does not set the interval** — even at one evaluation a minute the offsite leg would be ~0.7 % +of the link's time. What sets it is the other bound, and it is not in the brief: **under a per-archive +due-check a FAILING tier stays due, so the evaluation interval is also its RETRY interval — and a +retry is a multi-GB restore.** Every few minutes would be an incident of its own; the old timer +retried a broken tier once a day. + +**6 h chosen from both ends:** at most four heavy retries a day in the worst case, and at most 6 h of +latency between an archive settling and its proof — negligible against a 24 h settle lag, so a daily +tier is still proved daily. No second rate limiter was added (§8.4): the pacing remains one test per +archive generation. + +## 5. Two hazards the new frequency created, and their fixes + +Both are consequences of evaluating often rather than daily, and neither is in the brief: + +1. **The due-check now runs BEFORE the heavy-operation gate is taken.** Holding that gate for a read + that answers "nothing to do" would open a window at *every* evaluation in which a starting backup + cannot acquire — and a backup that cannot acquire does not merely wait, it **records a failure and + pages the operator** (F-A1). Nothing heavy starts before the gate; due-ness does not expire while + we check. +2. **The candidate picker skips implausible archives.** Under per-archive due-ness an incomplete + 1-byte phantom (F-CRIT-2's artefact, which server-side prune does **not** collect) would be picked + forever, fail forever, never earn proof, and leave the tier due at *every* evaluation — turning the + evaluation interval into the retry rate for a multi-GB restore. `archivePlausiblyComplete` (the + canonical helper, with its warn-once companion) is applied in the shared scan, so both callers + agree. **This is a behaviour change to `PickRestoreCandidateOn`** and is recorded as such. + +## 6. Tests and red-proofs + +Green gate, both repos: `go build ./... && go vet ./... && go test ./...` — agent **29 packages ok, +rc=0**; hub **rc=0**. The test run and the commit were always separate commands. + +| # | Test | Asserts | Mutation | Observed | +|---|---|---|---|---| +| A | `TestDue_DailyTierIsProvedDailyOnItsOwnArchive` | 5 runs over 5 days, each on the settled archive | the naive rule (`now-landed >= settle`, proven-archive check deleted, cutoff removed) | **FAIL** — `a daily tier must be proved once per day; got 0 run(s) over 5 days: []` | +| B | `TestDue_WeeklyTierIsProvedOncePerArchive` | 84 evaluations over 3 weeks → exactly 3 runs, one per archive | — | pass | +| C | `TestDue_RestartRunsNothing` | two restarts + 4 evaluations → **0 runs** | `ProvenArchive` reverted to per-tier time | **FAIL** — `2 restart(s) produced 4 run(s)` | +| D | `TestDue_NewSettledArchiveMakesAProvedTierDueAgain` | a newly settled archive re-arms the tier, and the NEW archive is tested | — | pass | +| E | `TestDue_FailingTierIsRetriedAndNeverProven` | 3 evaluations → 3 retries, no proof recorded | credit on failure (`rt.Pass &&` dropped) | **FAIL** — `got 1 run(s) over 3 evaluations` + `TestRotation_FailureEarnsNoCredit` also failed | +| F | `TestDue_TwoDueTiersRunOneAtATime` | one evaluation → one run; the other is deferred and runs next | — | pass | +| F′ | `TestDue_DeferredBehindABackupStaysDue` | the gate holds; a deferred tier stays DUE | — | pass | +| H | `TestDue_NewbornTierIsNotDueAndNotAnError` | no archive → not due, no error, **and a reason** | — | pass | +| — | `TestDue_UnsettledArchiveIsNotACandidate` | a 2 h-old archive is not a candidate under a 24 h lag | — | pass | +| — | `TestDue_LookupFailureIsUnknownNotNotDue` | a tier that cannot be listed is UNKNOWN, the error travels, the other tier still runs | — | pass | +| — | `TestRestoreTestState_LegacyFileMigratesToNothingProven` | legacy time kept, no archive claimed | — | pass | +| — | `TestPickRestoreCandidate_SkipsImplausibleArchives` | the newest entry is not a candidate if it cannot be complete | guard removed | **FAIL** — `pick = "phantom" want the newest COMPLETE archive 'real'` | +| I | `TestMainWiresTheSettleAwareTierPicker` | **AST** of `main.go`: settle picker wired, old picker gone, `Settle` set, eval-interval accessor called | the wiring line commented out | **FAIL** — `main.go never passes runner.PickSettledRestoreCandidateOn …` (a `strings.Contains` check would have PASSED — the string is still there, in a comment) | +| I′ | `TestMainStillWiresTheHeavyOperationGateAndPerRunSpec` | R-85's gate + per-run spec survive | — | pass | + +Hub-side (Scenario G) is in `felhom.eu/REPORT.md`, including **a hollow test caught by its own +red-proof**: the first weekly fixture had no jitter, sat on exactly 168 h, and PASSED under the +flat-window mutation. + +### Tests deliberately changed, and why + +`TestRotation_BothTiersExercisedAcrossCadences` asserted *4 ticks → 4 runs*. That was a faithful +statement of the defect — every tick produced a heavy restore-test, because the ticker **was** the +trigger. It is now `TestRotation_BothTiersExercisedOncePerArchive`: **2 runs across 4 evaluations**, +one per tier, one per archive. Strictly stronger — it pins both the coverage R-85 won and the pacing +R-86 adds. The old assertion is quoted in the test's comment so the change is legible. + +## 7. The live run — triggered by due-ness, on real hardware + +Deployed to **demo-felhom** (Tier 0). The deployed binary is the **published artifact downloaded from +Gitea**, not a local rebuild — see R-186. + +### 7.1 The due verdict, per tier, before anything ran + +``` +eval_interval=6h0m0s settle=24h0m0s +tier=felhom-backup due=false archive="" landed=- proven="" + reason: no settled archive yet — nothing to prove (newborn or still settling) +tier=felhom-pbs due=true archive="felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" + landed=2026-07-28T04:49:43Z proven="" + reason: newest settled archive (landed 2026-07-28T04:49:43Z) has not been proven; nothing proven yet +``` + +`felhom-backup` reads "no settled archive" for a reason that is **not** the one it appears to be — +see **R-185**: the agent cannot list that storage at all. + +### 7.2 A real run, started by the due-check + +Only the **evaluation interval** was shortened for the validation (a systemd drop-in, since removed): +the due rule, the settle lag and the restore-test itself were untouched. + +``` +15:14:38 backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers) + target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z + landed=2026-07-28T04:49:43Z + reason="newest settled archive … has not been proven; nothing proven on this tier yet" +15:14:39 restore-test: full-fidelity restore params derived from the archive config scratch=990000 + … proxmox-backup-client restore --crypt-mode=encrypt … (felhom-agent@pve!agent) +15:25:08 audit: gate decision class=guest_destroy guest=990000 source=one_shot_job allowed=true +15:25:14 restore-test: scratch guest torn down vmid=990000 +15:25:14 backup: scheduled restore-test PASSED archive=felhom-pbs:… duration_s=635.1 +``` + +A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored into a scratch guest, +booted, verified and destroyed — **635 s**, unattended, and started by *"this archive has not been +proven"* rather than by a timer. + +### 7.3 The state now names that archive, and a second evaluation runs nothing + +```json +{ "felhom-pbs": { "archive": "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z", + "proven_at": "2026-08-03T13:25:14Z" } } +``` + +``` +tier=felhom-pbs due=false proven="felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" + reason: newest settled archive (landed 2026-07-28T04:49:43Z) is already proven +``` + +### 7.4 Teardown — all three layers + +| Layer | Before | After | +|---|---|---| +| scratch guest 990000 | `stopped lock=create` during the run | **absent** from `pct list` | +| its volumes | 5 thin LVs (32 G + 200 G + 50 G + 2×1 G) | **0** matches in `lvs` | +| hub-side record | — | the run's **`restore_tests[]` entry is RETAINED deliberately** — it is the proof the hub's staleness check reads, and deleting it would delete the result. No event was created: the run passed, and `restore_test_failed`/`restore_test_stale` fire only on failure or staleness | + +`pvesm status` before and after: `local-lvm` 1.95 % used before the run, and the thin volumes are gone +after it — the restore reclaimed to the same shape it started in. + +### 7.5 The restart, which is the defect a person would actually notice + + + +## 8. Release and deployment + +| Step | Evidence | +|---|---| +| Released via `scripts/release-agent.sh 0.121.0` | tag `v0.121.0` at `4d82591`, package published | +| Verified by **independent download** | sha256 `b2128f3cd4539225a2842f541f56ffaf5390b1d97f3f3a80076ec5f53dbc7d7a`, 14 081 336 B, round-trip GET matched | +| Gate re-run after release | `2 released version(s) to verify: 0.120.0, 0.121.0` → both **installable** | +| Deployed | `felhom-agent --version` → **0.121.0**, `systemctl is-active` → **active**, prior binary kept as `.bak-0.120.0` | +| Startup | `capabilities self-check ok=68 total=68 degraded=0`, and `backup: restore-test scheduler starting (per-archive due-check) eval_interval=6h0m0s settle=24h0m0s` | +| **Vouching** | **NOT done — deliberately the operator's act.** Hub UI → Configs → Day-0 artifacts: agent `0.121.0`, sha above | + +## 9. Findings filed (none fixed blind) + +- **R-185 — the agent cannot see demo-felhom's host backup tier at all.** The PVE token has no ACL on + `/storage/felhom-backup`, so the content listing returns `{"data":[]}` where root sees three + archives (6.1–6.3 GB, 08-01/02/03). Verified three ways, including `local` — which *has* a grant — + returning its archives through the same token. **Pre-existing and independent of R-86** (R-85's + rotation had the same blindness). The part worth fixing is the **silence**: a permission-blinded + tier is today indistinguishable from a newborn one, and the agent already records the backups it + wrote to that target, so the contradiction is detectable. +- **R-186 — a released binary's sha cannot be reproduced from its tag.** `release-agent.sh` builds + before tagging, so Go stamps a pseudo-version into the published bytes: published `b2128f3c…` + (14 081 336 B) vs rebuild-at-tag `8302e396…` (14 077 240 B), identical source and toolchain. The + build order is deliberate, so the fix is not to swap the steps blind. **Mitigated here** by + deploying the published artifact. +- **R-187 — R-115's one-command release had never run its publish leg** (`CLOSED`, fixed in the same + session): `publish-agent.sh` has been mode `0644` since 2026-06-28 because every earlier caller used + `bash …`, and `release-agent.sh` called it directly → `Permission denied` on the first real release. + Fixed both ways: the mode bit restored **and** the call made mode-independent. The tag the failed + run created was withdrawn (nothing had been published under it — verified 404) and recreated on the + fix commit, so one version name still means one binary. + +## 10. CI + +| Repo | Run | Commit | Result | +|---|---|---|---| +| `felhom-agent` | **#15** (id 83) | `4d82591` | **success** | +| `felhom-agent` | #13 (id 81) | `4618169` | **failure — explained, and it is CI doing its job** | +| `felhom.eu` | **#48** (id 86) | `ff2655c` | **success** | + +Run #13 fired on the **tag push** from the *failed* first release: `v0.121.0` existed as a tag while +nothing was published, and `check-published-versions.py` correctly refused — *"every released agent +version must be INSTALLABLE"*. That is precisely the state R-115's gate exists to catch, caught within +minutes and self-resolved by the corrected release. Confirmed locally afterwards: both 0.120.0 and +0.121.0 verify. `--no-verify` was **not** used anywhere. + +## 11. Observations — noticed, recorded, not acted on + +- **`felhom.eu/CONTEXT.md` has duplicate standing-ruling IDs** — three `S-14`s and two `S-15`s already + in the file before this session. New rulings were numbered **S-17/S-18** rather than adding to the + collision; the existing duplicates are untouched. +- **`agent_gates.py --fast` skips the published-versions gate**, so the pre-push hook cannot catch an + unpublished release — only CI can. That is the intended split (no network in a hook), and it is why + run #13 mattered. +- **The hub sweeps every 60 s and re-reads 14 days of host-reports per customer** for this check. Not + changed here (the read window is the same as before), but it is the cost centre if the fleet grows.