Compare commits
5 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 7581f8140a | |||
| 3d0a1d615d | |||
| 77e2cc4583 | |||
| cd1b087db7 | |||
| 53d0c6bfc4 |
+106
@@ -1,3 +1,109 @@
|
||||
## v0.122.0 — three ways the signals lied about themselves (2026-08-03, R-189 · R-188 · R-186)
|
||||
|
||||
All three are the reporting and release path misreporting its own work. **No customer machine, no
|
||||
backup, no restore, no disk layout, no data.** The restore-test itself and when it runs are unchanged
|
||||
from v0.121.1.
|
||||
|
||||
### R-189 — a passing restore-test no longer vanishes on a restart
|
||||
|
||||
`restore_tests[]` came only from the in-memory `backup.Store`, whose own comment read *"lost on
|
||||
restart; the cadence re-populates"*. That was true under a timer. It stopped being true when R-86 made
|
||||
the agent refuse to re-test an archive it has already proven: a proof lost to a restart is not
|
||||
repeated for a whole archive generation — **a week on the offsite tier** — and the hub calls the tier
|
||||
unproven for all of it.
|
||||
|
||||
**Observed, not predicted (2026-08-03):** a real 14.5 GB offsite restore-test PASSED at 15:25:14, the
|
||||
agent was restarted 2 m 43 s later for a deploy, and the hub logged `0 restore-tests` on the next two
|
||||
host-reports.
|
||||
|
||||
The durable proof already existed — `RestoreTestState`, on disk, per tier, with the archive since
|
||||
R-86 — and `Snapshot()` had carried the doc comment *"for the host-report gauge"* since the day it was
|
||||
written **with no caller at all**: a seam built, documented, and never connected. It now carries the
|
||||
`tier` and what was `verified` as well (stored at proof time, when they are known for certain, rather
|
||||
than derived later by a storage lookup that can fail), and `ProvenRestoreTests` renders them as report
|
||||
entries which the collector merges.
|
||||
|
||||
- **Merge rule: one entry per tier, newest by `TestedAt` wins.** A fresh failure beats a stored
|
||||
success — the failure is the news and lives nowhere else; a stored success beats a stale in-memory
|
||||
entry after a restart; a tier never appears twice, which the hub would read as two tests. An
|
||||
unparseable timestamp counts as older, so a malformed entry cannot displace a good one.
|
||||
- **It refuses to lie.** A record missing the archive or the tier produces NO entry, and run mechanics
|
||||
(scratch VMID, duration) are not re-invented — an absent duration is not a claim, a fabricated one
|
||||
would be. An unproven tier reading as proven would be worse than the defect being fixed.
|
||||
- **Only successes are persisted, and that asymmetry is now written down where it will be read:** a
|
||||
success suppresses future work, so losing it leaves the system quietly less tested than it believes;
|
||||
a failure causes future work and heals itself at the next evaluation.
|
||||
- The `Store` comment that stopped being true is corrected in place rather than left to mislead.
|
||||
|
||||
### R-188 — a correct release no longer emails a failure
|
||||
|
||||
`on: [push]` fires the gates workflow on the **tag** push, and the release pushed its tag *before*
|
||||
publishing, so CI ran the published-versions gate in the seconds before the package existed and
|
||||
correctly reported it missing. Measured across two releases in one session: runs 12/13 and 17/18, same
|
||||
sha each time, opposite results — a race, not a rule. R-168 made that mail the thing that cannot be
|
||||
missed; one that is wrong half the time is one you stop reading.
|
||||
|
||||
**Only the tag PUSH moved** (build → tag locally → publish → push tag). The tag is still created before
|
||||
anything is published, so the build and the tag still describe the same commit; it simply becomes
|
||||
*visible* — to CI, and to any `raw/tag/…` fetch — once the package is downloadable.
|
||||
|
||||
The invariant the old order protected is **not traded away**: `check-published-versions.py` now asserts
|
||||
the converse directly — **no published version may be missing its tag** — as a bounded probe of the
|
||||
frontier (where a failed tag push leaves an orphan) and of patch gaps, printing its probe set every
|
||||
run because a check whose coverage is invisible reads as a guarantee it is not making. The package
|
||||
listing api still answers **401** without a token (re-measured), so absence cannot be enumerated, and
|
||||
the script says so.
|
||||
|
||||
A publish that succeeds and a tag push that then fails now **dies loudly**, printing the one-line
|
||||
recovery; and a publish that *fails* removes the local-only tag so the release can simply be retried
|
||||
instead of colliding with itself.
|
||||
|
||||
### R-186 — a released binary can now be verified by rebuilding it
|
||||
|
||||
`go build` stamps a module version derived from VCS state, so a build made before the tag existed and
|
||||
a rebuild made after it were different binaries. Measured at one commit, same source, same toolchain:
|
||||
|
||||
```
|
||||
default flags, no tag yet ... 18f4a495… 14 085 464 B (mod v0.121.2-0.2026…-3d0a1d61)
|
||||
default flags, tagged ....... 4a38f394… 14 085 440 B (mod v0.121.99)
|
||||
-trimpath -buildvcs=false ... 7ffcdf1d… 14 064 574 B IDENTICAL both ways
|
||||
```
|
||||
|
||||
The stamp is removed rather than sequenced around — nothing in this repo reads it (no `ReadBuildInfo`
|
||||
caller) and the version comes from the explicit `-X main.version` ldflag. `-trimpath` additionally
|
||||
makes a rebuild from a different checkout directory match.
|
||||
|
||||
**A second discrepancy fell out of it:** `publish-agent.sh`'s fallback build forced `CGO_ENABLED=0` and
|
||||
therefore produced a binary **74 KB smaller** than the release path built for the same version — one
|
||||
version name, two binaries, decided by which entry point was used. Both now build identically.
|
||||
|
||||
`CLAUDE.md` records the exact command an operator can run to verify a published binary independently.
|
||||
|
||||
## v0.121.1 — "nothing is due" must be AUDIBLE (2026-08-03, R-86 + standing rule 3)
|
||||
|
||||
**Found while live-validating v0.121.0, and it is this project's own rule pointed at the change that
|
||||
had just shipped.** Before R-86 every tick ran a heavy restore-test, so the scheduler was audible by
|
||||
construction. After it, *"nothing is due"* is the NORMAL outcome — and it was logged at **DEBUG**,
|
||||
which journald drops. An empty journal would then have been equally consistent with a healthy loop
|
||||
and with a dead goroutine: the exact shape the R-88 watcher was retired for, re-created in a new
|
||||
place by making the quiet path the common one.
|
||||
|
||||
A not-due evaluation now logs one **INFO** line naming every tier's verdict:
|
||||
|
||||
```
|
||||
backup: restore-test evaluated — nothing due
|
||||
verdicts="felhom-pbs: newest settled archive (landed 2026-07-28T04:49:43Z) is already proven;
|
||||
felhom-backup: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
```
|
||||
|
||||
Four lines a day at the 6 h default, and the answer to *"why did nothing run last night?"* is in the
|
||||
log instead of being re-derived. A tier whose storage cannot be listed reads `UNKNOWN` with its error
|
||||
in the same line, so a lookup failure can never present as "nothing due".
|
||||
|
||||
Red-proved by reverting to the bare `Debug` line: the test asserts what the SCHEDULER emits on a real
|
||||
`tick`, not what the helper returns — a helper-level test would have passed against a tick that never
|
||||
called it.
|
||||
|
||||
## v0.121.0 — a restore-test proves each BACKUP, not the clock (2026-08-03, R-86)
|
||||
|
||||
**The trigger changed; the restore-test did not.** `Scheduler.Run` still has a ticker, but it is now
|
||||
|
||||
@@ -72,6 +72,32 @@ internal/storage/ storage observer + durable ids + role/claim classifiers + S
|
||||
> verifies by an **independent download** rather than trusting the publish step's own output.
|
||||
> `scripts/publish-agent.sh` still exists and is still correct — the release script CALLS it rather
|
||||
> than reimplementing it.
|
||||
>
|
||||
> **THE ORDER IS build → tag LOCALLY → publish → push tag, and each step protects something (R-188,
|
||||
> R-186).** The tag is created before the publish so the build and the tag describe the same commit;
|
||||
> it is *pushed* after, because the push is what wakes CI (`on: [push]`) and a tag visible before its
|
||||
> package makes the published-versions gate correctly fail a correct release — it did, on roughly
|
||||
> every second release, and R-168 sends that failure to you by mail. The invariant the old order
|
||||
> protected is asserted directly instead: the gate now also refuses a **published version with no
|
||||
> tag**. If the push fails after a successful publish the script says so loudly and prints the
|
||||
> one-line recovery; if the *publish* fails it removes the local-only tag so a retry is clean.
|
||||
>
|
||||
> **A RELEASED BINARY IS INDEPENDENTLY VERIFIABLE (R-186).** The build uses `-trimpath
|
||||
> -buildvcs=false` so the same source produces the same bytes whether or not the tag exists yet —
|
||||
> before this, a rebuild could not reproduce the sha you were vouching. To check any published
|
||||
> version yourself:
|
||||
>
|
||||
> ```bash
|
||||
> V=0.122.0
|
||||
> git checkout "v$V" && go build -trimpath -buildvcs=false -ldflags "-X main.version=$V" \
|
||||
> -o /tmp/felhom-agent-check ./cmd/felhom-agent
|
||||
> sha256sum /tmp/felhom-agent-check
|
||||
> curl -fsSL "https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/$V/felhom-agent" | sha256sum
|
||||
> ```
|
||||
>
|
||||
> The two hashes must match. `publish-agent.sh`'s fallback build uses the **same** flags — it used to
|
||||
> force `CGO_ENABLED=0` and produce a 74 KB-smaller binary for the same version; if either build line
|
||||
> ever changes, change both or one version name means two binaries again.
|
||||
|
||||
| Step | Where | One-liner |
|
||||
|---|---|---|
|
||||
@@ -79,6 +105,7 @@ internal/storage/ storage observer + durable ids + role/claim classifiers + S
|
||||
| Copy | local → felhom-pve | `scp /tmp/felhom-agent-<v> felhom-pve:/tmp/` (one hop) |
|
||||
| Deploy | felhom-pve | backup `.bak-<old>` → `install -m0755` → `systemctl restart felhom-agent` (non-root `felhom-agent` user, config `/etc/felhom-agent/agent.json`) |
|
||||
| Ship configs | felhom-pve | sudoers (`/etc/sudoers.d/felhom-agent`) + guarded-mkfs wrapper WITH the binary when `configs/` changed |
|
||||
| **Verify** (anyone, any time) | anywhere with the repo + Go | `git checkout v<ver> && go build -trimpath -buildvcs=false -ldflags "-X main.version=<ver>" -o /tmp/a ./cmd/felhom-agent && sha256sum /tmp/a` — must equal `curl -fsSL <pkg-url> \| sha256sum` |
|
||||
| **Vouch** | hub operator UI | Configs → Day-0 artifacts. **Deliberately NOT automated** — vouching is what points machines at a version, and it stays your act (prove-then-vouch) |
|
||||
| Verify | felhom-pve | `felhom-agent --version` + journal (clean ReassertGuestBinds, no capability degradation) |
|
||||
|
||||
|
||||
+33
@@ -5,6 +5,31 @@
|
||||
|
||||
## Current
|
||||
|
||||
- **2026-08-03 — v0.122.0 (R-189 · R-188 · R-186): three signals that lied about their own work.**
|
||||
None touches data; all three cost attention, which every other signal depends on.
|
||||
- **R-189 — a passing restore-test no longer vanishes on a restart.** `restore_tests[]` came only
|
||||
from the in-memory `backup.Store` (*"lost on restart; the cadence re-populates"* — true under a
|
||||
timer, FALSE since R-86, because the agent will not re-test a proven archive). **Observed live:**
|
||||
a 14.5 GB offsite PASS at 15:25:14, agent restarted 2 m 43 s later, hub logged `0 restore-tests`
|
||||
twice. `RestoreTestState` now stores `tier` + `verified` beside the archive (v3 shape; v1/v2
|
||||
still read, and a record missing archive-or-tier is NOT reported), exposes
|
||||
`ProvenRestoreTests`, and `Collector.SetProvenRestoreTests` merges it — **one entry per tier,
|
||||
newest by `TestedAt` wins**, so a fresh failure beats a stored success and a tier never appears
|
||||
twice. Wiring pinned by an AST test: the method this replaces (`Snapshot`) claimed a
|
||||
"host-report gauge" in its doc comment and had **no caller** for weeks.
|
||||
- **ONLY SUCCESSES ARE PERSISTED, and the reason is now in the code:** a success *suppresses*
|
||||
future work (a proven archive is never re-tested, so a lost proof leaves the box quietly less
|
||||
tested than it believes); a failure *causes* future work and heals itself at the next evaluation.
|
||||
- **R-188 — the release stopped emailing false failures.** Only the tag PUSH moved (build → tag
|
||||
locally → publish → push tag): the push is what wakes CI, and a tag visible before its package
|
||||
made the gate correctly fail a correct release ~half the time. The old order's invariant is now
|
||||
asserted directly — `check-published-versions.py` refuses a **published version with no tag**, as
|
||||
a bounded, printed probe (the package listing api is still 401 without a token, re-measured).
|
||||
- **R-186 — a released binary is verifiable.** `-trimpath -buildvcs=false`: same source → same
|
||||
bytes whether or not the tag exists. Measured. `publish-agent.sh`'s fallback also forced
|
||||
`CGO_ENABLED=0` and built a **74 KB different** binary for the same version — both paths now
|
||||
identical. The verification command is in `CLAUDE.md`.
|
||||
|
||||
- **2026-08-03 — v0.121.0 (R-86): the restore-test follows the BACKUP, not the clock.** The ticker is
|
||||
now only the **evaluation interval**; a tier is **DUE** when its newest archive that has settled for
|
||||
`settle` (default 24 h) **has not been proven**. Daily tier → proved daily on yesterday's archive;
|
||||
@@ -26,6 +51,14 @@
|
||||
able to make a starting backup record a failure — F-A1), and the candidate picker skips archives
|
||||
failing `archivePlausiblyComplete` (a phantom would be due forever and fail forever).
|
||||
- New read-only `--selftest=restore-test-due` prints the per-tier verdict + its cost.
|
||||
- **v0.121.1 — a quiet evaluation is AUDIBLE.** "Nothing is due" is now the NORMAL outcome, and at
|
||||
DEBUG it was silent: an empty journal would have been equally consistent with a healthy loop and
|
||||
a dead goroutine (standing rule 3 — the shape the R-88 watcher was retired for). A not-due
|
||||
evaluation logs ONE INFO line naming every tier's verdict; an unlistable tier reads `UNKNOWN`
|
||||
with its error in that same line.
|
||||
- **PROVEN LIVE 2026-08-03 on demo-felhom:** due-triggered offsite restore-test of a 14.5 GB
|
||||
encrypted PBS archive — restored, booted, verified, scratch destroyed, **635 s**; the state then
|
||||
named that archive, a second evaluation ran nothing, and an agent restart ran nothing.
|
||||
- **R-185 (filed, NOT fixed here):** on demo-felhom the agent token has no ACL on
|
||||
`/storage/felhom-backup`, so its content listing comes back EMPTY (root sees 3 archives) — the
|
||||
host tier has never been restore-testable there, and the due-check cannot distinguish that from
|
||||
|
||||
@@ -1,100 +1,317 @@
|
||||
# REPORT — releasing publishes, and an unreleasable version fails CI (R-115, R-183)
|
||||
# REPORT — R-86: a restore-test proves each BACKUP, not the clock
|
||||
|
||||
**Date:** 2026-08-03 · **Repo:** `felhom-agent` · **NO VERSION BUMP** — the agent stays **v0.120.0**,
|
||||
no Go code changed, nothing was built or deployed.
|
||||
**Date:** 2026-08-03 · **Repo:** `felhom-agent` **v0.120.0 → v0.121.0 → v0.121.1**
|
||||
(`4618169`, `4d82591`, `53d0c6b`) ·
|
||||
released, published, verified by independent download, deployed to demo-felhom and **proven live**.
|
||||
Sibling half: `felhom.eu` hub **v0.91.0 → v0.91.1** — the two ship together.
|
||||
|
||||
## What changed
|
||||
---
|
||||
|
||||
| File | |
|
||||
|---|---|
|
||||
| `scripts/release-agent.sh` | **new** — THE release path: build → tag → publish → verify by independent download |
|
||||
| `scripts/check-published-versions.py` | **new** — the R-115 gate |
|
||||
| `scripts/agent_gates.py` | registers the gate as **not `--fast`** (it needs network) |
|
||||
| `.gitea/workflows/gates.yml` | CI now runs the **full** gate set, not `--fast` |
|
||||
| `CLAUDE.md` | the raw `go build` line is replaced by the release script; a **Vouch** row replaces the old Publish row |
|
||||
## 1. Baselines, re-read on arrival
|
||||
|
||||
## Why
|
||||
| Repo | `main` @ commit | Version | Matched §1? |
|
||||
|---|---|---|---|
|
||||
| `felhom-agent` | `1b14cfd0b48b` | `v0.120.0` | **yes** |
|
||||
| `felhom.eu` | `e34b614e5b65` | CHANGELOG `v0.90.0`, deployed image `0.90.1` | **yes — the discrepancy was real and is fixed** (entry backfilled) |
|
||||
|
||||
Publishing was a step someone had to remember and was **forgotten three times in five days** —
|
||||
R-111's seventeen stranded releases, 0.114.0, and 0.120.0, which sat deployed on both demo hosts and
|
||||
undownloadable, so a documented-path reinstall would have silently downgraded them to the pre-merge
|
||||
agent **while reporting success**. R-111's own closing line named this leg and closed SHIPPED without
|
||||
it; it recurred the same afternoon. A note is not a mechanism.
|
||||
Highest register ID in use was **R-184**; `R-185`–`R-187` established free by grep across all four
|
||||
repos and `documentation/`.
|
||||
|
||||
The script also **tags**, because `felhom-host-install.sh` now fetches the agent's sixteen config
|
||||
files from `raw/tag/v<version>/` (R-183). A released version with no tag 404s a box mid-install, as
|
||||
root, on a virgin machine. Tag and package are two halves of one release.
|
||||
## 2. The rule, in one sentence — and the trap it avoids
|
||||
|
||||
It **verifies by downloading what it just published** and comparing the sha to what it built. The
|
||||
publish step's own success is a report on its own write; a fetch returning the right bytes is a
|
||||
different claim, and it is the one that matters.
|
||||
> Let **A** be the newest archive on a tier that has settled for at least the settle lag (24 h).
|
||||
> The tier is **DUE** when A exists and **A has not already been proven**.
|
||||
|
||||
It **does not vouch** — that points machines at a version and stays the operator's act.
|
||||
The literal reading of R-86's own wording — *"due when the newest archive is ≥ 24 h old"* — is
|
||||
**never true on a daily tier**, because a new archive lands each day and resets the newest-archive age
|
||||
to zero long before it reaches the lag. It would have switched restore-testing **off** for the tier
|
||||
that matters most, silently, while looking like the row was implemented.
|
||||
|
||||
## The gate's invariant — not the one specified, and the reason was measured
|
||||
**Evidence that a daily tier does become due**, at three levels:
|
||||
|
||||
The task's §8.4 asked for *"the version the hub tells machines to install must be downloadable"*.
|
||||
**CI cannot see that**, measured rather than assumed (P-C):
|
||||
1. **Unit, time-driven** — `TestDue_DailyTierIsProvedDailyOnItsOwnArchive`: five simulated days, one
|
||||
archive a day, evaluated hourly (120 evaluations) → **exactly 5 runs**, and run *i* tests day
|
||||
*i−1*'s archive, never the still-settling one.
|
||||
2. **The red-proof of the naive rule** — implemented and observed failing at **0 runs over 5 days**
|
||||
(§6), which is the trap made visible rather than argued about.
|
||||
3. **Live** — the offsite tier on demo-felhom became due on its own archive and ran (§7).
|
||||
|
||||
| Endpoint | Anonymous |
|
||||
|---|---|
|
||||
| Gitea package **download** | **200** (and **404** for a fake version — it discriminates) |
|
||||
| Gitea **tags** api | **200** |
|
||||
| Gitea package **listing** api | **401** — token required |
|
||||
| Hub `/api/v1/artifacts/<customer>` | **401** — per-customer passphrase required |
|
||||
## 3. What changed
|
||||
|
||||
So a credential-free gate can ask *"is this version installable"* but not *"which version is
|
||||
vouched"*. Adding an operator credential to CI to close that is the operator's call, not a gate
|
||||
author's. The implemented invariant — **every `v<semver>` tag must have a downloadable package and a
|
||||
tag tree that serves the agent's configs** — needs no credential and **catches all three recorded
|
||||
instances**, because the release script creates the tag and publishes in one act.
|
||||
|
||||
**What it does not catch, stated rather than assumed away:** the hub vouching a version that was
|
||||
never released at all. Nothing here can see that; it belongs at vouch time in the hub. → **R-184**.
|
||||
|
||||
## Proof
|
||||
|
||||
| Check | Result |
|
||||
|---|---|
|
||||
| `go build ./... && go vet ./...` | OK |
|
||||
| `go test ./...` | **29 packages ok, rc=0** (read separately from any commit) |
|
||||
| `agent_gates.py --fast` | `published` correctly **SKIPPED** — the pre-push hook must not fail because Gitea blinked |
|
||||
| `agent_gates.py` (full) | `reuse-refs` OK, `published` OK |
|
||||
| release script: re-release guard | `ERROR: tag v0.120.0 already exists — releasing over it would make one version name two binaries`, rc=1 |
|
||||
| release script: clean-tree guard | `ERROR: working tree is dirty — commit and push first`, rc=1 |
|
||||
|
||||
### Red-proof F — both directions
|
||||
|
||||
- **A tagged-but-unpublished version** (`v9.9.9` created for the purpose): gate **rc=1**,
|
||||
`binary NOT downloadable (HTTP 404 …)`. This is the R-115 shape exactly.
|
||||
- **The gate deregistered from the entry point**, same bad state: `agent_gates.py` → **rc=0, "all
|
||||
agent gates OK"**. Restored → **rc=1, CONVICTED: published**. The guard is what catches it, not
|
||||
something else.
|
||||
|
||||
### Scenario F measured on REAL CI, not inferred
|
||||
|
||||
Runs **69** and **70** are on the **same commit** `0db7766`:
|
||||
|
||||
| run | state of the repo | CI |
|
||||
| Piece | File | Change |
|
||||
|---|---|---|
|
||||
| 69 | no `v9.9.9` | **success** |
|
||||
| 70 | `v9.9.9` tagged, not published | **failure** |
|
||||
| the due-check | `internal/backup/restoretest_due.go` (new) | `EvaluateDue` / `evaluateTier` — per-tier verdict + the reason, ordered oldest-proven first |
|
||||
| the trigger | `internal/backup/schedule.go` | the ticker is now the **evaluation interval**; `pickForThisRun` answers *"is anything due?"*, and "nothing" is a normal answer |
|
||||
| the state | `internal/backup/restoretest_state.go` | records **which archive** was proven, with migration |
|
||||
| the picker | `internal/backup/runner.go` | `PickSettledRestoreCandidateOn(ctx, target, notAfter)`; `PickRestoreCandidateOn` is a one-line call into it |
|
||||
| the knobs | `internal/config/config.go` | `restore_test_eval_interval_seconds` + `restore_test_settle_seconds`; the old key deprecated, not repurposed |
|
||||
| the wiring | `cmd/felhom-agent/main.go` | settle-aware picker + `Settle`; deprecation WARN; new `--selftest=restore-test-due` |
|
||||
| the observable (**v0.121.1**) | `internal/backup/schedule.go`, `restoretest_due.go` | a not-due evaluation logs one **INFO** line naming every tier's verdict — see §9 |
|
||||
|
||||
Same code, same workflow, one variable — so the gate demonstrably RUNS in CI and fails for exactly
|
||||
the R-115 condition. This also retrospectively explains runs 67/68, which were red in the window when
|
||||
`v9.9.9` first existed. **One deliberate CI failure e-mail reached the operator — that was this
|
||||
proof, not an incident.**
|
||||
### The state records the archive (§8.2)
|
||||
|
||||
I could not read CI's own step log to attribute those runs directly: the Gitea jobs endpoint requires
|
||||
an API token, and the only credential available on this host (`~/.docker/config.json`) is a registry
|
||||
password, which the API rejects. The controlled before/after above replaced that log rather than an
|
||||
assumption standing in for it.
|
||||
A timestamp cannot answer *"have we proven **this** archive"* — it is the same class as the workspace
|
||||
rule that a timestamp records an *attempt*, not a *result*: here it records a result, but not **which**
|
||||
result. `RestoreTestState` now holds `{archive, proven_at}` per tier.
|
||||
|
||||
`v9.9.9` was deleted afterwards; `git ls-remote --tags` shows only `v0.120.0`.
|
||||
**Migration:** a pre-R-86 file (`{"target": "<RFC3339>"}`) keeps its **time** — rotation ordering
|
||||
survives a deploy, which is why the file exists at all — and yields **no proven archive**, so each
|
||||
tier is due exactly once after the upgrade. One extra test per tier, once, is the safe direction;
|
||||
reading a legacy time as proof of whatever archive is current would invent a guarantee.
|
||||
|
||||
## Tag convention
|
||||
### The config (§8.3) — and a correction to the spec
|
||||
|
||||
`v<semver>`, at the commit the binary was built from. `v0.120.0` was created retroactively at
|
||||
`cd6e267` — the commit that produced the published binary (sha `a7763d31b55b5ce7…`). `configs/` is
|
||||
byte-identical between that commit and `main`, so nothing about the sixteen fetched files depends on
|
||||
the choice; `cd6e267` is tagged because it is the honest one.
|
||||
The spec said *"`0` must keep meaning disabled"*. **In the code as it stands, `0` means *use the
|
||||
default* and NEGATIVE means disabled** (`RestoreTestCadence`, pre-existing). Making `0` disable would
|
||||
have switched restore-testing off on every box that leaves the key unset — the worst possible reading
|
||||
— so the actual semantics were preserved and this is flagged rather than silently followed.
|
||||
|
||||
- `restore_test_eval_interval_seconds` — how often due-ness is **asked**. Default **6 h**.
|
||||
- `restore_test_settle_seconds` — how long an archive must sit. Default **24 h**.
|
||||
- `restore_test_cadence_seconds` — **deprecated**. Negative still **disables**, verbatim. A positive
|
||||
value seeds the **settle lag** (the quantity a person setting it was expressing: how long may pass
|
||||
between a backup and confidence that it restores), and the daemon logs one start-up WARN naming
|
||||
both replacements. It is deliberately **not** carried into the evaluation interval: a box that set
|
||||
72 h to spare a weak endpoint would otherwise get a 72-hour-latency due-check, whereas what it
|
||||
wanted — fewer heavy restores — is what per-archive due-ness already gives it.
|
||||
|
||||
## 4. Part 1.4 — the measurement, and the interval chosen from it
|
||||
|
||||
Measured on demo-felhom, 2026-08-03, via `--selftest=restore-test-due` and by timing the underlying
|
||||
API call directly (3 runs each):
|
||||
|
||||
| tier | what it is | one due-check |
|
||||
|---|---|---|
|
||||
| `felhom-backup` (dir) | local, on-box | **18 ms** (18.7 / 18.3 / 18.5) |
|
||||
| `felhom-pbs` | offsite, **WAN to ep0** | **392 ms** (375 / 378 / 424) |
|
||||
| both together | one full evaluation | **430 ms** |
|
||||
|
||||
**Cost does not set the interval** — even at one evaluation a minute the offsite leg would be ~0.7 %
|
||||
of the link's time. What sets it is the other bound, and it is not in the brief: **under a per-archive
|
||||
due-check a FAILING tier stays due, so the evaluation interval is also its RETRY interval — and a
|
||||
retry is a multi-GB restore.** Every few minutes would be an incident of its own; the old timer
|
||||
retried a broken tier once a day.
|
||||
|
||||
**6 h chosen from both ends:** at most four heavy retries a day in the worst case, and at most 6 h of
|
||||
latency between an archive settling and its proof — negligible against a 24 h settle lag, so a daily
|
||||
tier is still proved daily. No second rate limiter was added (§8.4): the pacing remains one test per
|
||||
archive generation.
|
||||
|
||||
## 5. Two hazards the new frequency created, and their fixes
|
||||
|
||||
Both are consequences of evaluating often rather than daily, and neither is in the brief:
|
||||
|
||||
1. **The due-check now runs BEFORE the heavy-operation gate is taken.** Holding that gate for a read
|
||||
that answers "nothing to do" would open a window at *every* evaluation in which a starting backup
|
||||
cannot acquire — and a backup that cannot acquire does not merely wait, it **records a failure and
|
||||
pages the operator** (F-A1). Nothing heavy starts before the gate; due-ness does not expire while
|
||||
we check.
|
||||
2. **The candidate picker skips implausible archives.** Under per-archive due-ness an incomplete
|
||||
1-byte phantom (F-CRIT-2's artefact, which server-side prune does **not** collect) would be picked
|
||||
forever, fail forever, never earn proof, and leave the tier due at *every* evaluation — turning the
|
||||
evaluation interval into the retry rate for a multi-GB restore. `archivePlausiblyComplete` (the
|
||||
canonical helper, with its warn-once companion) is applied in the shared scan, so both callers
|
||||
agree. **This is a behaviour change to `PickRestoreCandidateOn`** and is recorded as such.
|
||||
|
||||
## 6. Tests and red-proofs
|
||||
|
||||
Green gate, both repos: `go build ./... && go vet ./... && go test ./...` — agent **29 packages ok,
|
||||
rc=0**; hub **rc=0**. The test run and the commit were always separate commands.
|
||||
|
||||
| # | Test | Asserts | Mutation | Observed |
|
||||
|---|---|---|---|---|
|
||||
| A | `TestDue_DailyTierIsProvedDailyOnItsOwnArchive` | 5 runs over 5 days, each on the settled archive | the naive rule (`now-landed >= settle`, proven-archive check deleted, cutoff removed) | **FAIL** — `a daily tier must be proved once per day; got 0 run(s) over 5 days: []` |
|
||||
| B | `TestDue_WeeklyTierIsProvedOncePerArchive` | 84 evaluations over 3 weeks → exactly 3 runs, one per archive | — | pass |
|
||||
| C | `TestDue_RestartRunsNothing` | two restarts + 4 evaluations → **0 runs** | `ProvenArchive` reverted to per-tier time | **FAIL** — `2 restart(s) produced 4 run(s)` |
|
||||
| D | `TestDue_NewSettledArchiveMakesAProvedTierDueAgain` | a newly settled archive re-arms the tier, and the NEW archive is tested | — | pass |
|
||||
| E | `TestDue_FailingTierIsRetriedAndNeverProven` | 3 evaluations → 3 retries, no proof recorded | credit on failure (`rt.Pass &&` dropped) | **FAIL** — `got 1 run(s) over 3 evaluations` + `TestRotation_FailureEarnsNoCredit` also failed |
|
||||
| F | `TestDue_TwoDueTiersRunOneAtATime` | one evaluation → one run; the other is deferred and runs next | — | pass |
|
||||
| F′ | `TestDue_DeferredBehindABackupStaysDue` | the gate holds; a deferred tier stays DUE | — | pass |
|
||||
| H | `TestDue_NewbornTierIsNotDueAndNotAnError` | no archive → not due, no error, **and a reason** | — | pass |
|
||||
| — | `TestDue_UnsettledArchiveIsNotACandidate` | a 2 h-old archive is not a candidate under a 24 h lag | — | pass |
|
||||
| — | `TestDue_LookupFailureIsUnknownNotNotDue` | a tier that cannot be listed is UNKNOWN, the error travels, the other tier still runs | — | pass |
|
||||
| — | `TestRestoreTestState_LegacyFileMigratesToNothingProven` | legacy time kept, no archive claimed | — | pass |
|
||||
| — | `TestPickRestoreCandidate_SkipsImplausibleArchives` | the newest entry is not a candidate if it cannot be complete | guard removed | **FAIL** — `pick = "phantom" want the newest COMPLETE archive 'real'` |
|
||||
| I | `TestMainWiresTheSettleAwareTierPicker` | **AST** of `main.go`: settle picker wired, old picker gone, `Settle` set, eval-interval accessor called | the wiring line commented out | **FAIL** — `main.go never passes runner.PickSettledRestoreCandidateOn …` (a `strings.Contains` check would have PASSED — the string is still there, in a comment) |
|
||||
| I′ | `TestMainStillWiresTheHeavyOperationGateAndPerRunSpec` | R-85's gate + per-run spec survive | — | pass |
|
||||
|
||||
Hub-side (Scenario G) is in `felhom.eu/REPORT.md`, including **a hollow test caught by its own
|
||||
red-proof**: the first weekly fixture had no jitter, sat on exactly 168 h, and PASSED under the
|
||||
flat-window mutation.
|
||||
|
||||
### Tests deliberately changed, and why
|
||||
|
||||
`TestRotation_BothTiersExercisedAcrossCadences` asserted *4 ticks → 4 runs*. That was a faithful
|
||||
statement of the defect — every tick produced a heavy restore-test, because the ticker **was** the
|
||||
trigger. It is now `TestRotation_BothTiersExercisedOncePerArchive`: **2 runs across 4 evaluations**,
|
||||
one per tier, one per archive. Strictly stronger — it pins both the coverage R-85 won and the pacing
|
||||
R-86 adds. The old assertion is quoted in the test's comment so the change is legible.
|
||||
|
||||
## 7. The live run — triggered by due-ness, on real hardware
|
||||
|
||||
Deployed to **demo-felhom** (Tier 0). The deployed binary is the **published artifact downloaded from
|
||||
Gitea**, not a local rebuild — see R-186.
|
||||
|
||||
### 7.1 The due verdict, per tier, before anything ran
|
||||
|
||||
```
|
||||
eval_interval=6h0m0s settle=24h0m0s
|
||||
tier=felhom-backup due=false archive="" landed=- proven=""
|
||||
reason: no settled archive yet — nothing to prove (newborn or still settling)
|
||||
tier=felhom-pbs due=true archive="felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z"
|
||||
landed=2026-07-28T04:49:43Z proven=""
|
||||
reason: newest settled archive (landed 2026-07-28T04:49:43Z) has not been proven; nothing proven yet
|
||||
```
|
||||
|
||||
`felhom-backup` reads "no settled archive" for a reason that is **not** the one it appears to be —
|
||||
see **R-185**: the agent cannot list that storage at all.
|
||||
|
||||
### 7.2 A real run, started by the due-check
|
||||
|
||||
Only the **evaluation interval** was shortened for the validation (a systemd drop-in, since removed):
|
||||
the due rule, the settle lag and the restore-test itself were untouched.
|
||||
|
||||
```
|
||||
15:14:38 backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)
|
||||
target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z
|
||||
landed=2026-07-28T04:49:43Z
|
||||
reason="newest settled archive … has not been proven; nothing proven on this tier yet"
|
||||
15:14:39 restore-test: full-fidelity restore params derived from the archive config scratch=990000
|
||||
… proxmox-backup-client restore --crypt-mode=encrypt … (felhom-agent@pve!agent)
|
||||
15:25:08 audit: gate decision class=guest_destroy guest=990000 source=one_shot_job allowed=true
|
||||
15:25:14 restore-test: scratch guest torn down vmid=990000
|
||||
15:25:14 backup: scheduled restore-test PASSED archive=felhom-pbs:… duration_s=635.1
|
||||
```
|
||||
|
||||
A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored into a scratch guest,
|
||||
booted, verified and destroyed — **635 s**, unattended, and started by *"this archive has not been
|
||||
proven"* rather than by a timer.
|
||||
|
||||
### 7.3 The state now names that archive, and a second evaluation runs nothing
|
||||
|
||||
```json
|
||||
{ "felhom-pbs": { "archive": "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z",
|
||||
"proven_at": "2026-08-03T13:25:14Z" } }
|
||||
```
|
||||
|
||||
```
|
||||
tier=felhom-pbs due=false proven="felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z"
|
||||
reason: newest settled archive (landed 2026-07-28T04:49:43Z) is already proven
|
||||
```
|
||||
|
||||
### 7.4 Teardown — all three layers
|
||||
|
||||
| Layer | Before | After |
|
||||
|---|---|---|
|
||||
| scratch guest 990000 | `stopped lock=create` during the run | **absent** from `pct list` |
|
||||
| its volumes | 5 thin LVs (32 G + 200 G + 50 G + 2×1 G) | **0** matches in `lvs` |
|
||||
| hub-side record | — | the run's **`restore_tests[]` entry is RETAINED deliberately** — it is the proof the hub's staleness check reads, and deleting it would delete the result. No event was created: the run passed, and `restore_test_failed`/`restore_test_stale` fire only on failure or staleness |
|
||||
|
||||
`pvesm status` before and after: `local-lvm` 1.95 % used before the run, and the thin volumes are gone
|
||||
after it — the restore reclaimed to the same shape it started in.
|
||||
|
||||
### 7.5 The restart, which is the defect a person would actually notice
|
||||
|
||||
Under the old scheduler every agent deploy restarted the ticker, so a restore-test ran one interval
|
||||
after each deploy regardless of what had been proven. The proof of the fix cannot be *"nothing
|
||||
appeared in the log"* — that is the absent-line trap this project has a standing rule about — so
|
||||
v0.121.1 makes a quiet evaluation say what it decided, and the evaluation interval was shortened to
|
||||
2 min for the validation so evaluations are **observable**, not assumed:
|
||||
|
||||
```
|
||||
15:32:10 felhom-agent daemon starting version=0.121.1 ← the restart
|
||||
15:32:11 backup: restore-test scheduler starting (per-archive due-check) eval_interval=2m0s settle=24h0m0s
|
||||
15:34:12 backup: restore-test evaluated — nothing due
|
||||
verdicts="felhom-backup: no settled archive yet — nothing to prove (newborn or still settling);
|
||||
felhom-pbs: newest settled archive (landed 2026-07-28T04:49:43Z) is already proven"
|
||||
15:36:12 backup: restore-test evaluated — nothing due (same verdicts)
|
||||
|
||||
runs since the restart: 0
|
||||
```
|
||||
|
||||
Evaluations demonstrably **happened** and demonstrably **decided**; nothing ran. The 6 h default was
|
||||
restored afterwards (§8).
|
||||
|
||||
|
||||
## 8. Release and deployment
|
||||
|
||||
| Step | Evidence |
|
||||
|---|---|
|
||||
| Released via `scripts/release-agent.sh 0.121.0` | tag `v0.121.0` at `4d82591`, package published |
|
||||
| Verified by **independent download** | sha256 `b2128f3cd4539225a2842f541f56ffaf5390b1d97f3f3a80076ec5f53dbc7d7a`, 14 081 336 B, round-trip GET matched |
|
||||
| Gate re-run after release | `2 released version(s) to verify: 0.120.0, 0.121.0` → both **installable** |
|
||||
| Deployed | `felhom-agent --version` → **0.121.0**, `systemctl is-active` → **active**, prior binary kept as `.bak-0.120.0` |
|
||||
| Startup | `capabilities self-check ok=68 total=68 degraded=0`, and `backup: restore-test scheduler starting (per-archive due-check) eval_interval=6h0m0s settle=24h0m0s` |
|
||||
| Second release, same path | **v0.121.1** — tag `v0.121.1`, sha256 `afaeeb509d1ed70d6e6bebac0393a3cd5be59d51e3db9ff96ef8524bd78546d7`, round-trip verified |
|
||||
| Deployed (published bytes again) | `felhom-agent --version` → **0.121.1**, `active`; prior kept as `.bak-0.121.0` |
|
||||
| Validation config removed | the 2-min drop-in deleted; the daemon back on **`eval_interval=6h0m0s settle=24h0m0s`** |
|
||||
| Fleet | demo-hp still runs **0.120.0** — deliberate: pointing machines at a version is what **vouching** does |
|
||||
| **Vouching** | **NOT done — deliberately the operator's act.** Hub UI → Configs → Day-0 artifacts: agent **`0.121.1`**, sha `afaeeb50…` (0.121.0 also published, sha `b2128f3c…`) |
|
||||
| Config compatibility checked on both boxes | `restore_test_cadence_seconds = 0` on demo-felhom **and** demo-hp, and the installer writes `0` — so no box is on the deprecation path, and a fresh install gets the new defaults with no installer change |
|
||||
|
||||
## 9. Findings filed (none fixed blind)
|
||||
|
||||
- **R-185 — the agent cannot see demo-felhom's host backup tier at all.** The PVE token has no ACL on
|
||||
`/storage/felhom-backup`, so the content listing returns `{"data":[]}` where root sees three
|
||||
archives (6.1–6.3 GB, 08-01/02/03). Verified three ways, including `local` — which *has* a grant —
|
||||
returning its archives through the same token. **Pre-existing and independent of R-86** (R-85's
|
||||
rotation had the same blindness). The part worth fixing is the **silence**: a permission-blinded
|
||||
tier is today indistinguishable from a newborn one, and the agent already records the backups it
|
||||
wrote to that target, so the contradiction is detectable.
|
||||
- **R-186 — a released binary's sha cannot be reproduced from its tag.** `release-agent.sh` builds
|
||||
before tagging, so Go stamps a pseudo-version into the published bytes: published `b2128f3c…`
|
||||
(14 081 336 B) vs rebuild-at-tag `8302e396…` (14 077 240 B), identical source and toolchain. The
|
||||
build order is deliberate, so the fix is not to swap the steps blind. **Mitigated here** by
|
||||
deploying the published artifact.
|
||||
- **R-187 — R-115's one-command release had never run its publish leg** (`CLOSED`, fixed in the same
|
||||
session): `publish-agent.sh` has been mode `0644` since 2026-06-28 because every earlier caller used
|
||||
`bash …`, and `release-agent.sh` called it directly → `Permission denied` on the first real release.
|
||||
Fixed both ways: the mode bit restored **and** the call made mode-independent. The tag the failed
|
||||
run created was withdrawn (nothing had been published under it — verified 404) and recreated on the
|
||||
fix commit, so one version name still means one binary.
|
||||
|
||||
- **R-188 — a correct agent release emails a CI failure about half the time.** `on: [push]` fires the
|
||||
gates workflow on the **tag** push too, and `release-agent.sh` pushes the tag before publishing
|
||||
(deliberately). CI can therefore run the published-versions gate inside the window where the tag
|
||||
exists and the package does not, and correctly report *"every released agent version must be
|
||||
INSTALLABLE"* for a release that completes seconds later. **Measured across two releases in one
|
||||
session:** v0.121.0 → runs #12 success / #13 failure on the same sha; v0.121.1 → #17 failure / #18
|
||||
success on the same sha; and one pair both green — a race, not a rule. It matters because R-168
|
||||
made CI email on failure so a red gate cannot be missed; a signal that cries wolf on every second
|
||||
correct release is how that mail becomes something you archive unread.
|
||||
- **R-189 — a passing restore-test can be invisible to the hub, and R-86 widened that window.**
|
||||
Observed live: **the 15:25:14 PASS reached no host-report at all**. `restore_tests[]` comes from an
|
||||
**in-memory** store (*"lost on restart; the cadence re-populates"*) and the report interval is
|
||||
900 s; the agent was restarted 2 m 43 s after the run for the v0.121.1 deploy. That used to
|
||||
self-heal within 24 h because the next cadence re-tested the tier — **under per-archive due-ness
|
||||
the agent will not re-test a proven archive**, so the hub can stay ignorant until the next archive
|
||||
generation. The persisted proof already exists: `RestoreTestState.Snapshot()` is documented *"for
|
||||
the host-report gauge"* and has **no production caller** — a seam built and never wired, and an
|
||||
invariant asserted only in a comment, in one method. Bounded, not over-ranked: the hub scans its
|
||||
retained window and the offsite tier's window (12 d) is wider than its archive rhythm (7 d), so one
|
||||
lost report is tolerated. Filed, not fixed — it is a report-contract change.
|
||||
|
||||
## 10. CI
|
||||
|
||||
| Repo | Run | Commit | Result |
|
||||
|---|---|---|---|
|
||||
| `felhom-agent` | **#15** (id 83) | `4d82591` | **success** |
|
||||
| `felhom-agent` | #13 (id 81) | `4618169` | **failure — explained, and it is CI doing its job** |
|
||||
| `felhom.eu` | **#48** (id 86) | `ff2655c` | **success** |
|
||||
|
||||
Run #13 fired on the **tag push** from the *failed* first release: `v0.121.0` existed as a tag while
|
||||
nothing was published, and `check-published-versions.py` correctly refused — *"every released agent
|
||||
version must be INSTALLABLE"*. That is precisely the state R-115's gate exists to catch, caught within
|
||||
minutes and self-resolved by the corrected release. Confirmed locally afterwards: both 0.120.0 and
|
||||
0.121.0 verify. `--no-verify` was **not** used anywhere.
|
||||
|
||||
## 11. Observations — noticed, recorded, not acted on
|
||||
|
||||
- **`felhom.eu/CONTEXT.md` has duplicate standing-ruling IDs** — three `S-14`s and two `S-15`s already
|
||||
in the file before this session. New rulings were numbered **S-17/S-18** rather than adding to the
|
||||
collision; the existing duplicates are untouched.
|
||||
- **`agent_gates.py --fast` skips the published-versions gate**, so the pre-push hook cannot catch an
|
||||
unpublished release — only CI can. That is the intended split (no network in a hook), and it is why
|
||||
run #13 mattered.
|
||||
- **The hub sweeps every 60 s and re-reads 14 days of host-reports per customer** for this check. Not
|
||||
changed here (the read window is the same as before), but it is the cost centre if the fleet grows.
|
||||
|
||||
@@ -148,7 +148,8 @@
|
||||
| `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go |
|
||||
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
|
||||
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
|
||||
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,t)` / `ProvenArchive(target)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. |
|
||||
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
|
||||
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
|
||||
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
|
||||
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
|
||||
| `localapi.StaleLockController` | internal/localapi/stalelock.go | `*staleLockController` (Client + Runner + pool) | `fakeStaleLock` (Server-level) stalelock_test.go; `fakeStaleLockAPI` (controller-level, tests the A1 pool intersect) stalelock_pool_test.go |
|
||||
|
||||
@@ -662,6 +662,12 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
|
||||
heavyOps := &backup.InFlight{}
|
||||
scheduler := buildRestoreTestScheduler(cfg, px, engine, backupStore, rtState, heavyOps, logger)
|
||||
// R-189: the host report's restore_tests[] must survive an agent restart. The in-memory store
|
||||
// holds only this process's latest run, and under per-archive due-ness the agent will not
|
||||
// re-test an archive it has already proven — so without this the hub can report a tier unproven
|
||||
// for a whole archive generation after a deploy. Observed live on 2026-08-03: a passing 14.5 GB
|
||||
// offsite restore-test reached no host-report at all.
|
||||
collector.SetProvenRestoreTests(rtState)
|
||||
|
||||
// PBS verify loop (slice 6 Phase B): the fifth daemon goroutine. Cheap, key-free,
|
||||
// ciphertext-level integrity check on its own cadence (default 6h), reporting per-snapshot
|
||||
|
||||
@@ -111,3 +111,46 @@ func parseMainForWiring(t *testing.T) *ast.File {
|
||||
}
|
||||
return f
|
||||
}
|
||||
|
||||
// R-189 Scenario I — the DURABLE proof source must actually be wired into the collector.
|
||||
//
|
||||
// This test exists because the method it feeds is the project's own cautionary tale:
|
||||
// `RestoreTestState.Snapshot` carried the doc comment "for the host-report gauge" from the day it
|
||||
// was written and **had no caller at all** — a seam built, documented and never connected, found
|
||||
// only when a live restore-test's PASS reached no host-report. The fix must not become the next
|
||||
// instance, so the wiring is asserted rather than trusted.
|
||||
//
|
||||
// AST, not grep: a commented-out call still contains the string (proven yesterday, when commenting
|
||||
// out the tier-picker line failed this test while a `strings.Contains` check would have passed).
|
||||
func TestMainWiresTheDurableRestoreTestProof(t *testing.T) {
|
||||
f := parseMainForWiring(t)
|
||||
|
||||
var wired, feedsState bool
|
||||
ast.Inspect(f, func(n ast.Node) bool {
|
||||
call, ok := n.(*ast.CallExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
sel, ok := call.Fun.(*ast.SelectorExpr)
|
||||
if !ok || sel.Sel.Name != "SetProvenRestoreTests" {
|
||||
return true
|
||||
}
|
||||
wired = true
|
||||
// ...and it must be fed the PERSISTED state, not the in-memory store.
|
||||
if len(call.Args) == 1 {
|
||||
if id, ok := call.Args[0].(*ast.Ident); ok && id.Name == "rtState" {
|
||||
feedsState = true
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
|
||||
if !wired {
|
||||
t.Error("main.go never calls collector.SetProvenRestoreTests — the persisted proof would never " +
|
||||
"reach the hub, which is the R-189 defect exactly: a passing restore-test that vanishes on restart")
|
||||
}
|
||||
if wired && !feedsState {
|
||||
t.Error("collector.SetProvenRestoreTests is not fed rtState — the in-memory store is the thing " +
|
||||
"that does NOT survive a restart, so wiring it here would fix nothing")
|
||||
}
|
||||
}
|
||||
|
||||
@@ -140,3 +140,29 @@ func (s *Scheduler) evaluateTier(ctx context.Context, target string, cutoff time
|
||||
func (s *Scheduler) EvaluateDueTier(ctx context.Context, target string) DueVerdict {
|
||||
return s.evaluateTier(ctx, target, s.settleCutoff())
|
||||
}
|
||||
|
||||
// verdictSummary renders one compact line of per-tier verdicts for the "nothing due" log.
|
||||
//
|
||||
// It re-evaluates rather than threading the verdicts out of pickForThisRun, and that is a
|
||||
// deliberate trade: this runs only on the path where NOTHING is due, so the cost is one extra
|
||||
// storage listing per tier on an otherwise idle evaluation (measured 18 ms local / 392 ms offsite,
|
||||
// R-86 Part 1.4), and in exchange the logging path cannot drift from the deciding path by holding a
|
||||
// stale copy of it. If that cost ever matters, pass the verdicts in — do not let the two diverge.
|
||||
func (s *Scheduler) verdictSummary(ctx context.Context) string {
|
||||
out := ""
|
||||
for _, v := range s.EvaluateDue(ctx) {
|
||||
if out != "" {
|
||||
out += "; "
|
||||
}
|
||||
switch {
|
||||
case v.Err != nil:
|
||||
out += v.Target + ": UNKNOWN (" + v.Err.Error() + ")"
|
||||
default:
|
||||
out += v.Target + ": " + v.Reason
|
||||
}
|
||||
}
|
||||
if out == "" {
|
||||
return "no tiers configured"
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
@@ -4,8 +4,10 @@ import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
@@ -124,8 +126,8 @@ func dailyArchives(tier string, n int) []archiveStub {
|
||||
// COMPANION RED-PROOF (observed 2026-08-03). In Scheduler.evaluateTier, the per-archive comparison
|
||||
// was replaced by the naive age rule:
|
||||
//
|
||||
// - if ok && proven == archive { … not due … }
|
||||
// + if s.now().Sub(landed) < s.settle { … not due … } // and the proven-archive check deleted
|
||||
// - if ok && proven == archive { … not due … }
|
||||
// - if s.now().Sub(landed) < s.settle { … not due … } // and the proven-archive check deleted
|
||||
//
|
||||
// and the picker cutoff was removed (`cutoff := time.Time{}`), i.e. exactly "is the newest archive
|
||||
// old enough". Result:
|
||||
@@ -194,12 +196,13 @@ func TestDue_WeeklyTierIsProvedOncePerArchive(t *testing.T) {
|
||||
// COMPANION RED-PROOF (observed 2026-08-03): revert the state to per-tier TIME by making
|
||||
// ProvenArchive ignore the stored archive —
|
||||
//
|
||||
// - if !ok || p.Archive == "" { return "", false }
|
||||
// + return "", false // per-tier time only, the pre-R-86 state
|
||||
// - if !ok || p.Archive == "" { return "", false }
|
||||
// - return "", false // per-tier time only, the pre-R-86 state
|
||||
//
|
||||
// → --- FAIL: TestDue_RestartRunsNothing
|
||||
// restoretest_due_test.go:226: an agent restart must not trigger a restore-test; 2 restart(s)
|
||||
// produced 4 run(s)
|
||||
//
|
||||
// restoretest_due_test.go:226: an agent restart must not trigger a restore-test; 2 restart(s)
|
||||
// produced 4 run(s)
|
||||
//
|
||||
// Four: the same already-proven archive re-tested on EVERY evaluation after EVERY restart, which is
|
||||
// today's behaviour with the ticker's phase reset by the deploy. Restored.
|
||||
@@ -256,12 +259,13 @@ func TestDue_NewSettledArchiveMakesAProvedTierDueAgain(t *testing.T) {
|
||||
//
|
||||
// COMPANION RED-PROOF (observed 2026-08-03): give credit on failure in Scheduler.tick —
|
||||
//
|
||||
// - if rt.Pass && s.rtState != nil && target != "" {
|
||||
// + if s.rtState != nil && target != "" {
|
||||
// - if rt.Pass && s.rtState != nil && target != "" {
|
||||
// - if s.rtState != nil && target != "" {
|
||||
//
|
||||
// → --- FAIL: TestDue_FailingTierIsRetriedAndNeverProven
|
||||
// restoretest_due_test.go: a failing tier must keep being retried; got 1 run(s) over 3
|
||||
// evaluations
|
||||
//
|
||||
// restoretest_due_test.go: a failing tier must keep being retried; got 1 run(s) over 3
|
||||
// evaluations
|
||||
//
|
||||
// A single failure would have retired the archive as proven — a permanently broken DR tier looking
|
||||
// freshly verified, which is the loudest signal this system produces going silent. Restored.
|
||||
@@ -438,7 +442,7 @@ func TestRestoreTestState_ArchiveRoundTrips(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "rt.json")
|
||||
now := time.Now().UTC().Truncate(time.Second)
|
||||
st := NewRestoreTestState(path)
|
||||
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", now); err != nil {
|
||||
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", "pbs", "boot+running", now); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
re := NewRestoreTestState(path)
|
||||
@@ -456,3 +460,134 @@ func TestRestoreTestState_ArchiveRoundTrips(t *testing.T) {
|
||||
func writeFileForTest(path, content string) error {
|
||||
return os.WriteFile(path, []byte(content), 0o600)
|
||||
}
|
||||
|
||||
// Standing rule 3: an absent log line is not evidence. "Nothing is due" is now the NORMAL outcome of
|
||||
// an evaluation, so it must produce a POSITIVE observable naming each tier's verdict — otherwise a
|
||||
// quiet journal is equally consistent with a healthy loop and a dead goroutine.
|
||||
//
|
||||
// COMPANION RED-PROOF (observed 2026-08-03): drop the summary back to a bare
|
||||
// `s.logger.Debug("backup: restore-test not due this evaluation")` and this fails with
|
||||
// "a not-due evaluation must name each tier's verdict; got \"\"" — i.e. nothing at INFO at all.
|
||||
func TestDue_NothingDueStillNamesEveryTiersVerdict(t *testing.T) {
|
||||
ts := &tierStorage{archives: map[string][]archiveStub{
|
||||
"local": {{volid: "local:backup/a.tar.zst", landed: day0}},
|
||||
"felhom-pbs": nil, // no archive at all
|
||||
}}
|
||||
h := newDueHarness(t, day0.AddDate(0, 0, 1), 24*time.Hour, true, []string{"local", "felhom-pbs"}, ts)
|
||||
// Prove the local tier so NOTHING is due.
|
||||
if err := h.st.RecordSuccess("local", "local:backup/a.tar.zst", "local", "boot+running", h.clock); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// Assert what the SCHEDULER emits on a real evaluation, not what a helper returns — a helper
|
||||
// test would pass against a tick that never calls it.
|
||||
var logbuf strings.Builder
|
||||
h.s.logger = slog.New(slog.NewTextHandler(&logbuf, &slog.HandlerOptions{Level: slog.LevelInfo}))
|
||||
h.s.tick(context.Background())
|
||||
got := logbuf.String()
|
||||
for _, want := range []string{"local", "felhom-pbs", "already proven", "no settled archive"} {
|
||||
if !strings.Contains(got, want) {
|
||||
t.Fatalf("a not-due evaluation must name each tier's verdict; got %q (missing %q)", got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A tier whose storage cannot be listed must say UNKNOWN in that same line — a lookup failure that
|
||||
// reads as "nothing due" is the silence this rule exists to prevent.
|
||||
func TestDue_VerdictSummaryNamesAnUnknownTier(t *testing.T) {
|
||||
ts := &tierStorage{
|
||||
archives: map[string][]archiveStub{"local": nil},
|
||||
err: map[string]error{"felhom-pbs": errors.New("storage unreachable")},
|
||||
}
|
||||
h := newDueHarness(t, day0, 24*time.Hour, true, []string{"local", "felhom-pbs"}, ts)
|
||||
got := h.s.verdictSummary(context.Background())
|
||||
if !strings.Contains(got, "UNKNOWN") || !strings.Contains(got, "storage unreachable") {
|
||||
t.Fatalf("an unlistable tier must read as UNKNOWN with its error; got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// ── R-189 — the persisted proof must be REPORTABLE, and must refuse to lie ───────────────────
|
||||
//
|
||||
// A proof held only in the in-memory store dies with the process, and under per-archive due-ness the
|
||||
// agent will not repeat the work. So the persisted record has to be able to become a host-report
|
||||
// entry — without inventing anything it does not know.
|
||||
//
|
||||
// COMPANION RED-PROOF (observed 2026-08-03): drop the `reportable()` filter from
|
||||
// ProvenRestoreTests, so a pre-R-189 record (archive but no tier) is emitted →
|
||||
//
|
||||
// --- FAIL: TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe
|
||||
// restoretest_due_test.go: a record with no TIER must not be reported (the hub keys its
|
||||
// per-tier proof on it); got [{... SourceTier: ...}]
|
||||
//
|
||||
// Restored.
|
||||
func TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "rt.json")
|
||||
// v1 (a bare time), v2 (archive, no tier) and v3 (complete) side by side — every shape this
|
||||
// file has ever had, which is what a real box carries after two upgrades.
|
||||
legacy := `{
|
||||
"old-v1": "2026-07-30T02:11:07Z",
|
||||
"old-v2": {"archive":"felhom-backup:backup/vzdump-lxc-9201-a.tar.zst","proven_at":"2026-08-01T04:41:58Z"},
|
||||
"felhom-pbs": {"archive":"felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z","tier":"pbs","verified":"boot+running","proven_at":"2026-08-03T13:25:14Z"}
|
||||
}`
|
||||
if err := writeFileForTest(path, legacy); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
got := NewRestoreTestState(path).ProvenRestoreTests(context.Background())
|
||||
if len(got) != 1 {
|
||||
t.Fatalf("only the record that can be described honestly may be reported; got %d: %+v", len(got), got)
|
||||
}
|
||||
e := got[0]
|
||||
if e.SourceTier != "pbs" {
|
||||
t.Fatalf("a record with no TIER must not be reported (the hub keys its per-tier proof on it); got %+v", got)
|
||||
}
|
||||
if e.SourceArchive != "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" || !e.Pass {
|
||||
t.Fatalf("the reported entry must be the stored proof, unchanged; got %+v", e)
|
||||
}
|
||||
if e.TestedAt != "2026-08-03T13:25:14Z" {
|
||||
t.Fatalf("the entry must carry the time the run passed, not now(); got %q", e.TestedAt)
|
||||
}
|
||||
if e.Verified != "boot+running" {
|
||||
t.Fatalf("what the run verified must survive the round trip; got %q", e.Verified)
|
||||
}
|
||||
// Run mechanics are NOT invented: an absent duration is not a claim, a fabricated one would be.
|
||||
if e.DurationSeconds != 0 || e.ScratchVMID != 0 {
|
||||
t.Fatalf("the re-report must not invent run mechanics it never stored; got duration=%v scratch=%d",
|
||||
e.DurationSeconds, e.ScratchVMID)
|
||||
}
|
||||
// The legacy records still serve the DUE-check, which is a separate question from reporting.
|
||||
if _, ok := NewRestoreTestState(path).ProvenArchive("old-v2"); !ok {
|
||||
t.Fatal("a v2 record must still answer the due-check even though it cannot be reported")
|
||||
}
|
||||
}
|
||||
|
||||
// A tier proved through the SCHEDULER (not by hand) lands in the state complete enough to report —
|
||||
// the production path, not a hand-built fixture.
|
||||
func TestScheduler_ProofIsRecordedReportably(t *testing.T) {
|
||||
ts := &tierStorage{archives: map[string][]archiveStub{"felhom-pbs": {{volid: "felhom-pbs:backup/ct/9201/w0", landed: day0}}}}
|
||||
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(97*time.Minute), 24*time.Hour, true, []string{"felhom-pbs"}, ts)
|
||||
// The fake runner echoes the spec's tier; give the spec a tier the way main.go does.
|
||||
h.s.spec = func(_ context.Context, archive string) reconcile.RestoreTestSpec {
|
||||
return reconcile.RestoreTestSpec{RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009, SourceTier: "pbs"}
|
||||
}
|
||||
h.s.tick(context.Background())
|
||||
|
||||
got := h.st.ProvenRestoreTests(context.Background())
|
||||
if len(got) != 1 {
|
||||
t.Fatalf("a scheduled pass must leave a REPORTABLE proof; got %d: %+v", len(got), got)
|
||||
}
|
||||
if got[0].SourceTier != "pbs" || got[0].SourceArchive != "felhom-pbs:backup/ct/9201/w0" {
|
||||
t.Fatalf("the proof must name the tier and the archive the run used; got %+v", got[0])
|
||||
}
|
||||
}
|
||||
|
||||
// A FAILED run leaves nothing to report — the asymmetry of §8.1, asserted rather than assumed.
|
||||
func TestScheduler_AFailureLeavesNoPersistedProof(t *testing.T) {
|
||||
ts := &tierStorage{archives: map[string][]archiveStub{"felhom-pbs": {{volid: "felhom-pbs:backup/ct/9201/w0", landed: day0}}}}
|
||||
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(97*time.Minute), 24*time.Hour, false, []string{"felhom-pbs"}, ts)
|
||||
h.s.tick(context.Background())
|
||||
if got := h.st.ProvenRestoreTests(context.Background()); len(got) != 0 {
|
||||
t.Fatalf("a FAILED run must persist nothing — a failing tier is retried, and a stored failure "+
|
||||
"would outlive the fault; got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,12 +1,15 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// RestoreTestState persists the last SUCCESSFUL restore-test per backup tier.
|
||||
@@ -49,16 +52,45 @@ type RestoreTestState struct {
|
||||
last map[string]provenTier // target id → what was last PROVEN on that tier
|
||||
}
|
||||
|
||||
// provenTier is one tier's proof: the archive that passed, and when it passed.
|
||||
// provenTier is one tier's proof: the archive that passed, which tier it was, what was verified,
|
||||
// and when.
|
||||
//
|
||||
// R-189 added `Tier` and `Verified`. Until then this record could answer the DUE-check but could not
|
||||
// be REPORTED, and being reportable is what closes R-189: a proof held only in the in-memory result
|
||||
// store vanishes on restart, and under per-archive due-ness the box will not repeat the work, so the
|
||||
// hub can stay ignorant of a real success until the next archive generation.
|
||||
//
|
||||
// `Tier` is stored rather than derived because it is known for certain at proof time (the run's own
|
||||
// spec used it to choose the restore timeout) and deriving it later would need a storage-type lookup
|
||||
// at report-building time — a network call that can fail, on a path where failing means mis-labelling
|
||||
// a proof. Store what you knew when you knew it.
|
||||
type provenTier struct {
|
||||
Archive string // volid of the archive that PASSED; "" = a legacy record with no archive
|
||||
At time.Time // when that run passed (UTC)
|
||||
Archive string // volid of the archive that PASSED; "" = a legacy record with no archive
|
||||
Tier string // "local" | "pbs" — as the run reported it; "" = pre-R-189 record
|
||||
Verified string // what the run verified (e.g. "boot+running"); "" = pre-R-189 record
|
||||
At time.Time // when that run passed (UTC)
|
||||
}
|
||||
|
||||
// provenTierJSON is the on-disk shape (R-86). The legacy shape was a bare RFC3339 STRING per
|
||||
// target; both are read, only this one is written — see NewRestoreTestState.
|
||||
// reportable reports whether this record can be re-reported to the hub as a restore-test result.
|
||||
//
|
||||
// It needs BOTH the archive and the tier: the hub keys its edge-triggered failure state on the
|
||||
// archive and its per-tier proof lookup on the tier, so an entry missing either is not a usable
|
||||
// proof — and emitting one anyway would be a report the hub cannot act on, dressed as evidence.
|
||||
// A pre-R-189 record is therefore silently not reported; the tier's next real proof fills it in.
|
||||
func (p provenTier) reportable() bool { return p.Archive != "" && p.Tier != "" }
|
||||
|
||||
// provenTierJSON is the on-disk shape. Two older shapes are read and neither is written:
|
||||
//
|
||||
// v1 (pre-R-86) "<target>": "<RFC3339>" — a time, no archive
|
||||
// v2 (R-86) "<target>": {archive, proven_at} — due-check usable, not reportable
|
||||
// v3 (R-189) "<target>": {archive, tier, verified, …} — both
|
||||
//
|
||||
// Fields absent in an older file unmarshal to "", which is exactly the "no usable proof" signal the
|
||||
// readers above test for — the migration needs no version number because the absence IS the answer.
|
||||
type provenTierJSON struct {
|
||||
Archive string `json:"archive"`
|
||||
Tier string `json:"tier,omitempty"`
|
||||
Verified string `json:"verified,omitempty"`
|
||||
ProvenAt string `json:"proven_at"`
|
||||
}
|
||||
|
||||
@@ -99,21 +131,31 @@ func NewRestoreTestState(path string) *RestoreTestState {
|
||||
if perr != nil {
|
||||
continue
|
||||
}
|
||||
s.last[target] = provenTier{Archive: cur.Archive, At: t.UTC()}
|
||||
s.last[target] = provenTier{Archive: cur.Archive, Tier: cur.Tier, Verified: cur.Verified, At: t.UTC()}
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// RecordSuccess stamps a tier as proven at t, naming the ARCHIVE that passed. Only call this for a
|
||||
// PASSING restore-test — the archive is what makes the tier not-due, so recording one for a failed
|
||||
// run would retire the archive unproven.
|
||||
func (s *RestoreTestState) RecordSuccess(target, archive string, t time.Time) error {
|
||||
// RecordSuccess stamps a tier as proven at t, naming the ARCHIVE that passed, the TIER the run
|
||||
// reported, and what it verified. Only call this for a PASSING restore-test — the archive is what
|
||||
// makes the tier not-due, so recording one for a failed run would retire the archive unproven.
|
||||
//
|
||||
// ONLY SUCCESSES ARE PERSISTED, AND THE ASYMMETRY IS DELIBERATE (R-189 §8.1). Say it here because
|
||||
// the next reader will notice failures are absent and try to "fix" it:
|
||||
//
|
||||
// a SUCCESS suppresses future work — a proven archive is never re-tested, so a lost proof leaves
|
||||
// the system quietly less tested than it believes. It must survive a restart.
|
||||
//
|
||||
// a FAILURE causes future work — a failing tier stays due and is retried at the next evaluation,
|
||||
// so a lost failure heals itself within one interval. Persisting it would do the opposite of
|
||||
// helping: a healed tier would keep reporting a failure that is no longer true.
|
||||
func (s *RestoreTestState) RecordSuccess(target, archive, tier, verified string, t time.Time) error {
|
||||
if target == "" {
|
||||
return nil
|
||||
}
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
s.last[target] = provenTier{Archive: archive, At: t.UTC()}
|
||||
s.last[target] = provenTier{Archive: archive, Tier: tier, Verified: verified, At: t.UTC()}
|
||||
return s.saveLocked()
|
||||
}
|
||||
|
||||
@@ -138,7 +180,13 @@ func (s *RestoreTestState) ProvenArchive(target string) (string, bool) {
|
||||
return p.Archive, true
|
||||
}
|
||||
|
||||
// Snapshot returns a copy of the last-proven TIMES — for the host-report gauge.
|
||||
// Snapshot returns a copy of the last-proven TIMES.
|
||||
//
|
||||
// It carried the comment "for the host-report gauge" from the day it was written and **had no caller
|
||||
// at all** until R-189 — a seam built and never wired, and an invariant asserted in a comment with
|
||||
// nothing pinning it, in one method. The host report is now fed by ProvenRestoreTests below, which
|
||||
// carries the archive and the tier that a bare timestamp cannot. This stays for callers that want
|
||||
// only the times; if it acquires none, delete it rather than let it claim a purpose again.
|
||||
func (s *RestoreTestState) Snapshot() map[string]time.Time {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
@@ -149,6 +197,41 @@ func (s *RestoreTestState) Snapshot() map[string]time.Time {
|
||||
return out
|
||||
}
|
||||
|
||||
// ProvenRestoreTests renders the persisted proofs as host-report entries — the R-189 fix.
|
||||
//
|
||||
// It satisfies hub.RestoreTestReporter's shape, so the collector can merge these with the in-memory
|
||||
// results. What it emits is a RE-REPORT of a run that really happened, not a synthesis:
|
||||
//
|
||||
// - `Pass` is true because ONLY successes are stored (RecordSuccess is the sole writer);
|
||||
// - `SourceArchive`, `SourceTier`, `Verified` and `TestedAt` are the values that run reported;
|
||||
// - the run mechanics (scratch VMID, duration, warnings) are NOT re-invented. An absent duration
|
||||
// is not a claim; a fabricated one would be.
|
||||
//
|
||||
// A record that cannot be reported honestly is omitted rather than padded — see provenTier.reportable.
|
||||
// **A tier with no usable proof produces NO entry**: an unproven tier reading as proven would be a
|
||||
// worse defect than the one this fixes.
|
||||
func (s *RestoreTestState) ProvenRestoreTests(context.Context) []hub.RestoreTest {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
out := make([]hub.RestoreTest, 0, len(s.last))
|
||||
for _, p := range s.last {
|
||||
if !p.reportable() {
|
||||
continue
|
||||
}
|
||||
out = append(out, hub.RestoreTest{
|
||||
SourceArchive: p.Archive,
|
||||
SourceTier: p.Tier,
|
||||
Pass: true,
|
||||
Verified: p.Verified,
|
||||
TestedAt: p.At.UTC().Format(time.RFC3339),
|
||||
})
|
||||
}
|
||||
// Deterministic order: the report is compared byte-wise by the contract test, and Go's map
|
||||
// iteration is randomised.
|
||||
sort.Slice(out, func(i, j int) bool { return out[i].SourceTier < out[j].SourceTier })
|
||||
return out
|
||||
}
|
||||
|
||||
// OldestFirst orders targets by "least recently proven first"; never-proven sorts FIRST.
|
||||
//
|
||||
// This is the operator's 2026-07-26 ruling (Option 1): self-balancing, no new config knob, and it
|
||||
@@ -185,7 +268,10 @@ func (s *RestoreTestState) OldestFirst(targets []string) []string {
|
||||
func (s *RestoreTestState) saveLocked() error {
|
||||
raw := make(map[string]provenTierJSON, len(s.last))
|
||||
for target, p := range s.last {
|
||||
raw[target] = provenTierJSON{Archive: p.Archive, ProvenAt: p.At.UTC().Format(time.RFC3339)}
|
||||
raw[target] = provenTierJSON{
|
||||
Archive: p.Archive, Tier: p.Tier, Verified: p.Verified,
|
||||
ProvenAt: p.At.UTC().Format(time.RFC3339),
|
||||
}
|
||||
}
|
||||
data, err := json.MarshalIndent(raw, "", " ")
|
||||
if err != nil {
|
||||
|
||||
@@ -303,11 +303,11 @@ func TestOldestFirst_Ordering(t *testing.T) {
|
||||
t.Fatalf("unexpected: %v", got)
|
||||
}
|
||||
}
|
||||
_ = st.RecordSuccess("local", "local:backup/a.tar.zst", now)
|
||||
_ = st.RecordSuccess("local", "local:backup/a.tar.zst", "local", "boot+running", now)
|
||||
if got := st.OldestFirst([]string{"local", "felhom-pbs"}); got[0] != "felhom-pbs" {
|
||||
t.Fatalf("a never-proven tier must sort before a proven one; got %v", got)
|
||||
}
|
||||
_ = st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/b", now.Add(time.Hour))
|
||||
_ = st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/b", "pbs", "boot+running", now.Add(time.Hour))
|
||||
if got := st.OldestFirst([]string{"local", "felhom-pbs"}); got[0] != "local" {
|
||||
t.Fatalf("the least recently proven must sort first; got %v", got)
|
||||
}
|
||||
@@ -318,8 +318,8 @@ func TestOldestFirst_Ordering(t *testing.T) {
|
||||
func TestOldestFirst_DeterministicOnTies(t *testing.T) {
|
||||
st := NewRestoreTestState(filepath.Join(t.TempDir(), "rt.json"))
|
||||
now := time.Now().UTC()
|
||||
_ = st.RecordSuccess("b-tier", "b:archive", now)
|
||||
_ = st.RecordSuccess("a-tier", "a:archive", now)
|
||||
_ = st.RecordSuccess("b-tier", "b:archive", "local", "boot+running", now)
|
||||
_ = st.RecordSuccess("a-tier", "a:archive", "local", "boot+running", now)
|
||||
for i := 0; i < 20; i++ {
|
||||
if got := st.OldestFirst([]string{"b-tier", "a-tier"}); got[0] != "a-tier" {
|
||||
t.Fatalf("tie-break must be deterministic; iteration %d gave %v", i, got)
|
||||
@@ -334,7 +334,7 @@ func TestRestoreTestState_PersistenceAndCorruption(t *testing.T) {
|
||||
now := time.Now().UTC().Truncate(time.Second)
|
||||
|
||||
st := NewRestoreTestState(path)
|
||||
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", now); err != nil {
|
||||
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", "pbs", "boot+running", now); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
reopened := NewRestoreTestState(path)
|
||||
|
||||
@@ -180,7 +180,15 @@ func (s *Scheduler) tick(ctx context.Context) {
|
||||
return
|
||||
}
|
||||
if archive == "" {
|
||||
s.logger.Debug("backup: restore-test not due this evaluation")
|
||||
// A POSITIVE OBSERVABLE, at INFO, and this is not noise — it is standing rule 3.
|
||||
//
|
||||
// Before R-86 every tick ran a heavy restore-test, so the scheduler was audible by
|
||||
// construction. Now "nothing is due" is the NORMAL outcome, and at DEBUG it is silent: an
|
||||
// empty journal would be equally consistent with a healthy loop and with a dead goroutine,
|
||||
// which is the exact shape the R-88 watcher was retired for. One line per evaluation is four
|
||||
// lines a day at the 6h default, and it names each tier's verdict so the answer to "why did
|
||||
// nothing run last night?" is in the log rather than in a re-derivation.
|
||||
s.logger.Info("backup: restore-test evaluated — nothing due", "verdicts", s.verdictSummary(ctx))
|
||||
return
|
||||
}
|
||||
|
||||
@@ -210,7 +218,11 @@ func (s *Scheduler) tick(ctx context.Context) {
|
||||
if rt.Pass && s.rtState != nil && target != "" {
|
||||
// R-86: the ARCHIVE is recorded, not merely the time — that is what makes the tier
|
||||
// not-due until a NEWER archive settles, and what makes a proof survive a restart.
|
||||
if err := s.rtState.RecordSuccess(target, archive, s.now()); err != nil {
|
||||
// R-189: the TIER and what was VERIFIED go with it, so the proof can be RE-REPORTED after a
|
||||
// restart. Both come from the run's own result, never re-derived — `rt.SourceTier` is what
|
||||
// this run was actually judged as, and deriving it later would need a storage lookup that
|
||||
// can fail on the one path where failing means mislabelling a proof.
|
||||
if err := s.rtState.RecordSuccess(target, archive, rt.SourceTier, rt.Verified, s.now()); err != nil {
|
||||
s.logger.Warn("backup: could not persist the restore-test proof state", "target", target, "err", err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -10,8 +10,22 @@ import (
|
||||
// Store holds the agent's LATEST backup result per target and the latest restore-test
|
||||
// result — the point-in-time state the host-report surfaces. It is updated by the backup
|
||||
// runner + the restore-test scheduler/selftest and read by the collector via the hub
|
||||
// BackupReporter / RestoreTestReporter seams. In-memory (lost on restart; the cadence
|
||||
// re-populates) and mutex-guarded for the concurrent collector vs scheduler access.
|
||||
// BackupReporter / RestoreTestReporter seams. In-memory and mutex-guarded for the concurrent
|
||||
// collector vs scheduler access.
|
||||
//
|
||||
// **"lost on restart; the cadence re-populates" — that sentence used to be here and it is now
|
||||
// FALSE for restore-tests (R-189, 2026-08-03).** It was true while a timer re-tested every tier
|
||||
// daily. Under R-86's per-archive due-check the agent will NOT re-test an archive it has already
|
||||
// proven, so a proof lost to a restart is not repeated until the next archive generation — a week on
|
||||
// the offsite tier — and the hub reports that tier unproven throughout. Observed, not predicted: a
|
||||
// real 14.5 GB offsite restore passed, the agent was restarted 2 m 43 s later for a deploy, and two
|
||||
// consecutive host-reports carried `0 restore-tests`.
|
||||
//
|
||||
// The durable half is `RestoreTestState` (on disk, per tier, with the archive) and the collector
|
||||
// merges the two — see hub.ProvenRestoreTestReporter. This store remains the ONLY place a FAILURE is
|
||||
// recorded, and that asymmetry is deliberate: a failing tier stays due and is retried, so a lost
|
||||
// failure heals itself, while a lost success leaves the system quietly less tested than it believes.
|
||||
// Backups are unaffected — their freshness has a ground truth on the storage (R-84).
|
||||
type Store struct {
|
||||
mu sync.Mutex
|
||||
byTarget map[string]hub.Backup // latest backup per target id
|
||||
|
||||
+96
-5
@@ -47,6 +47,22 @@ type RestoreTestReporter interface {
|
||||
RestoreTests(ctx context.Context) []RestoreTest
|
||||
}
|
||||
|
||||
// ProvenRestoreTestReporter is the DURABLE half of the restore-test signal (R-189).
|
||||
//
|
||||
// RestoreTestReporter above is backed by an in-memory store whose own comment used to read "lost on
|
||||
// restart; the cadence re-populates". That was true while a timer re-tested every tier daily. It
|
||||
// stopped being true on 2026-08-03: under per-archive due-ness the agent will not re-test an archive
|
||||
// it has already proven, so a proof lost to a restart is not repeated for a whole archive generation
|
||||
// — a week on the offsite tier — and the hub reports the tier unproven the entire time.
|
||||
//
|
||||
// Observed, not predicted: a real 14.5 GB offsite restore passed at 15:25:14, the agent was restarted
|
||||
// 2 m 43 s later for a deploy, and the hub logged `0 restore-tests` on the next two reports.
|
||||
//
|
||||
// (*backup.RestoreTestState).ProvenRestoreTests satisfies this. nil → the merge is a no-op.
|
||||
type ProvenRestoreTestReporter interface {
|
||||
ProvenRestoreTests(ctx context.Context) []RestoreTest
|
||||
}
|
||||
|
||||
// PBSReporter is the slice-6-Phase-B seam the pbs verify loop plugs into (same pattern).
|
||||
// Returns the agent's latest-known PBS snapshot inventory + verify-state. nil → empty.
|
||||
type PBSReporter interface {
|
||||
@@ -79,6 +95,7 @@ type Collector struct {
|
||||
storage StorageObserver
|
||||
backups BackupReporter
|
||||
restoreTests RestoreTestReporter
|
||||
provenTests ProvenRestoreTestReporter
|
||||
pbs PBSReporter
|
||||
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
|
||||
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
|
||||
@@ -427,16 +444,90 @@ func (c *Collector) collectBackups(ctx context.Context) []Backup {
|
||||
return []Backup{}
|
||||
}
|
||||
|
||||
// collectRestoreTests merges the in-memory result with the PERSISTED per-tier proofs (R-189).
|
||||
//
|
||||
// The rule is ONE ENTRY PER TIER, NEWEST WINS, and it falls out of what each source means rather
|
||||
// than from a preference between them:
|
||||
//
|
||||
// - the in-memory store holds this process's latest run, pass OR fail. A failure exists nowhere
|
||||
// else and must always reach the hub — a failing tier is retried at the next evaluation, so its
|
||||
// record is short-lived by design;
|
||||
// - the persisted state holds the last SUCCESS per tier and survives a restart.
|
||||
//
|
||||
// Comparing by TestedAt gives the right answer in every case without special-casing: a fresh failure
|
||||
// beats an older stored success (the failure is the news), a stored success beats a stale in-memory
|
||||
// entry after a restart, and a tier proved twice never appears twice — two entries for one tier would
|
||||
// read at the hub as two tests.
|
||||
//
|
||||
// A tier with no usable persisted proof contributes NOTHING. Reporting an unproven tier as proven
|
||||
// would be a worse defect than the one this closes.
|
||||
func (c *Collector) collectRestoreTests(ctx context.Context) []RestoreTest {
|
||||
if c.restoreTests == nil {
|
||||
return []RestoreTest{}
|
||||
out := []RestoreTest{}
|
||||
if c.restoreTests != nil {
|
||||
if r := c.restoreTests.RestoreTests(ctx); r != nil {
|
||||
out = append(out, r...)
|
||||
}
|
||||
}
|
||||
if r := c.restoreTests.RestoreTests(ctx); r != nil {
|
||||
return r
|
||||
if c.provenTests == nil {
|
||||
return out
|
||||
}
|
||||
return []RestoreTest{}
|
||||
|
||||
// Index what we already have by tier, keeping the newest per tier.
|
||||
best := map[string]int{} // tier → index into out
|
||||
for i, rt := range out {
|
||||
if rt.SourceTier == "" {
|
||||
continue // untiered entry: never deduped, never overwritten — we cannot say what it is
|
||||
}
|
||||
if j, seen := best[rt.SourceTier]; !seen || newerRestoreTest(rt, out[j]) {
|
||||
best[rt.SourceTier] = i
|
||||
}
|
||||
}
|
||||
for _, p := range c.provenTests.ProvenRestoreTests(ctx) {
|
||||
if p.SourceTier == "" {
|
||||
continue // not usable as a per-tier proof; the state layer already filters these
|
||||
}
|
||||
i, seen := best[p.SourceTier]
|
||||
if !seen {
|
||||
out = append(out, p)
|
||||
best[p.SourceTier] = len(out) - 1
|
||||
continue
|
||||
}
|
||||
if newerRestoreTest(p, out[i]) {
|
||||
out[i] = p
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// newerRestoreTest reports whether a was tested after b. An unparseable or absent timestamp is
|
||||
// treated as OLDER, so a malformed entry can never displace a good one.
|
||||
func newerRestoreTest(a, b RestoreTest) bool {
|
||||
ta, aok := parseRestoreTestedAt(a.TestedAt)
|
||||
tb, bok := parseRestoreTestedAt(b.TestedAt)
|
||||
if !aok {
|
||||
return false
|
||||
}
|
||||
if !bok {
|
||||
return true
|
||||
}
|
||||
return ta.After(tb)
|
||||
}
|
||||
|
||||
func parseRestoreTestedAt(s string) (time.Time, bool) {
|
||||
t, err := time.Parse(time.RFC3339, s)
|
||||
if err != nil {
|
||||
return time.Time{}, false
|
||||
}
|
||||
return t.UTC(), true
|
||||
}
|
||||
|
||||
// SetProvenRestoreTests wires the durable proof source. It is a setter rather than a constructor
|
||||
// argument because the persisted state is opened later in main() than the collector is built; the
|
||||
// same shape as the other late-wired seams here. **The wiring is asserted by an AST test** — the
|
||||
// method it feeds carried a doc comment naming a "host-report gauge" for weeks with no caller at
|
||||
// all, and this fix must not become the next instance of that.
|
||||
func (c *Collector) SetProvenRestoreTests(p ProvenRestoreTestReporter) { c.provenTests = p }
|
||||
|
||||
// collectPBSSnapshots reads the latest PBS snapshot inventory via the seam (nil → empty).
|
||||
func (c *Collector) collectPBSSnapshots(ctx context.Context) []PBSSnapshot {
|
||||
if c.pbs == nil {
|
||||
|
||||
@@ -0,0 +1,213 @@
|
||||
package hub
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// R-189 — a passing restore-test must survive an agent restart and reach the hub.
|
||||
//
|
||||
// THE OBSERVATION THIS EXISTS FOR (2026-08-03, demo-felhom): a real 14.5 GB offsite restore-test
|
||||
// PASSED at 15:25:14; the agent was restarted 2 m 43 s later for a deploy; the hub logged
|
||||
// `0 restore-tests` on the next two host-reports. The in-memory store's own comment said "lost on
|
||||
// restart; the cadence re-populates", which was true under a timer and stopped being true when R-86
|
||||
// made the agent refuse to re-test an archive it has already proven.
|
||||
//
|
||||
// Timestamps here carry JITTER (odd minutes and seconds, not round hours) — yesterday a test was
|
||||
// hollow because a perfectly regular series landed exactly on a threshold and passed under the
|
||||
// mutation it was meant to catch.
|
||||
|
||||
type fakeLatest struct{ tests []RestoreTest }
|
||||
|
||||
func (f *fakeLatest) RestoreTests(context.Context) []RestoreTest { return f.tests }
|
||||
|
||||
type fakeProven struct{ tests []RestoreTest }
|
||||
|
||||
func (f *fakeProven) ProvenRestoreTests(context.Context) []RestoreTest { return f.tests }
|
||||
|
||||
func rt(tier, archive string, pass bool, at time.Time) RestoreTest {
|
||||
return RestoreTest{
|
||||
SourceArchive: archive, SourceTier: tier, Pass: pass,
|
||||
Verified: "boot+running", TestedAt: at.UTC().Format(time.RFC3339),
|
||||
}
|
||||
}
|
||||
|
||||
// mergeCollector builds a Collector with only the two restore-test seams wired — the merge is what
|
||||
// is under test, not the rest of the collection.
|
||||
func mergeCollector(latest, proven []RestoreTest) *Collector {
|
||||
c := &Collector{}
|
||||
if latest != nil {
|
||||
c.restoreTests = &fakeLatest{tests: latest}
|
||||
}
|
||||
if proven != nil {
|
||||
c.provenTests = &fakeProven{tests: proven}
|
||||
}
|
||||
return c
|
||||
}
|
||||
|
||||
func findTier(got []RestoreTest, tier string) (RestoreTest, int) {
|
||||
var hit RestoreTest
|
||||
n := 0
|
||||
for _, e := range got {
|
||||
if e.SourceTier == tier {
|
||||
hit, n = e, n+1
|
||||
}
|
||||
}
|
||||
return hit, n
|
||||
}
|
||||
|
||||
// ── SCENARIO A — a proof survives a restart and reaches the hub ──────────────────────────────
|
||||
//
|
||||
// COMPANION RED-PROOF (observed 2026-08-03): delete the `c.provenTests` merge from
|
||||
// collectRestoreTests (return the in-memory slice as it used to) →
|
||||
//
|
||||
// --- FAIL: TestMerge_ProofSurvivesARestart
|
||||
// restoretest_merge_test.go: after a restart the persisted proof must be reported; got 0 entr(ies)
|
||||
//
|
||||
// which is exactly the live observation: `0 restore-tests`. Restored.
|
||||
func TestMerge_ProofSurvivesARestart(t *testing.T) {
|
||||
provenAt := time.Date(2026, 8, 3, 13, 25, 14, 0, time.UTC) // the real run's timestamp
|
||||
// After a restart the in-memory store is EMPTY — this is the whole point.
|
||||
c := mergeCollector([]RestoreTest{}, []RestoreTest{
|
||||
rt("pbs", "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z", true, provenAt),
|
||||
})
|
||||
|
||||
got := c.collectRestoreTests(context.Background())
|
||||
if len(got) != 1 {
|
||||
t.Fatalf("after a restart the persisted proof must be reported; got %d entr(ies): %+v", len(got), got)
|
||||
}
|
||||
e := got[0]
|
||||
if e.SourceArchive != "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" {
|
||||
t.Fatalf("the entry must name the archive that was proven — the hub keys on it; got %q", e.SourceArchive)
|
||||
}
|
||||
if e.SourceTier != "pbs" || !e.Pass {
|
||||
t.Fatalf("the entry must be a PASS on the tier it was proven on; got tier=%q pass=%v", e.SourceTier, e.Pass)
|
||||
}
|
||||
if e.TestedAt != provenAt.Format(time.RFC3339) {
|
||||
t.Fatalf("the entry must carry the ORIGINAL test time, not now(); got %q", e.TestedAt)
|
||||
}
|
||||
}
|
||||
|
||||
// ── SCENARIO B — the report does not invent a pass ───────────────────────────────────────────
|
||||
//
|
||||
// COMPANION RED-PROOF (observed): make the state layer emit an entry for an unproven tier (drop the
|
||||
// `reportable()` filter in ProvenRestoreTests, so a legacy record with no archive is emitted) — the
|
||||
// equivalent at this layer is a proven-source that returns an entry for a tier nothing proved, which
|
||||
// this test injects directly and the assertion below rejects.
|
||||
func TestMerge_NeverInventsAPassForAnUnprovenTier(t *testing.T) {
|
||||
// Nothing proven anywhere: no in-memory result, no persisted proof.
|
||||
c := mergeCollector([]RestoreTest{}, []RestoreTest{})
|
||||
if got := c.collectRestoreTests(context.Background()); len(got) != 0 {
|
||||
t.Fatalf("a tier with no proof must produce NO entry — an unproven tier reading as proven is "+
|
||||
"worse than the defect being fixed; got %+v", got)
|
||||
}
|
||||
|
||||
// And an entry the state layer could not describe (no tier) is never promoted into a proof.
|
||||
c2 := mergeCollector([]RestoreTest{}, []RestoreTest{
|
||||
{SourceArchive: "local:backup/x.tar.zst", SourceTier: "", Pass: true,
|
||||
TestedAt: time.Date(2026, 8, 1, 4, 41, 58, 0, time.UTC).Format(time.RFC3339)},
|
||||
})
|
||||
if got := c2.collectRestoreTests(context.Background()); len(got) != 0 {
|
||||
t.Fatalf("a persisted record with no tier is not a usable proof and must be dropped; got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// ── SCENARIO C — a fresh in-memory result wins, and never duplicates ─────────────────────────
|
||||
//
|
||||
// COMPANION RED-PROOF (observed 2026-08-03): remove the de-duplication (append every persisted entry
|
||||
// unconditionally) →
|
||||
//
|
||||
// --- FAIL: TestMerge_NewerWinsAndNeverDuplicatesATier
|
||||
// restoretest_merge_test.go: one entry per tier; got 2 for "pbs" — the hub would read two tests
|
||||
//
|
||||
// Restored.
|
||||
func TestMerge_NewerWinsAndNeverDuplicatesATier(t *testing.T) {
|
||||
lastWeek := time.Date(2026, 7, 27, 19, 55, 41, 0, time.UTC) // jittered, from the real box
|
||||
fiveMinAgo := time.Date(2026, 8, 3, 13, 25, 14, 0, time.UTC)
|
||||
|
||||
c := mergeCollector(
|
||||
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/new", true, fiveMinAgo)},
|
||||
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/old", true, lastWeek)},
|
||||
)
|
||||
got := c.collectRestoreTests(context.Background())
|
||||
e, n := findTier(got, "pbs")
|
||||
if n != 1 {
|
||||
t.Fatalf("one entry per tier; got %d for \"pbs\" — the hub would read two tests: %+v", n, got)
|
||||
}
|
||||
if e.SourceArchive != "felhom-pbs:backup/ct/9201/new" {
|
||||
t.Fatalf("the NEWER result must win; got %q tested %q", e.SourceArchive, e.TestedAt)
|
||||
}
|
||||
|
||||
// ...and the older-in-memory / newer-persisted direction, which is the post-restart case.
|
||||
c2 := mergeCollector(
|
||||
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/old", true, lastWeek)},
|
||||
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/new", true, fiveMinAgo)},
|
||||
)
|
||||
e2, n2 := findTier(c2.collectRestoreTests(context.Background()), "pbs")
|
||||
if n2 != 1 || e2.SourceArchive != "felhom-pbs:backup/ct/9201/new" {
|
||||
t.Fatalf("newest must win regardless of which source it came from; got %d entr(ies), archive %q", n2, e2.SourceArchive)
|
||||
}
|
||||
}
|
||||
|
||||
// ── SCENARIO D — a failure still reaches the hub ─────────────────────────────────────────────
|
||||
//
|
||||
// The merge must not mask a failure with an older stored success. A failing tier is retried at the
|
||||
// next evaluation and its record lives ONLY in memory, so losing it here would silence the loudest
|
||||
// DR signal this system produces.
|
||||
func TestMerge_AFailureIsStillReported(t *testing.T) {
|
||||
provenLastWeek := time.Date(2026, 7, 27, 19, 55, 41, 0, time.UTC)
|
||||
failedJustNow := time.Date(2026, 8, 3, 13, 41, 7, 0, time.UTC)
|
||||
|
||||
c := mergeCollector(
|
||||
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/new", false, failedJustNow)},
|
||||
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/old", true, provenLastWeek)},
|
||||
)
|
||||
e, n := findTier(c.collectRestoreTests(context.Background()), "pbs")
|
||||
if n != 1 {
|
||||
t.Fatalf("one entry per tier; got %d: %+v", n, c.collectRestoreTests(context.Background()))
|
||||
}
|
||||
if e.Pass {
|
||||
t.Fatalf("a FAILURE newer than the stored proof must be what is reported — masking it would "+
|
||||
"silence the loudest DR signal there is; got pass=%v archive=%q", e.Pass, e.SourceArchive)
|
||||
}
|
||||
}
|
||||
|
||||
// Two different tiers are both reported — the merge is per tier, not a single slot.
|
||||
func TestMerge_BothTiersSurvive(t *testing.T) {
|
||||
c := mergeCollector(
|
||||
[]RestoreTest{rt("local", "felhom-backup:backup/vzdump-lxc-9201-a.tar.zst", true,
|
||||
time.Date(2026, 8, 3, 4, 44, 50, 0, time.UTC))},
|
||||
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/x", true,
|
||||
time.Date(2026, 8, 2, 5, 12, 33, 0, time.UTC))},
|
||||
)
|
||||
got := c.collectRestoreTests(context.Background())
|
||||
if _, n := findTier(got, "local"); n != 1 {
|
||||
t.Fatalf("the in-memory tier must survive the merge; got %+v", got)
|
||||
}
|
||||
if _, n := findTier(got, "pbs"); n != 1 {
|
||||
t.Fatalf("the persisted tier must survive the merge; got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A malformed timestamp must never displace a good entry — "unparseable" is not "newest".
|
||||
func TestMerge_MalformedTimestampNeverWins(t *testing.T) {
|
||||
good := rt("pbs", "felhom-pbs:backup/ct/9201/good", true, time.Date(2026, 8, 3, 13, 25, 14, 0, time.UTC))
|
||||
bad := RestoreTest{SourceArchive: "felhom-pbs:backup/ct/9201/bad", SourceTier: "pbs", Pass: true, TestedAt: "not-a-time"}
|
||||
|
||||
c := mergeCollector([]RestoreTest{good}, []RestoreTest{bad})
|
||||
e, n := findTier(c.collectRestoreTests(context.Background()), "pbs")
|
||||
if n != 1 || e.SourceArchive != "felhom-pbs:backup/ct/9201/good" {
|
||||
t.Fatalf("an unparseable timestamp must not displace a good entry; got %d entr(ies), archive %q", n, e.SourceArchive)
|
||||
}
|
||||
}
|
||||
|
||||
// A nil proven-source leaves the pre-R-189 behaviour exactly as it was.
|
||||
func TestMerge_NilProvenSourceIsANoOp(t *testing.T) {
|
||||
only := rt("local", "felhom-backup:backup/x.tar.zst", true, time.Date(2026, 8, 3, 4, 44, 50, 0, time.UTC))
|
||||
c := mergeCollector([]RestoreTest{only}, nil)
|
||||
got := c.collectRestoreTests(context.Background())
|
||||
if len(got) != 1 || got[0].SourceArchive != only.SourceArchive {
|
||||
t.Fatalf("a nil durable source must not change anything; got %+v", got)
|
||||
}
|
||||
}
|
||||
@@ -14,10 +14,32 @@ Nothing in the build, deploy or session-end path checked that either existed, so
|
||||
documented-path reinstall would have silently DOWNGRADED both boxes to the pre-merge
|
||||
agent — and would have *succeeded* while doing it.
|
||||
|
||||
THE INVARIANT, AND WHY IT IS THIS ONE.
|
||||
THE INVARIANTS — there are TWO now, and the second is R-188's price.
|
||||
|
||||
For every `v<semver>` git tag in this repo: the matching generic package must be DOWNLOADABLE,
|
||||
and the tag must serve the agent's configs.
|
||||
(1) For every `v<semver>` git tag in this repo: the matching generic package must be DOWNLOADABLE,
|
||||
and the tag must serve the agent's configs.
|
||||
|
||||
(2) No PUBLISHED version may be missing its tag.
|
||||
|
||||
Invariant (2) is new (R-188, 2026-08-03) and it exists because `release-agent.sh` now pushes the tag
|
||||
AFTER publishing. The old order pushed the tag first, and the old comment said why: a tag with no
|
||||
package is caught here, a package with no tag is invisible, because the Gitea package LISTING api
|
||||
needs a token this gate does not have. That reasoning was sound and the ordering was still wrong —
|
||||
the tag push is what wakes CI, so every correct release had a ~50% chance of running this gate in the
|
||||
seconds before its own package existed and mailing the operator a failure for a release that worked
|
||||
(measured across two releases: runs 12/13 and 17/18, same shas, opposite results).
|
||||
|
||||
Moving the push does not get to trade invariant (2) away, so it is asserted here instead — WITHOUT a
|
||||
token, and therefore as a BOUNDED PROBE rather than an enumeration:
|
||||
|
||||
* the FRONTIER — the versions immediately above the highest tag. This is the realistic failure the
|
||||
new ordering makes possible: publish succeeds, tag push fails, so the orphan is exactly one
|
||||
version beyond the newest tag.
|
||||
* the GAPS — patch versions that fall between two existing tags and have no tag of their own.
|
||||
|
||||
Re-measured 2026-08-03, not assumed: `GET /api/v1/packages/admin?type=generic` answers **401** with no
|
||||
token, so absence still cannot be proven. The probe set is PRINTED on every run, because a check whose
|
||||
coverage is invisible reads as a guarantee it is not making.
|
||||
|
||||
The task's §8.4 asked for a different one — *"the version the hub tells machines to install must be
|
||||
downloadable"* — and that is the better invariant in principle. **It is not implementable from CI,
|
||||
@@ -36,6 +58,8 @@ v0.120.0, which is published.
|
||||
|
||||
**What it does NOT catch, stated plainly:** the hub vouching a version that was never released at
|
||||
all (no tag, no package). Nothing here can see that; it belongs at vouch time, in the hub. → R-184.
|
||||
Nor does the converse probe prove that NO untagged package exists — only that none exists at the
|
||||
probed versions, which are printed. Closing that properly needs a read token in CI (→ R-184).
|
||||
|
||||
FAIL-CLOSED. A network error, an unparseable response or an unreachable Gitea is exit **2
|
||||
INCONCLUSIVE**, naming every URL tried — never a pass. "Cannot determine" is not "fine": that is the
|
||||
@@ -48,7 +72,7 @@ carrying python3 and git and nothing else, and an earlier workflow step died on
|
||||
|
||||
python3 scripts/check-published-versions.py
|
||||
|
||||
Exit: 0 every tag installable · 1 at least one is not · 2 could not be determined.
|
||||
Exit: 0 both invariants hold · 1 either is violated · 2 could not be determined.
|
||||
Env: GITEA_BASE overrides the Gitea root (CI sets the in-cluster service URL).
|
||||
"""
|
||||
import json
|
||||
@@ -94,6 +118,53 @@ def inconclusive(msg):
|
||||
sys.exit(2)
|
||||
|
||||
|
||||
def _pkg_exists(version):
|
||||
"""True iff the generic package for `version` is downloadable anonymously."""
|
||||
url = "%s/api/packages/%s/generic/%s/%s/%s" % (GITEA_BASE, OWNER, PKG, version, PKG)
|
||||
status, _ = _get(url)
|
||||
return status == 200, url
|
||||
|
||||
|
||||
def untagged_probe_set(versions):
|
||||
"""The versions to probe for invariant (2), as (version, why) pairs.
|
||||
|
||||
Bounded on purpose and printed by the caller: the package listing api needs a token (401,
|
||||
re-measured 2026-08-03), so absence cannot be enumerated. What CAN be done is to probe the
|
||||
places an orphan would actually land.
|
||||
|
||||
FRONTIER — a publish that succeeded followed by a tag push that failed leaves the orphan
|
||||
exactly one version past the newest tag. This is the failure mode the R-188
|
||||
reordering makes possible, so it is the one that must not be guesswork.
|
||||
GAPS — a patch number skipped between two consecutive tags. Bounded per gap so a typo'd
|
||||
tag (v0.130.0 after v0.121.1) cannot turn this into a thousand requests.
|
||||
"""
|
||||
parsed = sorted(tuple(int(p) for p in v.split(".")) for v in versions)
|
||||
have = set(parsed)
|
||||
out = []
|
||||
if not parsed:
|
||||
return out
|
||||
|
||||
hi = parsed[-1]
|
||||
for cand, why in (
|
||||
((hi[0], hi[1], hi[2] + 1), "next patch after the newest tag"),
|
||||
((hi[0], hi[1], hi[2] + 2), "second patch after the newest tag"),
|
||||
((hi[0], hi[1] + 1, 0), "next minor after the newest tag"),
|
||||
((hi[0] + 1, 0, 0), "next major after the newest tag"),
|
||||
):
|
||||
if cand not in have:
|
||||
out.append(("%d.%d.%d" % cand, why))
|
||||
|
||||
MAX_GAP_PROBES = 12
|
||||
for a, b in zip(parsed, parsed[1:]):
|
||||
if a[0] != b[0] or a[1] != b[1]:
|
||||
continue # a minor/major step is not a patch gap
|
||||
for patch in range(a[2] + 1, min(b[2], a[2] + 1 + MAX_GAP_PROBES)):
|
||||
cand = (a[0], a[1], patch)
|
||||
if cand not in have:
|
||||
out.append(("%d.%d.%d" % cand, "patch gap between v%d.%d.%d and v%d.%d.%d" % (a + b)))
|
||||
return out
|
||||
|
||||
|
||||
def main():
|
||||
print("check-published-versions — every released agent version must be INSTALLABLE")
|
||||
print(" gitea:", GITEA_BASE)
|
||||
@@ -144,14 +215,42 @@ def main():
|
||||
else:
|
||||
print(" ok v%s: binary downloadable + tag serves its configs" % v)
|
||||
|
||||
# ── invariant (2): no PUBLISHED version may be missing its tag (R-188) ──────────────────────
|
||||
probes = untagged_probe_set(versions)
|
||||
orphans = []
|
||||
print()
|
||||
print(" converse probe — a published version with no tag (bounded; the package listing api")
|
||||
print(" needs a token, so this cannot enumerate). Probing %d version(s):" % len(probes))
|
||||
for v, why in probes:
|
||||
try:
|
||||
exists, url = _pkg_exists(v)
|
||||
except Exception as e:
|
||||
inconclusive("network failure while probing v%s: %s" % (v, e))
|
||||
mark = "PUBLISHED — NO TAG" if exists else "absent (ok)"
|
||||
print(" %-10s %-42s %s" % (v, why, mark))
|
||||
if exists:
|
||||
orphans.append((v, url))
|
||||
|
||||
print()
|
||||
if bad or orphans:
|
||||
if orphans:
|
||||
print("check-published-versions: %d PUBLISHED VERSION(S) WITH NO TAG" % len(orphans))
|
||||
for v, url in orphans:
|
||||
print(" v%s is downloadable at %s but has no git tag." % (v, url))
|
||||
print(" A release publishes and then pushes its tag; a package with no tag means the")
|
||||
print(" push failed or was skipped. The local tag is probably still in the release")
|
||||
print(" clone — finish it with:")
|
||||
for v, _ in orphans:
|
||||
print(" git push origin v%s" % v)
|
||||
print(" (and if the tag is gone, re-create it on the released commit before pushing.)")
|
||||
if bad:
|
||||
print("check-published-versions: %d RELEASED VERSION(S) NOT INSTALLABLE" % len(bad))
|
||||
print(" A tagged version with no package is a release that was BUILT and never PUBLISHED —")
|
||||
print(" the R-115 defect, three times in five days. Publish it with:")
|
||||
print(" scripts/release-agent.sh <version>")
|
||||
if bad or orphans:
|
||||
return 1
|
||||
print("check-published-versions: ALL RELEASED VERSIONS INSTALLABLE")
|
||||
print("check-published-versions: ALL RELEASED VERSIONS INSTALLABLE, AND NONE UNTAGGED")
|
||||
return 0
|
||||
|
||||
|
||||
|
||||
@@ -51,7 +51,11 @@ if [[ -z "$BIN" ]]; then
|
||||
BIN="$(mktemp -t felhom-agent.XXXXXX)"
|
||||
CLEANUP_BIN="$BIN"
|
||||
log "building felhom-agent $VERSION from $REPO_ROOT …"
|
||||
( cd "$REPO_ROOT" && CGO_ENABLED=0 go build -ldflags "-X main.version=${VERSION}" -o "$BIN" ./cmd/felhom-agent )
|
||||
# These flags MUST match release-agent.sh's build exactly — see the long comment there (R-186).
|
||||
# They used to differ: this line forced CGO_ENABLED=0 and produced a binary 74 KB smaller than
|
||||
# the one the release path built for the same version. One version name must mean one binary
|
||||
# whichever entry point produced it.
|
||||
( cd "$REPO_ROOT" && go build -trimpath -buildvcs=false -ldflags "-X main.version=${VERSION}" -o "$BIN" ./cmd/felhom-agent )
|
||||
fi
|
||||
[[ -f "$BIN" ]] || die "binary not found: $BIN"
|
||||
trap '[[ -n "$CLEANUP_BIN" ]] && rm -f "$CLEANUP_BIN"' EXIT
|
||||
|
||||
@@ -74,7 +74,25 @@ existing="$(curl -fsS -o /dev/null -w '%{http_code}' \
|
||||
BIN="$(mktemp -t felhom-agent-XXXXXX)"
|
||||
trap 'rm -f "$BIN"' EXIT
|
||||
log "building $VERSION …"
|
||||
go build -ldflags "-X main.version=$VERSION" -o "$BIN" ./cmd/felhom-agent \
|
||||
# REPRODUCIBLE BY CONSTRUCTION (R-186). The sha printed below is the one the operator vouches, and
|
||||
# until now nobody could rebuild it to check: `go build` stamps a module version derived from VCS
|
||||
# state, so a build made BEFORE the tag exists and a rebuild made after it are different binaries.
|
||||
# Measured 2026-08-03 at this commit — same source, same toolchain, same ldflags:
|
||||
#
|
||||
# default flags, no tag yet .. 18f4a495… 14 085 464 B (mod v0.121.2-0.2026…-3d0a1d61)
|
||||
# default flags, tagged ...... 4a38f394… 14 085 440 B (mod v0.121.99)
|
||||
# -trimpath -buildvcs=false ... 7ffcdf1d… 14 064 574 B IDENTICAL both ways
|
||||
#
|
||||
# `-buildvcs=false` removes the stamp — nothing in this repo reads it (no `ReadBuildInfo` caller,
|
||||
# verified) and the version comes from the explicit ldflag below, which is where it belongs.
|
||||
# `-trimpath` removes absolute build paths, so a rebuild from a different checkout directory also
|
||||
# matches. Neither is a sequencing trick: the property no longer depends on WHEN the build happens.
|
||||
#
|
||||
# CGO is deliberately left at its default. publish-agent.sh's fallback build used to force
|
||||
# CGO_ENABLED=0 and therefore produced a DIFFERENT binary (13 990 236 B, 74 KB smaller) for the same
|
||||
# version — one version name, two binaries, by whichever entry point was used. Both now build the
|
||||
# same way; if that ever has to change, change it in BOTH or the guarantee is gone.
|
||||
go build -trimpath -buildvcs=false -ldflags "-X main.version=$VERSION" -o "$BIN" ./cmd/felhom-agent \
|
||||
|| die "go build failed"
|
||||
built_ver="$("$BIN" --version 2>/dev/null | awk '{print $2}')"
|
||||
[[ "$built_ver" == "$VERSION" ]] \
|
||||
@@ -82,10 +100,27 @@ built_ver="$("$BIN" --version 2>/dev/null | awk '{print $2}')"
|
||||
BUILT_SHA="$(sha256sum "$BIN" | awk '{print $1}')"
|
||||
log "built ok: sha256 $BUILT_SHA"
|
||||
|
||||
# ── 4. Tag (before publishing, so a published version always has a tag) ─────────────────────────
|
||||
# Order matters in this direction only: a tag with no package is caught by
|
||||
# scripts/check-published-versions.py on the next CI run; a package with no tag is invisible to it,
|
||||
# ── 4. Tag LOCALLY (the push comes after the publish — see step 6) ──────────────────────────────
|
||||
#
|
||||
# THE ORDER CHANGED, AND ONLY THE PUSH MOVED (R-188, 2026-08-03).
|
||||
#
|
||||
# It used to be tag → push tag → publish, and the reason written here was sound: a tag with no
|
||||
# package is caught by scripts/check-published-versions.py, a package with no tag is invisible to it,
|
||||
# because the Gitea package LISTING api needs a token the gate does not have.
|
||||
#
|
||||
# What that reasoning missed is that the tag PUSH is what wakes CI (`on: [push]`), so the gate ran in
|
||||
# the seconds between the tag becoming visible and the package existing — and correctly failed. Every
|
||||
# correct release had roughly a coin-flip chance of emailing the operator a failure for a release
|
||||
# that worked. Measured across two releases in one session: runs 12/13 (v0.121.0) and 17/18
|
||||
# (v0.121.1), same sha each time, opposite results. R-168 made that mail the thing that cannot be
|
||||
# missed; a mail that is wrong half the time is one you stop reading, and then the real one goes too.
|
||||
#
|
||||
# So the tag is still created HERE, before anything is published — the build and the tag still
|
||||
# describe the same commit, and a failed publish leaves a purely local tag that never misled anyone.
|
||||
# It simply becomes VISIBLE (to CI, and to any installer fetching raw/tag/…) only once the package
|
||||
# is downloadable. The invariant the old order protected is not traded away: it is asserted directly
|
||||
# by the gate's new converse probe (a published version with no tag FAILS), so both directions are
|
||||
# now checked rather than one being arranged for.
|
||||
log "tagging $TAG at $(git rev-parse --short HEAD) …"
|
||||
git tag -a "$TAG" -m "agent $TAG
|
||||
|
||||
@@ -94,7 +129,6 @@ sha256 of the published binary: $BUILT_SHA
|
||||
|
||||
felhom-host-install.sh fetches this version's config files from raw/tag/$TAG/configs/,
|
||||
so this tag is part of the released artifact, not a bookmark (R-183)."
|
||||
git push origin "$TAG" || die "tag push failed — refusing to publish an untagged version"
|
||||
|
||||
# ── 5. Publish (the existing script; deliberately not reimplemented) ────────────────────────────
|
||||
log "publishing …"
|
||||
@@ -104,9 +138,48 @@ log "publishing …"
|
||||
# R-115 exists to make unforgettable was, on its first use, unrunnable. The mode bit is restored in
|
||||
# the same commit; this line makes the release independent of it, because a file mode is exactly the
|
||||
# kind of thing that is lost again by a checkout, an archive, or a copy.
|
||||
bash "$REPO_ROOT/scripts/publish-agent.sh" "$VERSION" "$BIN" || die "publish failed"
|
||||
if ! bash "$REPO_ROOT/scripts/publish-agent.sh" "$VERSION" "$BIN"; then
|
||||
# The tag is LOCAL-ONLY at this point, so a failed publish must not leave one behind: the next
|
||||
# attempt would die at step 2's "tag $TAG already exists" and read as "this version is already
|
||||
# released", which would be exactly backwards. Only remove it if nothing was in fact published —
|
||||
# if a package DOES exist, the tag is wanted and must be pushed, not deleted.
|
||||
now_published="$(curl -fsS -o /dev/null -w '%{http_code}' \
|
||||
"$GITEA_BASE/api/packages/$GITEA_OWNER/generic/felhom-agent/$VERSION/felhom-agent" 2>/dev/null || true)"
|
||||
if [[ "$now_published" == "200" ]]; then
|
||||
log "publish reported failure but the package IS downloadable — keeping the local tag; push it with: git push origin $TAG"
|
||||
else
|
||||
git tag -d "$TAG" >/dev/null 2>&1 && log "removed the local-only tag $TAG so the release can be retried"
|
||||
fi
|
||||
die "publish failed"
|
||||
fi
|
||||
|
||||
# ── 6. Verify by an INDEPENDENT download ────────────────────────────────────────────────────────
|
||||
# ── 6. Push the tag, now that the package exists ────────────────────────────────────────────────
|
||||
# This is the step that makes the release VISIBLE — to CI, and to every `raw/tag/v<version>/` fetch
|
||||
# the installer makes. It runs last of the two so CI can never see a tag whose package is not there.
|
||||
#
|
||||
# If it fails, the release is HALF DONE and must be said so loudly: the package is published and the
|
||||
# tag exists only in this clone, which is precisely the orphan the gate's converse probe now catches.
|
||||
# The recovery is one line and it is printed rather than described.
|
||||
log "pushing $TAG …"
|
||||
if ! git push origin "$TAG"; then
|
||||
cat >&2 <<EOF
|
||||
|
||||
RELEASE HALF DONE — the package is PUBLISHED and its tag is NOT pushed.
|
||||
|
||||
version : $VERSION
|
||||
sha256 : $BUILT_SHA
|
||||
|
||||
The tag exists in this clone only. Nothing installs from an untagged version (the installer
|
||||
fetches this version's configs from raw/tag/$TAG/), and scripts/check-published-versions.py will
|
||||
FAIL on it as a published version with no tag. Finish the release with:
|
||||
|
||||
git push origin $TAG
|
||||
|
||||
EOF
|
||||
die "tag push failed after a successful publish — see above"
|
||||
fi
|
||||
|
||||
# ── 7. Verify by an INDEPENDENT download ────────────────────────────────────────────────────────
|
||||
# The publish step's own success is not proof: it reports on its own write. What matters is that a
|
||||
# box can now GET the bytes and that they are the bytes that were built. This is the same
|
||||
# presence-is-not-success rule the project earned twice — a step that says "done" and a fetch that
|
||||
|
||||
Reference in New Issue
Block a user