7 Commits

Author SHA1 Message Date
admin 4d82591052 release-agent.sh: the publish leg was unrunnable on its first real use
gates / gates (push) Successful in 7s
R-115's whole point is that publishing cannot be forgotten because it rides the
release script. On the first real release through it (v0.121.0, R-86) it died at
exactly that leg:

  scripts/release-agent.sh: line 101: .../scripts/publish-agent.sh: Permission denied

publish-agent.sh has been mode 0644 since it was created on 2026-06-28 — every
earlier caller ran it as `bash scripts/publish-agent.sh`, so nothing ever noticed,
and release-agent.sh (written the same day it was needed) called it directly.

Two fixes, both small and both wanted: restore the executable bit, and invoke it
through `bash` so the release no longer depends on a file mode — the kind of
thing a checkout, an archive or a copy loses again.

The v0.121.0 tag created by the failed run is withdrawn and recreated on this
commit; nothing was published under it (verified 404 on the package endpoint), so
one version name still means one binary.
2026-08-03 15:04:52 +02:00
admin 4618169036 R-86: restore-test follows the backup, not the clock (v0.121.0)
gates / gates (push) Failing after 7s
The ticker survives as the EVALUATION interval only. A tier is DUE when its
newest archive that has settled for `settle` (default 24h) has not been proven:
daily tier -> proved daily on yesterday's archive, weekly tier -> weekly on its
own, newborn -> UNKNOWN.

The trap avoided: the literal reading ("newest archive is >= 24h old") is NEVER
true on a daily tier, so it silently switches restore-testing off where it
matters most. Red-proved at 0 runs over 5 simulated days.

- state records WHICH archive was proven; legacy files keep their time and yield
  no proven archive (each tier due once after the upgrade, deliberately)
- two knobs replace one: restore_test_eval_interval_seconds (6h, measured) and
  restore_test_settle_seconds (24h). The old cadence key keeps its DISABLE
  meaning verbatim and now seeds the settle lag, with a start-up WARN.
- due-check runs BEFORE the heavy-op gate (a frequent poll must not make a
  starting backup record a failure, F-A1)
- candidate picker skips implausible archives (a phantom would be due forever)
- new read-only --selftest=restore-test-due prints the verdict + its cost
2026-08-03 14:54:57 +02:00
admin 1b14cfd0b4 REPORT: Scenario F measured on real CI (runs 69 vs 70, same commit)
gates / gates (push) Successful in 7s
2026-08-03 12:44:37 +02:00
admin 0db77666c6 REPORT: release path + published-versions gate (no version bump)
gates / gates (push) Failing after 7s
2026-08-03 12:37:09 +02:00
admin dd2d1feb6e release path publishes, and an unreleasable version fails CI (R-115, R-183)
gates / gates (push) Failing after 7s
NO VERSION BUMP and nothing built: no Go code changed. The agent stays v0.120.0.

scripts/release-agent.sh — THE way to release. build -> tag -> publish -> verify
by INDEPENDENT download. Publishing was a separate remembered step and was
forgotten three times in five days (R-111's 17 stranded releases, 0.114.0, and
0.120.0 — deployed to both demo hosts and undownloadable, so a documented-path
reinstall would have silently downgraded them WHILE REPORTING SUCCESS). R-111's
own closing line named this leg and closed SHIPPED without it; it recurred the
same afternoon, which is the evidence that a note is not a mechanism.

It tags because felhom-host-install.sh now fetches the sixteen agent config
files from raw/tag/v<version>/ (R-183): a released version with no tag 404s a
box mid-install, as root, on a virgin machine. It verifies by downloading what
it just published and comparing the sha to what it built — the publish step's
own success is a report on its own write; a fetch returning the right bytes is
a different claim. It refuses a dirty/unpushed tree and refuses to re-release an
existing version. It does NOT vouch: that points machines at a version and stays
the operator's act.

scripts/check-published-versions.py — the gate. Every v<semver> tag must have a
downloadable package AND a tag tree serving the agent's configs. Registered as
NOT --fast (needs network; a push must not fail because Gitea blinked), and the
CI workflow now runs the FULL gate set instead of --fast — otherwise the gate
would have been registered and never run, the built-but-never-wired failure this
project has shipped four times.

The invariant is not the one specified, and the reason was measured, not assumed:
the hub artifact manifest is 401 without a per-customer passphrase and Gitea's
package LISTING api is 401 without a token, while the package DOWNLOAD url and
the git TAGS api are anonymous. So CI cannot ask "what is vouched" without an
operator credential — whose addition is the operator's call. The tag-based
invariant needs none and catches all three recorded instances. What it does not
catch (the hub vouching a version never released at all) is filed as R-184.
2026-08-03 12:34:20 +02:00
admin 9dfd89cb94 docs: agent 0.120.0 published + vouched, proven on two reinstalled boxes (R-178)
gates / gates (push) Successful in 6s
v0.120.0 had been built, committed and deployed to both demo hosts but never
published: the Gitea generic package 404'd and the hub manifest vouched 0.119.0.
Installer step 5 skips only on an exact version match, so a documented-path
reinstall would have downgraded both boxes to the pre-merge agent -- and would
have succeeded, since step_grows passes -sysdata-grow 0 and 0.119.0's mp1 resize
never fires. Published from a clean tree (upload 201, round-trip GET verified,
sha a7763d31b55b5ce7...) and vouched; both reinstalls then fetched and verified
it over the real customer path.

Filed as the third instance of R-115. No version bump, nothing built.
Evidence: felhom.eu/REPORT.md
2026-08-03 09:33:50 +02:00
admin 4bb84fc3ca REPORT: v0.120.0 + golden 3.0.0 — built and proven at the bake, NOT proven on a box
gates / gates (push) Successful in 12s
States the scope reduction first: Phases 6-7 (reinstall both demo boxes and
prove one end to end) were NOT done, nothing was wiped, and the golden is
deliberately left unvouched as a result. Filed as R-178.

Also records the two instrument errors this session: a census truncated by
head -10 that gave the wrong answer about --sysdata-grow (the installer does
pass it), and a wait gated on a marker the bake prints before publishing.
2026-08-03 07:15:29 +02:00
20 changed files with 1759 additions and 156 deletions
+13 -1
View File
@@ -43,7 +43,19 @@ jobs:
- name: Run the gate entry point
# The ONLY thing CI runs. No go build, no go test, no linting, no deploy. The
# exit code IS the result: no `|| true`, no pipe that could swallow it.
run: cd ws/felhom-agent && python3 scripts/agent_gates.py --fast
#
# THE FULL SET, NOT `--fast` (R-115, 2026-08-03). `--fast` means "no network and no
# container runtime" and exists for `.githooks/pre-push`, where a push must not fail
# because Gitea blinked or because someone is on a train. CI is the opposite machine: it
# has the network, it is not in anyone's way, and it is the half that emails. The
# published-versions gate — the R-115 mechanism, which asks Gitea whether a released
# version can actually be downloaded — is network-bound and therefore runs ONLY here.
# Leaving `--fast` in place would have registered that gate and never run it, which is the
# built-but-never-wired failure this project has shipped four times.
env:
# In-cluster, so the check does not depend on public DNS or the ingress TLS chain.
GITEA_BASE: http://gitea.gitea-system.svc.cluster.local:3000
run: cd ws/felhom-agent && python3 scripts/agent_gates.py
- name: Alarm on failure
# THE POINT OF THE WHOLE THING. Probe P5 measured that a failed run produces NO mail, NO
+128
View File
@@ -1,3 +1,131 @@
## v0.121.0 — a restore-test proves each BACKUP, not the clock (2026-08-03, R-86)
**The trigger changed; the restore-test did not.** `Scheduler.Run` still has a ticker, but it is now
the **evaluation interval** — how often "is anything due?" is asked — and no longer the thing that
decides a test happens. What decides is a per-archive due-check
(`internal/backup/restoretest_due.go`):
> Let **A** = the newest archive on this tier that has settled for at least `settle` (default 24 h).
> The tier is **DUE** when **A** exists and **A has not already been proven**.
A daily tier is therefore proved once a day, on yesterday's archive; a weekly tier once a week, on
its own; a newborn tier is UNKNOWN and never a fault. Per-archive due-ness IS the pacing — one test
per archive generation and no more — so there is deliberately no second rate limiter on top of it.
**The trap this avoided, recorded because it is the version a reasonable person writes.** R-86's own
wording ("~24 h after its own newest archive") implemented literally is *"due when the newest archive
is ≥ 24 h old"* — and on a **daily** tier that is never true, because a new archive resets the
newest-archive age to zero long before it reaches 24 h. The literal rule silently switches
restore-testing OFF for the tier that matters most. It has its own red-proof, which was observed
failing with **0 runs over 5 simulated days**.
**What the fix rests on**
- **The state records WHICH archive was proven** (`restoretest_state.go`), not merely when a tier last
passed — a time cannot answer "have we proven *this* archive". A pre-R-86 state file keeps its time
(rotation ordering survives the upgrade) and yields **no** proven archive, so each tier is due
exactly once after the upgrade: one extra test per tier, once, which is the safe direction.
- **Two knobs replace one, and the old one is not silently repurposed.**
`restore_test_eval_interval_seconds` (how often due-ness is asked; default **6 h**) and
`restore_test_settle_seconds` (how long an archive must sit; default **24 h**). The deprecated
`restore_test_cadence_seconds` keeps its DISABLE meaning (negative) verbatim, and a positive value
now seeds the **settle lag** — with a start-up WARN naming both replacements.
- **6 h is bounded from both sides, not picked.** MEASURED cost of one evaluation on demo-felhom
(Part 1.4): local dir storage **18 ms**, the PBS tier over the WAN to ep0 **392 ms**, both together
**430 ms** — cheap enough for minutes, so cost is not the constraint. The **ceiling** is: a tier
whose restore-test keeps failing stays due, so the evaluation interval is also its RETRY interval,
and a retry is a multi-GB restore.
- **The due-check runs BEFORE the heavy-operation gate is taken.** Evaluations are frequent now, and
holding that gate for a read that answers "nothing to do" would open a window at every evaluation
in which a starting backup cannot acquire — and a backup that cannot acquire records a failure and
pages the operator (F-A1). Nothing heavy starts before the gate.
- **The candidate picker skips implausible archives.** Under per-archive due-ness an incomplete
1-byte phantom (F-CRIT-2's artefact, which server-side prune does not collect) would be picked
forever, fail forever, never earn proof, and leave the tier due at EVERY evaluation — turning the
evaluation interval into the retry rate for a multi-GB restore. `PickRestoreCandidateOn` now
delegates to the settle-aware picker, so both callers agree.
**Unchanged, deliberately:** the restore-test itself (restore → boot → verify → destroy the scratch),
its journal, crash recovery, the scratch VMID band, the one-heavy-operation gate, success-only proof
credit, and oldest-proven ordering — which survives as the tie-break **between due tiers**.
**New:** `--selftest=restore-test-due` — read-only; prints the per-tier due verdict the scheduler
would act on, with the measured cost of the lookup.
**Live finding, pre-existing and NOT caused by this change (filed as R-185):** on demo-felhom the
agent's PVE token has no ACL on `/storage/felhom-backup`, so the API returns an EMPTY content listing
for that storage (root sees three archives). The host tier has therefore never been restore-testable
on that box, and both R-85's rotation and R-86's due-check report it indistinguishably from "newborn"
("no settled archive yet"). Verified live against `local` (grant present → 3 archives) and
`felhom-backup` (no grant → `{"data":[]}`).
## Releasing publishes, and an unreleasable version cannot pass CI (2026-08-03, R-115 + R-183) — **NO VERSION BUMP**
**No Go code changed, so nothing is bumped and nothing was built.** This is the release path and a
gate; the agent stays **v0.120.0**.
**`scripts/release-agent.sh` — THE way to release.** Build → **tag** → publish → **verify by an
independent download**. Publishing used to be a separate remembered step and was **forgotten three
times in five days** (R-111's seventeen stranded releases, 0.114.0, and 0.120.0 — which sat deployed
on both demo hosts and undownloadable, so a documented-path reinstall would have silently downgraded
them to the pre-merge agent *while reporting success*). R-111's own closing line named this leg and
closed SHIPPED without it; it recurred the same afternoon, which is the evidence that a note is not a
mechanism.
- It **tags** because `felhom-host-install.sh` now fetches this version's sixteen config files from
`raw/tag/v<version>/` (R-183) — a released version without a tag 404s a box mid-install, as root.
- It **verifies by downloading what it just published** and comparing the sha to what it built. The
publish step's own success is a report on its own write; a fetch returning the right bytes is a
different claim, and it is the one that matters.
- It **refuses** a dirty or unpushed tree, and refuses to re-release an existing version — one
version name must never mean two binaries.
- It **does NOT vouch.** Vouching points machines at a version and stays the operator's act.
**`scripts/check-published-versions.py` — the gate (R-115 mechanism (b)).** Every `v<semver>` tag
must have a downloadable package AND a tag tree that serves the agent's configs. Registered in
`agent_gates.py` as **not `--fast`** (it needs network, and a push must not fail because Gitea
blinked), and **the CI workflow now runs the FULL gate set** rather than `--fast` — otherwise the
gate would have been registered and never run, which is the built-but-never-wired failure this
project has shipped four times.
**The invariant is NOT the one the task specified, and the reason was measured.** The task asked for
*"the version the hub tells machines to install must be downloadable"*. That is the better invariant
and CI cannot see it: the hub's artifact manifest answers **401** without a per-customer passphrase,
and Gitea's package LISTING api answers **401** without a token, while the package DOWNLOAD url and
the git TAGS api are both anonymous. Putting an operator credential into CI to close that gap is the
operator's call, not a gate author's. The tag-based invariant needs no credential and **catches all
three recorded instances**, because the release script creates the tag and publishes in one act.
**What it does not catch — the hub vouching a version that was never released at all — is recorded
as R-184 rather than assumed away.**
## docs — v0.120.0 PUBLISHED + vouched, and proven on two reinstalled boxes (2026-08-03, R-178) — **no version bump, nothing built**
**Nothing shipped in this entry.** It records an operational fact the version history could not
otherwise carry: **v0.120.0 had been built, committed at `cd6e267` and deployed to both demo hosts,
but never published.** `GET https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/0.120.0/felhom-agent`
returned **HTTP 404** (0.119.0 → 200), and the hub's Day-0 manifest accordingly vouched **0.119.0**.
**Why that mattered more than it looks.** `felhom-host-install.sh` step 5 skips the binary install
only when `installed == vouched` *exactly*, so a documented-path reinstall would have replaced the
merge-aware 0.120.0 with the pre-merge 0.119.0 — **and would have succeeded**, because the current
`step_grows` passes `-sysdata-grow 0` and 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on error)
therefore never fires. Both demo boxes would have come back on an agent that predates the merge while
every log line read green.
Published this session on an operator ruling, from a clean tree (`git status --porcelain` empty,
`HEAD == origin/main == 4bb84fc3`): `scripts/publish-agent.sh 0.120.0` → upload **HTTP 201**,
round-trip GET verified, `AGENT_SHA256=a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`;
then vouched in the hub manifest, which resolved the sha authoritatively from Gitea rather than
trusting the submitted value.
Both reinstalls then fetched and sha-verified it over the real customer path —
`verified sha256 a7763d31b55b5ce7… matches the hub manifest` — and 4b's single-volume grow was
exercised live on both boxes (`data +46G (->70G, ONE volume)` on demo-hp,
`+226G (->250G)` on demo-felhom), each producing `mp0` at `/var/lib/felhom` with **no `mp1`**.
**Filed as the third instance of R-115** (*"publishing is a remembered step"*), which has been
WAITING-ON-OPERATOR since 2026-07-29. Full evidence: `felhom.eu/REPORT.md`.
## v0.120.0 — one data volume (2026-08-03, R-165 · decision D-a · variant V-c) — `build-golden.sh` 2.1.0 → **3.0.0**
**The dedicated backup partition stops existing.** A golden built by `build-golden.sh` v3.0.0 ships a
+12 -2
View File
@@ -63,13 +63,23 @@ internal/storage/ storage observer + durable ids + role/claim classifiers + S
> stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from
> elsewhere).
> **RELEASING IS ONE COMMAND, AND IT PUBLISHES (R-115).** There used to be a raw `go build` line
> here and a *separate* "Publish" row, so publishing was a step someone had to remember — and it was
> **forgotten three times in five days**, the last leaving agent v0.120.0 deployed on both demo hosts
> and undownloadable, where a documented-path reinstall would have silently downgraded them while
> reporting success. Do not hand-roll the build: the script also creates the `v<version>` git TAG
> that `felhom-host-install.sh` fetches this version's sixteen config files from (R-183), and it
> verifies by an **independent download** rather than trusting the publish step's own output.
> `scripts/publish-agent.sh` still exists and is still correct — the release script CALLS it rather
> than reimplementing it.
| Step | Where | One-liner |
|---|---|---|
| Build | DooPlex (local) | `cd /mnt/5_hdd/felhom.eu/git/felhom-agent && git pull && go build -ldflags '-X main.version=<v>' -o /tmp/felhom-agent-<v> ./cmd/felhom-agent` |
| **Release** (build + tag + publish + verify) | DooPlex (local) | `GITEA_USER=admin GITEA_TOKEN=<tok> scripts/release-agent.sh <ver>` — refuses a dirty/unpushed tree and refuses to re-release an existing version |
| Copy | local → felhom-pve | `scp /tmp/felhom-agent-<v> felhom-pve:/tmp/` (one hop) |
| Deploy | felhom-pve | backup `.bak-<old>``install -m0755``systemctl restart felhom-agent` (non-root `felhom-agent` user, config `/etc/felhom-agent/agent.json`) |
| Ship configs | felhom-pve | sudoers (`/etc/sudoers.d/felhom-agent`) + guarded-mkfs wrapper WITH the binary when `configs/` changed |
| Publish | DooPlex (local) | `scripts/publish-agent.sh <ver> <bin>` (REGISTRY_* creds); hub Day-0 manifest vouch = operator follow-up |
| **Vouch** | hub operator UI | Configs → Day-0 artifacts. **Deliberately NOT automated** — vouching is what points machines at a version, and it stays your act (prove-then-vouch) |
| Verify | felhom-pve | `felhom-agent --version` + journal (clean ReassertGuestBinds, no capability degradation) |
## Proxmox model (the load-bearing rules)
+26
View File
@@ -5,6 +5,32 @@
## Current
- **2026-08-03 — v0.121.0 (R-86): the restore-test follows the BACKUP, not the clock.** The ticker is
now only the **evaluation interval**; a tier is **DUE** when its newest archive that has settled for
`settle` (default 24 h) **has not been proven**. Daily tier → proved daily on yesterday's archive;
weekly tier → weekly on its own; newborn → UNKNOWN. **The trap, so it is not reintroduced:** the
literal reading of R-86 — *"due when the newest archive is ≥ 24 h old"* — is NEVER true on a daily
tier (a new archive resets the age before it reaches the lag), so it switches restore-testing off
where it matters most. Red-proved at 0 runs over 5 simulated days.
- **The state now records WHICH archive was proven**, not just when a tier passed. A pre-R-86 file
keeps its time (ordering survives) and yields no proven archive → each tier is due once after the
upgrade, deliberately.
- **The old cadence key:** `restore_test_cadence_seconds` is DEPRECATED. Negative still DISABLES
(verbatim); a positive value now seeds the **settle lag** and the daemon WARNs once at start-up
naming `restore_test_eval_interval_seconds` (default 6 h) and `restore_test_settle_seconds`
(default 24 h). It is NOT carried into the evaluation interval.
- **6 h is bounded from both ends:** measured evaluation cost (local 18 ms, PBS-over-WAN 392 ms,
both 430 ms) says cost is irrelevant; the ceiling is that a FAILING tier stays due, so the
evaluation interval is also its retry interval for a multi-GB restore.
- The due-check now runs **before** the heavy-operation gate is taken (a frequent poll must not be
able to make a starting backup record a failure — F-A1), and the candidate picker skips archives
failing `archivePlausiblyComplete` (a phantom would be due forever and fail forever).
- New read-only `--selftest=restore-test-due` prints the per-tier verdict + its cost.
- **R-185 (filed, NOT fixed here):** on demo-felhom the agent token has no ACL on
`/storage/felhom-backup`, so its content listing comes back EMPTY (root sees 3 archives) — the
host tier has never been restore-testable there, and the due-check cannot distinguish that from
a newborn tier.
- **2026-07-28 — v0.107.0: F-REBOOT fixed — a guest rebooted mid-backup now comes back by itself.**
New `internal/localapi/guestpower.go`: a 60 s watchdog that starts a guest which is `onboot:1`,
stopped, unlocked, and has no vzdump in flight. It closes the two narrow gaps that let
+89 -30
View File
@@ -1,41 +1,100 @@
# REPORT — CI runs the gate entry point on every push (R-168, 2026-08-02)
# REPORT — releasing publishes, and an unreleasable version fails CI (R-115, R-183)
**Overwritten** per the standing rule. The prior contents (session 1's gate entry point, same day) have their durable record in `CHANGELOG.md`.
**No version bump, no build, no deploy.** This adds `.gitea/workflows/gates.yml` and a CHANGELOG
entry. Nothing compiled changed.
**Date:** 2026-08-03 · **Repo:** `felhom-agent` · **NO VERSION BUMP** — the agent stays **v0.120.0**,
no Go code changed, nothing was built or deployed.
## What changed
`.gitea/workflows/gates.yml` — on every push, a Gitea Actions runner obtains this repo at the
**exact pushed SHA** (shallow `git fetch`, no `uses:` step anywhere) and runs
`python3 scripts/agent_gates.py --fast` and nothing else. The exit code is the job's result: no `|| true`, no
pipe that could swallow it.
| File | |
|---|---|
| `scripts/release-agent.sh` | **new** — THE release path: build → tag → publish → verify by independent download |
| `scripts/check-published-versions.py` | **new** — the R-115 gate |
| `scripts/agent_gates.py` | registers the gate as **not `--fast`** (it needs network) |
| `.gitea/workflows/gates.yml` | CI now runs the **full** gate set, not `--fast` |
| `CLAUDE.md` | the raw `go build` line is replaced by the release script; a **Vouch** row replaces the old Publish row |
**It REPORTS, it cannot REFUSE**, and the workflow header says so in the pre-push hook's voice: this
repo pushes straight to `main` with no pull request, so there is no merge for a status check to stand
at. The refusing half is `.githooks/pre-push` (per-clone, `--no-verify`-able); this half notices when
that was skipped. Making CI blocking needs branch protection plus a PR workflow — an operator
decision, tracked as `felhom.eu` `OPEN-ITEMS.md` **R-169**.
## Why
**A failed run emails the operator.** Probe P5 measured that Gitea itself sends **nothing** on a
failed run — no mail, no notification row, no log line — so the workflow sends its own via Resend and
prints the provider's accepted id, making "a message left the machine" an observable. Demonstrated on
a real red run in `felhom.eu`: `RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab`.
Publishing was a step someone had to remember and was **forgotten three times in five days**
R-111's seventeen stranded releases, 0.114.0, and 0.120.0, which sat deployed on both demo hosts and
undownloadable, so a documented-path reinstall would have silently downgraded them to the pre-merge
agent **while reporting success**. R-111's own closing line named this leg and closed SHIPPED without
it; it recurred the same afternoon. A note is not a mechanism.
**CI reproduces the workspace's SIBLING LAYOUT on purpose.** This repo's entry point invokes the
shared `reuse_refs_check.py` that lives in the `felhom.eu` clone next door and is deliberately never
copied here, and this repo's `REUSE.md` cites `hub/internal/store/dr_recipe.go`, which lives in the hub. The workflow clones
`felhom.eu` as a sibling; without it the gate fails **closed** with `gate is MISSING` — correctly,
but for the wrong reason.
The script also **tags**, because `felhom-host-install.sh` now fetches the agent's sixteen config
files from `raw/tag/v<version>/` (R-183). A released version with no tag 404s a box mid-install, as
root, on a virgin machine. Tag and package are two halves of one release.
## Verification
It **verifies by downloading what it just published** and comparing the sha to what it built. The
publish step's own success is a report on its own write; a fetch returning the right bytes is a
different claim, and it is the one that matters.
First run: run #1, id=9, conclusion **success**, sha `eb991445`.
It **does not vouch** — that points machines at a version and stays the operator's act.
**CI and the local pre-push hook AGREE**, which is the check that mattered: `90 cited paths — exact 88, suffix 1, ambiguous 0, cross-repo 1, FAILED 0 (siblings searched: felhom.eu)` — identical to the local run.
## The gate's invariant — not the one specified, and the reason was measured
The runner is unprivileged host-mode, shared with the other three repos on a single owner-scoped
registration (measured: all four repos' tasks claimed by `felhom-gates-runner`). Full probe detail,
the security posture and the teardown: `felhom.eu/documentation/audits/SPIKE-ci-runner-2026-08-02.md`
and `felhom.eu/REPORT.md`.
The task's §8.4 asked for *"the version the hub tells machines to install must be downloadable"*.
**CI cannot see that**, measured rather than assumed (P-C):
| Endpoint | Anonymous |
|---|---|
| Gitea package **download** | **200** (and **404** for a fake version — it discriminates) |
| Gitea **tags** api | **200** |
| Gitea package **listing** api | **401** — token required |
| Hub `/api/v1/artifacts/<customer>` | **401** — per-customer passphrase required |
So a credential-free gate can ask *"is this version installable"* but not *"which version is
vouched"*. Adding an operator credential to CI to close that is the operator's call, not a gate
author's. The implemented invariant — **every `v<semver>` tag must have a downloadable package and a
tag tree that serves the agent's configs** — needs no credential and **catches all three recorded
instances**, because the release script creates the tag and publishes in one act.
**What it does not catch, stated rather than assumed away:** the hub vouching a version that was
never released at all. Nothing here can see that; it belongs at vouch time in the hub. → **R-184**.
## Proof
| Check | Result |
|---|---|
| `go build ./... && go vet ./...` | OK |
| `go test ./...` | **29 packages ok, rc=0** (read separately from any commit) |
| `agent_gates.py --fast` | `published` correctly **SKIPPED** — the pre-push hook must not fail because Gitea blinked |
| `agent_gates.py` (full) | `reuse-refs` OK, `published` OK |
| release script: re-release guard | `ERROR: tag v0.120.0 already exists — releasing over it would make one version name two binaries`, rc=1 |
| release script: clean-tree guard | `ERROR: working tree is dirty — commit and push first`, rc=1 |
### Red-proof F — both directions
- **A tagged-but-unpublished version** (`v9.9.9` created for the purpose): gate **rc=1**,
`binary NOT downloadable (HTTP 404 …)`. This is the R-115 shape exactly.
- **The gate deregistered from the entry point**, same bad state: `agent_gates.py` → **rc=0, "all
agent gates OK"**. Restored → **rc=1, CONVICTED: published**. The guard is what catches it, not
something else.
### Scenario F measured on REAL CI, not inferred
Runs **69** and **70** are on the **same commit** `0db7766`:
| run | state of the repo | CI |
|---|---|---|
| 69 | no `v9.9.9` | **success** |
| 70 | `v9.9.9` tagged, not published | **failure** |
Same code, same workflow, one variable — so the gate demonstrably RUNS in CI and fails for exactly
the R-115 condition. This also retrospectively explains runs 67/68, which were red in the window when
`v9.9.9` first existed. **One deliberate CI failure e-mail reached the operator — that was this
proof, not an incident.**
I could not read CI's own step log to attribute those runs directly: the Gitea jobs endpoint requires
an API token, and the only credential available on this host (`~/.docker/config.json`) is a registry
password, which the API rejects. The controlled before/after above replaced that log rather than an
assumption standing in for it.
`v9.9.9` was deleted afterwards; `git ls-remote --tags` shows only `v0.120.0`.
## Tag convention
`v<semver>`, at the commit the binary was built from. `v0.120.0` was created retroactively at
`cd6e267` — the commit that produced the published binary (sha `a7763d31b55b5ce7…`). `configs/` is
byte-identical between that commit and `main`, so nothing about the sixteen fetched files depends on
the choice; `cd6e267` is tagged because it is the honest one.
+2 -2
View File
@@ -148,8 +148,8 @@
| `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,t)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test rotation state, persisted (atomic tmp+rename) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target) (string,error)` | The per-run restore-test spec + per-tier candidate lookup (R-85) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", nil)`**`""` is NOT an error**, or every fresh box looks broken for its first week. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,t)` / `ProvenArchive(target)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename)**which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)`**`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
| `localapi.StaleLockController` | internal/localapi/stalelock.go | `*staleLockController` (Client + Runner + pool) | `fakeStaleLock` (Server-level) stalelock_test.go; `fakeStaleLockAPI` (controller-level, tests the A1 pool intersect) stalelock_pool_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
+84 -6
View File
@@ -165,7 +165,7 @@ func main() {
showVersion bool
)
flag.StringVar(&cfgPath, "config", envOr("FELHOM_AGENT_CONFIG", "/etc/felhom-agent/agent.json"), "path to the agent config file (JSON)")
flag.Var(&selftest, "selftest", "run a self-test and exit: bare/`read` = read-only queries; `task` = reversible mutating exercise (needs -vmid); `hub` = one collect+report; `storage` = observe storage (+ -watch); `backup` = one-shot backup of -vmid; `restore-test` = restore→boot→verify→teardown of -archive (or newest backup); `pbs-verify` = trigger a PBS verify + print snapshot records; `bring-up` = restore→reset identity→size→start link-up of -archive into -vmid (needs -mode/-archive/-vmid; optional -cores/-memory cap; tears down unless -keep); `provision` = full slice-8A chain: bring-up provision + mint token + populate bootstrap config mount (needs -archive/-vmid/-customer-id/-hub-password; optional -rootfs-grow/-datavol-grow/-cores/-memory (-sysdata-grow is deprecated: folded into -datavol-grow); keeps the guest)")
flag.Var(&selftest, "selftest", "run a self-test and exit: bare/`read` = read-only queries; `task` = reversible mutating exercise (needs -vmid); `hub` = one collect+report; `storage` = observe storage (+ -watch); `backup` = one-shot backup of -vmid; `restore-test` = restore→boot→verify→teardown of -archive (or newest backup); `restore-test-due` = READ-ONLY: print the per-tier due verdict the scheduler would act on, with its cost; `pbs-verify` = trigger a PBS verify + print snapshot records; `bring-up` = restore→reset identity→size→start link-up of -archive into -vmid (needs -mode/-archive/-vmid; optional -cores/-memory cap; tears down unless -keep); `provision` = full slice-8A chain: bring-up provision + mint token + populate bootstrap config mount (needs -archive/-vmid/-customer-id/-hub-password; optional -rootfs-grow/-datavol-grow/-cores/-memory (-sysdata-grow is deprecated: folded into -datavol-grow); keeps the guest)")
flag.IntVar(&vmid, "vmid", 0, "guest VMID for --selftest=task|backup|bring-up")
flag.DurationVar(&watch, "watch", 0, "for --selftest=storage: run the watchdog verbose for this duration (e.g. 3m) with the re-mount response live; 0 = observe pass only")
flag.StringVar(&archive, "archive", "", "for --selftest=restore-test|bring-up: the backup volid to restore (restore-test: default newest on the local target)")
@@ -238,6 +238,8 @@ func main() {
os.Exit(runSelftestBackup(context.Background(), cfg, logger, vmid))
case "restore-test":
os.Exit(runSelftestRestoreTest(context.Background(), cfg, logger, archive))
case "restore-test-due":
os.Exit(runSelftestRestoreTestDue(context.Background(), cfg, logger))
case "pbs-verify":
os.Exit(runSelftestPBSVerify(context.Background(), cfg, logger))
case "lanresolver":
@@ -1273,13 +1275,21 @@ func primaryBackupTargetOf(cfg config.Config) func() hub.ConfiguredBackupTarget
// scratch band / restore storage is invalid — a misconfig must not crash the daemon, and the
// machinery still works on-demand via --selftest=restore-test.
func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *reconcile.Engine, store *backup.Store, rtState *backup.RestoreTestState, inFlight *backup.InFlight, logger *slog.Logger) *backup.Scheduler {
cadence := cfg.Backup.RestoreTestCadence()
// R-86: this is the EVALUATION interval, not the trigger. What decides a test happens is the
// per-archive due-check in internal/backup/restoretest_due.go.
cadence := cfg.Backup.RestoreTestEvalInterval()
if cadence > 0 {
if err := cfg.Backup.ValidateForRestoreTest(); err != nil {
logger.Warn("daemon: restore-test cadence disabled (config invalid)", "err", err)
logger.Warn("daemon: restore-test disabled (config invalid)", "err", err)
cadence = 0
}
}
if cadence > 0 && cfg.Backup.RestoreTestLegacyCadenceInUse() {
// Said ONCE, at start-up, naming both replacements: a key whose meaning changed under a box
// without a word is the silent repurposing R-86 §8.3 forbids.
logger.Warn("daemon: backup.restore_test_cadence_seconds is DEPRECATED — R-86 replaced the interval trigger with a per-archive due-check; this value now seeds the SETTLE lag only. Set backup.restore_test_settle_seconds and backup.restore_test_eval_interval_seconds explicitly",
"settle", cfg.Backup.RestoreTestSettle(), "eval_interval", cadence)
}
min, max := cfg.Backup.ScratchBand()
target := cfg.Backup.BackupTarget()
runner := backup.NewBackupRunner(px, target, "", "felhom restore-test", "", logger)
@@ -1315,13 +1325,18 @@ func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *re
}
},
Cadence: cadence,
Logger: logger,
// R-86: the settle lag — how long an archive must have sat before it is a candidate. With
// the per-archive due-check, this plus the archive rhythm is the whole schedule.
Settle: cfg.Backup.RestoreTestSettle(),
Logger: logger,
// R-85: rotate across EVERY configured tier, oldest-proven first (operator ruling, Option 1).
// Before this the scheduler only ever saw cfg.Backup.BackupTarget(), so the offsite tier's
// archives were never candidates and the DR tier went unproven for its whole existence.
// R-86 demoted that ordering to the tie-break BETWEEN DUE TIERS and widened this picker to
// the settle-aware one, which is what makes due-ness per archive generation.
Tiers: tierIDs,
TierPick: runner.PickRestoreCandidateOn,
TierPick: runner.PickSettledRestoreCandidateOn,
State: rtState,
InFlight: inFlight,
})
@@ -1751,6 +1766,67 @@ func runSelftestBackup(ctx context.Context, cfg config.Config, logger *slog.Logg
// running → teardown) of -archive (or the newest backup on the local target) into a scratch
// guest. Standalone (no hub). Runs engine.Recover first so a leaked scratch from a prior
// crashed test is reaped before this run.
// runSelftestRestoreTestDue prints the per-tier DUE verdict the scheduler would act on, and what
// each evaluation COST — read-only, so it is safe on any box at any time.
//
// It exists for two reasons R-86 needed and could not get from a log line. First, the due-check's
// verdict is the whole schedule now: "why did nothing run last night?" is answerable only by asking
// the same question the scheduler asks, against the same storages, in the same order. Second, the
// evaluation interval had to be chosen from a MEASURED cost rather than a guess — an offsite tier's
// candidate lookup crosses the WAN, and a monitoring loop that costs more than it is worth is how a
// check becomes the load. It reuses the daemon's own construction path (buildRestoreTestScheduler),
// so what it prints is what the daemon would decide, not a re-derivation of it.
func runSelftestRestoreTestDue(ctx context.Context, cfg config.Config, logger *slog.Logger) int {
if err := cfg.Validate(); err != nil {
fmt.Fprintln(os.Stderr, "selftest: proxmox not configured:", err)
return 1
}
px, err := newProxmoxClient(cfg)
if err != nil {
fmt.Fprintln(os.Stderr, "selftest: proxmox client:", err)
return 1
}
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
sched := buildRestoreTestScheduler(cfg, px, nil, backup.NewStore(), rtState, &backup.InFlight{}, logger)
fmt.Printf("eval_interval=%s settle=%s\n", cfg.Backup.RestoreTestEvalInterval(), cfg.Backup.RestoreTestSettle())
start := time.Now()
verdicts := sched.EvaluateDue(ctx)
total := time.Since(start)
if len(verdicts) == 0 {
fmt.Println("no tiers configured for restore-testing (or rotation not wired)")
return 0
}
rc := 0
for _, v := range verdicts {
proven, _ := rtState.ProvenArchive(v.Target)
fmt.Printf("tier=%-16s due=%-5v archive=%q landed=%s proven=%q\n reason: %s\n",
v.Target, v.Due, v.Archive, formatOrDash(v.Landed), proven, v.Reason)
if v.Err != nil {
// A tier we could not list is UNKNOWN, and it is a non-zero exit: an unreadable tier is
// a real condition, not a quiet "nothing to do".
fmt.Printf(" ERROR: %v\n", v.Err)
rc = 3
}
}
// Per-tier timing, measured one tier at a time so the WAN leg is attributable (R-86 Part 1.4).
for _, v := range verdicts {
t0 := time.Now()
_ = sched.EvaluateDueTier(ctx, v.Target)
fmt.Printf("cost tier=%-16s one_lookup=%s\n", v.Target, time.Since(t0).Round(time.Millisecond))
}
fmt.Printf("cost all_tiers=%s\n", total.Round(time.Millisecond))
return rc
}
// formatOrDash renders a time, or "-" when it is zero (no archive).
func formatOrDash(t time.Time) string {
if t.IsZero() {
return "-"
}
return t.UTC().Format(time.RFC3339)
}
func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog.Logger, archive string) int {
if err := cfg.Validate(); err != nil {
fmt.Fprintln(os.Stderr, "selftest: proxmox not configured:", err)
@@ -2940,6 +3016,8 @@ func (f *selftestFlag) Set(v string) error {
f.mode = "backup"
case "restore-test":
f.mode = "restore-test"
case "restore-test-due":
f.mode = "restore-test-due"
case "pbs-verify":
f.mode = "pbs-verify"
case "lanresolver":
@@ -2957,7 +3035,7 @@ func (f *selftestFlag) Set(v string) error {
case "controller-swap":
f.mode = "controller-swap"
default:
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|pbs-verify|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap)", v)
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap)", v)
}
return nil
}
@@ -0,0 +1,113 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-86 Scenario I — the seam-discipline test for the due-check.
//
// A due-check is worth nothing if the daemon still wires the OLD picker: every unit test in
// internal/backup would stay green (they inject the seam directly), the scheduler would ask for the
// newest archive with no settle cutoff, and the per-archive rule would run against a candidate that
// changes every time a backup lands. That is the same shape as the v0.91.0 inert seam — built,
// tested, never called — and this repo has shipped it four times.
//
// It walks main.go's AST rather than grepping: a commented-out call still satisfies a substring
// match, and a comment is not a caller.
func TestMainWiresTheSettleAwareTierPicker(t *testing.T) {
f := parseMainForWiring(t)
var settlePicker, oldPicker, settleWired, evalInterval bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.SelectorExpr:
// runner.PickSettledRestoreCandidateOn passed as a value (not called).
switch node.Sel.Name {
case "PickSettledRestoreCandidateOn":
settlePicker = true
case "PickRestoreCandidateOn":
oldPicker = true
}
case *ast.KeyValueExpr:
key, ok := node.Key.(*ast.Ident)
if !ok {
return true
}
if key.Name == "Settle" {
settleWired = true
}
case *ast.CallExpr:
if sel, ok := node.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "RestoreTestEvalInterval" {
evalInterval = true
}
}
return true
})
if !settlePicker {
t.Error("main.go never passes runner.PickSettledRestoreCandidateOn as the scheduler's TierPick — " +
"the due-check would run without a settle cutoff, i.e. against an archive that may still be being written")
}
if oldPicker {
t.Error("main.go still wires the pre-R-86 PickRestoreCandidateOn as a tier picker — " +
"two pickers means the one under test is not the one running")
}
if !settleWired {
t.Error("main.go never sets SchedulerOptions.Settle — the settle lag would default to 0 in the daemon " +
"and every freshly-landed archive would be an immediate candidate")
}
if !evalInterval {
t.Error("main.go never calls cfg.Backup.RestoreTestEvalInterval() — the scheduler would be driven by " +
"the retired cadence knob")
}
}
// The two R-85 guarantees the due-check must not have quietly dropped: the spec is still built PER
// RUN, and the shared heavy-operation gate is still handed to the scheduler.
func TestMainStillWiresTheHeavyOperationGateAndPerRunSpec(t *testing.T) {
f := parseMainForWiring(t)
var inFlightWired, specIsAFunc bool
ast.Inspect(f, func(n ast.Node) bool {
kv, ok := n.(*ast.KeyValueExpr)
if !ok {
return true
}
key, ok := kv.Key.(*ast.Ident)
if !ok {
return true
}
switch key.Name {
case "InFlight":
inFlightWired = true
case "Spec":
// A FuncLit means it is evaluated per run; anything else is a frozen value.
if _, isFunc := kv.Value.(*ast.FuncLit); isFunc {
specIsAFunc = true
}
}
return true
})
if !inFlightWired {
t.Error("main.go no longer hands the scheduler the shared InFlight gate — a restore-test could pull a " +
"multi-GB archive over the same tunnel an offsite backup is pushing one over (Scenario F)")
}
if !specIsAFunc {
t.Error("SchedulerOptions.Spec is no longer a function literal — a frozen spec is the R-85 defect " +
"(the tier and its timeout evaluated once at daemon start, forever)")
}
}
func parseMainForWiring(t *testing.T) *ast.File {
t.Helper()
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
return f
}
+24 -3
View File
@@ -135,10 +135,11 @@ func TestBackup_VzdumpFailureReturnsFailedRecord(t *testing.T) {
}
func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
const big = 4 << 30 // a plausible whole-guest archive
api := &fakeBackupAPI{content: []proxmox.StorageContent{
{VolID: "a", Content: "backup", CTime: 10},
{VolID: "b", Content: "backup", CTime: 99},
{VolID: "iso", Content: "iso", CTime: 999}, // not a backup → ignored
{VolID: "a", Content: "backup", CTime: 10, Size: big},
{VolID: "b", Content: "backup", CTime: 99, Size: big},
{VolID: "iso", Content: "iso", CTime: 999, Size: big}, // not a backup → ignored
}}
r := NewBackupRunner(api, "local", "", "", "", quiet())
vol, err := r.PickRestoreCandidate(context.Background())
@@ -152,6 +153,26 @@ func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
}
}
// R-86: the NEWEST entry is not a candidate if it cannot be a complete archive. An incomplete
// artefact (F-CRIT-2's 1-byte phantom, which server-side prune does not collect) would otherwise be
// picked forever, fail its restore forever, never earn proof, and so leave the tier due at every
// evaluation — turning the evaluation interval into the retry rate for a multi-GB restore.
//
// COMPANION RED-PROOF (observed): drop the `archivePlausiblyComplete` guard from
// PickSettledRestoreCandidateOn and this fails with
// `pick = "phantom" want the newest COMPLETE archive 'real'`.
func TestPickRestoreCandidate_SkipsImplausibleArchives(t *testing.T) {
api := &fakeBackupAPI{content: []proxmox.StorageContent{
{VolID: "real", Content: "backup", CTime: 10, Size: 4 << 30},
{VolID: "phantom", Content: "backup", CTime: 99, Size: 1}, // newest, and impossible
}}
r := NewBackupRunner(api, "local", "", "", "", quiet())
vol, err := r.PickRestoreCandidate(context.Background())
if err != nil || vol != "real" {
t.Fatalf("pick = %q,%v want the newest COMPLETE archive 'real'", vol, err)
}
}
// --- scheduler ---
type fakeRTRunner struct {
+142
View File
@@ -0,0 +1,142 @@
package backup
import (
"context"
"fmt"
"time"
)
// R-86 — a restore-test follows the BACKUP, not the clock.
//
// ── WHAT WAS WRONG ───────────────────────────────────────────────────────────────────────────
//
// The trigger was `time.NewTicker(cadence)` started at daemon start, and the tier was chosen by
// oldest-proven rotation. Its phase was therefore the PROCESS'S UPTIME: agent deploys are routine,
// so the test drifted to an arbitrary time of day every week; a fresh archive could sit unproven
// while an older one was re-tested; and a weekly tier was tested on the same rhythm as a daily one,
// sometimes twice on the same archive.
//
// ── THE RULE, AND THE TRAP IN ITS OBVIOUS FORM ───────────────────────────────────────────────
//
// R-86's ask reads "test a tier ~24 h after its own newest archive". Implemented literally —
// *"due when the newest archive is at least `settle` old"* — a DAILY tier is NEVER due: a new
// archive lands every day, so the newest archive's age resets to zero long before it reaches 24 h.
// The naive rule silently switches restore-testing off for the tier that matters most, and it is
// the version a reasonable person would write. It has a red-proof of its own
// (TestDue_NaiveNewestArchiveAgeRuleNeverFiresOnADailyTier).
//
// The rule implemented here:
//
// Let A = the newest archive on this tier that is at least `settle` old.
// The tier is DUE when A exists and A HAS NOT ALREADY BEEN PROVEN.
//
// daily tier → A is yesterday's archive; a new one settles each day → proved once per day
// weekly tier → A is last week's until the next settles → proved once per week
// newborn tier → A does not exist → UNKNOWN, never a fault
//
// Per-archive due-ness IS the pacing: one test per archive generation and no more. There is
// deliberately no second rate limiter on top of it (§8.4) — two independent pacing mechanisms
// produce a cadence nobody can predict from either.
//
// ── WHAT DID NOT CHANGE ──────────────────────────────────────────────────────────────────────
//
// The one-heavy-operation gate, the success-only proof credit, the oldest-proven ordering (now the
// tie-break between two DUE tiers), the restore-test itself, its journal and its scratch band. Only
// the trigger changed.
// DueVerdict is one tier's due-ness, and the evidence for it. Every field is logged: a due-check
// that cannot say WHY is a quiet path, and quiet paths are what this monitor family keeps shipping.
type DueVerdict struct {
Target string // the tier's storage target id
// Due is true only when Archive is set and has not been proven.
Due bool
// Archive is the settled candidate A ("" when the tier holds none).
Archive string
// Landed is when A landed on the tier (zero when Archive is "").
Landed time.Time
// ProvenArchive is what the state says was last proven on this tier ("" = nothing/legacy).
ProvenArchive string
// Err is a candidate-lookup failure. A tier whose archives cannot be listed is UNKNOWN — it is
// NEVER reported as "not due", which would silently retire a tier the moment its storage
// stopped answering. Due stays false (we have no archive to test) and the error travels.
Err error
// Reason is the one-line human account of this verdict.
Reason string
}
// String renders a verdict for the operator log / selftest output.
func (v DueVerdict) String() string {
return fmt.Sprintf("tier=%s due=%v archive=%q reason=%s", v.Target, v.Due, v.Archive, v.Reason)
}
// EvaluateDue returns the due verdict for every configured tier, ordered oldest-proven first.
//
// Ordering is the R-85 rotation, demoted to a TIE-BREAK: it no longer decides whether a test
// happens (due-ness does), only which of several due tiers goes first. Keeping it means a tier can
// still never be starved — a tier that has waited longest is served first — and keeping it as the
// order rather than as the trigger is the whole of this change.
func (s *Scheduler) EvaluateDue(ctx context.Context) []DueVerdict {
if !s.rotating() {
return nil
}
order := s.tiers
if s.rtState != nil {
order = s.rtState.OldestFirst(s.tiers)
}
cutoff := s.settleCutoff()
out := make([]DueVerdict, 0, len(order))
for _, target := range order {
out = append(out, s.evaluateTier(ctx, target, cutoff))
}
return out
}
// settleCutoff is the newest landing time an archive may have and still count as settled.
func (s *Scheduler) settleCutoff() time.Time {
if s.settle <= 0 {
return time.Time{} // no settle requirement configured → any archive is a candidate
}
return s.now().Add(-s.settle)
}
// evaluateTier is the per-tier due-check. PURE given the picker and the state, so the rule is
// unit-tested directly rather than inferred from whether a fake runner happened to be called.
func (s *Scheduler) evaluateTier(ctx context.Context, target string, cutoff time.Time) DueVerdict {
v := DueVerdict{Target: target}
archive, landed, err := s.tierPick(ctx, target, cutoff)
if err != nil {
// UNKNOWN, never "not due", and never silent.
v.Err = err
v.Reason = fmt.Sprintf("candidate lookup FAILED (%v) — tier is unknown this evaluation, not proven and not dismissed", err)
return v
}
v.Archive, v.Landed = archive, landed
if archive == "" {
v.Reason = "no settled archive yet — nothing to prove (newborn or still settling)"
return v
}
proven, ok := "", false
if s.rtState != nil {
proven, ok = s.rtState.ProvenArchive(target)
}
v.ProvenArchive = proven
if ok && proven == archive {
v.Reason = fmt.Sprintf("newest settled archive (landed %s) is already proven", landed.Format(time.RFC3339))
return v
}
v.Due = true
switch {
case !ok && proven == "":
v.Reason = fmt.Sprintf("newest settled archive (landed %s) has not been proven; nothing proven on this tier yet", landed.Format(time.RFC3339))
default:
v.Reason = fmt.Sprintf("newest settled archive (landed %s) has not been proven (last proven archive was a different one)", landed.Format(time.RFC3339))
}
return v
}
// EvaluateDueTier is EvaluateDue for ONE named tier — the selftest's per-tier cost probe, so the
// WAN leg of an offsite lookup is attributable rather than buried in an aggregate.
func (s *Scheduler) EvaluateDueTier(ctx context.Context, target string) DueVerdict {
return s.evaluateTier(ctx, target, s.settleCutoff())
}
+458
View File
@@ -0,0 +1,458 @@
package backup
import (
"context"
"errors"
"fmt"
"os"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-86 — the restore-test follows the BACKUP, not the clock.
//
// Every test here DRIVES time (`s.now` is injected and stepped) rather than waiting for it. A test
// that slept could not say anything about a 24-hour rule in under 24 hours, and one that only
// asserted "no error" would pass against a scheduler that never ran anything at all — which is
// precisely the failure mode §8.1's trap produces. So the assertions are: did a test run, on WHICH
// archive, and did a second evaluation correctly run NOTHING.
// ── the fake tier storage ────────────────────────────────────────────────────────────────────
// archiveStub is one archive on a tier: its volid and when it landed.
type archiveStub struct {
volid string
landed time.Time
}
// tierStorage is a TierPicker over per-tier archive lists. It implements the SAME contract as the
// production picker (*BackupRunner).PickSettledRestoreCandidateOn — newest archive that landed at
// or before the cutoff — which is itself covered against a fake PVE API in backup_test.go, and
// end-to-end by the live run. Naming the seam explicitly: everything below is true up to this
// picker; that the real picker obeys the same rule is asserted there, not here.
type tierStorage struct {
archives map[string][]archiveStub
err map[string]error // target → lookup failure
}
func (ts *tierStorage) pick(_ context.Context, target string, notAfter time.Time) (string, time.Time, error) {
if e, ok := ts.err[target]; ok && e != nil {
return "", time.Time{}, e
}
var best archiveStub
for _, a := range ts.archives[target] {
if !notAfter.IsZero() && a.landed.After(notAfter) {
continue // not settled yet
}
if best.volid == "" || a.landed.After(best.landed) {
best = a
}
}
return best.volid, best.landed, nil
}
// dueHarness is a scheduler with a driven clock over a fake tier storage.
type dueHarness struct {
s *Scheduler
rr *rotRunner
st *RestoreTestState
ts *tierStorage
clock time.Time
path string
}
func newDueHarness(t *testing.T, start time.Time, settle time.Duration, pass bool, tiers []string, ts *tierStorage) *dueHarness {
t.Helper()
return newDueHarnessAt(t, filepath.Join(t.TempDir(), "rt.json"), start, settle, pass, tiers, ts)
}
func newDueHarnessAt(t *testing.T, statePath string, start time.Time, settle time.Duration, pass bool, tiers []string, ts *tierStorage) *dueHarness {
t.Helper()
h := &dueHarness{rr: &rotRunner{pass: pass}, ts: ts, clock: start, path: statePath}
h.st = NewRestoreTestState(statePath)
h.s = NewScheduler(SchedulerOptions{
Runner: h.rr,
Store: NewStore(),
Spec: func(_ context.Context, archive string) reconcile.RestoreTestSpec {
return reconcile.RestoreTestSpec{RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009}
},
Cadence: time.Hour,
Settle: settle,
Logger: quiet(),
Tiers: tiers,
TierPick: ts.pick,
State: h.st,
InFlight: &InFlight{},
})
h.s.now = func() time.Time { return h.clock }
return h
}
// advance steps the clock by step, evaluating once at every step — the scheduler's real shape.
func (h *dueHarness) advance(step, total time.Duration) {
for elapsed := time.Duration(0); elapsed < total; elapsed += step {
h.clock = h.clock.Add(step)
h.s.tick(context.Background())
}
}
var day0 = time.Date(2026, 8, 1, 2, 0, 0, 0, time.UTC)
// dailyArchives lands one archive a day at 02:00 for n days, starting at day0.
func dailyArchives(tier string, n int) []archiveStub {
out := make([]archiveStub, 0, n)
for d := 0; d < n; d++ {
out = append(out, archiveStub{
volid: fmt.Sprintf("%s:backup/vzdump-lxc-9201-day%d.tar.zst", tier, d),
landed: day0.AddDate(0, 0, d),
})
}
return out
}
// ── SCENARIO A — a daily tier is proved daily, on its own archive ────────────────────────────
//
// THE TRAP THIS PINS (§8.1). R-86 reads "trigger a tier ~24 h after its own newest archive", and
// the literal implementation of that — *due when the newest archive is at least `settle` old* — is
// NEVER true on a daily tier: a new archive lands every day, so the newest archive's age resets to
// zero long before it reaches 24 h. The literal reading silently switches restore-testing OFF for
// the tier that matters most.
//
// COMPANION RED-PROOF (observed 2026-08-03). In Scheduler.evaluateTier, the per-archive comparison
// was replaced by the naive age rule:
//
// - if ok && proven == archive { … not due … }
// + if s.now().Sub(landed) < s.settle { … not due … } // and the proven-archive check deleted
//
// and the picker cutoff was removed (`cutoff := time.Time{}`), i.e. exactly "is the newest archive
// old enough". Result:
//
// --- FAIL: TestDue_DailyTierIsProvedDailyOnItsOwnArchive
// restoretest_due_test.go: a daily tier must be proved once per day; got 0 run(s) over 5 days
//
// Zero runs — restore-testing off. Restored immediately afterwards.
func TestDue_DailyTierIsProvedDailyOnItsOwnArchive(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"local": dailyArchives("local", 6)}}
h := newDueHarness(t, day0.Add(time.Hour), 24*time.Hour, true, []string{"local"}, ts)
// Five days, evaluated hourly.
h.advance(time.Hour, 5*24*time.Hour)
got := h.rr.seen()
if len(got) != 5 {
t.Fatalf("a daily tier must be proved once per day; got %d run(s) over 5 days: %v", len(got), got)
}
// And each run must be on the archive that settled that day — day0's on day 1, and so on.
for i, a := range got {
want := fmt.Sprintf("local:backup/vzdump-lxc-9201-day%d.tar.zst", i)
if a != want {
t.Fatalf("run %d tested %q, want %q — the test is not following the archive", i+1, a, want)
}
}
// The newest archive is NEVER the one tested: it has not settled.
if last := got[len(got)-1]; last == "local:backup/vzdump-lxc-9201-day5.tar.zst" {
t.Fatal("the still-settling archive was tested — the settle cutoff is not being applied")
}
}
// ── SCENARIO B — a weekly tier is proved weekly, not every other day ─────────────────────────
func TestDue_WeeklyTierIsProvedOncePerArchive(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"felhom-pbs": {
{volid: "felhom-pbs:backup/ct/9201/w0", landed: day0},
{volid: "felhom-pbs:backup/ct/9201/w1", landed: day0.AddDate(0, 0, 7)},
{volid: "felhom-pbs:backup/ct/9201/w2", landed: day0.AddDate(0, 0, 14)},
}}}
h := newDueHarness(t, day0.Add(time.Hour), 24*time.Hour, true, []string{"felhom-pbs"}, ts)
// Three weeks, evaluated every 6 hours — 84 evaluations.
h.advance(6*time.Hour, 21*24*time.Hour)
got := h.rr.seen()
want := []string{
"felhom-pbs:backup/ct/9201/w0",
"felhom-pbs:backup/ct/9201/w1",
"felhom-pbs:backup/ct/9201/w2",
}
if len(got) != len(want) {
t.Fatalf("a weekly tier must be proved ONCE PER ARCHIVE (3 archives over 3 weeks); got %d run(s): %v", len(got), got)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("run %d tested %q, want %q", i+1, got[i], want[i])
}
}
}
// ── SCENARIO C — an agent restart does not change the schedule ───────────────────────────────
//
// This is the defect a person actually notices: today every deploy restarts the ticker, so a
// restore-test runs one interval after each deploy regardless of what has already been proven.
//
// COMPANION RED-PROOF (observed 2026-08-03): revert the state to per-tier TIME by making
// ProvenArchive ignore the stored archive —
//
// - if !ok || p.Archive == "" { return "", false }
// + return "", false // per-tier time only, the pre-R-86 state
//
// → --- FAIL: TestDue_RestartRunsNothing
// restoretest_due_test.go:226: an agent restart must not trigger a restore-test; 2 restart(s)
// produced 4 run(s)
//
// Four: the same already-proven archive re-tested on EVERY evaluation after EVERY restart, which is
// today's behaviour with the ticker's phase reset by the deploy. Restored.
func TestDue_RestartRunsNothing(t *testing.T) {
dir := t.TempDir()
path := filepath.Join(dir, "rt.json")
ts := &tierStorage{archives: map[string][]archiveStub{"local": dailyArchives("local", 2)}}
start := day0.AddDate(0, 0, 1).Add(time.Hour) // day 1, 03:00 — day0's archive has settled
h := newDueHarnessAt(t, path, start, 24*time.Hour, true, []string{"local"}, ts)
h.s.tick(context.Background())
if n := len(h.rr.seen()); n != 1 {
t.Fatalf("precondition: the settled archive should have been proved once; got %d run(s)", n)
}
// --- two restarts: brand-new scheduler + brand-new state object over the SAME file ---
total := 0
for i := 0; i < 2; i++ {
h2 := newDueHarnessAt(t, path, start.Add(time.Duration(i+1)*time.Hour), 24*time.Hour, true, []string{"local"}, ts)
h2.s.tick(context.Background())
h2.s.tick(context.Background())
total += len(h2.rr.seen())
}
if total != 0 {
t.Fatalf("an agent restart must not trigger a restore-test; 2 restart(s) produced %d run(s)", total)
}
}
// ── SCENARIO D — a new archive makes a tier due even if it was tested yesterday ──────────────
func TestDue_NewSettledArchiveMakesAProvedTierDueAgain(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"local": dailyArchives("local", 2)}}
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(time.Hour), 24*time.Hour, true, []string{"local"}, ts)
h.s.tick(context.Background()) // proves day0's archive
h.s.tick(context.Background()) // nothing new has settled → nothing
if n := len(h.rr.seen()); n != 1 {
t.Fatalf("want exactly 1 run before the new archive settles, got %d: %v", n, h.rr.seen())
}
// Day 2, 03:00 — day1's archive has now settled.
h.clock = day0.AddDate(0, 0, 2).Add(time.Hour)
h.s.tick(context.Background())
got := h.rr.seen()
if len(got) != 2 {
t.Fatalf("a newly settled archive must make the tier due again; got %v", got)
}
if got[1] != "local:backup/vzdump-lxc-9201-day1.tar.zst" {
t.Fatalf("the NEW archive must be the one tested; got %q", got[1])
}
}
// ── SCENARIO E — a failing tier keeps being retried, and earns no proof ──────────────────────
//
// COMPANION RED-PROOF (observed 2026-08-03): give credit on failure in Scheduler.tick —
//
// - if rt.Pass && s.rtState != nil && target != "" {
// + if s.rtState != nil && target != "" {
//
// → --- FAIL: TestDue_FailingTierIsRetriedAndNeverProven
// restoretest_due_test.go: a failing tier must keep being retried; got 1 run(s) over 3
// evaluations
//
// A single failure would have retired the archive as proven — a permanently broken DR tier looking
// freshly verified, which is the loudest signal this system produces going silent. Restored.
func TestDue_FailingTierIsRetriedAndNeverProven(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"local": dailyArchives("local", 1)}}
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(time.Hour), 24*time.Hour, false, []string{"local"}, ts)
for i := 0; i < 3; i++ {
h.s.tick(context.Background())
}
got := h.rr.seen()
if len(got) != 3 {
t.Fatalf("a failing tier must keep being retried; got %d run(s) over 3 evaluations: %v", len(got), got)
}
if _, ok := h.st.ProvenArchive("local"); ok {
t.Fatal("a FAILED restore-test must not record the archive as proven")
}
if _, ok := h.st.LastSuccess("local"); ok {
t.Fatal("a FAILED restore-test must not stamp the tier as proven")
}
}
// ── SCENARIO F — two tiers due at once do not run at once ────────────────────────────────────
func TestDue_TwoDueTiersRunOneAtATime(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{
"local": {{volid: "local:backup/a.tar.zst", landed: day0}},
"felhom-pbs": {{volid: "felhom-pbs:backup/ct/9201/a", landed: day0}},
}}
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(time.Hour), 24*time.Hour, true, []string{"local", "felhom-pbs"}, ts)
// Both tiers are due at this instant.
due := h.s.EvaluateDue(context.Background())
if len(due) != 2 || !due[0].Due || !due[1].Due {
t.Fatalf("precondition: both tiers should be due; got %v", due)
}
h.s.tick(context.Background())
if n := len(h.rr.seen()); n != 1 {
t.Fatalf("ONE evaluation must start ONE restore-test, never two multi-GB restores over one link; got %d: %v", n, h.rr.seen())
}
// The other tier was DEFERRED, not cancelled: it is still due and runs on the next evaluation.
h.s.tick(context.Background())
got := h.rr.seen()
if len(got) != 2 || got[0] == got[1] {
t.Fatalf("the deferred tier must run on the NEXT evaluation, on its own archive; got %v", got)
}
}
// The heavy-operation gate still holds, and a tier deferred behind a backup stays DUE.
func TestDue_DeferredBehindABackupStaysDue(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"local": {{volid: "local:backup/a.tar.zst", landed: day0}}}}
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(time.Hour), 24*time.Hour, true, []string{"local"}, ts)
gate := &InFlight{}
h.s.inFlight = gate
release, _, _ := gate.TryAcquire("backup:felhom-pbs")
h.s.tick(context.Background())
if n := len(h.rr.seen()); n != 0 {
t.Fatalf("the restore-test must DEFER while a backup holds the gate; got %d run(s)", n)
}
if due := h.s.EvaluateDue(context.Background()); !due[0].Due {
t.Fatal("a deferred tier must remain DUE — deferral is not dismissal")
}
release()
h.s.tick(context.Background())
if n := len(h.rr.seen()); n != 1 {
t.Fatalf("must resume once the gate frees; got %d run(s)", n)
}
}
// ── SCENARIO H — a newborn box is UNKNOWN, not stale and not a fault ─────────────────────────
func TestDue_NewbornTierIsNotDueAndNotAnError(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"felhom-pbs": nil}}
h := newDueHarness(t, day0, 24*time.Hour, true, []string{"felhom-pbs"}, ts)
due := h.s.EvaluateDue(context.Background())
if len(due) != 1 {
t.Fatalf("want one verdict, got %v", due)
}
v := due[0]
if v.Due || v.Err != nil || v.Archive != "" {
t.Fatalf("a tier with no archive is UNKNOWN — not due, not an error; got %+v", v)
}
if v.Reason == "" {
t.Fatal("every verdict must carry a reason — a due-check that cannot say why is a quiet path")
}
h.s.tick(context.Background())
if n := len(h.rr.seen()); n != 0 {
t.Fatalf("a newborn tier must not be restore-tested; got %d run(s)", n)
}
}
// An archive that exists but has NOT settled yet is not a candidate — and that is not an error.
func TestDue_UnsettledArchiveIsNotACandidate(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"local": {{volid: "local:backup/fresh.tar.zst", landed: day0}}}}
h := newDueHarness(t, day0.Add(2*time.Hour), 24*time.Hour, true, []string{"local"}, ts)
if v := h.s.EvaluateDue(context.Background())[0]; v.Due || v.Archive != "" {
t.Fatalf("an archive 2h old must not be a candidate under a 24h settle lag; got %+v", v)
}
h.s.tick(context.Background())
if n := len(h.rr.seen()); n != 0 {
t.Fatalf("nothing settled → no run; got %d", n)
}
}
// A tier whose archives cannot be LISTED is UNKNOWN — never silently "not due", and never silent.
// Treating a lookup failure as "not due" would retire a tier the moment its storage stopped
// answering, which is the same absence-is-not-evidence error this monitor family keeps making.
func TestDue_LookupFailureIsUnknownNotNotDue(t *testing.T) {
boom := errors.New("storage unreachable")
ts := &tierStorage{
archives: map[string][]archiveStub{"local": {{volid: "local:backup/a.tar.zst", landed: day0}}},
err: map[string]error{"felhom-pbs": boom},
}
h := newDueHarness(t, day0.AddDate(0, 0, 1), 24*time.Hour, true, []string{"local", "felhom-pbs"}, ts)
var pbs DueVerdict
for _, v := range h.s.EvaluateDue(context.Background()) {
if v.Target == "felhom-pbs" {
pbs = v
}
}
if pbs.Err == nil {
t.Fatal("a lookup failure must travel in the verdict, not be swallowed")
}
if pbs.Due {
t.Fatal("a tier we could not list must not be reported DUE — we have no archive to test")
}
if pbs.Reason == "" {
t.Fatal("the failure must be explained, not merely flagged")
}
// And the OTHER tier still runs: one tier's storage being unreadable must not cost the other
// tier its proof.
h.s.tick(context.Background())
if got := h.rr.seen(); len(got) != 1 || got[0] != "local:backup/a.tar.zst" {
t.Fatalf("the readable tier must still be proved; got %v", got)
}
}
// ── the state's migration (§8.2) ─────────────────────────────────────────────────────────────
// A pre-R-86 state file carries a TIME and no archive. It must keep its time (rotation ordering
// survives the upgrade) and yield NO proven archive, so each tier is due exactly once. Reading a
// legacy time as proof of the CURRENT archive would mark an unproven archive proven — a guarantee
// invented by a migration.
func TestRestoreTestState_LegacyFileMigratesToNothingProven(t *testing.T) {
path := filepath.Join(t.TempDir(), "rt.json")
legacy := `{"local":"2026-08-01T02:00:00Z","felhom-pbs":"2026-07-30T02:00:00Z"}`
if err := writeFileForTest(path, legacy); err != nil {
t.Fatal(err)
}
st := NewRestoreTestState(path)
if _, ok := st.ProvenArchive("local"); ok {
t.Fatal("a legacy record names no archive — it must NOT be read as proof of the current one")
}
at, ok := st.LastSuccess("local")
if !ok || !at.Equal(time.Date(2026, 8, 1, 2, 0, 0, 0, time.UTC)) {
t.Fatalf("the legacy TIME must survive (rotation ordering depends on it); got %v ok=%v", at, ok)
}
// Ordering still works off the legacy times.
if got := st.OldestFirst([]string{"local", "felhom-pbs"}); got[0] != "felhom-pbs" {
t.Fatalf("oldest-first must still order legacy records; got %v", got)
}
}
// The new shape round-trips, archive and all.
func TestRestoreTestState_ArchiveRoundTrips(t *testing.T) {
path := filepath.Join(t.TempDir(), "rt.json")
now := time.Now().UTC().Truncate(time.Second)
st := NewRestoreTestState(path)
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", now); err != nil {
t.Fatal(err)
}
re := NewRestoreTestState(path)
got, ok := re.ProvenArchive("felhom-pbs")
if !ok || got != "felhom-pbs:backup/ct/9201/x" {
t.Fatalf("the proven ARCHIVE must survive a restart; got %q ok=%v", got, ok)
}
at, ok := re.LastSuccess("felhom-pbs")
if !ok || !at.Equal(now) {
t.Fatalf("the proven TIME must survive too; got %v ok=%v", at, ok)
}
}
// writeFileForTest is a tiny helper so the legacy-migration fixture reads clearly above.
func writeFileForTest(path, content string) error {
return os.WriteFile(path, []byte(content), 0o600)
}
+83 -19
View File
@@ -28,42 +28,92 @@ import (
//
// Only SUCCESS is recorded. A failed run must not satisfy rotation, or a tier that fails every time
// would look freshly proven and stop being retried — the same "a failure satisfies the cadence"
// trap the backup due-check avoids.
// trap the backup due-check avoids. R-86 keeps that property unchanged and gives it a second job:
// the due-check reads this state, so a failure that recorded proof would ALSO stop the tier from
// ever becoming due again. The rule earns its keep twice now.
//
// R-86 (1.2) — WHICH ARCHIVE, not just when.
//
// A timestamp alone cannot answer the question the due-check asks. "This tier passed at 04:00" is
// consistent both with "yesterday's archive is proven" and with "an archive from a week ago is
// proven and nothing since has been looked at". Restore-testing is now per ARCHIVE GENERATION —
// a tier is due once it holds a settled archive that has not been proven — so the identity of the
// proven archive is the state, and the time is metadata (rotation ordering, operator reporting).
//
// This is the same class as the workspace rule "a timestamp records an ATTEMPT, not a RESULT":
// here it records a result, but not WHICH result, and that is just as unable to answer the question
// being asked of it.
type RestoreTestState struct {
path string
mu sync.Mutex
last map[string]time.Time // target id → last SUCCESSFUL restore-test (UTC)
last map[string]provenTier // target id → what was last PROVEN on that tier
}
// provenTier is one tier's proof: the archive that passed, and when it passed.
type provenTier struct {
Archive string // volid of the archive that PASSED; "" = a legacy record with no archive
At time.Time // when that run passed (UTC)
}
// provenTierJSON is the on-disk shape (R-86). The legacy shape was a bare RFC3339 STRING per
// target; both are read, only this one is written — see NewRestoreTestState.
type provenTierJSON struct {
Archive string `json:"archive"`
ProvenAt string `json:"proven_at"`
}
// NewRestoreTestState opens (or creates) the state at path. A missing or unreadable file is NOT an
// error: it degrades to "nothing proven yet", which is the correct starting point and keeps a
// corrupt file from wedging the daemon.
//
// MIGRATION (R-86). The pre-R-86 file is `{"<target>": "<RFC3339>"}` — a time and no archive. A
// legacy record keeps its TIME (rotation ordering survives a deploy, which is why the file exists
// at all) but yields NO proven archive, so every tier is due exactly once on first evaluation after
// the upgrade. One extra restore-test per tier, once, is the safe direction: the alternative is to
// read a legacy time as proof of whatever archive happens to be current, which would mark an
// unproven archive proven — inventing a guarantee out of a migration.
func NewRestoreTestState(path string) *RestoreTestState {
s := &RestoreTestState{path: path, last: map[string]time.Time{}}
s := &RestoreTestState{path: path, last: map[string]provenTier{}}
data, err := os.ReadFile(path)
if err != nil {
return s
}
var raw map[string]string
var raw map[string]json.RawMessage
if json.Unmarshal(data, &raw) != nil {
return s
}
for target, ts := range raw {
if t, perr := time.Parse(time.RFC3339, ts); perr == nil {
s.last[target] = t.UTC()
for target, msg := range raw {
// Legacy shape: a bare RFC3339 string.
var legacy string
if json.Unmarshal(msg, &legacy) == nil {
if t, perr := time.Parse(time.RFC3339, legacy); perr == nil {
s.last[target] = provenTier{At: t.UTC()} // no archive → due once, deliberately
}
continue
}
var cur provenTierJSON
if json.Unmarshal(msg, &cur) != nil {
continue // one unreadable entry must not lose the others
}
t, perr := time.Parse(time.RFC3339, cur.ProvenAt)
if perr != nil {
continue
}
s.last[target] = provenTier{Archive: cur.Archive, At: t.UTC()}
}
return s
}
// RecordSuccess stamps a tier as proven at t. Only call this for a PASSING restore-test.
func (s *RestoreTestState) RecordSuccess(target string, t time.Time) error {
// RecordSuccess stamps a tier as proven at t, naming the ARCHIVE that passed. Only call this for a
// PASSING restore-test — the archive is what makes the tier not-due, so recording one for a failed
// run would retire the archive unproven.
func (s *RestoreTestState) RecordSuccess(target, archive string, t time.Time) error {
if target == "" {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
s.last[target] = t.UTC()
s.last[target] = provenTier{Archive: archive, At: t.UTC()}
return s.saveLocked()
}
@@ -71,17 +121,30 @@ func (s *RestoreTestState) RecordSuccess(target string, t time.Time) error {
func (s *RestoreTestState) LastSuccess(target string) (time.Time, bool) {
s.mu.Lock()
defer s.mu.Unlock()
t, ok := s.last[target]
return t, ok
p, ok := s.last[target]
return p.At, ok
}
// Snapshot returns a copy of the whole map — for the host-report gauge.
// ProvenArchive returns the archive last PROVEN on this tier (ok=false = none — either never tested,
// or a legacy record carrying only a time). It is the due-check's whole question: an archive that is
// not this one has not been proven.
func (s *RestoreTestState) ProvenArchive(target string) (string, bool) {
s.mu.Lock()
defer s.mu.Unlock()
p, ok := s.last[target]
if !ok || p.Archive == "" {
return "", false
}
return p.Archive, true
}
// Snapshot returns a copy of the last-proven TIMES — for the host-report gauge.
func (s *RestoreTestState) Snapshot() map[string]time.Time {
s.mu.Lock()
defer s.mu.Unlock()
out := make(map[string]time.Time, len(s.last))
for k, v := range s.last {
out[k] = v
out[k] = v.At
}
return out
}
@@ -100,8 +163,9 @@ func (s *RestoreTestState) OldestFirst(targets []string) []string {
defer s.mu.Unlock()
out := append([]string(nil), targets...)
sort.SliceStable(out, func(i, j int) bool {
ti, oki := s.last[out[i]]
tj, okj := s.last[out[j]]
pi, oki := s.last[out[i]]
pj, okj := s.last[out[j]]
ti, tj := pi.At, pj.At
switch {
case !oki && !okj:
return out[i] < out[j] // both never proven → deterministic
@@ -119,9 +183,9 @@ func (s *RestoreTestState) OldestFirst(targets []string) []string {
}
func (s *RestoreTestState) saveLocked() error {
raw := make(map[string]string, len(s.last))
for target, t := range s.last {
raw[target] = t.UTC().Format(time.RFC3339)
raw := make(map[string]provenTierJSON, len(s.last))
for target, p := range s.last {
raw[target] = provenTierJSON{Archive: p.Archive, ProvenAt: p.At.UTC().Format(time.RFC3339)}
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
+36 -17
View File
@@ -39,10 +39,19 @@ func (r *rotRunner) seen() []string {
return append([]string(nil), r.archives...)
}
// testLanded is a landing time old enough to be settled under any cutoff these tests use. R-86
// widened the TierPicker seam with the archive's landing time; the rotation tests below are about
// tier ORDER and the heavy-operation gate, not about settling, so they hold it constant.
var testLanded = time.Date(2026, 7, 1, 0, 0, 0, 0, time.UTC)
// archiveFor is a TierPicker over a fixed map: target → archive ("" = that tier holds none).
func archiveFor(m map[string]string) TierPicker {
return func(_ context.Context, target string) (string, error) {
return m[target], nil
return func(_ context.Context, target string, _ time.Time) (string, time.Time, error) {
a := m[target]
if a == "" {
return "", time.Time{}, nil
}
return a, testLanded, nil
}
}
@@ -63,14 +72,24 @@ func rotScheduler(t *testing.T, rr *rotRunner, st *RestoreTestState, pick TierPi
})
}
// ── SCENARIO A — both tiers get tested across consecutive cadences ───────────────────────────
// ── SCENARIO A — both tiers get tested, each ONCE per archive ────────────────────────────────
//
// R-86 CHANGED THIS TEST'S CONTRACT, deliberately, and the old assertion is worth recording because
// it was a faithful statement of the defect. It read:
//
// 4 ticks → 4 runs, and consecutive runs must hit different tiers
//
// i.e. every tick produced a heavy restore-test, because the ticker WAS the trigger. Under R-86 a
// tick is an EVALUATION: both tiers are still exercised (rotation is intact), but a tier whose
// newest settled archive is already proven is not re-tested just because time passed. So the
// assertion is now 2 runs across 4 evaluations — one per tier, one per archive — which is a
// STRICTLY STRONGER statement: it pins both the coverage R-85 won and the pacing R-86 adds.
//
// COMPANION RED-PROOF (observed): restore the single-target picker — set `Tiers`/`TierPick` to nil
// so `pickForThisRun` falls back to `s.pick` on the primary runner — and this fails with
// "both tiers must be exercised across 4 cadences; got [local:… local:… local:… local:…]",
// i.e. the offsite tier never appears. That is today's behaviour, and it is why demo-hp's DR tier
// went unproven for its entire existence.
func TestRotation_BothTiersExercisedAcrossCadences(t *testing.T) {
// "both tiers must be exercised; got [local:…]", i.e. the offsite tier never appears. That is
// pre-R-85 behaviour, and it is why demo-hp's DR tier went unproven for its entire existence.
func TestRotation_BothTiersExercisedOncePerArchive(t *testing.T) {
rr := &rotRunner{pass: true}
st := NewRestoreTestState(filepath.Join(t.TempDir(), "rt.json"))
s := rotScheduler(t, rr, st, archiveFor(map[string]string{
@@ -94,14 +113,14 @@ func TestRotation_BothTiersExercisedAcrossCadences(t *testing.T) {
}
}
if !sawLocal || !sawPBS {
t.Fatalf("both tiers must be exercised across 4 cadences; got %v", got)
t.Fatalf("both tiers must be exercised; got %v", got)
}
// Oldest-first must ALTERNATE, not clump — otherwise one tier is starved between visits.
if len(got) != 4 {
t.Fatalf("want 4 runs, got %d: %v", len(got), got)
// Exactly one run per tier: the archives never changed, so nothing became due a second time.
if len(got) != 2 {
t.Fatalf("want 2 runs across 4 evaluations (one per archive generation), got %d: %v", len(got), got)
}
if got[0] == got[1] {
t.Fatalf("consecutive runs hit the same tier — oldest-first is not rotating: %v", got)
t.Fatalf("the two runs must be different tiers — oldest-first is not ordering due tiers: %v", got)
}
}
@@ -284,11 +303,11 @@ func TestOldestFirst_Ordering(t *testing.T) {
t.Fatalf("unexpected: %v", got)
}
}
_ = st.RecordSuccess("local", now)
_ = st.RecordSuccess("local", "local:backup/a.tar.zst", now)
if got := st.OldestFirst([]string{"local", "felhom-pbs"}); got[0] != "felhom-pbs" {
t.Fatalf("a never-proven tier must sort before a proven one; got %v", got)
}
_ = st.RecordSuccess("felhom-pbs", now.Add(time.Hour))
_ = st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/b", now.Add(time.Hour))
if got := st.OldestFirst([]string{"local", "felhom-pbs"}); got[0] != "local" {
t.Fatalf("the least recently proven must sort first; got %v", got)
}
@@ -299,8 +318,8 @@ func TestOldestFirst_Ordering(t *testing.T) {
func TestOldestFirst_DeterministicOnTies(t *testing.T) {
st := NewRestoreTestState(filepath.Join(t.TempDir(), "rt.json"))
now := time.Now().UTC()
_ = st.RecordSuccess("b-tier", now)
_ = st.RecordSuccess("a-tier", now)
_ = st.RecordSuccess("b-tier", "b:archive", now)
_ = st.RecordSuccess("a-tier", "a:archive", now)
for i := 0; i < 20; i++ {
if got := st.OldestFirst([]string{"b-tier", "a-tier"}); got[0] != "a-tier" {
t.Fatalf("tie-break must be deterministic; iteration %d gave %v", i, got)
@@ -315,7 +334,7 @@ func TestRestoreTestState_PersistenceAndCorruption(t *testing.T) {
now := time.Now().UTC().Truncate(time.Second)
st := NewRestoreTestState(path)
if err := st.RecordSuccess("felhom-pbs", now); err != nil {
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", now); err != nil {
t.Fatal(err)
}
reopened := NewRestoreTestState(path)
+42 -4
View File
@@ -260,21 +260,59 @@ func (r *BackupRunner) PickRestoreCandidate(ctx context.Context) (string, error)
// restore is not an error** — a brand-new offsite tier legitimately has nothing yet, and turning
// that into a failure would make every fresh box look broken for its first week.
func (r *BackupRunner) PickRestoreCandidateOn(ctx context.Context, target string) (string, error) {
archive, _, err := r.PickSettledRestoreCandidateOn(ctx, target, time.Time{})
return archive, err
}
// PickSettledRestoreCandidateOn is the R-86 due-check's picker: the newest archive on target that
// landed AT OR BEFORE notAfter (the settle cutoff), with the time it landed. A zero notAfter means
// "no cutoff" — that is the pre-R-86 behaviour, which is why PickRestoreCandidateOn is now a
// one-line call into this and its contract is untouched (one scan, one owner).
//
// WHY A CUTOFF AT ALL. An archive that landed minutes ago may still be settling — R-71a's
// settle-gate exists because the offsite tier's day-0 consume raced its own floor update — and
// restore-testing the archive a backup is still writing proves nothing about the backup that
// finished. The due-check therefore asks about the newest SETTLED archive, and §8.1's rule is built
// on that: the tier is due when a settled archive exists that has not been proven.
//
// The plausibility floor is applied here and not in the old path on purpose. Under R-86 the picked
// archive becomes the tier's due-ness: an incomplete 1-byte phantom (F-CRIT-2's artefact — server
// prune does NOT collect it) would be selected forever, fail its restore forever, never earn proof,
// and so make the tier due at EVERY evaluation. Skipping it is what keeps the retry rate bounded by
// the archive generation rather than by the evaluation interval.
//
// Contract preserved: ("", zero, nil) when the storage holds no eligible archive. **A tier with
// nothing to restore is not an error** — a brand-new offsite tier legitimately has nothing yet, and
// turning that into a failure would make every fresh box look broken for its first week.
func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target string, notAfter time.Time) (string, time.Time, error) {
if target == "" {
return "", nil
return "", time.Time{}, nil
}
contents, err := r.api.StorageContent(ctx, target)
if err != nil {
return "", err
return "", time.Time{}, err
}
var best string
var bestCTime int64 = -1
for _, e := range contents {
if e.Content == "backup" && e.CTime > bestCTime {
if e.Content != "backup" {
continue
}
if !notAfter.IsZero() && e.CTime > notAfter.Unix() {
continue // not settled yet — a newer archive is not a reason to re-prove an older one
}
if ok, why := archivePlausiblyComplete(e); !ok {
r.warnRejectedArchiveOnce(e, why)
continue
}
if e.CTime > bestCTime {
bestCTime, best = e.CTime, e.VolID
}
}
return best, nil
if best == "" {
return "", time.Time{}, nil
}
return best, time.Unix(bestCTime, 0).UTC(), nil
}
// latestArchive finds the newest backup archive volid + size for vmid on the target.
+107 -59
View File
@@ -32,22 +32,31 @@ type CandidatePicker func(ctx context.Context) (string, error)
// PBS archive was classified "local" and got the 10-minute local wait.
type SpecBuilder func(ctx context.Context, archive string) reconcile.RestoreTestSpec
// TierPicker resolves the newest archive on a NAMED tier, or "" when that tier holds none.
// (*BackupRunner).PickRestoreCandidateOn satisfies it. "" must NOT be an error — a brand-new
// offsite tier legitimately has nothing to restore yet.
type TierPicker func(ctx context.Context, target string) (string, error)
// TierPicker resolves the newest archive on a NAMED tier that landed AT OR BEFORE notAfter (the
// settle cutoff), together with when it landed. (*BackupRunner).PickSettledRestoreCandidateOn
// satisfies it. A zero notAfter means "no settle requirement".
//
// R-86 widened this seam from (target) → archive. The landing time is what makes the due-check's
// verdict explainable — "archive X, which landed at T, has not been proven" — and the cutoff is
// what makes the rule per-ARCHIVE-GENERATION instead of per-interval. "" must NOT be an error: a
// brand-new offsite tier legitimately has nothing to restore yet.
type TierPicker func(ctx context.Context, target string, notAfter time.Time) (archive string, landed time.Time, err error)
// Scheduler runs the self-restore-test on an agent-internal cadence. It is the fourth daemon
// goroutine; it does real restore→boot→destroy, so it only runs when the cadence is enabled
// AND a valid scratch band is configured (validated by the caller before construction).
type Scheduler struct {
runner RestoreTestRunner
pick CandidatePicker
store *Store
spec SpecBuilder // R-85: evaluated PER RUN, never frozen at construction
runner RestoreTestRunner
pick CandidatePicker
store *Store
spec SpecBuilder // R-85: evaluated PER RUN, never frozen at construction
// cadence is the EVALUATION interval (R-86) — how often "is anything due?" is asked. It is no
// longer the thing that decides a test happens; see restoretest_due.go.
cadence time.Duration
logger *slog.Logger
now func() time.Time
// settle is how long an archive must have sat before it is a candidate (R-86).
settle time.Duration
logger *slog.Logger
now func() time.Time
// R-85 tier rotation. All optional: without them the scheduler behaves exactly as before
// (single tier via `pick`), which keeps every existing caller and test working untouched.
@@ -64,9 +73,14 @@ type SchedulerOptions struct {
Store *Store
// Spec builds the run's spec (RestoreStorage, ScratchMin/Max, SourceTier, timeouts) from the
// picked archive. Called ONCE PER RUN — see SpecBuilder for why it is not a value.
Spec SpecBuilder
Cadence time.Duration // 0 → disabled
Logger *slog.Logger
Spec SpecBuilder
// Cadence is the EVALUATION interval — how often due-ness is asked, NOT how often a test runs
// (R-86). 0 → disabled.
Cadence time.Duration
// Settle is how long an archive must have sat before it is a restore-test candidate (R-86).
// 0 → no settle requirement (any archive is a candidate).
Settle time.Duration
Logger *slog.Logger
// R-85 (all optional — omit for the pre-R-85 single-tier behaviour):
// Tiers are the configured tier target ids (primary first); TierPick resolves an archive on a
@@ -89,6 +103,7 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
store: opts.Store,
spec: opts.Spec,
cadence: opts.Cadence,
settle: opts.Settle,
logger: logger,
now: func() time.Time { return time.Now().UTC() },
tiers: append([]string(nil), opts.Tiers...),
@@ -98,9 +113,19 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
}
}
// Run fires a restore-test on the cadence until ctx is cancelled. A 0 cadence disables it
// (the goroutine just waits for shutdown). It does NOT fire immediately on start (a restore
// is heavy; the first runs one interval in) — on-demand runs use the selftest harness.
// Run EVALUATES due-ness on the interval until ctx is cancelled, and runs a restore-test only when
// a tier is actually due (R-86). A 0 interval disables it (the goroutine just waits for shutdown).
//
// The ticker survives as the evaluation interval and nothing else. It is emphatically NOT the
// trigger any more: its phase is the process's uptime, and agent deploys reset it, which is exactly
// the defect R-86 removes. What decides that a test happens is `EvaluateDue`.
//
// It still does NOT evaluate immediately on start — the first evaluation is one interval in. That
// is an EARNED restraint, kept deliberately: a restore is heavy, agent restarts are routine, and a
// crash-loop that evaluated at start would hammer a permanently-failing tier as fast as it could
// restart. Due-ness does not expire while we wait, so the only cost is up to one interval of
// latency on a tier that just became due. On-demand runs use `--selftest=restore-test`.
//
// Returns nil on ctx cancellation.
func (s *Scheduler) Run(ctx context.Context) error {
if s.cadence <= 0 || s.runner == nil || s.spec == nil || (s.pick == nil && !s.rotating()) {
@@ -108,7 +133,8 @@ func (s *Scheduler) Run(ctx context.Context) error {
<-ctx.Done()
return nil
}
s.logger.Info("backup: restore-test scheduler starting", "cadence", s.cadence)
s.logger.Info("backup: restore-test scheduler starting (per-archive due-check)",
"eval_interval", s.cadence, "settle", s.settle)
t := time.NewTicker(s.cadence)
defer t.Stop()
for {
@@ -122,8 +148,12 @@ func (s *Scheduler) Run(ctx context.Context) error {
}
}
// tick runs one scheduled restore-test: pick a backup → run → record. No-ops cleanly when
// no backup exists yet. Deterministic given s.now — tests call it directly.
// tick is ONE EVALUATION: gate → due-check → run the first due tier → record which archive was
// proven. No-ops cleanly when nothing is due, when no backup exists yet, or when a heavy operation
// is already in flight. Deterministic given s.now — tests call it directly.
//
// One run per evaluation, by construction (Scenario F): a second due tier is left DUE and picked up
// by the next evaluation. Deferred, never cancelled, and never two multi-GB restores over one link.
func (s *Scheduler) tick(ctx context.Context) {
if s.spec == nil {
// Defensive: Run() already refuses to start without a SpecBuilder, but tick is also
@@ -132,27 +162,40 @@ func (s *Scheduler) tick(ctx context.Context) {
s.logger.Error("backup: restore-test has no spec builder — skipping (this is a wiring bug)")
return
}
// Scenario F: join the one-heavy-operation-at-a-time gate. A restore-test PULLS a multi-GB
// archive over the same tunnel an offsite backup PUSHES one; running both saturates the link and
// drives each toward its timeout, which is how a healthy tier gets recorded as failed. DEFER —
// never cancel what is already running: a deferred restore-test costs hours of coverage, a
// cancelled backup costs the backup.
release, busy, ok := s.inFlight.TryAcquire("restore-test")
if !ok {
s.logger.Info("backup: restore-test deferred — a heavy operation is already in flight", "busy", busy)
return
}
defer release()
// The due-check runs BEFORE the gate is taken, and that ORDER is load-bearing under R-86.
//
// It used to be the other way round, and correctly so: the gate was held for one heavy run a
// day, and the candidate lookup rode along inside it. Evaluations are now frequent, and the
// lookup is a storage listing that for the offsite tier crosses the WAN. Holding the
// one-heavy-operation gate for a read that answers "nothing to do" would open a small window at
// EVERY evaluation in which a starting backup cannot acquire — and a backup that cannot acquire
// does not merely wait, it records a failure and pages the operator (F-A1). A cheap poll must
// not be able to manufacture that.
//
// Nothing is lost by checking first: due-ness does not expire, and the gate is still taken
// before anything heavy begins.
archive, target, err := s.pickForThisRun(ctx)
if err != nil {
s.logger.Warn("backup: restore-test could not pick a candidate; skipping", "err", err)
return
}
if archive == "" {
s.logger.Info("backup: restore-test skipped; no backup available yet")
s.logger.Debug("backup: restore-test not due this evaluation")
return
}
// Scenario F: join the one-heavy-operation-at-a-time gate. A restore-test PULLS a multi-GB
// archive over the same tunnel an offsite backup PUSHES one; running both saturates the link and
// drives each toward its timeout, which is how a healthy tier gets recorded as failed. DEFER —
// never cancel what is already running: a deferred restore-test costs hours of coverage, a
// cancelled backup costs the backup. A deferred tier stays DUE, so the next evaluation retries it.
release, busy, ok := s.inFlight.TryAcquire("restore-test")
if !ok {
s.logger.Info("backup: restore-test deferred — a heavy operation is already in flight",
"busy", busy, "target", target, "archive", archive)
return
}
defer release()
// R-85: build the spec for THIS run, from THIS archive. Never a frozen value.
spec := s.spec(ctx, archive)
spec.Archive = archive
@@ -165,8 +208,10 @@ func (s *Scheduler) tick(ctx context.Context) {
// Rotation credit is given ONLY on success. A failing tier must keep sorting first, or a tier
// that fails every time would look freshly proven and quietly stop being retried.
if rt.Pass && s.rtState != nil && target != "" {
if err := s.rtState.RecordSuccess(target, s.now()); err != nil {
s.logger.Warn("backup: could not persist the restore-test rotation state", "target", target, "err", err)
// R-86: the ARCHIVE is recorded, not merely the time — that is what makes the tier
// not-due until a NEWER archive settles, and what makes a proof survive a restart.
if err := s.rtState.RecordSuccess(target, archive, s.now()); err != nil {
s.logger.Warn("backup: could not persist the restore-test proof state", "target", target, "err", err)
}
}
switch {
@@ -189,46 +234,49 @@ func (s *Scheduler) tick(ctx context.Context) {
// rotating reports whether multi-tier rotation is wired.
func (s *Scheduler) rotating() bool { return len(s.tiers) > 0 && s.tierPick != nil }
// pickForThisRun chooses the tier and its newest archive.
// pickForThisRun chooses the tier to test THIS evaluation: the first DUE tier, in oldest-proven
// order.
//
// OLDEST-FIRST (operator ruling 2026-07-26, Option 1): the tier whose last SUCCESSFUL restore-test
// is oldest goes first, never-proven first of all. Self-balancing, no config knob, and it naturally
// prioritises a tier that has never been proven — which on this fleet was the offsite tier, unproven
// for its entire existence while reporting `applied`.
// R-86 changed what this answers. It used to answer "whose turn is it?", and the answer was always
// somebody's — the ticker had fired, so a test was going to happen. It now answers "is anything
// due?", and "nothing" is a normal, frequent and correct answer.
//
// A tier with no archives is SKIPPED, not failed, and the next tier is tried. Skipping to a testable
// tier is strictly better than burning the whole cadence: a brand-new offsite tier has nothing to
// restore yet, and that is normal, not broken. It cannot starve the empty tier either — as soon as
// it has an archive it still sorts first, because it is still the least recently proven.
// OLDEST-FIRST (operator ruling 2026-07-26, Option 1) survives as the ORDER among due tiers: the
// tier whose last successful restore-test is oldest goes first, never-proven first of all. It is
// self-balancing, needs no config knob, and it still cannot starve a tier — but it no longer decides
// that a test happens at all.
//
// Returns ("", "", nil) when nothing anywhere is testable.
// A tier with no settled archive is SKIPPED, not failed — a brand-new offsite tier has nothing to
// restore yet, and that is normal, not broken. A tier whose archives cannot be LISTED is likewise
// skipped, loudly, and its error is returned only when no other tier was testable: one tier's
// storage being unreadable must not cost the other tier its proof, and must not be silent either.
//
// Returns ("", "", nil) when nothing anywhere is due.
func (s *Scheduler) pickForThisRun(ctx context.Context) (archive, target string, err error) {
if !s.rotating() {
// Pre-R-85 single-tier path (tests and any caller that wires only `Pick`): there is no tier
// identity and no persisted proof here, so there is nothing to compare an archive against
// and no due-check is possible. It runs on every evaluation, exactly as it always did.
a, perr := s.pick(ctx)
return a, "", perr // pre-R-85 single-tier path; no rotation credit to record
}
order := s.tiers
if s.rtState != nil {
order = s.rtState.OldestFirst(s.tiers)
return a, "", perr
}
var firstErr error
for _, t := range order {
a, perr := s.tierPick(ctx, t)
if perr != nil {
// One tier's storage being unreadable must not block the others.
for _, v := range s.EvaluateDue(ctx) {
if v.Err != nil {
s.logger.Warn("backup: restore-test candidate lookup failed for a tier; trying the next",
"target", t, "err", perr)
"target", v.Target, "err", v.Err)
if firstErr == nil {
firstErr = perr
firstErr = v.Err
}
continue
}
if a == "" {
s.logger.Debug("backup: restore-test tier has no archive yet; trying the next", "target", t)
if !v.Due {
s.logger.Debug("backup: restore-test tier is not due", "target", v.Target, "reason", v.Reason)
continue
}
s.logger.Info("backup: restore-test tier selected (oldest-proven first)", "target", t, "archive", a)
return a, t, nil
s.logger.Info("backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)",
"target", v.Target, "archive", v.Archive, "landed", v.Landed.Format(time.RFC3339), "reason", v.Reason)
return v.Archive, v.Target, nil
}
if firstErr != nil {
return "", "", firstErr
+95 -13
View File
@@ -333,9 +333,23 @@ type BackupConfig struct {
LocalBackupTarget string `json:"local_backup_target"`
// RestoreStorage is where a restore-test's restored rootfs lands, e.g. "local-lvm".
RestoreStorage string `json:"restore_storage"`
// RestoreTestCadenceSeconds is the self-restore-test interval; 0 → default (24h).
// Set negative to DISABLE the automatic cadence (on-demand selftest still works).
// RestoreTestCadenceSeconds is the LEGACY restore-test knob, retained for one meaning only:
// NEGATIVE still DISABLES the automatic restore-test entirely (on-demand selftest still works),
// and 0 still means "use the default". It no longer sets how often a test runs — R-86 replaced
// the interval trigger with a per-archive due-check — so a positive value now seeds
// RestoreTestSettleSeconds instead (see RestoreTestSettle). Prefer the two explicit keys below.
RestoreTestCadenceSeconds int `json:"restore_test_cadence_seconds"`
// RestoreTestEvalIntervalSeconds is how often the scheduler ASKS whether any tier is due
// (R-86); 0 → default. It is not how often a test runs: a tier is tested once per archive
// generation no matter how often it is asked. This interval sets two things — the latency
// between an archive settling and its proof, and the retry rate of a tier whose restore-test
// keeps failing. See defaultRestoreTestEvalInterval for the measurement it was chosen from.
RestoreTestEvalIntervalSeconds int `json:"restore_test_eval_interval_seconds"`
// RestoreTestSettleSeconds is how long an archive must have sat on its tier before it is a
// restore-test candidate (R-86); 0 → default (24h), negative → 0 (no settle requirement).
// Restore-testing an archive a backup is still writing proves nothing about the backup that
// finished — this is the same settle discipline R-71a's gate applies to the offsite consume.
RestoreTestSettleSeconds int `json:"restore_test_settle_seconds"`
// ScratchVMIDMin/Max bound the throwaway restore-test scratch-guest VMID band. The
// restore-test refuses to run unless this is a valid band (min>0, max>=min); 9999 is
// always excluded. Defaults to 990000990009.
@@ -545,26 +559,92 @@ func (b BackupConfig) BackupTarget() string {
return defaultBackupTarget
}
// Default scratch VMID band + restore-test cadence.
// Default scratch VMID band + the two R-86 restore-test knobs.
const (
defaultScratchVMIDMin = 990000
defaultScratchVMIDMax = 990009
defaultRestoreTestCadence = 24 * time.Hour
defaultScratchVMIDMin = 990000
defaultScratchVMIDMax = 990009
// defaultRestoreTestEvalInterval is how often due-ness is ASKED. It is bounded from BOTH sides,
// and neither bound alone would have picked it:
//
// FLOOR — what one evaluation costs. MEASURED on demo-felhom, 2026-08-03 (R-86 Part 1.4), via
// --selftest=restore-test-due and by timing the underlying API call directly. One evaluation
// is one storage-content listing per tier:
//
// local dir storage (3 archives) ....... 18 ms (18.7 / 18.3 / 18.5)
// PBS tier, WAN to ep0 (2 snapshots) ... 392 ms (375 / 378 / 424)
// both tiers together .................. 430 ms
//
// So cost does NOT set this: even at one evaluation a minute the offsite leg would be ~0.7 %
// of a WAN link's time and ~9 minutes of ep0's day. Worth writing down anyway, because the
// number that would have forbidden a frequent poll is the one nobody measures.
//
// CEILING — the retry rate of a FAILING tier. Under a per-archive due-check a tier whose
// restore-test keeps failing stays due, so the evaluation interval IS its retry interval, and
// a retry is a multi-GB restore. Every few minutes would be an incident of its own; the old
// timer retried a broken tier once a day.
//
// 6h sits between them: four heavy retries a day at the very worst, latency from settle to
// proof of at most 6h against a 24h settle lag (so a daily tier is still proved daily), and no
// second rate limiter anywhere — the pacing remains one test per archive generation.
defaultRestoreTestEvalInterval = 6 * time.Hour
// defaultRestoreTestSettle is how long an archive must sit before it may be restore-tested.
// 24h is R-86's own figure ("~24 h after its own newest archive") and it is what makes the
// candidate on a daily tier YESTERDAY's archive rather than the one still being written.
defaultRestoreTestSettle = 24 * time.Hour
)
// RestoreTestCadence returns the configured restore-test interval: a positive value as-is,
// 0 → 24h default, negative → 0 (disabled).
func (b BackupConfig) RestoreTestCadence() time.Duration {
// RestoreTestEvalInterval returns how often the scheduler evaluates due-ness (R-86): a positive
// value as-is, 0 → the measured default, negative → 0 (disabled).
//
// The LEGACY `restore_test_cadence_seconds` keeps exactly one power here, the one a box may be
// relying on: a NEGATIVE value still disables the automatic restore-test outright. It no longer
// sets the interval, because the interval no longer decides that a test happens.
func (b BackupConfig) RestoreTestEvalInterval() time.Duration {
if b.RestoreTestCadenceSeconds < 0 {
return 0 // legacy DISABLE — preserved verbatim
}
switch {
case b.RestoreTestCadenceSeconds > 0:
return time.Duration(b.RestoreTestCadenceSeconds) * time.Second
case b.RestoreTestCadenceSeconds < 0:
case b.RestoreTestEvalIntervalSeconds > 0:
return time.Duration(b.RestoreTestEvalIntervalSeconds) * time.Second
case b.RestoreTestEvalIntervalSeconds < 0:
return 0 // disabled
default:
return defaultRestoreTestCadence
return defaultRestoreTestEvalInterval
}
}
// RestoreTestSettle returns how long an archive must have sat before it is a restore-test
// candidate (R-86): a positive value as-is, negative → 0 (no settle requirement), 0 → the default.
//
// WHAT HAPPENED TO THE OLD KEY. A box that set `restore_test_cadence_seconds` to a positive value
// was expressing "how long may pass between a backup and the confidence that it restores". That
// quantity survives R-86 as the SETTLE LAG, so a positive legacy value seeds this rather than being
// dropped or silently repurposed as the evaluation interval — and the daemon says so at start-up
// (see RestoreTestLegacyCadenceInUse). It is deliberately not carried into the evaluation interval:
// a box that set 72h to spare a weak endpoint would otherwise get a 72h-latency due-check, whereas
// what it actually wanted — fewer heavy restores — is what per-archive due-ness already gives it.
func (b BackupConfig) RestoreTestSettle() time.Duration {
switch {
case b.RestoreTestSettleSeconds > 0:
return time.Duration(b.RestoreTestSettleSeconds) * time.Second
case b.RestoreTestSettleSeconds < 0:
return 0 // explicitly no settle requirement
case b.RestoreTestCadenceSeconds > 0:
return time.Duration(b.RestoreTestCadenceSeconds) * time.Second // legacy seeding
default:
return defaultRestoreTestSettle
}
}
// RestoreTestLegacyCadenceInUse reports whether the deprecated key is what is deciding the settle
// lag, so the daemon can name both replacements ONCE at start-up. A config key that changed meaning
// without saying so is exactly the silent repurposing §8.3 forbids.
func (b BackupConfig) RestoreTestLegacyCadenceInUse() bool {
return b.RestoreTestCadenceSeconds > 0 && b.RestoreTestSettleSeconds == 0
}
// PBSVerifyCadence returns the verify-loop interval: positive as-is, 0 → 6h default,
// negative → 0 (disabled).
func (b BackupConfig) PBSVerifyCadence() time.Duration {
@@ -810,6 +890,8 @@ func applyEnv(cfg *Config) {
cfg.Backup.RestoreStorage = v
}
cfg.Backup.RestoreTestCadenceSeconds = envInt("FELHOM_AGENT_BACKUP_RESTORE_TEST_CADENCE_SECONDS", cfg.Backup.RestoreTestCadenceSeconds)
cfg.Backup.RestoreTestEvalIntervalSeconds = envInt("FELHOM_AGENT_BACKUP_RESTORE_TEST_EVAL_INTERVAL_SECONDS", cfg.Backup.RestoreTestEvalIntervalSeconds)
cfg.Backup.RestoreTestSettleSeconds = envInt("FELHOM_AGENT_BACKUP_RESTORE_TEST_SETTLE_SECONDS", cfg.Backup.RestoreTestSettleSeconds)
}
// envInt overlays an int env var, keeping cur (with a stderr warning) on parse
+7
View File
@@ -9,6 +9,12 @@
Gates (all must pass; **non-zero exit on any failure**):
1. reuse-refs every path cited by this repo's REUSE.md still resolves
2. published every `v<semver>` tag has a downloadable package AND a tag tree that serves
the agent's configs (R-115). NEEDS NETWORK, so it is **not** in `--fast` and
the pre-push hook does not run it a push must not fail because Gitea blinked
or because someone is offline on a train. CI runs the FULL set for exactly this
reason: it is the machine that can afford a network check, and it is the half
that emails when something is wrong.
WHY THIS FILE EXISTS, WITH ONE GATE (2026-08-02, R-29 leg (b)).
@@ -39,6 +45,7 @@ SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reus
# (label, absolute script path, args, fast)
GATES = [
("reuse-refs", SHARED_REUSE, [ROOT], True),
("published", os.path.join(ROOT, "scripts", "check-published-versions.py"), [], False),
]
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}
+159
View File
@@ -0,0 +1,159 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""check-published-versions.py — a released agent version must be INSTALLABLE, not merely built.
R-115. A box installs the agent from a Gitea generic package the hub vouches, never from git, and
since R-110/R-183 it also fetches the agent's sixteen config files from `raw/tag/v<version>/`.
Nothing in the build, deploy or session-end path checked that either existed, so "deployed" and
"installable" were independent states that drifted silently. **Three instances in five days:**
* R-111 (2026-07-29) 17 releases v0.97.0-v0.113.0 built and never published a new customer
would have installed without the whole R-82 tiered-backup arc, F-CRIT-2 and F-REBOOT.
* 0.114.0 (same afternoon) built, deployed to felhom-pve, never published.
* 0.120.0 (2026-08-03) built, committed, deployed to BOTH demo hosts, never published. A
documented-path reinstall would have silently DOWNGRADED both boxes to the pre-merge
agent and would have *succeeded* while doing it.
THE INVARIANT, AND WHY IT IS THIS ONE.
For every `v<semver>` git tag in this repo: the matching generic package must be DOWNLOADABLE,
and the tag must serve the agent's configs.
The task's §8.4 asked for a different one — *"the version the hub tells machines to install must be
downloadable"* — and that is the better invariant in principle. **It is not implementable from CI,
and that was measured rather than assumed:** the hub's artifact manifest
(`GET /api/v1/artifacts/<customer>`) answers **401** without a per-customer retrieval passphrase,
and the Gitea package LISTING api answers **401** without a token, while the package DOWNLOAD url
and the git TAGS api are both anonymous. So a credential-free gate can ask *"is this version
installable"* but not *"which version is vouched"*. Putting an operator credential into CI to close
that gap is the operator's call, not a gate author's it is recorded as a limitation below and as a
backlog row rather than quietly assumed away.
**What this invariant does catch: all three instances above.** `release-agent.sh` creates the tag and
publishes in one act, so a release whose publish was skipped, failed, or was forgotten leaves a tag
with no package which is exactly what this refuses. It needs no version floor: tags begin at
v0.120.0, which is published.
**What it does NOT catch, stated plainly:** the hub vouching a version that was never released at
all (no tag, no package). Nothing here can see that; it belongs at vouch time, in the hub. R-184.
FAIL-CLOSED. A network error, an unparseable response or an unreachable Gitea is exit **2
INCONCLUSIVE**, naming every URL tried never a pass. "Cannot determine" is not "fine": that is the
standing rule this project earned twice, and a gate that green-lights on its own blindness is worse
than no gate, because it looks like coverage.
Pure python3 + urllib, NO curl and no third-party module: the CI runner is a host-mode container
carrying python3 and git and nothing else, and an earlier workflow step died on
`curl: command not found`.
python3 scripts/check-published-versions.py
Exit: 0 every tag installable · 1 at least one is not · 2 could not be determined.
Env: GITEA_BASE overrides the Gitea root (CI sets the in-cluster service URL).
"""
import json
import os
import re
import sys
import urllib.error
import urllib.request
GITEA_BASE = os.environ.get("GITEA_BASE", "https://gitea.dooplex.hu").rstrip("/")
OWNER = "admin"
REPO = "felhom-agent"
PKG = "felhom-agent"
TIMEOUT = 25
# One config the installer fetches. Its presence proves the TAG's tree carries the configs the
# sixteen `fetch_raw` calls will ask for — a tag that exists but predates them would 404 a box
# mid-install, on a virgin machine, as root.
PROBE_CONFIG = "configs/felhom-agent.service"
TAG_RE = re.compile(r"^v(\d+\.\d+\.\d+)$")
tried = []
def _get(url, want_body=False):
"""GET a URL. Returns (status, body_or_None). Network failure raises."""
tried.append(url)
req = urllib.request.Request(url, method="GET")
try:
with urllib.request.urlopen(req, timeout=TIMEOUT) as r:
body = r.read() if want_body else None
return r.status, body
except urllib.error.HTTPError as e:
return e.code, None
def inconclusive(msg):
print("INCONCLUSIVE:", msg)
print(" URLs tried (a 'no access' claim must name its attempts):")
for u in tried:
print(" ", u)
sys.exit(2)
def main():
print("check-published-versions — every released agent version must be INSTALLABLE")
print(" gitea:", GITEA_BASE)
tags_url = "%s/api/v1/repos/%s/%s/tags?limit=200" % (GITEA_BASE, OWNER, REPO)
try:
status, body = _get(tags_url, want_body=True)
except Exception as e:
inconclusive("cannot reach Gitea to list tags: %s" % e)
if status != 200 or not body:
inconclusive("tags api returned HTTP %s — cannot enumerate releases" % status)
try:
tags = [t["name"] for t in json.loads(body.decode("utf-8"))]
except Exception as e:
inconclusive("tags api response is not the expected JSON: %s" % e)
versions = sorted(m.group(1) for m in (TAG_RE.match(t) for t in tags) if m)
if not versions:
# Not a failure: a repo legitimately has no version tags before its first release. Say so
# loudly rather than reporting a vacuous pass — an empty check that prints OK is how an
# unexamined thing becomes a documented-clean one.
print(" no v<semver> tags in this repo yet — nothing to check, and nothing proven")
print("\ncheck-published-versions: NOTHING TO CHECK")
return 0
print(" %d released version(s) to verify: %s" % (len(versions), ", ".join(versions)))
bad = []
for v in versions:
pkg_url = "%s/api/packages/%s/generic/%s/%s/%s" % (GITEA_BASE, OWNER, PKG, v, PKG)
raw_url = "%s/%s/%s/raw/tag/v%s/%s" % (GITEA_BASE, OWNER, REPO, v, PROBE_CONFIG)
try:
pkg_status, _ = _get(pkg_url)
raw_status, _ = _get(raw_url)
except Exception as e:
inconclusive("network failure while checking v%s: %s" % (v, e))
problems = []
if pkg_status != 200:
problems.append("binary NOT downloadable (HTTP %s at %s)" % (pkg_status, pkg_url))
if raw_status != 200:
problems.append("tag does not serve %s (HTTP %s) — a box would 404 mid-install"
% (PROBE_CONFIG, raw_status))
if problems:
bad.append((v, problems))
print(" FAIL v%s:" % v)
for p in problems:
print(" -", p)
else:
print(" ok v%s: binary downloadable + tag serves its configs" % v)
print()
if bad:
print("check-published-versions: %d RELEASED VERSION(S) NOT INSTALLABLE" % len(bad))
print(" A tagged version with no package is a release that was BUILT and never PUBLISHED —")
print(" the R-115 defect, three times in five days. Publish it with:")
print(" scripts/release-agent.sh <version>")
return 1
print("check-published-versions: ALL RELEASED VERSIONS INSTALLABLE")
return 0
if __name__ == "__main__":
sys.exit(main())
Regular → Executable
View File
+139
View File
@@ -0,0 +1,139 @@
#!/usr/bin/env bash
# release-agent.sh — THE way to release a felhom-agent version. One act: build → tag → publish →
# verify by independent download.
#
# WHY THIS EXISTS (R-115). Publishing used to be a step someone had to remember, and it was
# forgotten THREE TIMES IN FIVE DAYS:
#
# * R-111 (2026-07-29) 17 releases v0.97.0-v0.113.0 built and never published, so a new customer
# would have installed without the whole R-82 tiered-backup arc, F-CRIT-2 and F-REBOOT.
# * 0.114.0 (same afternoon) built, deployed to felhom-pve, never published.
# * 0.120.0 (2026-08-03) built, committed and deployed to BOTH demo hosts, never published. A
# documented-path reinstall would have silently DOWNGRADED both boxes to the pre-merge
# agent — and would have *succeeded* while doing it, because the current `step_grows`
# sets SYSDATA_GROW=0 so the older agent's fatal mp1 resize never fires.
#
# R-111's own closing line said publishing should join the release train rather than stay a
# remembered step. It closed SHIPPED without that leg, and the leg recurred the same afternoon —
# which is the evidence that a note is not a mechanism. This file is the mechanism. The
# documentation now points here instead of at a raw `go build` line, so there is ONE documented way
# to release and it cannot complete without publishing.
#
# WHY IT TAGS (R-183). Since felhom-host-install.sh pins its sixteen agent-config fetches to
# `raw/tag/v<version>`, a released version without a git tag 404s a box mid-install, as root, on a
# virgin machine. The tag and the package are two halves of one release and are created together.
#
# WHY IT DOES NOT VOUCH. Vouching is what points machines at a version, and it stays the operator's
# deliberate act — the same prove-then-vouch principle that governed the golden two sessions ago.
# This script prints the version and sha to vouch; a human decides when.
#
# Usage:
# GITEA_USER=admin GITEA_TOKEN=<token> ./scripts/release-agent.sh <version>
#
# Env: GITEA_USER/GITEA_TOKEN (package write) — same credentials publish-agent.sh already takes.
# GITEA_BASE / GITEA_OWNER override the defaults.
# RELEASE_ALLOW_DIRTY=1 skips the clean-tree gate (for a rehearsal; never for a real release).
set -euo pipefail
GITEA_BASE="${GITEA_BASE:-https://gitea.dooplex.hu}"
GITEA_OWNER="${GITEA_OWNER:-admin}"
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
die() { echo "[release-agent] ERROR: $*" >&2; exit 1; }
log() { echo "[release-agent] $*" >&2; }
VERSION="${1:-}"
[[ -n "$VERSION" ]] || die "version required (usage: GITEA_USER=.. GITEA_TOKEN=.. $0 <version>)"
[[ "$VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]] || die "version must be bare semver X.Y.Z (got '$VERSION')"
TAG="v$VERSION"
cd "$REPO_ROOT"
# ── 1. Clean-tree gate ──────────────────────────────────────────────────────────────────────────
# An unpushed change does not exist. Releasing a dirty tree publishes a binary whose source nobody
# else can obtain, and tags a commit that does not contain what was built.
if [[ "${RELEASE_ALLOW_DIRTY:-0}" != "1" ]]; then
[[ -z "$(git status --porcelain)" ]] || die "working tree is dirty — commit and push first"
local_head="$(git rev-parse HEAD)"
git fetch -q origin main
[[ "$local_head" == "$(git rev-parse origin/main)" ]] \
|| die "HEAD != origin/main — push first (an unpushed change does not exist)"
fi
# ── 2. Refuse to re-release a version that already exists ───────────────────────────────────────
# Silently overwriting a published artifact is how "the same version" comes to mean two different
# binaries on two different boxes.
if git rev-parse -q --verify "refs/tags/$TAG" >/dev/null; then
die "tag $TAG already exists — releasing over it would make one version name two binaries"
fi
existing="$(curl -fsS -o /dev/null -w '%{http_code}' \
"$GITEA_BASE/api/packages/$GITEA_OWNER/generic/felhom-agent/$VERSION/felhom-agent" 2>/dev/null || true)"
[[ "$existing" != "200" ]] || die "version $VERSION is ALREADY PUBLISHED — bump the version instead"
# ── 3. Build ────────────────────────────────────────────────────────────────────────────────────
BIN="$(mktemp -t felhom-agent-XXXXXX)"
trap 'rm -f "$BIN"' EXIT
log "building $VERSION"
go build -ldflags "-X main.version=$VERSION" -o "$BIN" ./cmd/felhom-agent \
|| die "go build failed"
built_ver="$("$BIN" --version 2>/dev/null | awk '{print $2}')"
[[ "$built_ver" == "$VERSION" ]] \
|| die "the built binary reports '$built_ver', not '$VERSION' — the ldflag did not take"
BUILT_SHA="$(sha256sum "$BIN" | awk '{print $1}')"
log "built ok: sha256 $BUILT_SHA"
# ── 4. Tag (before publishing, so a published version always has a tag) ─────────────────────────
# Order matters in this direction only: a tag with no package is caught by
# scripts/check-published-versions.py on the next CI run; a package with no tag is invisible to it,
# because the Gitea package LISTING api needs a token the gate does not have.
log "tagging $TAG at $(git rev-parse --short HEAD)"
git tag -a "$TAG" -m "agent $TAG
Released by scripts/release-agent.sh.
sha256 of the published binary: $BUILT_SHA
felhom-host-install.sh fetches this version's config files from raw/tag/$TAG/configs/,
so this tag is part of the released artifact, not a bookmark (R-183)."
git push origin "$TAG" || die "tag push failed — refusing to publish an untagged version"
# ── 5. Publish (the existing script; deliberately not reimplemented) ────────────────────────────
log "publishing …"
# Invoked through `bash` DELIBERATELY, not as an executable. On 2026-08-03 the first real release
# through this script died here — `publish-agent.sh` has been mode 0644 since it was created on
# 2026-06-28, because every earlier caller ran it as `bash scripts/publish-agent.sh`. So the one leg
# R-115 exists to make unforgettable was, on its first use, unrunnable. The mode bit is restored in
# the same commit; this line makes the release independent of it, because a file mode is exactly the
# kind of thing that is lost again by a checkout, an archive, or a copy.
bash "$REPO_ROOT/scripts/publish-agent.sh" "$VERSION" "$BIN" || die "publish failed"
# ── 6. Verify by an INDEPENDENT download ────────────────────────────────────────────────────────
# The publish step's own success is not proof: it reports on its own write. What matters is that a
# box can now GET the bytes and that they are the bytes that were built. This is the same
# presence-is-not-success rule the project earned twice — a step that says "done" and a fetch that
# returns the right sha are different claims.
log "verifying by independent download …"
DL="$(mktemp -t felhom-agent-dl-XXXXXX)"
trap 'rm -f "$BIN" "$DL"' EXIT
curl -fsS -o "$DL" "$GITEA_BASE/api/packages/$GITEA_OWNER/generic/felhom-agent/$VERSION/felhom-agent" \
|| die "round-trip GET failed — the version is NOT installable"
DL_SHA="$(sha256sum "$DL" | awk '{print $1}')"
[[ "$DL_SHA" == "$BUILT_SHA" ]] \
|| die "published sha $DL_SHA != built sha $BUILT_SHA — the artifact is not what was built"
# The tag must also serve the configs the installer will fetch from it.
cfg_code="$(curl -fsS -o /dev/null -w '%{http_code}' \
"$GITEA_BASE/$GITEA_OWNER/felhom-agent/raw/tag/$TAG/configs/felhom-agent.service" 2>/dev/null || true)"
[[ "$cfg_code" == "200" ]] \
|| die "tag $TAG does not serve configs/felhom-agent.service (HTTP $cfg_code) — a box would 404 mid-install"
cat <<EOF
RELEASED — and installable, verified by download, not by this script's own say-so.
version : $VERSION
tag : $TAG
sha256 : $BUILT_SHA
NOT VOUCHED. Vouching is what points machines at this version and stays your deliberate act:
hub operator UI → Configs → Day-0 artifacts. Until then boxes keep installing the previous one.
EOF