Scenario A proven on demo-felhom against the exact observation that filed R-189: a 675 s offsite restore-test passed, the agent was restarted 11 seconds later (inside the reporting window), and the hub's very next report carried '1 restore-tests' where the same sequence produced 0 this morning. The hub's own database holds the archive, the tier, the pass and the ORIGINAL test time, with the run mechanics deliberately zero. Also records the property the validation surfaced: the state holds one proof per tier, so proving an older archive re-arms a newer one — confirmed live after the defaults were restored.
17 KiB
REPORT — R-189 · R-188 · R-186: three ways the signals lied about themselves
Date: 2026-08-03 · Repo: felhom-agent v0.121.1 → v0.122.0 (7581f81) · released,
published, verified by an independent download and by rebuilding it, deployed to demo-felhom.
felhom.eu: register + docs only, no hub change and no hub bump — the hub already reads
restore_tests[]; the defect was that the agent stopped sending them.
1. Baselines, re-read on arrival
| Repo | main @ commit |
Version | Matched §1? |
|---|---|---|---|
felhom-agent |
3d0a1d615d11 |
v0.121.1 |
yes |
felhom.eu |
c9a3e48b2106 |
hub v0.91.1 |
yes |
Highest register ID in use was R-189; no new IDs were needed — all three rows already existed. (Grep confirmed R-190+ free, in case one had been.)
2. Scenario H — the reproducibility measurement (Part 3, done first on purpose)
Before, one commit, same source, same toolchain, same ldflags — the only difference is whether the tag existed when the build ran:
| build | sha256 | size | embedded module version |
|---|---|---|---|
| default flags, no tag yet | 18f4a495… |
14 085 464 B | v0.121.2-0.20260803133646-3d0a1d61 |
| default flags, tagged | 4a38f394… |
14 085 440 B | v0.121.99 |
-trimpath -buildvcs=false, either way |
7ffcdf1d… |
14 064 574 B | (none) |
After, on the real release (v0.122.0) — the three values §15.2 asks for:
| artifact | sha256 | size |
|---|---|---|
| published, downloaded from Gitea | d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf |
14 076 649 B |
| rebuild at the tag, #1 | d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf |
14 076 649 B |
| rebuild at the tag, #2 | d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf |
14 076 649 B |
All three identical. The property is removed-cause, not sequenced-around: -buildvcs=false drops
a stamp nothing reads (no ReadBuildInfo caller, verified by grep), and the version still comes from
the explicit -X main.version ldflag. -trimpath additionally makes a rebuild from a different
checkout directory match.
A second discrepancy fell out of the measurement and is fixed with it: publish-agent.sh's
fallback build forced CGO_ENABLED=0 and produced 13 990 236 B against the release path's
14 064 574 B — a 74 KB difference, i.e. one version name meaning two binaries depending on which
entry point ran. Both paths now use identical flags, each commented with a pointer to the other.
3. Scenario E — a correct release no longer emails a failure (Part 2)
Only the tag PUSH moved. The order is now build → tag locally → publish → push tag. The tag is
still created before anything is published, so the build and the tag describe the same commit; it
becomes visible — to CI (on: [push]) and to any raw/tag/… fetch — only once the package is
downloadable.
The old ordering's invariant is asserted directly rather than arranged for.
check-published-versions.py now carries two invariants: every tag has an installable package (as
before) and no published version is missing its tag (new). The second is a bounded probe —
the frontier past the newest tag, where a failed tag push leaves an orphan, plus patch gaps — and it
prints its probe set on every run, because a check whose coverage is invisible reads as a
guarantee it is not making. The package listing api was re-measured, not assumed: 401 without a
token, so absence still cannot be enumerated, and the script says so in its own output.
Live result — this release is the test:
| release | CI runs on the release commit | outcome |
|---|---|---|
| v0.121.0 (yesterday) | #12 / #13 | one green, one red |
| v0.121.1 (yesterday) | #17 / #18 | one red, one green |
| v0.122.0 (this one) | #21 (task id 96) / #22 (task id 97) | both green |
Failure modes made loud rather than tidy: a publish that succeeds followed by a tag push that
fails now dies printing git push origin v<ver> — the local tag is already there, so recovery is one
line — and a publish that fails deletes the local-only tag so a retry is clean instead of colliding
with step 2's re-release guard.
4. Scenarios F and G — both directions demonstrated, then cleaned up
G — a published version with no tag must FAIL. A real fixture: version 0.121.2 (the frontier, exactly where a failed tag push lands) published to the live registry with no tag.
0.121.2 next patch after the newest tag PUBLISHED — NO TAG
check-published-versions: 1 PUBLISHED VERSION(S) WITH NO TAG
v0.121.2 is downloadable at … but has no git tag.
git push origin v0.121.2
EXIT=1
Red-proof (observed): with the fixture still live, replacing the probe set with [] produced
ALL RELEASED VERSIONS INSTALLABLE, AND NONE UNTAGGED, exit 0 — a green run over a published
orphan. Restored.
Teardown: fixture deleted (HTTP 204), absence independently re-verified (GET → 404), gate green
again (exit 0). No scratch tag was ever pushed.
F — a tag with no package must still FAIL. Demonstrated against a local stand-in serving two tags where only one has a package:
ok v0.120.0: binary downloadable + tag serves its configs
FAIL v0.199.0:
- binary NOT downloadable (HTTP 404 …)
- tag does not serve configs/felhom-agent.service (HTTP 404) — a box would 404 mid-install
EXIT=1
Why not a real pushed tag: pushing one wakes CI and would have emailed the operator a true alarm about a fixture — the same attention cost R-188 exists to remove. F is also already demonstrated in the wild: CI runs #13 and #17 failed for exactly this reason yesterday.
5. R-189 — a proof that survives a restart reaches the hub (Part 1)
restore_tests[] came only from the in-memory backup.Store, whose comment read "lost on restart;
the cadence re-populates". True under a timer; false since R-86, because the agent refuses to re-test
an archive it has already proven — so a lost proof is not repeated for a whole archive generation.
What changed
RestoreTestStatestores the tier and what was verified beside the archive (v3 shape). Both are recorded at proof time from the run's own result — deriving them later would need a storage-type lookup at report-building time, a network call that can fail on the one path where failing means mislabelling a proof.ProvenRestoreTestsrenders the stored proofs as report entries;Collector.SetProvenRestoreTestsmerges them with the in-memory result.- Merge rule: one entry per tier, newest by
TestedAtwins. It falls out of what each source means rather than from a preference: a fresh failure beats a stored success (the failure is the news and lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice — the hub would read that as two tests. An unparseable timestamp counts as older, so a malformed entry cannot displace a good one. - It refuses to lie. A record missing the archive or the tier produces no entry, and run mechanics (scratch VMID, duration) are not re-invented: an absent duration is not a claim, a fabricated one would be.
- The asymmetry is now in the code (§8.1): a success suppresses future work so it must be durable; a failure causes future work and heals itself, and persisting one would make a healed tier keep reporting a fault.
Store's comment is corrected in place — leaving it is how the next reader concludes this is handled.
Migration, and it is visible on the live box: a pre-R-189 record has an archive but no tier, so it
is not reportable. Upgrading does not retroactively make an old proof visible; the tier's next
real proof fills it in. Confirmed immediately after the deploy — still 0 restore-tests, with the
v2 record sitting on disk.
6. Scenario A, live on demo-felhom — against the observation that filed R-189
The observation being replaced (2026-08-03, 15:25): a real offsite restore-test PASSED, the agent
was restarted 2 m 43 s later, and the hub logged 0 restore-tests on the next two host-reports.
The same sequence, on v0.122.0:
16:44:06 restore-test tier is DUE target=felhom-pbs
archive=felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z
reason="newest settled archive … has not been proven (last proven archive was a different one)"
16:55:21 restore-test: scratch guest torn down vmid=990000
16:55:21 backup: scheduled restore-test PASSED archive=felhom-pbs:…2026-07-27T19:55:41Z duration_s=675.1
16:55:32 systemctl restart felhom-agent ← INSIDE the 15-minute reporting window
16:55:36 hub: host-report from demo-felhom-8363b5 (… 1 restore-tests …) ← was 0
The proof on disk (v3 — the tier is what the old shape lacked):
{"felhom-pbs": {"archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
"tier": "pbs", "verified": "boot+running", "proven_at": "2026-08-03T14:55:21Z"}}
What the HUB stored — read from its own database (copied with its -wal, freshness confirmed by
the newest row's received_at = 2026-08-03 14:55:36 UTC, matching the ingest line):
{ "source_archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
"source_tier": "pbs", "pass": true, "verified": "boot+running",
"tested_at": "2026-08-03T14:55:21Z", "scratch_vmid": 0, "duration_seconds": 0 }
That report was built after the restart, when the in-memory store was empty — so the entry can only have come from the persisted state. The archive, the tier, and the original test time survived; the run mechanics are zero because they are deliberately not re-invented.
A 14.5 GB encrypted offsite archive, restored, booted, verified and destroyed in 675 s — and this time the proof outlived the process that produced it.
Teardown, all three layers: scratch guest absent from pct list (0), its volumes gone from lvs
(0), the validation drop-in removed and the daemon back on its defaults
(eval_interval=6h0m0s settle=24h0m0s). The hub-side restore_tests[] record is retained
deliberately — it is the proof the staleness check reads, so deleting it would delete the result.
No restore_test_* event was raised, because nothing failed and nothing is stale.
7. Tests and red-proofs
Green gate: go build ./... && go vet ./... && go test ./... — 29 packages ok, rc=0, plus
python3 scripts/agent_gates.py (reuse-refs + published-versions) all OK. The test run and the commit
were always separate commands.
| # | Test | Asserts | Mutation | Observed |
|---|---|---|---|---|
| A | TestMerge_ProofSurvivesARestart |
an empty in-memory store + a persisted proof → the proof is reported, with its archive and its original time | the persisted merge deleted (the pre-R-189 body) | FAIL — after a restart the persisted proof must be reported; got 0 entr(ies): [] — the live observation exactly |
| B | TestMerge_NeverInventsAPassForAnUnprovenTier |
no proof → no entry; a tier-less record → no entry | — (its state-layer twin below carries the mutation) | pass |
| B′ | TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe |
v1 + v2 + v3 records side by side → only the describable one is reported | the reportable() filter dropped |
FAIL — got 3 entries, two with an empty SourceTier/SourceArchive |
| C | TestMerge_NewerWinsAndNeverDuplicatesATier |
one entry per tier, newest wins, in both directions | de-duplication removed | FAIL — one entry per tier; got 2 for "pbs" — the hub would read two tests |
| D | TestMerge_AFailureIsStillReported |
a fresh failure beats an older stored success | (same mutation) | FAIL — 2 entries, i.e. the failure no longer the single answer for that tier |
| — | TestMerge_MalformedTimestampNeverWins |
unparseable ≠ newest | — | pass |
| — | TestMerge_NilProvenSourceIsANoOp |
pre-R-189 behaviour unchanged when unwired | — | pass |
| — | TestScheduler_ProofIsRecordedReportably |
a pass through the scheduler leaves a reportable proof | — | pass |
| — | TestScheduler_AFailureLeavesNoPersistedProof |
§8.1's asymmetry, asserted not assumed | — | pass |
| G | the gate's converse assertion | a published version with no tag fails | probe set → [] |
FAIL (green over a live orphan) |
| I | TestMainWiresTheDurableRestoreTestProof |
AST: SetProvenRestoreTests is called and fed rtState |
the call commented out | FAIL — main.go never calls collector.SetProvenRestoreTests (a strings.Contains check would have passed — the string is still there) |
| H | reproducibility | three identical sha256 | — (measurement, §2) | pass |
Jitter, per §10: every timestamp fixture uses odd minutes and seconds (13:25:14, 19:55:41,
13:41:07, 04:41:58) — several taken from the real box — rather than round hours. Yesterday a test
was hollow because a perfectly regular series landed exactly on a threshold and survived its own
mutation.
8. Files changed
internal/backup/restoretest_state.go (v3 record + ProvenRestoreTests), internal/backup/store.go
(the comment that had become false), internal/backup/schedule.go (record tier + verified),
internal/hub/collect.go (the seam + the merge), cmd/felhom-agent/main.go (wiring),
scripts/release-agent.sh (ordering, reproducible build, loud half-done release),
scripts/publish-agent.sh (identical build flags), scripts/check-published-versions.py (the
converse invariant), plus REUSE.md, CHANGELOG.md, CONTEXT.md, CLAUDE.md and three test files.
Commits — felhom-agent: 7581f81 (v0.122.0). felhom.eu: see §10.
9. The independent-verification command (Part 3, recorded in CLAUDE.md)
V=0.122.0
git checkout "v$V" && go build -trimpath -buildvcs=false -ldflags "-X main.version=$V" \
-o /tmp/felhom-agent-check ./cmd/felhom-agent
sha256sum /tmp/felhom-agent-check
curl -fsSL "https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/$V/felhom-agent" | sha256sum
Both print d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf (§2).
10. Registers
- R-189 → CLOSED (shipped + proven live), R-188 → CLOSED (shipped), R-186 → CLOSED (shipped + measured).
- R-185 remains OPEN and untouched — it is a missing
/storage/felhom-backupACL on demo-felhom, a permission defect, not a reporting one. Nothing in this session changed it, and the priority list says so explicitly. - No new IDs minted.
ROADMAP.mdcontains none of these three rows, so there was nothing to collapse. 00-capability-map.md's restore-proof row now records that the evidence path itself had a gap and what closed it;CONTEXT.mdgains S-19 (the proof/failure asymmetry and the merge rule) and S-20 (the release ordering and what each step protects);STATUS.mdrewritten for the operator and trimmed to 83 lines.
11. Observations — noticed, recorded, NOT acted on
- The proof state holds ONE record per tier, so proving an OLDER archive re-arms a newer one —
CONFIRMED after the validation, not merely predicted. With the defaults restored, the due-check
reads:
tier=felhom-pbs due=true archive="…2026-07-28T04:49:43Z" proven="…2026-07-27T19:55:41Z" — newest settled archive has not been proven (last proven archive was a different one). Surfaced by this session's own validation method: to get a fresh proof without waiting a week, the settle lag was widened so the older, unproven offsite archive became the candidate. That overwrote the record for the newer archive, so once the default 24 h settle returns, the newest settled archive is no longer the recorded proof and the tier becomes due once more. Consequence, stated rather than left to surprise: demo-felhom will run one further unattended offsite restore-test within 6 h, after which the newest archive is the recorded proof — the correct steady state. In normal operation this cannot arise, because the candidate only ever moves forward. RestoreTestState.Snapshot()has no caller again. The host report is now fed byProvenRestoreTests, which carries what a bare timestamp cannot. The method's doc comment says in as many words that it should be deleted if it does not acquire one — deliberately not deleted in this session, because removing an exported method is a change with no bearing on the three rows.- Ten files in this repo are not
gofmt-clean and were already so on arrival (internal/capability/probe.go,internal/escrow/consume.go,internal/mgmtplane/mgmtplane.go,internal/reconcile/bringup.go,internal/signedjobs/runner.go,internal/storage/{candidates, intent}.goand three test files). Every file this session touched is clean; the others are untouched, and no gate checks formatting. - The hub sweeps every 60 s and re-reads 14 days of host-reports per customer for the restore-test staleness check. Unchanged here and not a defect at this fleet size; it is the cost centre if the fleet grows, and it is the reason the window read was deliberately left at 14 days yesterday.
felhom.eu/CONTEXT.mdstill carries duplicate standing-ruling IDs (threeS-14s, twoS-15s) from before yesterday. New rulings continue to be numbered above the collision (S-19, S-20) rather than adding to it; renumbering the existing ones is a separate, purely editorial change.