Files
felhom-agent/REPORT.md
T
admin 0b28eae7bb
gates / gates (push) Successful in 6s
REPORT: R-189/R-188/R-186 — live evidence, the three sha values, and the observations
Scenario A proven on demo-felhom against the exact observation that filed R-189:
a 675 s offsite restore-test passed, the agent was restarted 11 seconds later
(inside the reporting window), and the hub's very next report carried
'1 restore-tests' where the same sequence produced 0 this morning. The hub's own
database holds the archive, the tier, the pass and the ORIGINAL test time, with
the run mechanics deliberately zero.

Also records the property the validation surfaced: the state holds one proof per
tier, so proving an older archive re-arms a newer one — confirmed live after the
defaults were restored.
2026-08-03 16:58:35 +02:00

17 KiB
Raw Blame History

REPORT — R-189 · R-188 · R-186: three ways the signals lied about themselves

Date: 2026-08-03 · Repo: felhom-agent v0.121.1 → v0.122.0 (7581f81) · released, published, verified by an independent download and by rebuilding it, deployed to demo-felhom. felhom.eu: register + docs only, no hub change and no hub bump — the hub already reads restore_tests[]; the defect was that the agent stopped sending them.


1. Baselines, re-read on arrival

Repo main @ commit Version Matched §1?
felhom-agent 3d0a1d615d11 v0.121.1 yes
felhom.eu c9a3e48b2106 hub v0.91.1 yes

Highest register ID in use was R-189; no new IDs were needed — all three rows already existed. (Grep confirmed R-190+ free, in case one had been.)

2. Scenario H — the reproducibility measurement (Part 3, done first on purpose)

Before, one commit, same source, same toolchain, same ldflags — the only difference is whether the tag existed when the build ran:

build sha256 size embedded module version
default flags, no tag yet 18f4a495… 14 085 464 B v0.121.2-0.20260803133646-3d0a1d61
default flags, tagged 4a38f394… 14 085 440 B v0.121.99
-trimpath -buildvcs=false, either way 7ffcdf1d… 14 064 574 B (none)

After, on the real release (v0.122.0) — the three values §15.2 asks for:

artifact sha256 size
published, downloaded from Gitea d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf 14 076 649 B
rebuild at the tag, #1 d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf 14 076 649 B
rebuild at the tag, #2 d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf 14 076 649 B

All three identical. The property is removed-cause, not sequenced-around: -buildvcs=false drops a stamp nothing reads (no ReadBuildInfo caller, verified by grep), and the version still comes from the explicit -X main.version ldflag. -trimpath additionally makes a rebuild from a different checkout directory match.

A second discrepancy fell out of the measurement and is fixed with it: publish-agent.sh's fallback build forced CGO_ENABLED=0 and produced 13 990 236 B against the release path's 14 064 574 B — a 74 KB difference, i.e. one version name meaning two binaries depending on which entry point ran. Both paths now use identical flags, each commented with a pointer to the other.

3. Scenario E — a correct release no longer emails a failure (Part 2)

Only the tag PUSH moved. The order is now build → tag locally → publish → push tag. The tag is still created before anything is published, so the build and the tag describe the same commit; it becomes visible — to CI (on: [push]) and to any raw/tag/… fetch — only once the package is downloadable.

The old ordering's invariant is asserted directly rather than arranged for. check-published-versions.py now carries two invariants: every tag has an installable package (as before) and no published version is missing its tag (new). The second is a bounded probe — the frontier past the newest tag, where a failed tag push leaves an orphan, plus patch gaps — and it prints its probe set on every run, because a check whose coverage is invisible reads as a guarantee it is not making. The package listing api was re-measured, not assumed: 401 without a token, so absence still cannot be enumerated, and the script says so in its own output.

Live result — this release is the test:

release CI runs on the release commit outcome
v0.121.0 (yesterday) #12 / #13 one green, one red
v0.121.1 (yesterday) #17 / #18 one red, one green
v0.122.0 (this one) #21 (task id 96) / #22 (task id 97) both green

Failure modes made loud rather than tidy: a publish that succeeds followed by a tag push that fails now dies printing git push origin v<ver> — the local tag is already there, so recovery is one line — and a publish that fails deletes the local-only tag so a retry is clean instead of colliding with step 2's re-release guard.

4. Scenarios F and G — both directions demonstrated, then cleaned up

G — a published version with no tag must FAIL. A real fixture: version 0.121.2 (the frontier, exactly where a failed tag push lands) published to the live registry with no tag.

0.121.2    next patch after the newest tag            PUBLISHED — NO TAG
check-published-versions: 1 PUBLISHED VERSION(S) WITH NO TAG
  v0.121.2 is downloadable at … but has no git tag.
      git push origin v0.121.2
EXIT=1

Red-proof (observed): with the fixture still live, replacing the probe set with [] produced ALL RELEASED VERSIONS INSTALLABLE, AND NONE UNTAGGED, exit 0 — a green run over a published orphan. Restored.

Teardown: fixture deleted (HTTP 204), absence independently re-verified (GET → 404), gate green again (exit 0). No scratch tag was ever pushed.

F — a tag with no package must still FAIL. Demonstrated against a local stand-in serving two tags where only one has a package:

  ok   v0.120.0: binary downloadable + tag serves its configs
  FAIL v0.199.0:
        - binary NOT downloadable (HTTP 404 …)
        - tag does not serve configs/felhom-agent.service (HTTP 404) — a box would 404 mid-install
EXIT=1

Why not a real pushed tag: pushing one wakes CI and would have emailed the operator a true alarm about a fixture — the same attention cost R-188 exists to remove. F is also already demonstrated in the wild: CI runs #13 and #17 failed for exactly this reason yesterday.

5. R-189 — a proof that survives a restart reaches the hub (Part 1)

restore_tests[] came only from the in-memory backup.Store, whose comment read "lost on restart; the cadence re-populates". True under a timer; false since R-86, because the agent refuses to re-test an archive it has already proven — so a lost proof is not repeated for a whole archive generation.

What changed

  • RestoreTestState stores the tier and what was verified beside the archive (v3 shape). Both are recorded at proof time from the run's own result — deriving them later would need a storage-type lookup at report-building time, a network call that can fail on the one path where failing means mislabelling a proof.
  • ProvenRestoreTests renders the stored proofs as report entries; Collector.SetProvenRestoreTests merges them with the in-memory result.
  • Merge rule: one entry per tier, newest by TestedAt wins. It falls out of what each source means rather than from a preference: a fresh failure beats a stored success (the failure is the news and lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice — the hub would read that as two tests. An unparseable timestamp counts as older, so a malformed entry cannot displace a good one.
  • It refuses to lie. A record missing the archive or the tier produces no entry, and run mechanics (scratch VMID, duration) are not re-invented: an absent duration is not a claim, a fabricated one would be.
  • The asymmetry is now in the code (§8.1): a success suppresses future work so it must be durable; a failure causes future work and heals itself, and persisting one would make a healed tier keep reporting a fault.
  • Store's comment is corrected in place — leaving it is how the next reader concludes this is handled.

Migration, and it is visible on the live box: a pre-R-189 record has an archive but no tier, so it is not reportable. Upgrading does not retroactively make an old proof visible; the tier's next real proof fills it in. Confirmed immediately after the deploy — still 0 restore-tests, with the v2 record sitting on disk.

6. Scenario A, live on demo-felhom — against the observation that filed R-189

The observation being replaced (2026-08-03, 15:25): a real offsite restore-test PASSED, the agent was restarted 2 m 43 s later, and the hub logged 0 restore-tests on the next two host-reports.

The same sequence, on v0.122.0:

16:44:06  restore-test tier is DUE   target=felhom-pbs
          archive=felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z
          reason="newest settled archive … has not been proven (last proven archive was a different one)"
16:55:21  restore-test: scratch guest torn down  vmid=990000
16:55:21  backup: scheduled restore-test PASSED  archive=felhom-pbs:…2026-07-27T19:55:41Z  duration_s=675.1
16:55:32  systemctl restart felhom-agent        ← INSIDE the 15-minute reporting window
16:55:36  hub: host-report from demo-felhom-8363b5 (… 1 restore-tests …)      ← was 0

The proof on disk (v3 — the tier is what the old shape lacked):

{"felhom-pbs": {"archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
                "tier": "pbs", "verified": "boot+running", "proven_at": "2026-08-03T14:55:21Z"}}

What the HUB stored — read from its own database (copied with its -wal, freshness confirmed by the newest row's received_at = 2026-08-03 14:55:36 UTC, matching the ingest line):

{ "source_archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
  "source_tier": "pbs", "pass": true, "verified": "boot+running",
  "tested_at": "2026-08-03T14:55:21Z", "scratch_vmid": 0, "duration_seconds": 0 }

That report was built after the restart, when the in-memory store was empty — so the entry can only have come from the persisted state. The archive, the tier, and the original test time survived; the run mechanics are zero because they are deliberately not re-invented.

A 14.5 GB encrypted offsite archive, restored, booted, verified and destroyed in 675 s — and this time the proof outlived the process that produced it.

Teardown, all three layers: scratch guest absent from pct list (0), its volumes gone from lvs (0), the validation drop-in removed and the daemon back on its defaults (eval_interval=6h0m0s settle=24h0m0s). The hub-side restore_tests[] record is retained deliberately — it is the proof the staleness check reads, so deleting it would delete the result. No restore_test_* event was raised, because nothing failed and nothing is stale.

7. Tests and red-proofs

Green gate: go build ./... && go vet ./... && go test ./...29 packages ok, rc=0, plus python3 scripts/agent_gates.py (reuse-refs + published-versions) all OK. The test run and the commit were always separate commands.

# Test Asserts Mutation Observed
A TestMerge_ProofSurvivesARestart an empty in-memory store + a persisted proof → the proof is reported, with its archive and its original time the persisted merge deleted (the pre-R-189 body) FAILafter a restart the persisted proof must be reported; got 0 entr(ies): [] — the live observation exactly
B TestMerge_NeverInventsAPassForAnUnprovenTier no proof → no entry; a tier-less record → no entry — (its state-layer twin below carries the mutation) pass
B TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe v1 + v2 + v3 records side by side → only the describable one is reported the reportable() filter dropped FAILgot 3 entries, two with an empty SourceTier/SourceArchive
C TestMerge_NewerWinsAndNeverDuplicatesATier one entry per tier, newest wins, in both directions de-duplication removed FAILone entry per tier; got 2 for "pbs" — the hub would read two tests
D TestMerge_AFailureIsStillReported a fresh failure beats an older stored success (same mutation) FAIL — 2 entries, i.e. the failure no longer the single answer for that tier
TestMerge_MalformedTimestampNeverWins unparseable ≠ newest pass
TestMerge_NilProvenSourceIsANoOp pre-R-189 behaviour unchanged when unwired pass
TestScheduler_ProofIsRecordedReportably a pass through the scheduler leaves a reportable proof pass
TestScheduler_AFailureLeavesNoPersistedProof §8.1's asymmetry, asserted not assumed pass
G the gate's converse assertion a published version with no tag fails probe set → [] FAIL (green over a live orphan)
I TestMainWiresTheDurableRestoreTestProof AST: SetProvenRestoreTests is called and fed rtState the call commented out FAILmain.go never calls collector.SetProvenRestoreTests (a strings.Contains check would have passed — the string is still there)
H reproducibility three identical sha256 — (measurement, §2) pass

Jitter, per §10: every timestamp fixture uses odd minutes and seconds (13:25:14, 19:55:41, 13:41:07, 04:41:58) — several taken from the real box — rather than round hours. Yesterday a test was hollow because a perfectly regular series landed exactly on a threshold and survived its own mutation.

8. Files changed

internal/backup/restoretest_state.go (v3 record + ProvenRestoreTests), internal/backup/store.go (the comment that had become false), internal/backup/schedule.go (record tier + verified), internal/hub/collect.go (the seam + the merge), cmd/felhom-agent/main.go (wiring), scripts/release-agent.sh (ordering, reproducible build, loud half-done release), scripts/publish-agent.sh (identical build flags), scripts/check-published-versions.py (the converse invariant), plus REUSE.md, CHANGELOG.md, CONTEXT.md, CLAUDE.md and three test files.

Commitsfelhom-agent: 7581f81 (v0.122.0). felhom.eu: see §10.

9. The independent-verification command (Part 3, recorded in CLAUDE.md)

V=0.122.0
git checkout "v$V" && go build -trimpath -buildvcs=false -ldflags "-X main.version=$V" \
    -o /tmp/felhom-agent-check ./cmd/felhom-agent
sha256sum /tmp/felhom-agent-check
curl -fsSL "https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/$V/felhom-agent" | sha256sum

Both print d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf (§2).

10. Registers

  • R-189 → CLOSED (shipped + proven live), R-188 → CLOSED (shipped), R-186 → CLOSED (shipped + measured).
  • R-185 remains OPEN and untouched — it is a missing /storage/felhom-backup ACL on demo-felhom, a permission defect, not a reporting one. Nothing in this session changed it, and the priority list says so explicitly.
  • No new IDs minted. ROADMAP.md contains none of these three rows, so there was nothing to collapse.
  • 00-capability-map.md's restore-proof row now records that the evidence path itself had a gap and what closed it; CONTEXT.md gains S-19 (the proof/failure asymmetry and the merge rule) and S-20 (the release ordering and what each step protects); STATUS.md rewritten for the operator and trimmed to 83 lines.

11. Observations — noticed, recorded, NOT acted on

  • The proof state holds ONE record per tier, so proving an OLDER archive re-arms a newer one — CONFIRMED after the validation, not merely predicted. With the defaults restored, the due-check reads: tier=felhom-pbs due=true archive="…2026-07-28T04:49:43Z" proven="…2026-07-27T19:55:41Z" — newest settled archive has not been proven (last proven archive was a different one). Surfaced by this session's own validation method: to get a fresh proof without waiting a week, the settle lag was widened so the older, unproven offsite archive became the candidate. That overwrote the record for the newer archive, so once the default 24 h settle returns, the newest settled archive is no longer the recorded proof and the tier becomes due once more. Consequence, stated rather than left to surprise: demo-felhom will run one further unattended offsite restore-test within 6 h, after which the newest archive is the recorded proof — the correct steady state. In normal operation this cannot arise, because the candidate only ever moves forward.
  • RestoreTestState.Snapshot() has no caller again. The host report is now fed by ProvenRestoreTests, which carries what a bare timestamp cannot. The method's doc comment says in as many words that it should be deleted if it does not acquire one — deliberately not deleted in this session, because removing an exported method is a change with no bearing on the three rows.
  • Ten files in this repo are not gofmt-clean and were already so on arrival (internal/capability/probe.go, internal/escrow/consume.go, internal/mgmtplane/mgmtplane.go, internal/reconcile/bringup.go, internal/signedjobs/runner.go, internal/storage/{candidates, intent}.go and three test files). Every file this session touched is clean; the others are untouched, and no gate checks formatting.
  • The hub sweeps every 60 s and re-reads 14 days of host-reports per customer for the restore-test staleness check. Unchanged here and not a defect at this fleet size; it is the cost centre if the fleet grows, and it is the reason the window read was deliberately left at 14 days yesterday.
  • felhom.eu/CONTEXT.md still carries duplicate standing-ruling IDs (three S-14s, two S-15s) from before yesterday. New rulings continue to be numbered above the collision (S-19, S-20) rather than adding to it; renumbering the existing ones is a separate, purely editorial change.