7581f8140a
gates / gates (push) Successful in 7s
All three are the reporting and release path misreporting its own work. No
customer machine, no backup, no restore, no data. The restore-test itself and
when it runs are unchanged.
R-189 — a passing restore-test no longer vanishes on a restart. restore_tests[]
came only from the in-memory store, whose comment ("lost on restart; the cadence
re-populates") was true under a timer and stopped being true when R-86 made the
agent refuse to re-test a proven archive: the proof is then not repeated for a
whole archive generation. Observed live — a 14.5 GB offsite PASS reached no
host-report because the agent was restarted 2m43s later. RestoreTestState now
carries tier + verified beside the archive and renders reportable entries; the
collector merges them, one per tier, newest by TestedAt. It refuses to lie: a
record missing archive-or-tier produces no entry, and run mechanics are not
re-invented. Only successes are persisted, and the asymmetry is now written where
it will be read.
R-188 — a correct release stops emailing a failure. Only the tag PUSH moved
(build -> tag locally -> publish -> push tag): the push wakes CI, and a tag
visible before its package made the gate correctly fail a correct release about
half the time. The old order's invariant is asserted directly instead — the gate
now refuses a published version with no tag, as a bounded probe that prints its
own coverage, because the package listing api is still 401 without a token.
R-186 — a released binary can be verified by rebuilding it. -trimpath
-buildvcs=false: same source, same bytes, tag or no tag. Measured. publish-agent's
fallback also forced CGO_ENABLED=0 and produced a 74 KB different binary for the
same version; both paths now build identically. CLAUDE.md records the command.
76 lines
2.8 KiB
Go
76 lines
2.8 KiB
Go
package backup
|
|
|
|
import (
|
|
"context"
|
|
"sync"
|
|
|
|
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
|
)
|
|
|
|
// Store holds the agent's LATEST backup result per target and the latest restore-test
|
|
// result — the point-in-time state the host-report surfaces. It is updated by the backup
|
|
// runner + the restore-test scheduler/selftest and read by the collector via the hub
|
|
// BackupReporter / RestoreTestReporter seams. In-memory and mutex-guarded for the concurrent
|
|
// collector vs scheduler access.
|
|
//
|
|
// **"lost on restart; the cadence re-populates" — that sentence used to be here and it is now
|
|
// FALSE for restore-tests (R-189, 2026-08-03).** It was true while a timer re-tested every tier
|
|
// daily. Under R-86's per-archive due-check the agent will NOT re-test an archive it has already
|
|
// proven, so a proof lost to a restart is not repeated until the next archive generation — a week on
|
|
// the offsite tier — and the hub reports that tier unproven throughout. Observed, not predicted: a
|
|
// real 14.5 GB offsite restore passed, the agent was restarted 2 m 43 s later for a deploy, and two
|
|
// consecutive host-reports carried `0 restore-tests`.
|
|
//
|
|
// The durable half is `RestoreTestState` (on disk, per tier, with the archive) and the collector
|
|
// merges the two — see hub.ProvenRestoreTestReporter. This store remains the ONLY place a FAILURE is
|
|
// recorded, and that asymmetry is deliberate: a failing tier stays due and is retried, so a lost
|
|
// failure heals itself, while a lost success leaves the system quietly less tested than it believes.
|
|
// Backups are unaffected — their freshness has a ground truth on the storage (R-84).
|
|
type Store struct {
|
|
mu sync.Mutex
|
|
byTarget map[string]hub.Backup // latest backup per target id
|
|
lastTest *hub.RestoreTest
|
|
}
|
|
|
|
// NewStore builds an empty Store.
|
|
func NewStore() *Store {
|
|
return &Store{byTarget: map[string]hub.Backup{}}
|
|
}
|
|
|
|
// RecordBackup stores the latest backup for its target.
|
|
func (s *Store) RecordBackup(b hub.Backup) {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
s.byTarget[b.TargetID] = b
|
|
}
|
|
|
|
// RecordRestoreTest stores the latest restore-test result.
|
|
func (s *Store) RecordRestoreTest(r hub.RestoreTest) {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
cp := r
|
|
s.lastTest = &cp
|
|
}
|
|
|
|
// Backups implements hub.BackupReporter — the latest backup per target (stable order by
|
|
// target id is not guaranteed; the hub does not depend on order).
|
|
func (s *Store) Backups(context.Context) []hub.Backup {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
out := make([]hub.Backup, 0, len(s.byTarget))
|
|
for _, b := range s.byTarget {
|
|
out = append(out, b)
|
|
}
|
|
return out
|
|
}
|
|
|
|
// RestoreTests implements hub.RestoreTestReporter — the latest restore-test result (0 or 1).
|
|
func (s *Store) RestoreTests(context.Context) []hub.RestoreTest {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
if s.lastTest == nil {
|
|
return []hub.RestoreTest{}
|
|
}
|
|
return []hub.RestoreTest{*s.lastTest}
|
|
}
|