hub v0.77.0 — R-85 Part 2: a restore-test result becomes a SIGNAL

A failed restore-test was a [WARN] line in the ingest handler and nothing else —
no event, no notification, no gauge. True for the LOCAL tier that was already
being tested, so the loudest DR signal this system produces was inaudible.
Rotating tiers without this would only mean two tiers can fail silently
instead of one.

Two signals, deliberately NOT merged:
  restore_test_failed (error)   — a run completed and did NOT pass
  restore_test_stale (warning)  — a tier not PROVEN within its interval
Merging them collapses 'your DR is broken' into 'your DR is unverified', and
the second is the one that quietly becomes the first. The staleness wording
says 'unverified, not known-broken' and a test asserts that phrasing.

Anchored per R-81, not re-derived: a never-proven tier on a newborn box is
UNKNOWN, not FAILED, until the window elapses. This family has made the
opposite mistake three times; this monitor was written straight after the third,
so it copies R-81's structure rather than inventing a fourth shape.

restoreProvenStaleAfter = 7d is derived: oldest-first over two tiers at a 24h
cadence proves each ~every 2 days, so 7d tolerates ~3 missed opportunities and
sits inside the 2-week offsite retention.

Per-tier proof comes from the hub's retained host-report window — the agent
reports only its latest run, so the latest report alone cannot answer 'when was
the OTHER tier last proven?'. Reused R-81's mechanism instead of a wire change.

Both types registered in allowedEventTypes (R-77's inert-seam lesson) and
operator-tier only — no customerMessages entry.

FIXED a time bomb I introduced in Slice C: the restart-blind-window test
hard-coded 2026-07-18T18:31:06Z while comparing against the real clock. Harmless
under one 26h threshold; once the offsite tier got an 8-day limit it passed all
day and began failing at 18:31 UTC, exactly 8 days later. Now relative.

Red-proofs B and D observed. Full suite green (17 packages, rc=0).
This commit is contained in:
Claude Code
2026-07-26 21:08:02 +02:00
parent b6d93e3fcb
commit ce4e03dcd8
6 changed files with 671 additions and 3 deletions
+58
View File
@@ -1,5 +1,63 @@
# Felhom Hub — Changelog
## v0.77.0 — R-85 Part 2: a restore-test result becomes a SIGNAL (2026-07-26)
Until now a failed restore-test was a `[WARN]` line in the ingest handler and nothing else — no
event, no notification, no gauge. That was true for the **local tier that was already being tested**,
so the loudest DR signal this system produces was, in practice, inaudible. Rotating tiers (agent
v0.104.0) without this would only mean two tiers can fail silently instead of one.
### Two signals, deliberately NOT merged
| event | meaning | severity |
|---|---|---|
| `restore_test_failed` | a run completed and did **not** pass — something is broken NOW | error |
| `restore_test_stale` | a tier has not been **proven** within its interval — nothing has necessarily broken; we no longer know | warning |
Merging them would collapse *"your DR is broken"* into *"your DR is unverified"*, and the second is
the one that quietly becomes the first. The staleness wording says **"unverified, not known-broken"**
and a test asserts that phrasing.
### Anchored, per R-81 — not re-derived
A tier never proven on a newborn box is **UNKNOWN**, not FAILED, until an anchored window elapses.
This monitor family has made the opposite mistake three times (hub v0.12.0, v0.73.0, R-81); this is a
NEW monitor written straight after the third, so it copies R-81's verdict structure and boundary-test
discipline rather than inventing a fourth shape. The deferral is logged once — a quiet check must
never be indistinguishable from one that did not run.
`restoreProvenStaleAfter = 7d` is derived, not guessed: a 24 h cadence rotating oldest-first across
two tiers proves each about every 2 days, so 7 days tolerates ~3 consecutive missed opportunities
before alarming — and sits comfortably inside the 2-week offsite retention, so a tier is never called
stale against an archive that is about to be pruned.
### How per-tier proof is recovered
The agent reports only its **latest** restore-test and its store is in-memory, so the latest report
alone cannot answer *"when was the OTHER tier last proven?"*. The hub's **retained host-report
window** can — the same R-81 mechanism, reused rather than re-solved with a wire change.
### Also
- Both types registered in `allowedEventTypes`. They are hub-generated, but that map is the
project's single register of legitimate event types, and R-77's lesson was that a type missing
from it ships as an inert seam.
- **Operator-tier only** — neither has a `customerMessages` entry, so the dispatcher cannot route it
to a customer. A customer can take no action on a failed restore-test, and *"a visszaállítási teszt
nem sikerült"* would frighten without informing. A persistently unproven DR tier may eventually
warrant a customer-visible statement; that needs copy review, not a side effect of this task.
- A tier the box does not HAVE is never reported stale (the Slice-C gate, same reasoning).
### Fixed — a time bomb introduced in Slice C
`TestCheckBackupDeadlines_RestartBlindWindow_NoEvent` hard-coded the literal incident timestamp
`2026-07-18T18:31:06Z` while comparing against the REAL clock. Harmless while one 26 h threshold
covered every tier; once Slice C gave the offsite tier an 8-day limit it became a bomb — the test
passed all day on 2026-07-26 and began failing at **18:31 UTC**, exactly 8 days after that instant.
Now relative. **A test that passes at commit time and fails hours later is worse than one that fails
immediately**, because it lands on whoever is next in the file.
### Tests
+10, full suite green (17 packages, `rc=0`, vet unpiped). Red-proofs observed:
- **B** — log-and-stop (the pre-R-85 shape) yields `a FAILED restore-test must EMIT an operator event; got 0 event(s)`. The assertion is that a NOTIFICATION IS EMITTED; the hollow version checks for a log line, which passes against exactly the code this replaces.
- **D** — removing the anchor yields `a newborn box must NOT alarm; got restore_test_stale: local tier: NEVER successfully restore-proven in 0s of watching`.
## v0.76.0 — R-82 Slice C: tier-aware backup thresholds (2026-07-26)
R-81 merged every backup signal into one "newest" and judged it against a single 26 h limit. That
+4
View File
@@ -550,6 +550,9 @@ func main() {
// (90/95% of quota_gb) + staleness (enabled+escrowed but no run >48h — the silently-stuck detector;
// run FAILURES already alert via backup_failed). Nil-safe on pre-v0.109 reports. Same sweep.
offsiteChecker := monitor.NewOffsiteChecker(dataStore, 0, dispatcher.ProcessEvent, logger)
// R-85: the restore-test result becomes a signal instead of a log line. Two distinct events
// (failure vs staleness), operator-tier only — neither has a customerMessages entry.
restoreTestChecker := monitor.NewRestoreTestChecker(dataStore, dispatcher.ProcessEvent, logger)
// R-70 + R-71c: the delivery-state checker — surfaces the burned-credential shape as
// offsite_delivery_stuck (warning, 24h/customer) and self-heals it via the Re-issue path
// (offsite_credential_restaged, one restage/customer/24h, R-39(a)-guarded). Cooldowns are
@@ -572,6 +575,7 @@ func main() {
hostMgmtPlaneChecker.Check()
hostOOBChecker.Check()
offsiteChecker.Check()
restoreTestChecker.Check() // R-85: restore-test failure + per-tier staleness
offsiteDeliveryChecker.Check()
if offsiteBoxChecker != nil {
offsiteBoxChecker.Check() // R-5: restic pool-box aggregate (fetch-throttled internally)
+7
View File
@@ -1548,6 +1548,13 @@ func (h *Handler) handleClaimResetRequest(w http.ResponseWriter, r *http.Request
// allowedEventTypes lists all valid event_type values the Hub accepts.
var allowedEventTypes = map[string]bool{
// R-85: restore-test signals. Hub-GENERATED (source "hub"), but listed here on purpose —
// allowedEventTypes is the project's single register of legitimate event types, and R-77's
// lesson was that an event type missing from it ships as an inert seam. Neither has a
// customerMessages entry, so both stay operator-tier.
"restore_test_failed": true,
"restore_test_stale": true,
// Controller-pushed events
"controller_started": true,
"claim_lockout": true, // v0.50.0 — claim/reset code brute-force lockout tripped
+15 -3
View File
@@ -374,15 +374,27 @@ func TestCheckBackupDeadlines_RestartBlindWindow_NoEvent(t *testing.T) {
}{{rfc(-20 * time.Hour), true}}
// Pre-restart report: carries the vzdump. Back-dated so it is not the latest.
pre := hostReportJSON(t, [][2]string{{"2026-07-18T18:31:06Z", "ok"}}, okVz)
//
// The offsite snapshot is RELATIVE and FRESH on purpose. It used to be the literal
// `2026-07-18T18:31:06Z` from the incident, which was harmless while ONE 26h threshold covered
// every tier — but Slice C gave the offsite tier its own 8-day limit, and this test compares
// against the REAL clock (runDeadline uses time.Now). So the fixture quietly became a TIME BOMB:
// it passed all day on 2026-07-26 and began failing at 18:31 UTC, exactly 8 days after the
// hard-coded instant. A test that passes at commit time and fails hours later is worse than one
// that fails immediately, because it lands on whoever is next in the file.
//
// This test is about the HOST tier's restart blind window; the offsite tier must be healthy so
// it cannot contribute to the verdict.
pre := hostReportJSON(t, [][2]string{{rfc(-2 * time.Hour), "ok"}}, okVz)
if err := st.SaveHostReport("h1", "c1", []byte(pre), store.HostReportDenorm{}); err != nil {
t.Fatal(err)
}
if err := st.SetHostReportsReceivedAtForTest("c1", sqliteAgo(15*time.Hour)); err != nil {
t.Fatal(err)
}
// Post-restart report: PBS only, 176h stale — exactly demo-felhom's 07-26 shape.
post := hostReportJSON(t, [][2]string{{"2026-07-18T18:31:06Z", "ok"}}, nil)
// Post-restart report: offsite only (the agent's in-memory `backups` was wiped by the restart)
// — demo-felhom's 07-26 shape, with the offsite snapshot kept fresh per the note above.
post := hostReportJSON(t, [][2]string{{rfc(-2 * time.Hour), "ok"}}, nil)
if err := st.SaveHostReport("h1", "c1", []byte(post), store.HostReportDenorm{}); err != nil {
t.Fatal(err)
}
+302
View File
@@ -0,0 +1,302 @@
package monitor
import (
"encoding/json"
"fmt"
"log"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
)
// R-85 Part 2 — the restore-test result becomes a SIGNAL.
//
// Until now a failed restore-test was a `[WARN]` log line in the ingest handler and nothing else:
// no event, no notification, no staleness gauge. That was true for the LOCAL tier that was already
// being tested — so the loudest DR signal this system produces was, in practice, inaudible.
//
// TWO SIGNALS, DELIBERATELY NOT MERGED. They mean different things and warrant different urgency:
//
// restore_test_failed — a run completed and did NOT pass. Something is broken NOW.
// restore_test_stale — a tier has not been PROVEN within its expected interval. Nothing has
// necessarily broken; we simply no longer know whether it works.
//
// Merging them would collapse "your DR is broken" into "your DR is unverified", and the second is
// the one that quietly becomes the first.
//
// ── THE INVARIANT, inherited from R-81 ────────────────────────────────────────────────────────
//
// Absence of a signal is UNKNOWN, and becomes a fault only once it has outlived an ANCHORED window.
// This monitor family has made the opposite mistake three times (hub v0.12.0, v0.73.0, R-81); this
// is a NEW monitor written straight after the third, so it copies R-81's verdict structure rather
// than re-deriving it. A tier never proven on a newborn box is UNKNOWN, never FAILED.
// restoreProvenStaleAfter is how long a tier may go unproven before it is called stale.
//
// Derivation, not a guess: the restore-test cadence is 24h and rotation is oldest-first across two
// tiers, so each tier is proven roughly every 2 days. 7 days therefore tolerates ~3 consecutive
// missed opportunities before alarming — loud enough to matter, quiet enough not to fire on one
// skipped cycle (a deferral behind a long backup is normal, not a fault). It is also comfortably
// inside the 2-week offsite retention (operator ruling 2026-07-26), so a tier is never reported
// stale against an archive that is about to be pruned anyway.
const restoreProvenStaleAfter = 7 * 24 * time.Hour
// Event types. Operator-tier only — see the dispatcher note in RestoreTestChecker.
const (
EventRestoreTestFailed = "restore_test_failed"
EventRestoreTestStale = "restore_test_stale"
)
// hostReportRestoreTests is the slice of a host-report this checker reads.
type hostReportRestoreTests struct {
RestoreTests []struct {
SourceArchive string `json:"source_archive"`
SourceTier string `json:"source_tier"`
Pass bool `json:"pass"`
Error string `json:"error"`
TestedAt string `json:"tested_at"`
} `json:"restore_tests"`
StorageTargets []struct {
Name string `json:"name"`
Type string `json:"type"`
Content string `json:"content"`
} `json:"storage_targets"`
}
// RestoreTestChecker watches restore-test outcomes per customer.
type RestoreTestChecker struct {
store *store.Store
logger *log.Logger
onEvent EventNotifyFunc
now func() time.Time
mu sync.Mutex
// failStates is edge-trigger state: customerID|archive → already reported. A failing tier is
// re-reported only when the FAILING RUN changes, so a permanently broken tier does not emit
// every 60s sweep — the flapping-spam rule every checker here follows.
failStates map[string]bool
// staleStates is customerID|tier → "ok"|"stale", so the stale signal is edge-triggered too.
staleStates map[string]string
}
// NewRestoreTestChecker builds the checker. onEvent may be nil.
func NewRestoreTestChecker(s *store.Store, onEvent EventNotifyFunc, logger *log.Logger) *RestoreTestChecker {
return &RestoreTestChecker{
store: s, logger: logger, onEvent: onEvent,
now: func() time.Time { return time.Now().UTC() },
failStates: map[string]bool{},
staleStates: map[string]string{},
}
}
// Check sweeps every active customer. Safe to call on the shared monitor ticker.
func (c *RestoreTestChecker) Check() {
ids, err := c.store.GetActiveCustomerIDs()
if err != nil {
c.logger.Printf("[WARN] restore-test check: failed to list customers: %v", err)
return
}
now := c.now()
for _, id := range ids {
latest, rerr := c.store.GetLatestHostReportJSON(id)
if rerr != nil || latest == "" {
continue // no agent host-report → nothing to judge (the R-81 legacy-shape branch)
}
c.checkFailure(id, latest)
c.checkStaleness(id, latest, now)
}
}
// checkFailure raises restore_test_failed when the latest report carries a FAILED run.
func (c *RestoreTestChecker) checkFailure(customerID, reportJSON string) {
var hr hostReportRestoreTests
if json.Unmarshal([]byte(reportJSON), &hr) != nil {
return
}
for _, rt := range hr.RestoreTests {
if rt.Pass {
// A pass CLEARS the edge state for that archive so a later failure re-reports.
c.mu.Lock()
delete(c.failStates, customerID+"|"+rt.SourceArchive)
c.mu.Unlock()
continue
}
key := customerID + "|" + rt.SourceArchive
c.mu.Lock()
already := c.failStates[key]
c.failStates[key] = true
c.mu.Unlock()
if already {
continue
}
tier := rt.SourceTier
if tier == "" {
tier = "unknown"
}
msg := fmt.Sprintf("Restore-test FAILED on the %s tier: archive %s could not be restored+booted (%s)",
tier, rt.SourceArchive, rt.Error)
c.emit(customerID, EventRestoreTestFailed, "error", msg)
}
}
// checkStaleness raises restore_test_stale for a tier not PROVEN within restoreProvenStaleAfter.
//
// "Proven" means a PASSING run for that tier, found across the hub's retained host-report window —
// the R-81 mechanism. The window is load-bearing here: the agent reports only its LATEST
// restore-test and its store is in-memory, so the latest report alone cannot answer "when was the
// OTHER tier last proven?". The hub's own history can.
func (c *RestoreTestChecker) checkStaleness(customerID, latestJSON string, now time.Time) {
var latest hostReportRestoreTests
if json.Unmarshal([]byte(latestJSON), &latest) != nil {
return
}
tiers := expectedRestoreTiers(latest)
if len(tiers) == 0 {
return
}
rows, err := c.store.GetHostReportsSince(customerID, now.Add(-2*restoreProvenStaleAfter))
if err != nil {
c.logger.Printf("[WARN] restore-test check: window read failed for %s: %v", customerID, err)
return
}
proven := lastProvenPerTier(rows)
first, ferr := c.store.GetFirstHostReportAt(customerID)
if ferr != nil {
c.logger.Printf("[WARN] restore-test check: first-contact read failed for %s: %v", customerID, ferr)
return
}
for _, tier := range tiers {
v := assessRestoreProven(tier, proven[tier], first, now)
key := customerID + "|" + tier
c.mu.Lock()
prev := c.staleStates[key]
state := "ok"
if v.verdict == verdictMissed {
state = "stale"
}
c.staleStates[key] = state
c.mu.Unlock()
switch {
case state == "stale" && prev != "stale":
c.emit(customerID, EventRestoreTestStale, "warning", v.reason)
case v.verdict == verdictUnknown && prev == "":
// Make the deferral VISIBLE exactly once, the v0.73.0 / R-81 precedent: a quiet check
// must never be indistinguishable from a check that did not run.
c.logger.Printf("[INFO] restore-test staleness: %s %s", customerID, v.reason)
}
}
}
// assessRestoreProven is the per-tier verdict. PURE (now injected) so the policy is unit-tested —
// the property that made R-81 provable, kept deliberately.
//
// no proof, anchor NOT elapsed → UNKNOWN (newborn box; never an alarm)
// no proof, anchor elapsed → MISSED
// proof older than the window → MISSED
// otherwise → OK
func assessRestoreProven(tier string, provenAt, firstReportAt, now time.Time) backupAssessment {
if provenAt.IsZero() {
if firstReportAt.IsZero() {
return backupAssessment{verdict: verdictMissed,
reason: fmt.Sprintf("%s tier: never restore-proven, and no first-contact anchor to defer against", tier)}
}
watched := now.Sub(firstReportAt)
if watched <= restoreProvenStaleAfter {
return backupAssessment{verdict: verdictUnknown,
reason: fmt.Sprintf("%s tier: not restore-proven yet, but only watching for %s (grace %s since first contact %s) — newborn, not a fault",
tier, watched.Round(time.Hour), restoreProvenStaleAfter, firstReportAt.Format(time.RFC3339))}
}
return backupAssessment{verdict: verdictMissed,
reason: fmt.Sprintf("%s tier: NEVER successfully restore-proven in %s of watching (limit %s) — the tier is unverified, not known-broken",
tier, watched.Round(time.Hour), restoreProvenStaleAfter)}
}
if age := now.Sub(provenAt); age > restoreProvenStaleAfter {
return backupAssessment{verdict: verdictMissed,
reason: fmt.Sprintf("%s tier: last successful restore-test was %s ago (limit %s) — the tier is unverified, not known-broken",
tier, age.Round(time.Hour), restoreProvenStaleAfter)}
}
return backupAssessment{verdict: verdictOK}
}
// expectedRestoreTiers names the tiers this box actually HAS, so a box without an offsite tier is
// never reported stale for one. Same gate as Slice C's `expected`, and for the same reason: without
// it every box lacking a tier would alarm once the anchor elapsed — absence-is-not-failure,
// re-introduced one level down.
func expectedRestoreTiers(hr hostReportRestoreTests) []string {
var out []string
seenHost, seenOffsite := false, false
for _, st := range hr.StorageTargets {
isPBS := st.Type == "pbs"
if isPBS && !seenOffsite {
out = append(out, "pbs")
seenOffsite = true
continue
}
if !isPBS && !seenHost && containsBackupContent(st.Content) {
out = append(out, "local")
seenHost = true
}
}
return out
}
func containsBackupContent(content string) bool {
for i := 0; i+6 <= len(content); i++ {
if content[i:i+6] == "backup" {
return true
}
}
return false
}
// lastProvenPerTier walks the retained window for the newest PASSING run per tier.
func lastProvenPerTier(rows []store.HostReportRow) map[string]time.Time {
out := map[string]time.Time{}
for _, r := range rows {
var hr hostReportRestoreTests
if json.Unmarshal([]byte(r.ReportJSON), &hr) != nil {
continue // one malformed retained report must not blind the scan
}
for _, rt := range hr.RestoreTests {
if !rt.Pass || rt.SourceTier == "" {
continue
}
t, err := time.Parse(time.RFC3339, rt.TestedAt)
if err != nil {
continue
}
if cur, ok := out[rt.SourceTier]; !ok || t.After(cur) {
out[rt.SourceTier] = t.UTC()
}
}
}
return out
}
// emit saves the event and notifies. OPERATOR-TIER ONLY: neither type has a `customerMessages`
// entry, so the dispatcher cannot route it to a customer. That is deliberate — a customer can take
// no action on a failed restore-test, and „a visszaállítási teszt nem sikerült" would frighten
// without informing. A PERSISTENTLY unproven DR tier may eventually warrant a customer-visible
// statement, but that needs copy review, not a side effect of this task.
func (c *RestoreTestChecker) emit(customerID, eventType, severity, msg string) {
if _, err := c.store.SaveEvent(customerID, eventType, severity, msg, "{}", "hub"); err != nil {
c.logger.Printf("[WARN] restore-test check: failed to save %s for %s: %v", eventType, customerID, err)
return
}
c.logger.Printf("[%s] restore-test: %s %s", severityLabel(severity), customerID, msg)
if c.onEvent != nil {
c.onEvent(customerID, eventType, severity, msg, "{}", "hub")
}
}
func severityLabel(sev string) string {
if sev == "error" {
return "ERROR"
}
return "WARN"
}
+285
View File
@@ -0,0 +1,285 @@
package monitor
import (
"encoding/json"
"io"
"log"
"path/filepath"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
)
// R-85 Part 2 — the restore-test result must be HEARD.
//
// Before this, a failed restore-test was a `[WARN]` line in the ingest handler and nothing else —
// no event, no notification, no gauge. That was true for the LOCAL tier that was already being
// tested, which means the loudest DR signal this system produces was in practice inaudible.
type capturedEvent struct {
customerID, eventType, severity, message string
}
func rtStore(t *testing.T) *store.Store {
t.Helper()
st, err := store.New(filepath.Join(t.TempDir(), "rt.db"), log.New(io.Discard, "", 0))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { st.Close() })
if err := st.SaveCustomerConfig(&store.CustomerConfig{CustomerID: "c1", APIKey: "k", RetrievalPassword: "p"}); err != nil {
t.Fatal(err)
}
if err := st.UpsertHost(&store.Host{HostID: "h1", CustomerID: "c1", APIKey: "k1"}); err != nil {
t.Fatal(err)
}
return st
}
// rtReport builds a host-report carrying storage targets and zero or one restore-test.
func rtReport(t *testing.T, tier string, pass bool, testedAt time.Time, errMsg string) string {
t.Helper()
type rt struct {
SourceArchive string `json:"source_archive"`
SourceTier string `json:"source_tier"`
Pass bool `json:"pass"`
Error string `json:"error,omitempty"`
TestedAt string `json:"tested_at"`
}
type stg struct {
Name string `json:"name"`
Type string `json:"type"`
Content string `json:"content"`
}
payload := struct {
RestoreTests []rt `json:"restore_tests"`
StorageTargets []stg `json:"storage_targets"`
}{
StorageTargets: []stg{
{Name: "local", Type: "local", Content: "backup,iso"},
{Name: "felhom-pbs", Type: "pbs", Content: "backup"},
},
}
if tier != "" {
payload.RestoreTests = []rt{{
SourceArchive: tier + ":backup/ct/9201/x", SourceTier: tier,
Pass: pass, Error: errMsg, TestedAt: testedAt.UTC().Format(time.RFC3339),
}}
}
b, err := json.Marshal(payload)
if err != nil {
t.Fatal(err)
}
return string(b)
}
func rtChecker(st *store.Store, got *[]capturedEvent, logTo io.Writer) *RestoreTestChecker {
if logTo == nil {
logTo = io.Discard
}
return NewRestoreTestChecker(st, func(cid, et, sev, msg, _, _ string) {
*got = append(*got, capturedEvent{cid, et, sev, msg})
}, log.New(logTo, "", 0))
}
// ── SCENARIO B — a failed restore-test is HEARD ──────────────────────────────────────────────
//
// The assertion is that a NOTIFICATION IS EMITTED, not that a handler ran. The hollow version of
// this test checks for a log line, which passes against exactly the code this task replaces.
//
// COMPANION RED-PROOF (observed): make checkFailure `return` before emit (the pre-R-85 shape — log
// and stop) and this fails with "a FAILED restore-test must EMIT an operator event; got 0 event(s)".
// Restored.
func TestRestoreTest_FailureEmitsAnOperatorEvent(t *testing.T) {
st := rtStore(t)
now := time.Now().UTC()
if err := st.SaveHostReport("h1", "c1",
[]byte(rtReport(t, "pbs", false, now, "restore task: context deadline exceeded")),
store.HostReportDenorm{}); err != nil {
t.Fatal(err)
}
var got []capturedEvent
rtChecker(st, &got, nil).Check()
var fail *capturedEvent
for i := range got {
if got[i].eventType == EventRestoreTestFailed {
fail = &got[i]
}
}
if fail == nil {
t.Fatalf("a FAILED restore-test must EMIT an operator event; got %d event(s): %+v", len(got), got)
}
if fail.severity != "error" {
t.Fatalf("a failed restore-test is an error, got severity %q", fail.severity)
}
if !strings.Contains(fail.message, "pbs") {
t.Fatalf("the TIER must be named — 'a restore-test failed' without saying which tier is not actionable; got %q", fail.message)
}
}
// Edge-triggered: a permanently failing tier must not emit on every sweep.
func TestRestoreTest_FailureDoesNotSpam(t *testing.T) {
st := rtStore(t)
now := time.Now().UTC()
st.SaveHostReport("h1", "c1", []byte(rtReport(t, "pbs", false, now, "boom")), store.HostReportDenorm{})
var got []capturedEvent
c := rtChecker(st, &got, nil)
c.Check()
c.Check()
c.Check()
n := 0
for _, e := range got {
if e.eventType == EventRestoreTestFailed {
n++
}
}
if n != 1 {
t.Fatalf("a persistently failing tier must report ONCE per failing run, got %d", n)
}
}
// A PASS must not emit a failure event.
func TestRestoreTest_PassEmitsNoFailure(t *testing.T) {
st := rtStore(t)
now := time.Now().UTC()
st.SaveHostReport("h1", "c1", []byte(rtReport(t, "pbs", true, now, "")), store.HostReportDenorm{})
var got []capturedEvent
rtChecker(st, &got, nil).Check()
for _, e := range got {
if e.eventType == EventRestoreTestFailed {
t.Fatalf("a PASSING restore-test must not emit a failure; got %+v", e)
}
}
}
// ── SCENARIO D — a newborn box does not alarm (the R-81 lesson) ─────────────────────────────
//
// COMPANION RED-PROOF (observed): remove the anchor branch from assessRestoreProven (treat "never
// proven" as MISSED outright) and this fails with "a newborn box must NOT alarm; got
// restore_test_stale ...". That is the fourth instance of no-signal-as-bad-signal, and it would
// have been written by the same hand that just fixed the third.
func TestRestoreTest_NewbornDoesNotAlarm(t *testing.T) {
st := rtStore(t)
now := time.Now().UTC()
// A report with NO restore-test at all, and first contact one hour ago.
st.SaveHostReport("h1", "c1", []byte(rtReport(t, "", false, now, "")), store.HostReportDenorm{})
var got []capturedEvent
var logbuf strings.Builder
rtChecker(st, &got, &logbuf).Check()
for _, e := range got {
if e.eventType == EventRestoreTestStale {
t.Fatalf("a newborn box must NOT alarm; got %s: %s", e.eventType, e.message)
}
}
// ...but the deferral must be VISIBLE, or quiet is indistinguishable from not-checked.
if !strings.Contains(logbuf.String(), "not restore-proven yet") {
t.Fatalf("the deferred verdict must be logged; log:\n%s", logbuf.String())
}
}
// The boundary, pinned by name so a refactor has to delete an obviously-named contract.
func TestRestoreTest_Contract_UnprovenIsUnknownUntilTheAnchorElapses(t *testing.T) {
now := time.Date(2026, 7, 26, 12, 0, 0, 0, time.UTC)
cases := []struct {
name string
watched time.Duration
wantMissed bool
}{
{"newborn, 1h", time.Hour, false},
{"just inside", restoreProvenStaleAfter - time.Minute, false},
{"exactly at the limit", restoreProvenStaleAfter, false},
{"just outside", restoreProvenStaleAfter + time.Minute, true},
{"long past", 30 * 24 * time.Hour, true},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
got := assessRestoreProven("pbs", time.Time{}, now.Add(-c.watched), now)
if got.missed() != c.wantMissed {
t.Fatalf("CONTRACT VIOLATED: unproven for %s (limit %s) → missed=%v, want %v (reason %q)",
c.watched, restoreProvenStaleAfter, got.missed(), c.wantMissed, got.reason)
}
if !c.wantMissed && got.verdict != verdictUnknown {
t.Fatalf("a deferred tier must be UNKNOWN (visible), not OK; got verdict=%d", got.verdict)
}
})
}
}
// ── SCENARIO C — an unproven tier becomes visible, and is DISTINCT from a failure ────────────
func TestRestoreTest_StaleIsSeparateFromFailure(t *testing.T) {
now := time.Date(2026, 7, 26, 12, 0, 0, 0, time.UTC)
stale := assessRestoreProven("pbs", now.Add(-9*24*time.Hour), now.Add(-60*24*time.Hour), now)
if !stale.missed() {
t.Fatalf("a tier last proven 9 days ago (limit %s) must be stale; got %q", restoreProvenStaleAfter, stale.reason)
}
// The wording must not read as "broken" — that is the other signal.
if !strings.Contains(stale.reason, "unverified, not known-broken") {
t.Fatalf("staleness must say UNVERIFIED, not broken — merging the two is the thing this avoids; got %q", stale.reason)
}
fresh := assessRestoreProven("pbs", now.Add(-2*24*time.Hour), now.Add(-60*24*time.Hour), now)
if fresh.verdict != verdictOK {
t.Fatalf("a tier proven 2 days ago is fine; got verdict=%d reason=%q", fresh.verdict, fresh.reason)
}
}
// The two signals must be DIFFERENT event types — merging them would collapse "your DR is broken"
// into "your DR is unverified", and the second is the one that quietly becomes the first.
func TestRestoreTest_EventTypesAreDistinct(t *testing.T) {
if EventRestoreTestFailed == EventRestoreTestStale {
t.Fatal("failure and staleness must be distinct event types")
}
}
// A tier the box does not HAVE is never reported stale — the Slice-C gate, for the same reason:
// without it, absence-is-not-failure would be re-introduced one level down.
func TestRestoreTest_TierNotPresentIsNeverStale(t *testing.T) {
var hr hostReportRestoreTests
if err := json.Unmarshal([]byte(`{"storage_targets":[{"name":"local","type":"local","content":"backup"}]}`), &hr); err != nil {
t.Fatal(err)
}
tiers := expectedRestoreTiers(hr)
for _, tr := range tiers {
if tr == "pbs" {
t.Fatalf("a box with no PBS storage must not expect a pbs tier; got %v", tiers)
}
}
if len(tiers) != 1 || tiers[0] != "local" {
t.Fatalf("want just the local tier; got %v", tiers)
}
}
// The window scan recovers per-tier proof even though each report carries only the LATEST run.
func TestRestoreTest_LastProvenPerTierAcrossTheWindow(t *testing.T) {
now := time.Now().UTC()
rows := []store.HostReportRow{
{ReceivedAt: now.Add(-1 * time.Hour), ReportJSON: rtReportRaw("pbs", true, now.Add(-2*time.Hour))},
{ReceivedAt: now.Add(-3 * time.Hour), ReportJSON: rtReportRaw("local", true, now.Add(-4*time.Hour))},
{ReceivedAt: now.Add(-5 * time.Hour), ReportJSON: `{{{malformed`},
}
got := lastProvenPerTier(rows)
if _, ok := got["pbs"]; !ok {
t.Fatalf("the pbs tier's proof must be recovered from the window; got %v", got)
}
if _, ok := got["local"]; !ok {
t.Fatalf("the local tier's proof must be recovered even though a LATER report shows only pbs; got %v", got)
}
}
func rtReportRaw(tier string, pass bool, at time.Time) string {
return `{"restore_tests":[{"source_archive":"` + tier + `:x","source_tier":"` + tier +
`","pass":` + boolStr(pass) + `,"tested_at":"` + at.UTC().Format(time.RFC3339) + `"}]}`
}
func boolStr(b bool) string {
if b {
return "true"
}
return "false"
}