package backup import ( "context" "fmt" "os" "path/filepath" "sort" "strings" "time" "gitea.dooplex.hu/admin/felhom-controller/internal/settings" ) // ── R-87 — the box proves its own off-site copy still HOLDS something ──────────────────────────── // // WHAT THIS ANSWERS THAT NOTHING ELSE DOES. `restic check --read-data-subset=100%` (R-359/R-399) runs // weekly and proves the stored bytes are the bytes we stored. **It cannot tell us we stored the wrong // thing.** A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks // cleanly at full depth, restores cleanly, and gives the customer nothing back. R-403 measured that // shape on this fleet on 2026-08-31: 120 082 104 B became 7 036 B in one nightly run, recorded as a // success. Until this job, no tier and no cadence asked the question. // // WHAT IT DELIBERATELY DOES NOT ANSWER, said here because the green tick will be read as covering it // otherwise: **it does not prove a restore puts data back into a running app.** It restores to a // throwaway directory, looks, and deletes. Putting data back is drill work, and the spike said so — // of the five restore-path defects human drills found in the six days to 2026-08-31, an unattended // scratch restore would have caught ONE (R-356). `07-backup-architecture.md` §8 matrix row 4 does not // move on the strength of this job. // // THE SHAPE IS `CheckOffboxIntegrity`'S, deliberately — same file conventions, same guards, same // three-outcome honesty — because a second scheduled off-site job that reasons differently about the // single-writer flag is how the two come to disagree about repository safety. // // ── THE HAZARD, AND WHY THIS JOB TAKES A FLAG THAT ITS RESTORE DOES NOT ───────────────────────── // // `resticStep` self-heals a crash lock by running `unlock --remove-all`, and its own doc comment // records that this is safe only because *"the in-process single-flight mutex (held by every caller // of this method) proves no sibling operation is live"*. // // `offbox_integrity.go`'s header states that invariant as universal — *"Every off-site operation // takes acquireRunning for exactly that reason"* — and **it is not true today**: // `RestoreOffboxScratch` does not take it (R-408, measured 2026-08-31; `restore_wizard.go` records // the same fact independently and the UI works around it with a separate display flag). This job // therefore takes the flag ITSELF rather than waiting for R-408 to be fixed, and SKIPS rather than // waits — a skipped proof simply runs tomorrow, whereas waiting would pin a nightly backup behind it. // // ── AND IT NEVER WRITES TO THE REPOSITORY, FOR REAL ──────────────────────────────────────────── // // R-95's constraint is that the off-site credential can delete, so a scheduled job must not be able // to write. The spike established what "read-only" actually costs here, by observation rather than by // reading the code: // // - `restic restore` in 0.14.0 takes NO repository lock (measured on demo-hp with a lock sampler // that was positively controlled first: it saw a lock appear and vanish across a real `check`, and // zero across two restores); // - `restic check` DOES take one — so the comment claiming the check never writes is wrong (R-407); // - the PRODUCT's restore path writes anyway, because `unlockStale` runs `restic unlock` — a delete // verb against `locks/` — before every restore (`offbox_restore.go`). // // So this job runs its restore with `--no-lock` and WITHOUT `unlockStale`, through `m.runner()` // rather than `resticStep`, so the `unlock --remove-all` escalation is not merely unlikely but // unreachable. `TestR87_NeverWritesToTheRepository` asserts that as a NON-EFFECT — the argv, not the // absence of an error. Using the seam rather than replacing `resticStep` is REUSE.md's own rule: // replacing it would hide the escalation from exactly the assertion that must see it. // // `restic restore --verify` is NOT used as a correctness check and must never be: the spike measured // it passing a byte-level corruption with size and mtime preserved, on a 213 MB tree, in 131 ms. It // is a size-and-mtime reconciliation. The correctness question is answered by `JudgeRestoredUnit`. // proofRestoreTimeout bounds one proof restore. // // CHOSEN from measurement, not inherited: on demo-hp 2026-08-31 a unit restore took 2.3–4.0 s per app // regardless of size (185 KB → 2.25 s, 213 MB → 3.20 s — the cost is per-snapshot round-trip plus // roughly 1 s per 200 MB), and all eight apps back to back took 25 s. Ten minutes is ~150x the // slowest single app measured, so the failure mode of a wedged SFTP mount is a released flag and a // retry tomorrow, never a box whose backups stop because a proof never returned. It is deliberately // well under `integrityCheckTimeout` (30 min): this job downloads ONE app's unit, not the store. const proofRestoreTimeout = 10 * time.Minute // ProofResult is what one nightly proof did, so the log, the event, the report and any future debug // endpoint state the same facts from one place. // // Skipped is a first-class outcome and NOT a failure — a proof that yielded to a running backup did // the right thing. NoSnapshot is separated from a failure for the same reason R-185 separated "I // could not look" from "I looked and it is broken": an app with no off-site snapshot yet is a // staleness question, which R-339 and the hub's tier deadline already own, and alarming here would be // a second alarm for a fact that already has one. type ProofResult struct { RanAt time.Time Duration time.Duration Stack string SnapshotID string Judgement UnitProofResult Skipped bool SkipReason string NoSnapshot bool // Err is a restore/plumbing failure — NOT a verdict about the backup. It never becomes the // "intact but empty" alarm, because a download that did not finish has seen nothing. Err error // CannotRun (R-414) is set when the proof could not START on this box at all — today the only // cause is that there is nowhere to put a scratch. // // IT IS SEPARATE FROM Err ON PURPOSE. An Err is transient: a download that failed says nothing // about the backup and nothing about the box, so it is right that it records nothing and retries. // This is a STANDING PROPERTY of the machine — it will be just as true tomorrow — and leaving it // unrecorded is what R-414 measured: `demo-felhom` failed this way every night with only a WARN, // and because nothing was written, `last_proof_result` stayed ABSENT — which is also what a // controller older than v0.231.0 sends. The hub could not tell "cannot run here" from "too old to // have the feature". That is the StatsKnown trap one level up, and this field is what closes it. CannotRun bool CannotWhy string } // Verdict renders the outcome for the persisted record and the wire. "" is reserved for "no verdict // was reached", which is what a skip, a missing snapshot and a restore error all are. func (r ProofResult) Verdict() string { if r.CannotRun { // A recorded, distinguishable state — never "" , which the hub reads as "not recorded". return ProofResultCannotRun } if r.Skipped || r.NoSnapshot || r.Err != nil { return "" } return string(r.Judgement.Verdict) } // ProofResultCannotRun is the persisted value for R-414's outcome. It is deliberately NOT one of the // UnitProofVerdict values: those judge a UNIT, and here no unit was ever looked at. const ProofResultCannotRun = "cannot_run" // ProveOffboxUnit runs ONE nightly proof: pick the due app, restore its newest snapshot read-only to // a throwaway directory, judge it, delete the copy, and record WHICH SNAPSHOT was proved. // // The order below is not arrangeable. The first three guards run before anything is downloaded, and // the scratch is removed on EVERY path including failure — a failed judgement that leaves a copy // behind is a slow disk-filler and, worse, a copy nobody knows the provenance of. func (m *Manager) ProveOffboxUnit(ctx context.Context) ProofResult { res := ProofResult{RanAt: time.Now()} if !m.OffboxConfigured() { // A box with no off-site tier has nothing to prove. Silent, and not an alarm. res.Skipped, res.SkipReason = true, "no off-site target configured" return res } if err := m.acquireRunning(); err != nil { res.Skipped, res.SkipReason = true, "a backup or restore is already running" m.logger.Printf("[INFO] [offbox] proof: skipped — %s; due-ness is NOT advanced, so this retries on the next run", res.SkipReason) return res } defer m.releaseRunning() stack, id, unitPath, ok := m.nextProofTarget(ctx) if !ok { res.NoSnapshot = true m.logger.Printf("[INFO] [offbox] proof: nothing due — every deployed app's newest off-site snapshot has already been proved, or none has one yet") return res } res.Stack, res.SnapshotID = stack, id scratch, nsRoot, derr := m.offboxProofScratchDir(stack) if derr != nil { // R-414: a standing property of this box, not a transient failure — so it is RECORDED, not // merely warned about. See ProofResult.CannotRun. res.CannotRun, res.CannotWhy = true, derr.Error() m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to, and this box will fail the same way every night until a data drive is registered: %v", stack, derr) return res } // The headroom gate runs BEFORE the download, through the same helper and the same Hungarian // wording the customer's restore uses (`unitOnlyHeadroom`). A gate after the download is not a // gate. if herr := unitOnlyHeadroom(m.offboxFree()(nsRoot)); herr != nil { res.Err = herr m.logger.Printf("[WARN] [offbox] proof: %s refused before any download: %v", stack, herr) return res } // Whatever happens from here, the copy goes. Registered before the restore so an early return // cannot leak one, and it removes a stale copy from an interrupted previous run on the way in. m.removeProofScratch(stack, scratch) defer m.removeProofScratch(stack, scratch) start := time.Now() if rerr := m.restoreUnitReadOnly(ctx, stack, id, unitPath, scratch); rerr != nil { res.Duration = time.Since(start) res.Err = rerr m.logger.Printf("[ERROR] [offbox] proof: %s could not be restored for proving (nothing is concluded about the backup): %v", stack, rerr) return res } res.Duration = time.Since(start) // The unit sits inside the scratch at its own absolute snapshot path. res.Judgement = JudgeRestoredUnit(filepath.Join(scratch, strings.TrimPrefix(unitPath, string(filepath.Separator)))) switch res.Judgement.Verdict { case UnitProofPass: m.logger.Printf("[INFO] [offbox] proof: %s PASSED on snapshot %s in %s — the backup holds what this app should have", stack, id, res.Duration.Round(time.Millisecond)) case UnitProofCannotJudge: m.logger.Printf("[WARN] [offbox] proof: %s on snapshot %s could NOT be judged (%s) — recorded as such, never as a pass", stack, id, res.Judgement.Reason) default: m.logger.Printf("[ERROR] [offbox] proof: %s on snapshot %s is READABLE AND EMPTY (%s%s) — the store is not damaged; the backup does not contain this app's data", stack, id, res.Judgement.Reason, missingSuffix(res.Judgement.Missing)) } return res } // missingSuffix renders the Missing list for a log line without ever printing a path outside the unit. func missingSuffix(missing []string) string { if len(missing) == 0 { return "" } return ": " + strings.Join(missing, ", ") } // RecordProofVerdict persists the outcome and advances due-ness. The caller owns WHEN a verdict // counts, exactly as `RecordIntegrityVerdict` does. // // A SKIP, A MISSING SNAPSHOT AND A RESTORE ERROR DO NOT REACH A VERDICT and must not advance // due-ness — tomorrow tries again. A `cannot_judge` DOES advance it: re-downloading the same // unjudgeable snapshot every night is load with no new information, the outcome is recorded where a // surface can read it, and a new snapshot makes the app due again by itself. func (m *Manager) RecordProofVerdict(res ProofResult) { v := res.Verdict() if v == "" { return } if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) { if v == ProofResultCannotRun { // RECORDED so the hub can see the state, but due-ness is deliberately NOT advanced: no // snapshot was proved, and marking one proved would stop the app ever being retried once a // drive is finally registered. A status report, not a proof. o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339) o.LastProofStack = res.Stack o.LastProofSnapshot = "" o.LastProofResult = v o.LastProofReason = res.CannotWhy return } if o.ProvedSnapshots == nil { o.ProvedSnapshots = map[string]string{} } o.ProvedSnapshots[res.Stack] = res.SnapshotID o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339) o.LastProofStack = res.Stack o.LastProofSnapshot = res.SnapshotID o.LastProofResult = v o.LastProofReason = string(res.Judgement.Reason) }); err != nil { m.logger.Printf("[ERROR] [offbox] proof: could not persist the outcome: %v — the proof RAN and its verdict for %s on %s was %q, but due-ness did not advance, so it will run again tomorrow", err, res.Stack, res.SnapshotID, v) } } // nextProofTarget picks ONE app: the deployed app whose newest off-site snapshot has not been proved. // // DUE-NESS IS PER SNAPSHOT, NOT PER CLOCK (R-86's model, 07 §3). An app is due when the ID of its // newest snapshot differs from the ID last proved for it. A timestamp would re-prove the same // snapshot forever and say nothing about the newest one — which is the defect R-341 recorded in a // different costume. // // Apps are considered in a stable sorted order and the FIRST due one wins, so eight apps are covered // in eight nights and the rotation cannot starve one: an app stops being due only once its newest // snapshot has actually been proved. func (m *Manager) nextProofTarget(ctx context.Context) (stack, id, unitPath string, ok bool) { if m.stackProvider == nil { return "", "", "", false } proved := map[string]string{} if t := m.settings.GetOffboxTarget(); t != nil && t.ProvedSnapshots != nil { proved = t.ProvedSnapshots } var names []string for _, st := range m.stackProvider.ListDeployedStacks() { names = append(names, st.Name) } sort.Strings(names) for _, name := range names { if !isSafeStackName(name) { continue } sid, paths, err := m.offboxLatestSnapshot(ctx, name) if err != nil || strings.TrimSpace(sid) == "" { // No snapshot yet for this app is NOT a failure of this test — staleness belongs to // R-339 and the hub's tier deadline. Logged, never alarmed. continue } if proved[name] == sid { continue // already proved on exactly this snapshot } up := offboxUnitPathOf(paths, name) if up == "" { // The snapshot exists and carries no recovery unit. That IS the strongest form of // "nothing recoverable", so it is a target rather than a skip — the judgement below will // find no manifest and fail closed. m.logger.Printf("[WARN] [offbox] proof: %s snapshot %s lists no recovery-unit path", name, sid) continue } return name, sid, up, true } return "", "", "", false } // restoreUnitReadOnly downloads ONE recovery unit into scratch WITHOUT writing anything to the // repository. // // Three differences from `RestoreOffboxScratch`, each deliberate and each named: // - `--no-lock`, so restic does not create a lock file (it does not for `restore` in 0.14.0 anyway // — measured — but the flag makes that a guarantee rather than an observed behaviour of one // version); // - no `unlockStale`, so no delete verb is issued against `locks/`; // - `m.runner()` rather than `resticStep`, so the `unlock --remove-all` escalation is unreachable // rather than merely unlikely. // // It deliberately does NOT write the R-358 completion marker: that marker certifies a scratch for // PLACEMENT into a live app, and a proof copy must never be placeable. Its absence is what keeps this // directory inert even if the two roots were ever confused. func (m *Manager) restoreUnitReadOnly(ctx context.Context, stack, id, unitPath, scratch string) error { if err := os.MkdirAll(scratch, 0o755); err != nil { return fmt.Errorf("proof scratch dir: %w", err) } t := m.settings.GetOffboxTarget() base, env := m.offboxBaseArgs(t) rctx, cancel := context.WithTimeout(ctx, proofRestoreTimeout) defer cancel() args := append(append([]string{}, base...), "restore", id, "--target", scratch, "--include", unitPath, "--no-lock") out, err := m.runner()(rctx, env, args...) if err != nil { return fmt.Errorf("offbox proof restore %s: %w: %s", stack, err, truncate(out)) } return nil } // removeProofScratch deletes a proof copy, refusing loudly if the path is not strictly inside a // proof root — the same prefix instinct as `DeleteOffsiteRestoreCopy`, because reaching here with an // out-of-sandbox path means a helper above is wrong and a best-effort skip would hide that. func (m *Manager) removeProofScratch(stack, scratch string) { clean := filepath.Clean(scratch) // R-414: the accepted roots must include the SYSTEM DATA PATH, because that is now where a // unit-only scratch lands on a box with no registered drive. // // CAUGHT BY LIVE VALIDATION ON demo-felhom, 2026-09-01, and it was a leak I introduced with the // fallback itself: `offsiteRestoreDriveRoots` is built from registered drives and schedulable // paths, so on a driveless box it is EMPTY — the scratch resolved fine and then this refused to // delete it ("refusing to remove … it is not inside a proof root"). Every nightly proof would have // left a copy behind, growing forever, on exactly the boxes the fallback exists for. The unit // tests did not see it because they register a drive. roots := append(m.offsiteRestoreDriveRoots(), strings.TrimSpace(m.cfg.Paths.SystemDataPath)) for _, drive := range roots { if strings.TrimSpace(drive) == "" { continue } root := filepath.Clean(m.offsiteProofRootFor(drive)) + string(filepath.Separator) if strings.HasPrefix(clean+string(filepath.Separator), root) { if err := os.RemoveAll(clean); err != nil { m.logger.Printf("[WARN] [offbox] proof: could not remove the proof copy %s: %v", clean, err) } return } } m.logger.Printf("[WARN] [offbox] proof: refusing to remove %s — it is not inside a proof root", clean) }