Files
felhom-controller/controller/internal/backup/offbox_proof.go
T
admin fcef8e069c
gates / gates (push) Successful in 12s
one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:

  RestoreOffboxScratch      - the reported one
  OffboxRestorePrepareFull  - the SECOND request in the customer's own two-step full-restore
                              flow, and the one that actually shells `restic stats`. The UI
                              reaches it FIRST, so flagging only the restore would have left
                              the collision reachable by the ordinary path.
  RestoreSharesScratch      - R-411's exact shape on the shares tier: unlockStale + resticStep,
                              a live web caller, and its sibling PlaceSharesRestore has always
                              taken the flag.
  RestoreOffbox             - no production caller today, but the same dangerous pattern.
                              Flagged rather than left for a future caller to inherit.

OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.

THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.

R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.

R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.

The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.

And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.

R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.

16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
2026-09-01 10:19:36 +02:00

352 lines
18 KiB
Go
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
package backup
import (
"context"
"fmt"
"os"
"path/filepath"
"sort"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// ── R-87 — the box proves its own off-site copy still HOLDS something ────────────────────────────
//
// WHAT THIS ANSWERS THAT NOTHING ELSE DOES. `restic check --read-data-subset=100%` (R-359/R-399) runs
// weekly and proves the stored bytes are the bytes we stored. **It cannot tell us we stored the wrong
// thing.** A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks
// cleanly at full depth, restores cleanly, and gives the customer nothing back. R-403 measured that
// shape on this fleet on 2026-08-31: 120 082 104 B became 7 036 B in one nightly run, recorded as a
// success. Until this job, no tier and no cadence asked the question.
//
// WHAT IT DELIBERATELY DOES NOT ANSWER, said here because the green tick will be read as covering it
// otherwise: **it does not prove a restore puts data back into a running app.** It restores to a
// throwaway directory, looks, and deletes. Putting data back is drill work, and the spike said so —
// of the five restore-path defects human drills found in the six days to 2026-08-31, an unattended
// scratch restore would have caught ONE (R-356). `07-backup-architecture.md` §8 matrix row 4 does not
// move on the strength of this job.
//
// THE SHAPE IS `CheckOffboxIntegrity`'S, deliberately — same file conventions, same guards, same
// three-outcome honesty — because a second scheduled off-site job that reasons differently about the
// single-writer flag is how the two come to disagree about repository safety.
//
// ── THE HAZARD, AND WHY THIS JOB TAKES A FLAG THAT ITS RESTORE DOES NOT ─────────────────────────
//
// `resticStep` self-heals a crash lock by running `unlock --remove-all`, and its own doc comment
// records that this is safe only because *"the in-process single-flight mutex (held by every caller
// of this method) proves no sibling operation is live"*.
//
// `offbox_integrity.go`'s header states that invariant as universal — *"Every off-site operation
// takes acquireRunning for exactly that reason"* — and **it is not true today**:
// `RestoreOffboxScratch` does not take it (R-408, measured 2026-08-31; `restore_wizard.go` records
// the same fact independently and the UI works around it with a separate display flag). This job
// therefore takes the flag ITSELF rather than waiting for R-408 to be fixed, and SKIPS rather than
// waits — a skipped proof simply runs tomorrow, whereas waiting would pin a nightly backup behind it.
//
// ── AND IT NEVER WRITES TO THE REPOSITORY, FOR REAL ────────────────────────────────────────────
//
// R-95's constraint is that the off-site credential can delete, so a scheduled job must not be able
// to write. The spike established what "read-only" actually costs here, by observation rather than by
// reading the code:
//
// - `restic restore` in 0.14.0 takes NO repository lock (measured on demo-hp with a lock sampler
// that was positively controlled first: it saw a lock appear and vanish across a real `check`, and
// zero across two restores);
// - `restic check` DOES take one — so the comment claiming the check never writes is wrong (R-407);
// - the PRODUCT's restore path writes anyway, because `unlockStale` runs `restic unlock` — a delete
// verb against `locks/` — before every restore (`offbox_restore.go`).
//
// So this job runs its restore with `--no-lock` and WITHOUT `unlockStale`, through `m.runner()`
// rather than `resticStep`, so the `unlock --remove-all` escalation is not merely unlikely but
// unreachable. `TestR87_NeverWritesToTheRepository` asserts that as a NON-EFFECT — the argv, not the
// absence of an error. Using the seam rather than replacing `resticStep` is REUSE.md's own rule:
// replacing it would hide the escalation from exactly the assertion that must see it.
//
// `restic restore --verify` is NOT used as a correctness check and must never be: the spike measured
// it passing a byte-level corruption with size and mtime preserved, on a 213 MB tree, in 131 ms. It
// is a size-and-mtime reconciliation. The correctness question is answered by `JudgeRestoredUnit`.
// proofRestoreTimeout bounds one proof restore.
//
// CHOSEN from measurement, not inherited: on demo-hp 2026-08-31 a unit restore took 2.3–4.0 s per app
// regardless of size (185 KB → 2.25 s, 213 MB → 3.20 s — the cost is per-snapshot round-trip plus
// roughly 1 s per 200 MB), and all eight apps back to back took 25 s. Ten minutes is ~150x the
// slowest single app measured, so the failure mode of a wedged SFTP mount is a released flag and a
// retry tomorrow, never a box whose backups stop because a proof never returned. It is deliberately
// well under `integrityCheckTimeout` (30 min): this job downloads ONE app's unit, not the store.
const proofRestoreTimeout = 10 * time.Minute
// ProofResult is what one nightly proof did, so the log, the event, the report and any future debug
// endpoint state the same facts from one place.
//
// Skipped is a first-class outcome and NOT a failure — a proof that yielded to a running backup did
// the right thing. NoSnapshot is separated from a failure for the same reason R-185 separated "I
// could not look" from "I looked and it is broken": an app with no off-site snapshot yet is a
// staleness question, which R-339 and the hub's tier deadline already own, and alarming here would be
// a second alarm for a fact that already has one.
type ProofResult struct {
RanAt time.Time
Duration time.Duration
Stack string
SnapshotID string
Judgement UnitProofResult
Skipped bool
SkipReason string
NoSnapshot bool
// Err is a restore/plumbing failure — NOT a verdict about the backup. It never becomes the
// "intact but empty" alarm, because a download that did not finish has seen nothing.
Err error
// CannotRun (R-414) is set when the proof could not START on this box at all — today the only
// cause is that there is nowhere to put a scratch.
//
// IT IS SEPARATE FROM Err ON PURPOSE. An Err is transient: a download that failed says nothing
// about the backup and nothing about the box, so it is right that it records nothing and retries.
// This is a STANDING PROPERTY of the machine — it will be just as true tomorrow — and leaving it
// unrecorded is what R-414 measured: `demo-felhom` failed this way every night with only a WARN,
// and because nothing was written, `last_proof_result` stayed ABSENT — which is also what a
// controller older than v0.231.0 sends. The hub could not tell "cannot run here" from "too old to
// have the feature". That is the StatsKnown trap one level up, and this field is what closes it.
CannotRun bool
CannotWhy string
}
// Verdict renders the outcome for the persisted record and the wire. "" is reserved for "no verdict
// was reached", which is what a skip, a missing snapshot and a restore error all are.
func (r ProofResult) Verdict() string {
if r.CannotRun {
// A recorded, distinguishable state — never "" , which the hub reads as "not recorded".
return ProofResultCannotRun
}
if r.Skipped || r.NoSnapshot || r.Err != nil {
return ""
}
return string(r.Judgement.Verdict)
}
// ProofResultCannotRun is the persisted value for R-414's outcome. It is deliberately NOT one of the
// UnitProofVerdict values: those judge a UNIT, and here no unit was ever looked at.
const ProofResultCannotRun = "cannot_run"
// ProveOffboxUnit runs ONE nightly proof: pick the due app, restore its newest snapshot read-only to
// a throwaway directory, judge it, delete the copy, and record WHICH SNAPSHOT was proved.
//
// The order below is not arrangeable. The first three guards run before anything is downloaded, and
// the scratch is removed on EVERY path including failure — a failed judgement that leaves a copy
// behind is a slow disk-filler and, worse, a copy nobody knows the provenance of.
func (m *Manager) ProveOffboxUnit(ctx context.Context) ProofResult {
res := ProofResult{RanAt: time.Now()}
if !m.OffboxConfigured() {
// A box with no off-site tier has nothing to prove. Silent, and not an alarm.
res.Skipped, res.SkipReason = true, "no off-site target configured"
return res
}
if err := m.acquireRunning(); err != nil {
res.Skipped, res.SkipReason = true, "a backup or restore is already running"
m.logger.Printf("[INFO] [offbox] proof: skipped — %s; due-ness is NOT advanced, so this retries on the next run", res.SkipReason)
return res
}
defer m.releaseRunning()
stack, id, unitPath, ok := m.nextProofTarget(ctx)
if !ok {
res.NoSnapshot = true
m.logger.Printf("[INFO] [offbox] proof: nothing due — every deployed app's newest off-site snapshot has already been proved, or none has one yet")
return res
}
res.Stack, res.SnapshotID = stack, id
scratch, nsRoot, derr := m.offboxProofScratchDir(stack)
if derr != nil {
// R-414: a standing property of this box, not a transient failure — so it is RECORDED, not
// merely warned about. See ProofResult.CannotRun.
res.CannotRun, res.CannotWhy = true, derr.Error()
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to, and this box will fail the same way every night until a data drive is registered: %v", stack, derr)
return res
}
// The headroom gate runs BEFORE the download, through the same helper and the same Hungarian
// wording the customer's restore uses (`unitOnlyHeadroom`). A gate after the download is not a
// gate.
if herr := unitOnlyHeadroom(m.offboxFree()(nsRoot)); herr != nil {
res.Err = herr
m.logger.Printf("[WARN] [offbox] proof: %s refused before any download: %v", stack, herr)
return res
}
// Whatever happens from here, the copy goes. Registered before the restore so an early return
// cannot leak one, and it removes a stale copy from an interrupted previous run on the way in.
m.removeProofScratch(stack, scratch)
defer m.removeProofScratch(stack, scratch)
start := time.Now()
if rerr := m.restoreUnitReadOnly(ctx, stack, id, unitPath, scratch); rerr != nil {
res.Duration = time.Since(start)
res.Err = rerr
m.logger.Printf("[ERROR] [offbox] proof: %s could not be restored for proving (nothing is concluded about the backup): %v", stack, rerr)
return res
}
res.Duration = time.Since(start)
// The unit sits inside the scratch at its own absolute snapshot path.
res.Judgement = JudgeRestoredUnit(filepath.Join(scratch, strings.TrimPrefix(unitPath, string(filepath.Separator))))
switch res.Judgement.Verdict {
case UnitProofPass:
m.logger.Printf("[INFO] [offbox] proof: %s PASSED on snapshot %s in %s — the backup holds what this app should have",
stack, id, res.Duration.Round(time.Millisecond))
case UnitProofCannotJudge:
m.logger.Printf("[WARN] [offbox] proof: %s on snapshot %s could NOT be judged (%s) — recorded as such, never as a pass",
stack, id, res.Judgement.Reason)
default:
m.logger.Printf("[ERROR] [offbox] proof: %s on snapshot %s is READABLE AND EMPTY (%s%s) — the store is not damaged; the backup does not contain this app's data",
stack, id, res.Judgement.Reason, missingSuffix(res.Judgement.Missing))
}
return res
}
// missingSuffix renders the Missing list for a log line without ever printing a path outside the unit.
func missingSuffix(missing []string) string {
if len(missing) == 0 {
return ""
}
return ": " + strings.Join(missing, ", ")
}
// RecordProofVerdict persists the outcome and advances due-ness. The caller owns WHEN a verdict
// counts, exactly as `RecordIntegrityVerdict` does.
//
// A SKIP, A MISSING SNAPSHOT AND A RESTORE ERROR DO NOT REACH A VERDICT and must not advance
// due-ness — tomorrow tries again. A `cannot_judge` DOES advance it: re-downloading the same
// unjudgeable snapshot every night is load with no new information, the outcome is recorded where a
// surface can read it, and a new snapshot makes the app due again by itself.
func (m *Manager) RecordProofVerdict(res ProofResult) {
v := res.Verdict()
if v == "" {
return
}
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) {
if v == ProofResultCannotRun {
// RECORDED so the hub can see the state, but due-ness is deliberately NOT advanced: no
// snapshot was proved, and marking one proved would stop the app ever being retried once a
// drive is finally registered. A status report, not a proof.
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
o.LastProofStack = res.Stack
o.LastProofSnapshot = ""
o.LastProofResult = v
o.LastProofReason = res.CannotWhy
return
}
if o.ProvedSnapshots == nil {
o.ProvedSnapshots = map[string]string{}
}
o.ProvedSnapshots[res.Stack] = res.SnapshotID
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
o.LastProofStack = res.Stack
o.LastProofSnapshot = res.SnapshotID
o.LastProofResult = v
o.LastProofReason = string(res.Judgement.Reason)
}); err != nil {
m.logger.Printf("[ERROR] [offbox] proof: could not persist the outcome: %v — the proof RAN and its verdict for %s on %s was %q, but due-ness did not advance, so it will run again tomorrow",
err, res.Stack, res.SnapshotID, v)
}
}
// nextProofTarget picks ONE app: the deployed app whose newest off-site snapshot has not been proved.
//
// DUE-NESS IS PER SNAPSHOT, NOT PER CLOCK (R-86's model, 07 §3). An app is due when the ID of its
// newest snapshot differs from the ID last proved for it. A timestamp would re-prove the same
// snapshot forever and say nothing about the newest one — which is the defect R-341 recorded in a
// different costume.
//
// Apps are considered in a stable sorted order and the FIRST due one wins, so eight apps are covered
// in eight nights and the rotation cannot starve one: an app stops being due only once its newest
// snapshot has actually been proved.
func (m *Manager) nextProofTarget(ctx context.Context) (stack, id, unitPath string, ok bool) {
if m.stackProvider == nil {
return "", "", "", false
}
proved := map[string]string{}
if t := m.settings.GetOffboxTarget(); t != nil && t.ProvedSnapshots != nil {
proved = t.ProvedSnapshots
}
var names []string
for _, st := range m.stackProvider.ListDeployedStacks() {
names = append(names, st.Name)
}
sort.Strings(names)
for _, name := range names {
if !isSafeStackName(name) {
continue
}
sid, paths, err := m.offboxLatestSnapshot(ctx, name)
if err != nil || strings.TrimSpace(sid) == "" {
// No snapshot yet for this app is NOT a failure of this test — staleness belongs to
// R-339 and the hub's tier deadline. Logged, never alarmed.
continue
}
if proved[name] == sid {
continue // already proved on exactly this snapshot
}
up := offboxUnitPathOf(paths, name)
if up == "" {
// The snapshot exists and carries no recovery unit. That IS the strongest form of
// "nothing recoverable", so it is a target rather than a skip — the judgement below will
// find no manifest and fail closed.
m.logger.Printf("[WARN] [offbox] proof: %s snapshot %s lists no recovery-unit path", name, sid)
continue
}
return name, sid, up, true
}
return "", "", "", false
}
// restoreUnitReadOnly downloads ONE recovery unit into scratch WITHOUT writing anything to the
// repository.
//
// Three differences from `RestoreOffboxScratch`, each deliberate and each named:
// - `--no-lock`, so restic does not create a lock file (it does not for `restore` in 0.14.0 anyway
// — measured — but the flag makes that a guarantee rather than an observed behaviour of one
// version);
// - no `unlockStale`, so no delete verb is issued against `locks/`;
// - `m.runner()` rather than `resticStep`, so the `unlock --remove-all` escalation is unreachable
// rather than merely unlikely.
//
// It deliberately does NOT write the R-358 completion marker: that marker certifies a scratch for
// PLACEMENT into a live app, and a proof copy must never be placeable. Its absence is what keeps this
// directory inert even if the two roots were ever confused.
func (m *Manager) restoreUnitReadOnly(ctx context.Context, stack, id, unitPath, scratch string) error {
if err := os.MkdirAll(scratch, 0o755); err != nil {
return fmt.Errorf("proof scratch dir: %w", err)
}
t := m.settings.GetOffboxTarget()
base, env := m.offboxBaseArgs(t)
rctx, cancel := context.WithTimeout(ctx, proofRestoreTimeout)
defer cancel()
args := append(append([]string{}, base...),
"restore", id, "--target", scratch, "--include", unitPath, "--no-lock")
out, err := m.runner()(rctx, env, args...)
if err != nil {
return fmt.Errorf("offbox proof restore %s: %w: %s", stack, err, truncate(out))
}
return nil
}
// removeProofScratch deletes a proof copy, refusing loudly if the path is not strictly inside a
// proof root — the same prefix instinct as `DeleteOffsiteRestoreCopy`, because reaching here with an
// out-of-sandbox path means a helper above is wrong and a best-effort skip would hide that.
func (m *Manager) removeProofScratch(stack, scratch string) {
clean := filepath.Clean(scratch)
for _, drive := range m.offsiteRestoreDriveRoots() {
root := filepath.Clean(m.offsiteProofRootFor(drive)) + string(filepath.Separator)
if strings.HasPrefix(clean+string(filepath.Separator), root) {
if err := os.RemoveAll(clean); err != nil {
m.logger.Printf("[WARN] [offbox] proof: could not remove the proof copy %s: %v", clean, err)
}
return
}
}
m.logger.Printf("[WARN] [offbox] proof: refusing to remove %s — it is not inside a proof root", clean)
}