fcef8e069c
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:
RestoreOffboxScratch - the reported one
OffboxRestorePrepareFull - the SECOND request in the customer's own two-step full-restore
flow, and the one that actually shells `restic stats`. The UI
reaches it FIRST, so flagging only the restore would have left
the collision reachable by the ordinary path.
RestoreSharesScratch - R-411's exact shape on the shares tier: unlockStale + resticStep,
a live web caller, and its sibling PlaceSharesRestore has always
taken the flag.
RestoreOffbox - no production caller today, but the same dangerous pattern.
Flagged rather than left for a future caller to inherit.
OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.
THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.
R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.
R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.
The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.
And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.
R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.
16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
352 lines
18 KiB
Go
352 lines
18 KiB
Go
package backup
|
||
|
||
import (
|
||
"context"
|
||
"fmt"
|
||
"os"
|
||
"path/filepath"
|
||
"sort"
|
||
"strings"
|
||
"time"
|
||
|
||
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||
)
|
||
|
||
// ── R-87 — the box proves its own off-site copy still HOLDS something ────────────────────────────
|
||
//
|
||
// WHAT THIS ANSWERS THAT NOTHING ELSE DOES. `restic check --read-data-subset=100%` (R-359/R-399) runs
|
||
// weekly and proves the stored bytes are the bytes we stored. **It cannot tell us we stored the wrong
|
||
// thing.** A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks
|
||
// cleanly at full depth, restores cleanly, and gives the customer nothing back. R-403 measured that
|
||
// shape on this fleet on 2026-08-31: 120 082 104 B became 7 036 B in one nightly run, recorded as a
|
||
// success. Until this job, no tier and no cadence asked the question.
|
||
//
|
||
// WHAT IT DELIBERATELY DOES NOT ANSWER, said here because the green tick will be read as covering it
|
||
// otherwise: **it does not prove a restore puts data back into a running app.** It restores to a
|
||
// throwaway directory, looks, and deletes. Putting data back is drill work, and the spike said so —
|
||
// of the five restore-path defects human drills found in the six days to 2026-08-31, an unattended
|
||
// scratch restore would have caught ONE (R-356). `07-backup-architecture.md` §8 matrix row 4 does not
|
||
// move on the strength of this job.
|
||
//
|
||
// THE SHAPE IS `CheckOffboxIntegrity`'S, deliberately — same file conventions, same guards, same
|
||
// three-outcome honesty — because a second scheduled off-site job that reasons differently about the
|
||
// single-writer flag is how the two come to disagree about repository safety.
|
||
//
|
||
// ── THE HAZARD, AND WHY THIS JOB TAKES A FLAG THAT ITS RESTORE DOES NOT ─────────────────────────
|
||
//
|
||
// `resticStep` self-heals a crash lock by running `unlock --remove-all`, and its own doc comment
|
||
// records that this is safe only because *"the in-process single-flight mutex (held by every caller
|
||
// of this method) proves no sibling operation is live"*.
|
||
//
|
||
// `offbox_integrity.go`'s header states that invariant as universal — *"Every off-site operation
|
||
// takes acquireRunning for exactly that reason"* — and **it is not true today**:
|
||
// `RestoreOffboxScratch` does not take it (R-408, measured 2026-08-31; `restore_wizard.go` records
|
||
// the same fact independently and the UI works around it with a separate display flag). This job
|
||
// therefore takes the flag ITSELF rather than waiting for R-408 to be fixed, and SKIPS rather than
|
||
// waits — a skipped proof simply runs tomorrow, whereas waiting would pin a nightly backup behind it.
|
||
//
|
||
// ── AND IT NEVER WRITES TO THE REPOSITORY, FOR REAL ────────────────────────────────────────────
|
||
//
|
||
// R-95's constraint is that the off-site credential can delete, so a scheduled job must not be able
|
||
// to write. The spike established what "read-only" actually costs here, by observation rather than by
|
||
// reading the code:
|
||
//
|
||
// - `restic restore` in 0.14.0 takes NO repository lock (measured on demo-hp with a lock sampler
|
||
// that was positively controlled first: it saw a lock appear and vanish across a real `check`, and
|
||
// zero across two restores);
|
||
// - `restic check` DOES take one — so the comment claiming the check never writes is wrong (R-407);
|
||
// - the PRODUCT's restore path writes anyway, because `unlockStale` runs `restic unlock` — a delete
|
||
// verb against `locks/` — before every restore (`offbox_restore.go`).
|
||
//
|
||
// So this job runs its restore with `--no-lock` and WITHOUT `unlockStale`, through `m.runner()`
|
||
// rather than `resticStep`, so the `unlock --remove-all` escalation is not merely unlikely but
|
||
// unreachable. `TestR87_NeverWritesToTheRepository` asserts that as a NON-EFFECT — the argv, not the
|
||
// absence of an error. Using the seam rather than replacing `resticStep` is REUSE.md's own rule:
|
||
// replacing it would hide the escalation from exactly the assertion that must see it.
|
||
//
|
||
// `restic restore --verify` is NOT used as a correctness check and must never be: the spike measured
|
||
// it passing a byte-level corruption with size and mtime preserved, on a 213 MB tree, in 131 ms. It
|
||
// is a size-and-mtime reconciliation. The correctness question is answered by `JudgeRestoredUnit`.
|
||
|
||
// proofRestoreTimeout bounds one proof restore.
|
||
//
|
||
// CHOSEN from measurement, not inherited: on demo-hp 2026-08-31 a unit restore took 2.3–4.0 s per app
|
||
// regardless of size (185 KB → 2.25 s, 213 MB → 3.20 s — the cost is per-snapshot round-trip plus
|
||
// roughly 1 s per 200 MB), and all eight apps back to back took 25 s. Ten minutes is ~150x the
|
||
// slowest single app measured, so the failure mode of a wedged SFTP mount is a released flag and a
|
||
// retry tomorrow, never a box whose backups stop because a proof never returned. It is deliberately
|
||
// well under `integrityCheckTimeout` (30 min): this job downloads ONE app's unit, not the store.
|
||
const proofRestoreTimeout = 10 * time.Minute
|
||
|
||
// ProofResult is what one nightly proof did, so the log, the event, the report and any future debug
|
||
// endpoint state the same facts from one place.
|
||
//
|
||
// Skipped is a first-class outcome and NOT a failure — a proof that yielded to a running backup did
|
||
// the right thing. NoSnapshot is separated from a failure for the same reason R-185 separated "I
|
||
// could not look" from "I looked and it is broken": an app with no off-site snapshot yet is a
|
||
// staleness question, which R-339 and the hub's tier deadline already own, and alarming here would be
|
||
// a second alarm for a fact that already has one.
|
||
type ProofResult struct {
|
||
RanAt time.Time
|
||
Duration time.Duration
|
||
Stack string
|
||
SnapshotID string
|
||
Judgement UnitProofResult
|
||
|
||
Skipped bool
|
||
SkipReason string
|
||
NoSnapshot bool
|
||
// Err is a restore/plumbing failure — NOT a verdict about the backup. It never becomes the
|
||
// "intact but empty" alarm, because a download that did not finish has seen nothing.
|
||
Err error
|
||
// CannotRun (R-414) is set when the proof could not START on this box at all — today the only
|
||
// cause is that there is nowhere to put a scratch.
|
||
//
|
||
// IT IS SEPARATE FROM Err ON PURPOSE. An Err is transient: a download that failed says nothing
|
||
// about the backup and nothing about the box, so it is right that it records nothing and retries.
|
||
// This is a STANDING PROPERTY of the machine — it will be just as true tomorrow — and leaving it
|
||
// unrecorded is what R-414 measured: `demo-felhom` failed this way every night with only a WARN,
|
||
// and because nothing was written, `last_proof_result` stayed ABSENT — which is also what a
|
||
// controller older than v0.231.0 sends. The hub could not tell "cannot run here" from "too old to
|
||
// have the feature". That is the StatsKnown trap one level up, and this field is what closes it.
|
||
CannotRun bool
|
||
CannotWhy string
|
||
}
|
||
|
||
// Verdict renders the outcome for the persisted record and the wire. "" is reserved for "no verdict
|
||
// was reached", which is what a skip, a missing snapshot and a restore error all are.
|
||
func (r ProofResult) Verdict() string {
|
||
if r.CannotRun {
|
||
// A recorded, distinguishable state — never "" , which the hub reads as "not recorded".
|
||
return ProofResultCannotRun
|
||
}
|
||
if r.Skipped || r.NoSnapshot || r.Err != nil {
|
||
return ""
|
||
}
|
||
return string(r.Judgement.Verdict)
|
||
}
|
||
|
||
// ProofResultCannotRun is the persisted value for R-414's outcome. It is deliberately NOT one of the
|
||
// UnitProofVerdict values: those judge a UNIT, and here no unit was ever looked at.
|
||
const ProofResultCannotRun = "cannot_run"
|
||
|
||
// ProveOffboxUnit runs ONE nightly proof: pick the due app, restore its newest snapshot read-only to
|
||
// a throwaway directory, judge it, delete the copy, and record WHICH SNAPSHOT was proved.
|
||
//
|
||
// The order below is not arrangeable. The first three guards run before anything is downloaded, and
|
||
// the scratch is removed on EVERY path including failure — a failed judgement that leaves a copy
|
||
// behind is a slow disk-filler and, worse, a copy nobody knows the provenance of.
|
||
func (m *Manager) ProveOffboxUnit(ctx context.Context) ProofResult {
|
||
res := ProofResult{RanAt: time.Now()}
|
||
|
||
if !m.OffboxConfigured() {
|
||
// A box with no off-site tier has nothing to prove. Silent, and not an alarm.
|
||
res.Skipped, res.SkipReason = true, "no off-site target configured"
|
||
return res
|
||
}
|
||
if err := m.acquireRunning(); err != nil {
|
||
res.Skipped, res.SkipReason = true, "a backup or restore is already running"
|
||
m.logger.Printf("[INFO] [offbox] proof: skipped — %s; due-ness is NOT advanced, so this retries on the next run", res.SkipReason)
|
||
return res
|
||
}
|
||
defer m.releaseRunning()
|
||
|
||
stack, id, unitPath, ok := m.nextProofTarget(ctx)
|
||
if !ok {
|
||
res.NoSnapshot = true
|
||
m.logger.Printf("[INFO] [offbox] proof: nothing due — every deployed app's newest off-site snapshot has already been proved, or none has one yet")
|
||
return res
|
||
}
|
||
res.Stack, res.SnapshotID = stack, id
|
||
|
||
scratch, nsRoot, derr := m.offboxProofScratchDir(stack)
|
||
if derr != nil {
|
||
// R-414: a standing property of this box, not a transient failure — so it is RECORDED, not
|
||
// merely warned about. See ProofResult.CannotRun.
|
||
res.CannotRun, res.CannotWhy = true, derr.Error()
|
||
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to, and this box will fail the same way every night until a data drive is registered: %v", stack, derr)
|
||
return res
|
||
}
|
||
// The headroom gate runs BEFORE the download, through the same helper and the same Hungarian
|
||
// wording the customer's restore uses (`unitOnlyHeadroom`). A gate after the download is not a
|
||
// gate.
|
||
if herr := unitOnlyHeadroom(m.offboxFree()(nsRoot)); herr != nil {
|
||
res.Err = herr
|
||
m.logger.Printf("[WARN] [offbox] proof: %s refused before any download: %v", stack, herr)
|
||
return res
|
||
}
|
||
|
||
// Whatever happens from here, the copy goes. Registered before the restore so an early return
|
||
// cannot leak one, and it removes a stale copy from an interrupted previous run on the way in.
|
||
m.removeProofScratch(stack, scratch)
|
||
defer m.removeProofScratch(stack, scratch)
|
||
|
||
start := time.Now()
|
||
if rerr := m.restoreUnitReadOnly(ctx, stack, id, unitPath, scratch); rerr != nil {
|
||
res.Duration = time.Since(start)
|
||
res.Err = rerr
|
||
m.logger.Printf("[ERROR] [offbox] proof: %s could not be restored for proving (nothing is concluded about the backup): %v", stack, rerr)
|
||
return res
|
||
}
|
||
res.Duration = time.Since(start)
|
||
|
||
// The unit sits inside the scratch at its own absolute snapshot path.
|
||
res.Judgement = JudgeRestoredUnit(filepath.Join(scratch, strings.TrimPrefix(unitPath, string(filepath.Separator))))
|
||
switch res.Judgement.Verdict {
|
||
case UnitProofPass:
|
||
m.logger.Printf("[INFO] [offbox] proof: %s PASSED on snapshot %s in %s — the backup holds what this app should have",
|
||
stack, id, res.Duration.Round(time.Millisecond))
|
||
case UnitProofCannotJudge:
|
||
m.logger.Printf("[WARN] [offbox] proof: %s on snapshot %s could NOT be judged (%s) — recorded as such, never as a pass",
|
||
stack, id, res.Judgement.Reason)
|
||
default:
|
||
m.logger.Printf("[ERROR] [offbox] proof: %s on snapshot %s is READABLE AND EMPTY (%s%s) — the store is not damaged; the backup does not contain this app's data",
|
||
stack, id, res.Judgement.Reason, missingSuffix(res.Judgement.Missing))
|
||
}
|
||
return res
|
||
}
|
||
|
||
// missingSuffix renders the Missing list for a log line without ever printing a path outside the unit.
|
||
func missingSuffix(missing []string) string {
|
||
if len(missing) == 0 {
|
||
return ""
|
||
}
|
||
return ": " + strings.Join(missing, ", ")
|
||
}
|
||
|
||
// RecordProofVerdict persists the outcome and advances due-ness. The caller owns WHEN a verdict
|
||
// counts, exactly as `RecordIntegrityVerdict` does.
|
||
//
|
||
// A SKIP, A MISSING SNAPSHOT AND A RESTORE ERROR DO NOT REACH A VERDICT and must not advance
|
||
// due-ness — tomorrow tries again. A `cannot_judge` DOES advance it: re-downloading the same
|
||
// unjudgeable snapshot every night is load with no new information, the outcome is recorded where a
|
||
// surface can read it, and a new snapshot makes the app due again by itself.
|
||
func (m *Manager) RecordProofVerdict(res ProofResult) {
|
||
v := res.Verdict()
|
||
if v == "" {
|
||
return
|
||
}
|
||
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) {
|
||
if v == ProofResultCannotRun {
|
||
// RECORDED so the hub can see the state, but due-ness is deliberately NOT advanced: no
|
||
// snapshot was proved, and marking one proved would stop the app ever being retried once a
|
||
// drive is finally registered. A status report, not a proof.
|
||
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
|
||
o.LastProofStack = res.Stack
|
||
o.LastProofSnapshot = ""
|
||
o.LastProofResult = v
|
||
o.LastProofReason = res.CannotWhy
|
||
return
|
||
}
|
||
if o.ProvedSnapshots == nil {
|
||
o.ProvedSnapshots = map[string]string{}
|
||
}
|
||
o.ProvedSnapshots[res.Stack] = res.SnapshotID
|
||
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
|
||
o.LastProofStack = res.Stack
|
||
o.LastProofSnapshot = res.SnapshotID
|
||
o.LastProofResult = v
|
||
o.LastProofReason = string(res.Judgement.Reason)
|
||
}); err != nil {
|
||
m.logger.Printf("[ERROR] [offbox] proof: could not persist the outcome: %v — the proof RAN and its verdict for %s on %s was %q, but due-ness did not advance, so it will run again tomorrow",
|
||
err, res.Stack, res.SnapshotID, v)
|
||
}
|
||
}
|
||
|
||
// nextProofTarget picks ONE app: the deployed app whose newest off-site snapshot has not been proved.
|
||
//
|
||
// DUE-NESS IS PER SNAPSHOT, NOT PER CLOCK (R-86's model, 07 §3). An app is due when the ID of its
|
||
// newest snapshot differs from the ID last proved for it. A timestamp would re-prove the same
|
||
// snapshot forever and say nothing about the newest one — which is the defect R-341 recorded in a
|
||
// different costume.
|
||
//
|
||
// Apps are considered in a stable sorted order and the FIRST due one wins, so eight apps are covered
|
||
// in eight nights and the rotation cannot starve one: an app stops being due only once its newest
|
||
// snapshot has actually been proved.
|
||
func (m *Manager) nextProofTarget(ctx context.Context) (stack, id, unitPath string, ok bool) {
|
||
if m.stackProvider == nil {
|
||
return "", "", "", false
|
||
}
|
||
proved := map[string]string{}
|
||
if t := m.settings.GetOffboxTarget(); t != nil && t.ProvedSnapshots != nil {
|
||
proved = t.ProvedSnapshots
|
||
}
|
||
var names []string
|
||
for _, st := range m.stackProvider.ListDeployedStacks() {
|
||
names = append(names, st.Name)
|
||
}
|
||
sort.Strings(names)
|
||
for _, name := range names {
|
||
if !isSafeStackName(name) {
|
||
continue
|
||
}
|
||
sid, paths, err := m.offboxLatestSnapshot(ctx, name)
|
||
if err != nil || strings.TrimSpace(sid) == "" {
|
||
// No snapshot yet for this app is NOT a failure of this test — staleness belongs to
|
||
// R-339 and the hub's tier deadline. Logged, never alarmed.
|
||
continue
|
||
}
|
||
if proved[name] == sid {
|
||
continue // already proved on exactly this snapshot
|
||
}
|
||
up := offboxUnitPathOf(paths, name)
|
||
if up == "" {
|
||
// The snapshot exists and carries no recovery unit. That IS the strongest form of
|
||
// "nothing recoverable", so it is a target rather than a skip — the judgement below will
|
||
// find no manifest and fail closed.
|
||
m.logger.Printf("[WARN] [offbox] proof: %s snapshot %s lists no recovery-unit path", name, sid)
|
||
continue
|
||
}
|
||
return name, sid, up, true
|
||
}
|
||
return "", "", "", false
|
||
}
|
||
|
||
// restoreUnitReadOnly downloads ONE recovery unit into scratch WITHOUT writing anything to the
|
||
// repository.
|
||
//
|
||
// Three differences from `RestoreOffboxScratch`, each deliberate and each named:
|
||
// - `--no-lock`, so restic does not create a lock file (it does not for `restore` in 0.14.0 anyway
|
||
// — measured — but the flag makes that a guarantee rather than an observed behaviour of one
|
||
// version);
|
||
// - no `unlockStale`, so no delete verb is issued against `locks/`;
|
||
// - `m.runner()` rather than `resticStep`, so the `unlock --remove-all` escalation is unreachable
|
||
// rather than merely unlikely.
|
||
//
|
||
// It deliberately does NOT write the R-358 completion marker: that marker certifies a scratch for
|
||
// PLACEMENT into a live app, and a proof copy must never be placeable. Its absence is what keeps this
|
||
// directory inert even if the two roots were ever confused.
|
||
func (m *Manager) restoreUnitReadOnly(ctx context.Context, stack, id, unitPath, scratch string) error {
|
||
if err := os.MkdirAll(scratch, 0o755); err != nil {
|
||
return fmt.Errorf("proof scratch dir: %w", err)
|
||
}
|
||
t := m.settings.GetOffboxTarget()
|
||
base, env := m.offboxBaseArgs(t)
|
||
rctx, cancel := context.WithTimeout(ctx, proofRestoreTimeout)
|
||
defer cancel()
|
||
args := append(append([]string{}, base...),
|
||
"restore", id, "--target", scratch, "--include", unitPath, "--no-lock")
|
||
out, err := m.runner()(rctx, env, args...)
|
||
if err != nil {
|
||
return fmt.Errorf("offbox proof restore %s: %w: %s", stack, err, truncate(out))
|
||
}
|
||
return nil
|
||
}
|
||
|
||
// removeProofScratch deletes a proof copy, refusing loudly if the path is not strictly inside a
|
||
// proof root — the same prefix instinct as `DeleteOffsiteRestoreCopy`, because reaching here with an
|
||
// out-of-sandbox path means a helper above is wrong and a best-effort skip would hide that.
|
||
func (m *Manager) removeProofScratch(stack, scratch string) {
|
||
clean := filepath.Clean(scratch)
|
||
for _, drive := range m.offsiteRestoreDriveRoots() {
|
||
root := filepath.Clean(m.offsiteProofRootFor(drive)) + string(filepath.Separator)
|
||
if strings.HasPrefix(clean+string(filepath.Separator), root) {
|
||
if err := os.RemoveAll(clean); err != nil {
|
||
m.logger.Printf("[WARN] [offbox] proof: could not remove the proof copy %s: %v", clean, err)
|
||
}
|
||
return
|
||
}
|
||
}
|
||
m.logger.Printf("[WARN] [offbox] proof: refusing to remove %s — it is not inside a proof root", clean)
|
||
}
|