v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s

R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.

THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.

IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.

THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".

THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.

THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".

IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.

IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.

DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.

ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.

SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.

NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.

33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.

A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
This commit is contained in:
2026-08-31 20:55:34 +02:00
parent 2d802d75e8
commit e43b5ec07d
16 changed files with 1973 additions and 33 deletions
+318
View File
@@ -0,0 +1,318 @@
package backup
import (
"context"
"fmt"
"os"
"path/filepath"
"sort"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// ── R-87 — the box proves its own off-site copy still HOLDS something ────────────────────────────
//
// WHAT THIS ANSWERS THAT NOTHING ELSE DOES. `restic check --read-data-subset=100%` (R-359/R-399) runs
// weekly and proves the stored bytes are the bytes we stored. **It cannot tell us we stored the wrong
// thing.** A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks
// cleanly at full depth, restores cleanly, and gives the customer nothing back. R-403 measured that
// shape on this fleet on 2026-08-31: 120 082 104 B became 7 036 B in one nightly run, recorded as a
// success. Until this job, no tier and no cadence asked the question.
//
// WHAT IT DELIBERATELY DOES NOT ANSWER, said here because the green tick will be read as covering it
// otherwise: **it does not prove a restore puts data back into a running app.** It restores to a
// throwaway directory, looks, and deletes. Putting data back is drill work, and the spike said so —
// of the five restore-path defects human drills found in the six days to 2026-08-31, an unattended
// scratch restore would have caught ONE (R-356). `07-backup-architecture.md` §8 matrix row 4 does not
// move on the strength of this job.
//
// THE SHAPE IS `CheckOffboxIntegrity`'S, deliberately — same file conventions, same guards, same
// three-outcome honesty — because a second scheduled off-site job that reasons differently about the
// single-writer flag is how the two come to disagree about repository safety.
//
// ── THE HAZARD, AND WHY THIS JOB TAKES A FLAG THAT ITS RESTORE DOES NOT ─────────────────────────
//
// `resticStep` self-heals a crash lock by running `unlock --remove-all`, and its own doc comment
// records that this is safe only because *"the in-process single-flight mutex (held by every caller
// of this method) proves no sibling operation is live"*.
//
// `offbox_integrity.go`'s header states that invariant as universal — *"Every off-site operation
// takes acquireRunning for exactly that reason"* — and **it is not true today**:
// `RestoreOffboxScratch` does not take it (R-408, measured 2026-08-31; `restore_wizard.go` records
// the same fact independently and the UI works around it with a separate display flag). This job
// therefore takes the flag ITSELF rather than waiting for R-408 to be fixed, and SKIPS rather than
// waits — a skipped proof simply runs tomorrow, whereas waiting would pin a nightly backup behind it.
//
// ── AND IT NEVER WRITES TO THE REPOSITORY, FOR REAL ────────────────────────────────────────────
//
// R-95's constraint is that the off-site credential can delete, so a scheduled job must not be able
// to write. The spike established what "read-only" actually costs here, by observation rather than by
// reading the code:
//
// - `restic restore` in 0.14.0 takes NO repository lock (measured on demo-hp with a lock sampler
// that was positively controlled first: it saw a lock appear and vanish across a real `check`, and
// zero across two restores);
// - `restic check` DOES take one — so the comment claiming the check never writes is wrong (R-407);
// - the PRODUCT's restore path writes anyway, because `unlockStale` runs `restic unlock` — a delete
// verb against `locks/` — before every restore (`offbox_restore.go`).
//
// So this job runs its restore with `--no-lock` and WITHOUT `unlockStale`, through `m.runner()`
// rather than `resticStep`, so the `unlock --remove-all` escalation is not merely unlikely but
// unreachable. `TestR87_NeverWritesToTheRepository` asserts that as a NON-EFFECT — the argv, not the
// absence of an error. Using the seam rather than replacing `resticStep` is REUSE.md's own rule:
// replacing it would hide the escalation from exactly the assertion that must see it.
//
// `restic restore --verify` is NOT used as a correctness check and must never be: the spike measured
// it passing a byte-level corruption with size and mtime preserved, on a 213 MB tree, in 131 ms. It
// is a size-and-mtime reconciliation. The correctness question is answered by `JudgeRestoredUnit`.
// proofRestoreTimeout bounds one proof restore.
//
// CHOSEN from measurement, not inherited: on demo-hp 2026-08-31 a unit restore took 2.3–4.0 s per app
// regardless of size (185 KB → 2.25 s, 213 MB → 3.20 s — the cost is per-snapshot round-trip plus
// roughly 1 s per 200 MB), and all eight apps back to back took 25 s. Ten minutes is ~150x the
// slowest single app measured, so the failure mode of a wedged SFTP mount is a released flag and a
// retry tomorrow, never a box whose backups stop because a proof never returned. It is deliberately
// well under `integrityCheckTimeout` (30 min): this job downloads ONE app's unit, not the store.
const proofRestoreTimeout = 10 * time.Minute
// ProofResult is what one nightly proof did, so the log, the event, the report and any future debug
// endpoint state the same facts from one place.
//
// Skipped is a first-class outcome and NOT a failure — a proof that yielded to a running backup did
// the right thing. NoSnapshot is separated from a failure for the same reason R-185 separated "I
// could not look" from "I looked and it is broken": an app with no off-site snapshot yet is a
// staleness question, which R-339 and the hub's tier deadline already own, and alarming here would be
// a second alarm for a fact that already has one.
type ProofResult struct {
RanAt time.Time
Duration time.Duration
Stack string
SnapshotID string
Judgement UnitProofResult
Skipped bool
SkipReason string
NoSnapshot bool
// Err is a restore/plumbing failure — NOT a verdict about the backup. It never becomes the
// "intact but empty" alarm, because a download that did not finish has seen nothing.
Err error
}
// Verdict renders the outcome for the persisted record and the wire. "" is reserved for "no verdict
// was reached", which is what a skip, a missing snapshot and a restore error all are.
func (r ProofResult) Verdict() string {
if r.Skipped || r.NoSnapshot || r.Err != nil {
return ""
}
return string(r.Judgement.Verdict)
}
// ProveOffboxUnit runs ONE nightly proof: pick the due app, restore its newest snapshot read-only to
// a throwaway directory, judge it, delete the copy, and record WHICH SNAPSHOT was proved.
//
// The order below is not arrangeable. The first three guards run before anything is downloaded, and
// the scratch is removed on EVERY path including failure — a failed judgement that leaves a copy
// behind is a slow disk-filler and, worse, a copy nobody knows the provenance of.
func (m *Manager) ProveOffboxUnit(ctx context.Context) ProofResult {
res := ProofResult{RanAt: time.Now()}
if !m.OffboxConfigured() {
// A box with no off-site tier has nothing to prove. Silent, and not an alarm.
res.Skipped, res.SkipReason = true, "no off-site target configured"
return res
}
if err := m.acquireRunning(); err != nil {
res.Skipped, res.SkipReason = true, "a backup or restore is already running"
m.logger.Printf("[INFO] [offbox] proof: skipped — %s; due-ness is NOT advanced, so this retries on the next run", res.SkipReason)
return res
}
defer m.releaseRunning()
stack, id, unitPath, ok := m.nextProofTarget(ctx)
if !ok {
res.NoSnapshot = true
m.logger.Printf("[INFO] [offbox] proof: nothing due — every deployed app's newest off-site snapshot has already been proved, or none has one yet")
return res
}
res.Stack, res.SnapshotID = stack, id
scratch, nsRoot, derr := m.offboxProofScratchDir(stack)
if derr != nil {
res.Err = derr
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to: %v", stack, derr)
return res
}
// The headroom gate runs BEFORE the download, through the same helper and the same Hungarian
// wording the customer's restore uses (`unitOnlyHeadroom`). A gate after the download is not a
// gate.
if herr := unitOnlyHeadroom(m.offboxFree()(nsRoot)); herr != nil {
res.Err = herr
m.logger.Printf("[WARN] [offbox] proof: %s refused before any download: %v", stack, herr)
return res
}
// Whatever happens from here, the copy goes. Registered before the restore so an early return
// cannot leak one, and it removes a stale copy from an interrupted previous run on the way in.
m.removeProofScratch(stack, scratch)
defer m.removeProofScratch(stack, scratch)
start := time.Now()
if rerr := m.restoreUnitReadOnly(ctx, stack, id, unitPath, scratch); rerr != nil {
res.Duration = time.Since(start)
res.Err = rerr
m.logger.Printf("[ERROR] [offbox] proof: %s could not be restored for proving (nothing is concluded about the backup): %v", stack, rerr)
return res
}
res.Duration = time.Since(start)
// The unit sits inside the scratch at its own absolute snapshot path.
res.Judgement = JudgeRestoredUnit(filepath.Join(scratch, strings.TrimPrefix(unitPath, string(filepath.Separator))))
switch res.Judgement.Verdict {
case UnitProofPass:
m.logger.Printf("[INFO] [offbox] proof: %s PASSED on snapshot %s in %s — the backup holds what this app should have",
stack, id, res.Duration.Round(time.Millisecond))
case UnitProofCannotJudge:
m.logger.Printf("[WARN] [offbox] proof: %s on snapshot %s could NOT be judged (%s) — recorded as such, never as a pass",
stack, id, res.Judgement.Reason)
default:
m.logger.Printf("[ERROR] [offbox] proof: %s on snapshot %s is READABLE AND EMPTY (%s%s) — the store is not damaged; the backup does not contain this app's data",
stack, id, res.Judgement.Reason, missingSuffix(res.Judgement.Missing))
}
return res
}
// missingSuffix renders the Missing list for a log line without ever printing a path outside the unit.
func missingSuffix(missing []string) string {
if len(missing) == 0 {
return ""
}
return ": " + strings.Join(missing, ", ")
}
// RecordProofVerdict persists the outcome and advances due-ness. The caller owns WHEN a verdict
// counts, exactly as `RecordIntegrityVerdict` does.
//
// A SKIP, A MISSING SNAPSHOT AND A RESTORE ERROR DO NOT REACH A VERDICT and must not advance
// due-ness — tomorrow tries again. A `cannot_judge` DOES advance it: re-downloading the same
// unjudgeable snapshot every night is load with no new information, the outcome is recorded where a
// surface can read it, and a new snapshot makes the app due again by itself.
func (m *Manager) RecordProofVerdict(res ProofResult) {
v := res.Verdict()
if v == "" {
return
}
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) {
if o.ProvedSnapshots == nil {
o.ProvedSnapshots = map[string]string{}
}
o.ProvedSnapshots[res.Stack] = res.SnapshotID
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
o.LastProofStack = res.Stack
o.LastProofSnapshot = res.SnapshotID
o.LastProofResult = v
o.LastProofReason = string(res.Judgement.Reason)
}); err != nil {
m.logger.Printf("[ERROR] [offbox] proof: could not persist the outcome: %v — the proof RAN and its verdict for %s on %s was %q, but due-ness did not advance, so it will run again tomorrow",
err, res.Stack, res.SnapshotID, v)
}
}
// nextProofTarget picks ONE app: the deployed app whose newest off-site snapshot has not been proved.
//
// DUE-NESS IS PER SNAPSHOT, NOT PER CLOCK (R-86's model, 07 §3). An app is due when the ID of its
// newest snapshot differs from the ID last proved for it. A timestamp would re-prove the same
// snapshot forever and say nothing about the newest one — which is the defect R-341 recorded in a
// different costume.
//
// Apps are considered in a stable sorted order and the FIRST due one wins, so eight apps are covered
// in eight nights and the rotation cannot starve one: an app stops being due only once its newest
// snapshot has actually been proved.
func (m *Manager) nextProofTarget(ctx context.Context) (stack, id, unitPath string, ok bool) {
if m.stackProvider == nil {
return "", "", "", false
}
proved := map[string]string{}
if t := m.settings.GetOffboxTarget(); t != nil && t.ProvedSnapshots != nil {
proved = t.ProvedSnapshots
}
var names []string
for _, st := range m.stackProvider.ListDeployedStacks() {
names = append(names, st.Name)
}
sort.Strings(names)
for _, name := range names {
if !isSafeStackName(name) {
continue
}
sid, paths, err := m.offboxLatestSnapshot(ctx, name)
if err != nil || strings.TrimSpace(sid) == "" {
// No snapshot yet for this app is NOT a failure of this test — staleness belongs to
// R-339 and the hub's tier deadline. Logged, never alarmed.
continue
}
if proved[name] == sid {
continue // already proved on exactly this snapshot
}
up := offboxUnitPathOf(paths, name)
if up == "" {
// The snapshot exists and carries no recovery unit. That IS the strongest form of
// "nothing recoverable", so it is a target rather than a skip — the judgement below will
// find no manifest and fail closed.
m.logger.Printf("[WARN] [offbox] proof: %s snapshot %s lists no recovery-unit path", name, sid)
continue
}
return name, sid, up, true
}
return "", "", "", false
}
// restoreUnitReadOnly downloads ONE recovery unit into scratch WITHOUT writing anything to the
// repository.
//
// Three differences from `RestoreOffboxScratch`, each deliberate and each named:
// - `--no-lock`, so restic does not create a lock file (it does not for `restore` in 0.14.0 anyway
// — measured — but the flag makes that a guarantee rather than an observed behaviour of one
// version);
// - no `unlockStale`, so no delete verb is issued against `locks/`;
// - `m.runner()` rather than `resticStep`, so the `unlock --remove-all` escalation is unreachable
// rather than merely unlikely.
//
// It deliberately does NOT write the R-358 completion marker: that marker certifies a scratch for
// PLACEMENT into a live app, and a proof copy must never be placeable. Its absence is what keeps this
// directory inert even if the two roots were ever confused.
func (m *Manager) restoreUnitReadOnly(ctx context.Context, stack, id, unitPath, scratch string) error {
if err := os.MkdirAll(scratch, 0o755); err != nil {
return fmt.Errorf("proof scratch dir: %w", err)
}
t := m.settings.GetOffboxTarget()
base, env := m.offboxBaseArgs(t)
rctx, cancel := context.WithTimeout(ctx, proofRestoreTimeout)
defer cancel()
args := append(append([]string{}, base...),
"restore", id, "--target", scratch, "--include", unitPath, "--no-lock")
out, err := m.runner()(rctx, env, args...)
if err != nil {
return fmt.Errorf("offbox proof restore %s: %w: %s", stack, err, truncate(out))
}
return nil
}
// removeProofScratch deletes a proof copy, refusing loudly if the path is not strictly inside a
// proof root — the same prefix instinct as `DeleteOffsiteRestoreCopy`, because reaching here with an
// out-of-sandbox path means a helper above is wrong and a best-effort skip would hide that.
func (m *Manager) removeProofScratch(stack, scratch string) {
clean := filepath.Clean(scratch)
for _, drive := range m.offsiteRestoreDriveRoots() {
root := filepath.Clean(m.offsiteProofRootFor(drive)) + string(filepath.Separator)
if strings.HasPrefix(clean+string(filepath.Separator), root) {
if err := os.RemoveAll(clean); err != nil {
m.logger.Printf("[WARN] [offbox] proof: could not remove the proof copy %s: %v", clean, err)
}
return
}
}
m.logger.Printf("[WARN] [offbox] proof: refusing to remove %s — it is not inside a proof root", clean)
}