one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s

THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:

  RestoreOffboxScratch      - the reported one
  OffboxRestorePrepareFull  - the SECOND request in the customer's own two-step full-restore
                              flow, and the one that actually shells `restic stats`. The UI
                              reaches it FIRST, so flagging only the restore would have left
                              the collision reachable by the ordinary path.
  RestoreSharesScratch      - R-411's exact shape on the shares tier: unlockStale + resticStep,
                              a live web caller, and its sibling PlaceSharesRestore has always
                              taken the flag.
  RestoreOffbox             - no production caller today, but the same dangerous pattern.
                              Flagged rather than left for a future caller to inherit.

OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.

THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.

R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.

R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.

The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.

And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.

R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.

16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
This commit is contained in:
2026-09-01 10:19:36 +02:00
parent 9aea86cd48
commit fcef8e069c
9 changed files with 903 additions and 14 deletions
+35 -2
View File
@@ -99,17 +99,37 @@ type ProofResult struct {
// Err is a restore/plumbing failure — NOT a verdict about the backup. It never becomes the
// "intact but empty" alarm, because a download that did not finish has seen nothing.
Err error
// CannotRun (R-414) is set when the proof could not START on this box at all — today the only
// cause is that there is nowhere to put a scratch.
//
// IT IS SEPARATE FROM Err ON PURPOSE. An Err is transient: a download that failed says nothing
// about the backup and nothing about the box, so it is right that it records nothing and retries.
// This is a STANDING PROPERTY of the machine — it will be just as true tomorrow — and leaving it
// unrecorded is what R-414 measured: `demo-felhom` failed this way every night with only a WARN,
// and because nothing was written, `last_proof_result` stayed ABSENT — which is also what a
// controller older than v0.231.0 sends. The hub could not tell "cannot run here" from "too old to
// have the feature". That is the StatsKnown trap one level up, and this field is what closes it.
CannotRun bool
CannotWhy string
}
// Verdict renders the outcome for the persisted record and the wire. "" is reserved for "no verdict
// was reached", which is what a skip, a missing snapshot and a restore error all are.
func (r ProofResult) Verdict() string {
if r.CannotRun {
// A recorded, distinguishable state — never "" , which the hub reads as "not recorded".
return ProofResultCannotRun
}
if r.Skipped || r.NoSnapshot || r.Err != nil {
return ""
}
return string(r.Judgement.Verdict)
}
// ProofResultCannotRun is the persisted value for R-414's outcome. It is deliberately NOT one of the
// UnitProofVerdict values: those judge a UNIT, and here no unit was ever looked at.
const ProofResultCannotRun = "cannot_run"
// ProveOffboxUnit runs ONE nightly proof: pick the due app, restore its newest snapshot read-only to
// a throwaway directory, judge it, delete the copy, and record WHICH SNAPSHOT was proved.
//
@@ -141,8 +161,10 @@ func (m *Manager) ProveOffboxUnit(ctx context.Context) ProofResult {
scratch, nsRoot, derr := m.offboxProofScratchDir(stack)
if derr != nil {
res.Err = derr
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to: %v", stack, derr)
// R-414: a standing property of this box, not a transient failure — so it is RECORDED, not
// merely warned about. See ProofResult.CannotRun.
res.CannotRun, res.CannotWhy = true, derr.Error()
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to, and this box will fail the same way every night until a data drive is registered: %v", stack, derr)
return res
}
// The headroom gate runs BEFORE the download, through the same helper and the same Hungarian
@@ -205,6 +227,17 @@ func (m *Manager) RecordProofVerdict(res ProofResult) {
return
}
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) {
if v == ProofResultCannotRun {
// RECORDED so the hub can see the state, but due-ness is deliberately NOT advanced: no
// snapshot was proved, and marking one proved would stop the app ever being retried once a
// drive is finally registered. A status report, not a proof.
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
o.LastProofStack = res.Stack
o.LastProofSnapshot = ""
o.LastProofResult = v
o.LastProofReason = res.CannotWhy
return
}
if o.ProvedSnapshots == nil {
o.ProvedSnapshots = map[string]string{}
}