one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:
RestoreOffboxScratch - the reported one
OffboxRestorePrepareFull - the SECOND request in the customer's own two-step full-restore
flow, and the one that actually shells `restic stats`. The UI
reaches it FIRST, so flagging only the restore would have left
the collision reachable by the ordinary path.
RestoreSharesScratch - R-411's exact shape on the shares tier: unlockStale + resticStep,
a live web caller, and its sibling PlaceSharesRestore has always
taken the flag.
RestoreOffbox - no production caller today, but the same dangerous pattern.
Flagged rather than left for a future caller to inherit.
OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.
THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.
R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.
R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.
The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.
And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.
R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.
16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
This commit is contained in:
@@ -99,17 +99,37 @@ type ProofResult struct {
|
||||
// Err is a restore/plumbing failure — NOT a verdict about the backup. It never becomes the
|
||||
// "intact but empty" alarm, because a download that did not finish has seen nothing.
|
||||
Err error
|
||||
// CannotRun (R-414) is set when the proof could not START on this box at all — today the only
|
||||
// cause is that there is nowhere to put a scratch.
|
||||
//
|
||||
// IT IS SEPARATE FROM Err ON PURPOSE. An Err is transient: a download that failed says nothing
|
||||
// about the backup and nothing about the box, so it is right that it records nothing and retries.
|
||||
// This is a STANDING PROPERTY of the machine — it will be just as true tomorrow — and leaving it
|
||||
// unrecorded is what R-414 measured: `demo-felhom` failed this way every night with only a WARN,
|
||||
// and because nothing was written, `last_proof_result` stayed ABSENT — which is also what a
|
||||
// controller older than v0.231.0 sends. The hub could not tell "cannot run here" from "too old to
|
||||
// have the feature". That is the StatsKnown trap one level up, and this field is what closes it.
|
||||
CannotRun bool
|
||||
CannotWhy string
|
||||
}
|
||||
|
||||
// Verdict renders the outcome for the persisted record and the wire. "" is reserved for "no verdict
|
||||
// was reached", which is what a skip, a missing snapshot and a restore error all are.
|
||||
func (r ProofResult) Verdict() string {
|
||||
if r.CannotRun {
|
||||
// A recorded, distinguishable state — never "" , which the hub reads as "not recorded".
|
||||
return ProofResultCannotRun
|
||||
}
|
||||
if r.Skipped || r.NoSnapshot || r.Err != nil {
|
||||
return ""
|
||||
}
|
||||
return string(r.Judgement.Verdict)
|
||||
}
|
||||
|
||||
// ProofResultCannotRun is the persisted value for R-414's outcome. It is deliberately NOT one of the
|
||||
// UnitProofVerdict values: those judge a UNIT, and here no unit was ever looked at.
|
||||
const ProofResultCannotRun = "cannot_run"
|
||||
|
||||
// ProveOffboxUnit runs ONE nightly proof: pick the due app, restore its newest snapshot read-only to
|
||||
// a throwaway directory, judge it, delete the copy, and record WHICH SNAPSHOT was proved.
|
||||
//
|
||||
@@ -141,8 +161,10 @@ func (m *Manager) ProveOffboxUnit(ctx context.Context) ProofResult {
|
||||
|
||||
scratch, nsRoot, derr := m.offboxProofScratchDir(stack)
|
||||
if derr != nil {
|
||||
res.Err = derr
|
||||
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to: %v", stack, derr)
|
||||
// R-414: a standing property of this box, not a transient failure — so it is RECORDED, not
|
||||
// merely warned about. See ProofResult.CannotRun.
|
||||
res.CannotRun, res.CannotWhy = true, derr.Error()
|
||||
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to, and this box will fail the same way every night until a data drive is registered: %v", stack, derr)
|
||||
return res
|
||||
}
|
||||
// The headroom gate runs BEFORE the download, through the same helper and the same Hungarian
|
||||
@@ -205,6 +227,17 @@ func (m *Manager) RecordProofVerdict(res ProofResult) {
|
||||
return
|
||||
}
|
||||
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) {
|
||||
if v == ProofResultCannotRun {
|
||||
// RECORDED so the hub can see the state, but due-ness is deliberately NOT advanced: no
|
||||
// snapshot was proved, and marking one proved would stop the app ever being retried once a
|
||||
// drive is finally registered. A status report, not a proof.
|
||||
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
|
||||
o.LastProofStack = res.Stack
|
||||
o.LastProofSnapshot = ""
|
||||
o.LastProofResult = v
|
||||
o.LastProofReason = res.CannotWhy
|
||||
return
|
||||
}
|
||||
if o.ProvedSnapshots == nil {
|
||||
o.ProvedSnapshots = map[string]string{}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user