one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:
RestoreOffboxScratch - the reported one
OffboxRestorePrepareFull - the SECOND request in the customer's own two-step full-restore
flow, and the one that actually shells `restic stats`. The UI
reaches it FIRST, so flagging only the restore would have left
the collision reachable by the ordinary path.
RestoreSharesScratch - R-411's exact shape on the shares tier: unlockStale + resticStep,
a live web caller, and its sibling PlaceSharesRestore has always
taken the flag.
RestoreOffbox - no production caller today, but the same dangerous pattern.
Flagged rather than left for a future caller to inherit.
OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.
THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.
R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.
R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.
The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.
And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.
R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.
16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
This commit is contained in:
@@ -1343,7 +1343,21 @@ func (m *Manager) runOffboxInternal(ctx context.Context, apps, base, env []strin
|
||||
continue
|
||||
}
|
||||
res.backedUp++
|
||||
m.logger.Printf("[INFO] [offbox] backed up %s (%s, %d mandatory path(s))", stack, src, len(extra))
|
||||
// R-412 leg 1 — A PUSH THAT CARRIED NOTHING MUST NOT READ AS A PLAIN SUCCESS.
|
||||
//
|
||||
// Measured on demo-hp 2026-08-31: a recovery unit was destroyed mid-run, the capture rebuilt it
|
||||
// empty, and this line printed `backed up opengist (…, 0 mandatory path(s))` over a snapshot
|
||||
// holding none of the app's data. Every word of it was true and the whole of it was misleading.
|
||||
//
|
||||
// WORDING ONLY. This adds no guard and does NOT touch the capture — `07` §8.2 records why
|
||||
// guarding the capture would make the manifest lie. The reader of a log gets the fact; whether
|
||||
// the push should refuse is R-412 leg 2 and is still open.
|
||||
if unitIsHollow(src) {
|
||||
m.logger.Printf("[WARN] [offbox] backed up %s (%s, %d mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's data; the next run with a dump leg will replace it",
|
||||
stack, src, len(extra))
|
||||
} else {
|
||||
m.logger.Printf("[INFO] [offbox] backed up %s (%s, %d mandatory path(s))", stack, src, len(extra))
|
||||
}
|
||||
}
|
||||
// R-7b: the SHARES leg runs AFTER the per-app loop and BEFORE retention, so `forget --group-by
|
||||
// host,tags` covers the `_shares` group for free. It is placed BEFORE the firstErr return on
|
||||
@@ -1788,6 +1802,14 @@ func (m *Manager) offboxRecordStats(ctx context.Context, base, env []string) int
|
||||
// RestoreOffbox restores an app's latest off-box snapshot to destDir (a scratch/verify location — it does
|
||||
// NOT overwrite live data). Returns an error on failure (checks restic's own exit code).
|
||||
func (m *Manager) RestoreOffbox(ctx context.Context, stackName, destDir string) error {
|
||||
// R-411 — also found by the R-408 walk. This one has NO production caller today (only two tests
|
||||
// reach it), but it runs `unlockStale` + `resticStep restore`, which is precisely the pattern that
|
||||
// deleted a live restore's lock on demo-hp. Dead code gets resurrected; flagging it now costs
|
||||
// nothing and means a future caller inherits the guard rather than the defect.
|
||||
if err := m.acquireRunning(); err != nil {
|
||||
return err
|
||||
}
|
||||
defer m.releaseRunning()
|
||||
if !m.OffboxConfigured() {
|
||||
return fmt.Errorf("off-box backup not configured")
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user