one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s

THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:

  RestoreOffboxScratch      - the reported one
  OffboxRestorePrepareFull  - the SECOND request in the customer's own two-step full-restore
                              flow, and the one that actually shells `restic stats`. The UI
                              reaches it FIRST, so flagging only the restore would have left
                              the collision reachable by the ordinary path.
  RestoreSharesScratch      - R-411's exact shape on the shares tier: unlockStale + resticStep,
                              a live web caller, and its sibling PlaceSharesRestore has always
                              taken the flag.
  RestoreOffbox             - no production caller today, but the same dangerous pattern.
                              Flagged rather than left for a future caller to inherit.

OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.

THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.

R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.

R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.

The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.

And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.

R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.

16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
This commit is contained in:
2026-09-01 10:19:36 +02:00
parent 9aea86cd48
commit fcef8e069c
9 changed files with 903 additions and 14 deletions
+33 -8
View File
@@ -27,6 +27,18 @@ import (
// *"the in-process single-flight mutex (held by every caller of this method) proves no sibling
// operation is live"*. Every off-site operation takes `acquireRunning` for exactly that reason.
//
// **THAT SENTENCE WAS FALSE FROM v0.227.0 UNTIL v0.232.0, AND NOTHING CHECKED IT.** `RestoreOffboxScratch`,
// `OffboxRestorePrepareFull`, `RestoreSharesScratch` and `RestoreOffbox` all issued restic commands
// without the flag. R-411 measured the consequence on demo-hp 2026-08-31: a customer full-restore's
// `restic stats` held a repository lock, this check was therefore not blocked, met that lock, and
// `resticStep` removed it with `unlock --remove-all` while logging *"a stale exclusive lock left by a
// previous crash"*. There was no crash.
//
// The sentence is TRUE again as of v0.232.0, and it is now **pinned by
// `TestR408_EveryOffsiteEntryPointTakesTheFlagOrIsRegistered`** — an AST walk over this package that
// fails when a new off-site entry point takes neither the flag nor a registered exemption. Do not
// trust this paragraph; the test is what holds it.
//
// **An integrity check that did not take the flag would break that invariant.** It could meet the
// lock of a `forget --prune` running from this same box, remove it, and retry over the top of a live
// prune. So this check TAKES THE FLAG, and **skips rather than waits** when it cannot get it: waiting
@@ -252,8 +264,21 @@ func (m *Manager) recordIntegrityOutcome(at time.Time, ok bool, depth string) {
}
}
// CheckOffboxIntegrity runs one off-site integrity check. It NEVER writes to the repository: `check`
// is a read verb, and nothing here prunes, forgets, unlocks or backs up.
// CheckOffboxIntegrity runs one off-site integrity check. It prunes nothing, forgets nothing, unlocks
// nothing and backs up nothing.
//
// **IT DOES, HOWEVER, WRITE A LOCK FILE, and the earlier wording here — "It NEVER writes to the
// repository" — was wrong (R-407).** Observed on demo-hp 2026-08-31 with a lock sampler that was
// positively controlled first: across a real check the repository went `locks=0` -> `locks=1` for nine
// consecutive samples -> `locks=0`. `restic check` takes a lock like most restic verbs; only `--no-lock`
// avoids it, and this check deliberately does not pass it, because the lock is what makes the check
// exclusive against a prune.
//
// **AND THE FACT NOBODY HAD, which is what made R-411 possible: `restic stats` ALSO takes a lock.**
// Measured in a clean room the same night — nothing else running, four invocations, `locks=1`. That is
// why a customer full-restore, whose size probe shells `stats`, could collide with this check at all.
// The sentence is corrected in place rather than deleted, because R-360's rule is that a comment
// claiming a guard is exactly why nobody looks for the missing one.
//
// Due-ness is NOT consulted here — the caller decides. The scheduled job asks IntegrityDue first; the
// operator's debug button deliberately does not, because "run it now" is the whole point of a button.
@@ -355,14 +380,14 @@ func looksLikeRepositoryDamage(out []byte) bool {
// Every phrase below is restic's own error wording, taken from the 2026-08-30 damaged-pack run on
// demo-hp or from restic's check source — never paraphrased.
for _, sig := range []string{
"does not match", // "Pack ID does not match, want <id>, got <id>" — the measured one
"not found in index", // a blob or pack the index promises and the store lacks
"blob not found", //
"size mismatch", //
"does not match", // "Pack ID does not match, want <id>, got <id>" — the measured one
"not found in index", // a blob or pack the index promises and the store lacks
"blob not found", //
"size mismatch", //
"ciphertext verification failed",
"integrity error",
"repository contains errors", // restic's own summary verdict
"failed to load index", //
"repository contains errors", // restic's own summary verdict
"failed to load index", //
} {
if strings.Contains(s, sig) {
return true