a70c398dee15d06c563c335f20fb884997f8240a
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fcef8e069c |
one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:
RestoreOffboxScratch - the reported one
OffboxRestorePrepareFull - the SECOND request in the customer's own two-step full-restore
flow, and the one that actually shells `restic stats`. The UI
reaches it FIRST, so flagging only the restore would have left
the collision reachable by the ordinary path.
RestoreSharesScratch - R-411's exact shape on the shares tier: unlockStale + resticStep,
a live web caller, and its sibling PlaceSharesRestore has always
taken the flag.
RestoreOffbox - no production caller today, but the same dangerous pattern.
Flagged rather than left for a future caller to inherit.
OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.
THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.
R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.
R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.
The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.
And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.
R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.
16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
|
||
|
|
3c49dc8ea4 |
v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged without changing its size made plain `restic check` report "no errors were found" on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store: 35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means not-configured, therefore the default; a malformed value falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs a WARN naming the duration, the depth and R-401 — operator log only, no hub event, no depth change. The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT RECORDED, never "structure"). R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on page LOAD, so those panels were permanently blank. backup/crossdrive implemented; backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both storage/simulate-* deleted with their panels and JavaScript. scripts/debug_route_gate.py fails in both directions and is registered after the seven were resolved. 18 referenced, 18 dispatched, none orphaned. Corrections: the dead-field warning in report/types.go said the controller runs no integrity check and the notifiers are called from nowhere — both false since v0.227.0. controller.yaml.example gains its missing integrity: block. integrityCheckTimeout's "ships OFF" comment rewritten. |
||
|
|
45770f2282 |
v0.227.1: the damage classifier matched restic's ordinary progress output
gates / gates (push) Successful in 11s
A patch and not a rebuilt 0.227.0: that tag was already running on demo-hp, and re-pushing changed bytes under a live tag is the :latest hazard with extra steps. looksLikeRepositoryDamage matched bare "pack ", "tree ", "snapshot ", "blob ". A HEALTHY restic check prints "check all packs" and "check snapshots, trees and blobs" -- so any check that failed for a NON-damage reason, a connection dropped mid-run for instance, would have been classified as a corrupted repository and told the customer their backups may be damaged. That is the false alarm that teaches an operator to ignore the true one. Caught by the NEGATIVE control, built from the real bytes of a real passing check on demo-hp. The spec made the negative control mandatory and this is what it was for: a control that has only ever seen the failing case proves nothing. Signatures are now phrases from restic's own error wording. Also in this commit: CONTEXT.md records the three rulings (take the flag and skip, due-ness not a weekday, publish on OffboxReportStatus not the R-331 dead fields) plus the measurement a future session would otherwise assume wrongly -- THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. README documents the job, the route and the config, and corrects a line that listed four debug backup routes when only two exist. REUSE gains three rows, including one that records R-398 was my own mistake so nobody re-files it. |
||
|
|
0d52a42c17 |
R-359 + R-397: the off-site store gets checked, and the advertised check becomes real
gates / gates (push) Successful in 12s
Nothing ever verified that the off-site copies are still readable. The whole-guest tier has verify jobs; the tier holding the customer's documents and photos had none -- the complete set of restic verbs this controller used contained no `check`. We would have found out at restore time, with a customer waiting. On 2026-08-21 a deliberately damaged pack was caught at once by plain `restic check`; we had never run it. R-397: NotifyIntegrityOK/NotifyIntegrityFailed existed with no caller, the hub allowlists both event types and carries the Hungarian text for both, the settings checkbox exists, and the debug button posts to /api/debug/backup/ integrity. Everything was built except the part that runs. SIXTH instance of that shape in this project. THE HAZARD SHAPES THE WHOLE DESIGN. resticStep self-heals a crash lock by running `unlock --remove-all` and retrying, and its own comment records why that is safe: every caller holds the in-process single-flight mutex, so any lock it meets is stale. A check that did not take that flag could meet a LIVE prune's lock from this same box, remove it, and retry over the top of it. So the check TAKES THE FLAG and SKIPS rather than waits -- waiting would pin the nightly backup behind it, and a skip costs nothing because due-ness makes tomorrow try again. TestR359_SkipsWhenRunningFlagHeld asserts the NON-EFFECTS: restic never invoked, `unlock` never in any argv. Its red-proof prints the real thing -- restic running `check` while the flag was held. DUE-NESS, NOT A WEEKDAY. Daily job, weekly behaviour: "is the last successful check older than 7 days?" not "is it Sunday?". R-341 is exactly the other shape, a dated check quietly missed and never caught up. No Weekly primitive added. THREE OUTCOMES, NOT TWO. Skipped, Unreachable and failed are different facts. "I could not look" is not "I looked and it is broken" -- R-339 already owns reachability, and a second alarm for the same fact trains the operator to discount the one alarm that means the backups are damaged. A timeout is unreachable, never damage. A failure advances due-ness (a broken store must not be re-checked nightly); a skip and an unreachable store do not. Success is severity `info`, which severityNotifies DROPS -- it mails NOBODY, by design. A weekly success e-mail is how people stop reading their alerts. The customer gets a SENTENCE; restic's words go to the log, truncated (R-379: 615 bytes of raw database text reached a customer once). read-data-subset ships OFF and a malformed value is refused at read time rather than handed to restic, where one typo would fail the whole check. Published on OffboxReportStatus, NOT on report.BackupReport's IntegrityOK -- those were retired by R-331 YESTERDAY and TestBackupReport_DeadFieldsStayZero still passes unmodified. Also: the monitoring page stopped promising a Sunday job that never existed, and the debug button got its dispatch case. PART 0 WAS NOT BUILT, AND R-398 WAS MY OWN MISTAKE. The seam it asked for already exists: offboxRunner/SetOffboxRunner/m.runner() has been injectable since the off-site tier shipped, and other tests drive restic-backed paths through it. A resticStepFn seam would have been WORSE here -- it would replace the `unlock --remove-all` escalation and hide it from the assertions that must see it. R-358's AST ordering test is converted to a real execution test instead, which immediately surfaced something the AST walk could not: unlockStale legitimately runs before the restore. Four red-proofs, each printing the pre-fix behaviour. Green gate: 28 packages, rc 0. All 12 controller gates OK. |