6479933e8ef98f116f2213ec076a22959dbe4ed5
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8b55de734c |
R-414: the fallback scratch must also be DELETABLE - caught by live validation
gates / gates (push) Successful in 12s
The system-data fallback resolved a scratch fine and removeProofScratch then refused to delete it: its accepted-roots list is built from REGISTERED drives, and a driveless box has none. Observed on demo-felhom: 'refusing to remove ... it is not inside a proof root', with the copy still on disk. Every nightly proof would have left one behind, growing forever, on exactly the boxes the fallback exists for. My defect, introduced with the fallback in the same session. The unit tests missed it because every one of them registers a drive; the new pair deliberately does not, and the second asserts the guard still REFUSES a path outside every proof root, so the fix is not a widening into uselessness. |
||
|
|
fcef8e069c |
one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:
RestoreOffboxScratch - the reported one
OffboxRestorePrepareFull - the SECOND request in the customer's own two-step full-restore
flow, and the one that actually shells `restic stats`. The UI
reaches it FIRST, so flagging only the restore would have left
the collision reachable by the ordinary path.
RestoreSharesScratch - R-411's exact shape on the shares tier: unlockStale + resticStep,
a live web caller, and its sibling PlaceSharesRestore has always
taken the flag.
RestoreOffbox - no production caller today, but the same dangerous pattern.
Flagged rather than left for a future caller to inherit.
OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.
THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.
R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.
R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.
The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.
And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.
R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.
16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
|
||
|
|
e43b5ec07d |
v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged. THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does, nightly, on one app. IT DOES NOT prove a restore puts data back into a running app. That stays drill work and 07 section 8 matrix row 4 is NOT moved. THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So: (1) everything declared is present, AND (2) the manifest declares what the app is supposed to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test read verdict "pass". THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate the app's shape, and GetDockerVolumes describes the running app. Database half is DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are <project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose file's parent dir, which inside a unit is the literal string "compose". Measured on all eight real units on demo-hp the counts match exactly and the naming held every time - but "held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule that is invented. THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes that test read verdict "fail". IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect: --no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test fail on "unlock" appearing in the argv. IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch does not take it (R-408) while offbox_integrity.go states that invariant as universal. DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's app. ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app> would mean a nightly background job deleting the verification copy a CUSTOMER is looking at. It is also invisible to placement, so a proof copy can never be pushed into a live app. SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's behaviour is unchanged. NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is sound and the content is absent: different cause, different action. The hub half shipped FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted type is 400'd and vanishes. 33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller gates OK. Five red-proofs run and recorded in REPORT.md. A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242). |