v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged. THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does, nightly, on one app. IT DOES NOT prove a restore puts data back into a running app. That stays drill work and 07 section 8 matrix row 4 is NOT moved. THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So: (1) everything declared is present, AND (2) the manifest declares what the app is supposed to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test read verdict "pass". THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate the app's shape, and GetDockerVolumes describes the running app. Database half is DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are <project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose file's parent dir, which inside a unit is the literal string "compose". Measured on all eight real units on demo-hp the counts match exactly and the naming held every time - but "held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule that is invented. THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes that test read verdict "fail". IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect: --no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test fail on "unlock" appearing in the argv. IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch does not take it (R-408) while offbox_integrity.go states that invariant as universal. DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's app. ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app> would mean a nightly background job deleting the verification copy a CUSTOMER is looking at. It is also invisible to placement, so a proof copy can never be pushed into a live app. SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's behaviour is unchanged. NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is sound and the content is absent: different cause, different action. The hub half shipped FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted type is 400'd and vanishes. 33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller gates OK. Five red-proofs run and recorded in REPORT.md. A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
This commit is contained in:
@@ -55,6 +55,9 @@
|
||||
| `unitRestoreOutcomeMsg` + its three message constants (R-353, v0.226.0) | controller/internal/web/handlers.go | `(app string, res backup.UnitRestoreResult) string` | THE customer sentence for a completed LOCAL restore | Twin of `reconstituteOutcomeMsg`; copy its SHAPE (clauses earned by having done the thing, no filesystem path, base names only), never its text. The constants are named because `r353_unit_outcome_test.go` asserts them verbatim — a silent edit is how an honest message drifts back into a comforting one, which is the documented history of the sentence it replaces |
|
||||
| `offsiteNoSpaceMsgFmt` + `offsiteSizeUnknownMsg` (R-357, v0.226.0) | controller/internal/backup/offbox_restore.go | two consts | EVERY headroom refusal on the off-site restore surface | **All four gates share these** (prepare, scratch, place, and the destructive reconstitute). A customer meeting one wording on one path and a different one on another has to work out whether it is the same problem. **The reconstitute gate uses NO ×1.1 margin** — it copies a measured tree; `OffboxRestorePrepareFull`'s ×1.1 predicts a download. **Fail closed when either probe reads ≤ 0**: `free < need` with `need == 0` is FALSE, so an unmeasurable input sails through — a gate present and inert |
|
||||
| `Manager.OffboxFullScratchReady` + the scratch marker (R-358, v0.226.0) | controller/internal/backup/offbox_restore.go | `(stack) bool`; `.felhom-restore-complete.json` | THE gate for place-to-live and reconstitute | **It answers "did the run FINISH and was it FULL", not "are there files".** The old non-empty check passed a part-copy from a failed restic run, and the old doc comment ("PlaceOffsiteRestore re-validates per-path completeness") is what made it look adequate — that call stats top-level placements, not files. Marker written 0600 atomically AFTER restic returns nil; stale one cleared BEFORE it starts; both orders pinned by an AST test because `resticStep` is not a seam. Anything else — absent, unreadable, wrong schema, `full:false` — is NOT ready, with a WARN naming which. **Unit-only and full restores write the SAME directory**, so `full` is the only separator (R-396) |
|
||||
| `JudgeRestoredUnit` + `UnitProofResult` (R-87, v0.231.0) | controller/internal/backup/r403_hollow.go | `(unitDir string) UnitProofResult` | THE question "does this app's backup contain what THIS APP should have" | **The rule has TWO parts and part 1 alone is the trap:** "everything declared is present" passes a HOLLOW unit, because a hollow unit declares nothing — the exact shape it exists to catch. Part 2 is the expectation, and it comes from the unit's **own** captured compose file (`UnitComposeDir(unitDir)` + the compose filename), NEVER from live Docker (`GetDockerVolumes` describes the running app; the snapshot may predate it). Database half is `DBServiceNames`, the same discriminator `RestoreFromRecoveryUnit` uses, so this cannot disagree with the restore path about what an app is. **The volume half is an EXISTENCE check, not a name match** — `ResolveDockerVolumeNames` derives the project from the compose file's parent dir, which inside a unit is the literal string `compose`. **THREE outcomes:** pass / fail / **cannot judge**, and the third is never collapsed. Size is never consulted (`TestR87_SizeIsNeverConsulted`) |
|
||||
| `Manager.ProveOffboxUnit` + `ProofResult` (R-87, v0.231.0) | controller/internal/backup/offbox_proof.go | `(ctx) ProofResult` | THE nightly off-site content proof — one app, its newest snapshot, restored read-only and judged | **It NEVER writes to the repository and that is asserted on the ARGV:** `--no-lock`, no `unlockStale`, and `m.runner()` rather than `resticStep` so the `unlock --remove-all` escalation is unreachable. **It takes `acquireRunning` ITSELF** because `RestoreOffboxScratch` does not (R-408) — do not remove that. **Due-ness is per SNAPSHOT** (`ProvedSnapshots[stack]`), never a timestamp: a timestamp re-proves the same snapshot forever AND breaks the rotation. **Its scratch is a SEPARATE root** (`offsiteProofRootFor`, `backups/offsite-proof`) — sharing the customer's `offsite-restore` root would let a nightly job delete a copy the customer is looking at. A skip, a missing snapshot and a restore error reach NO verdict and do not advance due-ness |
|
||||
| `unitOnlyHeadroom` + `offboxScratchDirIn` (R-87, v0.231.0) | controller/internal/backup/offbox_restore.go | `(free int64) error`; `(stack, rootFor)` | The unit-only free-space gate and the drive-preference resolver, **shared** by the customer restore and the nightly proof | Extracted rather than forked so the two paths cannot drift on the parts that must not differ — the floor, the Hungarian refusal (`offsiteNoSpaceMsgFmt`), the network-storage refusal and the R-252 wording. `rootFor` is the ONLY difference between the two scratch paths. **Fail-closed on an unmeasurable probe:** `offboxFree` returns 0 when it cannot read, and `0 < floor` refuses — the inverse of the R-357 shape where a gate went inert |
|
||||
| `Manager.CheckOffboxIntegrity` + `IntegrityResult` + `IntegrityDue` (R-359, v0.227.0) | controller/internal/backup/offbox_integrity.go | `(ctx) IntegrityResult`; `(now) (due bool, last time.Time)` | THE off-site integrity check, and the only place `restic check` is run | **It TAKES `acquireRunning` and SKIPS rather than waits — never remove that guard.** `resticStep` escalates to `unlock --remove-all` on a lock error and is only safe because every caller holds the single-flight mutex; a check without it can strip a LIVE prune's lock. **THREE outcomes, not two:** `Skipped`, `Unreachable` and failed are different facts — a timeout is unreachable, NEVER damage, and only a failure notifies. A skip and an unreachable repo do **not** advance due-ness; a failure does. **DUE-NESS, NOT A WEEKDAY** (R-341). **`looksLikeRepositoryDamage` matches PHRASES, not words** — bare `pack `/`tree `/`snapshot ` appear in restic's ordinary progress output and made a healthy run look corrupt |
|
||||
| `runOffsiteIntegrityCheck` + `integrityFailedMsg` / `integrityOKMsg` (R-397, v0.227.0) | controller/cmd/controller/main.go | `(ctx, mgr, notifier, logger, force) backup.IntegrityResult` | THE one caller of the check — the scheduled job AND the debug button both go through it | ONE function so the hand-run cannot drift from the scheduled one; `force` skips due-ness and **nothing else**. Wired to `NotifyIntegrityOK` / `NotifyIntegrityFailed`, which existed with no caller since the notifier did (sixth built-but-never-wired instance). **`ok` is severity `info` and therefore mails NOBODY by design** — a weekly success e-mail is how alerts stop being read. The customer gets a sentence; restic's output goes to the log truncated (R-379) |
|
||||
| **`SetOffboxRunner` / `m.runner()` — the restic exec seam, and it has ALWAYS existed** | controller/internal/backup/offbox.go (~L52, ~L386) | `offboxRunner func(ctx, env, args...) ([]byte, error)` | Driving ANY restic-backed path under test | **R-398 claimed there was no such seam and was WRONG — do not re-file it.** `resticStep` is not itself overridable, but the layer it calls is, and tests have driven restic paths through it since the off-site tier shipped. **Use this rather than adding a `resticStepFn`:** replacing `resticStep` would hide its `unlock --remove-all` escalation from exactly the assertions that must see it (R-359's lock-safety tests assert `unlock` never appears in any argv) |
|
||||
|
||||
Reference in New Issue
Block a user