v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged. THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does, nightly, on one app. IT DOES NOT prove a restore puts data back into a running app. That stays drill work and 07 section 8 matrix row 4 is NOT moved. THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So: (1) everything declared is present, AND (2) the manifest declares what the app is supposed to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test read verdict "pass". THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate the app's shape, and GetDockerVolumes describes the running app. Database half is DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are <project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose file's parent dir, which inside a unit is the literal string "compose". Measured on all eight real units on demo-hp the counts match exactly and the naming held every time - but "held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule that is invented. THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes that test read verdict "fail". IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect: --no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test fail on "unlock" appearing in the argv. IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch does not take it (R-408) while offbox_integrity.go states that invariant as universal. DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's app. ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app> would mean a nightly background job deleting the verification copy a CUSTOMER is looking at. It is also invisible to placement, so a proof copy can never be pushed into a live app. SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's behaviour is unchanged. NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is sound and the content is absent: different cause, different action. The hub half shipped FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted type is 400'd and vanishes. 33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller gates OK. Five red-proofs run and recorded in REPORT.md. A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
This commit is contained in:
@@ -1089,6 +1089,52 @@ backups/primary/<app>/
|
||||
maradtak."). **Every one is a claim about the BACKUP, never about the app** — see CONTEXT.md's ruling
|
||||
and 07-backup-architecture §6.3.
|
||||
|
||||
### Off-site content proof (v0.231.0, R-87)
|
||||
|
||||
**What it is:** every night the box restores ONE app's newest off-site snapshot into a throwaway
|
||||
folder, asks whether that backup still contains the app's actual data, records which snapshot it
|
||||
proved, and deletes the copy.
|
||||
|
||||
**The question it answers, and why the integrity check beside it cannot.** `restic check` proves the
|
||||
stored bytes are the bytes we stored. **It cannot tell us we stored the wrong thing.** A hollow
|
||||
recovery unit — no database dump, no volume tar — backs up cleanly, checks cleanly at 100 % depth,
|
||||
restores cleanly, and gives the customer nothing back. Measured on `demo-hp` 2026-08-31 (R-403):
|
||||
120 082 104 B became 7 036 B in one nightly run, recorded as a success.
|
||||
|
||||
> **WHAT IT DOES NOT PROVE, said plainly because the green tick invites the other reading:** it does
|
||||
> **not** prove a restore puts data back into a running app. It restores to a throwaway folder, looks,
|
||||
> and deletes. It never touches a live app. Putting data back is drill work, and
|
||||
> `07-backup-architecture.md` §8 matrix row 4 does not move on the strength of this job.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| job | `offsite-proof`, `sched.Daily` at **05:30** |
|
||||
| cadence | **per SNAPSHOT, not per clock** (R-86's model) — an app is due when its newest snapshot's ID differs from the ID last proved for it. One app per run; eight apps are covered in eight nights, and a NEW backup makes an app due again immediately |
|
||||
| acceptance rule | **two parts, and both are needed.** (1) everything the manifest declares is present in the restored unit, AND (2) **the manifest declares what the app is supposed to have.** Part 1 alone passes a hollow unit, which is the shape this exists to catch |
|
||||
| where the expectation comes from | the unit's **own** `compose/docker-compose.yml`, never the live box — the snapshot may predate the app's current shape. Database: `DBServiceNames` (the same discriminator the restore path uses). Volumes: `ParseComposeNamedVolumes`, as an **existence** check, not a name match |
|
||||
| outcomes | **three:** pass, fail (readable and empty), and **cannot judge**. An app that legitimately has no database and no named volumes **passes** |
|
||||
| repository writes | **none.** `--no-lock`, no `unlockStale`, and the exec seam rather than `resticStep`, so the `unlock --remove-all` escalation is unreachable. Asserted on the argv as a non-effect |
|
||||
| guard | takes the single-writer flag itself and **SKIPS rather than waits** (`RestoreOffboxScratch` does not take it — R-408) |
|
||||
| scratch | `backups/offsite-proof/<app>` — a **separate root** from the customer's `backups/offsite-restore/`, so the nightly delete can never reach a copy the customer made, and a proof copy can never be offered for placement |
|
||||
| timeout | 10 min (`proofRestoreTimeout`) — ~150x the slowest single app measured |
|
||||
| cost, measured on demo-hp | **one app 2.3–4.0 s**, all eight back to back **25 s**, peak scratch = that app's logical size (213 MB largest). The weekly check beside it takes 40.3 s |
|
||||
| result | persisted on `settings.OffboxTarget` (`proved_snapshots`, `last_proof_*`) and published on `OffboxReportStatus`. `last_proof_result` absent = **NOT RECORDED** (a pre-0.231.0 controller), never "failed" |
|
||||
|
||||
**Notifications.** A failure emits **one** `offsite_proof_empty`, severity `error`, operator-only. A
|
||||
pass, a skip, a "cannot judge" and a restore error emit **nothing** — a nightly success mail is how
|
||||
people stop reading their alerts, and alarming on our own blind spot trains the operator to discount
|
||||
the one alarm that matters.
|
||||
|
||||
> **⚠ IT IS DELIBERATELY NOT `backup_integrity_failed`.** That type means **the store is damaged** and
|
||||
> carries a hub-side Hungarian template saying so. Here the store is sound and the CONTENT is absent —
|
||||
> a different cause and a different action. The message says the backup is *readable* and does *not*
|
||||
> contain the app's data, and explicitly that the store is not damaged.
|
||||
|
||||
> **`restic restore --verify` is NOT used as a correctness check and must not be.** Measured on
|
||||
> `demo-hp` 2026-08-31: a byte changed in place in a restored 160 MB tar, with size and mtime
|
||||
> preserved, **passed clean**; verify took 131 ms on a 213 MB tree, which cannot be hashing. It is a
|
||||
> size-and-mtime reconciliation.
|
||||
|
||||
### Off-site integrity check (v0.227.0, R-359/R-397)
|
||||
|
||||
**What it is:** a `restic check` against the off-site repository, run by the controller itself. Until
|
||||
|
||||
Reference in New Issue
Block a user