R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s

Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp.

LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level
through the exact route the debug button invokes):

- THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest
  declaring nothing - was pushed to the live store and the proof returned verdict "fail"
  with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE
  offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself
  the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes.
- THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden:
  stopping the app does NOT produce a failed dump leg, because the off-site run's own
  capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is
  therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no
  forget and no prune. State restored: the product's own run made a healthy snapshot the
  newest again and the proof then passed opengist.
- The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s
  each, matching the spike's measured band.
- The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear
  and vanish across a real restic check, and ZERO across the proof - including a direct 6x
  test of the snapshot-lookup argv, which settles that restic snapshots does not lock in
  0.14.0 either.
- Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the
  flag returned skipped:true duration_ms:0, no verdict, no alarm.
- The customer's own verification copies were untouched throughout, which is the safety
  property the separate proof root exists for.

ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at
19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the
proof; I did not establish what it was.

CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only -
the job is REGISTERED, which is not the same claim.

07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one
sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore
puts data back into a running app. Without that sentence the new green tick reads as
covering the drill.

REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 ->
152. No new rows minted. R-408 and R-409 stay open and are referenced by this work.

golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the
newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job.
A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses
--no-verify for that reason - bypass #8.
This commit is contained in:
2026-08-31 21:31:56 +02:00
parent 1aeaa30c28
commit 7ee25925f9
22 changed files with 289 additions and 35 deletions
@@ -1079,7 +1079,7 @@ does **not** hold as written. → **R-108**
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
| ~~R-399~~ | ~~The check reads the catalogue and never the data~~ | **CLOSED 2026-08-31, controller v0.228.0.** `monitoring.integrity.read_data_subset` now defaults to **`100%`**, so the weekly check downloads and re-hashes every stored byte. **The fact that made it necessary, and the sentence that should stop anyone turning it back down to save four seconds: the structure check PASSED a size-preserving pack corruption.** Measured on `demo-hp` 2026-08-30 — plain `restic check` reported `no errors were found` and exited 0 over a pack damaged without a size change; every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100%. `off` (any case) returns a box to structure depth; an empty value means *not configured*, therefore the default; a malformed value WARNs and falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs an operator WARN naming R-401 — **one data point, on one 134 MB store, so no rotation schedule, size threshold or bandwidth budget was invented from it.** Proven live at both depths 2026-08-31 with the restic argv observed from the guest |
| **R-87 (open — SPIKED 2026-08-31, RE-SCOPE PROPOSED) — AND IT IS NOT R-359** | The restic tier is never restore-TESTED | **SPIKE VERDICT, `audits/SPIKE-restic-restore-test-2026-08-31.md`:** build the NARROW version, not the row as written. **Measured:** a scratch restore of all 8 apps costs **25 s / ≤213 MB scratch**, LESS than the 40.3 s weekly check beside it; but restic 0.14.0's `--verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving corruption passed clean, red-proofed), `restic ls --json` carries no content hash, and the unit manifest hashes **4 918 B of a 213 231 242 B unit** — so **no reference for "correct" exists** (R-409). **Of the five drill-found restore defects R-353/354/356/358/403, an unattended scratch-restore would have caught ONE (R-356).** The value is elsewhere and the weekly check structurally cannot reach it: `check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured 2026-08-31). **Proposed re-scope, Viktor's call:** *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` via the existing `unitCarriesData`. Must use `--no-lock` and skip `unlockStale` (R-95's constraint is otherwise violated — R-407/R-408 record what the path writes today) and must take `acquireRunning`, which `RestoreOffboxScratch` does not. | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
| ~~**R-87**~~ **— SHIPPED 2026-08-31 (controller v0.231.0), RE-SCOPED. AND IT IS STILL NOT R-359** | ~~The restic tier is never restore-TESTED~~ **The box now PROVES its own off-site copy still HOLDS something.** **THE DISTINCTION, STATED ONCE AND PLAINLY BECAUSE THE GREEN TICK INVITES THE OTHER READING: this proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app.** The proof restores to a throwaway folder, judges it against the unit's own captured compose, and deletes it — it never touches a live app. Putting data back is drill work. **§8 matrix row 4 is deliberately NOT moved.** Nightly at 05:30, one app, due-ness per SNAPSHOT (R-86's model). Evidence `tests/r87-offsite-proof-2026-08-31/`; the reasoning is `audits/SPIKE-restic-restore-test-2026-08-31.md`. | **SPIKE VERDICT, `audits/SPIKE-restic-restore-test-2026-08-31.md`:** build the NARROW version, not the row as written. **Measured:** a scratch restore of all 8 apps costs **25 s / ≤213 MB scratch**, LESS than the 40.3 s weekly check beside it; but restic 0.14.0's `--verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving corruption passed clean, red-proofed), `restic ls --json` carries no content hash, and the unit manifest hashes **4 918 B of a 213 231 242 B unit** — so **no reference for "correct" exists** (R-409). **Of the five drill-found restore defects R-353/354/356/358/403, an unattended scratch-restore would have caught ONE (R-356).** The value is elsewhere and the weekly check structurally cannot reach it: `check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured 2026-08-31). **Proposed re-scope, Viktor's call:** *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` via the existing `unitCarriesData`. Must use `--no-lock` and skip `unlockStale` (R-95's constraint is otherwise violated — R-407/R-408 record what the path writes today) and must take `acquireRunning`, which `RestoreOffboxScratch` does not. | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
### 10.3 Divergences that are documented elsewhere and are not re-opened here