R-411/408/407, R-414, R-412a leg 1, R-410, R-406 CLOSED; determination + live evidence
gates / gates (push) Failing after 18s
gates / gates (push) Failing after 18s
Part 2.1's determination is the first artifact: the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded on state-only grounds - established from R-356's own commit 08eb1a6, whose tests say 'the prepared scratch still resolves ... only the DESTINATION moves'. So 07 section 6.3's rule applies and now has a FOURTH consumer, and the section says so. Live evidence: the collision rerun on demo-hp with the sampler positively controlled first (12 locks=1 across a real check, 4 locks=0 quiet), showing unlock --remove-all 0 times where the drill saw it twice; and the proof reaching verdict pass on demo-felhom - the box that could not run it at all - recorded where last_proof_result had been ABSENT every night. Capability map: the off-site proof row now records that the nightly firing IS proven (it ran unattended at 05:30 on demo-hp) and that a driveless box can now be proved. Register: six rows closed and compressed. OPEN 176 -> 170, CLOSED 152 -> 158.
This commit is contained in:
@@ -153,7 +153,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
|
||||
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore (unit-tested only). ~~Tier-2's own cross-drive copy of a secret-bearing unit~~ — **EXERCISED LIVE 2026-08-31 (controller v0.229.0):** docmost restored from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit` with the guest's `app.yaml` **moved aside** AND the primary unit moved aside, `secrets recovered=2/2` (`APP_SECRET`, `DB_PASSWORD`) taken from the MIRRORED unit's `compose/app.yaml`; the guest's `app.yaml` was rebuilt from it at 0600 and the app then read its own rows over TCP with its own credential. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/phase8-scenarioD-guest-appyaml-aside.log`. Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
||||
| **A restore SAYS what it returned, and refuses what it cannot do** — the four restore-surface truth defects from the 2026-08-21 drill | controller **v0.226.0** (R-353, R-357, R-358, R-360, R-396) | **PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only** | `audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`, controller `CHANGELOG.md` v0.226.0 + `REPORT.md`. Driven on `demo-hp` through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). **R-353:** the sentence read off the customer's own wizard page — `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). **R-360:** in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and **a planted canary file survived**. **R-358/R-396:** a `mode=unit` restore wrote `{"schema":1,…,"full":false}` at mode 0600 with no `.tmp` left, and the gate logged `scratch holds a UNIT-ONLY restore … place-to-live stays closed` | **WHAT IS AND IS NOT CLAIMED, split deliberately.** **R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE** — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`) whose central assertion is that `StopStack` was never called. **R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either**, and the reason is stated rather than glossed: no app on `demo-hp` still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. **This row is about the MESSAGE and the REFUSALS, not the recovery mechanism** — `07-backup-architecture.md` §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so |
|
||||
| **The box PROVES its own off-site copy still HOLDS something — a backup that is intact and EMPTY is caught without a person** | controller **v0.231.0** (R-87) + hub **v0.110.0** | **PROVEN-LIVE (2026-08-31) for the judgement, the alarm, the cleanup, the rotation, the read-only guarantee and the skip-if-busy hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r87-offsite-proof-2026-08-31/`. Driven on `demo-hp` through the endpoint the debug button invokes. **The failing case was produced and caught:** a hollow unit — compose declaring `opengist_data`, manifest declaring nothing — was pushed to the live store, and the proof returned `verdict:"fail"` with `volumes_expected_none_captured: opengist_data`, emitted **exactly one** `offsite_proof_empty` at severity `error`, accepted by the hub **HTTP 200** (which is itself the proof the allowlist entry landed — an unallowlisted type is 400'd and vanishes), and deleted its scratch. **The passing case was proven five times over** (bookstack, calibre-web, docmost, kimai, opengist), 2.2–4.0 s each, and the customer's own verification copies were untouched throughout. **The read-only guarantee was measured with a positively-controlled lock sampler** — it saw a lock appear and vanish across a real `restic check`, and **zero** across the proof, including a direct 6× test of the snapshot-lookup argv. **The skip-if-busy control fired live and unplanned:** a proof launched while the off-site backup run held the flag returned `skipped:true, duration_ms:0` with no verdict and no alarm | **⚠ WHAT A PASS MEANS, AND WHAT IT DOES NOT.** It means the newest off-site snapshot of ONE app contains what that app is supposed to have — judged from the unit's own captured compose, not from the live box. **It does NOT mean a restore puts data back into a running app**: the proof restores to a throwaway folder, looks, and deletes, and `07` §8 matrix row 4 is deliberately NOT moved. **It also does not vouch for the BYTES** — nothing available can: restic 0.14.0's `restore --verify` passed a byte-level corruption with size and mtime preserved (measured, 131 ms on a 213 MB tree), and the unit manifest hashes 4 918 B of a 213 231 242 B unit (R-409). **The nightly firing at 05:30 is IMPLEMENTED only** — the job is confirmed REGISTERED on `demo-hp` (`Daily job offsite-proof scheduled for 2026-09-01 05:30 CEST`), which is not the same claim, and the fleet is on 0.230.0 until a golden carries 0.231.0 |
|
||||
| **The box PROVES its own off-site copy still HOLDS something — a backup that is intact and EMPTY is caught without a person** | controller **v0.231.0** (R-87) + hub **v0.110.0** | **PROVEN-LIVE (2026-08-31) for the judgement, the alarm, the cleanup, the rotation, the read-only guarantee and the skip-if-busy hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r87-offsite-proof-2026-08-31/`. Driven on `demo-hp` through the endpoint the debug button invokes. **The failing case was produced and caught:** a hollow unit — compose declaring `opengist_data`, manifest declaring nothing — was pushed to the live store, and the proof returned `verdict:"fail"` with `volumes_expected_none_captured: opengist_data`, emitted **exactly one** `offsite_proof_empty` at severity `error`, accepted by the hub **HTTP 200** (which is itself the proof the allowlist entry landed — an unallowlisted type is 400'd and vanishes), and deleted its scratch. **The passing case was proven five times over** (bookstack, calibre-web, docmost, kimai, opengist), 2.2–4.0 s each, and the customer's own verification copies were untouched throughout. **The read-only guarantee was measured with a positively-controlled lock sampler** — it saw a lock appear and vanish across a real `restic check`, and **zero** across the proof, including a direct 6× test of the snapshot-lookup argv. **The skip-if-busy control fired live and unplanned:** a proof launched while the off-site backup run held the flag returned `skipped:true, duration_ms:0` with no verdict and no alarm | **⚠ WHAT A PASS MEANS, AND WHAT IT DOES NOT.** It means the newest off-site snapshot of ONE app contains what that app is supposed to have — judged from the unit's own captured compose, not from the live box. **It does NOT mean a restore puts data back into a running app**: the proof restores to a throwaway folder, looks, and deletes, and `07` §8 matrix row 4 is deliberately NOT moved. **It also does not vouch for the BYTES** — nothing available can: restic 0.14.0's `restore --verify` passed a byte-level corruption with size and mtime preserved (measured, 131 ms on a 213 MB tree), and the unit manifest hashes 4 918 B of a 213 231 242 B unit (R-409). **2026-09-01 (R-414, controller v0.232.0): it can now run on a box with NO registered data drive.** The first unattended firing, on `demo-felhom`, REFUSED — *"nowhere to restore to"* — because that box has `storage_paths: []` and the scratch resolver never consulted the system data path. A **unit-only** restore now falls back there (where a driveless app's unit already lives, `07` §7); a **full** restore still refuses, because the SSD is a state-only tier. And a proof that cannot start now records `cannot_run` instead of nothing, so `last_proof_result` is never ABSENT — absent already means *a controller too old to have the feature*. **PROVEN LIVE on `demo-felhom` 2026-09-01:** `verdict:"pass"` on `61e9cf30` in 2.117 s, recorded, and the scratch deleted. **The nightly firing IS now proven** — it ran unattended on `demo-hp` at 05:30 on 2026-09-01 (`bentopdf PASSED on 9d002b38 in 2.315s`), which this row previously listed as implemented-only. The job is confirmed REGISTERED on `demo-hp` (`Daily job offsite-proof scheduled for 2026-09-01 05:30 CEST`), which is not the same claim, and the fleet is on 0.230.0 until a golden carries 0.231.0 |
|
||||
| **The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable** | controller **v0.228.0** (R-359, R-397, R-399) | **PROVEN-LIVE (2026-08-30, re-proven at FULL DEPTH 2026-08-31) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r359-integrity-2026-08-30/`. Driven on `demo-hp` through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (**negative control first**), then one pack corrupted; the live store was checked read-only in **35.0 s**; and the notifier fired end to end — `Event pushed: backup_integrity_ok (info)`. The hazard control was observed live: a second check fired while the first held the single-writer flag returned `skipped:true, duration_ms:0` — **it never ran restic at all** | **⚠ WHAT AN `ok` MEANS — CHANGED 2026-08-31 (R-399, controller v0.228.0): the check now RE-READS THE DATA.** The default is `--read-data-subset=100%`, so an `ok` means every stored byte was downloaded and re-hashed, not merely that the catalogue hangs together. **The reason is measured, and it is why the default must not be turned back down to save four seconds:** a pack corrupted WITHOUT a size change made a structure check return `no errors were found`, exit 0, while every `--read-data*` form caught it. Cost curve on 134.3 MB: structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s — **and those do NOT extrapolate**, which is why v0.228.0 ships a slow-check WARN (R-401) rather than a rotation schedule. `off` returns a box to structure depth. **PROVEN-LIVE at the new depth 2026-08-31 on `demo-hp`**, endpoint-level, with the restic argv observed from the guest: default → `… check --read-data-subset=100%`, 38.7 s; `off` → `… check`, 34.7 s. **The weekly firing at the new depth is IMPLEMENTED only** — the job is confirmed REGISTERED on BOTH demo boxes (`Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST`), which is not the same claim. `demo-felhom` reached 0.228.0 by SELF-UPDATE on the 2026-08-31 floor raise and re-registered the job itself, so the depth change is on the fleet and not only on the box that was deployed to by hand. **This is a readability check and NOT a restore-test** — R-87 remains open and the two are routinely conflated because their register rows are adjacent |
|
||||
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
|
||||
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
|
||||
|
||||
Reference in New Issue
Block a user