R-411/408/407, R-414, R-412a leg 1, R-410, R-406 CLOSED; determination + live evidence
gates / gates (push) Failing after 18s
gates / gates (push) Failing after 18s
Part 2.1's determination is the first artifact: the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded on state-only grounds - established from R-356's own commit 08eb1a6, whose tests say 'the prepared scratch still resolves ... only the DESTINATION moves'. So 07 section 6.3's rule applies and now has a FOURTH consumer, and the section says so. Live evidence: the collision rerun on demo-hp with the sampler positively controlled first (12 locks=1 across a real check, 4 locks=0 quiet), showing unlock --remove-all 0 times where the drill saw it twice; and the proof reaching verdict pass on demo-felhom - the box that could not run it at all - recorded where last_proof_result had been ABSENT every night. Capability map: the off-site proof row now records that the nightly firing IS proven (it ran unattended at 05:30 on demo-hp) and that a driveless box can now be proved. Register: six rows closed and compressed. OPEN 176 -> 170, CLOSED 152 -> 158.
This commit is contained in:
@@ -153,7 +153,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
|
||||
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore (unit-tested only). ~~Tier-2's own cross-drive copy of a secret-bearing unit~~ — **EXERCISED LIVE 2026-08-31 (controller v0.229.0):** docmost restored from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit` with the guest's `app.yaml` **moved aside** AND the primary unit moved aside, `secrets recovered=2/2` (`APP_SECRET`, `DB_PASSWORD`) taken from the MIRRORED unit's `compose/app.yaml`; the guest's `app.yaml` was rebuilt from it at 0600 and the app then read its own rows over TCP with its own credential. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/phase8-scenarioD-guest-appyaml-aside.log`. Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
||||
| **A restore SAYS what it returned, and refuses what it cannot do** — the four restore-surface truth defects from the 2026-08-21 drill | controller **v0.226.0** (R-353, R-357, R-358, R-360, R-396) | **PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only** | `audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`, controller `CHANGELOG.md` v0.226.0 + `REPORT.md`. Driven on `demo-hp` through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). **R-353:** the sentence read off the customer's own wizard page — `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). **R-360:** in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and **a planted canary file survived**. **R-358/R-396:** a `mode=unit` restore wrote `{"schema":1,…,"full":false}` at mode 0600 with no `.tmp` left, and the gate logged `scratch holds a UNIT-ONLY restore … place-to-live stays closed` | **WHAT IS AND IS NOT CLAIMED, split deliberately.** **R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE** — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`) whose central assertion is that `StopStack` was never called. **R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either**, and the reason is stated rather than glossed: no app on `demo-hp` still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. **This row is about the MESSAGE and the REFUSALS, not the recovery mechanism** — `07-backup-architecture.md` §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so |
|
||||
| **The box PROVES its own off-site copy still HOLDS something — a backup that is intact and EMPTY is caught without a person** | controller **v0.231.0** (R-87) + hub **v0.110.0** | **PROVEN-LIVE (2026-08-31) for the judgement, the alarm, the cleanup, the rotation, the read-only guarantee and the skip-if-busy hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r87-offsite-proof-2026-08-31/`. Driven on `demo-hp` through the endpoint the debug button invokes. **The failing case was produced and caught:** a hollow unit — compose declaring `opengist_data`, manifest declaring nothing — was pushed to the live store, and the proof returned `verdict:"fail"` with `volumes_expected_none_captured: opengist_data`, emitted **exactly one** `offsite_proof_empty` at severity `error`, accepted by the hub **HTTP 200** (which is itself the proof the allowlist entry landed — an unallowlisted type is 400'd and vanishes), and deleted its scratch. **The passing case was proven five times over** (bookstack, calibre-web, docmost, kimai, opengist), 2.2–4.0 s each, and the customer's own verification copies were untouched throughout. **The read-only guarantee was measured with a positively-controlled lock sampler** — it saw a lock appear and vanish across a real `restic check`, and **zero** across the proof, including a direct 6× test of the snapshot-lookup argv. **The skip-if-busy control fired live and unplanned:** a proof launched while the off-site backup run held the flag returned `skipped:true, duration_ms:0` with no verdict and no alarm | **⚠ WHAT A PASS MEANS, AND WHAT IT DOES NOT.** It means the newest off-site snapshot of ONE app contains what that app is supposed to have — judged from the unit's own captured compose, not from the live box. **It does NOT mean a restore puts data back into a running app**: the proof restores to a throwaway folder, looks, and deletes, and `07` §8 matrix row 4 is deliberately NOT moved. **It also does not vouch for the BYTES** — nothing available can: restic 0.14.0's `restore --verify` passed a byte-level corruption with size and mtime preserved (measured, 131 ms on a 213 MB tree), and the unit manifest hashes 4 918 B of a 213 231 242 B unit (R-409). **The nightly firing at 05:30 is IMPLEMENTED only** — the job is confirmed REGISTERED on `demo-hp` (`Daily job offsite-proof scheduled for 2026-09-01 05:30 CEST`), which is not the same claim, and the fleet is on 0.230.0 until a golden carries 0.231.0 |
|
||||
| **The box PROVES its own off-site copy still HOLDS something — a backup that is intact and EMPTY is caught without a person** | controller **v0.231.0** (R-87) + hub **v0.110.0** | **PROVEN-LIVE (2026-08-31) for the judgement, the alarm, the cleanup, the rotation, the read-only guarantee and the skip-if-busy hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r87-offsite-proof-2026-08-31/`. Driven on `demo-hp` through the endpoint the debug button invokes. **The failing case was produced and caught:** a hollow unit — compose declaring `opengist_data`, manifest declaring nothing — was pushed to the live store, and the proof returned `verdict:"fail"` with `volumes_expected_none_captured: opengist_data`, emitted **exactly one** `offsite_proof_empty` at severity `error`, accepted by the hub **HTTP 200** (which is itself the proof the allowlist entry landed — an unallowlisted type is 400'd and vanishes), and deleted its scratch. **The passing case was proven five times over** (bookstack, calibre-web, docmost, kimai, opengist), 2.2–4.0 s each, and the customer's own verification copies were untouched throughout. **The read-only guarantee was measured with a positively-controlled lock sampler** — it saw a lock appear and vanish across a real `restic check`, and **zero** across the proof, including a direct 6× test of the snapshot-lookup argv. **The skip-if-busy control fired live and unplanned:** a proof launched while the off-site backup run held the flag returned `skipped:true, duration_ms:0` with no verdict and no alarm | **⚠ WHAT A PASS MEANS, AND WHAT IT DOES NOT.** It means the newest off-site snapshot of ONE app contains what that app is supposed to have — judged from the unit's own captured compose, not from the live box. **It does NOT mean a restore puts data back into a running app**: the proof restores to a throwaway folder, looks, and deletes, and `07` §8 matrix row 4 is deliberately NOT moved. **It also does not vouch for the BYTES** — nothing available can: restic 0.14.0's `restore --verify` passed a byte-level corruption with size and mtime preserved (measured, 131 ms on a 213 MB tree), and the unit manifest hashes 4 918 B of a 213 231 242 B unit (R-409). **2026-09-01 (R-414, controller v0.232.0): it can now run on a box with NO registered data drive.** The first unattended firing, on `demo-felhom`, REFUSED — *"nowhere to restore to"* — because that box has `storage_paths: []` and the scratch resolver never consulted the system data path. A **unit-only** restore now falls back there (where a driveless app's unit already lives, `07` §7); a **full** restore still refuses, because the SSD is a state-only tier. And a proof that cannot start now records `cannot_run` instead of nothing, so `last_proof_result` is never ABSENT — absent already means *a controller too old to have the feature*. **PROVEN LIVE on `demo-felhom` 2026-09-01:** `verdict:"pass"` on `61e9cf30` in 2.117 s, recorded, and the scratch deleted. **The nightly firing IS now proven** — it ran unattended on `demo-hp` at 05:30 on 2026-09-01 (`bentopdf PASSED on 9d002b38 in 2.315s`), which this row previously listed as implemented-only. The job is confirmed REGISTERED on `demo-hp` (`Daily job offsite-proof scheduled for 2026-09-01 05:30 CEST`), which is not the same claim, and the fleet is on 0.230.0 until a golden carries 0.231.0 |
|
||||
| **The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable** | controller **v0.228.0** (R-359, R-397, R-399) | **PROVEN-LIVE (2026-08-30, re-proven at FULL DEPTH 2026-08-31) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r359-integrity-2026-08-30/`. Driven on `demo-hp` through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (**negative control first**), then one pack corrupted; the live store was checked read-only in **35.0 s**; and the notifier fired end to end — `Event pushed: backup_integrity_ok (info)`. The hazard control was observed live: a second check fired while the first held the single-writer flag returned `skipped:true, duration_ms:0` — **it never ran restic at all** | **⚠ WHAT AN `ok` MEANS — CHANGED 2026-08-31 (R-399, controller v0.228.0): the check now RE-READS THE DATA.** The default is `--read-data-subset=100%`, so an `ok` means every stored byte was downloaded and re-hashed, not merely that the catalogue hangs together. **The reason is measured, and it is why the default must not be turned back down to save four seconds:** a pack corrupted WITHOUT a size change made a structure check return `no errors were found`, exit 0, while every `--read-data*` form caught it. Cost curve on 134.3 MB: structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s — **and those do NOT extrapolate**, which is why v0.228.0 ships a slow-check WARN (R-401) rather than a rotation schedule. `off` returns a box to structure depth. **PROVEN-LIVE at the new depth 2026-08-31 on `demo-hp`**, endpoint-level, with the restic argv observed from the guest: default → `… check --read-data-subset=100%`, 38.7 s; `off` → `… check`, 34.7 s. **The weekly firing at the new depth is IMPLEMENTED only** — the job is confirmed REGISTERED on BOTH demo boxes (`Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST`), which is not the same claim. `demo-felhom` reached 0.228.0 by SELF-UPDATE on the 2026-08-31 floor raise and re-registered the job itself, so the depth change is on the fleet and not only on the box that was deployed to by hand. **This is a readability check and NOT a restore-test** — R-87 remains open and the two are routinely conflated because their register rows are adjacent |
|
||||
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
|
||||
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
|
||||
|
||||
@@ -438,6 +438,22 @@ destination.** The drive if the app declares one (`HDD_PATH`), the system data p
|
||||
`Manager.GetAppDrivePath`, one expression, used by `CaptureRecoveryUnit` and, since controller
|
||||
**v0.219.0**, by `ReconstituteFromOffsite` and `PlaceOffsiteRestore` too.
|
||||
|
||||
**[FACT] 2026-09-01 — the rule now has a FOURTH consumer and the "one expression" sentence is
|
||||
true (R-414, controller v0.232.0).** `offboxRestoreScratchDir` — where a restore is UNPACKED, as
|
||||
distinct from where it LANDS — never consulted `systemDataPath`, so on a box with no registered
|
||||
storage path it refused, and the nightly proof could not run at all. **It was not excluded on
|
||||
purpose:** R-356's own commit (`08eb1a6`) records in its tests that *"the prepared scratch still
|
||||
resolves to the registered storage path … only the DESTINATION moves"* — it was out of scope, and
|
||||
every fixture assumed a registered path exists. The one comment about a `systemDataPath` fallback
|
||||
belonged to `PlaceOffsiteRestore`, concerned bulk **userdata**, and R-356 deliberately overruled it.
|
||||
|
||||
**But the fallback is SCOPED, and the scoping is the point.** The two callers ask different
|
||||
questions, and one predicate answering both is R-356's own defect: a **unit-only** restore (the
|
||||
R-87 proof) may fall back to the system data path, because §7 records as `[FACT]` that a driveless
|
||||
app's unit already lives there indefinitely and that the same-device placement is *"intended, not
|
||||
a defect"*; a **full** restore keeps the R-252 refusal, because it pulls the app's bulk userdata
|
||||
onto what §2.2 makes a **state-only** tier.
|
||||
|
||||
The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351) applies
|
||||
to apps that **have a drive to get wrong**. It used to be reached by testing `HDD_PATH == ""`, which
|
||||
also answered "is this app installed?" — one predicate for two questions. Measured in the catalogue at
|
||||
|
||||
@@ -0,0 +1,30 @@
|
||||
=== the collision, three overlap offsets: full restore FIRST, then the check into it ===
|
||||
--- restore, then check +1s ---
|
||||
t_restore=08:24:09
|
||||
[http=302 wall=0.010349s]
|
||||
t_check=08:24:10
|
||||
"duration_ms":0 "ok":false "skip_reason":"a backup or restore is already running" "skipped":true "ok":true
|
||||
--- restore, then check +2s ---
|
||||
t_restore=08:24:22
|
||||
[http=302 wall=0.010791s]
|
||||
t_check=08:24:24
|
||||
"duration_ms":39978 "ok":true "skip_reason":"" "skipped":false "ok":true
|
||||
--- restore, then check +3s ---
|
||||
t_restore=08:25:19
|
||||
[http=302 wall=0.011009s]
|
||||
t_check=08:25:22
|
||||
"duration_ms":0 "ok":false "skip_reason":"a backup or restore is already running" "skipped":true "ok":true
|
||||
|
||||
=== THE NON-EFFECT: unlock --remove-all must appear in NO argv ===
|
||||
'unlock --remove-all' in the sampler: 0
|
||||
any 'unlock' at all: 0
|
||||
'cleared a stale exclusive lock' in the log: 0
|
||||
|
||||
=== and the check SKIPPED rather than colliding ===
|
||||
integrity: check PASSED in 1m4s (structure, index, and 100% of the pack data re-read)
|
||||
integrity: skipped — a backup or restore is already running; due-ness is NOT advanced, so this retries on the next run
|
||||
off-box restore refused for kimai (mode=full): another backup/restore op is already running
|
||||
off-box restore kimai completed (full=true, async)
|
||||
integrity: check PASSED in 40s (structure, index, and 100% of the pack data re-read)
|
||||
integrity: skipped — a backup or restore is already running; due-ness is NOT advanced, so this retries on the next run
|
||||
off-box restore kimai completed (full=true, async)
|
||||
@@ -0,0 +1,7 @@
|
||||
=== demo-felhom BEFORE: still zero registered storage paths? ===
|
||||
registered storage paths: 0
|
||||
last_proof_result: '<ABSENT>'
|
||||
|
||||
=== deploy 0.232.0 ===
|
||||
deployed
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.232.0 Up 26 seconds (healthy)
|
||||
@@ -0,0 +1,19 @@
|
||||
=== THE PROOF ON A BOX WITH ZERO REGISTERED DRIVES ===
|
||||
{"error":"Nem bekötött","ok":false}
|
||||
|
||||
[http=501]
|
||||
|
||||
=== what it recorded (last night this was ABSENT) ===
|
||||
last_proof_result '<ABSENT>'
|
||||
last_proof_stack '<ABSENT>'
|
||||
last_proof_snapshot '<ABSENT>'
|
||||
last_proof_reason '<ABSENT>'
|
||||
last_proof_run '<ABSENT>'
|
||||
proved_snapshots '<ABSENT>'
|
||||
|
||||
=== the log line (last night: a WARN and nothing else) ===
|
||||
Daily job offsite-proof scheduled for 2026-09-02 05:30 CEST
|
||||
|
||||
=== where did the scratch resolve to, and is it gone? ===
|
||||
offsite-proof root: ls: cannot access '/mnt/sys_drive/felhom-data/backups/offsite-proof': No such file or directory
|
||||
(empty or absent = the copy was deleted, as it must be)
|
||||
@@ -0,0 +1,30 @@
|
||||
=== current logging level (to be restored) ===
|
||||
logging:
|
||||
file: ""
|
||||
level: info
|
||||
backup taken
|
||||
logging:
|
||||
file: ""
|
||||
level: debug
|
||||
=== THE PROOF ON A BOX WITH ZERO REGISTERED DRIVES ===
|
||||
{"data":{"duration_ms":2452,"missing":null,"no_snapshot":false,"reason":"","skip_reason":"","skipped":false,"snapshot":"851e6dce","stack":"opengist","verdict":"pass"},"message":"A mentés tartalmazza az alkalmazás adatait","ok":true}
|
||||
|
||||
[http=200]
|
||||
|
||||
=== what it recorded (last night this was ABSENT) ===
|
||||
last_proof_result 'pass'
|
||||
last_proof_stack 'opengist'
|
||||
last_proof_snapshot '851e6dce'
|
||||
last_proof_reason '<ABSENT>'
|
||||
last_proof_run '2026-09-01T08:28:00Z'
|
||||
proved_snapshots {'opengist': '851e6dce'}
|
||||
|
||||
=== the log line (last night: a WARN and nothing else) ===
|
||||
auth: valid session for POST /api/debug/backup/offsite-proof
|
||||
proof: refusing to remove /mnt/sys_drive/felhom-data/backups/offsite-proof/opengist — it is not inside a proof root
|
||||
proof: opengist PASSED on snapshot 851e6dce in 2.452s — the backup holds what this app should have
|
||||
proof: refusing to remove /mnt/sys_drive/felhom-data/backups/offsite-proof/opengist — it is not inside a proof root
|
||||
|
||||
=== where did the scratch resolve to, and is it gone? ===
|
||||
offsite-proof root: opengist
|
||||
(empty or absent = the copy was deleted, as it must be)
|
||||
@@ -0,0 +1,19 @@
|
||||
=== a fresh off-site backup makes the app due again (the real path, no hand-set state) ===
|
||||
offbox/run http=302
|
||||
|
||||
=== THE PROOF, on a box with zero registered drives ===
|
||||
{"data":{"duration_ms":2116,"missing":null,"no_snapshot":false,"reason":"","skip_reason":"","skipped":false,"snapshot":"61e9cf30","stack":"opengist","verdict":"pass"},"message":"A mentés tartalmazza az alkalmazás adatait","ok":true}
|
||||
|
||||
[http=200]
|
||||
|
||||
=== the verdict recorded ===
|
||||
last_proof_result 'pass'
|
||||
last_proof_stack 'opengist'
|
||||
last_proof_snapshot '61e9cf30'
|
||||
last_proof_run '2026-09-01T08:32:20Z'
|
||||
|
||||
=== AND THE SCRATCH MUST BE GONE (the leak live validation caught) ===
|
||||
(no LEFTOVER lines above = the copy was deleted)
|
||||
daily job offsite-proof: next run at 2026-09-02 05:30:00 CEST (waiting 18h59m40s)
|
||||
proof: nothing due — every deployed app's newest off-site snapshot has already been proved, or none has one yet
|
||||
proof: opengist PASSED on snapshot 61e9cf30 in 2.117s — the backup holds what this app should have
|
||||
@@ -0,0 +1,6 @@
|
||||
config restored from backup
|
||||
logging:
|
||||
file: ""
|
||||
level: info
|
||||
observer guest cleaned: [0]
|
||||
observer host cleaned: [0]
|
||||
@@ -0,0 +1,13 @@
|
||||
=== VALIDATION 3: an ordinary customer restore on demo-hp, unchanged ===
|
||||
--- unit restore ---
|
||||
[http=302 wall=0.010338s]
|
||||
off-box restore docmost completed (full=false, async)
|
||||
--- the scratch it produced ---
|
||||
124208399 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost
|
||||
marker: {"schema":1,"snapshot_id":"b2e059ee","full":false,"finished_at":"2026-09-01T08:33:33Z"}
|
||||
--- and the full two-step flow: prepare (which now takes the flag) then confirm ---
|
||||
[http=302 wall=5.495803s]
|
||||
off-box full-restore prepared for privatebin (size 2.0 MB) — awaiting the customer's confirm; no restore has started
|
||||
|
||||
--- R-412a live: does a hollow push now say so? (no hollow unit here, so this is the CONTROL) ---
|
||||
hollow-push warnings (0 expected, all units sound): 0
|
||||
@@ -238,4 +238,10 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-87** | **The restic tier was never restore-tested — RE-SCOPED by its own spike to "prove the off-site snapshot still CONTAINS a recoverable unit".** Shipped in controller **v0.231.0** + hub **v0.110.0**. Evidence: `tests/r87-offsite-proof-2026-08-31/`; reasoning: `audits/SPIKE-restic-restore-test-2026-08-31.md`. **Reasoning kept:** *The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing.* *The acceptance rule has TWO parts and the obvious one is a trap — "everything declared is present" passes a hollow unit, which is the shape it exists to catch.* *The expectation comes from INSIDE the unit, never the live box: the snapshot may predate the app's shape.* *The volume half is an EXISTENCE check and not a name match — the naming held on all eight real units, but "held on eight" is not "derivable" (R-355), and half a rule that is true beats a whole rule that is invented.* *THREE outcomes: pass, fail, and cannot-judge — collapsing the third hides a gap in one direction and alarms on our own blind spot in the other.* *It proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app — §8 matrix row 4 was deliberately NOT moved.* *The proof's scratch is a SEPARATE root because the job deletes on every path, and sharing the customer's root would mean a nightly job deleting a copy the customer is looking at.* | **CLOSED 2026-08-31 — SHIPPED + PROVEN-LIVE** (controller v0.231.0, hub v0.110.0) | full text: `git show 303129e:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-406** | **Two unrelated findings shared the identifier R-133.** Resolved 2026-09-01 by renumbering the hub-uniqueness finding to **R-415**. **Reasoning kept:** *citations were MEASURED before choosing — 3 for hub-uniqueness, 5 for the plaintext break-glass credential — and the FEWER-cited one moved*; *this is the opposite of the task's literal instruction, whose stated ground ("the older number has the longer reference trail") the measurement contradicts; the principle was followed and the letter was not*; *the within-register duplicate rule was deliberately NOT added in the commit that removed its only subject — a guard whose red-proof can only be a planted fixture is not this project's standard (R-416)*. | **CLOSED — RENUMBERED** (2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-410** | **`golden_currency_gate.py` was satisfied by a DIRECTORY NAME — `mkdir` turned it green with no bake behind it.** Shipped in `felhom.eu`, 2026-09-01. **Reasoning kept:** *a directory name is a label; `GOLDEN_SHA256=<64 hex>` is a fact only a completed publish produces*; *the self-test ships a POSITIVE CONTROL, without which "it fails on an empty directory" would be satisfied by a gate that fails on everything*; *directories that look right and hold nothing are printed by name rather than silently ignored, so a half-finished bake is visible*. | **CLOSED — SHIPPED** (`felhom.eu`, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-414** | **The nightly off-site proof was INERT on a box with no registered data drive, every night, with only a WARN.** Shipped in controller **v0.232.0**. Evidence: `audits/R411-R414-2026-09-01/`, determination in `00-part2.1-determination.md`. **Reasoning kept:** *the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded — its own tests say "the scratch still resolves … only the DESTINATION moves"*; *the fallback is SCOPED because the two callers ask different questions, and one predicate answering both is the R-356 defect itself* — unit-only may fall back (§7: a driveless app's unit already lives on the system data path, *"intended, not a defect"*), a full restore may not (§2.2: state-only tier); *absence on `last_proof_result` already means "controller too old", so a second meaning on one field is the `StatsKnown` trap one level up*; *a `cannot_run` is recorded but does NOT advance per-snapshot due-ness, or the app would never be retried once a drive is registered*. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-407** | **`restic check` DOES write a lock file, and the comment above it said it never writes.** Corrected in controller **v0.232.0**. **Reasoning kept:** *corrected in place, not deleted — R-360's rule is that a comment claiming a guard is why nobody looks for the missing one*; *the same paragraph now carries the fact that `restic stats` also takes a lock*. | **CLOSED — CORRECTED IN PLACE** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-408** | **`RestoreOffboxScratch` took no single-writer flag while a comment asserted every off-site operation did.** Shipped in controller **v0.232.0**. **Reasoning kept:** *the real deliverable is the WALK, not the acquire* — the sentence was false for months and nothing checked it, the ninth instance of this project's most-repeated class; *it is an AST pass and not `strings.Contains`, because a commented-out call still contains the string*; *adding a line to `offsiteExempt` is a deliberate act and belongs in the commit that adds it*; *the R-87 proof's exemption is kept HONEST by a second test that fails if that path ever gains `unlockStale`, routes through `resticStep`, or loses `--no-lock`*. | **CLOSED — SHIPPED** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-411** | **A background job deleted the lock of a live customer restore and logged it as a crash that did not happen.** Shipped in controller **v0.232.0**. Evidence: `audits/DRILL-soak-2026-08-31/phase1-lock-collision/`, `audits/R411-R414-2026-09-01/`. **Reasoning kept:** *`restic stats` TAKES a repository lock* — the fact nobody had, and the one that made the chain reachable; *`restic check` takes one too, `restic snapshots` and `restic list` do not*; *the fix was wider than the row — FOUR entry points were unflagged, three of them found by R-408's walk rather than by the report*; *the escalation in `resticStep` was NOT removed — real stale locks exist and it clears them; the defect was that a sibling could be live*. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -589,15 +589,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||
| **R-404** | **DECISION FOR VIKTOR — should a documentation-only push be subject to the golden-currency gate?** `git push --no-verify` has now been used **eight times** (the seventh and eighth on 2026-08-31: the R-87 spike's records-only Part 1 commit `6e550ae`, and R-87's own closing docs commit `baf52f3` — the latter red on a REAL new debt, a released v0.231.0 with no golden, which is the gate working correctly on a docs-only push), each with a recorded reason, because `repo_gates.py --fast` runs `golden_currency_gate.py` on every push to `felhom.eu` including pushes that touch only `documentation/`. **A guard that is correctly bypassed six times is training everyone to bypass it**, and the seventh bypass will be faster to reach for than the sixth | **OPEN — a DECISION, not work. Filed 2026-08-31, deliberately NOT acted on** | — | **FOR narrowing it:** a docs-only push cannot be the push that finishes a release, so scoping the gate to pushes that touch `controller/`, `agent/` or `hub/` code is arguably not a weakening at all — it would fire on exactly the pushes that can create the gap and on no others. It would also end the habit, which is the real cost being paid now. **AGAINST narrowing it:** the gate was earned by a real recurrence — v0.206.0 shipped while the vouched golden carried 0.205.0, and v0.204.0/v0.205.0 before it — and every narrowing of a guard risks the thing it was built for coming back. The gate is also deliberately `--fast` so that BOTH the pre-push hook and CI run it (R-29's census failure); a scoped version must stay in both or it runs in neither. **IF VIKTOR DOES NOTHING:** the bypass stays routine and the count keeps rising; nothing breaks, and the guard quietly stops being one. **Owner: Viktor decides, CC builds. This task did NOT change the gate.** | Viktor |
|
||||
| **R-405** | **R-87 sat in `CLOSED-ITEMS.md` for nine days while it was still open, and the register's own ranking paragraph ranked it fourth pointing at nothing.** Established from history, not inferred: it was moved by the 2026-08-22 compression sweep, commit `ef6ac6f` (*One register, enforced by a gate; closed work compressed into siblings, R-376..R-378*) — the same commit and the same defect class R-378 records. **R-378 caught six — R-123, R-190, R-214, R-264, R-295, R-352 — and missed a seventh.** R-87 escaped because its state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY` and the word `closed` later in the same cell describes a **different** row. **Count reproduced independently 2026-08-31, and the predicate decides the answer:** matching an open word anywhere in the state cell convicts **three** of 151 rows (R-87, plus R-224 and R-260, both genuinely closed with the words "open"/"OPEN" inside long prose verdicts); matching the **leading verdict** convicts exactly **one**, R-87; matching the whole row convicts **144**. **Fixed this session:** the row is back in `OPEN-ITEMS.md` verbatim from `ef6ac6f^`, next to R-95 where it sat before, and `scripts/closed_register_gate.py` is the 12th gate. Red-proofed both rules and negative-controlled against the pushed pre-fix files, where it convicts R-87 by name. **`R-398` was ALSO in both registers** — a deliberate cross-reference stub — and is now prose beneath the table rather than a row, because a row in both files is what rule 2 convicts on. | **CLOSED 2026-08-31 — corrected + gated in the same session** | R-378 | Nothing further. The gate's four residual holes are named in its docstring; hole 4 is R-406. | CC |
|
||||
| **R-406** | **Two unrelated findings in `OPEN-ITEMS.md` share the identifier R-133.** `OPEN-ITEMS.md:267` is *the hub enforces uniqueness on `customer_id` only* (`domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK); `OPEN-ITEMS.md:273` is *the vaulted break-glass console credential is PLAINTEXT AT REST*. Different subjects, different owners, one number. Found 2026-08-31 while measuring duplicate ids for R-405's gate — **this is the only such collision in either register** (measured: no id appears twice in `CLOSED-ITEMS.md`, and R-88a/R-88b and R-209/R-209a are distinct suffixed ids, not duplicates). **Why the gate does NOT check for it:** a within-register duplicate rule would fail on this pre-existing row, and a registered-but-failing gate refuses every push. **Renumbering is not obviously safe** — `R-133` is cited elsewhere and a blind renumber breaks whichever citation meant the other one. | **CLOSED 2026-09-01 — renumbered** | R-405 | **DONE.** Citations measured before choosing, which is what the row asked for: the hub-uniqueness finding had **3** references (all three inside `audits/RECON-subdomain-onboarding-2026-07-31.md`); the plaintext-break-glass finding had **5** (`CONTEXT.md:1544`, `runbooks/break-glass.md:113`, `hub/CHANGELOG.md:1448`, `00-capability-map.md:180`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md:478`). **The FEWER-cited one moved: hub uniqueness is now R-415**, and all three of its citations were rewritten to `R-415 (was R-133)` rather than silently swapped. **This is the OPPOSITE of the task's literal instruction** — it said renumber the second row, on the stated ground that *"the older number has the longer reference trail"*. Measured, that ground points the other way: the second row is the one with the longer trail. The principle was followed and the letter was not, which is recorded here rather than left as an unexplained diff. **The within-register duplicate rule was NOT added to `closed_register_gate.py`** — with the collision gone there is nothing to fail on, but adding a rule in the same commit that removes its only subject means shipping a gate whose red-proof cannot be run against real data; filed as R-416. | CC || CC |
|
||||
| **R-407** | **`restic check` DOES write a lock file to the repository, and the comment above it says it never writes.** `offbox_integrity.go:255` reads *"CheckOffboxIntegrity runs one off-site integrity check. It NEVER writes to the repository: `check` is a read verb, and nothing here prunes, forgets, unlocks or backs up."* The three named verbs are correct; the headline is not. **OBSERVED 2026-08-31 on demo-hp**, with a lock sampler that was positively controlled before it was believed: across the product's own integrity run (13:41:28→13:42:11) the repository went `locks=0` → `locks=1 id=81fd4d4200d848466e18cb7a9d8e0d43…` for nine consecutive samples → `locks=0`. The same instrument saw **zero** locks across two restores, so it is not reporting a constant. **Why it matters and why it is LOW rather than ignorable:** R-95's whole constraint is phrased as "must never be able to write to the repo", and a comment stating a guarantee the code does not provide is this project's most-repeated failure — nine instances. The lock itself is correct behaviour and there is no defect in the check; **the defect is the sentence.** | **OPEN — LOW (a comment, not behaviour)** | R-87, R-359 | Correct the sentence in place — say the check takes a repository LOCK and writes nothing else — and pin it with a test, or pass `--no-lock` and make the sentence true. Do NOT delete the sentence: R-360's rule is that a doc comment claiming a guard is why nobody looks for the missing guard. Evidence: `audits/evidence-spike-restic-restore-2026-08-31/16-q6-locks-full.txt`. | CC |
|
||||
| **R-408** | **`RestoreOffboxScratch` takes NO single-writer flag, and the file that depends on that invariant states it as universal.** `offbox_integrity.go:28` reads *"Every off-site operation takes `acquireRunning` for exactly that reason"* — the reason being that `resticStep` escalates to `unlock --remove-all` and is safe only because the in-process mutex proves no sibling operation is live. **Measured 2026-08-31:** `grep -rn 'acquireRunning()'` finds nine non-test callers and `RestoreOffboxScratch` (`offbox_restore.go:234`) is not among them. `restore_wizard.go:174` records the same fact independently — *"`RestoreOffboxScratch` never acquires it at all"* — and the UI works around it with a separate display flag (`opRunning`), so the gap is known at the web layer and unknown at the one that reasons about repository safety. **The exposure today is small and the exposure tomorrow is not:** the web handler's `restoreOpBlocked()` fences the only caller that exists, and restic's restore takes no lock (R-407's sampler), so nothing currently collides. **An unattended restore-test — R-87 — would be the first caller with no web handler in front of it.** | **OPEN — MEDIUM** | R-87, R-359 | Decide ONE way: either `RestoreOffboxScratch` takes `acquireRunning` (and every existing caller is re-checked for the double-acquire refusal `tier2_restore.go:183` warns about), or `offbox_integrity.go:28`'s sentence is corrected to name the exception. Whichever is chosen, **pin it with a test** — this is a comment asserting an invariant with nothing holding it. Do this BEFORE R-87 ships anything. | CC |
|
||||
| **R-409** | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC |
|
||||
| **R-410** | **`golden_currency_gate.py` is satisfied by a DIRECTORY NAME — `mkdir documentation/tests/golden-<VER>-<DATE>` turns it green with no bake behind it.** `EVIDENCE_RE = ^golden-(\d+)\.(\d+)\.(\d+)-\d{4}-\d{2}-\d{2}$` matched against `os.listdir` (`scripts/golden_currency_gate.py:89,123`) — it never opens the directory, never looks for a log, never checks a sha, and never asks Gitea whether the package exists. Noticed 2026-08-31 while the 0.230.0 bake turned it from red to green: **I created the directory before the bake finished, and the gate would have passed at that moment.** **What the gate DOES say honestly:** its own success line already reads *"this checks the BAKE, not the vouch"* — so the vouch hole is declared. **This one is not:** nothing tells a reader the bake check is a filename check. **Class:** an instrument that cannot distinguish the thing from a label for the thing — the same shape as R-378's whole-field match and R-233's un-matchable grep, aimed this time at the release gate. **Exposure today is low** because the runbook produces a real evidence directory and two bakes in a row have; it is the NEXT hurried session that pays. | **OPEN — LOW** | R-242 | Make it read something the bake alone can produce: the `GOLDEN_SHA256=` line in the directory's `bake.log`, or a HEAD against the Gitea package URL for that version. Prefer the log — it keeps the gate offline and `--fast`. Ship a red-proof: an empty `golden-9.9.9-2026-01-01/` directory must FAIL. | CC |
|
||||
| **R-411** | **A background job DELETES the lock of a live customer restore, and logs it as a crash that did not happen.** MEASURED on demo-hp 2026-08-31 during the overnight soak, through the product's own endpoints — this is R-408's consequence, which until tonight had only been reasoned about. **The chain, every step observed:** (1) a customer full-restore runs `OffboxRestorePrepareFull` → `restic stats`, and **`restic stats` TAKES A REPOSITORY LOCK** (clean-room test: nothing else running, 4x stats, sampler reads `locks=1`); (2) the restore holds `opRunning` but **NOT `acquireRunning`** (R-408), so the integrity check is not blocked and runs concurrently; (3) the check meets that lock, and `resticStep` escalates to **`unlock --remove-all`** — caught by the argv sampler at **20:50:51 with `restore 3c11059b --target …` and `unlock --remove-all` in the SAME sample**; (4) the log says *"cleared a stale exclusive lock left by a previous crash (single-writer repo)"* — **there was no crash**, and `resticStep` cannot know there was, because it fires on ANY `repository is already locked`. **THE CUSTOMER-FACING CONSEQUENCE WAS CONTAINED, and that is R-359's guard working:** the check returned `ok:false` in 6.7 s and was classified **Unreachable, NOT damage** — *"the check could not run to a verdict (other) — NOT reported as damage"* — so no `backup_integrity_failed` and no customer mail. Due-ness was not advanced either, so it retries. **What is NOT contained:** a live operation's lock is deleted by a background job; the single-writer premise `resticStep`'s own comment rests on is false in this pairing; and that night's integrity check silently did not verify the store, with only a WARN. **The opposite direction is FENCED and was measured too:** five restores fired into a running check at 5/15/25/35/40 s offsets were ALL refused by `restoreOpBlocked` (`offbox_handlers.go:358`), zero restic invoked — so the hazard is reachable only restore-FIRST. | **OPEN — MEDIUM** | R-408, R-407, R-359 | Decide ONE way, and R-408 is the same decision: either `RestoreOffboxScratch` (and the full-restore preparation) takes `acquireRunning`, or `resticStep`'s escalation stops claiming a crash it cannot verify and refuses instead of removing. **Pin whichever is chosen with a test that reproduces this pairing** — a unit test on `resticStep` alone cannot see it. Evidence: `audits/DRILL-soak-2026-08-31/phase1-lock-collision/`. | CC |
|
||||
| **R-412** | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** **CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it.** As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". **That is wrong, and measuring it overnight is what showed it:** the off-site run has its OWN pre-push dump leg — *"Stopping calibre-web for safe volume dump"*, *"Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB"* — so a unit that is hollow when a run starts is **REPAIRED before it is pushed**. Proven twice: `opengist` (2026-08-31 21:0x) and `calibre-web` (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (`6fee3b5a`) holds the volume tar and all 17 userdata files. **WHAT REMAINS REAL, and it is narrower:** the one hollow snapshot that DID reach the store (`35ba9fe7`, opengist) was created when the unit was destroyed **inside** a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged *"backed up opengist (… 0 mandatory path(s))"*, **a success line over a backup holding none of the app's data**. That race is real, it was observed, and the success wording is wrong either way. **The R-403 mirror guard holds throughout** — proven live: *"unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one"*, secondary byte-identical. | **OPEN — LOW (was HIGH; the correction is the reason)** | R-403, R-87, R-413 | Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. **Do NOT guard the capture** (08 §8.2). Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/` and `phase5-mutated-cycle/09-what-reached-the-store.txt`. | CC |
|
||||
| **R-412** | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** **CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it.** As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". **That is wrong, and measuring it overnight is what showed it:** the off-site run has its OWN pre-push dump leg — *"Stopping calibre-web for safe volume dump"*, *"Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB"* — so a unit that is hollow when a run starts is **REPAIRED before it is pushed**. Proven twice: `opengist` (2026-08-31 21:0x) and `calibre-web` (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (`6fee3b5a`) holds the volume tar and all 17 userdata files. **WHAT REMAINS REAL, and it is narrower:** the one hollow snapshot that DID reach the store (`35ba9fe7`, opengist) was created when the unit was destroyed **inside** a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged *"backed up opengist (… 0 mandatory path(s))"*, **a success line over a backup holding none of the app's data**. That race is real, it was observed, and the success wording is wrong either way. **The R-403 mirror guard holds throughout** — proven live: *"unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one"*, secondary byte-identical. | **LEG 1 CLOSED 2026-09-01 (controller v0.232.0) — LEG 2 STILL OPEN, LOW** | R-403, R-87, R-413 | **LEG 1 IS DONE:** a per-app push whose unit carried no database dump and no volume tar now logs at WARN and says what it did not carry, using the existing `unitIsHollow` predicate. Wording only — no guard, and the capture is untouched (08 §8.2). Pinned by `TestR412a_EmptyPushDoesNotReadAsAPlainSuccess`, which asserts the hollow line carries the words, the SOUND line does not, and neither is at INFO; red-proofed by restoring the single unconditional line. **LEG 2 IS STILL OPEN and is the remaining work on this row:** whether the push should RE-READ the unit it is about to send, or whether the race window is small enough to accept. Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. **Do NOT guard the capture** (08 §8.2). Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/` and `phase5-mutated-cycle/09-what-reached-the-store.txt`. | CC |
|
||||
| **R-413** | **R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended — the validation yesterday's session could only do with a declared hand-built fixture.** 2026-08-31 soak, demo-hp. After R-412's chain left `opengist`'s newest off-site snapshot hollow, the nightly proof rotated to it and returned **`verdict:"fail"`, `reason:"volumes_expected_none_captured"`, missing `opengist_data`**, logged *"READABLE AND EMPTY — the store is not damaged; the backup does not contain this app's data"*, and pushed **one** `offsite_proof_empty` at severity `error`. The four apps ahead of it in the rotation all passed, so the discrimination is real and not a constant fail. **This is recorded as a row rather than only as a report line because it upgrades a claim:** the capability map's R-87 row cites a CONSTRUCTED failing case; it can now cite a natural one. | **CLOSED 2026-08-31 — the claim it upgrades is recorded** | R-87, R-412 | Nothing to build. When the capability map is next touched, cite this instead of the constructed case. | CC |
|
||||
| **R-414** | **The nightly off-site PROOF is INERT on a box with no registered data drive, every night, and the only signal is a WARN in the log.** FOUND 2026-09-01 by the soak's UNTOUCHED observer, `demo-felhom`, on the first unattended run of the job — which is exactly what an untouched box was for. At 05:30 CEST the job fired, picked `opengist`, and refused: *"proof: opengist has nowhere to restore to: nincs regisztralt adatmeghajto, ezert nincs hova visszaallitani"* — `offboxRestoreScratchDir`'s R-252 refusal. **CAUSE ESTABLISHED, not inferred:** that box has **zero** registered storage paths (`storage_paths: []`), so there is no non-network schedulable path to put a scratch on. The same absence explains its `tier2-backup` completing in **3 ms** — a no-op with no second drive to mirror to. **The box is NOT unprotected:** its off-site backup ran normally in 46.9 s, because units live on the system data path, which needs no registration. **It is the PROOF that cannot run.** **WHY IT IS WORSE THAN A FAILED RUN:** the Err path reaches no verdict, so `RecordProofVerdict` is never called, so `last_proof_result` stays **ABSENT** — and absent is also what a controller older than v0.231.0 sends. **The hub therefore cannot tell "never ran" from "not deployed"**, which is the StatsKnown trap the field was explicitly designed to avoid, reappearing one level up. It will fail this way every night forever with nothing but a WARN. | **OPEN — MEDIUM** | R-87, R-402, R-252 | Decide what a box with no registered drive should do: fall back to the system data path for the scratch (it already holds the units), or record a distinct NOT-APPLICABLE verdict so the hub can tell it apart from never-ran. **Do not leave it as a WARN** — that is the shape R-397 and R-107 both cost a drill. Evidence: `audits/DRILL-soak-2026-08-31/phase6-observer/`. | CC |
|
||||
| **R-416** | **`closed_register_gate.py` still has no within-register duplicate-id rule.** R-406 closed by renumbering the only collision (R-133 → R-415), so the register is clean today and nothing stops the next one. The rule was deliberately NOT added in the same commit: with its only real subject removed, the red-proof would have had to be a planted fixture rather than the live defect, and this project's own standard is that a guard ships with a proof against something real. **Now that the register is clean it can be added safely** — a fresh duplicate would be the first thing it ever sees. | **OPEN — LOW** | R-406, R-405 | Add a third rule to `closed_register_gate.py`: no `R-` id may appear twice within either register. Ship it with a planted red-proof, and note that suffixed ids (R-88a/R-88b, R-209/R-209a) are distinct and must NOT be convicted. | CC |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
|
||||
Reference in New Issue
Block a user