R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s

07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild
rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2
skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the
shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question,
why the data legs are deliberately not guarded, and why the capture job is not guarded either.

00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left
ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what
the route can be relied on for.

Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a
DECISION and deliberately not acted on - should a documents-only push be subject to the
golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus
what happens if Viktor does nothing. The gate was NOT changed.

R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect
in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses
git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is
owed and is more urgent than the previous six.

STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now
says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with
its do-nothing outcome.

Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was
not installed in the guest and silently did nothing, and a session that expired mid-run so a POST
did nothing).
This commit is contained in:
2026-08-31 14:39:29 +02:00
parent 66156c619f
commit dddcc808be
10 changed files with 195 additions and 22 deletions
+34 -18
View File
@@ -13,21 +13,33 @@ delivered; both demo machines are on it and nothing is waiting on you about this
*This section is allowed to be longer than one screen, and each item says what happens if you do *This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.* nothing.*
1. **Nothing about delivery — the golden train is current.** Golden **0.229.0** was baked, published, 1. **A golden carrying 0.230.0 is owed.** Golden **0.229.0** is vouched and the fleet floor is 0.229.0
round-trip verified, vouched, and the fleet floor raised to 0.229.0 on 2026-08-31 (the second bake — **and 0.229.0 is the build that deletes a good copy** (R-403, measured today). Both demo machines
that day; 0.228.0 was the first). Both demo machines run it; **`demo-felhom` got there by itself** and any machine installed right now carry that defect. `demo-hp` has been updated to 0.230.0 by
and re-registered its jobs without anyone touching it. A machine installed today receives 0.229.0 hand; `demo-felhom` has not. **If you do nothing:** the fix stays on one machine and a newly
and everything shipped today, **including the restore from the second drive**. Evidence: installed box gets the defect. Baking and vouching 0.230.0 and raising the floor closes it — the
`documentation/tests/golden-0.229.0-2026-08-31/`. same three-field change as this morning.
2. **Nothing else about this release.** Everything in 0.229.0 is a fix to code that ships in the 2. **Nothing else about this release.** Everything in 0.229.0 is a fix to code that ships in the
controller image; no customer action, no data migration, no credential change. controller image; no customer action, no data migration, no credential change.
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. 3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
skipped that check **six times**, each time for a written reason: it runs on every push to the
website/documentation repository, including pushes that change nothing a machine installs.
**A guard we correctly skip six times is teaching us to skip it.**
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
**The case against:** the check was earned — a release went out while machines were still being
installed with the previous one, three times in three days — and narrowing a guard is how the thing
it was built for comes back.
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
I have NOT changed it; this is yours to decide and mine to build.
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one. it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is 5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour. as it misled one by an hour.
@@ -95,6 +107,20 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
database), and the undo copies no longer pile up forever — three per app, and they were being copied database), and the undo copies no longer pile up forever — three per app, and they were being copied
off-site permanently. off-site permanently.
- **An empty package can no longer wipe out a good one** (R-403, controller 0.230.0). Yesterday we
wrote this down as *suspected* and said plainly it had not been tested. **We tested it first, and it
was real.** On a demo machine, on yesterday's build: an app's copy on the second drive went from
**120 MB — four database backups and three data archives — to 7 KB, nothing left**, in a single
nightly run, and the run reported success. The cause was that the nightly job only asked *does the
folder exist* before copying over it, and an empty package is a folder that exists.
Now the nightly job refuses to replace a **complete** package with an **empty** one. It keeps what it
has, says so on the app's own backup page, and carries on with everything else. Proven on the same
machine, in the same state: **all seven files still there, byte for byte.**
Two more things came with it. The page no longer calls that copy fresh when the run did not refresh
it — it names the real date of the package instead. And after a restore from the second drive, the
first drive's package is filled back in immediately, so the empty state that started all this cannot
happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case
is fenced.
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0, - **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
@@ -151,16 +177,6 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
machine re-reads its whole store every week, however large it grows, and the first person to notice machine re-reads its whole store every week, however large it grows, and the first person to notice
would be a customer whose upload is busy. The warning is there so that does not happen. would be a customer whose upload is busy. The warning is there so that does not happen.
- **After a restore from the second drive, the first drive's package is rewritten EMPTY** (R-403,
found during the 0.229.0 drill, **not fixed**). Two seconds after the restore finished, the box's
five-minute housekeeping rebuilt the first drive's package from a drive that had no data files on
it, and wrote a package that lists nothing. The ordinary „Visszaállítás indítása" then read it and
said, correctly and uselessly, that the backup held only settings. **What we did not test:** the
nightly copy mirrors the first drive over the second one and deletes what is not there, so the
next night could plausibly overwrite the good copy with the empty one. That is a reading of the
code, not a measurement, and it is written down as unmeasured on purpose. **If you do nothing:**
a customer who recovers from their second drive may find, the next morning, that the copy they
recovered from has been replaced by an empty one. Settling it costs one test on a spare machine.
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we - **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
is **under a year** away on the corrected measurement, not two. is **under a year** away on the corrected measurement, not two.
@@ -126,7 +126,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|---|---|---|---|---| |---|---|---|---|---|
| Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded **DEGRADED** rather than silently normal (installer Case A/B) | agent v0.113, host-install v1.22.0 | **PROVEN-LIVE** | `E2D-fresh-vm-2026-07-29` C1 (real 1.22.0 install, rc=0, `Day-0 provision SUCCESS`) + C2 (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort) | Case A (a second drive already present at install) has never fired naturally — only Case B has | | Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded **DEGRADED** rather than silently normal (installer Case A/B) | agent v0.113, host-install v1.22.0 | **PROVEN-LIVE** | `E2D-fresh-vm-2026-07-29` C1 (real 1.22.0 install, rc=0, `Day-0 provision SUCCESS`) + C2 (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort) | Case A (a second drive already present at install) has never fired naturally — only Case B has |
| Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 **DB replay route → `07-backup-architecture.md` §8 row 3** | | Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 **DB replay route → `07-backup-architecture.md` §8 row 3** |
| Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 3b, 4, 5.** **R-102 CLOSED 2026-08-31 (controller v0.229.0):** the copy's `recovery-unit/` mirror is now restorable — „Teljes visszaállítás a másolatból" / `POST /backup/tier2/unit-restore` — and was proven live on `demo-hp` with the primary unit moved aside (`audits/DRILL-r102-tier2-unit-2026-08-31/`, §8 row 3b, 28.65 s). R-103 closed with it: the refusal that used to name a button on another page now offers the action on the row itself | | Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 3b, 4, 5.** **R-403 (controller v0.230.0, 2026-08-31) changes NO status on this row and that is stated rather than left ambiguous:** the nightly copy now refuses to replace a complete unit package with an empty one (`07` §8.2), which removes a way the route could be DESTROYED between uses — it does not change what the route can be relied on for, and every leg it covers is the same one it covered yesterday. Proven live on demo-hp: the same state that deleted 120 082 104 B on v0.229.0 preserved all 7 files on v0.230.0. **R-102 CLOSED 2026-08-31 (controller v0.229.0):** the copy's `recovery-unit/` mirror is now restorable — „Teljes visszaállítás a másolatból" / `POST /backup/tier2/unit-restore` — and was proven live on `demo-hp` with the primary unit moved aside (`audits/DRILL-r102-tier2-unit-2026-08-31/`, §8 row 3b, 28.65 s). R-103 closed with it: the refusal that used to name a button on another page now offers the action on the row itself |
| Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Route + RTO → `07-backup-architecture.md` §8 rows 4, 10, 12, 15** (incl. the R-95 delete exposure and the R-104 stale-lock defect). **2026-08-04 (R-193/R-197, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`) — the row is NOT overclaiming and the guard is not the gap; the CADENCE is.** This row already recorded that a recreated data volume orphans the repo, and the spike confirms the mechanism at source: `WriteOffboxSecrets` (`offbox.go:392`) mints a fresh 256-bit repo password whenever `<DataDir>/offbox/repo_password` is absent, and **no automatic path ever consults the escrowed one** — `InjectOffboxPassword` has exactly one caller, a web form a human pastes into. What was NOT recorded is that this now fires on an **ordinary, planned, unattended guest rebuild**, on every box: measured on BOTH demo boxes 2026-08-03/04 by comparing `host_escrow.restic_pw_sha256` against `host_escrow_superseded.restic_pw_sha256` (demo-hp `8e03eddf…`→`8a9e33aa…`, demo-felhom `48741892…`→`c60c8bc7…`), orphaning **15 snapshots / 40.9 MB** and **36 snapshots / 1.14 GB** respectively. **demo-felhom is the important half:** it kept its DELIVERY (a stale staged secret restored the target in 76 s) and lost its REPOSITORY anyway, with **nothing marking the escrow stale for 13 h** — `escrow_stale` is wired to `ReissueCredentials`, the one path that does NOT change the repo password (**R-196**), and absent from the rebuild path that does. Status unchanged: the classify-and-move-aside guard remains PROVEN-LIVE and correct, and is predicted (not yet measured) to refuse the 2026-08-05 run on both boxes rather than start a silent fresh history | | Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Route + RTO → `07-backup-architecture.md` §8 rows 4, 10, 12, 15** (incl. the R-95 delete exposure and the R-104 stale-lock defect). **2026-08-04 (R-193/R-197, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`) — the row is NOT overclaiming and the guard is not the gap; the CADENCE is.** This row already recorded that a recreated data volume orphans the repo, and the spike confirms the mechanism at source: `WriteOffboxSecrets` (`offbox.go:392`) mints a fresh 256-bit repo password whenever `<DataDir>/offbox/repo_password` is absent, and **no automatic path ever consults the escrowed one** — `InjectOffboxPassword` has exactly one caller, a web form a human pastes into. What was NOT recorded is that this now fires on an **ordinary, planned, unattended guest rebuild**, on every box: measured on BOTH demo boxes 2026-08-03/04 by comparing `host_escrow.restic_pw_sha256` against `host_escrow_superseded.restic_pw_sha256` (demo-hp `8e03eddf…`→`8a9e33aa…`, demo-felhom `48741892…`→`c60c8bc7…`), orphaning **15 snapshots / 40.9 MB** and **36 snapshots / 1.14 GB** respectively. **demo-felhom is the important half:** it kept its DELIVERY (a stale staged secret restored the target in 76 s) and lost its REPOSITORY anyway, with **nothing marking the escrow stale for 13 h** — `escrow_stale` is wired to `ReissueCredentials`, the one path that does NOT change the repo password (**R-196**), and absent from the rebuild path that does. Status unchanged: the classify-and-move-aside guard remains PROVEN-LIVE and correct, and is predicted (not yet measured) to refuse the 2026-08-05 run on both boxes rather than start a silent fresh history |
| Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live | controller v0.134/134.1/135 | **PROVEN-LIVE** (2026-07-20) | `CAMPAIGN-6D` accept legs (immich end-to-end from offsite alone) | **2026-07-19:** `audits/DIAG-immich-restore-2026-07-19.md` finds **no offsite path loads a DB dump** — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "**immich end-to-end from offsite alone**" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. **RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume** (`immich_postgres_data` is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" **overclaimed scope**: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → **PARTIAL**, scope-corrected. Evidence: `audits/DIAG-immich-restore-2026-07-19.md` (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). **2026-07-19, controller v0.148.0:** the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but **round 2 found it aborts against a running app** (`audits/DIAG-immich-restore-round2-2026-07-19.md`, H4: the replay races immich's own schema repair; `clip_index` recreated by the app 2 s before the dump's CREATE INDEX). **2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47)** — both restore paths now replay into a DB-ONLY window (`StartStackServices` brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. *(The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard `BackupStatus` fix. R-47 shipped in v0.153.0.)* **2026-07-20: the clean run HAPPENED** — endpoint-level supervised reconstitute of immich from snapshot `49e7cb46` (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no `already exists`, operation reported SUCCESS, immich's own DatabaseService logged `No schema drift detected` twice, 11 assets `active`, 4/4 containers healthy, 231 `public` indexes. **Operator confirmed the immich timeline renders correctly after the reconstitute** (screenshot held, 2026-07-20). Evidence: `felhom-controller/REPORT.md` §4b. **2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE.** The operator deleted the photos in immich own UI **and emptied the trash** (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. **`40 file(s) placed`** against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets `active`, `No schema drift detected`, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: `felhom-controller/REPORT.md` 4e **Route + RTO → `07-backup-architecture.md` §8 rows 3, 4** — the matrix also records that no offsite action unpacks the named-volume tars it captures (→ R-107) | | Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live | controller v0.134/134.1/135 | **PROVEN-LIVE** (2026-07-20) | `CAMPAIGN-6D` accept legs (immich end-to-end from offsite alone) | **2026-07-19:** `audits/DIAG-immich-restore-2026-07-19.md` finds **no offsite path loads a DB dump** — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "**immich end-to-end from offsite alone**" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. **RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume** (`immich_postgres_data` is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" **overclaimed scope**: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → **PARTIAL**, scope-corrected. Evidence: `audits/DIAG-immich-restore-2026-07-19.md` (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). **2026-07-19, controller v0.148.0:** the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but **round 2 found it aborts against a running app** (`audits/DIAG-immich-restore-round2-2026-07-19.md`, H4: the replay races immich's own schema repair; `clip_index` recreated by the app 2 s before the dump's CREATE INDEX). **2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47)** — both restore paths now replay into a DB-ONLY window (`StartStackServices` brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. *(The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard `BackupStatus` fix. R-47 shipped in v0.153.0.)* **2026-07-20: the clean run HAPPENED** — endpoint-level supervised reconstitute of immich from snapshot `49e7cb46` (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no `already exists`, operation reported SUCCESS, immich's own DatabaseService logged `No schema drift detected` twice, 11 assets `active`, 4/4 containers healthy, 231 `public` indexes. **Operator confirmed the immich timeline renders correctly after the reconstitute** (screenshot held, 2026-07-20). Evidence: `felhom-controller/REPORT.md` §4b. **2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE.** The operator deleted the photos in immich own UI **and emptied the trash** (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. **`40 file(s) placed`** against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets `active`, `No schema drift detected`, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: `felhom-controller/REPORT.md` 4e **Route + RTO → `07-backup-architecture.md` §8 rows 3, 4** — the matrix also records that no offsite action unpacks the named-volume tars it captures (→ R-107) |
| Manual `.fab` export/import: class-scoped capture, browser up/download, tunnel-proof chunking | controller v0.125/128/130/136 | **PROVEN-LIVE** | `CAMPAIGN-6D` P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking `CAMPAIGN-6B` P2 (100 MiB via real CF edge, 120 MiB→413) | Chunking proven at the real CF edge via `curl --resolve`; the **rendered browser file-picker** upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B *finding*; fix verified in 6D | | Manual `.fab` export/import: class-scoped capture, browser up/download, tunnel-proof chunking | controller v0.125/128/130/136 | **PROVEN-LIVE** | `CAMPAIGN-6D` P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking `CAMPAIGN-6B` P2 (100 MiB via real CF edge, 120 MiB→413) | Chunking proven at the real CF edge via `curl --resolve`; the **rendered browser file-picker** upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B *finding*; fix verified in 6D |
@@ -864,7 +864,7 @@ crosses the line — **R-158**.
| 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery | | 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery |
| 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** | | 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** |
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore | | 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore |
| 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go:359-393` | | 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go`. **The derived-copy rule is UNCHANGED by R-403 (controller v0.230.0) and the single exception is stated in §8.2 below — read it before "fixing" a skip you find in the code** |
| 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session | | 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session |
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | | 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
@@ -877,6 +877,47 @@ crosses the line — **R-158**.
| 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix | | 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix |
| 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** | | 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** |
### 8.2 The ONE exception to row 5's derived-copy rule (R-403, controller v0.230.0)
**[DESIGN] Row 5 stands: the secondary IS a derived copy and it IS rebuilt on the next run.** Nothing
below weakens that, and a future reader who finds `RunTier2` skipping a leg and does not find this
section will "fix" it back — which is why it is here and not only in a register row.
**[FACT] What was measured, on the shipped v0.229.0, on `demo-hp` 2026-08-31.** An app's Tier-2 copy
went from **120 082 104 B** (4 database dumps + 3 named-volume tars) to **7 036 B** (none of either)
in one nightly run, and the run recorded itself a success. The mechanism was three individually
correct lines: `RunTier2` guarded the unit leg with `os.Stat` alone — *does the folder exist* —
`rsyncMirror` is `rsync -a --delete`, and nothing between them compared source to destination. **An
empty recovery unit is a folder that exists.** Evidence:
`audits/DRILL-r403-tier2-delete-2026-08-31/`.
**Why it bites harder since 2026-08-30:** R-102 made that mirror a LIVE recovery route (§6.3, §8 row
3b). Deleting it used to cost a copy nobody could open; it now costs the route itself.
**[DESIGN] The exception, stated exactly.** The unit leg — and ONLY the unit leg — is skipped when the
SOURCE unit carries no data and the DESTINATION unit does. Everything else is unchanged:
| source unit | destination unit | behaviour |
|---|---|---|
| complete | complete | mirror, with `--delete`, as before |
| complete | hollow or absent | mirror (the normal first copy) |
| hollow | hollow | mirror — both sides agree, nothing is at risk |
| **hollow** | **complete** | **skip the unit leg, preserve the destination, warn, record for the surface** |
**"Hollow" is a MANIFEST question, never a size question** — the manifest lists no database dump and
no volume tar; absent or unparseable counts as hollow, fail closed. A unit with a fat compose capture
and no dumps is the dangerous shape; a 360-byte unit belonging to a tiny app is healthy.
**The data legs are NOT guarded and must not be.** A classified app's copy legitimately shrinks as
`export` drops out of its class set (`tier2.go` header), and fencing that would be calling this row's
own decision a defect.
**[DESIGN] And the CAUSE is closed at the other end.** The hollow primary was written by the 5-minute
capture job **two seconds** after a Tier-2 unit restore. Since v0.230.0 `RestoreTier2Unit` refills an
absent or hollow primary unit from the mirror **inside the call, before returning**, so no capture can
observe the hollow state. **The capture itself is deliberately not guarded:** a capture that describes
an empty drive as empty is correct, and guarding it would make the manifest lie.
### 8.1 The blank cells, listed explicitly ### 8.1 The blank cells, listed explicitly
Per the rule that a blank is a finding, here they are: Per the rule that a blank is a finding, here they are:
@@ -0,0 +1,18 @@
=== THE WARN, verbatim ===
2026/08/31 12:14:24 tier2.go:383: [WARN] [backup] Tier 2 docmost: unit leg SKIPPED — the recovery unit on the source drive lists no database dumps and no volume tars, while the existing copy at /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit does. The copy was PRESERVED rather than replaced with an empty one (R-403). The other legs continue.
2026/08/31 12:14:24 tier2.go:424: [INFO] [backup] Tier 2 copied docmost → /mnt/felhom-drives/hdd_1/backups/secondary/docmost (14.9 KB, 0 leg(s), 0s) [unit leg SKIPPED — existing package preserved, R-403]
=== the run summary line ===
2026/08/31 12:14:24 tier2.go:424: [INFO] [backup] Tier 2 copied docmost → /mnt/felhom-drives/hdd_1/backups/secondary/docmost (14.9 KB, 0 leg(s), 0s) [unit leg SKIPPED — existing package preserved, R-403]
=== AFTER the run ===
db-dumps: 4 volume-dumps: 3 size: 120082104
9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28 db-dumps/docmost-postgres.sql
73917ba6bc3072dfc7b5be9c6df4f8361da7e987230f5d56f7b62f397fe15ef1 db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
4c134c2ced74df26f49ef1910694cbbd25f2598549bb4b5bb145aac054935949 db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
13e5a864701966d9e4053b5bb7dd800cca3d77ebb07f4fd2f32c86389422af25 db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
f46a2fc3aa9a7ae2502d83b1c6ef27e503102f5ba71a0c6559246d9674c8e3b1 volume-dumps/docmost_docmost_postgres_data.tar
a8df17c444e41f54762e122ce1be998315c969015a1580bb8abdc7211cfa1a73 volume-dumps/docmost_docmost_redis_data.tar
88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d volume-dumps/docmost_docmost_storage.tar
VERDICT: PRESERVED — all 7 files still there (v0.229.0 left 0)
@@ -0,0 +1,23 @@
######## SCENARIO B LIVE — the same state, on the FIXED build ########
UTC 2026-08-31T12:04:11Z
controller under test: gitea.dooplex.hu/admin/felhom-controller:0.230.0
--- recreate the hollow primary (hand-made this time; the R-102 path was proven in phase 1b)
hollow primary manifest written: db_dumps=[] volume_dumps=null
primary tree:
compose/.felhom.yml
compose/app.yaml
compose/docker-compose.yml
manifest.json
=== BEFORE the run ===
secondary db-dumps: 4 volume-dumps: 3 size: 120082104
9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28 db-dumps/docmost-postgres.sql
73917ba6bc3072dfc7b5be9c6df4f8361da7e987230f5d56f7b62f397fe15ef1 db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
4c134c2ced74df26f49ef1910694cbbd25f2598549bb4b5bb145aac054935949 db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
13e5a864701966d9e4053b5bb7dd800cca3d77ebb07f4fd2f32c86389422af25 db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
f46a2fc3aa9a7ae2502d83b1c6ef27e503102f5ba71a0c6559246d9674c8e3b1 volume-dumps/docmost_docmost_postgres_data.tar
a8df17c444e41f54762e122ce1be998315c969015a1580bb8abdc7211cfa1a73 volume-dumps/docmost_docmost_redis_data.tar
88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d volume-dumps/docmost_docmost_storage.tar
=== POST /api/backup/tier2 ===
@@ -0,0 +1,10 @@
rows found: 8
app notice FIGYELEM skipped? package date in the confirm
bookstack False False False 2026-08-31 14:03
calibre-web False False False 2026-08-31 14:03
docmost True True True 2026-08-31 11:43
kimai False False False 2026-08-31 14:03
opengist False False False 2026-08-31 14:03
paperless-ngx False False False 2026-08-31 14:03
privatebin False False False 2026-08-31 14:03
romm False False False 2026-08-31 14:03
@@ -0,0 +1,19 @@
######## SCENARIO E LIVE — the primary is filled back in ########
UTC 2026-08-31T12:24:38Z
--- BEFORE: the primary is hollow (this is the state a Tier-2 unit restore leaves on 0.229.0)
created_at: 2026-08-31T12:08:49Z db_dumps: [] volume_dumps: None
compose/.felhom.yml
compose/app.yaml
compose/docker-compose.yml
manifest.json
--- the R-102 restore, through the real endpoint
302 https://127.0.0.1:443/backups/apps?flash=Teljes+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
{"ok":true,"data":{"running":false,"op":"tier2-unit-restore","stack":"docmost","started_at":"2026-08-31T12:24:38.178701187Z","last":{"op":"tier2-unit-restore","stack":"docmost","ok":true,"message":"A(z) docmost: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-08-31 14:23).","finished_at":"2026-08-31T12:25:07.772790659Z"},"last_recent":true}}
--- IMMEDIATELY after the call returned (no waiting): is the primary a real package?
created_at: 2026-08-31T09:43:41Z db_dumps: ['docmost-postgres.sql'] volume_dumps: ['docmost_docmost_postgres_data.tar', 'docmost_docmost_redis_data.tar', 'docmost_docmost_storage.tar']
volume tars on the app drive: 3 db dumps: 4
--- the rehydrate log line, verbatim
2026/08/31 12:25:07 tier2_restore.go:253: [INFO] [backup] docmost: primary unit refilled from the secondary mirror (R-403) — 3 volume tar(s), 4 database dump(s) now on the app's own drive
@@ -0,0 +1,45 @@
######## TEARDOWN / final state ########
UTC 2026-08-31T12:33:57Z
--- a normal Tier-2 run now that BOTH sides are complete (the unit leg must NOT be skipped)
200
2026/08/31 12:33:57 tier2.go:424: [INFO] [backup] Tier 2 copied docmost → /mnt/felhom-drives/hdd_1/backups/secondary/docmost (114.5 MB, 0 leg(s), 0s)
unit-leg skips in this run: 0
--- both copies
PRIMARY : 3 tars, 4 dumps, 120082104 bytes
SECONDARY : 3 tars, 4 dumps, 120082104 bytes
IDENTICAL volume-dumps/docmost_docmost_postgres_data.tar
IDENTICAL db-dumps/docmost-postgres.sql
--- the surface is back to normal (no stale notice anywhere)
adatcsomagja occurrences: 0
--- remove the drill safety net and the driver
safekeeping removed
.bash_history
.bashrc
.docker
.profile
.ssh
settings.json.drill-backup
--- every app healthy?
docmost Up 9 minutes (healthy)
docmost-redis Up 9 minutes (healthy)
docmost-postgres Up 9 minutes (healthy)
romm Up 3 hours (healthy)
romm-db Up 3 hours (healthy)
romm-redis Up 3 hours (healthy)
privatebin Up 3 hours (healthy)
paperless-webserver Up 3 hours (healthy)
paperless-postgres Up 3 hours (healthy)
paperless-redis Up 3 hours (healthy)
opengist Up 3 hours (healthy)
kimai Up 3 hours (healthy)
kimai-db Up 3 hours (healthy)
calibre-web Up 3 hours (healthy)
bookstack Up 3 hours (healthy)
bookstack-db Up 3 hours (healthy)
filebrowser Up 9 days (healthy)
cloudflared Up 9 days
traefik Up 9 days
+1
View File
@@ -237,3 +237,4 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` | | **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
| **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` | | **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
| **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` | | **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
File diff suppressed because one or more lines are too long