R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.
6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.
00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.
Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show 1623a4d5b5 for the original). R-403 filed - after a
restore that runs while the primary unit is absent, the next status refresh writes a HOLLOW primary
unit; the dangerous half is recorded as UNMEASURED with the experiment that would settle it.
R-242: sixth conviction of golden_currency_gate. THIS PUSH USES git push --no-verify, declared here
and in felhom-controller/REPORT.md - a BYPASS, not a waiver. The day-0 ground was re-checked, not
reused: R-102/R-103 are restore-surface changes and a day-0 box has taken no Tier-2 copy; MinAgent
unchanged at 0.129.0. OWED: bake a golden carrying 0.229.0, vouch it, raise the floor.
Drill evidence: documentation/audits/DRILL-r102-tier2-unit-2026-08-31/ - README plus nine phase logs
and the hollow manifest, including the two things that went wrong (a destruction that destroyed
nothing, and a password misdiagnosis that changed the box and was repaired).
This commit is contained in:
@@ -308,17 +308,33 @@ entry wins; else a `:ro` reader is *excluded*; else a writable bind is *mandator
|
||||
| **Tier-2** (`mandatory + optional`) | **7** — the 4 above plus audiobookshelf, komga, romm |
|
||||
| legacy resolver path (apps with no block) | **0** — no non-block template binds a namespace path |
|
||||
|
||||
> ### ⚠️ UNRESOLVED — two counts of the same thing disagree
|
||||
> ### ✔ RESOLVED 2026-08-31 (controller v0.229.0) — **A = 7 · B = 45 · C = 1**
|
||||
>
|
||||
> | source | count |
|
||||
> |---|---|
|
||||
> | C9-F1 Phase 0, as shipped | **A = 9 / B = 43 / C = 1** — `felhom.eu/REPORT.md:17-24`, restated `backlog/OPEN-ITEMS.md:31` |
|
||||
> | INV Part B.1, independent enumeration at catalog `4252121` | **A = 7 / B = 45 / C = 1** |
|
||||
> | source | count | verdict |
|
||||
> |---|---|---|
|
||||
> | C9-F1 Phase 0, as shipped | **A = 9 / B = 43 / C = 1** — `felhom.eu/REPORT.md:17-24` (since overwritten), restated `backlog/OPEN-ITEMS.md:31` | **WRONG by two apps** |
|
||||
> | INV Part B.1, independent enumeration at catalog `4252121` | **A = 7 / B = 45 / C = 1** | **CORRECT** |
|
||||
> | Measured at catalogue `459766cb16395fd1d1a66282f5cc6da59ead5924`, 2026-08-31 | **A = 7 / B = 45 / C = 1** | adopted |
|
||||
>
|
||||
> Both use the same definition ("templates whose Tier-2 copy can hold a readable file leg"). The
|
||||
> difference is two apps and **neither number is adopted here**. The Phase-0 enumeration is described
|
||||
> in prose but the script is not committed, so the two methods cannot be diffed from the repo.
|
||||
> **This must be resolved before either figure is used to size anything.**
|
||||
> **The method, so it can be re-run rather than re-argued.** A throwaway `main` inside the controller
|
||||
> module drove the PRODUCTION rule over all 53 template directories — `stacks.LoadMetadata` (the single
|
||||
> validation choke point, so a rejected `backup:` block degrades to legacy exactly as it does live) →
|
||||
> `stacks.ParseComposeClassifiableBinds` → `appbackup.ClassifyBinds` → `appbackup.ComputeCaptureSet` at
|
||||
> `TierSecondary`, with the legacy branch falling back to `AppDataBindsPresent` + `AppDataDirNames` as
|
||||
> `backup.tier2CaptureSet` does. **A = at least one leg survives that pipeline.** 13 templates carry a
|
||||
> valid `backup:` block; 40 are legacy and none of them binds a namespace path, so all 40 are B or C.
|
||||
>
|
||||
> **A (7):** audiobookshelf, calibre-web, immich, komga, nextcloud, paperless-ngx, romm.
|
||||
> **C (1):** bentopdf — it declares no `volumes:` key and no `${…_PATH}` bind at all.
|
||||
> **B (45):** everything else.
|
||||
>
|
||||
> **How the earlier disagreement arose, established rather than guessed.** Phase 0's own write-up
|
||||
> (controller `CHANGELOG.md`, v0.183.0) says four apps — plex, jellyfin, emby, navidrome — are in B
|
||||
> "only because their single bind is a `:ro` media mount, which `ClassifyBinds` correctly excludes". It
|
||||
> applied the `:ro` **default** rule. The two apps it therefore missed are **radarr and sonarr**: their
|
||||
> `${USERDATA_PATH}/media/*` and `${USERDATA_PATH}/downloads` binds are **writable**, so the `:ro` rule
|
||||
> does not reach them, and they are excluded by an **explicit** `class: excluded` entry instead. 9 − 2 =
|
||||
> 7 and 43 + 2 = 45, which is exactly the gap. No catalogue file was changed; this is a measurement.
|
||||
|
||||
**[FACT] The class-B consequence is real regardless of which count is right.** For an app whose data
|
||||
lives entirely in named volumes, the Tier-2 copy holds a full `recovery-unit/` and **no readable
|
||||
@@ -333,7 +349,7 @@ refuses **before** stopping the app and names the action that works.
|
||||
| tier | captured | read back by that tier's restore | gap |
|
||||
|---|---|---|---|
|
||||
| Tier-1 | unit incl. volume tars + DB dumps | all of it | none |
|
||||
| Tier-2 | unit mirror **+** file legs | `hdd/` and `userdata/` **only** (`tier2_restore.go:101-104`) | **the unit mirror is read by nothing** — `RecoveryUnitPath` resolves to `backups/primary/` (`appbackup/paths.go:46-48`) → **R-102** |
|
||||
| Tier-2 | unit mirror **+** file legs | the file restore reads `hdd/` and `userdata/`; **since controller v0.218.0's Tier-3 sibling and now v0.229.0, a SECOND action reads the unit mirror itself** (`RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`, `tier2_restore.go`) | **CLOSED — R-102** |
|
||||
| Tier-3 | unit (incl. volume tars) + mandatory legs | files + DB replay **+ the named-volume tars, replayed from the scratch unit** (`offbox_reconstitute.go` `volReplay`, controller **v0.218.0**); the unit itself is still **skipped** on the way to live (`offbox_reconstitute.go:284-289`; placed only if the live unit is absent, `offbox_restore.go:352-356`) | **CLOSED — R-107** |
|
||||
|
||||
**[FACT] 2026-08-22 — the Tier-3 row above was corrected; the Tier-2 row was NOT.** Until controller
|
||||
@@ -344,13 +360,25 @@ proven live on `demo-hp`. The old sentence is kept here, in the past tense, beca
|
||||
erases what was believed leaves the next reader no way to tell a fixed gap from one that was never
|
||||
noticed.
|
||||
|
||||
**R-102 — the Tier-2 half — is NOT closed and nothing in this correction touches it.** The Tier-2 row
|
||||
above stands exactly as written: the secondary unit mirror is still read by nothing. Do not read
|
||||
"R-107 closed" as covering both; they were always two register rows, and only one of them moved.
|
||||
**[FACT] 2026-08-31 — the Tier-2 half is now closed too (R-102, controller v0.229.0), and the old
|
||||
sentence is kept here in the past tense for the reason the paragraph above gives.** Until v0.229.0 this
|
||||
table said of Tier-2: *"the unit mirror is read by nothing — `RecoveryUnitPath` resolves to
|
||||
`backups/primary/` (`appbackup/paths.go:46-48`)"*, and that was true from the day Tier-2 shipped until
|
||||
2026-08-31. The mechanism was a hard-coded `primary` segment: every reader of a recovery unit could
|
||||
only name a path under it. `appbackup` now also exposes four **unit-directory-relative** primitives,
|
||||
`Manager.RestoreFromRecoveryUnitAt(stack, unitDir)` holds the restore body, and `RestoreTier2Unit`
|
||||
points it at `<dest>/backups/secondary/<app>/recovery-unit/`. Proven live on `demo-hp` **with the
|
||||
primary unit moved aside** — `audits/DRILL-r102-tier2-unit-2026-08-31/`.
|
||||
|
||||
**[FACT]** Tier-2's gap is the sharper one because of *when* it bites: Tier-2 exists for the case
|
||||
**Two things did NOT change, and both are load-bearing.** THE SOURCE MOVED; THE DESTINATION DID NOT —
|
||||
data still lands in the live named volumes and the live database container. And the FILE restore's
|
||||
reach is unchanged: `CanRestore()` still answers only "are there file legs?", and
|
||||
`tier2UnitNotCoveredMsg` is still appended where that restore runs, so a clean file result never reads
|
||||
as a clean bill of health for the database.
|
||||
|
||||
**[FACT]** Tier-2's gap was the sharper one because of *when* it bit: Tier-2 exists for the case
|
||||
where the primary drive is lost — and in exactly that case the primary unit is gone while this
|
||||
mirror survives on the second drive, unreachable by any customer action.
|
||||
mirror survives on the second drive. Until v0.229.0 it was unreachable by any customer action.
|
||||
|
||||
**[FACT] 2026-08-23 — taking the undo copy used to DESTROY the app's own database backup (R-361).**
|
||||
`writeSafetyDump` called `DumpOne` into the app's own unit directory and renamed the result to
|
||||
@@ -605,10 +633,20 @@ anywhere under the backup namespace. Full record: `audits/D5-drive-alone-restore
|
||||
|
||||
**[FACT]** Two instances, both current:
|
||||
|
||||
- **Tier-2 vs primary-drive loss.** Tier-2's stated purpose is surviving the loss of the primary
|
||||
drive. In that failure the primary recovery unit is gone; the surviving mirror on the second drive
|
||||
is `backups/secondary/<app>/recovery-unit/`, which **no code path reads** (§6.3). For the 45-or-43
|
||||
class-B apps the restore is a guaranteed no-op in exactly its designed scenario. → **R-102**
|
||||
- **~~Tier-2 vs primary-drive loss~~ — CLOSED 2026-08-31, controller v0.229.0 (R-102).** Tier-2's
|
||||
stated purpose is surviving the loss of the primary drive. It was true until v0.229.0 that in that
|
||||
failure the primary recovery unit is gone while the surviving mirror on the second drive —
|
||||
`backups/secondary/<app>/recovery-unit/` — was read by **no code path** (§6.3), so for the **45**
|
||||
class-B apps (§6.2, count settled the same day) the restore was a guaranteed no-op in exactly its
|
||||
designed scenario.
|
||||
**Tier-2 can now meet its prerequisite in the failure it exists for**, and that is stated plainly
|
||||
because it is the whole point: the restore was run on `demo-hp` **with the primary unit moved aside**
|
||||
and returned 3 volumes of 3 and 1 database of 1 in 28.65 s, byte-for-byte, with the app then reading
|
||||
its own row over TCP with its own credential — and again with the guest's `app.yaml` also moved
|
||||
aside, `secrets recovered=2/2` from the mirrored unit. Evidence:
|
||||
`audits/DRILL-r102-tier2-unit-2026-08-31/`.
|
||||
**What the drill did NOT cover, stated per §8 below:** the full drive-loss journey — a genuinely
|
||||
absent or replaced physical drive — was not run. Only the recovery-unit half was.
|
||||
- **Tier-3 vs guest loss — HALF of this closed.** Tier-3 holds the volume tars and the DB dump. It
|
||||
was true until controller v0.218.0 that the tars were unpacked only by the Tier-1 path; since
|
||||
v0.218.0 the reconstitution replays them itself (`volReplay`) → **R-107 CLOSED 2026-08-22**. What
|
||||
@@ -824,8 +862,8 @@ crosses the line — **R-158**.
|
||||
| 2 | **An app's data directory is destroyed** | the guest, the other tiers | same route | **customer** | **46 s** (43 files) | 24 h | **PROVEN** | CAMPAIGN-9 A3 — loss verified by a positive observable (doc download 200 → 404), then 16/16 documents usable |
|
||||
| 3 | **An app's DB and named volumes are lost** | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the **only** path that unpacks volume tars) | **customer** | **18.25 s** (path execution) · **27.6 s** (D5 drill, guest `app.yaml` absent) | 24 h | **PROVEN** (content recovery proven 2026-07-30) | CAMPAIGN-9 A2 proved the path executes; **D5 v0.188.0 closed the content gap** — after a restore with the guest's `app.yaml` moved aside, the app read the seeded row **over TCP with its own credential**, the pre-backup row returned and a post-backup row was gone (so the tar was really restored). No `.sql` dump in the unit ⇒ the DB came back from the volume tar. **2026-08-30 — this row KEEPS its PROVEN status and the reason is worth stating: R-353 was a defect in the MESSAGE, not in the mechanism.** The restore really did return what the unit held, every time; what it could not do was say so, because the count was discarded one call deep. Controller v0.226.0 fixed the sentence and changed nothing about the recovery path. A status that measures whether data comes back must not move because a status line was wrong |
|
||||
| 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery |
|
||||
| 3b | *same, for a class-B app via Tier-2* | — | **no route** — Tier-2 never reads the unit mirror | — | | | **NONE** | §6.3; **R-102** |
|
||||
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 or 9 of 53 apps); Tier-3 reconstitute for files + DB **+ the named-volume tars since controller v0.218.0** (`volReplay`); **the Tier-2 copy's volume tars remain unreachable** | **customer** (both) | | 24 h | **PARTIAL** | §7.2; **R-102** (open); **R-107 CLOSED 2026-08-22, v0.218.0** |
|
||||
| 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** |
|
||||
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore |
|
||||
| 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go:359-393` |
|
||||
| 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session |
|
||||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||||
@@ -845,8 +883,7 @@ Per the rule that a blank is a finding, here they are:
|
||||
|
||||
| row | blank | why |
|
||||
|---|---|---|
|
||||
| 3b | RTO, RPO | no route exists to time |
|
||||
| 4 | RTO | no drive-loss recovery has ever been timed |
|
||||
| 4 | RTO | no drive-loss recovery has ever been timed. **The Tier-2 unit restore inside it now is** — 28.65 s, row 3b, 2026-08-31 — but that is the route, not the journey: no drive has been removed or replaced under a recovery |
|
||||
| 5 | RTO | never timed; the rebuild is a normal Tier-2 run |
|
||||
| 6 | — | RTO present, but it is a **restore into a scratch guest on the same host**; a restore *to a different host* has never been timed |
|
||||
| 8 | RTO | a host has never been rebuilt as itself (INV Part D1) |
|
||||
|
||||
Reference in New Issue
Block a user