D5 SHIPPED: Tier-1/2 restore no longer depends on the whole-guest tier
Records controller v0.188.0 across the four coupled artifacts. 07-backup-architecture.md is the owning doc: - new 7.4 = the recovery chain AFTER D5 (7.1 leg 1 superseded; leg 2, the living-app dependency, explicitly unchanged so this is not read as more than it is) - 7.3 collapsed to history, with the correction that the target as written (data_key-only) was tested in Part 0 and rejected - 3 records that the two-lane split is now real, not just intended - matrix rows 3 / 3c (new) / 13; 10.1 D5 itself shipped Also: new capability-map row, D5 collapsed in ROADMAP + OPEN-ITEMS, and R-127 filed in both (data_key flag unreliable; O4 can regenerate a DB password that no longer matches the restored data directory). The audit is named D5-drive-alone-restore rather than "...secrets..." because .gitignore blocks *secret* -- a guard worth respecting, not forcing past.
This commit is contained in:
@@ -96,6 +96,7 @@
|
||||
| Data migration between drives (all / per-app), crash-safe | controller | **PROVEN-LIVE** | `CAMPAIGN-6C` 4P-5 (scope=app round-trip, byte-identical); `storage-lifecycle-acceptance-2026-06-15` (two migrate-all runs via dashboard UI, sha256 byte-identical) | (Cited `CAMPAIGN-2` T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS |
|
||||
| NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated | controller v0.113–117, agent v0.81/84/85 | **PROVEN-LIVE** | `SPIKE-nas-verify-2026-07-11`, `SPIKE-nas-dsm-2026-07-11`, `CAMPAIGN-3-2026-07-11` (boot/reassert fixes) | |
|
||||
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
|
||||
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore, and Tier-2's own cross-drive copy of a secret-bearing unit (both unit-tested only). Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
||||
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
|
||||
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
|
||||
| Boot ordering: automount + networking survive reboot; appliance self-heal watchdog | agent v0.85 | **PROVEN-LIVE** | `CAMPAIGN-4-2026-07-13` (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + `CAMPAIGN-6A-2026-07-14` 1D (re-arm reboot-survival across 9 guest + 1 host reboots) | (`CAMPAIGN-3` F10/F11/F12 were the CRITICAL/HIGH *failures*; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: `skip-active` on `pct reboot` carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A) |
|
||||
|
||||
@@ -140,6 +140,14 @@ credentials" — a hazard serious enough to have its own runbook page
|
||||
(`runbooks/RUNBOOK-manual-guest-restore.md:45-48`, F-OPS). Putting Lane 2 behind the operator is
|
||||
what lets Lane 1 be a button instead of a procedure.
|
||||
|
||||
**[FACT] And as of D5 (v0.188.0) the split is real, not just intended.** Until then Lane 1 only *looked*
|
||||
independent: an app's files sat on the customer's drive but could not be brought back without secrets
|
||||
that lived solely in the guest, so the fast customer lane was silently conditioned on the slow
|
||||
operator lane (§7.1 leg 1). The unit now carries those secrets, so **Lane 1 needs the drive and nothing
|
||||
else** for app rebuild (§7.4). The one honest caveat: Lane 1's *file-merge* surface (Tier-2 additive)
|
||||
still reads its destination from the guest's `settings.json`, so that surface remains guest-conditioned
|
||||
— D5 removed the secrets leg, not the living-app leg.
|
||||
|
||||
---
|
||||
|
||||
## 4. The three-part model (D4)
|
||||
@@ -307,7 +315,15 @@ here — §8 has rows where it is the actual state.
|
||||
|
||||
### 7.1 The dependency graph
|
||||
|
||||
**[FACT]** Every app-tier restore depends on the guest, for one of two reasons:
|
||||
> **[FACT] SUPERSEDED IN PART BY D5 — SHIPPED 2026-07-30 (controller v0.188.0).** Leg 1 below
|
||||
> (**secrets**) is **no longer true for Tier-1 and Tier-2**: the local recovery unit now carries the
|
||||
> portable secret class, so those two tiers need the **drive and nothing else**. The graph and leg 1
|
||||
> are kept as written because they are the model everything downstream was derived from, and because
|
||||
> leg 2 (**the living app**) is unchanged and still binds Tier-2's file leg and Tier-3. Read §7.4 for
|
||||
> the current chain; the corrected rows are 3, 4 and 9 in §8.
|
||||
|
||||
**[FACT — as of the R-108 survey, before D5]** Every app-tier restore depends on the guest, for one of
|
||||
two reasons:
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────┐
|
||||
@@ -357,7 +373,66 @@ here — §8 has rows where it is the actual state.
|
||||
|
||||
**[FACT] So all three app tiers are conditioned on the whole-guest tier**, which is Lane 2. The
|
||||
customer's own recovery lane rests on an operator-only tier, and today nothing tells the customer
|
||||
that.
|
||||
that. — **This sentence was the reason D5 exists. It is now false for Tier-1 and Tier-2; see §7.4.**
|
||||
|
||||
### 7.4 The chain AFTER D5 (controller v0.188.0, 2026-07-30) — the current model
|
||||
|
||||
**[FACT] Tier-1 and Tier-2 no longer depend on the whole-guest tier.** The unit's `compose/app.yaml`
|
||||
(mode 0600) carries the **portable secret class** — every catalog `type: secret` field, i.e. the
|
||||
declared `data_key`s, the 18 DB/root passwords and the internal signing secrets — so neither
|
||||
`encryption.key` nor the guest's `app.yaml` is required to rebuild an app:
|
||||
|
||||
```
|
||||
Tier-1 restore Tier-2 restore Tier-3 reconstitute Tier-3 scratch
|
||||
(rebuild the app) (fill in files) (overwrite + replay) (verification copy)
|
||||
│ │ │ │
|
||||
needs ONLY THE DRIVE needs the app needs the app needs only the
|
||||
(unit carries the running + the deployed + a DB repo password
|
||||
secrets; guest is recorded dest service identifiable
|
||||
consulted only for
|
||||
the withheld class) └────────── still conditioned on the guest (leg 2, unchanged) ──────┘
|
||||
│
|
||||
✅ INDEPENDENT of the whole-guest tier
|
||||
```
|
||||
|
||||
**[FACT] What still needs the guest, precisely** — so this is not read as more than it is:
|
||||
|
||||
1. **Leg 2 is untouched.** Tier-2's *file* restore still reads the destination from the guest's
|
||||
`settings.json` (`tier2_restore.go:114-116`) and Tier-3's reconstitution still refuses when the app
|
||||
is not deployed (`offbox_reconstitute.go:198-201`). D5 removed the **secrets** leg, not the
|
||||
**living-app** leg. So "Tier-2 no longer depends on the guest" is true of the *unit-restore* path it
|
||||
shares with Tier-1, and **false** of its additive file-merge path.
|
||||
2. **The withheld class.** The 7 `type: password` admin logins and `vaultwarden/ADMIN_TOKEN` are
|
||||
deliberately NOT on the drive; they come from the guest, or are regenerated (O4). Their absence
|
||||
costs a credential reset, never data.
|
||||
3. **The fail-closed gate remains.** A `data_key` in neither the unit nor the guest still refuses the
|
||||
restore outright. D5 makes it normally present; it does not soften the gate.
|
||||
|
||||
**[DESIGN] Why plaintext is the right answer here, and what it is coupled to.** Every travelling secret
|
||||
either decrypts data on the *same drive* or authenticates to a container on an internal compose network
|
||||
with no external listener — so possessing it adds nothing to possessing the drive, which is exactly D2's
|
||||
argument for keeping the DATA plaintext. That argument holds **only because** the internet-reachable
|
||||
class is withheld. **The exclusion is what licenses the plaintext; the two must not be relaxed
|
||||
independently.** `stacks.PortableSecretEnvVars` is the single boundary, with a code-level
|
||||
`nonPortableSecrets` register (not a catalog flag — a boundary a catalog push can move is not a
|
||||
boundary, cf. R-97a).
|
||||
|
||||
**[FACT] Precedence, because two sources now exist.** The **unit wins** over the guest. Not
|
||||
"newest wins": the unit's secrets are captured in the same run as the dumps beside them, so the unit's
|
||||
value matches *the data being restored*, whereas the guest's is merely the most recent. Preferring the
|
||||
guest is the data-loss bug — a rotated data key does not decrypt data encrypted with the old one, and a
|
||||
rotated DB password does not match the hash inside the restored data directory.
|
||||
|
||||
**[FACT] Proven live 2026-07-30**, on a scratch drill guest, through the real endpoints: AdventureLog
|
||||
(`SECRET_KEY` data_key + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** —
|
||||
`secrets recovered=2/2`, 27.6 s — and then **the application itself read the seeded customer row over
|
||||
TCP with its own credential**, which is the observable that matters (a restore that returns success onto
|
||||
unreadable data is a failure that looks like a pass). The discriminator held: the pre-backup row came
|
||||
back, a row added *after* the backup was gone, so the volume tar was genuinely restored. There was **no
|
||||
`.sql` dump** in that unit, so the DB came back from the volume tar — the exact case where a regenerated
|
||||
password would have failed. Separately, Grafana's `type: password` admin login was proven **withheld**:
|
||||
the sentinel value was live in the container and encrypted in the guest, and appeared in **0 files**
|
||||
anywhere under the backup namespace. Full record: `audits/D5-drive-alone-restore-2026-07-30.md`.
|
||||
|
||||
### 7.2 A tier whose prerequisites cannot be met in the failure it exists for
|
||||
|
||||
@@ -372,9 +447,21 @@ that.
|
||||
requires the guest's secrets. So offsite alone cannot rebuild an app onto a fresh guest.
|
||||
→ **R-107**
|
||||
|
||||
### 7.3 What D5 would change — and why it was blocked (PRECONDITION NOW MET, 2026-07-30)
|
||||
### 7.3 ~~What D5 would change — and why it was blocked~~ — **D5 SHIPPED 2026-07-30 (v0.188.0)**
|
||||
|
||||
**[DESIGN, TARGET — UNBLOCKED]** The intended fix is to make app secrets travel with the **local**
|
||||
> **This section is history.** D5 is implemented and proven live; the current chain is **§7.4**. Kept
|
||||
> because the reasoning below is why the precondition was required, and because the last paragraph
|
||||
> (`.fab` / R-126) is still open and still not part of D5.
|
||||
>
|
||||
> **One correction to the target as stated below.** It assumed the class that must travel is the
|
||||
> `data_key`-flagged secrets. Part 0 of the D5 task tested that and it did not survive: the flag is
|
||||
> unreliable (four encryption keys the catalog itself labels as such are unflagged → **R-127**), and a
|
||||
> DB password is **not resettable in practice** — `POSTGRES_PASSWORD` is ignored once PGDATA is
|
||||
> non-empty, so a regenerated value leaves the app unable to authenticate against its own restored rows
|
||||
> while the dump replay still reports success. The shipped boundary is therefore *all* `type: secret`
|
||||
> minus the register, not `data_key` alone.
|
||||
|
||||
**[DESIGN, TARGET — SUPERSEDED BY §7.4]** The intended fix is to make app secrets travel with the **local**
|
||||
recovery unit, so Tier-1 and Tier-2 restore work **without the guest and without R**. Offsite
|
||||
already encrypts everything, so secrets travelling offsite would be covered by the escrowed repo
|
||||
password. R would then be required for **offsite recovery and host identity only** — losing R would
|
||||
@@ -416,7 +503,8 @@ Until D5 is actually implemented, §7.1's chain stands as the model.
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 1 | **Customer deletes files** | everything else | „Fájlok visszaállítása" — Tier-2 additive merge | **customer** | **39 s** (6 files) | 24 h | **PROVEN** | CAMPAIGN-9 A1 — byte-identical returns, both non-destruction promises kept, bytes re-served by paperless's own API |
|
||||
| 2 | **An app's data directory is destroyed** | the guest, the other tiers | same route | **customer** | **46 s** (43 files) | 24 h | **PROVEN** | CAMPAIGN-9 A3 — loss verified by a positive observable (doc download 200 → 404), then 16/16 documents usable |
|
||||
| 3 | **An app's DB and named volumes are lost** | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the **only** path that unpacks volume tars) | **customer** | **18.25 s** (path execution) | 24 h | **PARTIAL** | CAMPAIGN-9 A2 proves the path executes and the app survives; content recovery after real loss **never demonstrated** (`CAMPAIGN-9…:825-827`) |
|
||||
| 3 | **An app's DB and named volumes are lost** | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the **only** path that unpacks volume tars) | **customer** | **18.25 s** (path execution) · **27.6 s** (D5 drill, guest `app.yaml` absent) | 24 h | **PROVEN** (content recovery proven 2026-07-30) | CAMPAIGN-9 A2 proved the path executes; **D5 v0.188.0 closed the content gap** — after a restore with the guest's `app.yaml` moved aside, the app read the seeded row **over TCP with its own credential**, the pre-backup row returned and a post-backup row was gone (so the tar was really restored). No `.sql` dump in the unit ⇒ the DB came back from the volume tar |
|
||||
| 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery |
|
||||
| 3b | *same, for a class-B app via Tier-2* | — | **no route** — Tier-2 never reads the unit mirror | — | | | **NONE** | §6.3; **R-102** |
|
||||
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 or 9 of 53 apps); Tier-3 reconstitute for files + DB; **the volume tars in either copy are unreachable** | **customer** (both) | | 24 h | **PARTIAL** | §7.2; **R-102**, **R-107** |
|
||||
| 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go:359-393` |
|
||||
@@ -428,7 +516,7 @@ Until D5 is actually implemented, §7.1's chain stands as the model.
|
||||
| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) |
|
||||
| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) |
|
||||
| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D |
|
||||
| 13 | **Customer loses R** | every byte, every tier, the Recipe | **on-premises recovery is unaffected** — Tier-1/2/3 and the whole-guest tiers all work while the guest and host live. What is lost: the ability to re-establish **host identity** and to open the escrowed offsite key material after a host loss | — | | | **NONE for host-loss** | R exists in **zero** system copies by design (INV Part C, row 24). Whether the operator should be able to recover is **open** → §11-A, §11-B |
|
||||
| 13 | **Customer loses R** | every byte, every tier, the Recipe | **on-premises recovery is unaffected** — Tier-1/2/3 and the whole-guest tiers all work while the guest and host live. What is lost: the ability to re-establish **host identity** and to open the escrowed offsite key material after a host loss | — | | | **NONE for host-loss** | R exists in **zero** system copies by design (INV Part C, row 24). Whether the operator should be able to recover is **open** → §11-A, §11-B. **D5 narrows this further:** Tier-1/2 app recovery needs only the drive, so losing R now costs the offsite route and host identity, never local app recovery (§7.4) |
|
||||
| 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix |
|
||||
| 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** |
|
||||
|
||||
@@ -477,7 +565,13 @@ uplink — **no customer has ever driven a restore**.
|
||||
|
||||
Every divergence between the model above and the system as it is, each with an ID.
|
||||
|
||||
### 10.1 ~~D5 is BLOCKED~~ — CLOSED 2026-07-30 by R-108 (controller v0.187.0)
|
||||
### 10.1 ~~D5 is BLOCKED~~ — precondition CLOSED by R-108 (v0.187.0); **D5 ITSELF SHIPPED (v0.188.0)**
|
||||
|
||||
> **D5 IS DONE — 2026-07-30, controller v0.188.0.** The precondition below was met by R-108, and D5 was
|
||||
> then implemented and proven live the same day: Tier-1/Tier-2 no longer depend on the whole-guest tier,
|
||||
> and a customer needs **the drive and nothing else**. The current recovery chain is **§7.4**; the
|
||||
> corrected matrix rows are 3, 3c and 13. This section's remaining value is the *precondition* analysis,
|
||||
> which is why the table below is kept.
|
||||
|
||||
> **D5's PRECONDITION IS MET.** An app's data namespace can no longer be placed on network storage, so
|
||||
> no `backups/` tree can exist inside FileBrowser's share-root bind, and the browsing surface therefore
|
||||
|
||||
Reference in New Issue
Block a user