a8f790c0c6
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
1407 lines
130 KiB
Markdown
1407 lines
130 KiB
Markdown
# 07 — The recovery model
|
||
|
||
> | | |
|
||
> |---|---|
|
||
> | **Status** | **NOT RATIFIED.** Ratification is Viktor's review, not an editor's. |
|
||
> | **Written** | 2026-07-28 (full rewrite; supersedes the 2026-07-14 DRAFT entirely) |
|
||
> | **Verified against** | controller **v0.183.0** · agent **v0.110.0** · hub **v0.80.0** · catalog `4252121` · repo HEADs `felhom.eu ff050cf`, `felhom-controller fd50a73`, `felhom-agent d5c7691` |
|
||
> | **Live fleet at verification** | demo-felhom + demo-hp, both guest 9201, both on the versions above |
|
||
> | **Freshness** | **CURRENT** as of 2026-07-28. Per standing ruling **S-2**, mark this STALE the moment it is more than a few trains behind. The document it replaces was DRAFT since 2026-07-14, verified against controller v0.132.0 — **fifty-one versions stale** — and was cited as authoritative throughout that time. |
|
||
>
|
||
> **How to read this document.** Two kinds of statement appear, and they are always labelled:
|
||
>
|
||
> - **[DESIGN]** — a decision taken in the 2026-07-28 architecture discussion (D1–D6, §3–§5, §9).
|
||
> Not derived from code; the code may not implement it yet. Where it does not, §10 says so.
|
||
> - **[FACT]** — an observed property, carrying a `file:line`, a live command output, or a citation
|
||
> to `_recovery-inventory-2026-07-28.md` (below: **INV**).
|
||
>
|
||
> Where the model is silent, this document says **OPEN** rather than filling the gap.
|
||
>
|
||
> **Primary input:** `_recovery-inventory-2026-07-28.md` (read-only inventory, 2026-07-28) — cited
|
||
> throughout as **INV Part n**. Every number in §6, §8 and §11 traces back to it.
|
||
|
||
---
|
||
|
||
## 1. Purpose and scope
|
||
|
||
This document describes **how a Felhom customer gets their system back**, and who can do it.
|
||
|
||
It replaces a document that described **where copies are written**. That was the wrong frame, and
|
||
§7 explains why: a set of copies is not a set of recovery routes, and the previous document's tier
|
||
table could be entirely satisfied while a real recovery was impossible.
|
||
|
||
**In scope:** the recovery model — trust boundary, lanes, the three-part model, the tiers as
|
||
*inputs to recovery*, the dependency chain between them, the failure→recovery matrix, and the
|
||
encryption policy.
|
||
|
||
**Out of scope, deliberately:** implementation specs (they live in task specs), the capture-set
|
||
algorithm (`internal/appbackup/captureset.go` and its tests are the source of truth), and per-tier
|
||
operational runbooks (`documentation/runbooks/`).
|
||
|
||
**Authority split.** This document is authoritative for the **failure→recovery matrix** (§8). The
|
||
capability map (`00-capability-map.md`) stays authoritative for **per-capability status**. Neither
|
||
restates the other; §8 rows are cited from the map, not copied into it.
|
||
|
||
---
|
||
|
||
## 2. The trust model
|
||
|
||
**[DESIGN] The operator holds root SSH on every box.** That is a fact of the product — the agent is
|
||
operator-tier, the host is operator-managed, and break-glass exists precisely so the operator can
|
||
get in when nothing else works. **"The operator cannot read customer data" was never the security
|
||
property**, and no part of this document may be read as claiming it.
|
||
|
||
**[FACT]** The mechanics that make this concrete:
|
||
|
||
- The operator's OOB SSH public key is pushed to every box from the hub
|
||
(`hub_settings.oob_operator_ssh_pubkey` → `/var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op`,
|
||
93 bytes, LIVE on both hosts — INV Part C, row 21).
|
||
- The break-glass `root@pam` console password for every host is stored in the hub and retrievable
|
||
with the operator's global key (`documentation/runbooks/break-glass.md:42-48`; LIVE:
|
||
`host_recovery` holds 3 rows — INV Part D2.1).
|
||
- Guest data is reachable from the host by definition: `pct exec`, and the data drives are bind
|
||
mounts on the host (`mp8 /mnt/felhom-drives`, LIVE `pct config 9201` on both hosts).
|
||
|
||
**[DESIGN] What R (the customer's recovery code) does provide, and must keep providing: the hub
|
||
alone is not enough.** A compromised hub yields blobs nobody can open. This property holds **only
|
||
while the operator's own key is not stored in the hub** — which is why §11-A is an open decision and
|
||
why its recommendation on record is "operator key held offline and never in the hub".
|
||
|
||
**[FACT]** The escrow is genuinely zero-knowledge today: `host_escrow` rows carry
|
||
`posture = zero_knowledge`, a 383-byte blob and a 572-byte identity blob, and the hub holds only a
|
||
`restic_pw_sha256` **hash** beside them (LIVE, both hosts — INV Part D2.1). R exists in **zero**
|
||
system copies by design (INV Part C, row 24).
|
||
|
||
So the honest statement of the property is:
|
||
|
||
> **The hub cannot open the escrow. The operator can reach a live box. Neither fact substitutes for
|
||
> the other, and R is what keeps the first true.**
|
||
|
||
**[DESIGN] Because R is the only key, the household is ASKED for it from the first login** (controller
|
||
v0.245.0, R-543). Off-site backup is enabled by default but does not RUN until the ceremony is done,
|
||
so the ask is not a nicety — it is the step that turns the default-on tier into an actual copy. The
|
||
volunteer guide asks for it immediately after the dashboard password and before the first app
|
||
(`runbooks/VOLUNTEER-first-hour.md` §6), and the product repeats the ask on every page until it is
|
||
done (§6.1).
|
||
|
||
---
|
||
|
||
## 3. The two lanes (D1)
|
||
|
||
**[DESIGN] Recovery is split into two lanes with different owners. This is a product decision, not a
|
||
limitation.**
|
||
|
||
### Lane 1 — the customer, unassisted: files and app data
|
||
|
||
The customer owns their data and can get it back themselves, through the „Visszaállítás" surfaces,
|
||
with nothing but their dashboard password. No operator, no ticket, no scheduling.
|
||
|
||
**[FACT] What Lane 1 contains today** (INV Part A.1 — seven paths, all behind the controller's
|
||
`RequireAuth` gate, which is the **customer-owned** password: `internal/web/auth.go:34-35` puts
|
||
`settings.json → password_hash` ahead of the operator-provisioned `controller.yaml` value):
|
||
|
||
| # | Surface (HU) | Endpoint | Semantics |
|
||
|---|---|---|---|
|
||
| 1 | „Visszaállítás indítása" | `POST /backup/restore` | destructive — rebuilds the app from its Tier-1 unit |
|
||
| 2 | „Fájlok visszaállítása" | `POST /backup/tier2/restore` | additive, missing-only |
|
||
| 3 | „Visszaállítás a távoli tárolóból" | `POST /backup/offbox/restore` | non-destructive, to a verification copy |
|
||
| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. `audits/SPIKE-restic-restore-test-2026-08-31.md`. |
|
||
| 5 | „Teljes visszaállítás (fájlok + adatbázis)" | `POST /backup/offbox/reconstitute` | destructive to files + DB, never deleting |
|
||
| 6 | „Megosztások visszaállítása" + place | `POST /backup/shares/{restore,place}` | additive |
|
||
| 7 | `.fab` import | `POST` → `apiImportStart` | destructive re-import of one app |
|
||
|
||
**[FACT] What is proven in Lane 1** (INV Part G.1): paths 2 and 6 are proven live end-to-end;
|
||
path 5 is proven live through a destructive drill (photos deleted, trash emptied, restored, timeline
|
||
confirmed); path 3 is proven for **bytes** only; path 7 is proven for the drive-to-drive circle but
|
||
its browser-upload leg is not; path 1 is proven to **execute** but its **content recovery after real
|
||
loss has never been demonstrated** — the single most valuable unproven item in the system
|
||
(`audits/CAMPAIGN-9-restore-proof-2026-07-28.md:825-827`); path 4 has never been exercised as a
|
||
distinct action (`audits/CAMPAIGN-8-backup-restore-2026-07-27.md:520`).
|
||
|
||
**[FACT] One honest residual on the whole lane:** every proven run was performed by the **operator**,
|
||
not by a customer. `00-capability-map.md:75` still carries "A customer (not the operator) performs a
|
||
restore via UI alone" as **MISSING as evidence** — by absence of the run, not by a product gap.
|
||
|
||
**[RULING 2026-09-17, operator — a design choice REVERSED] The restore record survives a restart
|
||
(R-550, controller v0.246.0).** The async restore status (`internal/backup/opstatus.go`) was in memory
|
||
by choice — „same precedent as notification cooldowns". Chaos night round 10 measured the cost: a
|
||
restore accepted, the box hard-reset four seconds later, and the status answering the Go zero value, so
|
||
the household could never learn whether it finished. **For the restore record only**, it is now written
|
||
atomically to `restore-status.json` in the controller's state directory at both ends of an op; at
|
||
startup a record still marked running becomes a failed, interrupted result („A visszaállítás megszakadt
|
||
(a doboz újraindult) — indítsd el újra."), shown per app on the restore page until that app's next
|
||
restore, and raised once as `restore_interrupted`. **Notification cooldowns stay in memory** — the
|
||
precedent is kept for what it was written for.
|
||
|
||
### Lane 2 — the operator: guest and host recovery
|
||
|
||
Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are **operator
|
||
work, by design and by contract**. They are not customer-facing and are not going to be.
|
||
|
||
**[FACT] What Lane 2 contains** (INV Part A.2): the scheduled restore-test, `pct restore`, raw
|
||
`proxmox-backup-client restore`, raw `restic restore`, the agent's DR bring-up
|
||
(`--selftest=bring-up -mode dr`), the host-loss plan builder, and the escrow-consume ceremony. Each
|
||
needs root on the host or a CLI flag; none is reachable from any customer surface.
|
||
|
||
**[FACT] What is proven in Lane 2** (INV Parts G.1, G.3): whole-guest `pct restore` from both tiers
|
||
is proven with exact mount parity; the unattended restore-test is proven and currently running on
|
||
both boxes; a corrupted snapshot is proven to fail cleanly. **The DR bring-up path has never been
|
||
executed** (`CAMPAIGN-8…:522`), the host-loss plan **executes nothing by construction**
|
||
(`felhom-agent/internal/dr/plan.go:1-4`), and **no host has ever been rebuilt as its former self**
|
||
(INV Part D1).
|
||
|
||
### Lane 2's restore-test is scheduled PER ARCHIVE GENERATION (R-86, 2026-08-03)
|
||
|
||
**[CONTRACT, changed 2026-08-03 — agent v0.121.0 + hub v0.91.0.]** The scheduled restore-test used to
|
||
fire on an interval started at daemon start. It no longer does. The rule is:
|
||
|
||
> Let **A** be the newest archive on a tier that has settled for at least the settle lag (24 h).
|
||
> The tier is **DUE** when **A** exists and **A has not already been proven**.
|
||
|
||
So a tier is proved **once per archive**, on its own archive, and the proof follows the backup rather
|
||
than the process's uptime:
|
||
|
||
| tier rhythm | what is proved, and when |
|
||
|---|---|
|
||
| daily (host tier) | yesterday's archive, once a day |
|
||
| weekly (offsite tier) | last week's archive, once a week |
|
||
| newborn (no archive yet) | nothing — **UNKNOWN, never a fault** |
|
||
|
||
**The trap in the obvious formulation, recorded so it is not reintroduced:** *"due when the newest
|
||
archive is ≥ 24 h old"* is never true on a **daily** tier — a new archive resets the newest-archive
|
||
age to zero long before it reaches the lag — so the literal reading silently switches restore-testing
|
||
off for the tier that matters most.
|
||
|
||
What survives unchanged: the restore-test itself (restore → boot → verify → destroy the scratch), its
|
||
journal and crash recovery, the scratch VMID band, the one-heavy-operation gate, proof credit only on
|
||
success, and oldest-proven ordering, which is now the tie-break **between due tiers**. A ticker
|
||
remains, but only as the **evaluation interval** (6 h by default, chosen from a measured cost: one
|
||
due-check is 18 ms on a local dir storage and 392 ms on the PBS tier over the WAN).
|
||
|
||
**The hub's half is not optional.** `restoreProvenStaleAfter` was a flat 7 days derived from the very
|
||
cadence this replaced, and a weekly tier proved weekly reaches a proof age of **exactly** one interval
|
||
just before its next proof — 168 h against a 168 h window. It sat ON the line, so any ordinary delay
|
||
tipped a healthy tier into a nightly alarm. The window is now per tier, from that tier's observed
|
||
archive interval, floored at the old 7 days, capped at 12 days (strictly inside the two-week offsite
|
||
retention), and falling back to the tier's declared rhythm when history is too short to observe one.
|
||
|
||
### Why the split is right, stated once
|
||
|
||
**[DESIGN]** A customer can reason about "my photos are gone". A customer cannot reason about
|
||
"restore the LXC and re-attach mp8 by durable_id without binding another guest's bootstrap
|
||
credentials" — a hazard serious enough to have its own runbook page
|
||
(`runbooks/RUNBOOK-manual-guest-restore.md:45-48`, F-OPS). Putting Lane 2 behind the operator is
|
||
what lets Lane 1 be a button instead of a procedure.
|
||
|
||
**[FACT] And as of D5 (v0.188.0) the split is real, not just intended.** Until then Lane 1 only *looked*
|
||
independent: an app's files sat on the customer's drive but could not be brought back without secrets
|
||
that lived solely in the guest, so the fast customer lane was silently conditioned on the slow
|
||
operator lane (§7.1 leg 1). The unit now carries those secrets, so **Lane 1 needs the drive and nothing
|
||
else** for app rebuild (§7.4). The one honest caveat: Lane 1's *file-merge* surface (Tier-2 additive)
|
||
still reads its destination from the guest's `settings.json`, so that surface remains guest-conditioned
|
||
— D5 removed the secrets leg, not the living-app leg.
|
||
|
||
---
|
||
|
||
## 4. The three-part model (D4)
|
||
|
||
**[DESIGN] Recovery needs three things, and they live in three different places.** Losing one is a
|
||
different problem from losing another, and the matrix in §8 is organised around that.
|
||
|
||
| Part | Holds | Where | Lost when |
|
||
|---|---|---|---|
|
||
| **Recipe** | scaffolding — guest sizing, drive inventory, PVE storage definitions, PBS coordinates, app inventory + storage bindings. **No secrets, no bytes.** | the hub | the hub is lost |
|
||
| **Escrow** | the host identity key material and the restic repo password | the hub, R-wrapped | the hub is lost, **or** R is lost |
|
||
| **Bytes** | Tier-1 / Tier-2 / Tier-3 / whole-guest | the tiers | the relevant medium is lost |
|
||
|
||
**[FACT] The Recipe exists and is current.** `dr_recipe` holds 6 rows; demo-felhom's was updated
|
||
`2026-07-28 17:31:13` and demo-hp's `17:25:31` — i.e. within one report cycle (LIVE, INV Part D2).
|
||
It carries guest sizing, `pve_storage[]`, and an app half with per-app `storage_bindings`.
|
||
|
||
**[FACT] The Escrow exists and is zero-knowledge.** `host_escrow`: 2 rows, `posture` `zero_knowledge`;
|
||
the identity bundle's shape is `{tunnel_token, pbs_token, wg_private_key, restic_repo_password}`
|
||
(`felhom-agent/internal/escrow/identity.go:26-39`).
|
||
|
||
**[FACT] Three parts of the Recipe are empty or wrong on the live fleet**, and they are exactly the
|
||
parts a host-loss recovery would read (INV Part D2.3): `hosts.dr_record_json` is `{}` on all three
|
||
hosts; `host_escrow.directive_json` is `{}` on both escrowed hosts; `dr_recipe.host_half.drives` is
|
||
`[]` on every customer including two with enrolled data drives; and `dr_recipe.host_half.pbs.namespace`
|
||
reads `"root"` while the real namespaces are `demo-felhom` / `demo-hp`. → **R-105**, **R-106**.
|
||
|
||
---
|
||
|
||
## 5. Encryption policy (D2)
|
||
|
||
**[DESIGN] Encryption follows the boundary, not the tier.**
|
||
|
||
### On the customer's own premises: plaintext
|
||
|
||
Tier-1 recovery units, Tier-2 cross-drive copies, the app data on the drives, and the local vzdump
|
||
archives on the host are **unencrypted, deliberately**.
|
||
|
||
The reasoning, recorded so nobody "hardens" this later:
|
||
|
||
1. **It buys nothing against a real threat.** Someone who can take the second drive can take the
|
||
first. Local encryption defends against a threat model — physical theft of *only* the backup
|
||
medium — that does not describe a home server where both media sit in the same box.
|
||
2. **It adds a key-loss path that turns a working backup into a brick.** Every local encryption key
|
||
is one more thing that must survive the disaster it exists for, and the system already has one
|
||
such dependency it is trying to reduce (§7).
|
||
3. **It would break browsing, which is a feature.** FileBrowser and SMB let the household see and
|
||
use their own files. An encrypted local copy is not browsable, and the „Megosztás" and
|
||
FileBrowser surfaces are part of the product, not an accident.
|
||
|
||
**[FACT]** The local plaintext posture is real and verifiable: `dir: local` in `/etc/pve/storage.cfg`
|
||
carries no `encryption-key` (LIVE, both hosts), so the daily whole-guest archive is a plain
|
||
`.tar.zst` — 5.93 GB on demo-felhom, 1.68 GB on demo-hp (LIVE, INV Part B.3). **That archive is the
|
||
single highest-value object on the box**: it contains `encryption.key`, the restic `repo_password`,
|
||
the offbox `ssh_key`, `settings.json` and `controller.yaml` (all under `/var/lib/docker` = `mp0`,
|
||
`backup=1`). Naming that plainly is part of the policy, not an argument against it.
|
||
|
||
### Leaving the premises: encrypted, and the provider must not be able to read it
|
||
|
||
**[FACT]** Both offsite tiers encrypt client-side:
|
||
|
||
- **restic (Tier-3):** repo password is a 256-bit hex value generated once on the box and never
|
||
logged (`internal/backup/offbox.go:392-410`); restic is invoked with
|
||
`RESTIC_PASSWORD_FILE` (`:608`).
|
||
- **PBS (whole-guest offsite):** `/etc/pve/storage.cfg` carries a **per-customer** `encryption-key`
|
||
and the fingerprints differ between demo-felhom and demo-hp (LIVE). The live restore command line
|
||
observed on demo-felhom carries `--crypt-mode=encrypt` (INV Part A.2.3). Per-tenant encryption is
|
||
why cross-customer dedup is impossible on the shared datastore — a cost consequence, recorded, not
|
||
a defect.
|
||
|
||
### The one exception, stated because it is not covered by either rule
|
||
|
||
**[FACT]** The `.fab` bundle writes `app.yaml` with **decrypted plaintext secrets, deliberately**,
|
||
and its password is **optional** (`internal/appexport/export.go:484,506-511`; the generated file's
|
||
first line is literally `# Exported by felhom-controller — plaintext secrets`, `:531`;
|
||
`Encrypted: req.Password != ""`, `:307`). A `.fab` can also be written to a registered **network**
|
||
share, because `storageDriveList()` does not filter network paths
|
||
(`internal/web/handler_export.go:377-387`). This is a portability artifact, not a tier — but it is
|
||
the one place where a customer action can put every secret of one app onto a NAS in plaintext.
|
||
Recorded here so the encryption policy is not read as covering it. → **R-108** (same root cause).
|
||
|
||
---
|
||
|
||
## 6. The tiers — what each captures
|
||
|
||
The tiers are **inputs to recovery**, not recovery routes. §7 and §8 say what they can actually do.
|
||
|
||
**[RULING 2026-09-16, operator] Tier 3 (off-site) is ON for every new customer — shared, 100 GB soft
|
||
quota.** Opting a customer out stays the per-customer exception, as DR-tier-by-default has been since
|
||
2026-07-12. **The reason is this document's own [FACT] two sections down:** the whole-guest tiers do
|
||
not carry `mp8 /mnt/felhom-drives`, and a Tier-1 unit has no file leg — so for the four class-A apps
|
||
a one-drive box with Tier 3 off keeps the household's files in NO tier at all. That was measured on a
|
||
fresh box the same day: five photos put into Nextcloud, deleted, restored from the box's own backup,
|
||
and none of them opened (`audits/DRILL-prove-fixes-0243-2026-09-16.md`, R-537/R-538). Implemented in
|
||
hub v0.116.0 (`handleConfigNewForm`). **What this ruling does NOT change: what any tier captures.**
|
||
|
||
|
||
**The app update's safety precondition accepts ANY tier** (operator ruling 2026-09-13, controller
|
||
v0.239.0, R-475): the first fresh copy in the order Tier 2, Tier 1, Tier 3; with none, it backs up
|
||
first. Design: `09-update-architecture.md` §3 decision 8. **What a Tier-1 route back restores is only
|
||
what the unit holds** — for an app whose data is a bind mount that is the definition, not the data
|
||
(R-479). **Removal and the tiers (controller v0.240.0):** removing an app with its backups KEPT keeps
|
||
the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás" still works afterwards
|
||
(R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and
|
||
never touches off-site snapshots (R-474). A removed app whose unit was kept is listed on both local backup pages with its restore since controller v0.242.0 (R-487): **the local lists are keyed on the drives, not on what is deployed** — the rule R-237 set for the off-site list — and the restore opens the unit where it sits, a data drive included. **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.**
|
||
|
||
**[FACT, measured from source 2026-09-24, controller v0.267.0/v0.268.0; the second-drive cell for file apps changed by v0.269.0] Which copy brings an app back WHOLE
|
||
— read from the restores' own refusals, not from what the tier stores** (`09` §3 decision 25, R-659):
|
||
|
||
| app | own unit (Tier 1) | second drive (Tier 2) | off-site (Tier 3) |
|
||
|---|---|---|---|
|
||
| no declared drive files (class B, 45 apps) | whole — the unit restore | whole — „Teljes visszaállítás" (the unit restore) | whole — the full restore |
|
||
| declared drive files (`DeclaredDriveFileLegs`: calibre-web, immich, nextcloud, paperless-ngx) | **not** — refused (R-538) | **whole since controller v0.269.0** — „Teljes visszaállítás" runs `RestoreTier2Whole`: the mirror's files by four rules (never delete a file; never write an older file over a newer one; bring back every missing one; an older, different live file is replaced and KEPT beside as `<name>.felhom-<UTC>`), then the unit from the mirror (decision 26, R-661; proven live 2026-09-24, `audits/night-2026-09-24/A1/`). Needs a proven, openable mirror WITH the file legs and 2 GB free; `rsyncMirror` (`--delete`) is never used for it. The unit-only restore is still refused (R-538). | whole — „Teljes visszaállítás (fájlok + adatbázis)" |
|
||
|
||
`backup.WholeOnTier` is this table; `TestR659_TruthTableAgreesWithTheRestoresRefusal` pins that it cannot drift
|
||
from the refusal. An app with only OPTIONAL legs (audiobookshelf, komga, romm) is class-B here: its files are
|
||
never moved by an update or an undo, and the unit restore accepts it.
|
||
|
||
### 6.1 The four tiers, as configured on the live fleet
|
||
|
||
| Tier | Location | Captures | Cadence (LIVE) | Retention (LIVE) | Encrypted |
|
||
|---|---|---|---|---|---|
|
||
| **Tier-1** recovery unit | `<nsRoot>/backups/primary/<app>/` on the app's own drive | compose + `.felhom.yml` + **secret-stripped** `app.yaml`, `db-dumps/*.sql`, `volume-dumps/*.tar`, `manifest.json` | nightly at W, plus a checksum-gated refresh on the periodic status pass | one point per app (`restore_points.go:14-18`) | no |
|
||
| **Tier-2** cross-drive | `<target nsRoot>/backups/secondary/<app>/` | **always** a full mirror of the unit (`tier2.go:368-369`) **plus** `mandatory + optional` file legs, v2 layout `hdd/<rel>` + `userdata/<rel>` | nightly at W+60m | mirror (rsync) | no |
|
||
| **Tier-3** restic offsite | Hetzner Storage Box over SFTP, one multi-path snapshot per app per run | the unit **plus** `mandatory` legs only; a separate `_shares` snapshot | nightly at W+105m | `--keep-daily 7 --keep-weekly 4 --keep-monthly 6` | **yes** (restic) |
|
||
| **Plane-2** whole-guest, local | `local:` → `/var/lib/vz/dump` on the host | rootfs + `mp0 /var/lib/docker` + `mp1 /mnt/sys_drive` | **24 h** (`backup_cadence_seconds: 0`) | `local_backup_retention: 3` | **no** |
|
||
| **Plane-2** whole-guest, offsite | `felhom-pbs:` → ep0 datastore `felhom-offsite`, per-customer namespace, over WireGuard | same contents | **7 days** (`604800`) | server-side prune on ep0, `keep-last 2` at `03:30` — **and the box asks for none** (R-191) | **yes** (per-customer key) |
|
||
|
||
> **[FACT] Tier-3 has a fifth state the table above does not show: PAUSED (R-543, controller
|
||
> v0.245.0).** Off-site is ON by default from hub v0.116.0, and a run does not start until the
|
||
> household has performed the escrow ceremony — `tier3State` calls this `escrow_pending` and the page
|
||
> says „Kulcsletétre vár". **This is the design, not a defect:** the escrow is zero-knowledge (§2),
|
||
> the household's recovery code is the only key, and a run started without one would write a copy
|
||
> nobody could ever open. What was wrong until v0.245.0 is that **nothing asked the household for the
|
||
> code**, so a fresh box could sit paused indefinitely while its Tier-1 row promised that the off-site
|
||
> copy protected the app's files. Since v0.245.0 every dashboard page carries the reminder (the R-241
|
||
> bar, second instance) and the Tier-1 sentence renders by state — „védené … szünetel" while paused.
|
||
> Measured on a fresh box 2026-09-16: zero snapshots, and the page said the files were protected.
|
||
|
||
> **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.
|
||
|
||
|
||
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
|
||
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
|
||
(`cmd/controller/main.go:604-607`). Both boxes run W = `02:30`.
|
||
|
||
**[FACT, controller v0.271.0, 2026-09-25] A fourth leg: automatic app updates, CHAINED after the off-site
|
||
leg** (`09` §3 decisions 11 and 20, §6.4 part 7). The `offbox-backup` job runs the update leg when its off-site
|
||
half ends, on every path (no target, ok, failed, panicked); a box without an off-site tier runs it at W+105m. The
|
||
leg starts no step at or after **W+5h**, and the whole-guest gate — open [W+2h, W+6h) — **defers while the leg
|
||
runs until W+5h**, then only for a step already in flight, never past W+5h30m (decision 31). So on an update
|
||
night the whole-guest backup starts later inside its own window; it is never skipped for an update. Live
|
||
numbers: `audits/DRILL-night-2026-09-25.md` Part D.
|
||
|
||
**[FACT, 2026-09-25] demo-hp's local whole-guest target is its host ROOT disk; retention there is now
|
||
keep-last=1** (operator ruling 2026-09-24, option A; R-684). PVE prunes AFTER a successful backup, so the target
|
||
must hold keep-last + 1 archives during the run — measured: a 9201 archive is 8.18 GB, the root disk 39 GB. The
|
||
warning that should precede such a failure is R-685.
|
||
|
||
**[FACT] What the whole-guest tiers do NOT carry.** `mp8 /mnt/felhom-drives` and
|
||
`mp9 /etc/felhom-bootstrap` are **host bind mounts** and are out of `vzdump` scope entirely (LIVE
|
||
`pct config 9201`, both hosts). So the whole-guest tiers carry the guest and **none of the customer's
|
||
data drives** — 916 GB on demo-felhom, 938 GB on demo-hp.
|
||
|
||
### 6.2 Coverage per app class — and an unresolved count
|
||
|
||
**[FACT]** Of 53 catalog templates, **52 keep data in Docker named volumes**; exactly **13** carry a
|
||
`backup:` block, and those 13 are exactly the templates that bind `${HDD_PATH}` / `${USERDATA_PATH}` /
|
||
`${IMPORT_PATH}` at all; **14** have a database service (INV Part B.1, computed against catalog
|
||
`4252121`).
|
||
|
||
Applying the classifier's documented two-level default (`internal/appbackup/classify.go`: explicit
|
||
entry wins; else a `:ro` reader is *excluded*; else a writable bind is *mandatory*):
|
||
|
||
| tier filter | templates with at least one file leg |
|
||
|---|---|
|
||
| **Tier-3** (`mandatory` only) | **4** — calibre-web, immich, nextcloud, paperless-ngx |
|
||
| **Tier-2** (`mandatory + optional`) | **7** — the 4 above plus audiobookshelf, komga, romm |
|
||
| legacy resolver path (apps with no block) | **0** — no non-block template binds a namespace path |
|
||
|
||
> ### ✔ RESOLVED 2026-08-31 (controller v0.229.0) — **A = 7 · B = 45 · C = 1**
|
||
>
|
||
> | source | count | verdict |
|
||
> |---|---|---|
|
||
> | C9-F1 Phase 0, as shipped | **A = 9 / B = 43 / C = 1** — `felhom.eu/REPORT.md:17-24` (since overwritten), restated `backlog/OPEN-ITEMS.md:31` | **WRONG by two apps** |
|
||
> | INV Part B.1, independent enumeration at catalog `4252121` | **A = 7 / B = 45 / C = 1** | **CORRECT** |
|
||
> | Measured at catalogue `459766cb16395fd1d1a66282f5cc6da59ead5924`, 2026-08-31 | **A = 7 / B = 45 / C = 1** | adopted |
|
||
>
|
||
> **The method, so it can be re-run rather than re-argued.** A throwaway `main` inside the controller
|
||
> module drove the PRODUCTION rule over all 53 template directories — `stacks.LoadMetadata` (the single
|
||
> validation choke point, so a rejected `backup:` block degrades to legacy exactly as it does live) →
|
||
> `stacks.ParseComposeClassifiableBinds` → `appbackup.ClassifyBinds` → `appbackup.ComputeCaptureSet` at
|
||
> `TierSecondary`, with the legacy branch falling back to `AppDataBindsPresent` + `AppDataDirNames` as
|
||
> `backup.tier2CaptureSet` does. **A = at least one leg survives that pipeline.** 13 templates carry a
|
||
> valid `backup:` block; 40 are legacy and none of them binds a namespace path, so all 40 are B or C.
|
||
>
|
||
> **A (7):** audiobookshelf, calibre-web, immich, komga, nextcloud, paperless-ngx, romm.
|
||
> **C (1):** bentopdf — it declares no `volumes:` key and no `${…_PATH}` bind at all.
|
||
> **B (45):** everything else.
|
||
>
|
||
> **How the earlier disagreement arose, established rather than guessed.** Phase 0's own write-up
|
||
> (controller `CHANGELOG.md`, v0.183.0) says four apps — plex, jellyfin, emby, navidrome — are in B
|
||
> "only because their single bind is a `:ro` media mount, which `ClassifyBinds` correctly excludes". It
|
||
> applied the `:ro` **default** rule. The two apps it therefore missed are **radarr and sonarr**: their
|
||
> `${USERDATA_PATH}/media/*` and `${USERDATA_PATH}/downloads` binds are **writable**, so the `:ro` rule
|
||
> does not reach them, and they are excluded by an **explicit** `class: excluded` entry instead. 9 − 2 =
|
||
> 7 and 43 + 2 = 45, which is exactly the gap. No catalogue file was changed; this is a measurement.
|
||
|
||
**[FACT] The class-B consequence is real regardless of which count is right.** For an app whose data
|
||
lives entirely in named volumes, the Tier-2 copy holds a full `recovery-unit/` and **no readable
|
||
leg** — LIVE on demo-felhom, `backups/secondary/bookstack/` and `.../docmost/` contain
|
||
`recovery-unit` and nothing else, at 156 MB and 86 MB (INV Part B.2). Since v0.183.0 the button
|
||
refuses **before** stopping the app and names the action that works.
|
||
|
||
### 6.3 Where capture and restore disagree
|
||
|
||
**[FACT]** Three asymmetries, each source-cited:
|
||
|
||
| tier | captured | read back by that tier's restore | gap |
|
||
|---|---|---|---|
|
||
| Tier-1 | unit incl. volume tars + DB dumps | all of it | none |
|
||
| Tier-2 | unit mirror **+** file legs | the file restore reads `hdd/` and `userdata/`; **since controller v0.218.0's Tier-3 sibling and now v0.229.0, a SECOND action reads the unit mirror itself** (`RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`, `tier2_restore.go`) | **CLOSED — R-102** |
|
||
| Tier-3 | unit (incl. volume tars) + mandatory legs | files + DB replay **+ the named-volume tars, replayed from the scratch unit** (`offbox_reconstitute.go` `volReplay`, controller **v0.218.0**); the unit itself is still **skipped** on the way to live (`offbox_reconstitute.go:284-289`; placed only if the live unit is absent, `offbox_restore.go:352-356`) | **CLOSED — R-107** |
|
||
|
||
**[FACT] 2026-08-22 — the Tier-3 row above was corrected; the Tier-2 row was NOT.** Until controller
|
||
**v0.218.0** this table said *"no offsite action unpacks the named-volume tars it captures"*, and that
|
||
was true from the day Tier-3 shipped until 2026-08-21. **R-107 closed in v0.218.0**: `volReplay`
|
||
(`offbox_reconstitute.go`) replays the scratch unit's `volume-dumps/` into the live named volumes,
|
||
proven live on `demo-hp`. The old sentence is kept here, in the past tense, because a correction that
|
||
erases what was believed leaves the next reader no way to tell a fixed gap from one that was never
|
||
noticed.
|
||
|
||
**[FACT] 2026-08-31 — the Tier-2 half is now closed too (R-102, controller v0.229.0), and the old
|
||
sentence is kept here in the past tense for the reason the paragraph above gives.** Until v0.229.0 this
|
||
table said of Tier-2: *"the unit mirror is read by nothing — `RecoveryUnitPath` resolves to
|
||
`backups/primary/` (`appbackup/paths.go:46-48`)"*, and that was true from the day Tier-2 shipped until
|
||
2026-08-31. The mechanism was a hard-coded `primary` segment: every reader of a recovery unit could
|
||
only name a path under it. `appbackup` now also exposes four **unit-directory-relative** primitives,
|
||
`Manager.RestoreFromRecoveryUnitAt(stack, unitDir)` holds the restore body, and `RestoreTier2Unit`
|
||
points it at `<dest>/backups/secondary/<app>/recovery-unit/`. Proven live on `demo-hp` **with the
|
||
primary unit moved aside** — `audits/DRILL-r102-tier2-unit-2026-08-31/`.
|
||
|
||
**Two things did NOT change, and both are load-bearing.** THE SOURCE MOVED; THE DESTINATION DID NOT —
|
||
data still lands in the live named volumes and the live database container. And the FILE restore's
|
||
reach is unchanged: `CanRestore()` still answers only "are there file legs?", and
|
||
`tier2UnitNotCoveredMsg` is still appended where that restore runs, so a clean file result never reads
|
||
as a clean bill of health for the database.
|
||
|
||
**[FACT]** Tier-2's gap was the sharper one because of *when* it bit: Tier-2 exists for the case
|
||
where the primary drive is lost — and in exactly that case the primary unit is gone while this
|
||
mirror survives on the second drive. Until v0.229.0 it was unreachable by any customer action.
|
||
|
||
**[FACT] 2026-08-23 — taking the undo copy used to DESTROY the app's own database backup (R-361).**
|
||
`writeSafetyDump` called `DumpOne` into the app's own unit directory and renamed the result to
|
||
`pre-restore-*` afterwards. `DumpOne` writes `<stack>-<dbtype>.sql` — the app's canonical dump, the
|
||
name the replay loop matches exactly — so **every safety dump overwrote the app's real backup and then
|
||
moved it away**, leaving the app with no database backup of its own until the next nightly run. A local
|
||
restore-from-unit in that window finds no `.sql` and tells the customer the app never had a database.
|
||
The comment beside it asserted the rename meant it *"can never overwrite the app's real dump"*; it was
|
||
false as written and stood for four months. **Measured before the fix on `demo-hp` 2026-08-22:**
|
||
`docmost` and `bookstack` each held only `pre-restore-*` files and no canonical dump.
|
||
|
||
Fixed in controller **v0.221.0**: `DumpOneTo` takes an explicit final path and derives its own `.tmp`
|
||
from it, so neither the destination nor the scratch file can collide with a nightly dump running
|
||
beside it. **Proven the only way it can be** — the canonical dump's sha256, unchanged across a
|
||
restore, on both engines.
|
||
|
||
**[DESIGN] 2026-08-23 — `db_dumps` lists the app's OWN dumps, not the undo copies.** They are local
|
||
material for a restore that went wrong, not part of the app's recovery set: nothing reads them from
|
||
the manifest, and three per app were being pushed off-site permanently for no recovery value. The
|
||
files are neither deleted nor hidden — their visibility is a separate recorded decision and it stands.
|
||
**One consequence, recorded because it bit within minutes:** a stable `db_dumps` lets
|
||
`CaptureRecoveryUnit`'s already-current early return fire, so anything that must happen on every
|
||
capture — such as bounding the undo copies — has to sit ABOVE that check, not after it.
|
||
|
||
**[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.**
|
||
Recorded here rather than only in a closed register row, because a decision that survives only inside
|
||
a closed work item is a decision nobody will find.
|
||
|
||
1. **Replay.** The snapshot's `.sql` is imported into the app's database, with only the DB service up
|
||
(R-47). A pre-restore copy of the LIVE database was taken first and is on disk (R-43); the restore
|
||
refuses outright if it could not be taken.
|
||
2. **Rollback (controller v0.220.0, R-379).** If the replay fails, the product re-applies that undo
|
||
copy itself. **The whole set for this run** — an app with two databases gets two undo files, and
|
||
restoring only the first would leave the other half-written — matched on **the run's own stamp**,
|
||
never on the `pre-restore-` prefix, because several runs' copies coexist in the same directory. It
|
||
runs with the DB service still up and before any restart, so the app never observes the half
|
||
state, and into a **re-discovered** container: the DB-only start re-creates it, so the id captured
|
||
at dump time is dead by rollback time (v0.220.1, found by the first live run). The app then starts
|
||
and the customer is told **both** that the restore failed and that their data is as it was.
|
||
3. **Hold (v0.220.0, operator ruling 2026-08-22).** If the rollback ALSO fails, the app is **held
|
||
stopped**, not started. A running app on a half-written database lets the customer type into it and
|
||
makes the damage permanent. The hold is persisted, every start path refuses it with a reason and a
|
||
route, the app-stop marker is ended so nothing auto-restarts it at the next boot, and the app reads
|
||
**red** rather than green. An operator clears it with `--clear-restore-hold`, which requires a
|
||
controller restart.
|
||
|
||
**Why a rollback and not an engine flag.** Postgres gained `--single-transaction` in the same release
|
||
and that does make its replay all-or-nothing — but **MariaDB's DDL is not transactional**, so a
|
||
partial apply there is unavoidable at the engine. Measured 2026-08-22: the same truncated dump left
|
||
Postgres emptied and crash-looping, and left MariaDB with its user data intact, its schema-version
|
||
table wiped to zero rows, and the app reporting `health=healthy, running=true, restarts=0`. The flag
|
||
is a belt; the rollback is the fix. **R-379, R-380.** Evidence:
|
||
`audits/DRILL-r379-rollback-2026-08-22/`.
|
||
|
||
**[DESIGN] 2026-08-22 — the restore destination is resolved by the same rule as the capture
|
||
destination.** The drive if the app declares one (`HDD_PATH`), the system data path otherwise —
|
||
`Manager.GetAppDrivePath`, one expression, used by `CaptureRecoveryUnit` and, since controller
|
||
**v0.219.0**, by `ReconstituteFromOffsite` and `PlaceOffsiteRestore` too.
|
||
|
||
**[FACT] 2026-09-01 — the rule now has a FOURTH consumer and the "one expression" sentence is
|
||
true (R-414, controller v0.232.0).** `offboxRestoreScratchDir` — where a restore is UNPACKED, as
|
||
distinct from where it LANDS — never consulted `systemDataPath`, so on a box with no registered
|
||
storage path it refused, and the nightly proof could not run at all. **It was not excluded on
|
||
purpose:** R-356's own commit (`08eb1a6`) records in its tests that *"the prepared scratch still
|
||
resolves to the registered storage path … only the DESTINATION moves"* — it was out of scope, and
|
||
every fixture assumed a registered path exists. The one comment about a `systemDataPath` fallback
|
||
belonged to `PlaceOffsiteRestore`, concerned bulk **userdata**, and R-356 deliberately overruled it.
|
||
|
||
**But the fallback is SCOPED, and the scoping is the point.** The two callers ask different
|
||
questions, and one predicate answering both is R-356's own defect: a **unit-only** restore (the
|
||
R-87 proof) may fall back to the system data path, because §7 records as `[FACT]` that a driveless
|
||
app's unit already lives there indefinitely and that the same-device placement is *"intended, not
|
||
a defect"*; a **full** restore keeps the R-252 refusal, because it pulls the app's bulk userdata
|
||
onto what §2.2 makes a **state-only** tier.
|
||
|
||
The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351) applies
|
||
to apps that **have a drive to get wrong**. It used to be reached by testing `HDD_PATH == ""`, which
|
||
also answered "is this app installed?" — one predicate for two questions. Measured in the catalogue at
|
||
`459766cb1639`: **53 templates, 13 declare `needs_hdd: true`, 40 declare `false`**, so for 40 apps that
|
||
test was permanently true and the off-site restore refused them forever, while they were running,
|
||
with a message telling the customer to reinstall them "in the same place" — a place those apps never
|
||
offer. **An app with no drive is not misconfigured** (§8 of `01-topology-and-trust.md` carries the
|
||
`[DESIGN]` marker); it is the majority case.
|
||
|
||
Since v0.219.0 the two questions are asked separately: *installed?* of `ListDeployedStacks()`, failing
|
||
CLOSED when there is no provider to ask; *where?* of `GetAppDrivePath`. A third refusal, with its own
|
||
sentence, covers installed-but-no-resolvable-data-root. **R-356**; reasoning also recorded in
|
||
`felhom-controller/CONTEXT.md`.
|
||
|
||
---
|
||
|
||
### 6.4 The customer's page speaks per tier, and a tier without storage is skipped (2026-09-15)
|
||
|
||
**What the page may claim** (controller v0.243.0 + agent v0.131.0, R-517). The whole-system tile used
|
||
to show the agent's single latest record — so a failed 0-byte PBS attempt read „Naprakész" and ticked
|
||
„Távoli rendszermentés — külön hardveren" (BIGNIGHT 2026-09-14). Now `GET /backup/status` returns per
|
||
tier the newest **success**, the last **attempt** kept apart, and whether the tier's storage exists.
|
||
The page shows each tier's newest success; a failed attempt under it as „sikertelen"; a tier whose
|
||
storage does not exist as „nincs beállítva". „Naprakész" and the remote tick are computed from
|
||
successes only. After an agent restart the success is read back from the tier's storage (size unknown).
|
||
|
||
**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before
|
||
anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped.
|
||
**Still open:** quiescing per tier, so a slow second tier does not keep every app down.
|
||
|
||
### 6.5 Kept data — what a removed app leaves on the drive (controller v0.274.0, `09` §3 decision 36)
|
||
|
||
**What it is.** "Remove the app, keep my data" leaves the app's drive folder (`<drive>/appdata/<app>`, or the folder
|
||
its definition binds through `${HDD_PATH}`) in place. A reinstall over it no longer runs silently into the old files
|
||
(R-657): the install asks „A megőrzött adataimat használom" / "Use my kept data" (the database from the newest copy of
|
||
THIS drive's install — the app's own unit, the second-drive mirror or (since v0.277.0) the off-site copy — loaded under the kept files, then the
|
||
template's `after_load:`, e.g. nextcloud's `occ files:scan --all`) or „Tiszta lappal kezdem" / "Start fresh".
|
||
|
||
**Where it lives.** "Start fresh" RENAMES the folder — same drive, never a copy, never across drives — to
|
||
`<drive>/kept/<app>/<YYYY-MM-DD_HHMMSS>/`, together with the removed app's unit when one sits on that drive (so a later
|
||
Load has its database). `<drive>/kept` is in `ProtectedHDDPaths`, never under `userdata/`, never bound by a live app.
|
||
The page „Megőrzött adatok" / "Kept data" lists every dated folder and every `appdata/<x>` no installed app binds.
|
||
|
||
**The off-site copy counts too (controller v0.277.0, R-691 (2)).** „Use my kept data" and Load consider the app's
|
||
newest OFF-SITE snapshot beside the local copies, and use it when it is NEWER than every local copy or the only one (a tie
|
||
goes local). The page names the copy and its date („távoli mentés, 2026-09-28 04:15" / "the off-site copy, …"). The
|
||
repository is asked once per page, bounded to 20 s; an unreachable one offers nothing. The load downloads the unit ALONE
|
||
(never the files — they are the kept ones), and refuses it before anything moves when it was taken of another drive,
|
||
holds no database, or **does not record its data version** (§6.6 — a unit written before v0.275.0 is never loaded blind);
|
||
a mixed or mismatched unit is refused by §6.6's own rule. Then a dated folder's files move back and the one unit restore
|
||
runs. The downloaded copy is removed on every path. A failed download or a refusal leaves the kept files as they were.
|
||
|
||
**It is NOT backed up.** No tier captures `<drive>/kept` or a leftover `appdata/<x>` of a removed app; the page says
|
||
so („Erről nem készül mentés." / "This is not backed up."). What brings it back into an app is the unit it carries
|
||
(or the app's own unit on the drive), listed per row as „Visszatölthető innen" / "Can be loaded from".
|
||
|
||
**Who deletes it.** Only the household, by Delete on that page with the app's name typed (a wrong name, or a path that
|
||
is not a listed item, is refused — proven live 2026-09-25). **The box never deletes kept data by itself** — operator ruling
|
||
2026-09-26 (`09` §3 decision 40, D3 option A): no age limit, no automatic clean-up. So kept data can fill a drive; the drive-full warning names the kept folders and their sizes as space the household can free.
|
||
|
||
**Read-only view.** The file browser shows each kept item under „Megőrzött adatok", one `:ro` bind per item, and
|
||
follows the list at the next sync (a write is refused: `Read-only file system`, proven live). Not yet readable there:
|
||
a folder its app owns with mode 0770 (nextcloud, `www-data`) — R-691 — **until controller v0.275.0.**
|
||
|
||
**[DESIGN] How the view reads a folder another user owns** — *decided by CC unattended 2026-09-26 — operator may
|
||
reverse.* *One sentence:* how does the read-only view open a 0770 kept folder owned by another user without touching
|
||
the household's files? **Options:** (a) a read-only ACL on the folder; (b) run the file browser as root; (c) a second
|
||
viewer container running as the owner; (d) the file browser joins the folder's OWNING GROUP (`group_add`).
|
||
**Costs:** (a) changes the household's files' metadata — nextcloud checks its data folder's mode after a Load, and
|
||
the brief forbids it; (b) the whole file browser (read-write on the household's userdata) as root; (c) a new container,
|
||
route and login for one view; (d) the group also applies to the view's other mounts — so root's group (0) and the
|
||
view's own (1000) are never added, and only a group-READABLE folder's group is. **Why (d):** the kept binds stay
|
||
`:ro` (a write is refused by the mount), nothing on disk changes, and it is one compose line the sync already writes.
|
||
Limit: a file inside that is owner-only (0600) stays unreadable. Also: a language switch now re-syncs the file browser
|
||
so the source's name („Megőrzött adatok" / "Kept data") follows the box's language. Controller v0.275.0,
|
||
`audits/version-travel-2026-09-26/D3/`.
|
||
|
||
Evidence: `audits/night-2026-09-26/E/` (E1 spike, E5 live proof).
|
||
|
||
### 6.6 Which version a restore brings back (controller v0.275.0, R-696, D4 option A)
|
||
|
||
**[DESIGN] The rule: a restore brings an app back at the version its DATA belongs to — never the data of
|
||
one version under the definition of another.** When the only copy holds data written by an older version,
|
||
the app comes back at that older version with its own data, and the normal guarded update then climbs the
|
||
ladder again, one tested step at a time (`09` §3 decision 14). The household's first sentence says so:
|
||
„A(z) %s visszaállt a(z) %s-i mentésből, a(z) %s verzióra. Elérhető frissítés: a doboz lépésenként hozza
|
||
naprakészre." / "%s is back from the backup of %s, at version %s. An update is available: the box brings it
|
||
up to date one step at a time." The version is every image as `name:tag` — a PostgreSQL step leaves the
|
||
app's own tag unchanged. D4 option B (restore into a temporary copy, update it there, move the data in) is
|
||
NOT built; it would be an add-on to this rule, not a replacement. *D4 is recorded in `STATUS.md`; until it is
|
||
ruled, A is in force.*
|
||
|
||
**[FACT] Why, measured before the fix (9202, 2026-09-26, `audits/version-travel-2026-09-26/A1/`).** The unit's
|
||
definition (`compose/`, `image_pins`) was re-captured by the five-minute status refresh as soon as an update
|
||
moved the pin, while its data (`db-dumps/`, `volume-dumps/`) is written only by the backup legs — so for up to
|
||
a day the unit paired the new definition with the old data. An app step (docmost 0.95.0 → 0.96.0) restored in
|
||
that window came back only because docmost migrated the old data forward at its first start — an update step
|
||
nobody guarded. An engine step (PostgreSQL 16 → 18) restored in that window poured the 16 datadir back,
|
||
`postgres:18` refused it, and **the app was left down**. The off-site restore never wrote a definition at all,
|
||
so after any update it mixed versions until the next night's off-site run.
|
||
|
||
**[DESIGN] How the data carries its version.** Every data file a backup leg writes (the database dump, each
|
||
volume tar) gets a stamp at the moment it is written — its size and mtime, the definition's image pins, and
|
||
what each service was running as `ref@digest` (`data-stamps.json` in the unit). The unit capture folds the
|
||
stamps into the manifest's `data` block (`at`, `image_pins`, `images`, per-file stamps) and **keeps the
|
||
definition its data belongs to**: when the pins moved since the data was written, `compose/` is not rewritten
|
||
until the next data run replaces the data. The manifest's `image_pins` stays the app's current pins, so it says
|
||
both. Files written under different pins (a failed leg) make the unit `mixed`.
|
||
|
||
**[DESIGN] What each restore does with it.**
|
||
|
||
| path | the unit's `data` block | what happens |
|
||
|---|---|---|
|
||
| own unit (Tier 1), second-drive mirror (Tier 2), kept-data Load | present, `compose/` names the data's pins | the data starts under that definition; if it differs from what ran, the restore pages' first sentence names the backup and the version (the kept-data Load has its own sentence and does not add this one) |
|
||
| same | present, `compose/` names OTHER pins, or `mixed` | **refused before anything is touched** (`ErrUnitVersionMismatch`): „…ezért a visszaállítás nem indult el. Az alkalmazás érintetlen" |
|
||
| off-site (Tier 3) | present, the snapshot's version differs from what runs | the snapshot unit's definition is written into the stack dir (and pinned) right after the stop, before any file, volume or database is touched; the database service is resolved from it |
|
||
| any | absent — a unit or snapshot written before v0.275.0, or with an unstamped file | restores **as before** (its captured definition; the off-site path the live one), with a WARN naming it |
|
||
|
||
**[DESIGN] What a restore keeps from the app it replaces (controller v0.276.0, R-697).** The restore writes a fresh `app.yaml` from the unit — env, locked fields, and the pin from the unit's definition — but keeps the records of the app's life on this box: the household's `desired_state`, the kept pre-conversion copies (`conversion_copy`, `earlier_conversion_copies` — a restore does not remove the volume, so it must not forget it), and the ladder history (`failed_update_step`, `last_update_undone`, `last_auto_update`). Not kept: `installed_images` (what ran before, maybe another version). **A drive move is not a restore:** it changes `HDD_PATH` and nothing else (R-700 — before v0.276.0 it dropped the pin, and the syncer then gave the app the catalog's newest version at its next start). Pinned by `internal/stacks/r700_records_carried_test.go`.
|
||
|
||
**[DESIGN] The jump guard.** An installed version that matches no ladder entry's `from` jumps to the catalog's
|
||
current definition when a PERSON presses Update (`ladder.go`, measured live on vikunja 2.5.0, `09` §6.4 part 5);
|
||
the AUTOMATIC leg never presses it (`LegSkipOlderThanLadder`, and a template without a ladder is
|
||
`LegSkipNoTestRecord`). So a restore to a version older than the ladder — or of an app with no ladder — gets
|
||
the other first sentence: „…Ettől a verziótól nem vezet kipróbált frissítési lépés, ezért a doboz magától nem
|
||
frissíti." / "…No tested update step leads on from this version, so the box does not update it by itself." A
|
||
jump is not a tested step, and the page does not promise one.
|
||
|
||
**[DESIGN] The times tell the truth (A4).** Every tier's "proven at" is the time of its DATA, never of a
|
||
manifest: Tier 1 = `data.at` (else the newest data file, the `pre-restore-*` undo copies excluded, else — for a
|
||
unit with no data file — the manifest); Tier 2 = the mirror's data time, capped by its copy time; Tier 3 = the
|
||
snapshot time, capped by the data time the box recorded when it pushed it (`settings.offsite_data_at`). The
|
||
update's precondition and the release of a kept pre-conversion copy read these times; the release additionally
|
||
requires a database dump written after the conversion whose recorded engine is the NEW major.
|
||
|
||
**[FACT] The limit — the image is not in the backup (R-698).** A unit stores image NAMES and digests; the
|
||
image is re-pulled on restore. A version its maker has deleted cannot start. Measured 2026-09-26: the 42 ladder
|
||
digests all resolve today (`audits/version-travel-2026-09-26/A7/`). Options are in R-698; nothing is decided.
|
||
|
||
## 7. The recovery chain (D3) — the reason this document exists
|
||
|
||
**[DESIGN] 3-2-1 describes copies. It does not describe recovery.**
|
||
|
||
Three copies on two media with one offsite is a statement about *bytes surviving*. It says nothing
|
||
about whether the bytes can be turned back into a working system, and a system can satisfy 3-2-1
|
||
completely while having no executable recovery route for a given failure. That is not a hypothetical
|
||
here — §8 has rows where it is the actual state.
|
||
|
||
### 7.0 What a customer can and cannot do ALONE — the four steps the drill found
|
||
|
||
> **[FACT] Added 2026-08-05.** The 2026-08-04 R-201 drill (`audits/DRILL-r201-night-run-2026-08-04.md`)
|
||
> is the first end-to-end proof that a customer's file survives a machine rebuild and comes back
|
||
> byte-identical. **It passed with an operator present, and four manual interventions stood between
|
||
> "the key is recoverable" and "the file is back" — none of which was in any design document.** They
|
||
> are recorded here because this is the section a future reader will use to answer *"can the customer
|
||
> do this alone?"*, and until 2026-08-05 the answer was no for reasons nothing wrote down.
|
||
|
||
| # | The step | Why it stopped a customer | Status |
|
||
|---|---|---|---|
|
||
| 1 | **The reset code was refused on the first try** | `--print-reset-code` runs as a SEPARATE process (`docker exec`); it persists a new code while the running controller keeps the old one in its cache, so the code the customer is told to type never matches. **The controller had to be restarted in between**, which nothing said. Two attempts failed during the drill before that was worked out. | **CLOSED — controller v0.198.0** (R-204 item 1). `effectiveClaimCode` reads through to the persisted state. Read-through, not a TTL: a TTL would leave a window in which a superseded code still works. Fails closed on an unreadable state. Proven live on demo-felhom 9201 with nothing restarted. |
|
||
| 2 | **Re-issuing the off-site credential marked the recovery escrow "stale"** | A stale escrow withholds `restic_pw_sha256` from the report ACK → the controller's auto-confirm cannot flip `pending→escrowed` → `OffboxRunnable()` is false → **every off-site backup refused**. The customer is then told to re-run the recovery ceremony, **which is the one act that would have destroyed the key just recovered.** The repository password had not changed at all. | **CLOSED — hub v0.95.0** (R-204 item 2 / R-196). The precautionary mark is gone. The real case is measured by the controller's Scenario-F hash re-check each ACK — which the mark was BLINDING by emptying that very hash — and by R-197's `offsite_repo_key_changed` at a supersession. |
|
||
| 3 | **The restore's default returned the wrong thing, silently** | `mode=unit` restores the recovery unit — the app's definition, configuration and database dumps — and **not the customer's files**; the userdata in the same snapshot is excluded by `--include`. The outcome message was one sentence for both modes and named neither scope. On the last step of a disaster recovery, the default quietly did not do what the person asked. | **CLOSED — controller v0.198.0** (R-204 item 3). The unit outcome names what came back, what did not, and the step that gets it; the wizard's intent card states its scope before the choice. **The size gate on `mode=full` is untouched**, and the default stays `unit`. |
|
||
| 4 | **A rebuilt box cannot obtain an off-site credential unaided** | The one-time provider password was spent by its predecessor, so the rebuilt guest has nothing to authenticate with and an **operator Re-issue** was required. | **CLOSED — controller v0.199.0 + hub v0.96.0** (2026-08-05, operator ruling: automate it, and **the trigger is a state the BOX DECLARES**). The box now reports `offsite.state=needs_credential` when two local facts hold together — a fresh data area AND a hub-held recovery package — and the hub's `internal/offsiteheal` re-arms the stored one-time secret, minting only when there is nothing to re-arm. **Deliberately NOT automated: the escrow ceremony.** A credential is replaceable; the recovery code is not. |
|
||
|
||
> **[FACT] The honest current answer, updated 2026-08-05 (controller v0.200.0): all four are gone AND
|
||
> the customer is now offered the recovery.** A full-page screen meets the owner of a rebuilt box while
|
||
> the hub holds a sealed package the box cannot open: it explains the situation, says plainly that
|
||
> **nobody can replace a lost recovery code**, takes the code, opens the repository and **lists what is
|
||
> in it** — apps, dates, sizes. No step requires an operator, and no step requires a command line.
|
||
>
|
||
> **WHERE THAT STOPS, and it is a real stop.** The screen **unlocks and only unlocks**. It restores
|
||
> nothing. Putting files back is per-app, lives in the backups area, and the step after the listing —
|
||
> a customer seeing what would change before anything is overwritten — is **not built** (→ R-213). A
|
||
> screen that unlocks and then offers to overwrite is two decisions wearing one button, which is why
|
||
> the operator ruled them separate.
|
||
>
|
||
> **What is still owed as evidence:** the final unlock has never been driven with a CORRECT code
|
||
> through the page (no code was kept for demo-felhom's orphaned history; demo-hp's is operator-held),
|
||
> so the live run exercised handler → agent → hub fetch → age KDF and stopped at the unseal. **And the
|
||
> whole journey has not been re-walked end to end since these fixes** — the closures are proven
|
||
> individually, not as one uninterrupted run. That re-walk is one more drill.
|
||
>
|
||
> **WHY THE TRIGGER FOR STEP 4 IS A DECLARATION, recorded here because it is the design and not an
|
||
> implementation detail:** from the hub, an ABSENT off-site object means *never configured*,
|
||
> *mid-restart*, *a transient config read failure* OR *rebuilt and stranded*, and the hub cannot
|
||
> distinguish them. The BOX can, from two local facts it holds with certainty. So the box states its
|
||
> condition and the hub acts on a stated request — never on a silence. Both facts are required:
|
||
> freshness alone is a box that never had off-site backups, and an escrow alone is a healthy box.
|
||
>
|
||
> **The split that is deliberate and must not be widened: CREDENTIAL AUTOMATIC, KEY CUSTOMER-PRESENT.**
|
||
> A credential is transport and is replaceable; the recovery code is not, because only the customer
|
||
> holds it. Nothing in this chain runs, or asks for, an escrow ceremony.
|
||
|
||
### 7.1 The dependency graph
|
||
|
||
> **[FACT] SUPERSEDED IN PART BY D5 — SHIPPED 2026-07-30 (controller v0.188.0).** Leg 1 below
|
||
> (**secrets**) is **no longer true for Tier-1 and Tier-2**: the local recovery unit now carries the
|
||
> portable secret class, so those two tiers need the **drive and nothing else**. The graph and leg 1
|
||
> are kept as written because they are the model everything downstream was derived from, and because
|
||
> leg 2 (**the living app**) is unchanged and still binds Tier-2's file leg and Tier-3. Read §7.4 for
|
||
> the current chain; the corrected rows are 3, 4 and 9 in §8.
|
||
|
||
**[FACT — as of the R-108 survey, before D5]** Every app-tier restore depends on the guest, for one of
|
||
two reasons:
|
||
|
||
```
|
||
┌──────────────────────────────────────────┐
|
||
│ the LIVE GUEST │
|
||
│ · settings.json (tier-2 destination) │
|
||
│ · encryption.key (32 B) │
|
||
│ · app.yaml (ENC: under that key) │
|
||
│ · the deployed app itself │
|
||
└───────────┬──────────────────────────────┘
|
||
│ required by
|
||
┌──────────────────────────┼──────────────────────────┬─────────────────────┐
|
||
│ │ │ │
|
||
Tier-1 restore Tier-2 restore Tier-3 reconstitute Tier-3 scratch
|
||
(rebuild the app) (fill in files) (overwrite + replay) (verification copy)
|
||
│ │ │ │
|
||
needs SECRETS needs the app needs the app needs only the
|
||
from app.yaml running + the deployed + a DB repo password
|
||
(restore_unit.go recorded dest service identifiable (also in the guest)
|
||
:130-132) (tier2_restore.go
|
||
:114-116)
|
||
│ │ │ │
|
||
└──────────────────────────┴──────────────────────────┴─────────────────────┘
|
||
│
|
||
┌───────────▼──────────────────────────────┐
|
||
│ WHOLE-GUEST TIER (local vzdump / PBS) │ ← Lane 2, operator only
|
||
└───────────┬──────────────────────────────┘
|
||
│ required by
|
||
┌───────────▼──────────────────────────────┐
|
||
│ THE HOST (agent.json, bootstrap.json, │ ← in NO backup at all
|
||
│ vmbr9, sudoers, /etc/pve/priv, WG) │ (INV Part D1)
|
||
└──────────────────────────────────────────┘
|
||
```
|
||
|
||
**[FACT] The two legs of the dependency, precisely:**
|
||
|
||
1. **Secrets.** The recovery unit is secret-free by design — *"It NEVER writes a secret value"*
|
||
(`internal/backup/recovery_unit.go:73`). `RestoreFromRecoveryUnit` recovers secrets **from the
|
||
guest, never from the unit** (`restore_unit.go:130-132`), and the policy is stated outright at
|
||
`:18-22`: *"Regenerate NOTHING… A missing DATA-ENCRYPTING key is FATAL: regenerating it would
|
||
render the restored data unreadable, so we refuse and tell the operator to do a PBS whole-guest
|
||
restore."* Those secrets are encrypted under `encryption.key`, a 32-byte file that exists **only
|
||
inside the guest** (LIVE, both boxes — INV Part C, row 9).
|
||
2. **The living app.** Tier-2's restore reads the destination recorded in the guest's `settings.json`
|
||
(`tier2_restore.go:114-116`), and Tier-3's reconstitution refuses outright when the app is not
|
||
deployed — *„a(z) %s nincs telepítve — előbb állítsd helyre az alkalmazást, utána az adatokat"*
|
||
(`offbox_reconstitute.go:198-201`).
|
||
|
||
**[FACT] So all three app tiers are conditioned on the whole-guest tier**, which is Lane 2. The
|
||
customer's own recovery lane rests on an operator-only tier, and today nothing tells the customer
|
||
that. — **This sentence was the reason D5 exists. It is now false for Tier-1 and Tier-2; see §7.4.**
|
||
|
||
### 7.4 The chain AFTER D5 (controller v0.188.0, 2026-07-30) — the current model
|
||
|
||
**[FACT] Tier-1 and Tier-2 no longer depend on the whole-guest tier.** The unit's `compose/app.yaml`
|
||
(mode 0600) carries the **portable secret class** — every catalog `type: secret` field, i.e. the
|
||
declared `data_key`s, the 18 DB/root passwords and the internal signing secrets — so neither
|
||
`encryption.key` nor the guest's `app.yaml` is required to rebuild an app:
|
||
|
||
```
|
||
Tier-1 restore Tier-2 restore Tier-3 reconstitute Tier-3 scratch
|
||
(rebuild the app) (fill in files) (overwrite + replay) (verification copy)
|
||
│ │ │ │
|
||
needs ONLY THE DRIVE needs the app needs the app needs only the
|
||
(unit carries the running + the deployed + a DB repo password
|
||
secrets; guest is recorded dest service identifiable
|
||
consulted only for
|
||
the withheld class) └────────── still conditioned on the guest (leg 2, unchanged) ──────┘
|
||
│
|
||
✅ INDEPENDENT of the whole-guest tier
|
||
```
|
||
|
||
**[FACT] What still needs the guest, precisely** — so this is not read as more than it is:
|
||
|
||
1. **Leg 2 is untouched.** Tier-2's *file* restore still reads the destination from the guest's
|
||
`settings.json` (`tier2_restore.go:114-116`) and Tier-3's reconstitution still refuses when the app
|
||
is not deployed (`offbox_reconstitute.go:198-201`). D5 removed the **secrets** leg, not the
|
||
**living-app** leg. So "Tier-2 no longer depends on the guest" is true of the *unit-restore* path it
|
||
shares with Tier-1, and **false** of its additive file-merge path.
|
||
2. **The withheld class.** The 7 `type: password` admin logins and `vaultwarden/ADMIN_TOKEN` are
|
||
deliberately NOT on the drive; they come from the guest, or are regenerated (O4). Their absence
|
||
costs a credential reset, never data.
|
||
3. **The fail-closed gate remains.** A `data_key` in neither the unit nor the guest still refuses the
|
||
restore outright. D5 makes it normally present; it does not soften the gate.
|
||
|
||
**[DESIGN] Why plaintext is the right answer here, and what it is coupled to.** Every travelling secret
|
||
either decrypts data on the *same drive* or authenticates to a container on an internal compose network
|
||
with no external listener — so possessing it adds nothing to possessing the drive, which is exactly D2's
|
||
argument for keeping the DATA plaintext. That argument holds **only because** the internet-reachable
|
||
class is withheld. **The exclusion is what licenses the plaintext; the two must not be relaxed
|
||
independently.** `stacks.PortableSecretEnvVars` is the single boundary, with a code-level
|
||
`nonPortableSecrets` register (not a catalog flag — a boundary a catalog push can move is not a
|
||
boundary, cf. R-97a).
|
||
|
||
**[FACT] Precedence, because two sources now exist.** The **unit wins** over the guest. Not
|
||
"newest wins": the unit's secrets are captured in the same run as the dumps beside them, so the unit's
|
||
value matches *the data being restored*, whereas the guest's is merely the most recent. Preferring the
|
||
guest is the data-loss bug — a rotated data key does not decrypt data encrypted with the old one, and a
|
||
rotated DB password does not match the hash inside the restored data directory.
|
||
|
||
**[FACT] Proven live 2026-07-30**, on a scratch drill guest, through the real endpoints: AdventureLog
|
||
(`SECRET_KEY` data_key + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** —
|
||
`secrets recovered=2/2`, 27.6 s — and then **the application itself read the seeded customer row over
|
||
TCP with its own credential**, which is the observable that matters (a restore that returns success onto
|
||
unreadable data is a failure that looks like a pass). The discriminator held: the pre-backup row came
|
||
back, a row added *after* the backup was gone, so the volume tar was genuinely restored. There was **no
|
||
`.sql` dump** in that unit, so the DB came back from the volume tar — the exact case where a regenerated
|
||
password would have failed. Separately, Grafana's `type: password` admin login was proven **withheld**:
|
||
the sentinel value was live in the container and encrypted in the guest, and appeared in **0 files**
|
||
anywhere under the backup namespace. Full record: `audits/D5-drive-alone-restore-2026-07-30.md`.
|
||
|
||
### 7.2 A tier whose prerequisites cannot be met in the failure it exists for
|
||
|
||
**[FACT]** Two instances, both current:
|
||
|
||
- **~~Tier-2 vs primary-drive loss~~ — CLOSED 2026-08-31, controller v0.229.0 (R-102).** Tier-2's
|
||
stated purpose is surviving the loss of the primary drive. It was true until v0.229.0 that in that
|
||
failure the primary recovery unit is gone while the surviving mirror on the second drive —
|
||
`backups/secondary/<app>/recovery-unit/` — was read by **no code path** (§6.3), so for the **45**
|
||
class-B apps (§6.2, count settled the same day) the restore was a guaranteed no-op in exactly its
|
||
designed scenario.
|
||
**Tier-2 can now meet its prerequisite in the failure it exists for**, and that is stated plainly
|
||
because it is the whole point: the restore was run on `demo-hp` **with the primary unit moved aside**
|
||
and returned 3 volumes of 3 and 1 database of 1 in 28.65 s, byte-for-byte, with the app then reading
|
||
its own row over TCP with its own credential — and again with the guest's `app.yaml` also moved
|
||
aside, `secrets recovered=2/2` from the mirrored unit. Evidence:
|
||
`audits/DRILL-r102-tier2-unit-2026-08-31/`.
|
||
**What the drill did NOT cover, stated per §8 below:** the full drive-loss journey — a genuinely
|
||
absent or replaced physical drive — was not run. Only the recovery-unit half was.
|
||
- **Tier-3 vs guest loss — HALF of this closed.** Tier-3 holds the volume tars and the DB dump. It
|
||
was true until controller v0.218.0 that the tars were unpacked only by the Tier-1 path; since
|
||
v0.218.0 the reconstitution replays them itself (`volReplay`) → **R-107 CLOSED 2026-08-22**. What
|
||
is **still** true, and is the part that was never R-107: reconstitution requires the app to be
|
||
**deployed** and skips the unit, so offsite alone cannot rebuild an app onto a fresh guest. That
|
||
requirement is a deliberate decision (R-253) — the restore does not choose a customer's drive for
|
||
them — and since **v0.219.0** it means only what it says: an app that is not deployed. It no longer
|
||
catches the 40 driveless apps → **R-356**.
|
||
|
||
### 7.3 ~~What D5 would change — and why it was blocked~~ — **D5 SHIPPED 2026-07-30 (v0.188.0)**
|
||
|
||
> **This section is history.** D5 is implemented and proven live; the current chain is **§7.4**. Kept
|
||
> because the reasoning below is why the precondition was required, and because the last paragraph
|
||
> (`.fab` / R-126) is still open and still not part of D5.
|
||
>
|
||
> **One correction to the target as stated below.** It assumed the class that must travel is the
|
||
> `data_key`-flagged secrets. Part 0 of the D5 task tested that and it did not survive: the flag is
|
||
> unreliable (four encryption keys the catalog itself labels as such are unflagged → **R-127**), and a
|
||
> DB password is **not resettable in practice** — `POSTGRES_PASSWORD` is ignored once PGDATA is
|
||
> non-empty, so a regenerated value leaves the app unable to authenticate against its own restored rows
|
||
> while the dump replay still reports success. The shipped boundary is therefore *all* `type: secret`
|
||
> minus the register, not `data_key` alone.
|
||
|
||
**[DESIGN, TARGET — SUPERSEDED BY §7.4]** The intended fix is to make app secrets travel with the **local**
|
||
recovery unit, so Tier-1 and Tier-2 restore work **without the guest and without R**. Offsite
|
||
already encrypts everything, so secrets travelling offsite would be covered by the escrowed repo
|
||
password. R would then be required for **offsite recovery and host identity only** — losing R would
|
||
cost the offsite route, not local recovery.
|
||
|
||
**The premise D5 rests on is now established.** §2 of the task that produced this document required
|
||
it to be proven, not assumed: the backup tree must be unreachable from every browsing, download and
|
||
export surface. When this document was written it was **not** — the FileBrowser network-share bind
|
||
reached it. **R-108 closed that on 2026-07-30** (controller v0.187.0) by refusing app namespaces on
|
||
network storage, so no `backups/` tree can exist under the share-root bind; every other surface was
|
||
already clear (§10.1's table). **D5's precondition is therefore MET and D5 may be adopted.**
|
||
|
||
**Still true, and not part of D5's precondition:** a `.fab` bundle carries plaintext secrets by
|
||
design with an optional password, and `storageDriveList()` does not filter network paths, so a bundle
|
||
can be **exported onto** a NAS (§5, → **R-126**). That is an export destination the customer chooses
|
||
explicitly, not a browsing surface reaching a backup tree, and it is unchanged by D5 — D5 moves
|
||
secrets into the local recovery unit, not into `.fab`. It is tracked separately rather than folded in.
|
||
|
||
Until D5 is actually implemented, §7.1's chain stands as the model.
|
||
|
||
---
|
||
|
||
### 7.5 Where a recovery unit is RETAINED — and the size bound it puts on Lane 1
|
||
|
||
Added 2026-08-02 from `audits/SPIKE-recovery-unit-space-2026-08-02.md` and
|
||
`audits/CAMPAIGN-10-closeout-2026-08-02.md`. §6.1 says a Tier-1 unit lives *"on the app's own drive"*,
|
||
which is true and is the whole story for a drive-resident app. **It is not the whole story for an app
|
||
with no data drive, and that case was undocumented until now.**
|
||
|
||
**The unit is RETAINED, not staged.** `GetAppDrivePath` (`felhom-controller/internal/backup/backup.go:245-255`)
|
||
returns the app's `HDD_PATH` if it has one and **`systemDataPath` otherwise** — what
|
||
`internal/appbackup/paths.go:26-27` calls *"the SSD-only system-data fallback"*. Nothing deletes a unit
|
||
after it is copied onward: the only prune is F5 (`backup.go:1053-1112`), which removes residue on an
|
||
**old** drive when an app **moves**. So for every app without a data drive, `mp1` holds the kept copy
|
||
indefinitely.
|
||
|
||
**What a unit contains, which bounds the problem** (`internal/backup/recovery_unit.go:20-25`):
|
||
`compose/` + `db-dumps/` + `volume-dumps/` (named-volume tars) + `manifest.json`. **It never contains
|
||
`mp8` / `HDD_PATH` userdata** — a photo library on a 1 TB drive is not in a unit, and cannot make one
|
||
overflow.
|
||
|
||
**The resulting mismatch, on a default appliance:**
|
||
|
||
| | size | holds |
|
||
|---|---|---|
|
||
| `mp0` `/var/lib/docker` | **50 G** | every app's live volumes — the thing a unit copies |
|
||
| `mp1` `/mnt/sys_drive` | **20 G** | the retained units of every **driveless** app |
|
||
|
||
A DB-backed app's unit is up to **~2×** its data, because it carries the volume tar **and** the SQL
|
||
dump (measured: 21.1 GB of app data → a **40.2 GB** unit). A file-only app's is **1.00×** (measured:
|
||
homebox 2305 MB → 2305 MB). `--sysdata-grow` (`felhom-agent/cmd/felhom-agent/main.go:178`) defaults to
|
||
**0**, is not derived from the physical drive, and no observed box passes it — demo-hp's production
|
||
guest 9201 ships `mp0 50G / mp1 20G`.
|
||
|
||
**`mp1` gates the entire app-data chain, not just Tier 1.** Tier-2 mirrors the unit *"(always)"* from
|
||
`RecoveryUnitPath` (`internal/backup/tier2.go:302,368`) and Tier-3 carries it too, so a unit that
|
||
cannot be written to `mp1` leaves Tier-2 and Tier-3 with nothing to copy.
|
||
|
||
**The bound this places on Lane 1 (§3).** Lane 1's promise is that a customer can restore an app from
|
||
the drive alone, without the operator and without the guest — D5 (§7.4) made that true by putting the
|
||
portable secrets in the unit. **That independence is bounded by app size**, and the bound is:
|
||
|
||
> On a default appliance, an app can be restored by Lane 1 only while its recovery unit still fits the
|
||
> retained space — **≈ 19 GB of app data for a file-only app, ≈ 10 GB for a DB-backed one.**
|
||
> Past that the unit cannot be captured, and the app falls back to Lane 2's operator-driven
|
||
> whole-guest route.
|
||
|
||
**AS OF 2026-08-02 SOMETHING NOW WARNS, AND THE ALERTING IS PART OF THIS CONTRACT (R-167 / R-158,
|
||
decision D-c; controller v0.191.x + hub v0.89.0).** The last sentence of this section used to end
|
||
"nothing warns when an app crosses the line". Two signals now exist and both are PROVEN-LIVE:
|
||
|
||
- **To the CUSTOMER, before anything fails** — `internal/fillwatch` warns per FILESYSTEM (never per
|
||
app: one full disk holding ten apps would fire ten times) on **whichever trips first, used ≥ 85% or
|
||
free < 5 GiB**, critical at 95% / 2 GiB, clearing at 75% / 7 GiB. **Two terms, because a percentage
|
||
alone lies at both ends of the range this section itself documents:** 85% of a 20 G `mp1` leaves
|
||
3 G — less than one DB-backed app's unit — while 85% of a 4 TB drive leaves 600 G. It watches the
|
||
app-data volume, the system-data volume **and** every registered drive, which the previous
|
||
`health_degraded` signal did not. Edge-triggered against persisted state; the hub owns cooldown.
|
||
- **To the OPERATOR, when a capture actually fails** — `recovery_unit_capture_failed`, per app, with
|
||
the target filesystem's used/free bytes at the moment of failure, so the *why* needs no login. It is
|
||
**operator-tier** (`notify.operatorOnlyEvents`) and deliberately not `backup_failed`: a customer can
|
||
take no action on a capture failure.
|
||
|
||
### 7.5.1 — THE CEILING THIS SECTION DESCRIBES HAS BEEN REMOVED (2026-08-03, R-165 / decision D-a)
|
||
|
||
**Everything above describes the SPLIT layout, which is now the legacy shape.** A golden built by
|
||
`build-golden.sh` **v3.0.0** ships **one** data volume; `mp1` does not exist. Both consumer paths are
|
||
binds of subdirectories of it (variant **V-c**):
|
||
|
||
```
|
||
mp0 -> /var/lib/felhom ├─ docker/ --bind--> /var/lib/docker
|
||
└─ sys_drive/ --bind--> /mnt/sys_drive
|
||
```
|
||
|
||
**So the size bound below no longer applies to a box built from that golden.** A driveless app's
|
||
recovery unit is limited by the box's actual free space, not by a partition set at build time. The
|
||
mismatch table above (`mp0` 50 G vs `mp1` 20 G) describes what a merged box no longer has.
|
||
|
||
**R-175, fixed here rather than left standing.** The bound below was stated as the fleet's and was
|
||
**one box's**: it is derived from `mp1 = 20 G`, which is demo-hp exactly and never was demo-felhom
|
||
(`mp0 200G / mp1 50G`, where the same arithmetic gives ≈ 49 GB / ≈ 24 GB), nor the golden (`16 G / 8 G`
|
||
before provision grew them). **Read it as a function of `mp1`, and only for a box still on the split
|
||
layout.** Measured: `audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1.
|
||
|
||
**What replaced the partition's second job — the reserve.** `mp1` was also a BULKHEAD: an overflow was
|
||
refused per app with the last good unit byte-identical, and it **could not reach `/var/lib/docker`**,
|
||
because that was a different filesystem. On a merged box it can. Decision **B2**, shipped in controller
|
||
**v0.192.0**, is that bulkhead made deliberate — a two-term reserve (97% used or 1 GiB free) in
|
||
`internal/fillwatch`'s shape, sitting beyond its critical band so the customer is always warned first.
|
||
It **refuses per app and never deletes**: nothing on this filesystem is generational, so pruning could
|
||
only destroy a different app's only local copy.
|
||
|
||
**WHAT IS RECORDED, WHAT IS E-MAILED, AND HOW OFTEN (controller v0.194.0 + hub v0.90.x, R-182).**
|
||
The two are deliberately different mechanisms, because conflating them is how seven failures went
|
||
missing on 2026-08-03 without leaving a trace.
|
||
|
||
| | Record | Notification |
|
||
|---|---|---|
|
||
| what | `recovery_unit_capture_failed`, one per failed app | `backup_run_failures`, one per RUN |
|
||
| when | every time, unconditionally | at the end of a run, **only if something failed** |
|
||
| gated by | nothing — not cooldowns, preferences or delivery | the hub's operator cooldown |
|
||
| where it lands | the events table **and** `notification_log` (status `recorded`) | the operator's inbox |
|
||
|
||
- **A clean run e-mails nothing.** Silence means the run finished and found nothing wrong — and that
|
||
is only safe because the hub's daily deadline check raises `expected_backup_missed` from the box's
|
||
REPORT freshness, independent of any mail the box sends. That check is load-bearing for this
|
||
design; weakening it re-opens a silent-failure path.
|
||
- **A suppressed operator notification leaves a `suppressed` row** naming the key that suppressed it.
|
||
Deciding not to tell someone is itself an event worth recording.
|
||
- **Deliberate skips are not failures** and never appear in the digest — a disconnected or
|
||
decommissioned drive has its own alert, and a nightly digest about an unplugged drive is one the
|
||
operator stops reading.
|
||
- **Cadence:** a nightly run gives at most one mail a day. A manual run always reports, even within
|
||
the hour, because someone pressing the button is actively trying to get a backup. The periodic
|
||
capture sweep is capped by the ordinary hourly cooldown.
|
||
|
||
**THE CONTRACT, stated as what the code provides (controller v0.193.0, R-181).** The reserve is a
|
||
**per-app, per-run ADMISSION decision, not a capture check.** It is taken once for an app, immediately
|
||
before that app's FIRST write of the run, and it covers **all three write legs — the database dump, the
|
||
volume dump and the recovery-unit capture**. Those three write under one per-app root
|
||
(`backups/primary/<app>`), which is what makes one verdict able to cover them honestly.
|
||
|
||
- **What it guarantees.** A refused app has **nothing written for it in that run**, its previous unit
|
||
is **byte-identical**, it is **not stopped**, nothing anywhere is deleted, and the operator gets
|
||
**exactly one** alert naming the app, the term that bound and the disk figures.
|
||
- **Two terms, two questions.** *Headroom*: is the filesystem already below the reserve? *Size*: would
|
||
THIS app's write take it below? The size estimate is the app's previous `.sql` + `.tar` on disk;
|
||
with no history the decision degrades to headroom alone, deliberately — otherwise the first backup
|
||
is the one that can never happen.
|
||
- **Why it is decided lazily and not once per run.** Space changes during a run: app A's dump can put
|
||
app B under the reserve, so a verdict taken at run start reads a disk that no longer exists.
|
||
- **Why it is never re-decided between an app's own legs.** That is precisely the shape v0.192.0 had —
|
||
the two dump legs unguarded and only the capture refused — under which the reserve was consumed by
|
||
the very write it exists to bound, and the refusal's *"the previous unit is untouched"* was measured
|
||
false. Proven live on demo-hp 2026-08-03 (R-181), fixed the same day, and re-proven by filling the
|
||
box for each of the two terms.
|
||
- **It sits ahead of `DumpAppVolumesSafe`**, which stops the stack as its first act — a refusal
|
||
decided inside it would already have bounced the app it is refusing to back up.
|
||
|
||
**Status caveat, deliberately explicit:** every box in the field that has not been reinstalled is still
|
||
on the split layout and everything above still describes them exactly. This subsection describes what a
|
||
box built from golden ≥ 0.192.0 gets. Both demo boxes were reinstalled from it on 2026-08-03 (R-178).
|
||
|
||
Two things are deliberately **not** recorded here. **The sizing ratio is the operator's ruling**
|
||
(**R-163**) — this section states the constraint, not a number. And **the same-device placement is
|
||
intended, not a defect**: a driveless app's unit sits on the same SSD as its volumes, but Tier-2's
|
||
cross-drive mirror, the whole-guest tier (which covers `rootfs + mp0 + mp1`, §6.1) and Tier-3 all cover
|
||
device loss. What was missing was that this case existed at all, and that nothing warns when an app
|
||
crosses the line — **R-158**.
|
||
|
||
## 8. The failure → recovery matrix
|
||
|
||
> ### THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01)
|
||
>
|
||
> **Closed at controller v0.232.0 / hub v0.111.1.** Everything a **customer** does for themselves is
|
||
> finished and proven live — rows **1, 2, 3, 3b, 3c, 6, 7, 14**. Everything only an **operator** does
|
||
> is **deliberately deferred until after beta** — rows **4, 8, 9, 10, 11 (+11b), 12**, each carrying
|
||
> the marker **`[BETA-DEFERRED]`** in its status cell so the set greps as a group.
|
||
>
|
||
> **The blanks in the RTO column below are now blank ON PURPOSE, and that is the whole difference.**
|
||
> Eleven rows carry a blank RTO — 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15 — and only six of those are
|
||
> deferred work. **Row 5** has no recovery to time (a derived copy), **row 13** has no route by
|
||
> design, **row 14** is proven and merely never stopwatched, **row 11b** is a note not a row, and
|
||
> **row 15 is an open DEFECT (R-104) that this line does NOT cover.**
|
||
>
|
||
> **NO STATUS MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day.** A stopping
|
||
> line that promotes a row is a stopping line that lies. Full text, and the conditions that reopen
|
||
> this: `documentation/backlog/OPEN-ITEMS.md`, *"DECIDED — the backup and restore arc is CLOSED FOR
|
||
> BETA"*.
|
||
|
||
**This is the core artifact.** One row per failure. It is authoritative for recovery routes;
|
||
`00-capability-map.md` stays authoritative for per-capability status.
|
||
|
||
**How to read the numbers.**
|
||
|
||
- **RTO** — **only measured durations** from INV Part F appear here. A blank cell means *nobody has
|
||
ever measured it*, and a blank is a finding, not an omission.
|
||
- **RPO** — **no RPO has ever been measured from an incident.** These cells carry the **configured
|
||
cadence that bounds RPO**, read live from the box, labelled `(cadence)`. A blank means no cadence
|
||
governs the row.
|
||
- **Status** — `PROVEN` (live, cited) · `PARTIAL` (some legs proven) · `IMPLEMENTED` (code + tests,
|
||
never exercised) · `NONE` (no route exists).
|
||
|
||
| # | Failure | What survives | Recovery route | Invocable by | RTO (measured) | RPO (cadence) | Status | Evidence |
|
||
|---|---|---|---|---|---|---|---|---|
|
||
| 1 | **Customer deletes files** | everything else | „Fájlok visszaállítása" — Tier-2 additive merge | **customer** | **39 s** (6 files) | 24 h | **PROVEN** | CAMPAIGN-9 A1 — byte-identical returns, both non-destruction promises kept, bytes re-served by paperless's own API |
|
||
| 2 | **An app's data directory is destroyed** | the guest, the other tiers | same route | **customer** | **46 s** (43 files) | 24 h | **PROVEN** | CAMPAIGN-9 A3 — loss verified by a positive observable (doc download 200 → 404), then 16/16 documents usable |
|
||
| 3 | **An app's DB and named volumes are lost** | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the **only** path that unpacks volume tars) | **customer** | **18.25 s** (path execution) · **27.6 s** (D5 drill, guest `app.yaml` absent) | 24 h | **PROVEN** (content recovery proven 2026-07-30) | CAMPAIGN-9 A2 proved the path executes; **D5 v0.188.0 closed the content gap** — after a restore with the guest's `app.yaml` moved aside, the app read the seeded row **over TCP with its own credential**, the pre-backup row returned and a post-backup row was gone (so the tar was really restored). No `.sql` dump in the unit ⇒ the DB came back from the volume tar. **2026-08-30 — this row KEEPS its PROVEN status and the reason is worth stating: R-353 was a defect in the MESSAGE, not in the mechanism.** The restore really did return what the unit held, every time; what it could not do was say so, because the count was discarded one call deep. Controller v0.226.0 fixed the sentence and changed nothing about the recovery path. A status that measures whether data comes back must not move because a status line was wrong |
|
||
| 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery |
|
||
| 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** |
|
||
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** **`[BETA-DEFERRED]`** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore |
|
||
| 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go`. **The derived-copy rule is UNCHANGED by R-403 (controller v0.230.0) and the single exception is stated in §8.2 below — read it before "fixing" a skip you find in the code** |
|
||
| 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session |
|
||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** **`[BETA-DEFERRED]`** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
||
| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** **`[BETA-DEFERRED]`** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) |
|
||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** **`[BETA-DEFERRED]`** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. |
|
||
| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** **`[BETA-DEFERRED]`** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) |
|
||
| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** **`[BETA-DEFERRED]`** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) |
|
||
| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** **`[BETA-DEFERRED]`** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D |
|
||
| 13 | **Customer loses R** | every byte, every tier, the Recipe | **on-premises recovery is unaffected** — Tier-1/2/3 and the whole-guest tiers all work while the guest and host live. What is lost: the ability to re-establish **host identity** and to open the escrowed offsite key material after a host loss | — | | | **NONE for host-loss** | R exists in **zero** system copies by design (INV Part C, row 24). Whether the operator should be able to recover is **open** → §11-A, §11-B. **D5 narrows this further:** Tier-1/2 app recovery needs only the drive, so losing R now costs the offsite route and host identity, never local app recovery (§7.4) |
|
||
| 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix |
|
||
| 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** |
|
||
|
||
### 8.2 The ONE exception to row 5's derived-copy rule (R-403, controller v0.230.0)
|
||
|
||
**[DESIGN] Row 5 stands: the secondary IS a derived copy and it IS rebuilt on the next run.** Nothing
|
||
below weakens that, and a future reader who finds `RunTier2` skipping a leg and does not find this
|
||
section will "fix" it back — which is why it is here and not only in a register row.
|
||
|
||
**[FACT] What was measured, on the shipped v0.229.0, on `demo-hp` 2026-08-31.** An app's Tier-2 copy
|
||
went from **120 082 104 B** (4 database dumps + 3 named-volume tars) to **7 036 B** (none of either)
|
||
in one nightly run, and the run recorded itself a success. The mechanism was three individually
|
||
correct lines: `RunTier2` guarded the unit leg with `os.Stat` alone — *does the folder exist* —
|
||
`rsyncMirror` is `rsync -a --delete`, and nothing between them compared source to destination. **An
|
||
empty recovery unit is a folder that exists.** Evidence:
|
||
`audits/DRILL-r403-tier2-delete-2026-08-31/`.
|
||
|
||
**Why it bites harder since 2026-08-30:** R-102 made that mirror a LIVE recovery route (§6.3, §8 row
|
||
3b). Deleting it used to cost a copy nobody could open; it now costs the route itself.
|
||
|
||
**[DESIGN] The exception, stated exactly.** The unit leg — and ONLY the unit leg — is skipped when the
|
||
SOURCE unit carries no data and the DESTINATION unit does. Everything else is unchanged:
|
||
|
||
| source unit | destination unit | behaviour |
|
||
|---|---|---|
|
||
| complete | complete | mirror, with `--delete`, as before |
|
||
| complete | hollow or absent | mirror (the normal first copy) |
|
||
| hollow | hollow | mirror — both sides agree, nothing is at risk |
|
||
| **hollow** | **complete** | **skip the unit leg, preserve the destination, warn, record for the surface** |
|
||
|
||
**"Hollow" is a MANIFEST question, never a size question** — the manifest lists no database dump and
|
||
no volume tar; absent or unparseable counts as hollow, fail closed. A unit with a fat compose capture
|
||
and no dumps is the dangerous shape; a 360-byte unit belonging to a tiny app is healthy.
|
||
|
||
**The data legs are NOT guarded and must not be.** A classified app's copy legitimately shrinks as
|
||
`export` drops out of its class set (`tier2.go` header), and fencing that would be calling this row's
|
||
own decision a defect.
|
||
|
||
**[DESIGN] And the CAUSE is closed at the other end.** The hollow primary was written by the 5-minute
|
||
capture job **two seconds** after a Tier-2 unit restore. Since v0.230.0 `RestoreTier2Unit` refills an
|
||
absent or hollow primary unit from the mirror **inside the call, before returning**, so no capture can
|
||
observe the hollow state. **The capture itself is deliberately not guarded:** a capture that describes
|
||
an empty drive as empty is correct, and guarding it would make the manifest lie.
|
||
|
||
### 8.1 The blank cells, listed explicitly
|
||
|
||
Per the rule that a blank is a finding, here they are:
|
||
|
||
| row | blank | why |
|
||
|---|---|---|
|
||
| 4 | RTO | no drive-loss recovery has ever been timed. **The Tier-2 unit restore inside it now is** — 28.65 s, row 3b, 2026-08-31 — but that is the route, not the journey: no drive has been removed or replaced under a recovery |
|
||
| 5 | RTO | never timed; the rebuild is a normal Tier-2 run |
|
||
| 6 | — | RTO present, but it is a **restore into a scratch guest on the same host**; a restore *to a different host* has never been timed |
|
||
| 8 | RTO | a host has never been rebuilt as itself (INV Part D1) |
|
||
| 9 | RTO | the composed whole-box path has never been run |
|
||
| 10 | RTO | no ransomware-shaped recovery has ever been run |
|
||
| 11 | RTO | a hub restore has never been performed |
|
||
| 12 | RTO, RPO | no provider-loss recovery has ever been run |
|
||
| 13 | RTO, RPO | not a timed recovery; a capability loss |
|
||
| 14 | RTO | the break-glass path is proven but was never timed |
|
||
| 15 | RTO | the manual unlock was performed but not timed |
|
||
|
||
**Also unmeasured, and not representable as a row** (INV Part F.3): any restore larger than 155.5 MB
|
||
from restic; any restore of a customer data drive (no tier holds one whole); the weekly PBS
|
||
incremental; `.fab` import wall-clock; and time-to-first-byte for a customer restore over a home
|
||
uplink — **no customer has ever driven a restore**.
|
||
|
||
---
|
||
|
||
## 9. What the model implies for the tiers (recorded, not new design)
|
||
|
||
**[DESIGN]** Three consequences follow from §3–§7 and are stated so they are not re-derived:
|
||
|
||
1. **Tier-2 is a *drive-loss* tier, not a second chance at Tier-1.** Its job is to survive one drive
|
||
dying. That is why it mirrors the unit at all — and why the unit mirror being unreadable (§6.3)
|
||
defeats the tier rather than degrading it.
|
||
2. **Tier-3 is a *premises-loss* tier.** It is the only copy that survives fire, theft and
|
||
ransomware-with-host-access, which is why it is the only tier whose credential exposure is ranked
|
||
as the largest open data risk (R-95).
|
||
3. **The whole-guest tiers are the *system* tier, and they sit under everything else** (§7.1).
|
||
Treating them as "one more copy" is precisely the 3-2-1 framing this document rejects.
|
||
|
||
---
|
||
|
||
## 10. Known gaps
|
||
|
||
Every divergence between the model above and the system as it is, each with an ID.
|
||
|
||
### 10.1 ~~D5 is BLOCKED~~ — precondition CLOSED by R-108 (v0.187.0); **D5 ITSELF SHIPPED (v0.188.0)**
|
||
|
||
> **D5 IS DONE — 2026-07-30, controller v0.188.0.** The precondition below was met by R-108, and D5 was
|
||
> then implemented and proven live the same day: Tier-1/Tier-2 no longer depend on the whole-guest tier,
|
||
> and a customer needs **the drive and nothing else**. The current recovery chain is **§7.4**; the
|
||
> corrected matrix rows are 3, 3c and 13. This section's remaining value is the *precondition* analysis,
|
||
> which is why the table below is kept.
|
||
|
||
> **D5's PRECONDITION IS MET.** An app's data namespace can no longer be placed on network storage, so
|
||
> no `backups/` tree can exist inside FileBrowser's share-root bind, and the browsing surface therefore
|
||
> cannot reach the backup tree on ANY storage class. Every other read surface in the table below was
|
||
> already NO. **D5 may be adopted** — nothing in this section blocks it.
|
||
>
|
||
> **The fix inverted the obvious one, and that is the durable lesson here.** The share-root bind was
|
||
> not narrowed, because it **cannot** be: (a) the `:rslave` share-ROOT bind is load-bearing — a
|
||
> Phase-0 probe (2026-07-22) proved an in-container access through it wakes the idle automount
|
||
> trigger, so narrowing it breaks NAS access itself; (b) there is no `userdata/` layer to scope to,
|
||
> since apps on a share store at `<share>/<app>`; and (c) creating one would write Felhom's directory
|
||
> convention onto a customer's own NAS, which R-67 forbids outright. The browsing surface being
|
||
> immovable is precisely *why* the backup tree must never be placed under it. Tier 2 had already
|
||
> reached the same conclusion for its own targets (`F-6C-1`); R-108 closes the PRIMARY namespace,
|
||
> which was the last remaining route.
|
||
>
|
||
> **Operator ruling 2026-07-30: refuse the placement, keep the browse bind.** `RefuseAsAppNamespace`
|
||
> (`internal/settings/settings.go`) is the single predicate; all placement surfaces consult it. It
|
||
> **fails closed** — `/mnt/felhom-drives` holds both storage kinds in-guest, so a path prefix cannot
|
||
> classify and `Kind` exists only on a REGISTERED path; an unregistered path under that root is
|
||
> therefore un-classifiable and is refused rather than assumed to be a drive.
|
||
>
|
||
> **Nothing was stranded:** zero apps on network storage across all six hub customers including Peti.
|
||
> R-67's browse capability is byte-for-byte unchanged (verified by diffing demo-hp's generated
|
||
> compose before and after the deploy). This also **supersedes** the controller README's "NAS backup
|
||
> locality — decision A" (v0.118.0), which deliberately kept a NAS-resident app's Tier-1 artifacts on
|
||
> the NAS: that case can no longer arise.
|
||
>
|
||
> Evidence: `audits/R108-network-app-namespace-2026-07-30.md`. The pre-fix analysis below is retained
|
||
> verbatim as the record of what was wrong.
|
||
|
||
#### 10.1 (historical) The exposure as it stood before v0.187.0
|
||
|
||
**[FACT] The verification and its result.** Every surface that can read a file was checked:
|
||
|
||
| surface | can it reach `backups/`? | evidence |
|
||
|---|---|---|
|
||
| SMB share creation | **NO** | `sharingResolvePath` (`internal/web/sharing_handlers.go:52-81`) resolves symlinks *before* containment, then refuses any path within `SharingDeniedRoots(root)`; that set covers `<root>/backups` **and** the legacy `<root>/felhom-data` + `<root>/felhom-data/backups` (`internal/stacks/samba.go:46-65`, derived from `ProtectedHDDPaths`, `delete.go:59-78`) — **both namespace shapes** |
|
||
| SMB browse (folder picker) | **NO** | same deny set applied per child (`sharing_handlers.go:521-537`) |
|
||
| SMB `ensureImportShare` (the store-direct bypass) | **NO** | writes one controller-generated constant, `GetImportRoot()` = `<system ns>/userdata/import` (`sharing_handlers.go:564-587`) |
|
||
| FileBrowser — **local drives** | **NO** | the bind is `appbackup.UserdataDir(sp.Path)` only, and the comment says why (`internal/web/handlers.go:2450-2460`) |
|
||
| **FileBrowser — network shares** | ~~**YES**~~ → **NO** (R-108, v0.187.0) | the bind is still the share **ROOT** (`- %s:/srv/%s:rslave`, `handlers.go:2437`) and deliberately so — but **no app namespace, hence no `backups/` tree, can exist on a share**, so the root bind reaches only the customer's own files. The reachability is closed at the PLACEMENT, not at the bind |
|
||
| `.fab` import path validation | **NO** | confined to `<root>/exports` (`handler_export.go:400-408`, `estimate.go:215-217`) |
|
||
| `.fab` browser download | **NO** | name-pattern + parent-must-be-the-staging-dir double guard (`handler_export_download.go:36-45,120-140`) |
|
||
| `/api/debug/*` | **NO** | no file-serving branch (`handler_debug.go:46-92`) |
|
||
| `http.ServeFile` (3 sites) | **NO** | assets only, `filepath.Base`-normalised (`server.go:720-771`) |
|
||
| log bundles | **NO producer found** | no filesystem-walk bundle producer exists on the box; searched `felhom-controller/internal`, `felhom-agent/internal`, `cmd/` |
|
||
| registering the backup dir as a drive | **NO** | the manual add requires `system.IsMountPoint(path)` (`handlers.go:2091-2095`); `<drive>/backups` is not a mount point |
|
||
|
||
**The exposure, end to end.** All six links are source-cited and the precondition is live today:
|
||
|
||
1. A NAS share is registerable as a storage path and lands `Schedulable: true` — **LIVE on demo-hp**:
|
||
`{"path":"/mnt/felhom-drives/Felhom-Share","kind":"network","schedulable":true}`.
|
||
2. `GetSchedulableStoragePaths()` has **no `IsNetwork()` filter** (`internal/settings/settings.go:904-914`),
|
||
so that share appears in the **deploy** dropdown (`handlers.go:462-473`).
|
||
3. The **per-app migrate** target list filters only current / decommissioned / disconnected /
|
||
schedulable — also no network filter (`handlers.go:674-679`).
|
||
4. `handleStorageMigrateApp` does **not** call `refuseNetworkLifecycle`, unlike its whole-namespace
|
||
sibling which does (`storage_handlers.go:397` vs `:410-424`), and `startMigration` has no guard
|
||
either (`internal/stacks/migrate.go:214-255`).
|
||
5. With `HDD_PATH` on the share, `namespaceRoot()` returns it as-is under Model A
|
||
(`internal/backup/backup.go:262-263`), so the app's Tier-1 unit is written to
|
||
`<share>/backups/primary/<app>/compose/app.yaml`.
|
||
6. FileBrowser binds that share at its **root** and serves it with `download: true`
|
||
(`internal/infra/infra.go:326`).
|
||
|
||
**Observed live on demo-hp, in the generated compose — the asymmetry is visible, not inferred:**
|
||
|
||
```
|
||
- /mnt/felhom-drives/nvme-1tb/userdata:/srv/nvme-1tb <- drive: userdata-scoped
|
||
- /mnt/felhom-drives/Felhom-Share:/srv/Felhom-Share:rslave <- network: ROOT-bound
|
||
- /mnt/sys_drive/felhom-data/userdata/import:/srv/beolvasas
|
||
```
|
||
|
||
**Today this is not a secret leak**, because the unit's `app.yaml` is secret-stripped
|
||
(`recovery_unit.go:73`). **D5 would make it one.** That is exactly the test §2 set, and D5 therefore
|
||
does **not** hold as written. → **R-108**
|
||
|
||
> **CLOSED (v0.187.0).** Links 2, 3 and 4 of the chain above are now guarded, and a **fifth** surface
|
||
> the chain did not list was found and guarded too: `handleStorageDecommission` mode=`migrate` checked
|
||
> only `req.Where` (the SOURCE) via `refuseNetworkLifecycle`, so a whole namespace could be
|
||
> decommissioned ONTO a NAS. Link 2's framing also understated the problem — the deploy **dropdown** is
|
||
> only a UI list; the boundary is the deploy **POST** (`internal/api/router.go`), which accepts any
|
||
> caller-supplied `HDD_PATH` and whose only other validation is `os.Stat` existence
|
||
> (`internal/stacks/deploy.go`). Filtering the list alone would have left the surface open. Links 5 and
|
||
> 6 are unchanged and still true — they simply can no longer be reached.
|
||
|
||
### 10.2 The gap register
|
||
|
||
| ID | Gap | Consequence |
|
||
|---|---|---|
|
||
| **R-102** | Tier-2 writes a full `recovery-unit/` mirror on every run and **no code path reads it** | Tier-2 is defeated in the drive-loss scenario it exists for (§7.2). Was C9-F4 |
|
||
| **R-103** | The Tier-2 no-coverage refusal **names** the working action but does not **route** to it | 45-or-43 of 53 apps dead-end at a message. Was C9-F1b |
|
||
| **R-104** | An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach, reported as *„ismeretlen okból"* | the offsite tier stays dead until a human unlocks. Was C9-F3 |
|
||
| **R-105** | Three hub-held DR records are empty on the whole live fleet: `hosts.dr_record_json`, `host_escrow.directive_json`, `dr_recipe.host_half.drives` | the Recipe (§4) is incomplete in exactly the fields host-loss recovery reads. Causes may differ per field |
|
||
| **R-106** | `dr_recipe.host_half.pbs.namespace` records `"root"` on every box | the recorded restore coordinate is wrong; real namespaces are per-customer |
|
||
| **R-107** | ~~No offsite action unpacks the named-volume tars Tier-3 captures on every run~~ — **CLOSED, controller v0.218.0, 2026-08-22** (`volReplay`, proven live on `demo-hp`). True from the day Tier-3 shipped until 2026-08-21. | was: offsite alone cannot rebuild a named-volume app (§7.2). Now: the tars replay; what remains is that the app must be **deployed** (R-253), which is not R-107 |
|
||
| **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) |
|
||
| ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` |
|
||
| **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) |
|
||
| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. **RE-SCOPED 2026-09-01:** the box can delete the LIVE repository but **cannot write to the daily snapshots of it** (measured, both boxes) — so the exposure is at most one day's data plus an operator-driven per-file recovery, not open-ended loss. Detection shipped hub v0.111.0 (R-431). **DRILL 2026-09-01: the re-scope's recovery clause is WITHDRAWN — no snapshot is reachable from a sub-account by any name (R-433, 777,600 names, controlled), so the exposure is NOT bounded by an operator-driven per-file recovery. R-436 is the new cheap lead: the provider already offers `rclone serve restic --stdio` server-side and restic 0.14.0 speaks `rclone:` (measured) — but the client supplies the server command line, so ask the vendor whether `--append-only` is pinned BEFORE building anything.** |
|
||
| ~~R-86~~ | ~~Restore-tests are interval-scheduled, not backup-aligned~~ | **CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0.** Restore-testing is now **per archive generation**: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) |
|
||
| ~~R-353~~ | ~~A local unit restore reported a bare completion whether it returned an entire dataset or nothing~~ | **CLOSED 2026-08-30, controller v0.226.0.** The volume count already existed (`restoreDockerVolumesFrom`) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. `RestoreFromRecoveryUnit` now returns `UnitRestoreResult` and the sentence has three cases, keyed on replayed-vs-**listed** — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no `SafetyDump` discriminator, and §6.3 is why that is not pedantry. Proven live on `demo-hp` |
|
||
| ~~R-357~~ | ~~The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths~~ | **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `writeSafetyDump` and `StopStack`, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. **Not live-validated** — filling a filesystem is a drill step |
|
||
| ~~R-358~~ | ~~`OffboxFullScratchReady` asked "non-empty directory", which is exactly what a failed restic run leaves~~ | **CLOSED 2026-08-30, controller v0.226.0.** A completion marker written only after restic returns nil, stale ones cleared before it starts, both orders pinned by an AST test. Both handlers refuse server-side — a hidden button is not a guard. Proven live on `demo-hp`. **See R-396: the bad state needed only a SUCCESSFUL safe restore, not a failed one** |
|
||
| ~~R-360~~ | ~~The verification-copy delete gated on the concurrency flag, which a verification restore never holds~~ | **CLOSED 2026-08-30, controller v0.226.0.** `restoreOpBlocked()`, matching the five sibling handlers. The doc comment that claimed it already did this is corrected in place — that sentence is why nobody looked. Proven live on `demo-hp` in the exact flag state that produced it |
|
||
| ~~R-396~~ | ~~A unit-only verification restore unlocked the destructive full restore~~ | **CLOSED 2026-08-30, controller v0.226.0** (by R-358's marker). Found while answering R-358's open question. Both restore modes write the SAME scratch directory, and one boolean (`ScratchReady`) drove three different intents — "is there a scratch", "may we place", "may we destructively restore". The safest action on the page unlocked the most dangerous one |
|
||
| ~~R-359~~ | ~~The off-site restic store is never verified by anything, ever~~ | **CLOSED 2026-08-30, controller v0.227.0/v0.227.1.** A daily `offsite-integrity` job on **due-ness, not a weekday**; it takes the single-writer flag and SKIPS rather than waits (`resticStep` escalates to `unlock --remove-all` and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. **⚠ The depth that ships ON does NOT catch silent corruption:** measured, a pack corrupted without a size change returned `no errors were found`, exit 0; only `--read-data*` caught it. Choosing the depth is **R-399** |
|
||
| ~~R-397~~ | ~~`NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller and the product advertised a weekly check that did not exist~~ | **CLOSED 2026-08-30, controller v0.227.0.** Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. `ok` is severity `info` and mails nobody by design |
|
||
| ~~R-399~~ | ~~The check reads the catalogue and never the data~~ | **CLOSED 2026-08-31, controller v0.228.0.** `monitoring.integrity.read_data_subset` now defaults to **`100%`**, so the weekly check downloads and re-hashes every stored byte. **The fact that made it necessary, and the sentence that should stop anyone turning it back down to save four seconds: the structure check PASSED a size-preserving pack corruption.** Measured on `demo-hp` 2026-08-30 — plain `restic check` reported `no errors were found` and exited 0 over a pack damaged without a size change; every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100%. `off` (any case) returns a box to structure depth; an empty value means *not configured*, therefore the default; a malformed value WARNs and falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs an operator WARN naming R-401 — **one data point, on one 134 MB store, so no rotation schedule, size threshold or bandwidth budget was invented from it.** Proven live at both depths 2026-08-31 with the restic argv observed from the guest |
|
||
| ~~**R-87**~~ **— SHIPPED 2026-08-31 (controller v0.231.0), RE-SCOPED. AND IT IS STILL NOT R-359** | ~~The restic tier is never restore-TESTED~~ **The box now PROVES its own off-site copy still HOLDS something.** **THE DISTINCTION, STATED ONCE AND PLAINLY BECAUSE THE GREEN TICK INVITES THE OTHER READING: this proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app.** The proof restores to a throwaway folder, judges it against the unit's own captured compose, and deletes it — it never touches a live app. Putting data back is drill work. **§8 matrix row 4 is deliberately NOT moved.** Nightly at 05:30, one app, due-ness per SNAPSHOT (R-86's model). Evidence `tests/r87-offsite-proof-2026-08-31/`; the reasoning is `audits/SPIKE-restic-restore-test-2026-08-31.md`. | **SPIKE VERDICT, `audits/SPIKE-restic-restore-test-2026-08-31.md`:** build the NARROW version, not the row as written. **Measured:** a scratch restore of all 8 apps costs **25 s / ≤213 MB scratch**, LESS than the 40.3 s weekly check beside it; but restic 0.14.0's `--verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving corruption passed clean, red-proofed), `restic ls --json` carries no content hash, and the unit manifest hashes **4 918 B of a 213 231 242 B unit** — so **no reference for "correct" exists** (R-409). **Of the five drill-found restore defects R-353/354/356/358/403, an unattended scratch-restore would have caught ONE (R-356).** The value is elsewhere and the weekly check structurally cannot reach it: `check` proves the stored bytes are the stored bytes, never that we stored the RIGHT thing — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured 2026-08-31). **Proposed re-scope, Viktor's call:** *prove the off-site snapshot still CONTAINS a recoverable unit* — one app a night, restored to scratch, checked against its own `manifest.json` via the existing `unitCarriesData`. Must use `--no-lock` and skip `unlockStale` (R-95's constraint is otherwise violated — R-407/R-408 record what the path writes today) and must take `acquireRunning`, which `RestoreOffboxScratch` does not. | matrix row 4's route has no unattended proof. **Stated explicitly because the two rows sit next to each other and are easy to conflate: a `check` proves the STORE IS READABLE; a restore-test proves DATA COMES BACK OUT.** v0.227.0 did the first and nothing else. This row is untouched by it |
|
||
|
||
### 10.3 Divergences that are documented elsewhere and are not re-opened here
|
||
|
||
**[FACT]** `06-offsite-connectivity.md:19-21` describes the operator's public edge as a Cloudflare
|
||
Tunnel and states DooPlex has no public IP. Live DNS resolves `hub.felhom.eu` through a no-ip
|
||
DynDNS CNAME straight to DooPlex's own public address (INV Part E.3). That is a topology-document
|
||
issue, not a recovery-model one; recorded so the discrepancy is not lost.
|
||
|
||
---
|
||
|
||
## 11. Open decisions — for the operator
|
||
|
||
**Recorded, deliberately not answered.**
|
||
|
||
**A. Escrow custody.** Split custody (R **or** an offline operator key) versus a 2-of-3 threshold
|
||
across customer / hub / box drives.
|
||
*Recommendation on record:* **split custody, operator key held offline and never in the hub.* Note
|
||
the constraint from §2: the moment the operator key lives in the hub, "a compromised hub yields
|
||
blobs nobody can open" stops being true.
|
||
|
||
**B. Lost-R policy.** Under split custody the operator **can** recover. Is that the stated policy —
|
||
and if so, what identity check gates it? Today matrix row 13 has no route at all, and the customer
|
||
is not told that losing R costs them the host-loss route.
|
||
|
||
**C. RTO / RPO targets per scenario.** **None have ever been stated.** Without them §8 cannot judge
|
||
whether, for example, offsite-only recovery is acceptable for drive loss, or whether the 7-day PBS
|
||
cadence is adequate for row 8. The measured column is the input; the target column does not exist.
|
||
|
||
**D. Hetzner as a single failure domain.** restic (Storage Box `u629488`, sub-accounts per customer)
|
||
and PBS (Cloud server ep0 + a Cloud Volume) are both Hetzner. Whether they share an **account,
|
||
login and payment method** is unverified (INV Unknown U-1) — the question was deliberately not
|
||
answered by calling the provider API with the production token. Accept explicitly, or mitigate.
|
||
|
||
> **The fence was reached again, and held, 2026-09-01 (SPIKE R-95).** The R-95 spike needed one
|
||
> provider-side fact — **do daily snapshots exist on the Storage Box, and how many** — because the
|
||
> register calls that mitigation ARMED and the whole ranking of R-95's options turns on it. It is
|
||
> readable from the box: measured on BOTH machines over their own SFTP credential, with positive
|
||
> and negative controls, **no `.snapshots` is visible to either sub-account**, and the account is
|
||
> jailed at `/`. The register's own confirming field, `size_snapshots`, is an API field. **So the
|
||
> spike stopped here rather than answering it by acting**, and left it as R-429 — ten minutes in
|
||
> the Storage Box panel, and it re-ranks R-95. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1.
|
||
|
||
**E. `local` vzdump shares a physical device with the guest it backs up.** LIVE on both hosts:
|
||
`/var/lib/vz` (the archive target) and `local-lvm` (the guest's rootfs and both `backup=1`
|
||
mountpoints) are both on `/dev/sda3` → VG `pve`. The tier therefore protects against **corruption
|
||
and operator error only**, never against disk failure. Accept and name it honestly in the customer-
|
||
facing description, or move the target.
|
||
|
||
**F (added by §10.1, not in the original list). — ANSWERED 2026-07-30.** D5 could not be adopted
|
||
until R-108 closed. Closing R-108 *was* the intended path, and it is done (controller v0.187.0): the
|
||
operator ruled to refuse app namespaces on network storage rather than narrow the browsing surface,
|
||
because the share-root bind is load-bearing and cannot be scoped. **D5 is no longer blocked.** Whether
|
||
to now *implement* D5 remains an open scheduling decision, not a blocked one.
|
||
|
||
---
|
||
|
||
## 12. Evidence index
|
||
|
||
| Claim | Grade | Source |
|
||
|---|---|---|
|
||
| Tier-2 file restore, gap-fill and after total loss | **PROVEN-LIVE** | CAMPAIGN-9 A1/A3 |
|
||
| Tier-2 refuses without an outage for a no-coverage app | **PROVEN-LIVE** | v0.183.0 replay, `felhom.eu/REPORT.md:60` |
|
||
| Tier-1 unit restore **executes** | **PROVEN-LIVE** | CAMPAIGN-9 A2 |
|
||
| Tier-1 **content recovery after loss** | **UNPROVEN** | `CAMPAIGN-9…:825-827` |
|
||
| restic restore of app data (bytes) | **PROVEN-LIVE** | CAMPAIGN-8 R-87, 155.5 MB, 6/7 byte-identical |
|
||
| offsite reconstitution of a DB-indexed app | **PROVEN-LIVE** | destructive immich drill 2026-07-20, `00-capability-map.md:66,75` |
|
||
| offsite **place-to-live** as a distinct action | **UNPROVEN** | `CAMPAIGN-8…:520` |
|
||
| shares restore (files + definitions + credential) | **PROVEN-LIVE** | 2026-07-18, `00-capability-map.md:96` |
|
||
| `.fab` drive-to-drive round trip | **PROVEN-LIVE** | CAMPAIGN-6D P-FAB, 1.7 GB |
|
||
| `.fab` **browser upload** leg | **UNPROVEN** | `00-capability-map.md:67` |
|
||
| whole-guest restore, local and PBS, exact mount parity | **PROVEN-LIVE** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C |
|
||
| corrupted PBS snapshot fails cleanly | **PROVEN-LIVE** | CAMPAIGN-8 fault 17 |
|
||
| the box cannot delete its own **PBS** snapshots | **PROVEN-LIVE** | CAMPAIGN-8, R-89 |
|
||
| unattended restore-test across tiers | **IMPLEMENTED**; **per-archive due-ness PROVEN-LIVE 2026-08-03** (agent v0.121.0) | `00-capability-map.md:41`; the due verdict + a real offsite run on demo-felhom (§3, Lane 2's per-archive rule) |
|
||
| guest-power watchdog | **PROVEN-LIVE** | agent v0.107.0, 120 s |
|
||
| quiesce crash recovery | **PROVEN-LIVE** | CAMPAIGN-8 fault 10, 1 s, by SIGKILL |
|
||
| break-glass | **PROVEN-LIVE** | `runbooks/break-glass.md` |
|
||
| agent DR bring-up (`ModeDRGuestLoss`) | **NEVER EXECUTED** | `CAMPAIGN-8…:522` |
|
||
| host-loss plan → an actual restore | **EXECUTES NOTHING BY CONSTRUCTION** | `felhom-agent/internal/dr/plan.go:1-4` |
|
||
| host rebuilt as its former self | **NEVER DONE** | INV Part D1 |
|
||
| escrow **consume** in a real recovery | **SPIKE-LEVEL ONLY** | `06-offsite-connectivity.md:327` |
|
||
| hub DB restore from its Longhorn backup | **NEVER DONE** | INV Part D2.2 |
|
||
| a **customer** performing a restore unassisted | **MISSING AS EVIDENCE** | `00-capability-map.md:75` |
|
||
|
||
---
|
||
|
||
## 13. What this document deliberately does not do
|
||
|
||
- It does not restate capability status — §8 cites the map, the map cites §8.
|
||
- It does not resolve the 7/53-vs-9/43/1 count (§6.2). Both are on the record; neither is adopted.
|
||
- It does not estimate a single RTO or RPO. Every blank in §8 is a real gap.
|
||
- It does not answer §11. Those are the operator's.
|
||
- It does not claim ratification.
|
||
| 5 | **A customer had no way to BEGIN — the only route was a command line** | Everything above was self-service, and nothing told the owner of a rebuilt box that a sealed package was waiting or how to open it. | **CLOSED — controller v0.200.0** (R-193). A full-page recovery screen, shown while the hub holds a package this box cannot open; it unlocks and lists, and **restores nothing** (→ R-213 for the put-back). |
|