Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
148 KiB
07 — The recovery model
Status NOT RATIFIED. Ratification is Viktor's review, not an editor's. Written 2026-07-28 (full rewrite; supersedes the 2026-07-14 DRAFT entirely) Verified against controller v0.183.0 · agent v0.110.0 · hub v0.80.0 · catalog 4252121· repo HEADsfelhom.eu ff050cf,felhom-controller fd50a73,felhom-agent d5c7691Live fleet at verification demo-felhom + demo-hp, both guest 9201, both on the versions above Freshness CURRENT as of 2026-07-28. Per standing ruling S-2, mark this STALE the moment it is more than a few trains behind. The document it replaces was DRAFT since 2026-07-14, verified against controller v0.132.0 — fifty-one versions stale — and was cited as authoritative throughout that time. How to read this document. Two kinds of statement appear, and they are always labelled:
- [DESIGN] — a decision taken in the 2026-07-28 architecture discussion (D1–D6, §3–§5, §9). Not derived from code; the code may not implement it yet. Where it does not, §10 says so.
- [FACT] — an observed property, carrying a
file:line, a live command output, or a citation to_recovery-inventory-2026-07-28.md(below: INV).Where the model is silent, this document says OPEN rather than filling the gap.
Primary input:
_recovery-inventory-2026-07-28.md(read-only inventory, 2026-07-28) — cited throughout as INV Part n. Every number in §6, §8 and §11 traces back to it.
1. Purpose and scope
This document describes how a Felhom customer gets their system back, and who can do it.
It replaces a document that described where copies are written. That was the wrong frame, and §7 explains why: a set of copies is not a set of recovery routes, and the previous document's tier table could be entirely satisfied while a real recovery was impossible.
In scope: the recovery model — trust boundary, lanes, the three-part model, the tiers as inputs to recovery, the dependency chain between them, the failure→recovery matrix, and the encryption policy.
Out of scope, deliberately: implementation specs (they live in task specs), the capture-set
algorithm (internal/appbackup/captureset.go and its tests are the source of truth), and per-tier
operational runbooks (documentation/runbooks/).
Authority split. This document is authoritative for the failure→recovery matrix (§8). The
capability map (00-capability-map.md) stays authoritative for per-capability status. Neither
restates the other; §8 rows are cited from the map, not copied into it.
2. The trust model
[DESIGN] The operator holds root SSH on every box. That is a fact of the product — the agent is operator-tier, the host is operator-managed, and break-glass exists precisely so the operator can get in when nothing else works. "The operator cannot read customer data" was never the security property, and no part of this document may be read as claiming it.
[FACT] The mechanics that make this concrete:
- The operator's OOB SSH public key is pushed to every box from the hub
(
hub_settings.oob_operator_ssh_pubkey→/var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op, 93 bytes, LIVE on both hosts — INV Part C, row 21). - The break-glass
root@pamconsole password for every host is stored in the hub and retrievable with the operator's global key (documentation/runbooks/break-glass.md:42-48; LIVE:host_recoveryholds 3 rows — INV Part D2.1). - Guest data is reachable from the host by definition:
pct exec, and the data drives are bind mounts on the host (mp8 /mnt/felhom-drives, LIVEpct config 9201on both hosts).
[DESIGN] What R (the customer's recovery code) does provide, and must keep providing: the hub alone is not enough. A compromised hub yields blobs nobody can open. This property holds only while the operator's own key is not stored in the hub — which is why §11-A is an open decision and why its recommendation on record is "operator key held offline and never in the hub".
[FACT] The escrow is genuinely zero-knowledge today: host_escrow rows carry
posture = zero_knowledge, a 383-byte blob and a 572-byte identity blob, and the hub holds only a
restic_pw_sha256 hash beside them (LIVE, both hosts — INV Part D2.1). R exists in zero
system copies by design (INV Part C, row 24).
So the honest statement of the property is:
The hub cannot open the escrow. The operator can reach a live box. Neither fact substitutes for the other, and R is what keeps the first true.
[DESIGN] Because R is the only key, the household is ASKED for it from the first login (controller
v0.245.0, R-543). Off-site backup is enabled by default but does not RUN until the ceremony is done,
so the ask is not a nicety — it is the step that turns the default-on tier into an actual copy. The
volunteer guide asks for it immediately after the dashboard password and before the first app
(runbooks/VOLUNTEER-first-hour.md §6), and the product repeats the ask on every page until it is
done (§6.1).
3. The two lanes (D1)
[DESIGN] Recovery is split into two lanes with different owners. This is a product decision, not a limitation.
Lane 1 — the customer, unassisted: files and app data
The customer owns their data and can get it back themselves, through the „Visszaállítás" surfaces, with nothing but their dashboard password. No operator, no ticket, no scheduling.
[FACT] What Lane 1 contains today (INV Part A.1 — seven paths, all behind the controller's
RequireAuth gate, which is the customer-owned password: internal/web/auth.go:34-35 puts
settings.json → password_hash ahead of the operator-provisioned controller.yaml value):
| # | Surface (HU) | Endpoint | Semantics |
|---|---|---|---|
| 1 | „Visszaállítás indítása" | POST /backup/restore |
destructive — rebuilds the app from its Tier-1 unit |
| 2 | „Fájlok visszaállítása" | POST /backup/tier2/restore |
additive, missing-only |
| 3 | „Visszaállítás a távoli tárolóból" | POST /backup/offbox/restore |
non-destructive, to a verification copy |
| 4 | „Helyreállítás az élő adatok közé" | POST /backup/offbox/place |
additive, missing-only 2026-08-30 (R-359): the store is now VERIFIED on a cadence — a daily offsite-integrity job runs restic check when the last successful one is over 7 days old. This is a readability check, not a restore-test (R-87 stays open). The depth that ships ON now re-reads 100% of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. 2026-08-31 (SPIKE R-87): a scratch restore of every app on demo-hp was measured at 25 s for 8 snapshots / 774 MB, against 40.3 s for one weekly check — but restic 0.14.0's --verify checks size and mtime, not content, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. audits/SPIKE-restic-restore-test-2026-08-31.md. audits/SPIKE-restic-restore-test-2026-08-31.md. |
| 5 | „Teljes visszaállítás (fájlok + adatbázis)" | POST /backup/offbox/reconstitute |
destructive to files + DB, never deleting |
| 6 | „Megosztások visszaállítása" + place | POST /backup/shares/{restore,place} |
additive |
| 7 | .fab import |
POST → apiImportStart |
destructive re-import of one app |
[FACT] What is proven in Lane 1 (INV Part G.1): paths 2 and 6 are proven live end-to-end;
path 5 is proven live through a destructive drill (photos deleted, trash emptied, restored, timeline
confirmed); path 3 is proven for bytes only; path 7 is proven for the drive-to-drive circle but
its browser-upload leg is not; path 1 is proven to execute but its content recovery after real
loss has never been demonstrated — the single most valuable unproven item in the system
(audits/CAMPAIGN-9-restore-proof-2026-07-28.md:825-827); path 4 has never been exercised as a
distinct action (audits/CAMPAIGN-8-backup-restore-2026-07-27.md:520).
[FACT] One honest residual on the whole lane: every proven run was performed by the operator,
not by a customer. 00-capability-map.md:75 still carries "A customer (not the operator) performs a
restore via UI alone" as MISSING as evidence — by absence of the run, not by a product gap.
[RULING 2026-09-17, operator — a design choice REVERSED] The restore record survives a restart
(R-550, controller v0.246.0). The async restore status (internal/backup/opstatus.go) was in memory
by choice — „same precedent as notification cooldowns". Chaos night round 10 measured the cost: a
restore accepted, the box hard-reset four seconds later, and the status answering the Go zero value, so
the household could never learn whether it finished. For the restore record only, it is now written
atomically to restore-status.json in the controller's state directory at both ends of an op; at
startup a record still marked running becomes a failed, interrupted result („A visszaállítás megszakadt
(a doboz újraindult) — indítsd el újra."), shown per app on the restore page until that app's next
restore, and raised once as restore_interrupted. Notification cooldowns stay in memory — the
precedent is kept for what it was written for.
A backup RUN cut off by a power cut or a restart is said (R-519, controller v0.296.0, 09 decision 123). The
same shape for the app-data run: appdata-run.json beside the restore record, written at both ends of a run. A start
that finds it still running turns it into a notice on /backups and /backups/apps („A legutóbbi mentés (…) megszakadt,
mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every step OK. Measured BIGNIGHT F2 (2026-09-14):
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh .sql
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
Live: PROVEN 2026-10-05 on 9202 (operator ruling 126): a run cut by a controller restart between bookstack's config
dump and its database volume → both pages carry the notice, the restore point reads the older volume's time, the next
complete run clears it (audits/hub-db-offsite-2026-10-05/partD/r519/; R-519 closed).
Lane 2 — the operator: guest and host recovery
Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are operator work, by design and by contract. They are not customer-facing and are not going to be.
[FACT] What Lane 2 contains (INV Part A.2): the scheduled restore-test, pct restore, raw
proxmox-backup-client restore, raw restic restore, the agent's DR bring-up
(--selftest=bring-up -mode dr), the host-loss plan builder, and the escrow-consume ceremony. Each
needs root on the host or a CLI flag; none is reachable from any customer surface.
[DESIGN + FACT, 2026-10-04, R-834] A whole-guest restore BESIDE a live original never comes up as a second box.
A restore that does not replace the box's own guest on a replaced host ends with onboot 0 and no host-path bind
(mp8 = the household's drives, mp9 = the original's bootstrap); the DR route on a replaced host keeps its binds,
because there they are right. Per route: the restore-test sets onboot=0 at create and gives mp8/mp9 throwaway
volumes (measured live on demo-hp, every poll); the DR bring-up REFUSES when the archive's source guest, or any guest
binding the drives, is still on the host (agent v0.139.0, proven live); the hand route is
scripts/felhom-restore-beside.sh (proven live); provisioning restores the golden, which has no binds. Pinned by
TestRunBringUp_DRRefusesBesideALiveOriginal, TestRestoreTest_NoHostPathBindBesideTheOriginal and
scripts/test_felhom_restore_beside.py. audits/backup-close-2026-10-04/partA/.
[FACT] What is proven in Lane 2 (INV Parts G.1, G.3): whole-guest pct restore from both tiers
is proven with exact mount parity; the unattended restore-test is proven and currently running on
both boxes; a corrupted snapshot is proven to fail cleanly. The DR bring-up path has never been
executed (CAMPAIGN-8…:522), the host-loss plan executes nothing by construction
(felhom-agent/internal/dr/plan.go:1-4), and no host has ever been rebuilt as its former self
(INV Part D1).
Lane 2's restore-test is scheduled PER ARCHIVE GENERATION (R-86, 2026-08-03)
[CONTRACT, changed 2026-08-03 — agent v0.121.0 + hub v0.91.0.] The scheduled restore-test used to fire on an interval started at daemon start. It no longer does. The rule is:
Let A be the newest archive on a tier that has settled for at least the settle lag (24 h). The tier is DUE when A exists and A has not already been proven.
So a tier is proved once per archive, on its own archive, and the proof follows the backup rather than the process's uptime:
| tier rhythm | what is proved, and when |
|---|---|
| daily (host tier) | yesterday's archive, once a day |
| weekly (offsite tier) | last week's archive, once a week |
| newborn (no archive yet) | nothing — UNKNOWN, never a fault |
The trap in the obvious formulation, recorded so it is not reintroduced: "due when the newest archive is ≥ 24 h old" is never true on a daily tier — a new archive resets the newest-archive age to zero long before it reaches the lag — so the literal reading silently switches restore-testing off for the tier that matters most.
What survives unchanged: the restore-test itself (restore → boot → verify → destroy the scratch), its journal and crash recovery, the scratch VMID band, the one-heavy-operation gate, proof credit only on success, and oldest-proven ordering, which is now the tie-break between due tiers. A ticker remains, but only as the evaluation interval (6 h by default, chosen from a measured cost: one due-check is 18 ms on a local dir storage and 392 ms on the PBS tier over the WAN).
The hub's half is not optional. restoreProvenStaleAfter was a flat 7 days derived from the very
cadence this replaced, and a weekly tier proved weekly reaches a proof age of exactly one interval
just before its next proof — 168 h against a 168 h window. It sat ON the line, so any ordinary delay
tipped a healthy tier into a nightly alarm. The window is now per tier, from that tier's observed
archive interval, floored at the old 7 days, capped at 12 days (strictly inside the two-week offsite
retention), and falling back to the tier's declared rhythm when history is too short to observe one.
Why the split is right, stated once
[DESIGN] A customer can reason about "my photos are gone". A customer cannot reason about
"restore the LXC and re-attach mp8 by durable_id without binding another guest's bootstrap
credentials" — a hazard serious enough to have its own runbook page
(runbooks/RUNBOOK-manual-guest-restore.md:45-48, F-OPS). Putting Lane 2 behind the operator is
what lets Lane 1 be a button instead of a procedure.
[FACT] And as of D5 (v0.188.0) the split is real, not just intended. Until then Lane 1 only looked
independent: an app's files sat on the customer's drive but could not be brought back without secrets
that lived solely in the guest, so the fast customer lane was silently conditioned on the slow
operator lane (§7.1 leg 1). The unit now carries those secrets, so Lane 1 needs the drive and nothing
else for app rebuild (§7.4). The one honest caveat: Lane 1's file-merge surface (Tier-2 additive)
still reads its destination from the guest's settings.json, so that surface remains guest-conditioned
— D5 removed the secrets leg, not the living-app leg.
4. The three-part model (D4)
[DESIGN] Recovery needs three things, and they live in three different places. Losing one is a different problem from losing another, and the matrix in §8 is organised around that.
| Part | Holds | Where | Lost when |
|---|---|---|---|
| Recipe | scaffolding — guest sizing, drive inventory, PVE storage definitions, PBS coordinates, app inventory + storage bindings. No secrets, no bytes. | the hub | the hub is lost |
| Escrow | the host identity key material and the restic repo password | the hub, R-wrapped | the hub is lost, or R is lost |
| Bytes | Tier-1 / Tier-2 / Tier-3 / whole-guest | the tiers | the relevant medium is lost |
[FACT] The Recipe exists and is current. dr_recipe holds 6 rows; demo-felhom's was updated
2026-07-28 17:31:13 and demo-hp's 17:25:31 — i.e. within one report cycle (LIVE, INV Part D2).
It carries guest sizing, pve_storage[], and an app half with per-app storage_bindings.
[FACT] The Escrow exists and is zero-knowledge. host_escrow: 2 rows, posture zero_knowledge;
the identity bundle's shape is {tunnel_token, pbs_token, wg_private_key, restic_repo_password}
(felhom-agent/internal/escrow/identity.go:26-39).
[DESIGN] Off-site deletion custody (decisions 68–69, 2026-10-03) — BUILT hub v0.127.0 / controller v0.289.1, live on both demo boxes. The restic repository password stays on the box only. The box's off-site key is append-only (pinned in authorized_keys, written by the hub — the key registrar); the box never receives the sub-account password, which the hub stores encrypted at rest. Old snapshots are pruned by the box itself, only in a weekly window the hub opens, behind a fake-snapshot guard (R-822); until that ships nothing prunes. A Felhom-side pruner holding repository passwords is rejected.
[FACT] Three parts of the Recipe are empty or wrong on the live fleet, and they are exactly the
parts a host-loss recovery would read (INV Part D2.3): hosts.dr_record_json is {} on all three
hosts; host_escrow.directive_json is {} on both escrowed hosts; dr_recipe.host_half.drives is
[] on every customer including two with enrolled data drives; and dr_recipe.host_half.pbs.namespace
reads "root" while the real namespaces are demo-felhom / demo-hp. → R-105, R-106.
5. Encryption policy (D2)
[DESIGN] Encryption follows the boundary, not the tier.
On the customer's own premises: plaintext
Tier-1 recovery units, Tier-2 cross-drive copies, the app data on the drives, and the local vzdump archives on the host are unencrypted, deliberately.
The reasoning, recorded so nobody "hardens" this later:
- It buys nothing against a real threat. Someone who can take the second drive can take the first. Local encryption defends against a threat model — physical theft of only the backup medium — that does not describe a home server where both media sit in the same box.
- It adds a key-loss path that turns a working backup into a brick. Every local encryption key is one more thing that must survive the disaster it exists for, and the system already has one such dependency it is trying to reduce (§7).
- It would break browsing, which is a feature. FileBrowser and SMB let the household see and use their own files. An encrypted local copy is not browsable, and the „Megosztás" and FileBrowser surfaces are part of the product, not an accident.
[FACT] The local plaintext posture is real and verifiable: dir: local in /etc/pve/storage.cfg
carries no encryption-key (LIVE, both hosts), so the daily whole-guest archive is a plain
.tar.zst — 5.93 GB on demo-felhom, 1.68 GB on demo-hp (LIVE, INV Part B.3). That archive is the
single highest-value object on the box: it contains encryption.key, the restic repo_password,
the offbox ssh_key, settings.json and controller.yaml (all under /var/lib/docker = mp0,
backup=1). Naming that plainly is part of the policy, not an argument against it.
Leaving the premises: encrypted, and the provider must not be able to read it
[FACT] Both offsite tiers encrypt client-side:
- restic (Tier-3): repo password is a 256-bit hex value generated once on the box and never
logged (
internal/backup/offbox.go:392-410); restic is invoked withRESTIC_PASSWORD_FILE(:608). - PBS (whole-guest offsite):
/etc/pve/storage.cfgcarries a per-customerencryption-keyand the fingerprints differ between demo-felhom and demo-hp (LIVE). The live restore command line observed on demo-felhom carries--crypt-mode=encrypt(INV Part A.2.3). Per-tenant encryption is why cross-customer dedup is impossible on the shared datastore — a cost consequence, recorded, not a defect.
The one exception, stated because it is not covered by either rule
[FACT] The .fab bundle writes app.yaml with decrypted plaintext secrets, deliberately,
and its password is optional (internal/appexport/export.go:484,506-511; the generated file's
first line is literally # Exported by felhom-controller — plaintext secrets, :531;
Encrypted: req.Password != "", :307). A .fab can also be written to a registered network
share, because storageDriveList() does not filter network paths
(internal/web/handler_export.go:377-387). This is a portability artifact, not a tier — but it is
the one place where a customer action can put every secret of one app onto a NAS in plaintext.
Recorded here so the encryption policy is not read as covering it. → R-108 (same root cause).
6. The tiers — what each captures
The tiers are inputs to recovery, not recovery routes. §7 and §8 say what they can actually do.
[RULING 2026-09-16, operator] Tier 3 (off-site) is ON for every new customer — shared, 100 GB soft
quota. Opting a customer out stays the per-customer exception, as DR-tier-by-default has been since
2026-07-12. The reason is this document's own [FACT] two sections down: the whole-guest tiers do
not carry mp8 /mnt/felhom-drives, and a Tier-1 unit has no file leg — so for the four class-A apps
a one-drive box with Tier 3 off keeps the household's files in NO tier at all. That was measured on a
fresh box the same day: five photos put into Nextcloud, deleted, restored from the box's own backup,
and none of them opened (audits/DRILL-prove-fixes-0243-2026-09-16.md, R-537/R-538). Implemented in
hub v0.116.0 (handleConfigNewForm). What this ruling does NOT change: what any tier captures.
[RULING 2026-09-30, operator — 09 §3 decision 50] …and every app on such a box is IN the off-site copy by
default. The 2026-09-16 ruling turned the tier on per customer; a per-app switch that started OFF left a new
household's apps out of it (measured on a fresh box, audits/DRILL-new-household-2026-09-30.md, R-720). A newly
installed app on a box whose customer has off-site is switched on at install. Over the customer's quota the page names
the largest apps and the household chooses; the box never deletes off-site history to make room without that
choice. Apps installed before the release are offered one press, never switched by the release itself.
The app update's safety precondition accepts ANY tier (operator ruling 2026-09-13, controller
v0.239.0, R-475): the first fresh copy in the order Tier 2, Tier 1, Tier 3; with none, it backs up
first. Design: 09-update-architecture.md §3 decision 8. What a Tier-1 route back restores is only
what the unit holds — for an app whose data is a bind mount that is the definition, not the data
(R-479). Removal and the tiers (controller v0.240.0): removing an app with its backups KEPT keeps
the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás" still works afterwards
(R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and
never touches off-site snapshots (R-474). A removed app whose unit was kept is listed on both local backup pages with its restore since controller v0.242.0 (R-487): the local lists are keyed on the drives, not on what is deployed — the rule R-237 set for the off-site list — and the restore opens the unit where it sits, a data drive included. For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.
[FACT, measured from source 2026-09-24, controller v0.267.0/v0.268.0; the second-drive cell for file apps changed by v0.269.0] Which copy brings an app back WHOLE
— read from the restores' own refusals, not from what the tier stores (09 §3 decision 25, R-659):
| app | own unit (Tier 1) | second drive (Tier 2) | off-site (Tier 3) |
|---|---|---|---|
| no declared drive files (class B, 45 apps) | whole — the unit restore | whole — „Teljes visszaállítás" (the unit restore) | whole — the full restore |
declared drive files (DeclaredDriveFileLegs: calibre-web, immich, nextcloud, paperless-ngx) |
not — refused (R-538) | whole since controller v0.269.0 — „Teljes visszaállítás" runs RestoreTier2Whole: the mirror's files by four rules (never delete a file; never write an older file over a newer one; bring back every missing one; an older, different live file is replaced and KEPT beside as <name>.felhom-<UTC>), then the unit from the mirror (decision 26, R-661; proven live 2026-09-24, audits/night-2026-09-24/A1/). Needs a proven, openable mirror WITH the file legs and 2 GB free; rsyncMirror (--delete) is never used for it. The unit-only restore is still refused (R-538). |
whole — „Teljes visszaállítás (fájlok + adatbázis)" |
backup.WholeOnTier is this table; TestR659_TruthTableAgreesWithTheRestoresRefusal pins that it cannot drift
from the refusal. An app with only OPTIONAL legs (audiobookshelf, komga, romm) is class-B here: its files are
never moved by an update or an undo, and the unit restore accepts it.
6.1 The four tiers, as configured on the live fleet
| Tier | Location | Captures | Cadence (LIVE) | Retention (LIVE) | Encrypted |
|---|---|---|---|---|---|
| Tier-1 recovery unit | <nsRoot>/backups/primary/<app>/ on the app's own drive |
compose + .felhom.yml + secret-stripped app.yaml, db-dumps/*.sql, volume-dumps/*.tar, manifest.json |
nightly at W, plus a checksum-gated refresh on the periodic status pass | one point per app (restore_points.go:14-18) |
no |
| Tier-2 cross-drive | <target nsRoot>/backups/secondary/<app>/ |
always a full mirror of the unit (tier2.go:368-369) plus mandatory + optional file legs, v2 layout hdd/<rel> + userdata/<rel> |
nightly at W+60m | mirror (rsync) | no |
| Tier-3 restic offsite | Hetzner Storage Box over SFTP, one multi-path snapshot per app per run | the unit plus mandatory legs only; a separate _shares snapshot |
nightly at W+105m | --keep-daily 7 --keep-weekly 4 --keep-monthly 6 |
yes (restic) |
| Plane-2 whole-guest, local | local: → /var/lib/vz/dump on the host |
rootfs + mp0 /var/lib/docker + mp1 /mnt/sys_drive |
24 h (backup_cadence_seconds: 0) |
local_backup_retention: 3 |
no |
| Plane-2 whole-guest, offsite | felhom-pbs: → ep0 datastore felhom-offsite, per-customer namespace, over WireGuard |
same contents | 7 days (604800) |
server-side prune on ep0, keep-last 2 at 03:30 — and the box asks for none (R-191) |
yes (per-customer key) |
[FACT] Tier-3 has a fifth state the table above does not show: PAUSED (R-543, controller v0.245.0). Off-site is ON by default from hub v0.116.0, and a run does not start until the household has performed the escrow ceremony —
tier3Statecalls thisescrow_pendingand the page says „Kulcsletétre vár". This is the design, not a defect: the escrow is zero-knowledge (§2), the household's recovery code is the only key, and a run started without one would write a copy nobody could ever open. What was wrong until v0.245.0 is that nothing asked the household for the code, so a fresh box could sit paused indefinitely while its Tier-1 row promised that the off-site copy protected the app's files. Since v0.245.0 every dashboard page carries the reminder (the R-241 bar, second instance) and the Tier-1 sentence renders by state — „védené … szünetel" while paused. Measured on a fresh box 2026-09-16: zero snapshots, and the page said the files were protected.
R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force. The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's
keep_last: 2on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused —whole_guest_backup_failedin the operator's inbox about a backup that had already succeeded. Fixed in installer 1.25.0 (keep_last: 0) and on both live boxes; a gate now asserts it. Verified before changing it: ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.
[FACT, 2026-10-04 — agent v0.140.0, 11-os-updates.md §8.1] The OS leg closes the night. After the
whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent
waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a
backup or a restore-test; at most once per 20 h. The backup minutes old is the guest's undo (no snapshot is possible,
R-837). A failed or missed backup → no OS leg that night. [FACT, 2026-10-04 — agent v0.141.1, 11 §8.2] On an
appliance, the HOST step follows under the same gate: after a healthy guest step only (a failed or unhealthy guest
step skips it), Debian-origin fixes only, never a kernel, boot or firmware package, never a reboot. Measured: both
steps with nothing to install, 23–32 s; a 108-package host pass, 70 s. [FACT, 2026-10-04 — agent v0.142.0, 11 §5.8] On a RING-0 box a third step
follows, still under the gate: the guest's Docker engine set (live-restore turned on first, once, by reload; every
container must keep its id). Ring 1 never takes it in the night leg. Measured: the three steps with nothing to install,
38–51 s; a six-package Docker step, ~50 s.
[FACT] The three nightly legs derive from one customer-settable window start W at fixed
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
(cmd/controller/main.go:604-607). Both boxes run W = 02:30.
[FACT, controller v0.271.0, 2026-09-25] A fourth leg: automatic app updates, CHAINED after the off-site
leg (09 §3 decisions 11 and 20, §6.4 part 7). The offbox-backup job runs the update leg when its off-site
half ends, on every path (no target, ok, failed, panicked); a box without an off-site tier runs it at W+105m. The
leg starts no step at or after W+5h, and the whole-guest gate — open [W+2h, W+6h) — defers while the leg
runs until W+5h, then only for a step already in flight, never past W+5h30m (decision 31). So on an update
night the whole-guest backup starts later inside its own window; it is never skipped for an update. Live
numbers: audits/DRILL-night-2026-09-25.md Part D.
[FACT, 2026-09-25] demo-hp's local whole-guest target is its host ROOT disk; retention there is now keep-last=1 (operator ruling 2026-09-24, option A; R-684). PVE prunes AFTER a successful backup, so the target must hold keep-last + 1 archives during the run — measured: a 9201 archive is 8.18 GB, the root disk 39 GB. The warning that should precede such a failure is R-685.
[FACT] What the whole-guest tiers do NOT carry. mp8 /mnt/felhom-drives and
mp9 /etc/felhom-bootstrap are host bind mounts and are out of vzdump scope entirely (LIVE
pct config 9201, both hosts). So the whole-guest tiers carry the guest and none of the customer's
data drives — 916 GB on demo-felhom, 938 GB on demo-hp.
6.1.1 A box that is not always on — the catch-up, the banner, the alarms (R-871..R-874)
Why this section exists (R-871): until 2026-10-05 no architecture document covered a box that is OFF at its window W. Tester 2 is a laptop switched off at night (operator, 2026-10-05). Measured (
audits/night-fixes-2026-10-05/ partF/FINDINGS.md, and live on 9202audits/catchup-2026-10-05/partA-spike/): the controller's daily jobs always schedule the NEXT future time, so a missed 02:30/03:30/04:15 waited for the next night for ever; no alarm fired while the box was down at 05:00; the household was mailed "your server cannot be reached" every night.
[DESIGN — ruled 2026-10-05, 09 decision 109] A missed night runs ONCE when the box comes back. Built controller
v0.295.0 (internal/nightchain):
- What a box does when it was off at W. On a controller START and on a host RESUME, the controller asks its
night ledger (
<data>/night-ledger.json) which backup legs missed their last scheduled time. The ledger records when each leg last RAN TO ITS END — an attempt record, used only to decide "missed", never as evidence a backup exists (R-100's rule). A leg that ran and failed was NOT missed: failures have their own alarms. - What the catch-up runs: the three BACKUP legs, in the night's order — database dump, second copy, off-site copy
— through the SAME wrapped leg bodies the scheduled jobs run (
main.gowithLeg,catchUpLegs). - What it does not run: the app-update leg and every Docker step (they restart apps; they wait for a real night). The OS fast lane is the agent's and follows a whole-guest backup, not the catch-up (below).
- When: 15 minutes after the trigger (apps settle; a box switched on and off again at once does nothing). A leg whose own next scheduled time is under 30 minutes away is left to its normal run.
- Only once: several missed nights = one catch-up (the question is about the LAST scheduled time). A normal night followed by a daytime restart = none. A power cut in the middle of the chain = only the legs that did not end.
- Never two at once: every leg, scheduled or catch-up, holds one lock; a second trigger while one is pending does nothing.
- The whole-guest backup. It keeps its own agent-side behaviour (inside [W+2h, W+6h), or the 48 h safety valve —
about one every 2 days for an evening-only box). It CAN collide with a catch-up (the valve fires on the first
5-minute poll after a start), so each waits for the other: a scheduled quiesce defers while a catch-up runs
(
quiesceSetCatchUpFn), and a catch-up waits up to 2 h while a quiesce holds the apps. - A suspended host (a laptop lid): Go's timers run on CLOCK_MONOTONIC, which does not count suspended time, so
the 02:30 timer would fire hours late — and would start the app-update leg at noon. Two rules: a daily job whose
timer fires more than 60 minutes after its wall-clock time is SKIPPED (
scheduler.DailyLateLimit), and a resume watch (wall clock vs monotonic clock, checked every minute) triggers the catch-up. Reasoned from the Go and Linux clock semantics and unit-tested with injected clocks; not measured live (no suspend was allowed on a demo box). - The first start on this release seeds the ledger: nothing before it counts as missed (no surprise catch-up on every box at the upgrade).
- What the household sees: one timeline line — "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 02:30-kor, a
mentés most elkészült." / "Missed backup made now: the box was off at 02:30, so the backup ran when it came back on."
(event
backup_catchup_done, info: recorded, never mailed).
[DESIGN — ruled 2026-10-05, 09 decision 110] The banner (the operator's idea). On every page of a logged-in
household: when the last daily backup (the database dump, and the off-site copy when configured) is over 26 h old —
when the last one was, "the box was off at backup time (02:30)" when the box's own record says so, and a suggested
time. The record is the controller's system-metrics table (one sample a minute, kept 30 days): "on" at W = a sample
within 5 minutes of it; "usually on" in an hour = on in it on at least 5 of the last 7 days. It never changes the
time — a button opens the backup-time setting. The household can close it: it stays closed until the NEXT missed
backup time (durable, in the ledger) and disappears by itself after a successful night.
Decided by CC — operator may reverse (09 decisions 112–116): the 15-minute delay and 30-minute leave-to-normal
line; the 60-minute late-fire limit; 26 h as the banner's line (one night + the chain's two hours, = the hub's
backupStaleAfter); the suggestion = the LATEST hour H with H, H+1, H+2 usually on (a later evening disturbs least),
none when the current window is already usually on; the catch-up's line on the household's timeline but no mail.
The operator's side (hub v0.134.0, 08 §6.4). R-872: a box DOWN at the 05:00 deadline is no longer skipped — it
is judged on longer lines (48 h without a dump, 72 h without a whole-guest backup), so a box that died last night
raises only its staleness alarm, and a box off at every deadline raises expected_dbdump_missed /
expected_backup_missed. R-873: the household hears "your server cannot be reached" at most once per 7 days (the
operator still gets every edge). R-874 (agent v0.145.0): the restore-test's first due-check runs 30 minutes after the
agent starts, so a box with short power-on sessions is still restore-tested.
[FACT, 2026-10-05] Proven live (audits/catchup-2026-10-05/partA/): 9202 off across 09:35, on at 07:38 UTC → the
catch-up at 07:53:09 made the dump (21 s); a crash of its host mid-wait → at the next start ONE new catch-up made all
three legs (22 s); demo-felhom's controller off across 10:07 → the dump made at 08:25:03, 15 min after the start, and
the household's timeline line reached the hub. What the household may notice: the dump leg stops an app with a
volume for its copy — measured 1 s for opengist (182.5 KB), the same as at night; a large volume takes longer, in the
day (R-878).
Edge cases, stated. A box on only in the day with W at night: the catch-up makes the backups every day it is switched on (after 15 min); the banner suggests an evening time once a pattern exists. A box switched on during the chain: legs already past are made up, legs still ahead run normally. A box whose controller restarts in the day after a normal night: nothing. A box off for weeks: one catch-up when it returns; the hub's staleness alarm has been running all along. An upgrade from an older release: the ledger is seeded, the first missed night after it is the first one made up.
6.2 Coverage per app class — and an unresolved count
[FACT] Of 53 catalog templates, 52 keep data in Docker named volumes; exactly 13 carry a
backup: block, and those 13 are exactly the templates that bind ${HDD_PATH} / ${USERDATA_PATH} /
${IMPORT_PATH} at all; 14 have a database service (INV Part B.1, computed against catalog
4252121).
Applying the classifier's documented two-level default (internal/appbackup/classify.go: explicit
entry wins; else a :ro reader is excluded; else a writable bind is mandatory):
| tier filter | templates with at least one file leg |
|---|---|
Tier-3 (mandatory only) |
4 — calibre-web, immich, nextcloud, paperless-ngx |
Tier-2 (mandatory + optional) |
7 — the 4 above plus audiobookshelf, komga, romm |
| legacy resolver path (apps with no block) | 0 — no non-block template binds a namespace path |
✔ RESOLVED 2026-08-31 (controller v0.229.0) — A = 7 · B = 45 · C = 1
source count verdict C9-F1 Phase 0, as shipped A = 9 / B = 43 / C = 1 — felhom.eu/REPORT.md:17-24(since overwritten), restatedbacklog/OPEN-ITEMS.md:31WRONG by two apps INV Part B.1, independent enumeration at catalog 4252121A = 7 / B = 45 / C = 1 CORRECT Measured at catalogue 459766cb16395fd1d1a66282f5cc6da59ead5924, 2026-08-31A = 7 / B = 45 / C = 1 adopted The method, so it can be re-run rather than re-argued. A throwaway
maininside the controller module drove the PRODUCTION rule over all 53 template directories —stacks.LoadMetadata(the single validation choke point, so a rejectedbackup:block degrades to legacy exactly as it does live) →stacks.ParseComposeClassifiableBinds→appbackup.ClassifyBinds→appbackup.ComputeCaptureSetatTierSecondary, with the legacy branch falling back toAppDataBindsPresent+AppDataDirNamesasbackup.tier2CaptureSetdoes. A = at least one leg survives that pipeline. 13 templates carry a validbackup:block; 40 are legacy and none of them binds a namespace path, so all 40 are B or C.A (7): audiobookshelf, calibre-web, immich, komga, nextcloud, paperless-ngx, romm. C (1): bentopdf — it declares no
volumes:key and no${…_PATH}bind at all. B (45): everything else.How the earlier disagreement arose, established rather than guessed. Phase 0's own write-up (controller
CHANGELOG.md, v0.183.0) says four apps — plex, jellyfin, emby, navidrome — are in B "only because their single bind is a:romedia mount, whichClassifyBindscorrectly excludes". It applied the:rodefault rule. The two apps it therefore missed are radarr and sonarr: their${USERDATA_PATH}/media/*and${USERDATA_PATH}/downloadsbinds are writable, so the:rorule does not reach them, and they are excluded by an explicitclass: excludedentry instead. 9 − 2 = 7 and 43 + 2 = 45, which is exactly the gap. No catalogue file was changed; this is a measurement.
[FACT] The class-B consequence is real regardless of which count is right. For an app whose data
lives entirely in named volumes, the Tier-2 copy holds a full recovery-unit/ and no readable
leg — LIVE on demo-felhom, backups/secondary/bookstack/ and .../docmost/ contain
recovery-unit and nothing else, at 156 MB and 86 MB (INV Part B.2). Since v0.183.0 the button
refuses before stopping the app and names the action that works.
6.3 Where capture and restore disagree
[FACT] Three asymmetries, each source-cited:
| tier | captured | read back by that tier's restore | gap |
|---|---|---|---|
| Tier-1 | unit incl. volume tars + DB dumps | all of it | none |
| Tier-2 | unit mirror + file legs | the file restore reads hdd/ and userdata/; since controller v0.218.0's Tier-3 sibling and now v0.229.0, a SECOND action reads the unit mirror itself (RestoreTier2Unit → RestoreFromRecoveryUnitAt, tier2_restore.go) |
CLOSED — R-102 |
| Tier-3 | unit (incl. volume tars) + mandatory legs | files + DB replay + the named-volume tars, replayed from the scratch unit (offbox_reconstitute.go volReplay, controller v0.218.0); the unit itself is still skipped on the way to live (offbox_reconstitute.go:284-289; placed only if the live unit is absent, offbox_restore.go:352-356) |
CLOSED — R-107 |
[FACT] 2026-08-22 — the Tier-3 row above was corrected; the Tier-2 row was NOT. Until controller
v0.218.0 this table said "no offsite action unpacks the named-volume tars it captures", and that
was true from the day Tier-3 shipped until 2026-08-21. R-107 closed in v0.218.0: volReplay
(offbox_reconstitute.go) replays the scratch unit's volume-dumps/ into the live named volumes,
proven live on demo-hp. The old sentence is kept here, in the past tense, because a correction that
erases what was believed leaves the next reader no way to tell a fixed gap from one that was never
noticed.
[FACT] 2026-08-31 — the Tier-2 half is now closed too (R-102, controller v0.229.0), and the old
sentence is kept here in the past tense for the reason the paragraph above gives. Until v0.229.0 this
table said of Tier-2: "the unit mirror is read by nothing — RecoveryUnitPath resolves to
backups/primary/ (appbackup/paths.go:46-48)", and that was true from the day Tier-2 shipped until
2026-08-31. The mechanism was a hard-coded primary segment: every reader of a recovery unit could
only name a path under it. appbackup now also exposes four unit-directory-relative primitives,
Manager.RestoreFromRecoveryUnitAt(stack, unitDir) holds the restore body, and RestoreTier2Unit
points it at <dest>/backups/secondary/<app>/recovery-unit/. Proven live on demo-hp with the
primary unit moved aside — audits/DRILL-r102-tier2-unit-2026-08-31/.
Two things did NOT change, and both are load-bearing. THE SOURCE MOVED; THE DESTINATION DID NOT —
data still lands in the live named volumes and the live database container. And the FILE restore's
reach is unchanged: CanRestore() still answers only "are there file legs?", and
tier2UnitNotCoveredMsg is still appended where that restore runs, so a clean file result never reads
as a clean bill of health for the database.
[FACT] Tier-2's gap was the sharper one because of when it bit: Tier-2 exists for the case where the primary drive is lost — and in exactly that case the primary unit is gone while this mirror survives on the second drive. Until v0.229.0 it was unreachable by any customer action.
[FACT] 2026-08-23 — taking the undo copy used to DESTROY the app's own database backup (R-361).
writeSafetyDump called DumpOne into the app's own unit directory and renamed the result to
pre-restore-* afterwards. DumpOne writes <stack>-<dbtype>.sql — the app's canonical dump, the
name the replay loop matches exactly — so every safety dump overwrote the app's real backup and then
moved it away, leaving the app with no database backup of its own until the next nightly run. A local
restore-from-unit in that window finds no .sql and tells the customer the app never had a database.
The comment beside it asserted the rename meant it "can never overwrite the app's real dump"; it was
false as written and stood for four months. Measured before the fix on demo-hp 2026-08-22:
docmost and bookstack each held only pre-restore-* files and no canonical dump.
Fixed in controller v0.221.0: DumpOneTo takes an explicit final path and derives its own .tmp
from it, so neither the destination nor the scratch file can collide with a nightly dump running
beside it. Proven the only way it can be — the canonical dump's sha256, unchanged across a
restore, on both engines.
[DESIGN] 2026-08-23 — db_dumps lists the app's OWN dumps, not the undo copies. They are local
material for a restore that went wrong, not part of the app's recovery set: nothing reads them from
the manifest, and three per app were being pushed off-site permanently for no recovery value. The
files are neither deleted nor hidden — their visibility is a separate recorded decision and it stands.
One consequence, recorded because it bit within minutes: a stable db_dumps lets
CaptureRecoveryUnit's already-current early return fire, so anything that must happen on every
capture — such as bounding the undo copies — has to sit ABOVE that check, not after it.
[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold. Recorded here rather than only in a closed register row, because a decision that survives only inside a closed work item is a decision nobody will find.
- Replay. The snapshot's
.sqlis imported into the app's database, with only the DB service up (R-47). A pre-restore copy of the LIVE database was taken first and is on disk (R-43); the restore refuses outright if it could not be taken. - Rollback (controller v0.220.0, R-379). If the replay fails, the product re-applies that undo
copy itself. The whole set for this run — an app with two databases gets two undo files, and
restoring only the first would leave the other half-written — matched on the run's own stamp,
never on the
pre-restore-prefix, because several runs' copies coexist in the same directory. It runs with the DB service still up and before any restart, so the app never observes the half state, and into a re-discovered container: the DB-only start re-creates it, so the id captured at dump time is dead by rollback time (v0.220.1, found by the first live run). The app then starts and the customer is told both that the restore failed and that their data is as it was. - Hold (v0.220.0, operator ruling 2026-08-22). If the rollback ALSO fails, the app is held
stopped, not started. A running app on a half-written database lets the customer type into it and
makes the damage permanent. The hold is persisted, every start path refuses it with a reason and a
route, the app-stop marker is ended so nothing auto-restarts it at the next boot, and the app reads
red rather than green. An operator clears it with
--clear-restore-hold, which requires a controller restart.
Why a rollback and not an engine flag. Postgres gained --single-transaction in the same release
and that does make its replay all-or-nothing — but MariaDB's DDL is not transactional, so a
partial apply there is unavoidable at the engine. Measured 2026-08-22: the same truncated dump left
Postgres emptied and crash-looping, and left MariaDB with its user data intact, its schema-version
table wiped to zero rows, and the app reporting health=healthy, running=true, restarts=0. The flag
is a belt; the rollback is the fix. R-379, R-380. Evidence:
audits/DRILL-r379-rollback-2026-08-22/.
[DESIGN] 2026-08-22 — the restore destination is resolved by the same rule as the capture
destination. The drive if the app declares one (HDD_PATH), the system data path otherwise —
Manager.GetAppDrivePath, one expression, used by CaptureRecoveryUnit and, since controller
v0.219.0, by ReconstituteFromOffsite and PlaceOffsiteRestore too.
[FACT] 2026-09-01 — the rule now has a FOURTH consumer and the "one expression" sentence is
true (R-414, controller v0.232.0). offboxRestoreScratchDir — where a restore is UNPACKED, as
distinct from where it LANDS — never consulted systemDataPath, so on a box with no registered
storage path it refused, and the nightly proof could not run at all. It was not excluded on
purpose: R-356's own commit (08eb1a6) records in its tests that "the prepared scratch still
resolves to the registered storage path … only the DESTINATION moves" — it was out of scope, and
every fixture assumed a registered path exists. The one comment about a systemDataPath fallback
belonged to PlaceOffsiteRestore, concerned bulk userdata, and R-356 deliberately overruled it.
But the fallback is SCOPED, and the scoping is the point. The two callers ask different
questions, and one predicate answering both is R-356's own defect: a unit-only restore (the
R-87 proof) may fall back to the system data path, because §7 records as [FACT] that a driveless
app's unit already lives there indefinitely and that the same-device placement is "intended, not
a defect"; a full restore keeps the R-252 refusal, because it pulls the app's bulk userdata
onto what §2.2 makes a state-only tier.
The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351) applies
to apps that have a drive to get wrong. It used to be reached by testing HDD_PATH == "", which
also answered "is this app installed?" — one predicate for two questions. Measured in the catalogue at
459766cb1639: 53 templates, 13 declare needs_hdd: true, 40 declare false, so for 40 apps that
test was permanently true and the off-site restore refused them forever, while they were running,
with a message telling the customer to reinstall them "in the same place" — a place those apps never
offer. An app with no drive is not misconfigured (§8 of 01-topology-and-trust.md carries the
[DESIGN] marker); it is the majority case.
Since v0.219.0 the two questions are asked separately: installed? of ListDeployedStacks(), failing
CLOSED when there is no provider to ask; where? of GetAppDrivePath. A third refusal, with its own
sentence, covers installed-but-no-resolvable-data-root. R-356; reasoning also recorded in
felhom-controller/CONTEXT.md.
6.4 The customer's page speaks per tier, and a tier without storage is skipped (2026-09-15)
What the page may claim (controller v0.243.0 + agent v0.131.0, R-517). The whole-system tile used
to show the agent's single latest record — so a failed 0-byte PBS attempt read „Naprakész" and ticked
„Távoli rendszermentés — külön hardveren" (BIGNIGHT 2026-09-14). Now GET /backup/status returns per
tier the newest success, the last attempt kept apart, and whether the tier's storage exists.
The page shows each tier's newest success; a failed attempt under it as „sikertelen"; a tier whose
storage does not exist as „nincs beállítva". „Naprakész" and the remote tick are computed from
successes only. After an agent restart the success is read back from the tier's storage (size unknown).
What a run may do (R-518, cheap half). A tier the agent reports storage: absent is dropped before
anything is stopped, logged, and reported once as backup_tier_skipped; unknown is never skipped.
Still open: quiescing per tier, so a slow second tier does not keep every app down.
Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0): „Mentés most" stopped the apps at 09:19:08Z, the
local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the
longest stop 5 min 47 s, for the local tier alone. The button text and its confirm (v0.296.0) give both
measurements (≈6 min / 9 apps, ≈8 min / 12 apps), say "minutes, not seconds", and that the off-site copy in the same
run makes it longer.
6.5 Kept data — what a removed app leaves on the drive (controller v0.274.0, 09 §3 decision 36)
What it is. "Remove the app, keep my data" leaves the app's drive folder (<drive>/appdata/<app>, or the folder
its definition binds through ${HDD_PATH}) in place. A reinstall over it no longer runs silently into the old files
(R-657): the install asks „A megőrzött adataimat használom" / "Use my kept data" (the database from the newest copy of
THIS drive's install — the app's own unit, the second-drive mirror or (since v0.277.0) the off-site copy — loaded under the kept files, then the
template's after_load:, e.g. nextcloud's occ files:scan --all) or „Tiszta lappal kezdem" / "Start fresh".
The household's userdata stays on every remove (09 §3 decision 67, R-800). "Remove with data" deletes the app's
own data — its volumes, <drive>/appdata/<app> (the ${HDD_PATH} binds) and its backups. A folder the app binds through
${USERDATA_PATH} (userdata/media/grimmory, userdata/media/metube, …) is the household's: it is never deleted by a
remove, and the dialog and the result say so and name the folder (controller ≥ the release that carries R-800).
Where it lives. "Start fresh" RENAMES the folder — same drive, never a copy, never across drives — to
<drive>/kept/<app>/<YYYY-MM-DD_HHMMSS>/, together with the removed app's unit when one sits on that drive (so a later
Load has its database). <drive>/kept is in ProtectedHDDPaths, never under userdata/, never bound by a live app.
The page „Megőrzött adatok" / "Kept data" lists every dated folder and every appdata/<x> no installed app binds.
The off-site copy counts too (controller v0.277.0, R-691 (2)). „Use my kept data" and Load consider the app's newest OFF-SITE snapshot beside the local copies, and use it when it is NEWER than every local copy or the only one (a tie goes local). The page names the copy and its date („távoli mentés, 2026-09-28 04:15" / "the off-site copy, …"). The repository is asked once per page, bounded to 20 s; an unreachable one offers nothing. The load downloads the unit ALONE (never the files — they are the kept ones), and refuses it before anything moves when it was taken of another drive, holds no database, or does not record its data version (§6.6 — a unit written before v0.275.0 is never loaded blind); a mixed or mismatched unit is refused by §6.6's own rule. Then a dated folder's files move back and the one unit restore runs. The downloaded copy is removed on every path. A failed download or a refusal leaves the kept files as they were.
It is NOT backed up. No tier captures <drive>/kept or a leftover appdata/<x> of a removed app; the page says
so („Erről nem készül mentés." / "This is not backed up."). What brings it back into an app is the unit it carries
(or the app's own unit on the drive), listed per row as „Visszatölthető innen" / "Can be loaded from".
Who deletes it. Only the household, by Delete on that page with the app's name typed (a wrong name, or a path that
is not a listed item, is refused — proven live 2026-09-25). The box never deletes kept data by itself — operator ruling
2026-09-26 (09 §3 decision 40, D3 option A): no age limit, no automatic clean-up. So kept data can fill a drive; the drive-full warning names the kept folders and their sizes as space the household can free.
Read-only view. The file browser shows each kept item under „Megőrzött adatok", one :ro bind per item, and
follows the list at the next sync (a write is refused: Read-only file system, proven live). Not yet readable there:
a folder its app owns with mode 0770 (nextcloud, www-data) — R-691 — until controller v0.275.0.
[DESIGN] How the view reads a folder another user owns — decided by CC unattended 2026-09-26 — operator may
reverse. One sentence: how does the read-only view open a 0770 kept folder owned by another user without touching
the household's files? Options: (a) a read-only ACL on the folder; (b) run the file browser as root; (c) a second
viewer container running as the owner; (d) the file browser joins the folder's OWNING GROUP (group_add).
Costs: (a) changes the household's files' metadata — nextcloud checks its data folder's mode after a Load, and
the brief forbids it; (b) the whole file browser (read-write on the household's userdata) as root; (c) a new container,
route and login for one view; (d) the group also applies to the view's other mounts — so root's group (0) and the
view's own (1000) are never added, and only a group-READABLE folder's group is. Why (d): the kept binds stay
:ro (a write is refused by the mount), nothing on disk changes, and it is one compose line the sync already writes.
Limit: a file inside that is owner-only (0600) stays unreadable. Also: a language switch now re-syncs the file browser
so the source's name („Megőrzött adatok" / "Kept data") follows the box's language. Controller v0.275.0,
audits/version-travel-2026-09-26/D3/.
Evidence: audits/night-2026-09-26/E/ (E1 spike, E5 live proof).
6.6 Which version a restore brings back (controller v0.275.0, R-696, D4 option A)
[DESIGN] The rule: a restore brings an app back at the version its DATA belongs to — never the data of
one version under the definition of another. When the only copy holds data written by an older version,
the app comes back at that older version with its own data, and the normal guarded update then climbs the
ladder again, one tested step at a time (09 §3 decision 14). The household's first sentence says so:
„A(z) %s visszaállt a(z) %s-i mentésből, a(z) %s verzióra. Elérhető frissítés: a doboz lépésenként hozza
naprakészre." / "%s is back from the backup of %s, at version %s. An update is available: the box brings it
up to date one step at a time." The version is every image as name:tag — a PostgreSQL step leaves the
app's own tag unchanged. D4 option B (restore into a temporary copy, update it there, move the data in) is
NOT built; it would be an add-on to this rule, not a replacement. D4 is recorded in STATUS.md; until it is
ruled, A is in force.
[FACT] Why, measured before the fix (9202, 2026-09-26, audits/version-travel-2026-09-26/A1/). The unit's
definition (compose/, image_pins) was re-captured by the five-minute status refresh as soon as an update
moved the pin, while its data (db-dumps/, volume-dumps/) is written only by the backup legs — so for up to
a day the unit paired the new definition with the old data. An app step (docmost 0.95.0 → 0.96.0) restored in
that window came back only because docmost migrated the old data forward at its first start — an update step
nobody guarded. An engine step (PostgreSQL 16 → 18) restored in that window poured the 16 datadir back,
postgres:18 refused it, and the app was left down. The off-site restore never wrote a definition at all,
so after any update it mixed versions until the next night's off-site run.
[DESIGN] How the data carries its version. Every data file a backup leg writes (the database dump, each
volume tar) gets a stamp at the moment it is written — its size and mtime, the definition's image pins, and
what each service was running as ref@digest (data-stamps.json in the unit). The unit capture folds the
stamps into the manifest's data block (at, image_pins, images, per-file stamps) and keeps the
definition its data belongs to: when the pins moved since the data was written, compose/ is not rewritten
until the next data run replaces the data. The manifest's image_pins stays the app's current pins, so it says
both. Files written under different pins (a failed leg) make the unit mixed.
[DESIGN] What each restore does with it.
| path | the unit's data block |
what happens |
|---|---|---|
| own unit (Tier 1), second-drive mirror (Tier 2), kept-data Load | present, compose/ names the data's pins |
the data starts under that definition; if it differs from what ran, the restore pages' first sentence names the backup and the version (the kept-data Load has its own sentence and does not add this one) |
| same | present, compose/ names OTHER pins, or mixed |
refused before anything is touched (ErrUnitVersionMismatch): „…ezért a visszaállítás nem indult el. Az alkalmazás érintetlen" |
| off-site (Tier 3) | present, the snapshot's version differs from what runs | the snapshot unit's definition is written into the stack dir (and pinned) right after the stop, before any file, volume or database is touched; the database service is resolved from it |
| any | absent — a unit or snapshot written before v0.275.0, or with an unstamped file | restores as before (its captured definition; the off-site path the live one), with a WARN naming it |
[DESIGN] What a restore keeps from the app it replaces (controller v0.276.0, R-697). The restore writes a fresh app.yaml from the unit — env, locked fields, and the pin from the unit's definition — but keeps the records of the app's life on this box: the household's desired_state, the kept pre-conversion copies (conversion_copy, earlier_conversion_copies — a restore does not remove the volume, so it must not forget it), and the ladder history (failed_update_step, last_update_undone, last_auto_update). Not kept: installed_images (what ran before, maybe another version). A drive move is not a restore: it changes HDD_PATH and nothing else (R-700 — before v0.276.0 it dropped the pin, and the syncer then gave the app the catalog's newest version at its next start). Pinned by internal/stacks/r700_records_carried_test.go.
[DESIGN] The jump guard. An installed version that matches no ladder entry's from jumps to the catalog's
current definition when a PERSON presses Update (ladder.go, measured live on vikunja 2.5.0, 09 §6.4 part 5);
the AUTOMATIC leg never presses it (LegSkipOlderThanLadder, and a template without a ladder is
LegSkipNoTestRecord). So a restore to a version older than the ladder — or of an app with no ladder — gets
the other first sentence: „…Ettől a verziótól nem vezet kipróbált frissítési lépés, ezért a doboz magától nem
frissíti." / "…No tested update step leads on from this version, so the box does not update it by itself." A
jump is not a tested step, and the page does not promise one.
[DESIGN] The times tell the truth (A4). Every tier's "proven at" is the time of its DATA, never of a
manifest: Tier 1 = data.at (else the newest data file, the pre-restore-* undo copies excluded, else — for a
unit with no data file — the manifest); Tier 2 = the mirror's data time, capped by its copy time; Tier 3 = the
snapshot time, capped by the data time the box recorded when it pushed it (settings.offsite_data_at). The
update's precondition and the release of a kept pre-conversion copy read these times; the release additionally
requires a database dump written after the conversion whose recorded engine is the NEW major.
[FACT] The limit — the image is not in the backup (R-698). A unit stores image NAMES and digests; the
image is re-pulled on restore. A version its maker has deleted cannot start. Measured 2026-09-26: the 42 ladder
digests all resolve today (audits/version-travel-2026-09-26/A7/). Options are in R-698; nothing is decided.
Since decision 53 (09 §3, 2026-09-30 evening): a box keeps only the image each app service runs and the one before
it; older images of the app are deleted. A restore that needs an older version re-pulls it — as before; the limit
above is unchanged, and kept data (decision 40) is not touched by the image clean-up.
7. The recovery chain (D3) — the reason this document exists
[DESIGN] 3-2-1 describes copies. It does not describe recovery.
Three copies on two media with one offsite is a statement about bytes surviving. It says nothing about whether the bytes can be turned back into a working system, and a system can satisfy 3-2-1 completely while having no executable recovery route for a given failure. That is not a hypothetical here — §8 has rows where it is the actual state.
7.0 What a customer can and cannot do ALONE — the four steps the drill found
[FACT] Added 2026-08-05. The 2026-08-04 R-201 drill (
audits/DRILL-r201-night-run-2026-08-04.md) is the first end-to-end proof that a customer's file survives a machine rebuild and comes back byte-identical. It passed with an operator present, and four manual interventions stood between "the key is recoverable" and "the file is back" — none of which was in any design document. They are recorded here because this is the section a future reader will use to answer "can the customer do this alone?", and until 2026-08-05 the answer was no for reasons nothing wrote down.
| # | The step | Why it stopped a customer | Status |
|---|---|---|---|
| 1 | The reset code was refused on the first try | --print-reset-code runs as a SEPARATE process (docker exec); it persists a new code while the running controller keeps the old one in its cache, so the code the customer is told to type never matches. The controller had to be restarted in between, which nothing said. Two attempts failed during the drill before that was worked out. |
CLOSED — controller v0.198.0 (R-204 item 1). effectiveClaimCode reads through to the persisted state. Read-through, not a TTL: a TTL would leave a window in which a superseded code still works. Fails closed on an unreadable state. Proven live on demo-felhom 9201 with nothing restarted. |
| 2 | Re-issuing the off-site credential marked the recovery escrow "stale" | A stale escrow withholds restic_pw_sha256 from the report ACK → the controller's auto-confirm cannot flip pending→escrowed → OffboxRunnable() is false → every off-site backup refused. The customer is then told to re-run the recovery ceremony, which is the one act that would have destroyed the key just recovered. The repository password had not changed at all. |
CLOSED — hub v0.95.0 (R-204 item 2 / R-196). The precautionary mark is gone. The real case is measured by the controller's Scenario-F hash re-check each ACK — which the mark was BLINDING by emptying that very hash — and by R-197's offsite_repo_key_changed at a supersession. |
| 3 | The restore's default returned the wrong thing, silently | mode=unit restores the recovery unit — the app's definition, configuration and database dumps — and not the customer's files; the userdata in the same snapshot is excluded by --include. The outcome message was one sentence for both modes and named neither scope. On the last step of a disaster recovery, the default quietly did not do what the person asked. |
CLOSED — controller v0.198.0 (R-204 item 3). The unit outcome names what came back, what did not, and the step that gets it; the wizard's intent card states its scope before the choice. The size gate on mode=full is untouched, and the default stays unit. |
| 4 | A rebuilt box cannot obtain an off-site credential unaided | The one-time provider password was spent by its predecessor, so the rebuilt guest has nothing to authenticate with and an operator Re-issue was required. | CLOSED — controller v0.199.0 + hub v0.96.0 (2026-08-05, operator ruling: automate it, and the trigger is a state the BOX DECLARES). The box now reports offsite.state=needs_credential when two local facts hold together — a fresh data area AND a hub-held recovery package — and the hub's internal/offsiteheal re-arms the stored one-time secret, minting only when there is nothing to re-arm. Deliberately NOT automated: the escrow ceremony. A credential is replaceable; the recovery code is not. |
[FACT] The honest current answer, updated 2026-08-05 (controller v0.200.0): all four are gone AND the customer is now offered the recovery. A full-page screen meets the owner of a rebuilt box while the hub holds a sealed package the box cannot open: it explains the situation, says plainly that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — apps, dates, sizes. No step requires an operator, and no step requires a command line.
WHERE THAT STOPS, and it is a real stop. The screen unlocks and only unlocks. It restores nothing. Putting files back is per-app, lives in the backups area, and the step after the listing — a customer seeing what would change before anything is overwritten — is not built (→ R-213). A screen that unlocks and then offers to overwrite is two decisions wearing one button, which is why the operator ruled them separate.
What is still owed as evidence: the final unlock has never been driven with a CORRECT code through the page (no code was kept for demo-felhom's orphaned history; demo-hp's is operator-held), so the live run exercised handler → agent → hub fetch → age KDF and stopped at the unseal. And the whole journey has not been re-walked end to end since these fixes — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill.
WHY THE TRIGGER FOR STEP 4 IS A DECLARATION, recorded here because it is the design and not an implementation detail: from the hub, an ABSENT off-site object means never configured, mid-restart, a transient config read failure OR rebuilt and stranded, and the hub cannot distinguish them. The BOX can, from two local facts it holds with certainty. So the box states its condition and the hub acts on a stated request — never on a silence. Both facts are required: freshness alone is a box that never had off-site backups, and an escrow alone is a healthy box.
The split that is deliberate and must not be widened: CREDENTIAL AUTOMATIC, KEY CUSTOMER-PRESENT. A credential is transport and is replaceable; the recovery code is not, because only the customer holds it. Nothing in this chain runs, or asks for, an escrow ceremony.
7.1 The dependency graph
[FACT] SUPERSEDED IN PART BY D5 — SHIPPED 2026-07-30 (controller v0.188.0). Leg 1 below (secrets) is no longer true for Tier-1 and Tier-2: the local recovery unit now carries the portable secret class, so those two tiers need the drive and nothing else. The graph and leg 1 are kept as written because they are the model everything downstream was derived from, and because leg 2 (the living app) is unchanged and still binds Tier-2's file leg and Tier-3. Read §7.4 for the current chain; the corrected rows are 3, 4 and 9 in §8.
[FACT — as of the R-108 survey, before D5] Every app-tier restore depends on the guest, for one of two reasons:
┌──────────────────────────────────────────┐
│ the LIVE GUEST │
│ · settings.json (tier-2 destination) │
│ · encryption.key (32 B) │
│ · app.yaml (ENC: under that key) │
│ · the deployed app itself │
└───────────┬──────────────────────────────┘
│ required by
┌──────────────────────────┼──────────────────────────┬─────────────────────┐
│ │ │ │
Tier-1 restore Tier-2 restore Tier-3 reconstitute Tier-3 scratch
(rebuild the app) (fill in files) (overwrite + replay) (verification copy)
│ │ │ │
needs SECRETS needs the app needs the app needs only the
from app.yaml running + the deployed + a DB repo password
(restore_unit.go recorded dest service identifiable (also in the guest)
:130-132) (tier2_restore.go
:114-116)
│ │ │ │
└──────────────────────────┴──────────────────────────┴─────────────────────┘
│
┌───────────▼──────────────────────────────┐
│ WHOLE-GUEST TIER (local vzdump / PBS) │ ← Lane 2, operator only
└───────────┬──────────────────────────────┘
│ required by
┌───────────▼──────────────────────────────┐
│ THE HOST (agent.json, bootstrap.json, │ ← in NO backup at all
│ vmbr9, sudoers, /etc/pve/priv, WG) │ (INV Part D1)
└──────────────────────────────────────────┘
[FACT] The two legs of the dependency, precisely:
- Secrets. The recovery unit is secret-free by design — "It NEVER writes a secret value"
(
internal/backup/recovery_unit.go:73).RestoreFromRecoveryUnitrecovers secrets from the guest, never from the unit (restore_unit.go:130-132), and the policy is stated outright at:18-22: "Regenerate NOTHING… A missing DATA-ENCRYPTING key is FATAL: regenerating it would render the restored data unreadable, so we refuse and tell the operator to do a PBS whole-guest restore." Those secrets are encrypted underencryption.key, a 32-byte file that exists only inside the guest (LIVE, both boxes — INV Part C, row 9). - The living app. Tier-2's restore reads the destination recorded in the guest's
settings.json(tier2_restore.go:114-116), and Tier-3's reconstitution refuses outright when the app is not deployed — „a(z) %s nincs telepítve — előbb állítsd helyre az alkalmazást, utána az adatokat" (offbox_reconstitute.go:198-201).
[FACT] So all three app tiers are conditioned on the whole-guest tier, which is Lane 2. The customer's own recovery lane rests on an operator-only tier, and today nothing tells the customer that. — This sentence was the reason D5 exists. It is now false for Tier-1 and Tier-2; see §7.4.
7.4 The chain AFTER D5 (controller v0.188.0, 2026-07-30) — the current model
[FACT] Tier-1 and Tier-2 no longer depend on the whole-guest tier. The unit's compose/app.yaml
(mode 0600) carries the portable secret class — every catalog type: secret field, i.e. the
declared data_keys, the 18 DB/root passwords and the internal signing secrets — so neither
encryption.key nor the guest's app.yaml is required to rebuild an app:
Tier-1 restore Tier-2 restore Tier-3 reconstitute Tier-3 scratch
(rebuild the app) (fill in files) (overwrite + replay) (verification copy)
│ │ │ │
needs ONLY THE DRIVE needs the app needs the app needs only the
(unit carries the running + the deployed + a DB repo password
secrets; guest is recorded dest service identifiable
consulted only for
the withheld class) └────────── still conditioned on the guest (leg 2, unchanged) ──────┘
│
✅ INDEPENDENT of the whole-guest tier
[FACT] What still needs the guest, precisely — so this is not read as more than it is:
- Leg 2 is untouched. Tier-2's file restore still reads the destination from the guest's
settings.json(tier2_restore.go:114-116) and Tier-3's reconstitution still refuses when the app is not deployed (offbox_reconstitute.go:198-201). D5 removed the secrets leg, not the living-app leg. So "Tier-2 no longer depends on the guest" is true of the unit-restore path it shares with Tier-1, and false of its additive file-merge path. - The withheld class. The 7
type: passwordadmin logins andvaultwarden/ADMIN_TOKENare deliberately NOT on the drive; they come from the guest, or are regenerated (O4). Their absence costs a credential reset, never data. - The fail-closed gate remains. A
data_keyin neither the unit nor the guest still refuses the restore outright. D5 makes it normally present; it does not soften the gate.
[DESIGN] Why plaintext is the right answer here, and what it is coupled to. Every travelling secret
either decrypts data on the same drive or authenticates to a container on an internal compose network
with no external listener — so possessing it adds nothing to possessing the drive, which is exactly D2's
argument for keeping the DATA plaintext. That argument holds only because the internet-reachable
class is withheld. The exclusion is what licenses the plaintext; the two must not be relaxed
independently. stacks.PortableSecretEnvVars is the single boundary, with a code-level
nonPortableSecrets register (not a catalog flag — a boundary a catalog push can move is not a
boundary, cf. R-97a).
[FACT] Precedence, because two sources now exist. The unit wins over the guest. Not "newest wins": the unit's secrets are captured in the same run as the dumps beside them, so the unit's value matches the data being restored, whereas the guest's is merely the most recent. Preferring the guest is the data-loss bug — a rotated data key does not decrypt data encrypted with the old one, and a rotated DB password does not match the hash inside the restored data directory.
[FACT] Proven live 2026-07-30, on a scratch drill guest, through the real endpoints: AdventureLog
(SECRET_KEY data_key + DB_PASSWORD) restored with the guest's app.yaml moved aside —
secrets recovered=2/2, 27.6 s — and then the application itself read the seeded customer row over
TCP with its own credential, which is the observable that matters (a restore that returns success onto
unreadable data is a failure that looks like a pass). The discriminator held: the pre-backup row came
back, a row added after the backup was gone, so the volume tar was genuinely restored. There was no
.sql dump in that unit, so the DB came back from the volume tar — the exact case where a regenerated
password would have failed. Separately, Grafana's type: password admin login was proven withheld:
the sentinel value was live in the container and encrypted in the guest, and appeared in 0 files
anywhere under the backup namespace. Full record: audits/D5-drive-alone-restore-2026-07-30.md.
7.2 A tier whose prerequisites cannot be met in the failure it exists for
[FACT] Two instances, both current:
Tier-2 vs primary-drive loss— CLOSED 2026-08-31, controller v0.229.0 (R-102). Tier-2's stated purpose is surviving the loss of the primary drive. It was true until v0.229.0 that in that failure the primary recovery unit is gone while the surviving mirror on the second drive —backups/secondary/<app>/recovery-unit/— was read by no code path (§6.3), so for the 45 class-B apps (§6.2, count settled the same day) the restore was a guaranteed no-op in exactly its designed scenario. Tier-2 can now meet its prerequisite in the failure it exists for, and that is stated plainly because it is the whole point: the restore was run ondemo-hpwith the primary unit moved aside and returned 3 volumes of 3 and 1 database of 1 in 28.65 s, byte-for-byte, with the app then reading its own row over TCP with its own credential — and again with the guest'sapp.yamlalso moved aside,secrets recovered=2/2from the mirrored unit. Evidence:audits/DRILL-r102-tier2-unit-2026-08-31/. What the drill did NOT cover, stated per §8 below: the full drive-loss journey — a genuinely absent or replaced physical drive — was not run. Only the recovery-unit half was.- Tier-3 vs guest loss — HALF of this closed. Tier-3 holds the volume tars and the DB dump. It
was true until controller v0.218.0 that the tars were unpacked only by the Tier-1 path; since
v0.218.0 the reconstitution replays them itself (
volReplay) → R-107 CLOSED 2026-08-22. What is still true, and is the part that was never R-107: reconstitution requires the app to be deployed and skips the unit, so offsite alone cannot rebuild an app onto a fresh guest. That requirement is a deliberate decision (R-253) — the restore does not choose a customer's drive for them — and since v0.219.0 it means only what it says: an app that is not deployed. It no longer catches the 40 driveless apps → R-356.
7.3 What D5 would change — and why it was blocked — D5 SHIPPED 2026-07-30 (v0.188.0)
This section is history. D5 is implemented and proven live; the current chain is §7.4. Kept because the reasoning below is why the precondition was required, and because the last paragraph (
.fab/ R-126) is still open and still not part of D5.One correction to the target as stated below. It assumed the class that must travel is the
data_key-flagged secrets. Part 0 of the D5 task tested that and it did not survive: the flag is unreliable (four encryption keys the catalog itself labels as such are unflagged → R-127), and a DB password is not resettable in practice —POSTGRES_PASSWORDis ignored once PGDATA is non-empty, so a regenerated value leaves the app unable to authenticate against its own restored rows while the dump replay still reports success. The shipped boundary is therefore alltype: secretminus the register, notdata_keyalone.
[DESIGN, TARGET — SUPERSEDED BY §7.4] The intended fix is to make app secrets travel with the local recovery unit, so Tier-1 and Tier-2 restore work without the guest and without R. Offsite already encrypts everything, so secrets travelling offsite would be covered by the escrowed repo password. R would then be required for offsite recovery and host identity only — losing R would cost the offsite route, not local recovery.
The premise D5 rests on is now established. §2 of the task that produced this document required
it to be proven, not assumed: the backup tree must be unreachable from every browsing, download and
export surface. When this document was written it was not — the FileBrowser network-share bind
reached it. R-108 closed that on 2026-07-30 (controller v0.187.0) by refusing app namespaces on
network storage, so no backups/ tree can exist under the share-root bind; every other surface was
already clear (§10.1's table). D5's precondition is therefore MET and D5 may be adopted.
Still true, and not part of D5's precondition: a .fab bundle carries plaintext secrets by
design with an optional password, and storageDriveList() does not filter network paths, so a bundle
can be exported onto a NAS (§5, → R-126). That is an export destination the customer chooses
explicitly, not a browsing surface reaching a backup tree, and it is unchanged by D5 — D5 moves
secrets into the local recovery unit, not into .fab. It is tracked separately rather than folded in.
Until D5 is actually implemented, §7.1's chain stands as the model.
7.5 Where a recovery unit is RETAINED — and the size bound it puts on Lane 1
Added 2026-08-02 from audits/SPIKE-recovery-unit-space-2026-08-02.md and
audits/CAMPAIGN-10-closeout-2026-08-02.md. §6.1 says a Tier-1 unit lives "on the app's own drive",
which is true and is the whole story for a drive-resident app. It is not the whole story for an app
with no data drive, and that case was undocumented until now.
The unit is RETAINED, not staged. GetAppDrivePath (felhom-controller/internal/backup/backup.go:245-255)
returns the app's HDD_PATH if it has one and systemDataPath otherwise — what
internal/appbackup/paths.go:26-27 calls "the SSD-only system-data fallback". Nothing deletes a unit
after it is copied onward: the only prune is F5 (backup.go:1053-1112), which removes residue on an
old drive when an app moves. So for every app without a data drive, mp1 holds the kept copy
indefinitely.
What a unit contains, which bounds the problem (internal/backup/recovery_unit.go:20-25):
compose/ + db-dumps/ + volume-dumps/ (named-volume tars) + manifest.json. It never contains
mp8 / HDD_PATH userdata — a photo library on a 1 TB drive is not in a unit, and cannot make one
overflow.
The resulting mismatch, on a default appliance:
| size | holds | |
|---|---|---|
mp0 /var/lib/docker |
50 G | every app's live volumes — the thing a unit copies |
mp1 /mnt/sys_drive |
20 G | the retained units of every driveless app |
A DB-backed app's unit is up to ~2× its data, because it carries the volume tar and the SQL
dump (measured: 21.1 GB of app data → a 40.2 GB unit). A file-only app's is 1.00× (measured:
homebox 2305 MB → 2305 MB). --sysdata-grow (felhom-agent/cmd/felhom-agent/main.go:178) defaults to
0, is not derived from the physical drive, and no observed box passes it — demo-hp's production
guest 9201 ships mp0 50G / mp1 20G.
mp1 gates the entire app-data chain, not just Tier 1. Tier-2 mirrors the unit "(always)" from
RecoveryUnitPath (internal/backup/tier2.go:302,368) and Tier-3 carries it too, so a unit that
cannot be written to mp1 leaves Tier-2 and Tier-3 with nothing to copy.
The bound this places on Lane 1 (§3). Lane 1's promise is that a customer can restore an app from the drive alone, without the operator and without the guest — D5 (§7.4) made that true by putting the portable secrets in the unit. That independence is bounded by app size, and the bound is:
On a default appliance, an app can be restored by Lane 1 only while its recovery unit still fits the retained space — ≈ 19 GB of app data for a file-only app, ≈ 10 GB for a DB-backed one. Past that the unit cannot be captured, and the app falls back to Lane 2's operator-driven whole-guest route.
AS OF 2026-08-02 SOMETHING NOW WARNS, AND THE ALERTING IS PART OF THIS CONTRACT (R-167 / R-158, decision D-c; controller v0.191.x + hub v0.89.0). The last sentence of this section used to end "nothing warns when an app crosses the line". Two signals now exist and both are PROVEN-LIVE:
- To the CUSTOMER, before anything fails —
internal/fillwatchwarns per FILESYSTEM (never per app: one full disk holding ten apps would fire ten times) on whichever trips first, used ≥ 85% or free < 5 GiB, critical at 95% / 2 GiB, clearing at 75% / 7 GiB. Two terms, because a percentage alone lies at both ends of the range this section itself documents: 85% of a 20 Gmp1leaves 3 G — less than one DB-backed app's unit — while 85% of a 4 TB drive leaves 600 G. It watches the app-data volume, the system-data volume and every registered drive, which the previoushealth_degradedsignal did not. Edge-triggered against persisted state; the hub owns cooldown. - To the OPERATOR, when a capture actually fails —
recovery_unit_capture_failed, per app, with the target filesystem's used/free bytes at the moment of failure, so the why needs no login. It is operator-tier (notify.operatorOnlyEvents) and deliberately notbackup_failed: a customer can take no action on a capture failure.
7.5.1 — THE CEILING THIS SECTION DESCRIBES HAS BEEN REMOVED (2026-08-03, R-165 / decision D-a)
Everything above describes the SPLIT layout, which is now the legacy shape. A golden built by
build-golden.sh v3.0.0 ships one data volume; mp1 does not exist. Both consumer paths are
binds of subdirectories of it (variant V-c):
mp0 -> /var/lib/felhom ├─ docker/ --bind--> /var/lib/docker
└─ sys_drive/ --bind--> /mnt/sys_drive
So the size bound below no longer applies to a box built from that golden. A driveless app's
recovery unit is limited by the box's actual free space, not by a partition set at build time. The
mismatch table above (mp0 50 G vs mp1 20 G) describes what a merged box no longer has.
R-175, fixed here rather than left standing. The bound below was stated as the fleet's and was
one box's: it is derived from mp1 = 20 G, which is demo-hp exactly and never was demo-felhom
(mp0 200G / mp1 50G, where the same arithmetic gives ≈ 49 GB / ≈ 24 GB), nor the golden (16 G / 8 G
before provision grew them). Read it as a function of mp1, and only for a box still on the split
layout. Measured: audits/SPIKE-r165-mp1-merge-2026-08-02.md M1.
What replaced the partition's second job — the reserve. mp1 was also a BULKHEAD: an overflow was
refused per app with the last good unit byte-identical, and it could not reach /var/lib/docker,
because that was a different filesystem. On a merged box it can. Decision B2, shipped in controller
v0.192.0, is that bulkhead made deliberate — a two-term reserve (97% used or 1 GiB free) in
internal/fillwatch's shape, sitting beyond its critical band so the customer is always warned first.
It refuses per app and never deletes: nothing on this filesystem is generational, so pruning could
only destroy a different app's only local copy.
WHAT IS RECORDED, WHAT IS E-MAILED, AND HOW OFTEN (controller v0.194.0 + hub v0.90.x, R-182). The two are deliberately different mechanisms, because conflating them is how seven failures went missing on 2026-08-03 without leaving a trace.
| Record | Notification | |
|---|---|---|
| what | recovery_unit_capture_failed, one per failed app |
backup_run_failures, one per RUN |
| when | every time, unconditionally | at the end of a run, only if something failed |
| gated by | nothing — not cooldowns, preferences or delivery | the hub's operator cooldown |
| where it lands | the events table and notification_log (status recorded) |
the operator's inbox |
- A clean run e-mails nothing. Silence means the run finished and found nothing wrong — and that
is only safe because the hub's daily deadline check raises
expected_backup_missedfrom the box's REPORT freshness, independent of any mail the box sends. That check is load-bearing for this design; weakening it re-opens a silent-failure path. - A suppressed operator notification leaves a
suppressedrow naming the key that suppressed it. Deciding not to tell someone is itself an event worth recording. - Deliberate skips are not failures and never appear in the digest — a disconnected or decommissioned drive has its own alert, and a nightly digest about an unplugged drive is one the operator stops reading.
- Cadence: a nightly run gives at most one mail a day. A manual run always reports, even within the hour, because someone pressing the button is actively trying to get a backup. The periodic capture sweep is capped by the ordinary hourly cooldown.
THE CONTRACT, stated as what the code provides (controller v0.193.0, R-181). The reserve is a
per-app, per-run ADMISSION decision, not a capture check. It is taken once for an app, immediately
before that app's FIRST write of the run, and it covers all three write legs — the database dump, the
volume dump and the recovery-unit capture. Those three write under one per-app root
(backups/primary/<app>), which is what makes one verdict able to cover them honestly.
- What it guarantees. A refused app has nothing written for it in that run, its previous unit is byte-identical, it is not stopped, nothing anywhere is deleted, and the operator gets exactly one alert naming the app, the term that bound and the disk figures.
- Two terms, two questions. Headroom: is the filesystem already below the reserve? Size: would
THIS app's write take it below? The size estimate is the app's previous
.sql+.taron disk; with no history the decision degrades to headroom alone, deliberately — otherwise the first backup is the one that can never happen. - Why it is decided lazily and not once per run. Space changes during a run: app A's dump can put app B under the reserve, so a verdict taken at run start reads a disk that no longer exists.
- Why it is never re-decided between an app's own legs. That is precisely the shape v0.192.0 had — the two dump legs unguarded and only the capture refused — under which the reserve was consumed by the very write it exists to bound, and the refusal's "the previous unit is untouched" was measured false. Proven live on demo-hp 2026-08-03 (R-181), fixed the same day, and re-proven by filling the box for each of the two terms.
- It sits ahead of
DumpAppVolumesSafe, which stops the stack as its first act — a refusal decided inside it would already have bounced the app it is refusing to back up.
Status caveat, deliberately explicit: every box in the field that has not been reinstalled is still on the split layout and everything above still describes them exactly. This subsection describes what a box built from golden ≥ 0.192.0 gets. Both demo boxes were reinstalled from it on 2026-08-03 (R-178).
Two things are deliberately not recorded here. The sizing ratio is the operator's ruling
(R-163) — this section states the constraint, not a number. And the same-device placement is
intended, not a defect: a driveless app's unit sits on the same SSD as its volumes, but Tier-2's
cross-drive mirror, the whole-guest tier (which covers rootfs + mp0 + mp1, §6.1) and Tier-3 all cover
device loss. What was missing was that this case existed at all, and that nothing warns when an app
crosses the line — R-158.
8. The failure → recovery matrix
THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01)
Closed at controller v0.232.0 / hub v0.111.1. Everything a customer does for themselves is finished and proven live — rows 1, 2, 3, 3b, 3c, 6, 7, 14. Everything only an operator does is deliberately deferred until after beta — rows 4, 8, 9, 10, 11 (+11b), 12, each carrying the marker
[BETA-DEFERRED]in its status cell so the set greps as a group.The blanks in the RTO column below are now blank ON PURPOSE, and that is the whole difference. Eleven rows carry a blank RTO — 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15 — and only six of those are deferred work. Row 5 has no recovery to time (a derived copy), row 13 has no route by design, row 14 is proven and merely never stopwatched, row 11b is a note not a row, and row 15 is an open DEFECT (R-104) that this line does NOT cover.
NO STATUS MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day. A stopping line that promotes a row is a stopping line that lies. Full text, and the conditions that reopen this:
documentation/backlog/OPEN-ITEMS.md, "DECIDED — the backup and restore arc is CLOSED FOR BETA".
This is the core artifact. One row per failure. It is authoritative for recovery routes;
00-capability-map.md stays authoritative for per-capability status.
How to read the numbers.
- RTO — only measured durations from INV Part F appear here. A blank cell means nobody has ever measured it, and a blank is a finding, not an omission.
- RPO — no RPO has ever been measured from an incident. These cells carry the configured
cadence that bounds RPO, read live from the box, labelled
(cadence). A blank means no cadence governs the row. - Status —
PROVEN(live, cited) ·PARTIAL(some legs proven) ·IMPLEMENTED(code + tests, never exercised) ·NONE(no route exists).
| # | Failure | What survives | Recovery route | Invocable by | RTO (measured) | RPO (cadence) | Status | Evidence |
|---|---|---|---|---|---|---|---|---|
| 1 | Customer deletes files | everything else | „Fájlok visszaállítása" — Tier-2 additive merge | customer | 39 s (6 files) | 24 h | PROVEN | CAMPAIGN-9 A1 — byte-identical returns, both non-destruction promises kept, bytes re-served by paperless's own API |
| 2 | An app's data directory is destroyed | the guest, the other tiers | same route | customer | 46 s (43 files) | 24 h | PROVEN | CAMPAIGN-9 A3 — loss verified by a positive observable (doc download 200 → 404), then 16/16 documents usable |
| 3 | An app's DB and named volumes are lost | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the only path that unpacks volume tars) | customer | 18.25 s (path execution) · 27.6 s (D5 drill, guest app.yaml absent) |
24 h | PROVEN (content recovery proven 2026-07-30) | CAMPAIGN-9 A2 proved the path executes; D5 v0.188.0 closed the content gap — after a restore with the guest's app.yaml moved aside, the app read the seeded row over TCP with its own credential, the pre-backup row returned and a post-backup row was gone (so the tar was really restored). No .sql dump in the unit ⇒ the DB came back from the volume tar. 2026-08-30 — this row KEEPS its PROVEN status and the reason is worth stating: R-353 was a defect in the MESSAGE, not in the mechanism. The restore really did return what the unit held, every time; what it could not do was say so, because the count was discarded one call deep. Controller v0.226.0 fixed the sentence and changed nothing about the recovery path. A status that measures whether data comes back must not move because a status line was wrong |
| 3c | same, with the GUEST GONE (secrets unavailable) | the drive | Tier-1 unit restore — the unit carries the portable secrets | customer | 27.6 s | 24 h | PROVEN | D5, §7.4. Before v0.188.0 this row was NONE: the data-key gate refused and the customer's own copy of their own data was not a recovery |
| 3b | same, for a class-B app via Tier-2 | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — Tier-2 unit restore (POST /backup/tier2/unit-restore → RestoreTier2Unit → RestoreFromRecoveryUnitAt), controller v0.229.0 |
customer | 28.65 s (3 volumes, 1 database, 114.5 MB unit) | 24 h | PROVEN (2026-08-31) | audits/DRILL-r102-tier2-unit-2026-08-31/. docmost — a class-B app whose Tier-2 run reports 0 leg(s) — restored with the primary unit moved aside: 3 volumes of 3 and 1 database of 1, from /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit. The observable is the DATA: an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row over TCP with its own credential; the post-backup discriminator was gone, so the tar was genuinely replayed. Repeated with the guest's app.yaml also aside → secrets recovered=2/2 from the mirrored unit. R-102 CLOSED |
| 4 | Primary drive dies | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for file legs (7 of 53 apps) and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own; Tier-3 reconstitute for files + DB + the named-volume tars since v0.218.0 (volReplay) |
customer (all) | 24 h | PARTIAL [BETA-DEFERRED] |
§7.2. Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0). This row stays PARTIAL deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. No drive has ever actually died or been replaced under this recovery — the drill removed a unit directory, not a disk, so drive re-attachment by durable_id, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore |
|
| 5 | Secondary drive dies | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (07 §8 migration rule: "Migration = rebuild, not preserve") |
automatic | 24 h | PROVEN (by construction) | tier2 v2 layout marker + rebuild, internal/backup/tier2.go. The derived-copy rule is UNCHANGED by R-403 (controller v0.230.0) and the single exception is stated in §8.2 below — read it before "fixing" a skip you find in the code |
|
| 6 | Guest lost or corrupted | the host, both whole-guest tiers, the data drives (they are host binds) | pct restore from local: or felhom-pbs: |
operator (SSH) | 84–112 s local · 1101 s PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | PROVEN | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, unprivileged: 1 preserved); LIVE restore-tests on both boxes this session |
| 7 | Guest stopped and does not come back | everything | guest-power watchdog (60 s, onboot as the deliberate-stop discriminator) |
automatic | 120 s | PROVEN | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | |
| 8 | Host dies (hardware), drives intact | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | install a new host, then --selftest=bring-up -mode dr per guest, then re-attach drives by durable_id |
operator (SSH) | 7 d | IMPLEMENTED [BETA-DEFERRED] |
bring-up code exists and has never been executed (CAMPAIGN-8…:522); the drive half of the plan is empty on every box (R-105) |
|
| 9 | Whole box lost (fire/theft) — host and drives gone | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with R → restore guests from PBS → app data from Tier-3 | operator + customer (R) | 7 d (guest) · 24 h (app data) | IMPLEMENTED / UNPROVEN [BETA-DEFERRED] |
every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (06-offsite-connectivity.md:327) [FACT] 2026-10-03 (decision 70): ep0's datastore now has a nightly copy on DooPlex (PBS pull-sync ep0-felhom-offsite, remove-vanished false, weekly verify, failures mailed to the operator); first pull 201 s, 12 GB, 4 of 4 snapshots, matching ep0 namespace for namespace. Restore route: runbooks/ep0-datastore-copy.md — never walked. audits/offsite-lock-build-2026-10-03/partF/. [FACT] 2026-10-04: the restore route from the DooPlex copy was WALKED on demo-hp (scratch VMID): listed in 2 s, restored in 186 s (15 GB logical), data read by pct mount; the restored config carries onboot: 1 and the host's real drive binds — strip them on a restore beside the original (R-834). The copy now keeps 8 weekly copies (decision 71). runbooks/ep0-datastore-copy.md route 1. |
|
| 10 | Ransomware / malicious deletion inside the guest | PBS offsite (the box cannot delete its own snapshots); the restic repo is NOT protected the same way | whole-guest restore from PBS to a point before the event | operator (SSH) | 7 d | PARTIAL [BETA-DEFERRED] |
R-89 proved the box is refused when deleting its own PBS snapshot (CAMPAIGN-8…:514-515). R-95 (open, ranked #1): the restic credential can delete — readonly=False, forget --prune runs from the box, and SFTP cannot express append-only (OPEN-ITEMS.md:13) 2026-08-30 (R-359): the store is now VERIFIED on a cadence — a daily offsite-integrity job runs restic check when the last successful one is over 7 days old. This is a readability check, not a restore-test (R-87 stays open). The depth that ships ON now re-reads 100% of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. 2026-08-31 (SPIKE R-87): a scratch restore of every app on demo-hp was measured at 25 s for 8 snapshots / 774 MB, against 40.3 s for one weekly check — but restic 0.14.0's --verify checks size and mtime, not content, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. audits/SPIKE-restic-restore-test-2026-08-31.md. 2026-09-01 (SPIKE R-95, audits/SPIKE-r95-offsite-delete-2026-09-01.md) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified. Measured on BOTH boxes over their own SFTP credential, with controls: no .snapshots is visible to either sub-account, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). The PBS shape does NOT transfer: the sub-account API has one permission axis, readonly, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. What the spike removed as a fear: withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). What it added: unlock --remove-all reports success while deleting nothing (R-430). 2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG. It searched .snapshots; the vendor documents /.zfs/snapshot. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: (a) the LIVE repository is deletable by the box — unchanged, R-95 stands; (b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited: a write into /.zfs/snapshot is refused (dest open …: Failure) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). Two limits kept honest: a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. The status is NOT moved — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. Detection shipped hub v0.111.0 (R-431): an unexplained fall in the snapshot count is noticed within a day — but only a fall of MORE than half (R-435). 2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED. The re-scope said a deletion is recoverable per-file. It is not, from the box: no snapshot is reachable by ANY name. MEASURED on demo-hp, read-only, no delete verb issued: 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS (nine full days, second granularity) plus 126 alternative shapes — zero hits, against a control where the identical 600-name batch returns a path that exists (6/6). Structural cause: /home (u629488-sub3) is st_dev 0,82, /.zfs/snapshot is st_dev 0,276, and /home/.zfs does not exist — a snapshot under /.zfs/snapshot belongs to a different dataset than the one holding the repository. Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed. So this row's "recoverable per-file (vendor)" and its limit "per-file recovery is operator-only today (R-432)" both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed. audits/evidence-drill-r95-recovery-2026-09-01/. [FACT] 2026-10-03 (decisions 68–69, hub v0.127.0, controller v0.289.1): THE BOX CAN NO LONGER DELETE ITS RESTIC HISTORY. Both demo boxes run on an append-only key the hub pinned; a delete from the box is refused (403 Forbidden, measured on each, counts 13→13 and 100→100); the box never receives the sub-account password (the old endpoint answers 410, measured with demo-hp's own key); the hub reads every key file daily. Retention runs only in a hub-opened window behind a fake-snapshot guard — mechanics proven live (window 1 on demo-hp: opened, guard refused, closed in 3 s, operator mailed); a real prune inside a window is NOT yet proven (R-824). Residual: an add-only attacker can still fill the quota and plant past-dated snapshots (R-822). audits/offsite-lock-build-2026-10-03/. [FACT] 2026-10-04 (controller v0.290.0, hub v0.128.0): the guard now skips same-day-superseded young snapshots instead of refusing; window 2 on demo-hp ran without refusal and removed nothing (127→127 — every candidate was a young same-day copy); weekly windows are ON fleet-wide. A window that actually removes snapshots has still not been observed (first candidates age past 8 days around 2026-10-11). The household's set-aside deletion is the hub's after 7 days (decision 74), proven live on a planted scratch dir. [FACT] 2026-10-05 (controller v0.294.0, R-867 CLOSED): A WINDOW THAT REMOVES SNAPSHOTS IS OBSERVED, on both demo boxes. The guard's line is now built from the policy's own constants (keep-daily 7 CALENDAR days, 09 decision 104) — the 8-day age line sat inside the keep window and refused every honest window (demo-felhom 2026-10-05 02:15 UTC, an error mail). One window each, opened by the operator's one-shot grant and run by the household's "run now": demo-felhom 16 → 14, demo-hp 145 → 127 — the removed snapshots are exactly those the policy predicted; the hub's rows read pruned with the same counts, no drop event, no mail, the key files clean. The UNATTENDED weekly window has not yet been observed removing (next due ~2026-10-11/12). Residual unchanged: R-822 (past-dated gap-fills steer older keeps). audits/night-fixes-2026-10-05/partA/. |
|
| 11 | Hub lost | every box's data plane, every tier, every Lane-1 route | none needed for recovery of a box; the hub itself restores from its Longhorn volume backup | operator (kubectl) |
24 h + weekly (Longhorn RETAIN 1 each) |
UNPROVEN [BETA-DEFERRED] |
a hub restore has never been performed. The backup target is nfs://192.168.0.180 — DooPlex itself — and exactly 2 restore points exist (LIVE, INV Part D2.2) |
|
| 11b | consequences while the hub is gone | — | — | — | [FACT] [BETA-DEFERRED] |
day: nothing customer-visible breaks; events queue (settings.go:1466-1485). week: the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. permanently: escrow custody and break-glass credentials are gone (INV Part D2.4) |
||
| 12 | Offsite provider lost (Hetzner) | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | [FACT] [BETA-DEFERRED] |
restic and PBS share the provider (INV Part E.1). Whether they share an account and payment method is UNKNOWN → §11-D [FACT] 2026-10-03: the PBS (whole-guest) off-site copy now also exists OFF Hetzner — DooPlex pulls ep0's datastore nightly (decision 70). The restic tier is still Hetzner-only. |
||
| 13 | Customer loses R | every byte, every tier, the Recipe | on-premises recovery is unaffected — Tier-1/2/3 and the whole-guest tiers all work while the guest and host live. What is lost: the ability to re-establish host identity and to open the escrowed offsite key material after a host loss | — | NONE for host-loss | R exists in zero system copies by design (INV Part C, row 24). Whether the operator should be able to recover is open → §11-A, §11-B. D5 narrows this further: Tier-1/2 app recovery needs only the drive, so losing R now costs the offsite route and host identity, never local app recovery (§7.4) | ||
| 14 | Host SSH management plane dies | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted root@pam on the PVE console |
operator | PROVEN | runbooks/break-glass.md; the original incident and its fix |
||
| 15 | Interrupted offsite run leaves a stale lock | everything | manual restic unlock --remove-all |
operator (SSH) | DEFECT | the self-heal exists (offbox.go:634-648) but the probe fails first and classifyResticProbe has no lock case → fail-fast, operator told „ismeretlen okból" → R-104 |
8.2 The ONE exception to row 5's derived-copy rule (R-403, controller v0.230.0)
[DESIGN] Row 5 stands: the secondary IS a derived copy and it IS rebuilt on the next run. Nothing
below weakens that, and a future reader who finds RunTier2 skipping a leg and does not find this
section will "fix" it back — which is why it is here and not only in a register row.
[FACT] What was measured, on the shipped v0.229.0, on demo-hp 2026-08-31. An app's Tier-2 copy
went from 120 082 104 B (4 database dumps + 3 named-volume tars) to 7 036 B (none of either)
in one nightly run, and the run recorded itself a success. The mechanism was three individually
correct lines: RunTier2 guarded the unit leg with os.Stat alone — does the folder exist —
rsyncMirror is rsync -a --delete, and nothing between them compared source to destination. An
empty recovery unit is a folder that exists. Evidence:
audits/DRILL-r403-tier2-delete-2026-08-31/.
Why it bites harder since 2026-08-30: R-102 made that mirror a LIVE recovery route (§6.3, §8 row 3b). Deleting it used to cost a copy nobody could open; it now costs the route itself.
[DESIGN] The exception, stated exactly. The unit leg — and ONLY the unit leg — is skipped when the SOURCE unit carries no data and the DESTINATION unit does. Everything else is unchanged:
| source unit | destination unit | behaviour |
|---|---|---|
| complete | complete | mirror, with --delete, as before |
| complete | hollow or absent | mirror (the normal first copy) |
| hollow | hollow | mirror — both sides agree, nothing is at risk |
| hollow | complete | skip the unit leg, preserve the destination, warn, record for the surface |
"Hollow" is a MANIFEST question, never a size question — the manifest lists no database dump and no volume tar; absent or unparseable counts as hollow, fail closed. A unit with a fat compose capture and no dumps is the dangerous shape; a 360-byte unit belonging to a tiny app is healthy.
The data legs are NOT guarded and must not be. A classified app's copy legitimately shrinks as
export drops out of its class set (tier2.go header), and fencing that would be calling this row's
own decision a defect.
[DESIGN] And the CAUSE is closed at the other end. The hollow primary was written by the 5-minute
capture job two seconds after a Tier-2 unit restore. Since v0.230.0 RestoreTier2Unit refills an
absent or hollow primary unit from the mirror inside the call, before returning, so no capture can
observe the hollow state. The capture itself is deliberately not guarded: a capture that describes
an empty drive as empty is correct, and guarding it would make the manifest lie.
8.1 The blank cells, listed explicitly
Per the rule that a blank is a finding, here they are:
| row | blank | why |
|---|---|---|
| 4 | RTO | no drive-loss recovery has ever been timed. The Tier-2 unit restore inside it now is — 28.65 s, row 3b, 2026-08-31 — but that is the route, not the journey: no drive has been removed or replaced under a recovery |
| 5 | RTO | never timed; the rebuild is a normal Tier-2 run |
| 6 | — | RTO present, but it is a restore into a scratch guest on the same host; a restore to a different host has never been timed |
| 8 | RTO | a host has never been rebuilt as itself (INV Part D1) |
| 9 | RTO | the composed whole-box path has never been run |
| 10 | RTO | no ransomware-shaped recovery has ever been run |
| 11 | RTO | a hub restore has never been performed |
| 12 | RTO, RPO | no provider-loss recovery has ever been run |
| 13 | RTO, RPO | not a timed recovery; a capability loss |
| 14 | RTO | the break-glass path is proven but was never timed |
| 15 | RTO | the manual unlock was performed but not timed |
Also unmeasured, and not representable as a row (INV Part F.3): any restore larger than 155.5 MB
from restic; any restore of a customer data drive (no tier holds one whole); the weekly PBS
incremental; .fab import wall-clock; and time-to-first-byte for a customer restore over a home
uplink — no customer has ever driven a restore.
9. What the model implies for the tiers (recorded, not new design)
[DESIGN] Three consequences follow from §3–§7 and are stated so they are not re-derived:
- Tier-2 is a drive-loss tier, not a second chance at Tier-1. Its job is to survive one drive dying. That is why it mirrors the unit at all — and why the unit mirror being unreadable (§6.3) defeats the tier rather than degrading it.
- Tier-3 is a premises-loss tier. It is the only copy that survives fire, theft and ransomware-with-host-access, which is why it is the only tier whose credential exposure is ranked as the largest open data risk (R-95).
- The whole-guest tiers are the system tier, and they sit under everything else (§7.1). Treating them as "one more copy" is precisely the 3-2-1 framing this document rejects.
10. Known gaps
Every divergence between the model above and the system as it is, each with an ID.
10.1 D5 is BLOCKED — precondition CLOSED by R-108 (v0.187.0); D5 ITSELF SHIPPED (v0.188.0)
D5 IS DONE — 2026-07-30, controller v0.188.0. The precondition below was met by R-108, and D5 was then implemented and proven live the same day: Tier-1/Tier-2 no longer depend on the whole-guest tier, and a customer needs the drive and nothing else. The current recovery chain is §7.4; the corrected matrix rows are 3, 3c and 13. This section's remaining value is the precondition analysis, which is why the table below is kept.
D5's PRECONDITION IS MET. An app's data namespace can no longer be placed on network storage, so no
backups/tree can exist inside FileBrowser's share-root bind, and the browsing surface therefore cannot reach the backup tree on ANY storage class. Every other read surface in the table below was already NO. D5 may be adopted — nothing in this section blocks it.The fix inverted the obvious one, and that is the durable lesson here. The share-root bind was not narrowed, because it cannot be: (a) the
:rslaveshare-ROOT bind is load-bearing — a Phase-0 probe (2026-07-22) proved an in-container access through it wakes the idle automount trigger, so narrowing it breaks NAS access itself; (b) there is nouserdata/layer to scope to, since apps on a share store at<share>/<app>; and (c) creating one would write Felhom's directory convention onto a customer's own NAS, which R-67 forbids outright. The browsing surface being immovable is precisely why the backup tree must never be placed under it. Tier 2 had already reached the same conclusion for its own targets (F-6C-1); R-108 closes the PRIMARY namespace, which was the last remaining route.Operator ruling 2026-07-30: refuse the placement, keep the browse bind.
RefuseAsAppNamespace(internal/settings/settings.go) is the single predicate; all placement surfaces consult it. It fails closed —/mnt/felhom-drivesholds both storage kinds in-guest, so a path prefix cannot classify andKindexists only on a REGISTERED path; an unregistered path under that root is therefore un-classifiable and is refused rather than assumed to be a drive.Nothing was stranded: zero apps on network storage across all six hub customers including Peti. R-67's browse capability is byte-for-byte unchanged (verified by diffing demo-hp's generated compose before and after the deploy). This also supersedes the controller README's "NAS backup locality — decision A" (v0.118.0), which deliberately kept a NAS-resident app's Tier-1 artifacts on the NAS: that case can no longer arise.
Evidence:
audits/R108-network-app-namespace-2026-07-30.md. The pre-fix analysis below is retained verbatim as the record of what was wrong.
10.1 (historical) The exposure as it stood before v0.187.0
[FACT] The verification and its result. Every surface that can read a file was checked:
| surface | can it reach backups/? |
evidence |
|---|---|---|
| SMB share creation | NO | sharingResolvePath (internal/web/sharing_handlers.go:52-81) resolves symlinks before containment, then refuses any path within SharingDeniedRoots(root); that set covers <root>/backups and the legacy <root>/felhom-data + <root>/felhom-data/backups (internal/stacks/samba.go:46-65, derived from ProtectedHDDPaths, delete.go:59-78) — both namespace shapes |
| SMB browse (folder picker) | NO | same deny set applied per child (sharing_handlers.go:521-537) |
SMB ensureImportShare (the store-direct bypass) |
NO | writes one controller-generated constant, GetImportRoot() = <system ns>/userdata/import (sharing_handlers.go:564-587) |
| FileBrowser — local drives | NO | the bind is appbackup.UserdataDir(sp.Path) only, and the comment says why (internal/web/handlers.go:2450-2460) |
| FileBrowser — network shares | the bind is still the share ROOT (- %s:/srv/%s:rslave, handlers.go:2437) and deliberately so — but no app namespace, hence no backups/ tree, can exist on a share, so the root bind reaches only the customer's own files. The reachability is closed at the PLACEMENT, not at the bind |
|
.fab import path validation |
NO | confined to <root>/exports (handler_export.go:400-408, estimate.go:215-217) |
.fab browser download |
NO | name-pattern + parent-must-be-the-staging-dir double guard (handler_export_download.go:36-45,120-140) |
/api/debug/* |
NO | no file-serving branch (handler_debug.go:46-92) |
http.ServeFile (3 sites) |
NO | assets only, filepath.Base-normalised (server.go:720-771) |
| log bundles | NO producer found | no filesystem-walk bundle producer exists on the box; searched felhom-controller/internal, felhom-agent/internal, cmd/ |
| registering the backup dir as a drive | NO | the manual add requires system.IsMountPoint(path) (handlers.go:2091-2095); <drive>/backups is not a mount point |
The exposure, end to end. All six links are source-cited and the precondition is live today:
- A NAS share is registerable as a storage path and lands
Schedulable: true— LIVE on demo-hp:{"path":"/mnt/felhom-drives/Felhom-Share","kind":"network","schedulable":true}. GetSchedulableStoragePaths()has noIsNetwork()filter (internal/settings/settings.go:904-914), so that share appears in the deploy dropdown (handlers.go:462-473).- The per-app migrate target list filters only current / decommissioned / disconnected /
schedulable — also no network filter (
handlers.go:674-679). handleStorageMigrateAppdoes not callrefuseNetworkLifecycle, unlike its whole-namespace sibling which does (storage_handlers.go:397vs:410-424), andstartMigrationhas no guard either (internal/stacks/migrate.go:214-255).- With
HDD_PATHon the share,namespaceRoot()returns it as-is under Model A (internal/backup/backup.go:262-263), so the app's Tier-1 unit is written to<share>/backups/primary/<app>/compose/app.yaml. - FileBrowser binds that share at its root and serves it with
download: true(internal/infra/infra.go:326).
Observed live on demo-hp, in the generated compose — the asymmetry is visible, not inferred:
- /mnt/felhom-drives/nvme-1tb/userdata:/srv/nvme-1tb <- drive: userdata-scoped
- /mnt/felhom-drives/Felhom-Share:/srv/Felhom-Share:rslave <- network: ROOT-bound
- /mnt/sys_drive/felhom-data/userdata/import:/srv/beolvasas
Today this is not a secret leak, because the unit's app.yaml is secret-stripped
(recovery_unit.go:73). D5 would make it one. That is exactly the test §2 set, and D5 therefore
does not hold as written. → R-108
CLOSED (v0.187.0). Links 2, 3 and 4 of the chain above are now guarded, and a fifth surface the chain did not list was found and guarded too:
handleStorageDecommissionmode=migratechecked onlyreq.Where(the SOURCE) viarefuseNetworkLifecycle, so a whole namespace could be decommissioned ONTO a NAS. Link 2's framing also understated the problem — the deploy dropdown is only a UI list; the boundary is the deploy POST (internal/api/router.go), which accepts any caller-suppliedHDD_PATHand whose only other validation isos.Statexistence (internal/stacks/deploy.go). Filtering the list alone would have left the surface open. Links 5 and 6 are unchanged and still true — they simply can no longer be reached.
10.2 The gap register
| ID | Gap | Consequence |
|---|---|---|
| R-102 | Tier-2 writes a full recovery-unit/ mirror on every run and no code path reads it |
Tier-2 is defeated in the drive-loss scenario it exists for (§7.2). Was C9-F4 |
| R-103 | The Tier-2 no-coverage refusal names the working action but does not route to it | 45-or-43 of 53 apps dead-end at a message. Was C9-F1b |
| R-104 | An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach, reported as „ismeretlen okból" | the offsite tier stays dead until a human unlocks. Was C9-F3 |
| R-105 | Three hub-held DR records are empty on the whole live fleet: hosts.dr_record_json, host_escrow.directive_json, dr_recipe.host_half.drives |
the Recipe (§4) is incomplete in exactly the fields host-loss recovery reads. Causes may differ per field |
| R-106 | dr_recipe.host_half.pbs.namespace records "root" on every box |
the recorded restore coordinate is wrong; real namespaces are per-customer |
| R-107 | volReplay, proven live on demo-hp). True from the day Tier-3 shipped until 2026-08-21. |
was: offsite alone cannot rebuild a named-volume app (§7.2). Now: the tars replay; what remains is that the app must be deployed (R-253), which is not R-107 |
| R-356 | The off-site restore resolved its destination with the raw HDD_PATH and read an empty answer as "not installed" — CLOSED, controller v0.219.0, 2026-08-22 |
40 of 53 apps were refused permanently while running (§6.3 [DESIGN]) |
CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED. An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). audits/R108-network-app-namespace-2026-07-30.md |
||
| R-126 | A .fab bundle — plaintext secrets, optional password — can be exported ONTO a NAS: storageDriveList() (internal/web/handler_export.go) does not filter network paths |
split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) |
| R-95 (open) | The restic offsite credential can delete — the box can forget --prune its own repo, from two call sites (offbox.go:1388 retention and offbox.go:1759 over-quota) |
the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). SPIKE 2026-09-01: prevention needs a transport change (restic 0.14.0 does speak rest: — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. Detection is nearly free and is recommended first — snapshot_count already reaches the hub and the hub appends reports, so the comparison needs no box change. audits/SPIKE-r95-offsite-delete-2026-09-01.md. RE-SCOPED 2026-09-01: the box can delete the LIVE repository but cannot write to the daily snapshots of it (measured, both boxes) — so the exposure is at most one day's data plus an operator-driven per-file recovery, not open-ended loss. Detection shipped hub v0.111.0 (R-431). DRILL 2026-09-01: the re-scope's recovery clause is WITHDRAWN — no snapshot is reachable from a sub-account by any name (R-433, 777,600 names, controlled), so the exposure is NOT bounded by an operator-driven per-file recovery. R-436 is the new cheap lead: the provider already offers rclone serve restic --stdio server-side and restic 0.14.0 speaks rclone: (measured) — but the client supplies the server command line, so ask the vendor whether --append-only is pinned BEFORE building anything. |
| CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0. Restore-testing is now per archive generation: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) | ||
CLOSED 2026-08-30, controller v0.226.0. The volume count already existed (restoreDockerVolumesFrom) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. RestoreFromRecoveryUnit now returns UnitRestoreResult and the sentence has three cases, keyed on replayed-vs-listed — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no SafetyDump discriminator, and §6.3 is why that is not pedantry. Proven live on demo-hp |
||
CLOSED 2026-08-30, controller v0.226.0. The gate sits before writeSafetyDump and StopStack, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. Not live-validated — filling a filesystem is a drill step |
||
OffboxFullScratchReady asked "non-empty directory", which is exactly what a failed restic run leaves |
CLOSED 2026-08-30, controller v0.226.0. A completion marker written only after restic returns nil, stale ones cleared before it starts, both orders pinned by an AST test. Both handlers refuse server-side — a hidden button is not a guard. Proven live on demo-hp. See R-396: the bad state needed only a SUCCESSFUL safe restore, not a failed one |
|
CLOSED 2026-08-30, controller v0.226.0. restoreOpBlocked(), matching the five sibling handlers. The doc comment that claimed it already did this is corrected in place — that sentence is why nobody looked. Proven live on demo-hp in the exact flag state that produced it |
||
CLOSED 2026-08-30, controller v0.226.0 (by R-358's marker). Found while answering R-358's open question. Both restore modes write the SAME scratch directory, and one boolean (ScratchReady) drove three different intents — "is there a scratch", "may we place", "may we destructively restore". The safest action on the page unlocked the most dangerous one |
||
CLOSED 2026-08-30, controller v0.227.0/v0.227.1. A daily offsite-integrity job on due-ness, not a weekday; it takes the single-writer flag and SKIPS rather than waits (resticStep escalates to unlock --remove-all and is only safe while that flag is held). Three outcomes — skipped / unreachable / failed — because 'I could not look' is not 'I looked and it is broken'. ⚠ The depth that ships ON does NOT catch silent corruption: measured, a pack corrupted without a size change returned no errors were found, exit 0; only --read-data* caught it. Choosing the depth is R-399 |
||
NotifyIntegrityOK/NotifyIntegrityFailed had no caller and the product advertised a weekly check that did not exist |
CLOSED 2026-08-30, controller v0.227.0. Sixth built-but-never-wired instance: hub allowlist, Hungarian text, settings checkbox and debug button all existed; only the caller did not. ok is severity info and mails nobody by design |
|
CLOSED 2026-08-31, controller v0.228.0. monitoring.integrity.read_data_subset now defaults to 100%, so the weekly check downloads and re-hashes every stored byte. The fact that made it necessary, and the sentence that should stop anyone turning it back down to save four seconds: the structure check PASSED a size-preserving pack corruption. Measured on demo-hp 2026-08-30 — plain restic check reported no errors were found and exited 0 over a pack damaged without a size change; every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100%. off (any case) returns a box to structure depth; an empty value means not configured, therefore the default; a malformed value WARNs and falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs an operator WARN naming R-401 — one data point, on one 134 MB store, so no rotation schedule, size threshold or bandwidth budget was invented from it. Proven live at both depths 2026-08-31 with the restic argv observed from the guest |
||
tests/r87-offsite-proof-2026-08-31/; the reasoning is audits/SPIKE-restic-restore-test-2026-08-31.md. |
SPIKE VERDICT, audits/SPIKE-restic-restore-test-2026-08-31.md: build the NARROW version, not the row as written. Measured: a scratch restore of all 8 apps costs 25 s / ≤213 MB scratch, LESS than the 40.3 s weekly check beside it; but restic 0.14.0's --verify is a size-and-mtime reconciliation (a size-and-mtime-preserving corruption passed clean, red-proofed), restic ls --json carries no content hash, and the unit manifest hashes 4 918 B of a 213 231 242 B unit — so no reference for "correct" exists (R-409). Of the five drill-found restore defects R-353/354/356/358/403, an unattended scratch-restore would have caught ONE (R-356). The value is elsewhere and the weekly check structurally cannot reach it: check proves the stored bytes are the stored bytes, never that we stored the RIGHT thing — a hollow unit backs up, checks and restores cleanly and recovers nothing (R-403, measured 2026-08-31). Proposed re-scope, Viktor's call: prove the off-site snapshot still CONTAINS a recoverable unit — one app a night, restored to scratch, checked against its own manifest.json via the existing unitCarriesData. Must use --no-lock and skip unlockStale (R-95's constraint is otherwise violated — R-407/R-408 record what the path writes today) and must take acquireRunning, which RestoreOffboxScratch does not. |
10.3 Divergences that are documented elsewhere and are not re-opened here
[FACT] 06-offsite-connectivity.md:19-21 describes the operator's public edge as a Cloudflare
Tunnel and states DooPlex has no public IP. Live DNS resolves hub.felhom.eu through a no-ip
DynDNS CNAME straight to DooPlex's own public address (INV Part E.3). That is a topology-document
issue, not a recovery-model one; recorded so the discrepancy is not lost.
11. Open decisions — for the operator
Recorded, deliberately not answered.
A. Escrow custody. Split custody (R or an offline operator key) versus a 2-of-3 threshold across customer / hub / box drives. Recommendation on record: *split custody, operator key held offline and never in the hub. Note the constraint from §2: the moment the operator key lives in the hub, "a compromised hub yields blobs nobody can open" stops being true.
B. Lost-R policy. Under split custody the operator can recover. Is that the stated policy — and if so, what identity check gates it? Today matrix row 13 has no route at all, and the customer is not told that losing R costs them the host-loss route.
C. RTO / RPO targets per scenario. None have ever been stated. Without them §8 cannot judge whether, for example, offsite-only recovery is acceptable for drive loss, or whether the 7-day PBS cadence is adequate for row 8. The measured column is the input; the target column does not exist.
D. Hetzner as a single failure domain. restic (Storage Box u629488, sub-accounts per customer)
and PBS (Cloud server ep0 + a Cloud Volume) are both Hetzner. Whether they share an account,
login and payment method is unverified (INV Unknown U-1) — the question was deliberately not
answered by calling the provider API with the production token. Accept explicitly, or mitigate.
The fence was reached again, and held, 2026-09-01 (SPIKE R-95). The R-95 spike needed one provider-side fact — do daily snapshots exist on the Storage Box, and how many — because the register calls that mitigation ARMED and the whole ranking of R-95's options turns on it. It is readable from the box: measured on BOTH machines over their own SFTP credential, with positive and negative controls, no
.snapshotsis visible to either sub-account, and the account is jailed at/. The register's own confirming field,size_snapshots, is an API field. So the spike stopped here rather than answering it by acting, and left it as R-429 — ten minutes in the Storage Box panel, and it re-ranks R-95.audits/SPIKE-r95-offsite-delete-2026-09-01.md§Q1.
[FACT]2026-10-03 — append-only on the Storage Box, MEASURED (R-436). On the provider (u629488-sub4, scratch repo, since removed), a key pinned inauthorized_keystocommand="rclone serve restic --stdio --append-only <dir>",restrictbacks up, lists, restores and checks, and every delete is refused:blob not removed, server response: 403 Forbidden (403); the same command through an unpinned key deletes. The pinned key gets no shell, sftp, scp, rsync or port forward. But the PASSWORD logs in on ports 22 and 23 and can rewriteauthorized_keys, and a box can obtain that password from the hub's off-site self-heal at will — so the pin protects nothing until the box stops receiving the password (R-820). Port 22 accepts no OpenSSH-format key; 23 does. Locks: a crash lock blockscheck, notbackup;unlock --remove-allclears it through the pinned key. An add-only key can still plant future-dated snapshots that make the box's retention policy select every real snapshot (R-822). Who prunes is a PROPOSAL awaiting the operator, not design:audits/offsite-append-only-2026-10-03/DESIGN.md; evidence…/live/,…/lab/.
E. local vzdump shares a physical device with the guest it backs up. LIVE on both hosts:
/var/lib/vz (the archive target) and local-lvm (the guest's rootfs and both backup=1
mountpoints) are both on /dev/sda3 → VG pve. The tier therefore protects against corruption
and operator error only, never against disk failure. Accept and name it honestly in the customer-
facing description, or move the target.
F (added by §10.1, not in the original list). — ANSWERED 2026-07-30. D5 could not be adopted until R-108 closed. Closing R-108 was the intended path, and it is done (controller v0.187.0): the operator ruled to refuse app namespaces on network storage rather than narrow the browsing surface, because the share-root bind is load-bearing and cannot be scoped. D5 is no longer blocked. Whether to now implement D5 remains an open scheduling decision, not a blocked one.
Decided, and deliberately not built — carried here 2026-10-03 when their register rows moved to
backlog/CLOSED-ITEMS.md (the reasoning is in CONTEXT.md, "Rules carried out of rows closed 2026-10-03"):
- [DESIGN] No automatic abandon on a date (R-245, decided 2026-08-07). A customer who never decides is not auto-abandoned after 30 days. An automatic ending, if ever built, triggers on the HARM (quota — old history blocking new backups), with a dated warning, never on a date alone. Reopens on quota.
- [DESIGN]
markOrphanedkeeps no guard against an active abandon countdown (R-303, decided 2026-08-13). The co-render is made harmless, not impossible; suppressing the orphan card during a countdown would hide a real second fault. Reopens on a real-world sighting. - [DESIGN] No in-product route from the recovery screen to a set-aside store (R-312, decided 2026-08-13).
Retention is an operator-only capability, and nothing anywhere may promise the customer can perform it
themselves. Re-evaluate on a real customer request. Its fixture,
demo-felhom's unrecoverable set-aside store (R-313), is kept until R-312 ships or is abandoned.
12. Evidence index
| Claim | Grade | Source |
|---|---|---|
| Tier-2 file restore, gap-fill and after total loss | PROVEN-LIVE | CAMPAIGN-9 A1/A3 |
| Tier-2 refuses without an outage for a no-coverage app | PROVEN-LIVE | v0.183.0 replay, felhom.eu/REPORT.md:60 |
| Tier-1 unit restore executes | PROVEN-LIVE | CAMPAIGN-9 A2 |
| Tier-1 content recovery after loss | UNPROVEN | CAMPAIGN-9…:825-827 |
| restic restore of app data (bytes) | PROVEN-LIVE | CAMPAIGN-8 R-87, 155.5 MB, 6/7 byte-identical |
| offsite reconstitution of a DB-indexed app | PROVEN-LIVE | destructive immich drill 2026-07-20, 00-capability-map.md:66,75 |
| offsite place-to-live as a distinct action | UNPROVEN | CAMPAIGN-8…:520 |
| shares restore (files + definitions + credential) | PROVEN-LIVE | 2026-07-18, 00-capability-map.md:96 |
.fab drive-to-drive round trip |
PROVEN-LIVE | CAMPAIGN-6D P-FAB, 1.7 GB |
.fab browser upload leg |
UNPROVEN | 00-capability-map.md:67 |
| whole-guest restore, local and PBS, exact mount parity | PROVEN-LIVE | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C |
| corrupted PBS snapshot fails cleanly | PROVEN-LIVE | CAMPAIGN-8 fault 17 |
| the box cannot delete its own PBS snapshots | PROVEN-LIVE | CAMPAIGN-8, R-89 |
| unattended restore-test across tiers | IMPLEMENTED; per-archive due-ness PROVEN-LIVE 2026-08-03 (agent v0.121.0) | 00-capability-map.md:41; the due verdict + a real offsite run on demo-felhom (§3, Lane 2's per-archive rule) |
| guest-power watchdog | PROVEN-LIVE | agent v0.107.0, 120 s |
| quiesce crash recovery | PROVEN-LIVE | CAMPAIGN-8 fault 10, 1 s, by SIGKILL |
| break-glass | PROVEN-LIVE | runbooks/break-glass.md |
agent DR bring-up (ModeDRGuestLoss) |
NEVER EXECUTED | CAMPAIGN-8…:522 |
| host-loss plan → an actual restore | EXECUTES NOTHING BY CONSTRUCTION | felhom-agent/internal/dr/plan.go:1-4 |
| host rebuilt as its former self | NEVER DONE | INV Part D1 |
| escrow consume in a real recovery | SPIKE-LEVEL ONLY | 06-offsite-connectivity.md:327 |
| hub DB restore from its Longhorn backup | NEVER DONE | INV Part D2.2 |
| a customer performing a restore unassisted | MISSING AS EVIDENCE | 00-capability-map.md:75 |
13. What this document deliberately does not do
- It does not restate capability status — §8 cites the map, the map cites §8.
- It does not resolve the 7/53-vs-9/43/1 count (§6.2). Both are on the record; neither is adopted.
- It does not estimate a single RTO or RPO. Every blank in §8 is a real gap.
- It does not answer §11. Those are the operator's.
- It does not claim ratification. | 5 | A customer had no way to BEGIN — the only route was a command line | Everything above was self-service, and nothing told the owner of a rebuilt box that a sealed package was waiting or how to open it. | CLOSED — controller v0.200.0 (R-193). A full-page recovery screen, shown while the hub holds a package this box cannot open; it unlocks and lists, and restores nothing (→ R-213 for the put-back). |