Files
felhom.eu/documentation/architecture/07-backup-architecture.md
T
Claude Code 945b7818b5 docs(arch): 07 §9.1 — record measured PBS whole-guest capacity state (R-82 Phase 0)
Per the operator's 2026-07-26 ruling: datastore growth deferred, R-82 proceeds.
Records the measurements so the constraint is written down rather than carried
in a session: 37.2 GB total / 10.8 GB used, no cross-customer dedup (per-tenant
encryption keys), 80% warn reached at roughly the second additional customer,
and the pvesm 0/0/0 reporting artifact that means operators must read fill from
the hub gauge. Also records what the tier does and does not carry, and the
conditional on the P0.1 weekly verdict (Tier-3 offsite must be healthy).

Doc NOT marked ratified — that stays Viktor's review (R-83).
2026-07-26 12:11:42 +02:00

32 KiB
Raw Blame History

07 — Backup architecture: tiers × classes × targets

Status: DRAFT — awaiting Viktor's review (async one-line vetoes on the §10 list). Written 2026-07-14 per the architecture-doc-first gate (Viktor ruling #5). Every claim below was verified against live Gitea at commit felhom-controller 95f3180 (v0.132.0), felhom.eu deacee11, catalog 21e8df1, agent c040c18 (v0.88.0), hub v0.54.0. Line numbers are landmarks — reconfirm before editing.

Inputs: SPIKE-backup-classification-2026-07-14.md (790ec84), CAMPAIGN-6C-2026-07-14.md (deacee11, F-6C-1), Viktor's locked rulings of 2026-07-14. This document DECIDES; it does not survey. Where a decision rests on an unproven restic mechanism, it is marked SPIKE and listed in §6.4 — per the spike-first gate, those spikes precede the Task 3a spec.


0. Verified baseline — what the code does today (source, not memory)

Surface Behavior Evidence (v0.132.0)
Recovery unit (tier-1) compose/ (docker-compose.yml, .felhom.yml, secret-stripped app.yaml) + enumerated db-dumps/*.sql + volume-dumps/*.tar + manifest.json (SchemaVersion 1). No userdata, no HDD appdata. recovery_unit.go:91-106,131
Offsite (tier-3) One restic backup <unitdir> --tag felhom-offbox --tag <stack> per toggled app; unit dir discovered across nsRoots, newest-manifest tiebreak. Retention forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune, default grouping = host+paths. offbox.go:504-552 (discover), :575 (backup), :590-597 (forget)
Offsite quota Soft gate on RepoSizeBytes from restic stats --json with no mode flag → default restore-size over all snapshots; ≥100% refuses new backups (prune still runs), ≥80% warns. offbox.go:407-415,630-645,700-708
Offsite restore restic restore latest --tag <stack> --target <DataDir>/offbox-restore/<app> — scratch, non-destructive, whole snapshot. DataDir default = /opt/docker/felhom-controller/data — the guest rootfs. web/offbox_handlers.go:236-246, config/config.go:310
Tier-2 rsync mirror of unit → backups/secondary/<stack>/recovery-unit + resolver-driven appdata dir → .../appdata (flat; refuses N>1). rsync -a --delete. Target pick: pinned > other data drive > SSD (headroom-gated); zero fs-type awareness (F-6C-1). Size for headroom = dirSizeBytes(unit)+dirSizeBytes(appdata). tier2.go:170-249 (run), :92-155 (select), :437-445 (rsync), :198 (size)
Tier-2 restore Missing-files-only merge into live appdata: rsync -a --ignore-existing; refuses N>1. tier2_restore.go:15-44,127-137
Manual .fab Config + DB dump + volume tars + ExportDataMounts = ${HDD_PATH} binds the whole userdata root as one tar (SQ6 over-capture). stacks/delete.go:575, appexport/export.go (v0.130.0)
Classification (INERT) stacks.Manager.ClassifiedBinds(name)([]ClassifiedBind, hasClassification); whole-block-reject at LoadMetadata; two-level default (explicit > :ro-reader > writable-mandatory; no block → legacy). Wired seam: StackDataProvider.GetStackClassifiedBinds (adapter main.go ~L1106). 13 catalog apps carry blocks (21e8df1). metadata.go:285-294, appbackup/classify.go:97,165
Network storage metadata settings.StoragePath.Kind == "network" + IsNetwork(); Protocol nfs|smb. Drive lifecycle already refuses network paths. settings.go:181-219,1057

Two new findings from this verification pass (not in the spike or 6C):

  • F-A1 (HIGH for 3a): the offsite restore scratch dir is a rootfs-filler once snapshots carry userdata. offbox-restore/<app> lives under cfg.Paths.DataDir, default /opt/docker/felhom-controller/data on the ~8 GB guest rootfs. Today snapshots are unit-sized (MB1 GB); with mandatory userdata (nextcloud's data dir, immich's upload library) a restore latest can fill the rootfs and take the guest down. §7 decides the fix.
  • F-A2 (MEDIUM for 3a): quota accounting mode is wrong-by-design for the new shape. restic stats default restore-size aggregates across all retained snapshots (~17 per app at full retention). Exact cross-snapshot counting semantics are undocumented enough to matter: if a 50 GB mandatory library counts once per snapshot, the gate trips at ~17× reality. §9 decides the candidate fix (--mode raw-data); SPIKE SP-1 proves it.
  • Minor: tier2.go:167 (RunTier2 doc comment) still claims "recovery unit + userdata" — the last surviving F-S1 stale comment (fourth stale-comment instance this arc). 3b makes it true.

1. The class model (normative recap)

A class belongs to a bind (host path), not to an app (SQ2 rule, locked):

  • mandatory — referentially COUPLED to app state (DB rows reference the content; SQ3: restore without it = broken-not-empty). Never deselectable at any tier that carries userdata.
  • optional — DECOUPLED but precious (hand-curated, not re-acquirable). Customer-selectable.
  • excluded — DECOUPLED bulk/transient/re-derivable. Never captured automatically; .fab opt-in only, behind the two-number warning.
  • legacy (origin, not a class) — app has no backup: block → every tier behaves byte-identically to v0.132.0. The safety net fails toward today's cost, never toward a quota blow-up (SQ5).

Roots: ${HDD_PATH} (hdd: list; appdata stores) and ${USERDATA_PATH} (userdata: list; browsable tree). RelPaths are ${VAR}-relative and hddPath-invariant (survive migration flips — the Task 1 lesson). Resolution to absolute paths happens at capture time against the app's live HDD_PATH (Model A: HDD_PATH == nsRoot).


2. The decision table: tier × class × target

Tier legacy (no block) mandatory optional excluded Allowed targets Copy mechanism Restore path
Tier-1 recovery unit config + DB dumps + volume tars unchanged — the unit never carries userdata; classes ride inside it via the captured .felhom.yml app's own drive (backups/primary/<app>) direct writes RestoreApp (existing)
Tier-2 cross-drive unit + resolver-appdata (byte-identical to v0.131.0) unit + mandatory binds + optional binds never real local drives only (network paths excluded per F-6C-1 ruling — both auto and pinned); SSD fallback = unit + mandatory iff headroom fits, optional skipped with honest reason rsync -a --delete per capture-set path missing-only merge (--ignore-existing), per-path from the new layout (§8)
Tier-3 offsite restic unit only (today's shape) unit + mandatory binds — not deselectable never never Hetzner Storage Box (SFTP) one multi-path restic snapshot per app per run (§6) staged scratch on a data drive → missing-only merge to live (§7)
Manual .fab v0.130.0 full-root capture locked-in (not deselectable) checkbox, pre-selected opt-in, behind the two-number size warning + FileBrowser pointer download / chosen drive tar, exclusion-scoped root (SQ5 verdict; manifest v1 unchanged) existing import (old controllers import new bundles correctly)
PBS whole-guest rootfs + docker volumes; bind-mounted drives out of reach unchanged PBS on DooPlex vzdump whole-guest restore

Row-level decisions folded in:

  1. Classified apps' tier-2 appdata leg becomes fully class-driven — per-bind paths, not the whole appdata dir. Consequence: paperless-ngx's tier-2 copy shrinks (appdata/paperless/media mandatory in, appdata/paperless/export excluded out) — correct per the model, and the N>1 refusal disappears for classified apps because per-bind capture never needed the single-dir assumption. The F-S2 resolver remains only for legacy apps.
  2. The SSD fallback is a state-only tier. Its existing Hungarian reason already says so ("csak az adatbázis/konfiguráció fér a belső SSD-re"); the engine now enforces it: unit + mandatory if tier2FitsHeadroom passes for that reduced set; optional never goes to the SSD.
  3. Offsite never carries optional. Optional is the customer's local-copy tier; offsite cost is shared-model quota. (Viktor ruling #1, restated for the matrix.)
  4. Undeployed apps (unit discovered on a drive, stack not deployed): offsite pushes the legacy unit-only shape + WARN. Rationale: mandatory-path resolution requires the live HDD_PATH; guessing it from a stripped app.yaml risks capturing a stale or foreign tree. No protection regression vs today.
  5. Declared-but-absent mandatory path (compose/classes declare it, dir missing on disk): the F-S2 pattern — skip that path, loud WARN, capture the rest. The stat-filter before the restic invocation is MANDATORY and is the ONLY detection point: restic 0.14.0 (the pinned production binary) does NOT error on a nonexistent source path — it skips with a stderr warning, exits 0, and silently writes a partial snapshot (SPIKE-restic-snapshot-shape-2026-07-14.md, SP-3.4). 3a must never rely on a nonzero restic exit to catch a missing mandatory path; the controller stat-filters to detect the absence itself and raises the WARN from its own check. (If the container's restic is ever bumped, SP-3.4 must be re-run — later releases changed this exit behavior.)

3. Capture-set computation (Task 3-core)

One pure function, leaf package appbackup, consumed by all three engines (as-built, v0.133.0):

ComputeCaptureSet(binds []ClassifiedBind, hasClassification bool,
                  tier CaptureTier, hddPath string) CaptureSet
  → CaptureSet{HasClassification bool; Paths []CapturePath; Skipped []SkippedPath}
  • Resolves each bind's (Root, RelPath) against hddPath (HDD) / hddPath+"/userdata" (USERDATA) into absolute paths (slash algebra — path.Join, never filepath); filters by the tier column of §2 (TierOffsite = mandatory only; TierSecondary = mandatory + optional; excluded is silently dropped). Each CapturePath carries {Abs, Root, RelPath, Class} (the relpath is what the tier-2 layout and restore need).
  • The former UnitOnly field is replaced by HasClassification (it mirrors the classifier's existing bool); engines derive unit-only as !HasClassification || len(Paths)==0. When hasClassification == false the function short-circuits to {HasClassification:false} — nothing resolved, Paths/Skipped nil — so the legacy path stays byte-identical and testable (the SQ5 cost-regression guard: an unmigrated app never resolves a bind into an automatic tier).
  • Skipped carries would-be captures the tier filter selected but that are structurally unsafe, each with an English Reason, so the pure function stays log-free (engines log a skipped mandatory as a loud capture GAP). Three structural guards, evaluated after the tier filter: traversal (a .. path segment or an absolute RelPath — load-bearing because the compose parser path.Cleans but does NOT reject .., and ValidateBackupSpec only vets spec entries, so an unlisted writable ${HDD_PATH}/../x bind arrives classed mandatory), bare HDD drive-root (RelPath "" — would nest <hddPath>/backups into the capture), and the reserved backups/ zone; a bare userdata root (<hddPath>/userdata) is allowed (it does not nest backups/).
  • After filtering+guards: an equal-Abs collision from two binds collapses to one entry with mandatory winning over optional (mandatory semantics never degrade); then containment dedup drops any path whose ancestor is already in the set (keep the ancestor); output is sorted by Abs (deterministic despite map/slice input).
  • CrossAppOverlaps(map[app]CaptureSet) []Overlap is shipped here as a pure function (exact-Abs match across ≥2 apps' Paths; cross-app containment is legitimate per §4 and is NOT an overlap). Its WARN wiring lands in 3a/3b, not in 3-core — no log call sites exist yet.
  • Pure: no FS, no docker, no logging (the Tasks 1/2 leaf-package discipline). The stat-filter (decision 2.5) lives in the engines, not here.
  • Companion red-proofs: (a) sonarr with no block must produce unit-only (HasClassification:false, empty Paths) for offsite (the SQ5 cost-regression guard — the test that fails if the two-level default is miswired); (b) immich's explicit-optional :ro external library must appear in tier-2's set and NOT in offsite's; (c) paperless's export must appear in no automatic tier.

The engines consume the existing wired seam GetStackClassifiedBinds (main.go ~L1106) — no new seam. Per the F-S3 lesson, 3-core still ships an end-to-end wiring test through a real Manager.


4. Cross-app dedup and writing authority

Capture attribution rule (refined). SQ2's "the writing app is the classification authority" covers apps that write a path — but immich's external library (media/photos:ro, explicit optional) is populated by the customer via FileBrowser, not by any app. The generalized rule: capture attribution = the app whose classified capture set contains the path. Writing authority remains the catalog rule for who gets to class a path non-excluded.

Dedup decision: the capture engine dedups nothing across apps. Ground truth from the SQ2 inventory: the overlap set for non-excluded classes is empty today — radarr+sonarr's shared downloads/ is excluded on both; the streamers' whole-media binds are RO-excluded readers; calibre-web's mandatory media/books merely sits inside those excluded reader binds (containment, not conflict). Engineering cross-app dedup for an empty case buys risk, not value. Instead:

  1. Catalog convention (normative): at most one app may class a given host path non-excluded. Two writers to one path is a catalog bug.
  2. Advisory guard (3-core, log-only): a cheap cross-app pass at capture time WARNs if two deployed apps' sets contain the same absolute path — the convention's tripwire, no behavior.
  3. Cost note if the convention is ever violated: offsite double-capture is nearly free in real repo bytes (restic content-defined chunking dedups across snapshots); tier-2 pays disk twice. The WARN makes it visible before it costs.

Restore side. Per-app restore restores that app's captured paths to their recorded relpaths — attribution is positional, so the containment case is trivially correct: media/books restores under calibre-web's set whether or not plex exists, and restoring plex restores nothing of media/ (it captured nothing). A path captured by exactly one app is restored by exactly that app. The missing-only merge (§7) makes overlapping restores idempotent even in the violation case.


5. Ownership/permission fidelity per mechanism, and the NAS-as-target stance

The motivating failure is F-6C-1: rsync -a (-o -g) → chown on a root_squash NFS export → exit 23 → the whole tier-2 run fails. Fidelity matters because userdata carries the setgid convention (SharedContentGID=1000, mode 2775 — appbackup/userdata.go:19-24) and appdata stores carry app-specific uids; a wrong-owner restore into a uid-sensitive app is a silently broken restore.

Mechanism uid/gid mode incl. setgid xattrs Fails on root_squash target? Verdict
rsync -a --delete (tier-2) preserved (-og) preserved (-p) no (-X not set — acceptable, nothing in the tree needs xattrs) yes (the F-6C-1 chown) keep for local drives; never point it at network fs
restic backup/restore (offsite) stored in metadata; restored when run as root preserved yes no (metadata lives in the repo, not on the target fs) correct for the enlarged offsite shape as-is
tar (.fab) preserved in archive; --same-owner applies on root extract preserved GNU tar: with flags no correct; Task 4 verifies the extract flags at spec time

NAS-as-tier-2-target stance (ruling #4, recorded): excluded for now — auto AND pinned — with an honest Hungarian card reason. Mechanism: filter sp.IsNetwork() in selectTier2Target (both the pinned-validity check and the auto loop) — the metadata already exists (settings.go:219); no fs-probing. Dropping -o/-g was rejected because it converts a loud failure into a silently wrong-owner restore. Revisit condition: a tar-based tier-2 leg (ownership inside the archive) would make NAS viable; that is a separate future task with its own restore round-trip proof, not a Task 3 rider.


6. Offsite snapshot shape

Decision: one multi-path snapshot per app per run. restic backup <unitdir> <mandatory-abs-1> … --tag felhom-offbox --tag <stack>.

Why this shape over two-snapshots-per-app (unit vs userdata split):

  • Preserves the invariant every consumer relies on today: latest --tag <stack> = one consistent app-level point (offbox.go:732 restore, SnapshotCount, the hub report).
  • restic chunk-dedup makes unchanged media across daily runs cheap; the incremental cost of a run is the delta, not the library.
  • The snapshot's own paths metadata (visible in snapshots --json) records exactly what was captured — the restore side introspects shape from there, so no recovery-unit manifest change and no SchemaVersion bump (spike rec #6 considered; answered: the unit's content and shape are unchanged, classes already ride in the captured .felhom.yml, and snapshot-paths metadata is a better shape record than a manifest field that can drift from it).

Consequences that must be proven before the 3a spec (spike-first gate — these are unvalidated restic mechanisms):

  • SP-1 — quota stats semantics. Measure restic stats default (restore-size, all snapshots) vs --mode raw-data on a repo holding N retained snapshots of a large path. Expected: raw-data ≈ actual Storage Box disk; restore-size multiplies. Decides §9's accounting switch.
  • SP-2 — retention grouping across a shape change. Default forget grouping is host+paths; the first enlarged push changes the path set → old-shape snapshots strand in their own group and its keep-daily/weekly/monthly floor retains them forever (no new snapshots ever age them out) — zombie quota cost. Candidate fix: --group-by host,tags (the <stack> tag groups old and new shape together so old shape ages out naturally). Prove: seed old-shape snapshots, push new-shape, run forget with the flag, verify old-shape aging.
  • SP-3 — multi-path restore shape. Verify the --target tree layout for a multi-path snapshot (absolute source paths reconstructed under target), and that --include <unitpath> gives a unit-only selective restore. Feeds §7.

All three fit one nested-VM spike session against a scratch restic repo (no Hetzner needed — a local SFTP or even a local repo reproduces the semantics).


7. Offsite restore path (3a first-class scope)

Today's restore is dump-to-scratch-and-inspect. With userdata in snapshots it needs three changes:

  1. Scratch relocation + headroom gate (fixes F-A1). The scratch target moves from DataDir/offbox-restore (rootfs) to a data-drive location (<nsRoot>/backups/offsite-restore/<app> on the app's drive, or the tier-2 target-pick logic reused to find room). Before restoring, read the snapshot's logical size (restic stats <snap> / snapshots --json) and refuse with an honest Hungarian reason if the target lacks headroom — the SizeUnknown-never-renders-as-fits guard philosophy applies.
  2. Selective restore. Default operator action restores the unit only (--include the unit path — SP-3 proves the flag shape); "full restore" (unit + userdata) is an explicit second action showing the size first. Keeps the daily case cheap, makes the big case deliberate.
  3. Restore-to-live = staged missing-only merge. For the SQ3 acceptance ("immich restorable end-to-end from offsite alone"): restore to scratch → place mandatory paths into their live locations with the tier2_restore.go merge pattern (rsync -a --ignore-existing, never-delete) → then the existing unit restore (RestoreApp) brings the app up. Data before app start — the DB references the paths, so the paths must exist when the app first scans. Full-overwrite restore stays out of scope (PBS + DR tier own catastrophic recovery).

8. Tier-2 destination layout rework + migration

New layout (v2), relpath-mirroring:

backups/secondary/<stack>/
  .felhom-tier2-layout        # marker: "2"
  recovery-unit/              # unchanged leg
  hdd/<relpath>/              # e.g. hdd/appdata/paperless/media/
  userdata/<relpath>/         # e.g. userdata/media/books/
  • Relpath mirroring represents N>1 appdata dirs and nested binds natively (the v0.131.0 refusal is lifted structurally, not by special-casing) and makes restore position-derivable: dest relpath → live path under the app's current HDD_PATH — hddPath-invariant like the classifier itself.
  • Legacy apps use the same layout: the resolver's appdata dir(s) map to hdd/appdata/<name>/. Capture SET stays byte-identical (the SQ5 promise governs footprint and cost, not the internal arrangement of a derived copy); one layout means one restore reader.
  • Migration = rebuild, not preserve. Tier-2 copies are fully derived from live data. On the first v2 run per app: marker absent → delete the old flat appdata/ (and the old recovery-unit/ stays — it's layout-identical) → mirror fresh into v2 → write the marker last ("floor field LAST" family: the marker asserts a complete layout, so it is written only after all legs succeed).
  • Reconcile step (answers the v0.131.0 deferred pruning question): after mirroring, dest subdirs under hdd//userdata/ not present in the current capture set are removed — a bind removed from the compose (or re-classed excluded) stops occupying the secondary drive within one run.
  • Restore reads v2 only. Marker absent → honest refusal: "futtass előbb egy másodlagos mentést" — acceptable because tier-2 restore is missing-files recovery, and the source of truth (live data) still exists in that scenario; pre-migration copies stay on disk untouched until the first v2 run replaces them.
  • Headroom math: unitSize becomes unit + Σ(capture-set path sizes) via the existing du -sb loop; the SSD fallback computes it over the reduced (unit+mandatory) set per §2.2.

9. Quota & cost model

  • Accounting switch (pending SP-1): quota compares restic stats --mode raw-data — actual deduped+compressed repo bytes, i.e. what the customer's Storage Box really fills — instead of today's all-snapshots restore-size. Without this, ~17 retained snapshots of a mandatory library would multiply the measured usage and permanently trip the ≥100% gate. The existing stale-but-safe last-known-value behavior stays.
  • Pre-push gate (ruling #2 mechanized). Before an app's first enlarged push (and whenever its mandatory-set size estimate grows materially), compute the mandatory-set size with the existing per-mount du loop (estimate.go reuse — zero new measurement plumbing, SQ4). If last-known repo size + estimate crosses the quota: the enlarged push is blocked and the customer is notified with the two numbers — but the unit-only push continues (blocking it too would regress existing protection; the gate governs the enlargement, not the tier). The notification is the "size surfacing before the first enlarged push" from the ruling.
  • The existing ≥80% warn / ≥100% refuse-new-backups run gate stays as the backstop, re-based on raw-data numbers.
  • Tier-2 cost: headroom per §8; the two-number estimate (state+coupled vs +bulk) reuses the classified du split for the .fab warning (Task 4) and the tier-2 config panel display.
  • Hetzner reality check: BX11 = 1 TB shared-model. Mandatory ≈ DB-coupled stores (immich uploads, nextcloud data, paperless documents) — bounded per SQ4's model, but immich uploads are the one class member that grows like a media library. The pre-push gate + notification is what makes that growth a conversation instead of a surprise.

9.1 PBS whole-guest tier — measured capacity state (R-82 Phase 0, 2026-07-26)

Recorded per the operator's 2026-07-26 ruling: the datastore will be grown later; R-82 proceeds meanwhile. These are measurements, not projections-of-record — re-measure before relying on them. Full method + evidence: audits/SPIKE-r82-phase0-2026-07-26.md.

Fact Value Source
felhom-offsite datastore total (ep0) 37.2 GB ep0 df via the hub usage op (scripts/felhom-tenantsync.sh v1.2.0)
Used 10.8 GB (28.9%) same, 2026-07-26 12:02 UTC
Hub alert thresholds warn 80% (29.8 GB) / crit 90% monitor.PBSDRBoxChecker
Snapshots present 1 — demo-felhom ct/9201/2026-07-18T18:31:06Z, 9.74 GB logical, verify ok. demo-hp namespace exists with zero PBS API
First-snapshot compression 1:1 (9.74 GB logical ≈ 10.8 GB on disk) source is already-compressed Docker layers
Weekly incremental size UNMEASURED — no guest has ever had a second PBS snapshot measure at the first Slice-D weekly cycle

Three facts that govern the cost model here — all differ from the restic tier above:

  1. No cross-customer dedup. Backups are client-side encrypted with a per-customer key (encryption-key in storage.cfg), so PBS derives chunk digests under that key and chunks never dedup between tenants. Every customer's snapshots cost their full independent size. There is no fleet-scale dividend — the opposite of the shared-model assumption in §9 above.
  2. The datastore is small relative to the fleet. At weekly keep-last retention the current three boxes project to ≈1521 GB (4057%); each additional customer costs ≈510 GB, so the 80% warn is reached at roughly the second additional customer. This is the open constraint the operator deferred, not a solved problem.
  3. PVE cannot see this datastore's fill. pvesm status felhom-pbs reports 0/0/0 KiB while active — PBS answers the status call with HTTP 200 and zeroed usage because the per-customer token holds DatastoreBackup on its namespace, not Datastore.Audit on the datastore root. Operators must read PBS fill from the hub's PBS-DR gauge, never from pvesm/the PVE UI. Cosmetic, not a fault: writes work (the 07-18 snapshot is owned by that exact token).

What the PBS tier does and does not carry (confirmed live against pct config 9201 and the agent's own uncovered_volumes): rootfs + /var/lib/docker (mp0, backup=1) + /mnt/sys_drive (mp1, backup=1) are in; /mnt/felhom-drives (mp8) and /etc/felhom-bootstrap (mp9) are bind mounts and are out. The §2 tier table's "bind-mounted drives out of reach" is correct.

Why weekly is sufficient (R-82 P0.1 verdict). The only state exposed by a 7-day-old snapshot is the non-SMB half of settings.jsonstorage_paths, app_backup toggles, notification prefs, password_hash, launcher_share_token — all recoverable (drives are physical and re-enrollable; the password has a hub claim-reset path). Critically, encryption.key and the offbox credentials are stable files unchanged since first boot, so a week-old copy is byte-identical. Everything referentially coupled to app state rides the DAILY tiers.

⚠️ The verdict is conditional on Tier-3 offsite being enabled and healthy. On a box without it (drill-r50 reports offsite: null), PBS-weekly is the ONLY DR tier and the 7-day window would then cover app data and definitions too. Such a box needs offsite enabled first, or a shorter PBS cadence.


10. Decisions requiring Viktor confirmation (one-line vetoes)

The locked rulings are not re-litigated; these are consequences the doc surfaced:

  1. Blocked-enlargement semantics (§9): when the quota gate blocks an app's enlarged push, its unit-only push CONTINUES (no protection regression). Confirm. <--- Confirmed
  2. .fab optional default = pre-selected (spike rec #2 wording; ruling #1 said "checkboxes" without a default). Confirm pre-selected. <-- Conrifmed
  3. Offsite restore default = unit-only; full (unit+userdata) is an explicit second action with the size shown first (§7.2). Confirm. <-- Conrifmed
  4. Undeployed apps stay legacy unit-only offsite + WARN (§2.4). Confirm. <-- Conrifmed
  5. Tier-2 v2 layout migration = delete-and-rebuild of the derived secondary copy (§8), old copies replaced on the first v2 run per app. Confirm. <-- Conrifmed
  6. Quota accounting switches to raw-data bytes (§9), pending SP-1. Confirm direction. <-- Conrifmed
  7. SSD fallback becomes explicitly state-only (unit + mandatory-if-fits, never optional) (§2.2). Confirm. <-- Conrifmed

11. Explicitly out of scope (noted, not decided here)

  • Cluster-mode / HA implications (Peti's two-node roadmap item): tier-2 target selection and drive discovery assume a single host's drive set; agent-follows-guest will revisit.
  • .fab exclusion-scoping implementation details (Task 4 owns them; the SQ5 verdict — root tar minus excluded subtrees, manifest v1 unchanged — stands and is assumed by §2's .fab row).
  • NAS as a tier-2 target via a tar-based leg (§5 revisit condition).
  • Hub-side surfaces: nothing here forces a hub field; if the 3a report shape adds one (e.g. enlarged-push-blocked state), the spec flags it explicitly.

12. Observations (documented, not acted on)

  • tier2.go:167 RunTier2 doc comment still says "recovery unit + userdata" (stale F-S1 family) — 3b corrects it as a side effect of making it true.
  • offboxRecordStats reuses one probe-timeout context (sctx) for both snapshots and stats (offbox.go:686-702) — with a large repo the second call inherits whatever budget remains; worth a per-call timeout when 3a touches the file. Not a bug today.
  • The offsite restore handler backgrounds a 30-minute context (offbox_handlers.go:230-233 per the F4 fix); enlarged restores may need a bigger ceiling — spec-time decision in 3a, sized by SP-3 timing evidence.

Appendix — verified symbols Task 3 builds on

Symbol File (landmark) Note
Manager.ClassifiedBinds stacks/metadata.go:285 the seam; validated via LoadMetadata
appbackup.ClassifyBinds / ValidateBackupSpec appbackup/classify.go:165/:97 two-level default; whole-block-reject
appbackup.AppDataDirNames appbackup/paths.go:90 legacy resolver (stays for no-block apps)
StackDataProvider.GetStackClassifiedBinds adapter main.go:~1106 wired end-to-end (Task 2)
Manager.RunTier2 / selectTier2Target / rsyncMirror tier2.go:170/:92/:438 engine to rework; -a --delete
settings.StoragePath.IsNetwork settings/settings.go:219 the F-6C-1 filter, metadata-only
Manager.RunOffboxBackup / runOffboxInternal / discoverOffboxUnit offbox.go:376/:559/:508 push engine; unit discovery
Manager.RestoreOffbox offbox.go:722 restore-to-scratch; §7 rework
offboxQuotaState / offboxRecordStats offbox.go:632/:685 quota gate; stats-mode switch
restoreTier2Files merge pattern tier2_restore.go:127-137 --ignore-existing never-delete merge
Exporter.EstimateExport / duBytes appexport/estimate.go:33/:126 the two-number split reuse
RecoveryManifest (SchemaVersion 1) recovery_unit.go:31,131 unchanged by this design

Do NOT reuse rsyncMirror for any restore leg (--delete — the named trap); restore merges use the --ignore-existing pattern only.