Per the operator's 2026-07-26 ruling: datastore growth deferred, R-82 proceeds. Records the measurements so the constraint is written down rather than carried in a session: 37.2 GB total / 10.8 GB used, no cross-customer dedup (per-tenant encryption keys), 80% warn reached at roughly the second additional customer, and the pvesm 0/0/0 reporting artifact that means operators must read fill from the hub gauge. Also records what the tier does and does not carry, and the conditional on the P0.1 weekly verdict (Tier-3 offsite must be healthy). Doc NOT marked ratified — that stays Viktor's review (R-83).
32 KiB
07 — Backup architecture: tiers × classes × targets
Status: DRAFT — awaiting Viktor's review (async one-line vetoes on the §10 list). Written 2026-07-14 per the architecture-doc-first gate (Viktor ruling #5). Every claim below was verified against live Gitea at commit felhom-controller
95f3180(v0.132.0), felhom.eudeacee11, catalog21e8df1, agentc040c18(v0.88.0), hub v0.54.0. Line numbers are landmarks — reconfirm before editing.Inputs:
SPIKE-backup-classification-2026-07-14.md(790ec84),CAMPAIGN-6C-2026-07-14.md(deacee11, F-6C-1), Viktor's locked rulings of 2026-07-14. This document DECIDES; it does not survey. Where a decision rests on an unproven restic mechanism, it is marked SPIKE and listed in §6.4 — per the spike-first gate, those spikes precede the Task 3a spec.
0. Verified baseline — what the code does today (source, not memory)
| Surface | Behavior | Evidence (v0.132.0) |
|---|---|---|
| Recovery unit (tier-1) | compose/ (docker-compose.yml, .felhom.yml, secret-stripped app.yaml) + enumerated db-dumps/*.sql + volume-dumps/*.tar + manifest.json (SchemaVersion 1). No userdata, no HDD appdata. |
recovery_unit.go:91-106,131 |
| Offsite (tier-3) | One restic backup <unitdir> --tag felhom-offbox --tag <stack> per toggled app; unit dir discovered across nsRoots, newest-manifest tiebreak. Retention forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune, default grouping = host+paths. |
offbox.go:504-552 (discover), :575 (backup), :590-597 (forget) |
| Offsite quota | Soft gate on RepoSizeBytes from restic stats --json with no mode flag → default restore-size over all snapshots; ≥100% refuses new backups (prune still runs), ≥80% warns. |
offbox.go:407-415,630-645,700-708 |
| Offsite restore | restic restore latest --tag <stack> --target <DataDir>/offbox-restore/<app> — scratch, non-destructive, whole snapshot. DataDir default = /opt/docker/felhom-controller/data — the guest rootfs. |
web/offbox_handlers.go:236-246, config/config.go:310 |
| Tier-2 | rsync mirror of unit → backups/secondary/<stack>/recovery-unit + resolver-driven appdata dir → .../appdata (flat; refuses N>1). rsync -a --delete. Target pick: pinned > other data drive > SSD (headroom-gated); zero fs-type awareness (F-6C-1). Size for headroom = dirSizeBytes(unit)+dirSizeBytes(appdata). |
tier2.go:170-249 (run), :92-155 (select), :437-445 (rsync), :198 (size) |
| Tier-2 restore | Missing-files-only merge into live appdata: rsync -a --ignore-existing; refuses N>1. |
tier2_restore.go:15-44,127-137 |
Manual .fab |
Config + DB dump + volume tars + ExportDataMounts = ${HDD_PATH} binds ∪ the whole userdata root as one tar (SQ6 over-capture). |
stacks/delete.go:575, appexport/export.go (v0.130.0) |
| Classification (INERT) | stacks.Manager.ClassifiedBinds(name) → ([]ClassifiedBind, hasClassification); whole-block-reject at LoadMetadata; two-level default (explicit > :ro-reader > writable-mandatory; no block → legacy). Wired seam: StackDataProvider.GetStackClassifiedBinds (adapter main.go ~L1106). 13 catalog apps carry blocks (21e8df1). |
metadata.go:285-294, appbackup/classify.go:97,165 |
| Network storage metadata | settings.StoragePath.Kind == "network" + IsNetwork(); Protocol nfs|smb. Drive lifecycle already refuses network paths. |
settings.go:181-219,1057 |
Two new findings from this verification pass (not in the spike or 6C):
- F-A1 (HIGH for 3a): the offsite restore scratch dir is a rootfs-filler once snapshots carry
userdata.
offbox-restore/<app>lives undercfg.Paths.DataDir, default/opt/docker/felhom-controller/dataon the ~8 GB guest rootfs. Today snapshots are unit-sized (MB–1 GB); with mandatory userdata (nextcloud's data dir, immich's upload library) arestore latestcan fill the rootfs and take the guest down. §7 decides the fix. - F-A2 (MEDIUM for 3a): quota accounting mode is wrong-by-design for the new shape.
restic statsdefault restore-size aggregates across all retained snapshots (~17 per app at full retention). Exact cross-snapshot counting semantics are undocumented enough to matter: if a 50 GB mandatory library counts once per snapshot, the gate trips at ~17× reality. §9 decides the candidate fix (--mode raw-data); SPIKE SP-1 proves it. - Minor:
tier2.go:167(RunTier2 doc comment) still claims "recovery unit + userdata" — the last surviving F-S1 stale comment (fourth stale-comment instance this arc). 3b makes it true.
1. The class model (normative recap)
A class belongs to a bind (host path), not to an app (SQ2 rule, locked):
- mandatory — referentially COUPLED to app state (DB rows reference the content; SQ3: restore without it = broken-not-empty). Never deselectable at any tier that carries userdata.
- optional — DECOUPLED but precious (hand-curated, not re-acquirable). Customer-selectable.
- excluded — DECOUPLED bulk/transient/re-derivable. Never captured automatically;
.fabopt-in only, behind the two-number warning. - legacy (origin, not a class) — app has no
backup:block → every tier behaves byte-identically to v0.132.0. The safety net fails toward today's cost, never toward a quota blow-up (SQ5).
Roots: ${HDD_PATH} (hdd: list; appdata stores) and ${USERDATA_PATH} (userdata: list;
browsable tree). RelPaths are ${VAR}-relative and hddPath-invariant (survive migration flips —
the Task 1 lesson). Resolution to absolute paths happens at capture time against the app's live
HDD_PATH (Model A: HDD_PATH == nsRoot).
2. The decision table: tier × class × target
| Tier | legacy (no block) | mandatory | optional | excluded | Allowed targets | Copy mechanism | Restore path |
|---|---|---|---|---|---|---|---|
| Tier-1 recovery unit | config + DB dumps + volume tars | unchanged — the unit never carries userdata; classes ride inside it via the captured .felhom.yml |
— | — | app's own drive (backups/primary/<app>) |
direct writes | RestoreApp (existing) |
| Tier-2 cross-drive | unit + resolver-appdata (byte-identical to v0.131.0) | unit + mandatory binds | + optional binds | never | real local drives only (network paths excluded per F-6C-1 ruling — both auto and pinned); SSD fallback = unit + mandatory iff headroom fits, optional skipped with honest reason | rsync -a --delete per capture-set path |
missing-only merge (--ignore-existing), per-path from the new layout (§8) |
| Tier-3 offsite restic | unit only (today's shape) | unit + mandatory binds — not deselectable | never | never | Hetzner Storage Box (SFTP) | one multi-path restic snapshot per app per run (§6) | staged scratch on a data drive → missing-only merge to live (§7) |
Manual .fab |
v0.130.0 full-root capture | locked-in (not deselectable) | checkbox, pre-selected | opt-in, behind the two-number size warning + FileBrowser pointer | download / chosen drive | tar, exclusion-scoped root (SQ5 verdict; manifest v1 unchanged) | existing import (old controllers import new bundles correctly) |
| PBS whole-guest | rootfs + docker volumes; bind-mounted drives out of reach | unchanged | PBS on DooPlex | vzdump | whole-guest restore |
Row-level decisions folded in:
- Classified apps' tier-2 appdata leg becomes fully class-driven — per-bind paths, not the
whole appdata dir. Consequence: paperless-ngx's tier-2 copy shrinks (
appdata/paperless/mediamandatory in,appdata/paperless/exportexcluded out) — correct per the model, and the N>1 refusal disappears for classified apps because per-bind capture never needed the single-dir assumption. The F-S2 resolver remains only for legacy apps. - The SSD fallback is a state-only tier. Its existing Hungarian reason already says so
("csak az adatbázis/konfiguráció fér a belső SSD-re"); the engine now enforces it: unit +
mandatory if
tier2FitsHeadroompasses for that reduced set; optional never goes to the SSD. - Offsite never carries optional. Optional is the customer's local-copy tier; offsite cost is shared-model quota. (Viktor ruling #1, restated for the matrix.)
- Undeployed apps (unit discovered on a drive, stack not deployed): offsite pushes the legacy unit-only shape + WARN. Rationale: mandatory-path resolution requires the live HDD_PATH; guessing it from a stripped app.yaml risks capturing a stale or foreign tree. No protection regression vs today.
- Declared-but-absent mandatory path (compose/classes declare it, dir missing on disk): the F-S2 pattern — skip that path, loud WARN, capture the rest. The stat-filter before the restic invocation is MANDATORY and is the ONLY detection point: restic 0.14.0 (the pinned production binary) does NOT error on a nonexistent source path — it skips with a stderr warning, exits 0, and silently writes a partial snapshot (SPIKE-restic-snapshot-shape-2026-07-14.md, SP-3.4). 3a must never rely on a nonzero restic exit to catch a missing mandatory path; the controller stat-filters to detect the absence itself and raises the WARN from its own check. (If the container's restic is ever bumped, SP-3.4 must be re-run — later releases changed this exit behavior.)
3. Capture-set computation (Task 3-core)
One pure function, leaf package appbackup, consumed by all three engines (as-built, v0.133.0):
ComputeCaptureSet(binds []ClassifiedBind, hasClassification bool,
tier CaptureTier, hddPath string) CaptureSet
→ CaptureSet{HasClassification bool; Paths []CapturePath; Skipped []SkippedPath}
- Resolves each bind's
(Root, RelPath)againsthddPath(HDD) /hddPath+"/userdata"(USERDATA) into absolute paths (slash algebra —path.Join, neverfilepath); filters by the tier column of §2 (TierOffsite= mandatory only;TierSecondary= mandatory + optional; excluded is silently dropped). EachCapturePathcarries{Abs, Root, RelPath, Class}(the relpath is what the tier-2 layout and restore need). - The former
UnitOnlyfield is replaced byHasClassification(it mirrors the classifier's existing bool); engines derive unit-only as!HasClassification || len(Paths)==0. WhenhasClassification == falsethe function short-circuits to{HasClassification:false}— nothing resolved,Paths/Skippednil — so the legacy path stays byte-identical and testable (the SQ5 cost-regression guard: an unmigrated app never resolves a bind into an automatic tier). Skippedcarries would-be captures the tier filter selected but that are structurally unsafe, each with an EnglishReason, so the pure function stays log-free (engines log a skipped mandatory as a loud capture GAP). Three structural guards, evaluated after the tier filter: traversal (a..path segment or an absolute RelPath — load-bearing because the compose parser path.Cleans but does NOT reject.., andValidateBackupSpeconly vets spec entries, so an unlisted writable${HDD_PATH}/../xbind arrives classed mandatory), bare HDD drive-root (RelPath""— would nest<hddPath>/backupsinto the capture), and the reservedbackups/zone; a bare userdata root (<hddPath>/userdata) is allowed (it does not nestbackups/).- After filtering+guards: an equal-
Abscollision from two binds collapses to one entry with mandatory winning over optional (mandatory semantics never degrade); then containment dedup drops any path whose ancestor is already in the set (keep the ancestor); output is sorted byAbs(deterministic despite map/slice input). CrossAppOverlaps(map[app]CaptureSet) []Overlapis shipped here as a pure function (exact-Abs match across ≥2 apps'Paths; cross-app containment is legitimate per §4 and is NOT an overlap). Its WARN wiring lands in 3a/3b, not in 3-core — no log call sites exist yet.- Pure: no FS, no docker, no logging (the Tasks 1/2 leaf-package discipline). The stat-filter (decision 2.5) lives in the engines, not here.
- Companion red-proofs: (a) sonarr with no block must produce unit-only (
HasClassification:false, emptyPaths) for offsite (the SQ5 cost-regression guard — the test that fails if the two-level default is miswired); (b) immich's explicit-optional:roexternal library must appear in tier-2's set and NOT in offsite's; (c) paperless'sexportmust appear in no automatic tier.
The engines consume the existing wired seam GetStackClassifiedBinds (main.go ~L1106) — no new
seam. Per the F-S3 lesson, 3-core still ships an end-to-end wiring test through a real Manager.
4. Cross-app dedup and writing authority
Capture attribution rule (refined). SQ2's "the writing app is the classification authority"
covers apps that write a path — but immich's external library (media/photos:ro, explicit
optional) is populated by the customer via FileBrowser, not by any app. The generalized rule:
capture attribution = the app whose classified capture set contains the path. Writing authority
remains the catalog rule for who gets to class a path non-excluded.
Dedup decision: the capture engine dedups nothing across apps. Ground truth from the SQ2
inventory: the overlap set for non-excluded classes is empty today — radarr+sonarr's shared
downloads/ is excluded on both; the streamers' whole-media binds are RO-excluded readers;
calibre-web's mandatory media/books merely sits inside those excluded reader binds
(containment, not conflict). Engineering cross-app dedup for an empty case buys risk, not value.
Instead:
- Catalog convention (normative): at most one app may class a given host path non-excluded. Two writers to one path is a catalog bug.
- Advisory guard (3-core, log-only): a cheap cross-app pass at capture time WARNs if two deployed apps' sets contain the same absolute path — the convention's tripwire, no behavior.
- Cost note if the convention is ever violated: offsite double-capture is nearly free in real repo bytes (restic content-defined chunking dedups across snapshots); tier-2 pays disk twice. The WARN makes it visible before it costs.
Restore side. Per-app restore restores that app's captured paths to their recorded relpaths —
attribution is positional, so the containment case is trivially correct: media/books restores
under calibre-web's set whether or not plex exists, and restoring plex restores nothing of
media/ (it captured nothing). A path captured by exactly one app is restored by exactly that app.
The missing-only merge (§7) makes overlapping restores idempotent even in the violation case.
5. Ownership/permission fidelity per mechanism, and the NAS-as-target stance
The motivating failure is F-6C-1: rsync -a (-o -g) → chown on a root_squash NFS export →
exit 23 → the whole tier-2 run fails. Fidelity matters because userdata carries the setgid
convention (SharedContentGID=1000, mode 2775 — appbackup/userdata.go:19-24) and appdata stores
carry app-specific uids; a wrong-owner restore into a uid-sensitive app is a silently broken
restore.
| Mechanism | uid/gid | mode incl. setgid | xattrs | Fails on root_squash target? | Verdict |
|---|---|---|---|---|---|
rsync -a --delete (tier-2) |
preserved (-og) |
preserved (-p) |
no (-X not set — acceptable, nothing in the tree needs xattrs) |
yes (the F-6C-1 chown) | keep for local drives; never point it at network fs |
| restic backup/restore (offsite) | stored in metadata; restored when run as root | preserved | yes | no (metadata lives in the repo, not on the target fs) | correct for the enlarged offsite shape as-is |
tar (.fab) |
preserved in archive; --same-owner applies on root extract |
preserved | GNU tar: with flags | no | correct; Task 4 verifies the extract flags at spec time |
NAS-as-tier-2-target stance (ruling #4, recorded): excluded for now — auto AND pinned — with an
honest Hungarian card reason. Mechanism: filter sp.IsNetwork() in selectTier2Target (both the
pinned-validity check and the auto loop) — the metadata already exists (settings.go:219); no
fs-probing. Dropping -o/-g was rejected because it converts a loud failure into a silently
wrong-owner restore. Revisit condition: a tar-based tier-2 leg (ownership inside the archive)
would make NAS viable; that is a separate future task with its own restore round-trip proof, not a
Task 3 rider.
6. Offsite snapshot shape
Decision: one multi-path snapshot per app per run.
restic backup <unitdir> <mandatory-abs-1> … --tag felhom-offbox --tag <stack>.
Why this shape over two-snapshots-per-app (unit vs userdata split):
- Preserves the invariant every consumer relies on today: latest
--tag <stack>= one consistent app-level point (offbox.go:732restore, SnapshotCount, the hub report). - restic chunk-dedup makes unchanged media across daily runs cheap; the incremental cost of a run is the delta, not the library.
- The snapshot's own
pathsmetadata (visible insnapshots --json) records exactly what was captured — the restore side introspects shape from there, so no recovery-unit manifest change and no SchemaVersion bump (spike rec #6 considered; answered: the unit's content and shape are unchanged, classes already ride in the captured.felhom.yml, and snapshot-paths metadata is a better shape record than a manifest field that can drift from it).
Consequences that must be proven before the 3a spec (spike-first gate — these are unvalidated restic mechanisms):
- SP-1 — quota stats semantics. Measure
restic statsdefault (restore-size, all snapshots) vs--mode raw-dataon a repo holding N retained snapshots of a large path. Expected: raw-data ≈ actual Storage Box disk; restore-size multiplies. Decides §9's accounting switch. - SP-2 — retention grouping across a shape change. Default
forgetgrouping is host+paths; the first enlarged push changes the path set → old-shape snapshots strand in their own group and its keep-daily/weekly/monthly floor retains them forever (no new snapshots ever age them out) — zombie quota cost. Candidate fix:--group-by host,tags(the<stack>tag groups old and new shape together so old shape ages out naturally). Prove: seed old-shape snapshots, push new-shape, run forget with the flag, verify old-shape aging. - SP-3 — multi-path restore shape. Verify the
--targettree layout for a multi-path snapshot (absolute source paths reconstructed under target), and that--include <unitpath>gives a unit-only selective restore. Feeds §7.
All three fit one nested-VM spike session against a scratch restic repo (no Hetzner needed — a local SFTP or even a local repo reproduces the semantics).
7. Offsite restore path (3a first-class scope)
Today's restore is dump-to-scratch-and-inspect. With userdata in snapshots it needs three changes:
- Scratch relocation + headroom gate (fixes F-A1). The scratch target moves from
DataDir/offbox-restore(rootfs) to a data-drive location (<nsRoot>/backups/offsite-restore/<app>on the app's drive, or the tier-2 target-pick logic reused to find room). Before restoring, read the snapshot's logical size (resticstats <snap>/snapshots --json) and refuse with an honest Hungarian reason if the target lacks headroom — theSizeUnknown-never-renders-as-fits guard philosophy applies. - Selective restore. Default operator action restores the unit only (
--includethe unit path — SP-3 proves the flag shape); "full restore" (unit + userdata) is an explicit second action showing the size first. Keeps the daily case cheap, makes the big case deliberate. - Restore-to-live = staged missing-only merge. For the SQ3 acceptance ("immich restorable
end-to-end from offsite alone"): restore to scratch → place mandatory paths into their live
locations with the
tier2_restore.gomerge pattern (rsync -a --ignore-existing, never-delete) → then the existing unit restore (RestoreApp) brings the app up. Data before app start — the DB references the paths, so the paths must exist when the app first scans. Full-overwrite restore stays out of scope (PBS + DR tier own catastrophic recovery).
8. Tier-2 destination layout rework + migration
New layout (v2), relpath-mirroring:
backups/secondary/<stack>/
.felhom-tier2-layout # marker: "2"
recovery-unit/ # unchanged leg
hdd/<relpath>/ # e.g. hdd/appdata/paperless/media/
userdata/<relpath>/ # e.g. userdata/media/books/
- Relpath mirroring represents N>1 appdata dirs and nested binds natively (the v0.131.0 refusal is lifted structurally, not by special-casing) and makes restore position-derivable: dest relpath → live path under the app's current HDD_PATH — hddPath-invariant like the classifier itself.
- Legacy apps use the same layout: the resolver's appdata dir(s) map to
hdd/appdata/<name>/. Capture SET stays byte-identical (the SQ5 promise governs footprint and cost, not the internal arrangement of a derived copy); one layout means one restore reader. - Migration = rebuild, not preserve. Tier-2 copies are fully derived from live data. On the
first v2 run per app: marker absent → delete the old flat
appdata/(and the oldrecovery-unit/stays — it's layout-identical) → mirror fresh into v2 → write the marker last ("floor field LAST" family: the marker asserts a complete layout, so it is written only after all legs succeed). - Reconcile step (answers the v0.131.0 deferred pruning question): after mirroring, dest
subdirs under
hdd//userdata/not present in the current capture set are removed — a bind removed from the compose (or re-classed excluded) stops occupying the secondary drive within one run. - Restore reads v2 only. Marker absent → honest refusal: "futtass előbb egy másodlagos mentést" — acceptable because tier-2 restore is missing-files recovery, and the source of truth (live data) still exists in that scenario; pre-migration copies stay on disk untouched until the first v2 run replaces them.
- Headroom math:
unitSizebecomes unit + Σ(capture-set path sizes) via the existingdu -sbloop; the SSD fallback computes it over the reduced (unit+mandatory) set per §2.2.
9. Quota & cost model
- Accounting switch (pending SP-1): quota compares
restic stats --mode raw-data— actual deduped+compressed repo bytes, i.e. what the customer's Storage Box really fills — instead of today's all-snapshots restore-size. Without this, ~17 retained snapshots of a mandatory library would multiply the measured usage and permanently trip the ≥100% gate. The existing stale-but-safe last-known-value behavior stays. - Pre-push gate (ruling #2 mechanized). Before an app's first enlarged push (and whenever
its mandatory-set size estimate grows materially), compute the mandatory-set size with the
existing per-mount
duloop (estimate.goreuse — zero new measurement plumbing, SQ4). If last-known repo size + estimate crosses the quota: the enlarged push is blocked and the customer is notified with the two numbers — but the unit-only push continues (blocking it too would regress existing protection; the gate governs the enlargement, not the tier). The notification is the "size surfacing before the first enlarged push" from the ruling. - The existing ≥80% warn / ≥100% refuse-new-backups run gate stays as the backstop, re-based on raw-data numbers.
- Tier-2 cost: headroom per §8; the two-number estimate (state+coupled vs +bulk) reuses the
classified
dusplit for the.fabwarning (Task 4) and the tier-2 config panel display. - Hetzner reality check: BX11 = 1 TB shared-model. Mandatory ≈ DB-coupled stores (immich uploads, nextcloud data, paperless documents) — bounded per SQ4's model, but immich uploads are the one class member that grows like a media library. The pre-push gate + notification is what makes that growth a conversation instead of a surprise.
9.1 PBS whole-guest tier — measured capacity state (R-82 Phase 0, 2026-07-26)
Recorded per the operator's 2026-07-26 ruling: the datastore will be grown later; R-82 proceeds
meanwhile. These are measurements, not projections-of-record — re-measure before relying on them.
Full method + evidence: audits/SPIKE-r82-phase0-2026-07-26.md.
| Fact | Value | Source |
|---|---|---|
felhom-offsite datastore total (ep0) |
37.2 GB | ep0 df via the hub usage op (scripts/felhom-tenantsync.sh v1.2.0) |
| Used | 10.8 GB (28.9%) | same, 2026-07-26 12:02 UTC |
| Hub alert thresholds | warn 80% (29.8 GB) / crit 90% | monitor.PBSDRBoxChecker |
| Snapshots present | 1 — demo-felhom ct/9201/2026-07-18T18:31:06Z, 9.74 GB logical, verify ok. demo-hp namespace exists with zero |
PBS API |
| First-snapshot compression | ≈ 1:1 (9.74 GB logical ≈ 10.8 GB on disk) | source is already-compressed Docker layers |
| Weekly incremental size | UNMEASURED — no guest has ever had a second PBS snapshot | measure at the first Slice-D weekly cycle |
Three facts that govern the cost model here — all differ from the restic tier above:
- No cross-customer dedup. Backups are client-side encrypted with a per-customer key
(
encryption-keyinstorage.cfg), so PBS derives chunk digests under that key and chunks never dedup between tenants. Every customer's snapshots cost their full independent size. There is no fleet-scale dividend — the opposite of the shared-model assumption in §9 above. - The datastore is small relative to the fleet. At weekly
keep-lastretention the current three boxes project to ≈15–21 GB (40–57%); each additional customer costs ≈5–10 GB, so the 80% warn is reached at roughly the second additional customer. This is the open constraint the operator deferred, not a solved problem. - PVE cannot see this datastore's fill.
pvesm status felhom-pbsreports 0/0/0 KiB while active — PBS answers the status call with HTTP 200 and zeroed usage because the per-customer token holdsDatastoreBackupon its namespace, notDatastore.Auditon the datastore root. Operators must read PBS fill from the hub's PBS-DR gauge, never frompvesm/the PVE UI. Cosmetic, not a fault: writes work (the 07-18 snapshot is owned by that exact token).
What the PBS tier does and does not carry (confirmed live against pct config 9201 and the
agent's own uncovered_volumes): rootfs + /var/lib/docker (mp0, backup=1) + /mnt/sys_drive
(mp1, backup=1) are in; /mnt/felhom-drives (mp8) and /etc/felhom-bootstrap (mp9) are bind
mounts and are out. The §2 tier table's "bind-mounted drives out of reach" is correct.
Why weekly is sufficient (R-82 P0.1 verdict). The only state exposed by a 7-day-old snapshot is
the non-SMB half of settings.json — storage_paths, app_backup toggles, notification prefs,
password_hash, launcher_share_token — all recoverable (drives are physical and re-enrollable;
the password has a hub claim-reset path). Critically, encryption.key and the offbox credentials
are stable files unchanged since first boot, so a week-old copy is byte-identical. Everything
referentially coupled to app state rides the DAILY tiers.
⚠️ The verdict is conditional on Tier-3 offsite being enabled and healthy. On a box without it (drill-r50 reports
offsite: null), PBS-weekly is the ONLY DR tier and the 7-day window would then cover app data and definitions too. Such a box needs offsite enabled first, or a shorter PBS cadence.
10. Decisions requiring Viktor confirmation (one-line vetoes)
The locked rulings are not re-litigated; these are consequences the doc surfaced:
- Blocked-enlargement semantics (§9): when the quota gate blocks an app's enlarged push, its unit-only push CONTINUES (no protection regression). Confirm. <--- Confirmed
.faboptional default = pre-selected (spike rec #2 wording; ruling #1 said "checkboxes" without a default). Confirm pre-selected. <-- Conrifmed- Offsite restore default = unit-only; full (unit+userdata) is an explicit second action with the size shown first (§7.2). Confirm. <-- Conrifmed
- Undeployed apps stay legacy unit-only offsite + WARN (§2.4). Confirm. <-- Conrifmed
- Tier-2 v2 layout migration = delete-and-rebuild of the derived secondary copy (§8), old copies replaced on the first v2 run per app. Confirm. <-- Conrifmed
- Quota accounting switches to raw-data bytes (§9), pending SP-1. Confirm direction. <-- Conrifmed
- SSD fallback becomes explicitly state-only (unit + mandatory-if-fits, never optional) (§2.2). Confirm. <-- Conrifmed
11. Explicitly out of scope (noted, not decided here)
- Cluster-mode / HA implications (Peti's two-node roadmap item): tier-2 target selection and drive discovery assume a single host's drive set; agent-follows-guest will revisit.
.fabexclusion-scoping implementation details (Task 4 owns them; the SQ5 verdict — root tar minus excluded subtrees, manifest v1 unchanged — stands and is assumed by §2's.fabrow).- NAS as a tier-2 target via a tar-based leg (§5 revisit condition).
- Hub-side surfaces: nothing here forces a hub field; if the 3a report shape adds one (e.g. enlarged-push-blocked state), the spec flags it explicitly.
12. Observations (documented, not acted on)
tier2.go:167RunTier2 doc comment still says "recovery unit + userdata" (stale F-S1 family) — 3b corrects it as a side effect of making it true.offboxRecordStatsreuses one probe-timeout context (sctx) for bothsnapshotsandstats(offbox.go:686-702) — with a large repo the second call inherits whatever budget remains; worth a per-call timeout when 3a touches the file. Not a bug today.- The offsite restore handler backgrounds a 30-minute context (
offbox_handlers.go:230-233per the F4 fix); enlarged restores may need a bigger ceiling — spec-time decision in 3a, sized by SP-3 timing evidence.
Appendix — verified symbols Task 3 builds on
| Symbol | File (landmark) | Note |
|---|---|---|
Manager.ClassifiedBinds |
stacks/metadata.go:285 |
the seam; validated via LoadMetadata |
appbackup.ClassifyBinds / ValidateBackupSpec |
appbackup/classify.go:165/:97 |
two-level default; whole-block-reject |
appbackup.AppDataDirNames |
appbackup/paths.go:90 |
legacy resolver (stays for no-block apps) |
StackDataProvider.GetStackClassifiedBinds |
adapter main.go:~1106 |
wired end-to-end (Task 2) |
Manager.RunTier2 / selectTier2Target / rsyncMirror |
tier2.go:170/:92/:438 |
engine to rework; -a --delete |
settings.StoragePath.IsNetwork |
settings/settings.go:219 |
the F-6C-1 filter, metadata-only |
Manager.RunOffboxBackup / runOffboxInternal / discoverOffboxUnit |
offbox.go:376/:559/:508 |
push engine; unit discovery |
Manager.RestoreOffbox |
offbox.go:722 |
restore-to-scratch; §7 rework |
offboxQuotaState / offboxRecordStats |
offbox.go:632/:685 |
quota gate; stats-mode switch |
restoreTier2Files merge pattern |
tier2_restore.go:127-137 |
--ignore-existing never-delete merge |
Exporter.EstimateExport / duBytes |
appexport/estimate.go:33/:126 |
the two-number split reuse |
RecoveryManifest (SchemaVersion 1) |
recovery_unit.go:31,131 |
unchanged by this design |
Do NOT reuse rsyncMirror for any restore leg (--delete — the named trap); restore merges
use the --ignore-existing pattern only.