Files
felhom.eu/documentation/architecture/07-backup-architecture.md
T

383 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 07 — Backup architecture: tiers × classes × targets
> **Status: DRAFT — awaiting Viktor's review (async one-line vetoes on the §10 list).**
> Written 2026-07-14 per the architecture-doc-first gate (Viktor ruling #5). Every claim below was
> verified against live Gitea at commit **felhom-controller `95f3180` (v0.132.0)**, felhom.eu
> `deacee11`, catalog `21e8df1`, agent `c040c18` (v0.88.0), hub v0.54.0. Line numbers are
> landmarks — reconfirm before editing.
>
> Inputs: `SPIKE-backup-classification-2026-07-14.md` (`790ec84`), `CAMPAIGN-6C-2026-07-14.md`
> (`deacee11`, F-6C-1), Viktor's locked rulings of 2026-07-14. This document DECIDES; it does not
> survey. Where a decision rests on an unproven restic mechanism, it is marked **SPIKE** and listed
> in §6.4 — per the spike-first gate, those spikes precede the Task 3a spec.
---
## 0. Verified baseline — what the code does today (source, not memory)
| Surface | Behavior | Evidence (v0.132.0) |
|---|---|---|
| Recovery unit (tier-1) | `compose/` (docker-compose.yml, `.felhom.yml`, secret-stripped app.yaml) + enumerated `db-dumps/*.sql` + `volume-dumps/*.tar` + `manifest.json` (SchemaVersion **1**). No userdata, no HDD appdata. | `recovery_unit.go:91-106,131` |
| Offsite (tier-3) | One `restic backup <unitdir> --tag felhom-offbox --tag <stack>` per toggled app; unit dir discovered across nsRoots, newest-manifest tiebreak. Retention `forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune`, **default grouping = host+paths**. | `offbox.go:504-552` (discover), `:575` (backup), `:590-597` (forget) |
| Offsite quota | Soft gate on `RepoSizeBytes` from `restic stats --json` with **no mode flag → default restore-size over all snapshots**; ≥100% refuses new backups (prune still runs), ≥80% warns. | `offbox.go:407-415,630-645,700-708` |
| Offsite restore | `restic restore latest --tag <stack> --target <DataDir>/offbox-restore/<app>` — scratch, non-destructive, whole snapshot. **DataDir default = `/opt/docker/felhom-controller/data` — the guest rootfs.** | `web/offbox_handlers.go:236-246`, `config/config.go:310` |
| Tier-2 | rsync mirror of unit → `backups/secondary/<stack>/recovery-unit` + resolver-driven appdata dir → `.../appdata` (flat; **refuses N>1**). `rsync -a --delete`. Target pick: pinned > other data drive > SSD (headroom-gated); **zero fs-type awareness** (F-6C-1). Size for headroom = `dirSizeBytes(unit)+dirSizeBytes(appdata)`. | `tier2.go:170-249` (run), `:92-155` (select), `:437-445` (rsync), `:198` (size) |
| Tier-2 restore | Missing-files-only merge into live appdata: `rsync -a --ignore-existing`; refuses N>1. | `tier2_restore.go:15-44,127-137` |
| Manual `.fab` | Config + DB dump + volume tars + `ExportDataMounts` = `${HDD_PATH}` binds the **whole userdata root** as one tar (SQ6 over-capture). | `stacks/delete.go:575`, `appexport/export.go` (v0.130.0) |
| Classification (INERT) | `stacks.Manager.ClassifiedBinds(name)``([]ClassifiedBind, hasClassification)`; whole-block-reject at `LoadMetadata`; two-level default (explicit > `:ro`-reader > writable-mandatory; **no block → legacy**). Wired seam: `StackDataProvider.GetStackClassifiedBinds` (adapter main.go ~L1106). 13 catalog apps carry blocks (`21e8df1`). | `metadata.go:285-294`, `appbackup/classify.go:97,165` |
| Network storage metadata | `settings.StoragePath.Kind == "network"` + `IsNetwork()`; `Protocol` nfs\|smb. Drive lifecycle already refuses network paths. | `settings.go:181-219,1057` |
Two **new findings** from this verification pass (not in the spike or 6C):
- **F-A1 (HIGH for 3a): the offsite restore scratch dir is a rootfs-filler once snapshots carry
userdata.** `offbox-restore/<app>` lives under `cfg.Paths.DataDir`, default
`/opt/docker/felhom-controller/data` on the ~8 GB guest rootfs. Today snapshots are unit-sized
(MB1 GB); with mandatory userdata (nextcloud's data dir, immich's upload library) a
`restore latest` can fill the rootfs and take the guest down. §7 decides the fix.
- **F-A2 (MEDIUM for 3a): quota accounting mode is wrong-by-design for the new shape.**
`restic stats` default restore-size aggregates across **all retained snapshots** (~17 per app at
full retention). Exact cross-snapshot counting semantics are undocumented enough to matter:
if a 50 GB mandatory library counts once per snapshot, the gate trips at ~17× reality. §9 decides
the candidate fix (`--mode raw-data`); **SPIKE SP-1** proves it.
- Minor: `tier2.go:167` (RunTier2 doc comment) still claims "recovery unit + **userdata**" — the
last surviving F-S1 stale comment (fourth stale-comment instance this arc). 3b makes it true.
---
## 1. The class model (normative recap)
A class belongs to a **bind (host path)**, not to an app (SQ2 rule, locked):
- **mandatory** — referentially COUPLED to app state (DB rows reference the content; SQ3: restore
without it = broken-not-empty). Never deselectable at any tier that carries userdata.
- **optional** — DECOUPLED but precious (hand-curated, not re-acquirable). Customer-selectable.
- **excluded** — DECOUPLED bulk/transient/re-derivable. Never captured automatically; `.fab` opt-in
only, behind the two-number warning.
- **legacy** (origin, not a class) — app has no `backup:` block → every tier behaves byte-identically
to v0.132.0. The safety net fails toward *today's cost*, never toward a quota blow-up (SQ5).
Roots: `${HDD_PATH}` (`hdd:` list; appdata stores) and `${USERDATA_PATH}` (`userdata:` list;
browsable tree). RelPaths are `${VAR}`-relative and hddPath-invariant (survive migration flips —
the Task 1 lesson). Resolution to absolute paths happens at capture time against the app's live
`HDD_PATH` (Model A: HDD_PATH == nsRoot).
---
## 2. The decision table: tier × class × target
| Tier | legacy (no block) | mandatory | optional | excluded | Allowed targets | Copy mechanism | Restore path |
|---|---|---|---|---|---|---|---|
| **Tier-1** recovery unit | config + DB dumps + volume tars | *unchanged* — the unit never carries userdata; classes ride inside it via the captured `.felhom.yml` | — | — | app's own drive (`backups/primary/<app>`) | direct writes | `RestoreApp` (existing) |
| **Tier-2** cross-drive | unit + resolver-appdata (byte-identical to v0.131.0) | unit + mandatory binds | + optional binds | never | **real local drives only** (network paths excluded per F-6C-1 ruling — both auto and pinned); SSD fallback = unit + mandatory **iff** headroom fits, optional skipped with honest reason | `rsync -a --delete` per capture-set path | missing-only merge (`--ignore-existing`), per-path from the new layout (§8) |
| **Tier-3** offsite restic | unit only (today's shape) | unit + mandatory binds — **not deselectable** | never | never | Hetzner Storage Box (SFTP) | one multi-path restic snapshot per app per run (§6) | staged scratch on a data drive → missing-only merge to live (§7) |
| **Manual `.fab`** | v0.130.0 full-root capture | locked-in (not deselectable) | checkbox, **pre-selected** | opt-in, behind the two-number size warning + FileBrowser pointer | download / chosen drive | tar, **exclusion-scoped** root (SQ5 verdict; manifest v1 unchanged) | existing import (old controllers import new bundles correctly) |
| **PBS** whole-guest | rootfs + docker volumes; bind-mounted drives out of reach | *unchanged* | | | PBS on DooPlex | vzdump | whole-guest restore |
Row-level decisions folded in:
1. **Classified apps' tier-2 appdata leg becomes fully class-driven** — per-bind paths, not the
whole appdata dir. Consequence: paperless-ngx's tier-2 copy *shrinks* (`appdata/paperless/media`
mandatory in, `appdata/paperless/export` excluded out) — correct per the model, and the N>1
refusal disappears for classified apps because per-bind capture never needed the single-dir
assumption. The F-S2 resolver remains **only** for legacy apps.
2. **The SSD fallback is a state-only tier.** Its existing Hungarian reason already says so
("csak az adatbázis/konfiguráció fér a belső SSD-re"); the engine now enforces it: unit +
mandatory if `tier2FitsHeadroom` passes for that reduced set; optional never goes to the SSD.
3. **Offsite never carries optional.** Optional is the customer's local-copy tier; offsite cost is
shared-model quota. (Viktor ruling #1, restated for the matrix.)
4. **Undeployed apps** (unit discovered on a drive, stack not deployed): offsite pushes the
**legacy unit-only shape** + WARN. Rationale: mandatory-path resolution requires the live
HDD_PATH; guessing it from a stripped app.yaml risks capturing a stale or foreign tree. No
protection regression vs today.
5. **Declared-but-absent mandatory path** (compose/classes declare it, dir missing on disk): the
F-S2 pattern — skip that path, **loud WARN**, capture the rest. The stat-filter before the
restic invocation is MANDATORY and is the ONLY detection point: restic 0.14.0 (the pinned
production binary) does NOT error on a nonexistent source path — it skips with a stderr
warning, exits 0, and silently writes a partial snapshot
(SPIKE-restic-snapshot-shape-2026-07-14.md, SP-3.4). 3a must never rely on a nonzero restic
exit to catch a missing mandatory path; the controller stat-filters to detect the absence
itself and raises the WARN from its own check. (If the container's restic is ever bumped,
SP-3.4 must be re-run — later releases changed this exit behavior.)
---
## 3. Capture-set computation (Task 3-core)
One pure function, leaf package `appbackup`, consumed by all three engines (as-built, v0.133.0):
```
ComputeCaptureSet(binds []ClassifiedBind, hasClassification bool,
tier CaptureTier, hddPath string) CaptureSet
→ CaptureSet{HasClassification bool; Paths []CapturePath; Skipped []SkippedPath}
```
- Resolves each bind's `(Root, RelPath)` against `hddPath` (HDD) / `hddPath+"/userdata"`
(USERDATA) into absolute paths (slash algebra — `path.Join`, never `filepath`); filters by the
tier column of §2 (`TierOffsite` = mandatory only; `TierSecondary` = mandatory + optional;
excluded is silently dropped). Each `CapturePath` carries `{Abs, Root, RelPath, Class}` (the
relpath is what the tier-2 layout and restore need).
- The former `UnitOnly` field is replaced by **`HasClassification`** (it mirrors the classifier's
existing bool); engines derive unit-only as `!HasClassification || len(Paths)==0`. When
`hasClassification == false` the function short-circuits to `{HasClassification:false}` — nothing
resolved, `Paths`/`Skipped` nil — so the legacy path stays byte-identical and testable (the SQ5
cost-regression guard: an unmigrated app never resolves a bind into an automatic tier).
- **`Skipped`** carries would-be captures the tier filter selected but that are structurally unsafe,
each with an English `Reason`, so the pure function stays log-free (engines log a skipped
**mandatory** as a loud capture GAP). Three structural guards, evaluated after the tier filter:
traversal (a `..` path segment or an absolute RelPath — **load-bearing because the compose parser
path.Cleans but does NOT reject `..`, and `ValidateBackupSpec` only vets *spec* entries, so an
unlisted writable `${HDD_PATH}/../x` bind arrives classed mandatory**), bare HDD drive-root
(RelPath `""` — would nest `<hddPath>/backups` into the capture), and the reserved `backups/`
zone; a bare **userdata** root (`<hddPath>/userdata`) is allowed (it does not nest `backups/`).
- After filtering+guards: an equal-`Abs` collision from two binds collapses to one entry with
**mandatory winning** over optional (mandatory semantics never degrade); then **containment
dedup** drops any path whose ancestor is already in the set (keep the ancestor); output is sorted
by `Abs` (deterministic despite map/slice input).
- `CrossAppOverlaps(map[app]CaptureSet) []Overlap` is shipped here as a **pure** function (exact-Abs
match across ≥2 apps' `Paths`; cross-app *containment* is legitimate per §4 and is NOT an overlap).
Its WARN **wiring** lands in 3a/3b, not in 3-core — no log call sites exist yet.
- Pure: no FS, no docker, no logging (the Tasks 1/2 leaf-package discipline). The stat-filter
(decision 2.5) lives in the engines, not here.
- Companion red-proofs: (a) sonarr with no block must produce unit-only (`HasClassification:false`,
empty `Paths`) for offsite (the SQ5 cost-regression guard — the test that fails if the two-level
default is miswired); (b) immich's explicit-optional `:ro` external library must appear in
tier-2's set and NOT in offsite's; (c) paperless's `export` must appear in no automatic tier.
The engines consume the existing wired seam `GetStackClassifiedBinds` (main.go ~L1106) — no new
seam. Per the F-S3 lesson, 3-core still ships an end-to-end wiring test through a real Manager.
---
## 4. Cross-app dedup and writing authority
**Capture attribution rule (refined).** SQ2's "the writing app is the classification authority"
covers apps that write a path — but immich's external library (`media/photos:ro`, explicit
*optional*) is populated by the **customer via FileBrowser**, not by any app. The generalized rule:
**capture attribution = the app whose classified capture set contains the path.** Writing authority
remains the *catalog* rule for who gets to class a path non-excluded.
**Dedup decision: the capture engine dedups nothing across apps.** Ground truth from the SQ2
inventory: the overlap set for non-excluded classes is **empty today** — radarr+sonarr's shared
`downloads/` is excluded on both; the streamers' whole-`media` binds are RO-excluded readers;
calibre-web's mandatory `media/books` merely sits *inside* those excluded reader binds
(containment, not conflict). Engineering cross-app dedup for an empty case buys risk, not value.
Instead:
1. **Catalog convention (normative):** at most one app may class a given host path non-excluded.
Two writers to one path is a catalog bug.
2. **Advisory guard (3-core, log-only):** a cheap cross-app pass at capture time WARNs if two
deployed apps' sets contain the same absolute path — the convention's tripwire, no behavior.
3. Cost note if the convention is ever violated: offsite double-capture is nearly free in real repo
bytes (restic content-defined chunking dedups across snapshots); tier-2 pays disk twice. The WARN
makes it visible before it costs.
**Restore side.** Per-app restore restores *that app's* captured paths to their recorded relpaths —
attribution is positional, so the containment case is trivially correct: `media/books` restores
under calibre-web's set whether or not plex exists, and restoring plex restores nothing of
`media/` (it captured nothing). A path captured by exactly one app is restored by exactly that app.
The missing-only merge (§7) makes overlapping restores idempotent even in the violation case.
---
## 5. Ownership/permission fidelity per mechanism, and the NAS-as-target stance
The motivating failure is F-6C-1: `rsync -a` (`-o -g`) → chown on a root_squash NFS export →
exit 23 → the whole tier-2 run fails. Fidelity matters because userdata carries the setgid
convention (`SharedContentGID=1000`, mode 2775 — `appbackup/userdata.go:19-24`) and appdata stores
carry app-specific uids; a wrong-owner restore into a uid-sensitive app is a *silently broken*
restore.
| Mechanism | uid/gid | mode incl. setgid | xattrs | Fails on root_squash target? | Verdict |
|---|---|---|---|---|---|
| `rsync -a --delete` (tier-2) | preserved (`-og`) | preserved (`-p`) | no (`-X` not set — acceptable, nothing in the tree needs xattrs) | **yes** (the F-6C-1 chown) | keep for local drives; never point it at network fs |
| restic backup/restore (offsite) | stored in metadata; restored when run as root | preserved | yes | no (metadata lives in the repo, not on the target fs) | correct for the enlarged offsite shape as-is |
| tar (`.fab`) | preserved in archive; `--same-owner` applies on root extract | preserved | GNU tar: with flags | no | correct; Task 4 verifies the extract flags at spec time |
**NAS-as-tier-2-target stance (ruling #4, recorded):** excluded for now — auto AND pinned — with an
honest Hungarian card reason. Mechanism: filter `sp.IsNetwork()` in `selectTier2Target` (both the
pinned-validity check and the auto loop) — the metadata already exists (`settings.go:219`); no
fs-probing. Dropping `-o/-g` was rejected because it converts a loud failure into a silently
wrong-owner restore. **Revisit condition:** a tar-based tier-2 leg (ownership inside the archive)
would make NAS viable; that is a separate future task with its own restore round-trip proof, not a
Task 3 rider.
---
## 6. Offsite snapshot shape
**Decision: one multi-path snapshot per app per run.**
`restic backup <unitdir> <mandatory-abs-1> … --tag felhom-offbox --tag <stack>`.
Why this shape over two-snapshots-per-app (unit vs userdata split):
- Preserves the invariant every consumer relies on today: *latest `--tag <stack>` = one consistent
app-level point* (`offbox.go:732` restore, SnapshotCount, the hub report).
- restic chunk-dedup makes unchanged media across daily runs cheap; the incremental cost of a run
is the delta, not the library.
- The snapshot's own `paths` metadata (visible in `snapshots --json`) records exactly what was
captured — the restore side introspects shape from there, so **no recovery-unit manifest change
and no SchemaVersion bump** (spike rec #6 considered; answered: the unit's content and shape are
unchanged, classes already ride in the captured `.felhom.yml`, and snapshot-paths metadata is a
better shape record than a manifest field that can drift from it).
**Consequences that must be proven before the 3a spec (spike-first gate — these are unvalidated
restic mechanisms):**
- **SP-1 — quota stats semantics.** Measure `restic stats` default (restore-size, all snapshots)
vs `--mode raw-data` on a repo holding N retained snapshots of a large path. Expected: raw-data ≈
actual Storage Box disk; restore-size multiplies. Decides §9's accounting switch.
- **SP-2 — retention grouping across a shape change.** Default `forget` grouping is host+**paths**;
the first enlarged push changes the path set → old-shape snapshots strand in their own group and
its keep-daily/weekly/monthly floor retains them **forever** (no new snapshots ever age them
out) — zombie quota cost. Candidate fix: `--group-by host,tags` (the `<stack>` tag groups old and
new shape together so old shape ages out naturally). Prove: seed old-shape snapshots, push
new-shape, run forget with the flag, verify old-shape aging.
- **SP-3 — multi-path restore shape.** Verify the `--target` tree layout for a multi-path snapshot
(absolute source paths reconstructed under target), and that `--include <unitpath>` gives a
unit-only selective restore. Feeds §7.
All three fit one nested-VM spike session against a scratch restic repo (no Hetzner needed —
a local SFTP or even a local repo reproduces the semantics).
---
## 7. Offsite restore path (3a first-class scope)
Today's restore is dump-to-scratch-and-inspect. With userdata in snapshots it needs three changes:
1. **Scratch relocation + headroom gate (fixes F-A1).** The scratch target moves from
`DataDir/offbox-restore` (rootfs) to a data-drive location
(`<nsRoot>/backups/offsite-restore/<app>` on the app's drive, or the tier-2 target-pick logic
reused to find room). Before restoring, read the snapshot's logical size (restic
`stats <snap>` / `snapshots --json`) and refuse with an honest Hungarian reason if the target
lacks headroom — the `SizeUnknown`-never-renders-as-fits guard philosophy applies.
2. **Selective restore.** Default operator action restores the **unit only** (`--include` the unit
path — SP-3 proves the flag shape); "full restore" (unit + userdata) is an explicit second
action showing the size first. Keeps the daily case cheap, makes the big case deliberate.
3. **Restore-to-live = staged missing-only merge.** For the SQ3 acceptance ("immich restorable
end-to-end from offsite alone"): restore to scratch → place mandatory paths into their live
locations with the `tier2_restore.go` merge pattern (`rsync -a --ignore-existing`,
never-delete) → then the existing unit restore (`RestoreApp`) brings the app up. Data before
app start — the DB references the paths, so the paths must exist when the app first scans.
Full-overwrite restore stays out of scope (PBS + DR tier own catastrophic recovery).
---
## 8. Tier-2 destination layout rework + migration
**New layout (v2), relpath-mirroring:**
```
backups/secondary/<stack>/
.felhom-tier2-layout # marker: "2"
recovery-unit/ # unchanged leg
hdd/<relpath>/ # e.g. hdd/appdata/paperless/media/
userdata/<relpath>/ # e.g. userdata/media/books/
```
- Relpath mirroring represents N>1 appdata dirs and nested binds natively (the v0.131.0 refusal is
lifted structurally, not by special-casing) and makes restore position-derivable: dest relpath →
live path under the app's current HDD_PATH — hddPath-invariant like the classifier itself.
- **Legacy apps use the same layout**: the resolver's appdata dir(s) map to
`hdd/appdata/<name>/`. Capture SET stays byte-identical (the SQ5 promise governs footprint and
cost, not the internal arrangement of a derived copy); one layout means one restore reader.
- **Migration = rebuild, not preserve.** Tier-2 copies are fully derived from live data. On the
first v2 run per app: marker absent → delete the old flat `appdata/` (and the old
`recovery-unit/` stays — it's layout-identical) → mirror fresh into v2 → write the marker last
("floor field LAST" family: the marker asserts a complete layout, so it is written only after all
legs succeed).
- **Reconcile step (answers the v0.131.0 deferred pruning question):** after mirroring, dest
subdirs under `hdd/`/`userdata/` not present in the current capture set are removed — a bind
removed from the compose (or re-classed excluded) stops occupying the secondary drive within one
run.
- **Restore reads v2 only.** Marker absent → honest refusal: "futtass előbb egy másodlagos
mentést" — acceptable because tier-2 restore is missing-files recovery, and the source of truth
(live data) still exists in that scenario; pre-migration copies stay on disk untouched until the
first v2 run replaces them.
- Headroom math: `unitSize` becomes unit + Σ(capture-set path sizes) via the existing `du -sb`
loop; the SSD fallback computes it over the reduced (unit+mandatory) set per §2.2.
---
## 9. Quota & cost model
- **Accounting switch (pending SP-1): quota compares `restic stats --mode raw-data` — actual
deduped+compressed repo bytes, i.e. what the customer's Storage Box really fills** — instead of
today's all-snapshots restore-size. Without this, ~17 retained snapshots of a mandatory library
would multiply the measured usage and permanently trip the ≥100% gate. The existing
stale-but-safe last-known-value behavior stays.
- **Pre-push gate (ruling #2 mechanized).** Before an app's **first enlarged push** (and whenever
its mandatory-set size estimate grows materially), compute the mandatory-set size with the
existing per-mount `du` loop (`estimate.go` reuse — zero new measurement plumbing, SQ4). If
last-known repo size + estimate crosses the quota: **the enlarged push is blocked and the
customer is notified with the two numbers** — but the **unit-only push continues** (blocking it
too would *regress* existing protection; the gate governs the enlargement, not the tier).
The notification is the "size surfacing before the first enlarged push" from the ruling.
- The existing ≥80% warn / ≥100% refuse-new-backups run gate stays as the backstop, re-based on
raw-data numbers.
- **Tier-2 cost:** headroom per §8; the two-number estimate (state+coupled vs +bulk) reuses the
classified `du` split for the `.fab` warning (Task 4) and the tier-2 config panel display.
- Hetzner reality check: BX11 = 1 TB shared-model. Mandatory ≈ DB-coupled stores (immich uploads,
nextcloud data, paperless documents) — bounded per SQ4's model, but immich uploads are the one
class member that grows like a media library. The pre-push gate + notification is what makes
that growth a conversation instead of a surprise.
---
## 10. Decisions requiring Viktor confirmation (one-line vetoes)
The locked rulings are not re-litigated; these are **consequences the doc surfaced**:
1. **Blocked-enlargement semantics (§9):** when the quota gate blocks an app's enlarged push, its
unit-only push CONTINUES (no protection regression). Confirm. <--- Confirmed
2. **`.fab` optional default = pre-selected** (spike rec #2 wording; ruling #1 said "checkboxes"
without a default). Confirm pre-selected. <-- Conrifmed
3. **Offsite restore default = unit-only; full (unit+userdata) is an explicit second action with
the size shown first (§7.2).** Confirm. <-- Conrifmed
4. **Undeployed apps stay legacy unit-only offsite + WARN (§2.4).** Confirm. <-- Conrifmed
5. **Tier-2 v2 layout migration = delete-and-rebuild of the derived secondary copy (§8),** old
copies replaced on the first v2 run per app. Confirm. <-- Conrifmed
6. **Quota accounting switches to raw-data bytes (§9), pending SP-1.** Confirm direction. <-- Conrifmed
7. **SSD fallback becomes explicitly state-only (unit + mandatory-if-fits, never optional)
(§2.2).** Confirm. <-- Conrifmed
## 11. Explicitly out of scope (noted, not decided here)
- Cluster-mode / HA implications (Peti's two-node roadmap item): tier-2 target selection and
drive discovery assume a single host's drive set; agent-follows-guest will revisit.
- `.fab` exclusion-scoping implementation details (Task 4 owns them; the SQ5 verdict — root tar
minus excluded subtrees, manifest v1 unchanged — stands and is assumed by §2's `.fab` row).
- NAS as a tier-2 target via a tar-based leg (§5 revisit condition).
- Hub-side surfaces: nothing here forces a hub field; if the 3a report shape adds one (e.g.
enlarged-push-blocked state), the spec flags it explicitly.
## 12. Observations (documented, not acted on)
- `tier2.go:167` RunTier2 doc comment still says "recovery unit + userdata" (stale F-S1 family) —
3b corrects it as a side effect of making it true.
- `offboxRecordStats` reuses one probe-timeout context (`sctx`) for both `snapshots` and `stats`
(`offbox.go:686-702`) — with a large repo the second call inherits whatever budget remains; worth
a per-call timeout when 3a touches the file. Not a bug today.
- The offsite restore handler backgrounds a 30-minute context (`offbox_handlers.go:230-233` per the
F4 fix); enlarged restores may need a bigger ceiling — spec-time decision in 3a, sized by SP-3
timing evidence.
## Appendix — verified symbols Task 3 builds on
| Symbol | File (landmark) | Note |
|---|---|---|
| `Manager.ClassifiedBinds` | `stacks/metadata.go:285` | the seam; validated via `LoadMetadata` |
| `appbackup.ClassifyBinds` / `ValidateBackupSpec` | `appbackup/classify.go:165/:97` | two-level default; whole-block-reject |
| `appbackup.AppDataDirNames` | `appbackup/paths.go:90` | legacy resolver (stays for no-block apps) |
| `StackDataProvider.GetStackClassifiedBinds` | adapter `main.go:~1106` | wired end-to-end (Task 2) |
| `Manager.RunTier2` / `selectTier2Target` / `rsyncMirror` | `tier2.go:170/:92/:438` | engine to rework; `-a --delete` |
| `settings.StoragePath.IsNetwork` | `settings/settings.go:219` | the F-6C-1 filter, metadata-only |
| `Manager.RunOffboxBackup` / `runOffboxInternal` / `discoverOffboxUnit` | `offbox.go:376/:559/:508` | push engine; unit discovery |
| `Manager.RestoreOffbox` | `offbox.go:722` | restore-to-scratch; §7 rework |
| `offboxQuotaState` / `offboxRecordStats` | `offbox.go:632/:685` | quota gate; stats-mode switch |
| `restoreTier2Files` merge pattern | `tier2_restore.go:127-137` | `--ignore-existing` never-delete merge |
| `Exporter.EstimateExport` / `duBytes` | `appexport/estimate.go:33/:126` | the two-number split reuse |
| `RecoveryManifest` (SchemaVersion 1) | `recovery_unit.go:31,131` | unchanged by this design |
**Do NOT reuse** `rsyncMirror` for any restore leg (`--delete` — the named trap); restore merges
use the `--ignore-existing` pattern only.