Files
felhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md

406 lines
29 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPIKE — referential-coupling backup classification (2026-07-14)
**Class:** spike (findings only — no production code, no schema edits, no repo mutations beyond this doc).
**Versions at spike time (recorded live):** controller **0.130.0** on demo guest 9201 (healthy),
agent **0.88.0** on felhom-pve (active), catalog @ HEAD (55 templates). Offsite repo: Hetzner
Storage Box sub-account, 12 snapshots, last run 2026-07-14 02:15 UTC, status ok, all inspection
**read-only** (`--no-lock`; no prune/forget/write).
**Empirical core:** a scratch Immich (catalog-pinned images v2.5.5) deployed in an isolated
namespace on 9201, exercised through the DB-present/library-absent restore scenario, then fully
removed.
Product decision under test: classify each userdata bind by **referential coupling**
**mandatory** (COUPLED: DB/config references it; restoring state without it yields a *broken* app),
**optional** (DECOUPLED-precious: user-specific, not re-downloadable, app re-derives on restore →
empty-not-broken), **excluded** (DECOUPLED-bulk: large, replaceable, cross-app).
## Verdict up front
1. **Offsite carries NO userdata today — confirmed at source and live, no exception.** Every restic
snapshot path is `backups/primary/<app>` (unit only). The latent gap is real: **immich is
offbox-toggled on demo, yet its photo library is in none of its 5 snapshots** (§SQ1).
2. **The coupling model is VALIDATED by the Immich experiment** — DB-only restore is not merely
empty, it is broken at two escalating levels: the server **refuses to boot** (its own
folder-check), and once coaxed past that, the timeline still advertises every asset while **every
thumbnail/original fetch is HTTP 500**. No self-heal, no re-import path for managed uploads
(§SQ3).
3. **Two NEW capture-gap findings** beyond the task baselines: (F-S1) tier-2 does **not** copy bulk
userdata — the tier2.go header comment is stale; browsable HDD userdata has **zero automated
off-drive protection today**; (F-S2) tier-2's appdata leg keys on the **stack name**, so
paperless-ngx (`appdata/paperless`) is **silently skipped** — live-proven (§SQ1).
4. **Round-trip verdict: exclusion scoping** (tar the userdata root, skip excluded subtrees) is the
only shape that round-trips through the v1 `.fab` manifest without format changes; inclusion
scoping (per-bind tars) cannot carry nested paths in `HDDSubdirs` (CTRL-001 forbids separators)
(§SQ5).
5. **The migration-default premise as stated does NOT hold** and needs a two-level default —
detailed in §SQ5.
---
## SQ1 — Capture map per tier (source + live)
### Source-confirmed mechanics
| Tier | What it captures | Code evidence |
|---|---|---|
| Recovery unit (primary) | `compose/` (docker-compose.yml, .felhom.yml, secret-stripped app.yaml) + enumerated `db-dumps/*.sql` + `volume-dumps/*.tar` + manifest.json. **No userdata, no HDD appdata.** | `backup/recovery_unit.go:91-106` (capture list is exactly the two config files + built app.yaml; dumps only *enumerated* from dirs the dump flow wrote) |
| Offsite (restic) | The recovery unit dir **only**`discoverOffboxUnit` returns `backups/primary/<app>`; one `restic backup <unitdir>` per toggled app. | `backup/offbox.go:372` (`RunOffboxBackup`), `offbox.go:559-586` (`runOffboxInternal`: `src` = the unit path, tagged `felhom-offbox,<app>`) |
| Tier-2 (cross-drive rsync) | Unit → `backups/secondary/<stack>/recovery-unit` **plus** `appdata/<stack>``.../appdata`. **NOT `userdata/`** — the only two `rsyncMirror` call sites in the tree are these (`tier2.go:169,177`). | `backup/tier2.go:140-184`; `AppDataDir` = `<nsRoot>/appdata/<stack>` (`appbackup/paths.go:68`) |
| Manual `.fab` (v0.130.0) | Config + DB dump + volume tars + `ExportDataMounts` = `${HDD_PATH}` binds the **userdata ROOT** as one tar. | `stacks/delete.go:575` (`ExportDataMounts`), `appexport/export.go:594-638` |
| PBS whole-guest | Guest rootfs + docker volumes; bind-mounted drive content is out of reach (platform docs; not re-derived here). | — |
### F-S1 (NEW, HIGH for the redesign): tier-2 does not copy bulk userdata — the comment lies
`tier2.go:17` says *"recovery unit + bulk userdata … the ONLY off-drive protection that browsable
HDD userdata can get"*, and `cmd/controller/main.go:1360-1361` says the tier-2 path *"copies the
felhom-data namespace wholesale"*. **The code copies neither**: `RunTier2` rsyncs exactly
`RecoveryUnitPath` and `AppDataDir(nsRoot, stack)``UserdataDir` appears nowhere in the backup
package (grep: only the two comment lines). Consequence for the model: **the shared media tree has
no automated off-drive copy at any tier today**; its only protection is a manual `.fab` export.
The task baseline "Tier-2 DOES copy bulk userdata" traced to these stale comments, not to behavior.
Live proof (demo 9201): `backups/secondary/radarr/` and `backups/secondary/jellyfin/` on felhom-usb
contain **only** `recovery-unit/` — no media, while `nas-media/userdata/media/` holds 96 MB the
apps actively use.
### F-S2 (NEW): tier-2 appdata leg silently skips apps whose appdata dir ≠ stack name
`RunTier2` copies `appdata/<stackName>`; the catalog binds `${HDD_PATH}/appdata/paperless/...` for
stack **paperless-ngx** (slug `paperless-ngx`). `AppDataDir(nsRoot,"paperless-ngx")` doesn't exist →
the `os.Stat` guard (`tier2.go:176`) skips the copy **silently** (status still records ok).
Live contrast on demo: `secondary/nextcloud/` contains `appdata/` (name matches);
`secondary/paperless-ngx/` (felhom-flash) contains **only** `recovery-unit/` while
`appdata/paperless/` exists on the same source drive. Affected today: **paperless-ngx** (its
document originals+archive!). immich/nextcloud/romm happen to match. Any future app with a nested
or differently-named appdata bind inherits the gap. The per-bind classification metadata fixes this
class of bug structurally (capture what the metadata declares, not what a name-convention guesses).
### Live capture map (per catalog app; source-derived, spot-verified live where deployed)
Legend: unit = config+dumps; ✔=captured, ✘=not captured. "app-named HDD appdata" = `${HDD_PATH}/appdata/<x>` bind.
| app | needs_hdd | offsite (restic) | tier-2 (cross-drive) | manual .fab (v0.130.0) |
|---|---|---|---|---|
| audiobookshelf | true | unit ✔ / userdata ✘ | unit ✔ / userdata ✘ | unit ✔ + userdata ROOT ✔ |
| calibre-web | true | unit ✔ / userdata ✘ | unit ✔ / userdata ✘ | unit ✔ + userdata ROOT ✔ |
| emby, jellyfin, plex | true | unit ✔ / media tree ✘ | unit ✔ / media tree ✘ | unit ✔ + userdata ROOT ✔ (RO shared tree over-captured) |
| immich | true | unit ✔ / **photo library ✘** | unit ✔ + `appdata/immich` ✔ | unit ✔ + `appdata/immich` ✔ + userdata ROOT ✔ |
| komga, navidrome, radarr, sonarr, romm | true | unit ✔ / userdata ✘ | unit ✔ / userdata ✘ (romm: + `appdata/romm` ✔) | unit ✔ + userdata ROOT ✔ (romm: + appdata bind ✔) |
| nextcloud | true | unit ✔ / **data dir ✘** | unit ✔ + `appdata/nextcloud` ✔ | unit ✔ + `appdata/nextcloud` ✔ |
| paperless-ngx | true | unit ✔ / **documents ✘** | unit ✔ / **`appdata/paperless` ✘ (F-S2)** | unit ✔ + appdata binds ✔ + userdata ROOT ✔ |
| 42 volume-only apps | false | unit ✔ (volume dumps ride in the unit) | unit ✔ (F6: they flow through tier-2 too) | unit ✔ (volume tars) |
Live legs run: restic `snapshots` + `ls` (below); `find` on `backups/primary/radarr` (unit =
compose/ + volume-dumps/ + manifest.json exactly) and `backups/secondary/{radarr,jellyfin,nextcloud,paperless-ngx}`.
### The offsite "no userdata" verdict — live evidence
Offbox-toggled apps on demo: audiobookshelf, calibre-web, immich (settings.json `app_backup`).
`restic snapshots --no-lock`: 12 snapshots, **every** path is
`/mnt/felhom-drives/{felhom-usb,nas-media}/backups/primary/<app>`. Full `restic ls` of the newest
immich snapshot (`81927042`, 2026-07-14 02:15):
```
.../backups/primary/immich/compose/{.felhom.yml,app.yaml,docker-compose.yml}
.../backups/primary/immich/db-dumps/immich-postgres.sql
.../backups/primary/immich/manifest.json
```
That is the complete file list. **The photo library (`appdata/immich`, 129 MB on the same drive) is
absent** → an offsite-only Immich restore today reproduces exactly the broken state measured in SQ3.
This is the "looks backed up but isn't" gap the redesign exists to close: the customer toggled
immich offsite ON, the status page says ok, and the photos are not there.
---
## SQ2 — Per-bind classification inventory (all 55 catalog apps)
Catalog ground truth (full read of all 55 templates): **13 apps** have `${USERDATA_PATH}`/
`${HDD_PATH}` binds; **42 are named-volume-only** (classification moot — their state already rides
in the unit's volume dumps; the only host bind among them is homepage's RO docker.sock). No
`.felhom.yml` anywhere has a backup-classification field (grep across all 55: zero hits) —
greenfield confirmed. `needs_hdd` lives under `resources:`.
Proposed class per bind (reason = coupling argument):
| app | bind | class | reason |
|---|---|---|---|
| immich | `${HDD_PATH}/appdata/immich``/usr/src/app/upload` | **mandatory** | Managed upload library. DB-authoritative: asset rows + `mountChecks` state reference these paths; **SQ3 proves restore-without = broken, not empty** |
| immich | `${USERDATA_PATH}/media/photos``/external/photos:ro` | **optional** | External library: RO, registered post-deploy, re-scanned on demand → empty-not-broken if absent; but content is user-precious (not re-downloadable) → customer-selectable |
| paperless-ngx | `${HDD_PATH}/appdata/paperless/media` | **mandatory** | Document originals + archive; DB rows carry content hashes/paths → restore-without = broken documents (coupled by design; empirical leg deferred, see SQ3 note) |
| paperless-ngx | `${HDD_PATH}/appdata/paperless/export` | **excluded** | Transient export target, re-creatable |
| paperless-ngx | `${USERDATA_PATH}/import/paperless` | **excluded** | Consume inbox — transient by contract (paperless deletes after ingest) |
| nextcloud | `${HDD_PATH}/appdata/nextcloud``/var/www/html/data` | **mandatory** | THE user files + `oc_filecache`/shares reference them; not re-derivable content |
| calibre-web | `${USERDATA_PATH}/media/books``/calibre-library` | **mandatory** | The Calibre library dir contains `metadata.db` **inside** the tree — DB and book files are one referential unit; restore-without = broken library |
| calibre-web | `${USERDATA_PATH}/import/calibre` | **excluded** | Ingest inbox, transient |
| audiobookshelf | `${USERDATA_PATH}/media/audiobooks`, `media/podcasts` | **excluded** | ABS re-scans folders; DB keeps progress/metadata keyed by rescan-match; podcasts are re-downloadable by definition. (Judgment call on audiobooks — see shared-tree note) |
| komga | `${USERDATA_PATH}/media/comics``/data` | **optional** | Komga re-scans (empty-not-broken), but comic collections are typically hand-curated/irreplaceable → precious, customer-selectable |
| navidrome | `${USERDATA_PATH}/media/music:ro` | **excluded** | Classic re-rippable bulk; navidrome re-scans; shared tree |
| radarr | `${USERDATA_PATH}/media/movies` | **excluded** | Re-downloadable; radarr re-scans and re-acquires by design |
| radarr / sonarr | `${USERDATA_PATH}/downloads` (SAME host path, both apps) | **excluded** | Transient download queue; also cross-app shared — backing it up per-app would double-capture |
| sonarr | `${USERDATA_PATH}/media/tv` | **excluded** | As movies |
| emby / jellyfin / plex | `${USERDATA_PATH}/media` (whole tree) `:ro` | **excluded** | Pure READERS of the shared tree (watch-state/config in volumes). See consistency rule below |
| romm | `${USERDATA_PATH}/roms``/romm/library` | **optional** | Romm re-scans (empty-not-broken); curated ROM sets are often not re-acquirable → precious |
| romm | `${HDD_PATH}/appdata/romm/resources` | **excluded** | Scraped artwork/metadata cache, re-derivable from scrapers |
**Shared-media-tree consistency (the model survives, with one rule).** emby/jellyfin/plex mount the
**whole** `media/` tree RO while seven apps own subtrees of it, and radarr+sonarr bind the identical
`downloads/` path. A naive per-app class assignment contradicts itself (books are *mandatory* to
calibre-web but sit inside the streamers' *excluded* tree; a per-app tar of `downloads` would
double-capture). The rule that dissolves both: **a class belongs to the BIND (the host path), and
the app that owns/writes the path is the classification authority; read-only binds are never
backup-relevant for the mounting app** (the streamers back up nothing from `media/`; calibre-web's
`media/books` = mandatory regardless of who else reads it). Overlaps of the same host path across
apps must be deduplicated at capture time by path, not by app — this already matches
`ExportDataMounts`' containment-dedup philosophy.
**Flagged judgment calls** (not blockers, spec should confirm with the operator): audiobooks
(excluded vs optional), comics/roms (optional chosen — precious-decoupled). The mandatory class is
NOT in doubt anywhere — every mandatory above is DB-referenced content.
---
## SQ3 — Immich DB-only restore (the empirical core)
**Setup** (scratch, isolated): compose project `spikeimmich` in `/root/spike-immich/` on 9201's
rootfs — catalog-pinned images (`immich-server:v2.5.5`, `postgres:16-vectorchord0.3.0`,
`redis:7-alpine`), ML disabled, loopback port 127.0.0.1:12283 only, no traefik, library bind
`/root/spike-immich/library``/usr/src/app/upload`, postgres in `/root/spike-immich/pgdata`.
**No felhom namespace touched** (registered drives, real apps, controller state all untouched;
controller verified healthy before/after).
**Healthy baseline:** admin created + logged in via API; 3 generated PNGs (640×480/800×600,
0.81.2 MB) uploaded via `POST /api/assets` → 3 assets `status:created`; after job settle, all 3
thumbnails HTTP 200 (813 KB webp/jpeg) and originals HTTP 200. Library tree: `upload/` (3
originals), `thumbs/` (preview+thumbnail per asset), plus Immich's own `.immich` marker files
(timestamp content) in all six subdirs. Sizes: library 3.6 MB, pgdata 302 MB.
**The coupled-broken scenario** — stop server → `mv library library-hidden` → empty `library/`
→ start (DB fully intact = "restored state without userdata"):
- **Level 1 — refuses to boot.** Server exits 1 in a crash-loop. Its StorageService logs
`Verifying system mount folder checks, current state: {"mountChecks":{"thumbs":true,"upload":true,
"backups":true,"library":true,"profile":true,"encoded-video":true}}` — i.e. **the DB itself
remembers the folders existed** — then
`Failed to read (/usr/src/app/upload/encoded-video/.immich): ENOENT` and
`microservices worker exited with code 1` (pointer to
docs.immich.app/administration/system-integrity#folder-checks). The coupling is not incidental:
Immich *actively encodes* it (markers written on first run, presence recorded in the DB,
verified on every boot).
- **Level 2 — skeleton-recreated restore is a lying gallery.** Recreating the six subdirs + copying
the `.immich` markers back (what a well-meaning restore script would do) lets the server boot.
Then: `GET /api/timeline/buckets` still reports `[{"timeBucket":"2026-07-01","count":3}]` — the
UI advertises all 3 photos — while **all 3 thumbnail fetches and all 3 original fetches return
HTTP 500** (`ENOENT``GlobalExceptionFilter Unknown error` per request in the server log).
- **Worse-than-empty verdict: CONFIRMED.** A clean install boots quietly with an empty timeline. The
DB-only restore either doesn't boot at all, or boots into a gallery where every remembered photo
errors — "looks recoverable but isn't", plus error spam.
- **No self-heal / re-import:** no repair API at `/api/repair` (404); job queues idle-clean; the
admin "repair" surface in Immich identifies orphans but cannot resurrect managed uploads whose
bytes are gone (nothing exists to re-import *from* — the library WAS the source). Reverse
direction (files present, DB absent) is a different scenario and irrelevant to offsite restore.
- **Control leg:** restoring the hidden library and restarting → all 3 thumbnails + originals
HTTP 200 again. The files alone were the broken link → **mandatory=coupled validated**.
**Paperless second candidate: deferred** (time-boxed; Immich was decisive). Paperless's coupling is
design-documented (content-hash-indexed originals in `media/documents/originals` + archive
versions); recommend a 15-minute repeat of this protocol during the implementation spec's
acceptance run rather than blocking the model on it.
**Cleanup:** done and verified — see §Cleanup.
---
## SQ4 — Size split on live data
Measured on demo (du; drive namespaces on felhom-flash/felhom-usb/nas-media):
| quantity | example values |
|---|---|
| (a) app state = recovery unit (config + DB dump + volume dumps) | navidrome 548 K · jellyfin 1.0 M · radarr 4.0 M · paperless-ngx 71 M (usb) · immich 44 M · calibre-web 248 M · nextcloud 933 M |
| (b) coupled/precious data (HDD appdata) | immich `appdata/immich` 129 M · nextcloud `appdata/nextcloud` 37 M · paperless `appdata/paperless` 56 K (demo-tiny) |
| (c) bulk shared tree | nas-media `userdata/media` 96 M (movies 57 M, music 37 M, books 12 M) + `downloads` 8.1 M; felhom-flash `userdata/navidrome` 2.9 G (campaign canary .bins — debris, see Observations) |
Ratios even at demo scale: jellyfin state (1 M) vs the media tree it reads (96 M) ≈ **1:100**;
radarr state (4 M) vs movies (57 M) ≈ 1:14. Production libraries are 100 GB10 TB against
state+coupled in the 0.15 GB band → the two-number estimate ("state+coupled" vs "+bulk") separates
by 24 orders of magnitude and is exactly the number the opt-in warning needs. Counter-example that
justifies *measuring* rather than assuming: scratch Immich's postgres (302 MB) dwarfed its 3-photo
library (3.6 MB) — state is not always the small number; and nextcloud's unit (933 MB, volume dumps
carrying `nextcloud_html`+DB) exceeds its appdata (37 MB) on demo.
**Where the numbers come from (reuse seams):** tier-2 already sizes with `du -sb`
(`tier2.go dirSizeBytes`); the export estimate already does a **per-mount** du loop
(`appexport/estimate.go:49-57 duBytes`) + the v0.129.0 F-A container-view `volumeSizer` for named
volumes, with the `SizeUnknown`-never-renders-as-fits guard. A classification-driven two-number
estimate needs **zero new measurement plumbing** — split the existing per-mount loop by the bind's
class and sum into two buckets.
---
## SQ5 — Metadata schema + round-trip feasibility
### Round-trip mechanics as-built (code evidence)
- The `.fab` manifest keys HDD tars by **basename**: `HDDSubdirs []string`
(`appexport/manifest.go:27`); export tars each mount as `<basename>.tar` with a fatal
basename-collision guard (`export.go:611-619`).
- **CTRL-001 hard-blocks nested paths in the v1 field:** `UnmarshalManifest`
`validateManifestPaths``ValidateSegment` rejects any segment containing a path separator
(`validate.go:35-36`), and restore re-checks per subdir (`restore.go:604`). `hdd_subdirs:
["userdata/media/tv"]` fails the parse *by security design*.
- Import maps a basename → a resolved `${HDD_PATH}` mount (note: `resolveHDDMounts` matches only
literal `${HDD_PATH}` references — `restore.go:670` — never `${USERDATA_PATH}`), else **fallback**
`<HDD_PATH>/<basename>` (`restore.go:616-626`). The v0.130.0 userdata ROOT round-trips purely via
this fallback: basename `userdata``<HDD_PATH>/userdata` = `USERDATA_PATH`. Pinned by
`stacks/export_mounts_test.go:55-59`.
### Feasibility verdict: **exclusion scoping wins**
- **Inclusion scoping** (tar each mandatory bind separately) requires nested-path knowledge on the
import side → a **manifest v2** field (e.g. `hdd_entries: [{rel_path, tar}]`), a new multi-segment
validator (per-segment CTRL-001 checks, no `..`/absolute), import-side placement logic, and a
compatibility story for old-controller-imports-new-bundle (v1 readers would silently ignore the
new field → hollow restores — the exact C6B-F1 failure class resurrected). Feasible but expensive
and risk-shaped.
- **Exclusion scoping** (tar the userdata ROOT, **skip excluded-class subtrees**; per-class
variants: offsite/state-only captures nothing of userdata unless a mandatory bind exists) keeps
the manifest byte-identical in shape: still one `userdata` entry in `HDDSubdirs`, still restored
by the existing basename fallback, **zero import changes, old controllers import new bundles
correctly** (they just extract a root tar that happens to contain less). The only new code is an
exclude list on the export tar call — and mandatory-bind capture for the *other* tiers (offsite
unit enrichment) reuses the same per-bind path resolution `ParseComposeUserdataMounts` already
provides (`stacks/delete.go:513`). The v0.130.0 root-capture comment's dedup philosophy carries
over unchanged.
- Same-drive HDD-appdata binds (`appdata/immich` etc.) already round-trip per-bind today (their
basenames are direct `${HDD_PATH}` children) — classification of those needs no tar-shape change
at all.
### Proposed `.felhom.yml` addition
```yaml
backup:
userdata:
- path: "media/photos" # relative to USERDATA_PATH
class: optional
hdd: # optional block for ${HDD_PATH} binds (appdata/*)
- path: "appdata/immich" # relative to HDD_PATH
class: mandatory
```
Semantics: `class ∈ {mandatory, optional, excluded}`; a bind **listed in the compose but absent
from the block defaults to `mandatory`** (C6B-F1 lesson: back up rather than silently drop);
read-only binds (`:ro`) are exempt from that default (reader rule, §SQ2) and default `excluded`.
The `${HDD_PATH}/appdata/*` binds default `mandatory` even without a block — they are app-managed
stores (immich/nextcloud/paperless all mandatory; only romm's scraper cache warrants an explicit
`excluded`).
### The migration default — the task's premise needs one correction
The premise *"missing field → mandatory default means today's behavior (offsite=config-only) does
NOT auto-start shipping media"* is **internally contradictory as stated**: if a missing field
defaults each bind to mandatory AND the new offsite policy ships mandatory binds, then an
unmigrated sonarr ships its whole `media/tv` offsite on day one — a surprise size/cost regression
(demo quota is 50 GB; a modest TV library is 10100× that; the SLICE-4 soft-quota gate would refuse
runs and alert-spam). The reasoning **holds with a two-level default**:
- **No `backup:` block at all → legacy behavior per tier** (offsite = unit only, .fab = full root
capture as v0.130.0). No unmigrated app changes its offsite footprint or cost.
- **Block present → unlisted writable binds default `mandatory`** (within an app that has opted
into classification, the safe direction is capture).
Since all 13 hdd apps get explicit classes in the same catalog change that ships the feature, the
block-absent branch exists only as a safety net for stragglers/third-party templates — and it fails
toward *today's* behavior, not toward a quota blow-up. The 42 volume-only apps need no block.
---
## SQ6 — `.fab` over-capture (derived + quantified; live export leg blocked)
**Code-proven:** `ExportDataMounts` returns the userdata **ROOT** for any app with a
`${USERDATA_PATH}` bind (`delete.go:575-598`; pinned by `TestExportDataMounts_UserdataConvention`),
so a sonarr export tars every sibling's content in the shared namespace by construction — the
question is only *how much*.
**Quantified on the live demo tree (nas-media, the drive radarr/jellyfin/navidrome/calibre-web
share):** a radarr `.fab` today captures the whole `userdata/` root ≈ **105 MB** (media 96 MB +
downloads 8.1 MB + import/documents/roms), of which radarr's own binds (`media/movies` 57 MB +
`downloads` 8.1 MB) are ~65 MB → **~1.6× over-capture even on a small demo tree**, carrying
navidrome's music, calibre-web's books (with its `metadata.db`!), and jellyfin's whole view. Under
the proposed classes, radarr's default `.fab` = state only (both its binds are excluded); the
over-capture factor vs today is unbounded (grows with everyone else's library).
**Live byte-level leg (deploy sonarr+jellyfin one-drive, marker files, export, inspect tar): NOT
run.** Triggering a real export requires the dashboard session (CSRF-protected UI flow) and the
demo box's login is customer-owned (Viktor) — the same blocker as the open `.fab` full-circle
upload test. Protocol for when it runs: drop `SPIKE-marker-sonarr.txt` in `media/tv` and
`SPIKE-marker-jelly.txt` in `media/movies`, export sonarr, `tar tf` the bundle's
`data/hdd/userdata.tar` for both markers (expected today: both present). Suitable as the
before-picture acceptance test of the implementation spec.
---
## Recommendations for the implementation spec
1. **Schema:** the `backup:` block above (per-bind, path-relative, three classes; reader-rule for
`:ro`; two-level default). Ship explicit classes for all 13 hdd apps in the same catalog commit.
2. **Per-tier policy:**
- **Offsite (restic)** = unit + **mandatory** binds, not deselectable. Fixes the live-proven
immich gap. Size impact is bounded because mandatory ≈ DB-coupled stores, not media trees.
- **Tier-2** = unit + mandatory + **optional** (it's the browsable local copy; a 2nd HDD has the
room) — and fix F-S2 by driving the appdata leg from the declared binds instead of the
stack-name guess; fix the F-S1/stale comments either by implementing the claimed userdata copy
for mandatory+optional or by correcting the comments in the same change.
- **Manual `.fab`** = state + mandatory always; optional pre-selected; excluded opt-in behind the
two-number size warning ("state+coupled ≈ X MB / +bulk ≈ Y GB", reusing the estimate.go
per-mount du split). UI states plainly: bulk media's intended path is FileBrowser, coupled data
rides in the `.fab`.
3. **Round-trip scoping: exclusion** (root tar minus excluded subtrees) — no manifest change, no
import change, backward-compatible both directions (§SQ5). Do NOT introduce per-bind nested tars
under manifest v1.
4. **Capture the class at deploy/unit time** into the recovery-unit manifest (the unit already
carries `.felhom.yml` — the offsite runner can read the classes without new sync paths).
5. **Acceptance tests to carry over:** the SQ3 protocol (immich DB-only restore must become
restorable end-to-end from offsite alone), the SQ6 marker test (sonarr `.fab` must stop carrying
jellyfin's markers), a paperless repeat of SQ3, and an F-S2 red-proof (paperless-ngx tier-2 must
copy `appdata/paperless`).
6. **Migration:** catalog commit with explicit classes for the 13 → controller feature gates on
block presence → stragglers keep legacy behavior. No `MinAgent` coupling expected (all
controller-side), but the offsite unit-content change should bump the recovery-unit
`SchemaVersion` consideration in the spec.
## Cleanup checklist (all verified live)
- [x] Scratch Immich: `docker compose down -v` (containers + volumes + `spikeimmich_spikenet`
network removed; `docker ps -a --filter name=spike` → 0).
- [x] `/root/spike-immich/` (library, pgdata, photos, token), `/root/spike-immich-compose.yml`,
`/root/genpng.py` deleted.
- [x] Pulled images removed (`immich-server:v2.5.5`, `postgres:16-vectorchord0.3.0`);
`redis:7-alpine` left in place (cannot prove it wasn't pre-existing; no container uses it;
zero-risk residue, noted here for honesty).
- [x] No restic write/prune/forget — inspection used `--no-lock` reads only; repo status unchanged
(12 snapshots before and after).
- [x] No tier-2 target, drive namespace, controller setting, or deployed app touched; controller
0.130.0 healthy at spike end (`Up`, healthy).
- [x] No secret values recorded anywhere in this doc (scratch instance used throwaway credentials,
destroyed with the instance).
## Observations (out of scope, documented not acted on)
- **Userdata debris on demo drives:** felhom-usb `userdata/` contains non-skeleton per-app dirs
(`actualbudget`, `bookstack`, `mealie`, `navidrome`, `paperless-ngx`, `uptime-kuma`,
`vaultwarden`) and felhom-flash `userdata/navidrome/` holds 2.9 GB of campaign canary `.bin`
files — leftovers from earlier campaigns; harmless but they'd inflate any root-capture `.fab`
and the future size estimates. Candidate for a demo-hygiene pass.
- **audiobookshelf binds `media/podcasts`, which is not in `UserdataSkeleton()`** — the deploy belt
pre-creates it from the compose scan, so no bug, but skeleton and catalog have drifted.
- **immich's current recovery unit carries no volume-dumps** (db-dump only) — consistent with the
dump flow not having run for its named volumes since the last teardown; worth a look during the
spec's offsite enrichment work.
- calibre-web's tier-2 status records an error against destination `/mnt/felhom-drives/teszt_enroll`
(a stale pinned target from an old enroll test) — the pin-invalidation fallback may deserve a
glance.
- The task baseline described the catalog as "~13 apps"; it is **55** (13 hdd + 42 volume-only) —
the classification workload is smaller than "all apps" implies but the migration commit touches
13 files, not 55.