# SPIKE — R-82 Phase 0: the three gates before the backup target split (2026-07-26) **Class:** read-only Phase-0 gate for R-82. **No code written. No backup triggered. No config changed.** Baselines: felhom-agent `dfd5d73` v0.96.0, felhom-controller `47fda06` v0.173.0, felhom.eu `c73800c` hub v0.75.0. All three trees clean at `origin/main`. --- ## Verdicts — one gate returns STOP | Gate | Verdict | |---|---| | **P0.1** — what is exposed for seven days | **weekly CONFIRMED**, with the exposed set named and one conditional | | **P0.2** — the `pvesm status` 0/0/0 anomaly | **RESOLVED — benign PVE-side reporting artifact.** Not a blocker | | **P0.3** — capacity headroom | **🛑 STOP.** The 80% alert is reached at roughly the **second additional customer**, inside the alpha horizon | **R-82 stops here pending an operator ruling on P0.3.** P0.1 and P0.2 clear; the blocker is capacity, not correctness. --- ## P0.1 — What is exposed for seven days? ### Method The whole-guest vzdump covers what PVE will actually snapshot. Live guest config (`pct config 9201`, demo-felhom): ``` rootfs: local-lvm:vm-9201-disk-0,size=32G → IN the snapshot mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/docker,backup=1,size=200G → IN mp1: local-lvm:vm-9201-disk-2,mp=/mnt/sys_drive,backup=1,size=50G → IN mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives → BIND — NOT in the snapshot mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1 → BIND — NOT in ``` Corroborated by the agent's own report: `"uncovered_volumes": ["/etc/felhom-bootstrap", "/mnt/felhom-drives"]`. **The daily tiers are alive** — this is what makes the answer come out the way it does, so it was verified rather than assumed: | Customer | Tier-3 offsite (restic, daily) | Evidence | |---|---|---| | demo-felhom | **running** | `last_run 2026-07-26T02:17:56Z`, `last_status "ok"`, 28 snapshots, 1.08 GB / 50 GB | | demo-hp | **running** | `last_run 2026-07-26T02:15:50Z`, `last_status "ok"`, 4 snapshots, 29 MB / 50 GB | | drill-r50 | **`offsite: null` — not configured** | see the conditional below | ### The exposed-state table | State | Where it lives | In the vzdump? | Covered by a DAILY tier? | 7-day staleness cost | |---|---|---|---|---| | **`encryption.key`** (app-secret key) | controller data dir | ✅ | ❌ none | **NONE — the key is STABLE.** `mtime 2026-07-18 16:32`, unchanged since creation. A 7-day-old copy is byte-identical. *This was the biggest theoretical risk and it is clean.* | | **`settings.json` — SMB shares** (`smb`, `smb_shares`) | controller data dir | ✅ | ✅ **yes** — R-7b payload in tier-2 **and** offsite (`_shares-manifest.json` + `passdb.tar`, staged 2026-07-26 02:16) | none | | **`settings.json` — everything else** | controller data dir | ✅ | ❌ **none** | **THE EXPOSED SET.** `mtime 2026-07-26 10:02` — actively changing. Detailed below | | `controller.yaml` | controller data dir | ✅ | ❌ | **NONE in practice** — the hub is the source of truth (`config_pull` / bootstrap re-fetch overwrites it; it explicitly never touches `settings.json`) | | offbox credentials (`repo_password`, `ssh_key`, `known_hosts`) | controller data dir | ✅ | ❌ | **NONE — stable** (`mtime 2026-07-18 16:55`) *and* the restic password is escrowed hub-side | | App definitions (`/opt/docker/stacks/*`) | guest rootfs | ✅ | ✅ **yes** — Tier-1 recovery unit captures `compose/` per app, daily; tier-3 lifts it offsite | none | | App data (DB dumps, volume tars) | app drives | ✅ (volumes on mp0) | ✅ **yes** — Tier-1/2/3, daily | none | | Customer bulk data | `/mnt/felhom-drives` | ❌ bind | ✅ Tier-2/3 per class | n/a — never was PBS's job | | `bootstrap.json` | host, bind at mp9 | ❌ bind | ❌ | none — agent-side state, regenerated at provision | | `metrics.db` | controller data dir | ✅ | ❌ | metrics history only — cosmetic | | Guest rootfs / OS / packages | rootfs | ✅ | ❌ | 7 days of OS drift. Re-derivable from the golden image + host-install | ### What is actually lost at 7 days Only the non-SMB half of `settings.json`, and none of it is catastrophic: | Key | Consequence at 7 days | Recovery | |---|---|---| | `storage_paths` | drive enrollments made in the last week are lost from the registry | drives are **physical and still present** — re-run the drive wizard | | `app_backup` | per-app backup toggles revert | customer re-selects | | `notifications` | prefs revert | hub seeds prefs at claim (F11/F12); `customer_notifications` is hub-side | | `password_hash` | dashboard password reverts | hub-issued claim-reset code — the supported path | | `launcher_share_token` | shared launcher URLs break | regenerate + reshare (v0.165.0) | | `claim_*`, `db_validations`, `hub_*` | operational state | re-derived on next hub contact | ### Verdict: **weekly CONFIRMED** Nothing catastrophic is exposed. The two things that would have overturned this — the app-secret encryption key and the offsite credentials — are **stable files that have not changed since first boot**, so a 7-day-old copy is identical to a fresh one. Everything referentially coupled to app state (definitions, DB dumps, volumes, bulk data, SMB shares) is carried **daily** by Tier-1/2/3, and Tier-3 offsite was verified running and `ok` on both production boxes this morning. **⚠️ The conditional — and it is load-bearing.** This verdict depends on the daily offsite tier being alive. **drill-r50 reports `offsite: null`.** On a box with no Tier-3, PBS-weekly would be the *only* DR tier, and the 7-day window would then cover **app data and app definitions**, not just settings — a completely different risk. Recommended ruling to carry into Slice D: **PBS weekly is only sufficient where Tier-3 offsite is enabled and healthy; a box without it needs either offsite enabled first, or a shorter PBS cadence.** That is a cadence decision and belongs to the operator. --- ## P0.2 — The `pvesm status` 0/0/0 KiB anomaly: RESOLVED, benign **It is a PVE-side reporting artifact. The datastore is real, writable, and has capacity.** Reproduced, then settled from three independent directions: **1. The anomaly is real and consistent** — both the CLI and the API agree: ``` pvesm status --storage felhom-pbs → active, Total 0, Used 0, Available 0, 0.00% (rc=0, empty stderr) pvesh get /nodes/demo-felhom/storage/felhom-pbs/status → {"active":1,"avail":0,"content":"backup","enabled":1,"shared":1,"total":0,"type":"pbs","used":0} ``` **2. PBS itself returns zeros to this token — with HTTP 200, not 403.** Queried directly over the wg tunnel as `felhom@pbs!demo-felhom`: ``` GET /api2/json/admin/datastore/felhom-offsite/status → http=200 {"data":{"avail":0,"backend-type":"filesystem","total":0,"used":0}} GET /api2/json/status/datastore-usage → http=200 {"data":[{"backend-type":"filesystem","mount-status":"nonremovable","store":"felhom-offsite"}]} ← usage fields OMITTED GET /api2/json/admin/datastore/felhom-offsite/snapshots?ns=demo-felhom → http=200, real data ``` The token is **namespace-scoped** — `GET .../namespace` returns exactly `[{"ns":"demo-felhom"}]`. It holds `DatastoreBackup` on `/datastore/felhom-offsite/demo-felhom` (per the tenantsync `provision` op's dual-grant), **not** `Datastore.Audit` on the datastore root. PBS answers the status call rather than refusing it, and zeroes/omits the usage figures it will not disclose. PVE has no way to distinguish "zero" from "not permitted" and prints 0/0/0. **3. Ground truth from an independent path says the datastore is healthy.** The hub reads fill over the ep0 SSH `usage` op (a literal `df` on the datastore path — `scripts/felhom-tenantsync.sh` v1.2.0, read-only, no admin token): ``` 2026/07/26 11:46:44 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB) 2026/07/26 12:02:42 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB) ``` **4. Writes demonstrably work with this exact token.** The 2026-07-18 snapshot is owned by `felhom@pbs!demo-felhom`, is 9.74 GB, encrypted, and `verification.state == "ok"`. **Conclusion: cosmetic.** Scheduling recurring writes here is safe on this ground. Two follow-ups recorded, neither blocking: PVE's storage view will keep showing 0% for this storage (the operator must read fill from the hub's PBS-DR gauge, not `pvesm status`), and granting `Datastore.Audit` at the namespace would be the fix if the PVE-side number is ever wanted — a tenantsync `provision` change, deliberately **not** made here. --- ## P0.3 — Capacity: 🛑 STOP → 🟢 **CLEARED 2026-07-27** > **Resolution (2026-07-27).** The datastore was moved off ep0's 40 GB root disk onto a dedicated > **100 GB Hetzner Cloud Volume**. Total **37.2 GB → 98 GB**; used **28.9 % → 13 %**; headroom to the > 80 % warn **19 GB → ≈65 GB**. The "80 % at roughly the second additional customer" projection below > becomes **roughly the seventh to thirteenth**. Method, verification and the restore round-trip that > re-cleared the tier: `runbooks/RUNBOOK-ep0-datastore-volume-2026-07-27.md`. > > **The dedup fact below is unchanged and still governs the slope** — per-tenant encryption still > means no cross-customer dedup. The volume bought runway, not a better cost model. The weekly > incremental size remains **unmeasured**. **The analysis below is the original 2026-07-26 record, retained as written.** ### Measured | Fact | Value | Source | |---|---|---| | `felhom-offsite` datastore total | **37.2 GB** | ep0 `df`, via the hub's usage op | | Used now | **10.8 GB (28.9%)** | same, 2026-07-26 12:02 | | 80% alert threshold | **29.8 GB** | hub PBS-DR box checker (`fill warn=80% crit=90%`) | | **Headroom to the alert** | **19.0 GB** | | | demo-felhom snapshot, logical | **9.74 GB** (`root.pxar.didx` 9.739 GB) | PBS API | | demo-hp local vzdump, compressed | 1.48 GB | host-report | | demo-felhom local vzdump, compressed | 5.69 GB | host-report | | Snapshots in the datastore today | **1** (demo-felhom); demo-hp namespace exists, **0 snapshots** | PBS API | The single 9.74 GB logical snapshot accounts for essentially all of the 10.8 GB used — first-snapshot compression is ≈1:1 here, which is expected when the source is already-compressed Docker layers. ### The dedup fact that makes this worse **Backups are client-side encrypted per tenant** (`encryption-key` in `storage.cfg`, one key per customer). PBS derives chunk digests under the crypt key, so **chunks do not dedup across customers.** Every customer's snapshots cost their full independent size. There is no fleet-scale dedup dividend to lean on. ### Projection (weekly, `keep-last=3` = three weeks) Weekly *incremental* cost is **not measured** — there has never been a second PBS snapshot of any guest to measure it from. Bracketed at 5–20% of full per week and stated as a range rather than a point estimate: | Box | First snapshot | 3 retained weekly | Note | |---|---|---|---| | demo-felhom | 9.7 GB | **10.7 – 13.7 GB** | already on disk | | demo-hp | ~2.5 GB | **2.8 – 3.5 GB** | scaled from local-vzdump ratio 5.69→9.74 (×1.71) | | drill-r50 | ~1.5–3 GB | **1.7 – 4.0 GB** | never backed up; estimate only | | **Fleet total** | | **≈ 15 – 21 GB (40–57%)** | comfortably under 80% | The current three boxes fit. The problem is the next ones: ``` Steady state, 3 boxes: ~18 GB (48%) 80% alert: 29.8 GB Headroom: ~12 GB Per additional customer: ~5–10 GB (first snapshot + 2 retained weekly increments, no cross-tenant dedup) → the 80% alert fires at roughly the SECOND additional customer. ``` ### Why this is a STOP The alpha horizon is *"the first remote tester"* (ROADMAP pre-invite checklist) and beyond — that is squarely within one-to-two additional customers. The gate condition is met: **projected fill crosses the 80% alert threshold within the alpha horizon.** This is not a reason to abandon R-82 — it is a reason not to schedule recurring writes into a 37.2 GB datastore without first ruling on which lever moves: | Lever | Effect | Cost | |---|---|---| | **Grow the ep0 datastore** | the clean fix; 37.2 GB is small for fleet DR | Hetzner resize — money + an operator action | | **Retention `keep-last=2`** weekly | ~1 snapshot/guest saved, ≈15% | two weeks of recovery depth instead of three | | **Fortnightly PBS** | halves increment accrual | doubles the P0.1 exposure window to 14 days — would reopen P0.1 | | **Exclude mp1 `/mnt/sys_drive`** from the snapshot | 1.8 GB used of 50 GB today | needs a ruling on what is on it | | **Stage the rollout** | drill + demo-hp now, defer real customers | buys time, does not solve it | **Recommendation (operator's call):** grow the datastore, and set PBS retention to `keep-last=2` weekly for now. Together they roughly double the runway. But the sizing question deserves an explicit answer before recurring writes start, because the failure mode — a DR datastore that silently refuses writes at full — is exactly the class of "applied and empty" fault R-82 exists to fix. --- ## Not collected | Item | Why | Needed access | |---|---|---| | Whether the ep0 datastore filesystem is **dedicated** or shares the root disk | no direct ep0 SSH; the hub's `usage` op returns only total/used/avail for the datastore path | root on ep0 (167.233.158.164) | | Actual PBS **weekly incremental** size for these guests | no second snapshot has ever existed to diff against | resolved by Slice D's first weekly cycle — the projection should be re-measured then, not trusted | | Other namespaces' usage in `felhom-offsite` | the PVE token is namespace-scoped (sees only `demo-felhom`) | PBS admin token | | demo-hp / drill-r50 host-level state | no SSH key; break-glass not used, per standing scope | baked key or explicit break-glass authorisation | --- ## Observations 1. **drill-r50 has no offsite tier at all** (`offsite: null`) and no backups of any kind. It is the box the rollout plan puts *first*, which is right — but it is also the box where the P0.1 verdict does not hold, so its PBS cadence cannot simply inherit the fleet ruling. 2. **`last_success` is not a field** on the controller's offsite object — the shape is `last_run` + `last_status`. An early probe here asked for `last_success`, got `null`, and briefly looked like a failing daily tier. Recorded because it is the same shape as the R-80 false alarm: *absence of a field read as evidence of failure.* The real values are `last_status: "ok"` on both production boxes. 3. **The PVE storage view will permanently under-report this datastore at 0%.** Any future operator check of PBS fill must use the hub's PBS-DR gauge. Worth a line in the runbook when R-82 lands. 4. **`/mnt/sys_drive` (mp1, `backup=1`, 50 GB allocated, 1.8 GB used)** is inside the snapshot. Nobody asked for it to be; it is worth confirming during Slice D that its contents justify the space at weekly cadence. 5. `07-backup-architecture.md`'s tier table already carries a **PBS whole-guest** row saying *"bind-mounted drives out of reach"* — which the live `uncovered_volumes` confirms exactly. The doc is stale on versions but correct on this mechanism.