diff --git a/documentation/audits/SPIKE-r82-phase0-2026-07-26.md b/documentation/audits/SPIKE-r82-phase0-2026-07-26.md new file mode 100644 index 0000000..10116f9 --- /dev/null +++ b/documentation/audits/SPIKE-r82-phase0-2026-07-26.md @@ -0,0 +1,243 @@ +# SPIKE — R-82 Phase 0: the three gates before the backup target split (2026-07-26) + +**Class:** read-only Phase-0 gate for R-82. **No code written. No backup triggered. No config changed.** +Baselines: felhom-agent `dfd5d73` v0.96.0, felhom-controller `47fda06` v0.173.0, felhom.eu `c73800c` +hub v0.75.0. All three trees clean at `origin/main`. + +--- + +## Verdicts — one gate returns STOP + +| Gate | Verdict | +|---|---| +| **P0.1** — what is exposed for seven days | **weekly CONFIRMED**, with the exposed set named and one conditional | +| **P0.2** — the `pvesm status` 0/0/0 anomaly | **RESOLVED — benign PVE-side reporting artifact.** Not a blocker | +| **P0.3** — capacity headroom | **🛑 STOP.** The 80% alert is reached at roughly the **second additional customer**, inside the alpha horizon | + +**R-82 stops here pending an operator ruling on P0.3.** P0.1 and P0.2 clear; the blocker is +capacity, not correctness. + +--- + +## P0.1 — What is exposed for seven days? + +### Method + +The whole-guest vzdump covers what PVE will actually snapshot. Live guest config (`pct config 9201`, +demo-felhom): + +``` +rootfs: local-lvm:vm-9201-disk-0,size=32G → IN the snapshot +mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/docker,backup=1,size=200G → IN +mp1: local-lvm:vm-9201-disk-2,mp=/mnt/sys_drive,backup=1,size=50G → IN +mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives → BIND — NOT in the snapshot +mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1 → BIND — NOT in +``` + +Corroborated by the agent's own report: `"uncovered_volumes": ["/etc/felhom-bootstrap", +"/mnt/felhom-drives"]`. + +**The daily tiers are alive** — this is what makes the answer come out the way it does, so it was +verified rather than assumed: + +| Customer | Tier-3 offsite (restic, daily) | Evidence | +|---|---|---| +| demo-felhom | **running** | `last_run 2026-07-26T02:17:56Z`, `last_status "ok"`, 28 snapshots, 1.08 GB / 50 GB | +| demo-hp | **running** | `last_run 2026-07-26T02:15:50Z`, `last_status "ok"`, 4 snapshots, 29 MB / 50 GB | +| drill-r50 | **`offsite: null` — not configured** | see the conditional below | + +### The exposed-state table + +| State | Where it lives | In the vzdump? | Covered by a DAILY tier? | 7-day staleness cost | +|---|---|---|---|---| +| **`encryption.key`** (app-secret key) | controller data dir | ✅ | ❌ none | **NONE — the key is STABLE.** `mtime 2026-07-18 16:32`, unchanged since creation. A 7-day-old copy is byte-identical. *This was the biggest theoretical risk and it is clean.* | +| **`settings.json` — SMB shares** (`smb`, `smb_shares`) | controller data dir | ✅ | ✅ **yes** — R-7b payload in tier-2 **and** offsite (`_shares-manifest.json` + `passdb.tar`, staged 2026-07-26 02:16) | none | +| **`settings.json` — everything else** | controller data dir | ✅ | ❌ **none** | **THE EXPOSED SET.** `mtime 2026-07-26 10:02` — actively changing. Detailed below | +| `controller.yaml` | controller data dir | ✅ | ❌ | **NONE in practice** — the hub is the source of truth (`config_pull` / bootstrap re-fetch overwrites it; it explicitly never touches `settings.json`) | +| offbox credentials (`repo_password`, `ssh_key`, `known_hosts`) | controller data dir | ✅ | ❌ | **NONE — stable** (`mtime 2026-07-18 16:55`) *and* the restic password is escrowed hub-side | +| App definitions (`/opt/docker/stacks/*`) | guest rootfs | ✅ | ✅ **yes** — Tier-1 recovery unit captures `compose/` per app, daily; tier-3 lifts it offsite | none | +| App data (DB dumps, volume tars) | app drives | ✅ (volumes on mp0) | ✅ **yes** — Tier-1/2/3, daily | none | +| Customer bulk data | `/mnt/felhom-drives` | ❌ bind | ✅ Tier-2/3 per class | n/a — never was PBS's job | +| `bootstrap.json` | host, bind at mp9 | ❌ bind | ❌ | none — agent-side state, regenerated at provision | +| `metrics.db` | controller data dir | ✅ | ❌ | metrics history only — cosmetic | +| Guest rootfs / OS / packages | rootfs | ✅ | ❌ | 7 days of OS drift. Re-derivable from the golden image + host-install | + +### What is actually lost at 7 days + +Only the non-SMB half of `settings.json`, and none of it is catastrophic: + +| Key | Consequence at 7 days | Recovery | +|---|---|---| +| `storage_paths` | drive enrollments made in the last week are lost from the registry | drives are **physical and still present** — re-run the drive wizard | +| `app_backup` | per-app backup toggles revert | customer re-selects | +| `notifications` | prefs revert | hub seeds prefs at claim (F11/F12); `customer_notifications` is hub-side | +| `password_hash` | dashboard password reverts | hub-issued claim-reset code — the supported path | +| `launcher_share_token` | shared launcher URLs break | regenerate + reshare (v0.165.0) | +| `claim_*`, `db_validations`, `hub_*` | operational state | re-derived on next hub contact | + +### Verdict: **weekly CONFIRMED** + +Nothing catastrophic is exposed. The two things that would have overturned this — the app-secret +encryption key and the offsite credentials — are **stable files that have not changed since first +boot**, so a 7-day-old copy is identical to a fresh one. Everything referentially coupled to app +state (definitions, DB dumps, volumes, bulk data, SMB shares) is carried **daily** by Tier-1/2/3, and +Tier-3 offsite was verified running and `ok` on both production boxes this morning. + +**⚠️ The conditional — and it is load-bearing.** This verdict depends on the daily offsite tier being +alive. **drill-r50 reports `offsite: null`.** On a box with no Tier-3, PBS-weekly would be the *only* +DR tier, and the 7-day window would then cover **app data and app definitions**, not just settings — +a completely different risk. Recommended ruling to carry into Slice D: **PBS weekly is only +sufficient where Tier-3 offsite is enabled and healthy; a box without it needs either offsite +enabled first, or a shorter PBS cadence.** That is a cadence decision and belongs to the operator. + +--- + +## P0.2 — The `pvesm status` 0/0/0 KiB anomaly: RESOLVED, benign + +**It is a PVE-side reporting artifact. The datastore is real, writable, and has capacity.** + +Reproduced, then settled from three independent directions: + +**1. The anomaly is real and consistent** — both the CLI and the API agree: +``` +pvesm status --storage felhom-pbs → active, Total 0, Used 0, Available 0, 0.00% (rc=0, empty stderr) +pvesh get /nodes/demo-felhom/storage/felhom-pbs/status + → {"active":1,"avail":0,"content":"backup","enabled":1,"shared":1,"total":0,"type":"pbs","used":0} +``` + +**2. PBS itself returns zeros to this token — with HTTP 200, not 403.** Queried directly over the wg +tunnel as `felhom@pbs!demo-felhom`: +``` +GET /api2/json/admin/datastore/felhom-offsite/status → http=200 + {"data":{"avail":0,"backend-type":"filesystem","total":0,"used":0}} +GET /api2/json/status/datastore-usage → http=200 + {"data":[{"backend-type":"filesystem","mount-status":"nonremovable","store":"felhom-offsite"}]} ← usage fields OMITTED +GET /api2/json/admin/datastore/felhom-offsite/snapshots?ns=demo-felhom → http=200, real data +``` +The token is **namespace-scoped** — `GET .../namespace` returns exactly `[{"ns":"demo-felhom"}]`. It +holds `DatastoreBackup` on `/datastore/felhom-offsite/demo-felhom` (per the tenantsync `provision` +op's dual-grant), **not** `Datastore.Audit` on the datastore root. PBS answers the status call +rather than refusing it, and zeroes/omits the usage figures it will not disclose. PVE has no way to +distinguish "zero" from "not permitted" and prints 0/0/0. + +**3. Ground truth from an independent path says the datastore is healthy.** The hub reads fill over +the ep0 SSH `usage` op (a literal `df` on the datastore path — `scripts/felhom-tenantsync.sh` v1.2.0, +read-only, no admin token): +``` +2026/07/26 11:46:44 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB) +2026/07/26 12:02:42 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB) +``` + +**4. Writes demonstrably work with this exact token.** The 2026-07-18 snapshot is owned by +`felhom@pbs!demo-felhom`, is 9.74 GB, encrypted, and `verification.state == "ok"`. + +**Conclusion: cosmetic.** Scheduling recurring writes here is safe on this ground. Two follow-ups +recorded, neither blocking: PVE's storage view will keep showing 0% for this storage (the operator +must read fill from the hub's PBS-DR gauge, not `pvesm status`), and granting `Datastore.Audit` at +the namespace would be the fix if the PVE-side number is ever wanted — a tenantsync `provision` +change, deliberately **not** made here. + +--- + +## P0.3 — Capacity: 🛑 STOP + +### Measured + +| Fact | Value | Source | +|---|---|---| +| `felhom-offsite` datastore total | **37.2 GB** | ep0 `df`, via the hub's usage op | +| Used now | **10.8 GB (28.9%)** | same, 2026-07-26 12:02 | +| 80% alert threshold | **29.8 GB** | hub PBS-DR box checker (`fill warn=80% crit=90%`) | +| **Headroom to the alert** | **19.0 GB** | | +| demo-felhom snapshot, logical | **9.74 GB** (`root.pxar.didx` 9.739 GB) | PBS API | +| demo-hp local vzdump, compressed | 1.48 GB | host-report | +| demo-felhom local vzdump, compressed | 5.69 GB | host-report | +| Snapshots in the datastore today | **1** (demo-felhom); demo-hp namespace exists, **0 snapshots** | PBS API | + +The single 9.74 GB logical snapshot accounts for essentially all of the 10.8 GB used — first-snapshot +compression is ≈1:1 here, which is expected when the source is already-compressed Docker layers. + +### The dedup fact that makes this worse + +**Backups are client-side encrypted per tenant** (`encryption-key` in `storage.cfg`, one key per +customer). PBS derives chunk digests under the crypt key, so **chunks do not dedup across +customers.** Every customer's snapshots cost their full independent size. There is no fleet-scale +dedup dividend to lean on. + +### Projection (weekly, `keep-last=3` = three weeks) + +Weekly *incremental* cost is **not measured** — there has never been a second PBS snapshot of any +guest to measure it from. Bracketed at 5–20% of full per week and stated as a range rather than a +point estimate: + +| Box | First snapshot | 3 retained weekly | Note | +|---|---|---|---| +| demo-felhom | 9.7 GB | **10.7 – 13.7 GB** | already on disk | +| demo-hp | ~2.5 GB | **2.8 – 3.5 GB** | scaled from local-vzdump ratio 5.69→9.74 (×1.71) | +| drill-r50 | ~1.5–3 GB | **1.7 – 4.0 GB** | never backed up; estimate only | +| **Fleet total** | | **≈ 15 – 21 GB (40–57%)** | comfortably under 80% | + +The current three boxes fit. The problem is the next ones: + +``` +Steady state, 3 boxes: ~18 GB (48%) +80% alert: 29.8 GB +Headroom: ~12 GB +Per additional customer: ~5–10 GB (first snapshot + 2 retained weekly increments, + no cross-tenant dedup) +→ the 80% alert fires at roughly the SECOND additional customer. +``` + +### Why this is a STOP + +The alpha horizon is *"the first remote tester"* (ROADMAP pre-invite checklist) and beyond — that is +squarely within one-to-two additional customers. The gate condition is met: **projected fill crosses +the 80% alert threshold within the alpha horizon.** + +This is not a reason to abandon R-82 — it is a reason not to schedule recurring writes into a 37.2 GB +datastore without first ruling on which lever moves: + +| Lever | Effect | Cost | +|---|---|---| +| **Grow the ep0 datastore** | the clean fix; 37.2 GB is small for fleet DR | Hetzner resize — money + an operator action | +| **Retention `keep-last=2`** weekly | ~1 snapshot/guest saved, ≈15% | two weeks of recovery depth instead of three | +| **Fortnightly PBS** | halves increment accrual | doubles the P0.1 exposure window to 14 days — would reopen P0.1 | +| **Exclude mp1 `/mnt/sys_drive`** from the snapshot | 1.8 GB used of 50 GB today | needs a ruling on what is on it | +| **Stage the rollout** | drill + demo-hp now, defer real customers | buys time, does not solve it | + +**Recommendation (operator's call):** grow the datastore, and set PBS retention to `keep-last=2` +weekly for now. Together they roughly double the runway. But the sizing question deserves an explicit +answer before recurring writes start, because the failure mode — a DR datastore that silently +refuses writes at full — is exactly the class of "applied and empty" fault R-82 exists to fix. + +--- + +## Not collected + +| Item | Why | Needed access | +|---|---|---| +| Whether the ep0 datastore filesystem is **dedicated** or shares the root disk | no direct ep0 SSH; the hub's `usage` op returns only total/used/avail for the datastore path | root on ep0 (167.233.158.164) | +| Actual PBS **weekly incremental** size for these guests | no second snapshot has ever existed to diff against | resolved by Slice D's first weekly cycle — the projection should be re-measured then, not trusted | +| Other namespaces' usage in `felhom-offsite` | the PVE token is namespace-scoped (sees only `demo-felhom`) | PBS admin token | +| demo-hp / drill-r50 host-level state | no SSH key; break-glass not used, per standing scope | baked key or explicit break-glass authorisation | + +--- + +## Observations + +1. **drill-r50 has no offsite tier at all** (`offsite: null`) and no backups of any kind. It is the + box the rollout plan puts *first*, which is right — but it is also the box where the P0.1 verdict + does not hold, so its PBS cadence cannot simply inherit the fleet ruling. +2. **`last_success` is not a field** on the controller's offsite object — the shape is + `last_run` + `last_status`. An early probe here asked for `last_success`, got `null`, and briefly + looked like a failing daily tier. Recorded because it is the same shape as the R-80 false alarm: + *absence of a field read as evidence of failure.* The real values are `last_status: "ok"` on both + production boxes. +3. **The PVE storage view will permanently under-report this datastore at 0%.** Any future operator + check of PBS fill must use the hub's PBS-DR gauge. Worth a line in the runbook when R-82 lands. +4. **`/mnt/sys_drive` (mp1, `backup=1`, 50 GB allocated, 1.8 GB used)** is inside the snapshot. Nobody + asked for it to be; it is worth confirming during Slice D that its contents justify the space at + weekly cadence. +5. `07-backup-architecture.md`'s tier table already carries a **PBS whole-guest** row saying + *"bind-mounted drives out of reach"* — which the live `uncovered_volumes` confirms exactly. The + doc is stale on versions but correct on this mechanism.