docs(audits): R-82 Phase 0 — P0.1 weekly CONFIRMED, P0.2 benign, P0.3 STOP (capacity)
Read-only gate. No code, no backup triggered, no config changed. P0.1 weekly CONFIRMED: the only 7-day-exposed state is the non-SMB half of settings.json (storage_paths, app_backup toggles, notification prefs, password_hash, launcher token) — none catastrophic, all recoverable. The two items that would have overturned it, encryption.key and the offbox credentials, are STABLE files unchanged since first boot, so a 7-day-old copy is byte-identical. Everything referentially coupled to app state is carried daily by Tier-1/2/3, and Tier-3 offsite was verified running and ok on both production boxes. CONDITIONAL: drill-r50 has offsite:null, so the verdict does not hold there. P0.2 RESOLVED, benign: PBS returns HTTP 200 with zeroed usage to the namespace-scoped token (DatastoreBackup, not Datastore.Audit), so PVE prints 0/0/0. Ground truth via the hub ep0 df op: 28.9% of 37.2 GB. Writes work — the 07-18 snapshot is owned by this token, 9.74 GB, verify ok. P0.3 STOP: 37.2 GB datastore, 10.8 GB used, 80% alert at 29.8 GB. Encrypted per-tenant backups do not dedup across customers. Current 3 boxes project to ~15-21 GB weekly keep-3; each additional customer costs ~5-10 GB, so the 80% alert fires at roughly the SECOND additional customer — inside the alpha horizon. Needs a ruling on datastore size / retention before recurring writes.
This commit is contained in:
@@ -0,0 +1,243 @@
|
|||||||
|
# SPIKE — R-82 Phase 0: the three gates before the backup target split (2026-07-26)
|
||||||
|
|
||||||
|
**Class:** read-only Phase-0 gate for R-82. **No code written. No backup triggered. No config changed.**
|
||||||
|
Baselines: felhom-agent `dfd5d73` v0.96.0, felhom-controller `47fda06` v0.173.0, felhom.eu `c73800c`
|
||||||
|
hub v0.75.0. All three trees clean at `origin/main`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Verdicts — one gate returns STOP
|
||||||
|
|
||||||
|
| Gate | Verdict |
|
||||||
|
|---|---|
|
||||||
|
| **P0.1** — what is exposed for seven days | **weekly CONFIRMED**, with the exposed set named and one conditional |
|
||||||
|
| **P0.2** — the `pvesm status` 0/0/0 anomaly | **RESOLVED — benign PVE-side reporting artifact.** Not a blocker |
|
||||||
|
| **P0.3** — capacity headroom | **🛑 STOP.** The 80% alert is reached at roughly the **second additional customer**, inside the alpha horizon |
|
||||||
|
|
||||||
|
**R-82 stops here pending an operator ruling on P0.3.** P0.1 and P0.2 clear; the blocker is
|
||||||
|
capacity, not correctness.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P0.1 — What is exposed for seven days?
|
||||||
|
|
||||||
|
### Method
|
||||||
|
|
||||||
|
The whole-guest vzdump covers what PVE will actually snapshot. Live guest config (`pct config 9201`,
|
||||||
|
demo-felhom):
|
||||||
|
|
||||||
|
```
|
||||||
|
rootfs: local-lvm:vm-9201-disk-0,size=32G → IN the snapshot
|
||||||
|
mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/docker,backup=1,size=200G → IN
|
||||||
|
mp1: local-lvm:vm-9201-disk-2,mp=/mnt/sys_drive,backup=1,size=50G → IN
|
||||||
|
mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives → BIND — NOT in the snapshot
|
||||||
|
mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1 → BIND — NOT in
|
||||||
|
```
|
||||||
|
|
||||||
|
Corroborated by the agent's own report: `"uncovered_volumes": ["/etc/felhom-bootstrap",
|
||||||
|
"/mnt/felhom-drives"]`.
|
||||||
|
|
||||||
|
**The daily tiers are alive** — this is what makes the answer come out the way it does, so it was
|
||||||
|
verified rather than assumed:
|
||||||
|
|
||||||
|
| Customer | Tier-3 offsite (restic, daily) | Evidence |
|
||||||
|
|---|---|---|
|
||||||
|
| demo-felhom | **running** | `last_run 2026-07-26T02:17:56Z`, `last_status "ok"`, 28 snapshots, 1.08 GB / 50 GB |
|
||||||
|
| demo-hp | **running** | `last_run 2026-07-26T02:15:50Z`, `last_status "ok"`, 4 snapshots, 29 MB / 50 GB |
|
||||||
|
| drill-r50 | **`offsite: null` — not configured** | see the conditional below |
|
||||||
|
|
||||||
|
### The exposed-state table
|
||||||
|
|
||||||
|
| State | Where it lives | In the vzdump? | Covered by a DAILY tier? | 7-day staleness cost |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| **`encryption.key`** (app-secret key) | controller data dir | ✅ | ❌ none | **NONE — the key is STABLE.** `mtime 2026-07-18 16:32`, unchanged since creation. A 7-day-old copy is byte-identical. *This was the biggest theoretical risk and it is clean.* |
|
||||||
|
| **`settings.json` — SMB shares** (`smb`, `smb_shares`) | controller data dir | ✅ | ✅ **yes** — R-7b payload in tier-2 **and** offsite (`_shares-manifest.json` + `passdb.tar`, staged 2026-07-26 02:16) | none |
|
||||||
|
| **`settings.json` — everything else** | controller data dir | ✅ | ❌ **none** | **THE EXPOSED SET.** `mtime 2026-07-26 10:02` — actively changing. Detailed below |
|
||||||
|
| `controller.yaml` | controller data dir | ✅ | ❌ | **NONE in practice** — the hub is the source of truth (`config_pull` / bootstrap re-fetch overwrites it; it explicitly never touches `settings.json`) |
|
||||||
|
| offbox credentials (`repo_password`, `ssh_key`, `known_hosts`) | controller data dir | ✅ | ❌ | **NONE — stable** (`mtime 2026-07-18 16:55`) *and* the restic password is escrowed hub-side |
|
||||||
|
| App definitions (`/opt/docker/stacks/*`) | guest rootfs | ✅ | ✅ **yes** — Tier-1 recovery unit captures `compose/` per app, daily; tier-3 lifts it offsite | none |
|
||||||
|
| App data (DB dumps, volume tars) | app drives | ✅ (volumes on mp0) | ✅ **yes** — Tier-1/2/3, daily | none |
|
||||||
|
| Customer bulk data | `/mnt/felhom-drives` | ❌ bind | ✅ Tier-2/3 per class | n/a — never was PBS's job |
|
||||||
|
| `bootstrap.json` | host, bind at mp9 | ❌ bind | ❌ | none — agent-side state, regenerated at provision |
|
||||||
|
| `metrics.db` | controller data dir | ✅ | ❌ | metrics history only — cosmetic |
|
||||||
|
| Guest rootfs / OS / packages | rootfs | ✅ | ❌ | 7 days of OS drift. Re-derivable from the golden image + host-install |
|
||||||
|
|
||||||
|
### What is actually lost at 7 days
|
||||||
|
|
||||||
|
Only the non-SMB half of `settings.json`, and none of it is catastrophic:
|
||||||
|
|
||||||
|
| Key | Consequence at 7 days | Recovery |
|
||||||
|
|---|---|---|
|
||||||
|
| `storage_paths` | drive enrollments made in the last week are lost from the registry | drives are **physical and still present** — re-run the drive wizard |
|
||||||
|
| `app_backup` | per-app backup toggles revert | customer re-selects |
|
||||||
|
| `notifications` | prefs revert | hub seeds prefs at claim (F11/F12); `customer_notifications` is hub-side |
|
||||||
|
| `password_hash` | dashboard password reverts | hub-issued claim-reset code — the supported path |
|
||||||
|
| `launcher_share_token` | shared launcher URLs break | regenerate + reshare (v0.165.0) |
|
||||||
|
| `claim_*`, `db_validations`, `hub_*` | operational state | re-derived on next hub contact |
|
||||||
|
|
||||||
|
### Verdict: **weekly CONFIRMED**
|
||||||
|
|
||||||
|
Nothing catastrophic is exposed. The two things that would have overturned this — the app-secret
|
||||||
|
encryption key and the offsite credentials — are **stable files that have not changed since first
|
||||||
|
boot**, so a 7-day-old copy is identical to a fresh one. Everything referentially coupled to app
|
||||||
|
state (definitions, DB dumps, volumes, bulk data, SMB shares) is carried **daily** by Tier-1/2/3, and
|
||||||
|
Tier-3 offsite was verified running and `ok` on both production boxes this morning.
|
||||||
|
|
||||||
|
**⚠️ The conditional — and it is load-bearing.** This verdict depends on the daily offsite tier being
|
||||||
|
alive. **drill-r50 reports `offsite: null`.** On a box with no Tier-3, PBS-weekly would be the *only*
|
||||||
|
DR tier, and the 7-day window would then cover **app data and app definitions**, not just settings —
|
||||||
|
a completely different risk. Recommended ruling to carry into Slice D: **PBS weekly is only
|
||||||
|
sufficient where Tier-3 offsite is enabled and healthy; a box without it needs either offsite
|
||||||
|
enabled first, or a shorter PBS cadence.** That is a cadence decision and belongs to the operator.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P0.2 — The `pvesm status` 0/0/0 KiB anomaly: RESOLVED, benign
|
||||||
|
|
||||||
|
**It is a PVE-side reporting artifact. The datastore is real, writable, and has capacity.**
|
||||||
|
|
||||||
|
Reproduced, then settled from three independent directions:
|
||||||
|
|
||||||
|
**1. The anomaly is real and consistent** — both the CLI and the API agree:
|
||||||
|
```
|
||||||
|
pvesm status --storage felhom-pbs → active, Total 0, Used 0, Available 0, 0.00% (rc=0, empty stderr)
|
||||||
|
pvesh get /nodes/demo-felhom/storage/felhom-pbs/status
|
||||||
|
→ {"active":1,"avail":0,"content":"backup","enabled":1,"shared":1,"total":0,"type":"pbs","used":0}
|
||||||
|
```
|
||||||
|
|
||||||
|
**2. PBS itself returns zeros to this token — with HTTP 200, not 403.** Queried directly over the wg
|
||||||
|
tunnel as `felhom@pbs!demo-felhom`:
|
||||||
|
```
|
||||||
|
GET /api2/json/admin/datastore/felhom-offsite/status → http=200
|
||||||
|
{"data":{"avail":0,"backend-type":"filesystem","total":0,"used":0}}
|
||||||
|
GET /api2/json/status/datastore-usage → http=200
|
||||||
|
{"data":[{"backend-type":"filesystem","mount-status":"nonremovable","store":"felhom-offsite"}]} ← usage fields OMITTED
|
||||||
|
GET /api2/json/admin/datastore/felhom-offsite/snapshots?ns=demo-felhom → http=200, real data
|
||||||
|
```
|
||||||
|
The token is **namespace-scoped** — `GET .../namespace` returns exactly `[{"ns":"demo-felhom"}]`. It
|
||||||
|
holds `DatastoreBackup` on `/datastore/felhom-offsite/demo-felhom` (per the tenantsync `provision`
|
||||||
|
op's dual-grant), **not** `Datastore.Audit` on the datastore root. PBS answers the status call
|
||||||
|
rather than refusing it, and zeroes/omits the usage figures it will not disclose. PVE has no way to
|
||||||
|
distinguish "zero" from "not permitted" and prints 0/0/0.
|
||||||
|
|
||||||
|
**3. Ground truth from an independent path says the datastore is healthy.** The hub reads fill over
|
||||||
|
the ep0 SSH `usage` op (a literal `df` on the datastore path — `scripts/felhom-tenantsync.sh` v1.2.0,
|
||||||
|
read-only, no admin token):
|
||||||
|
```
|
||||||
|
2026/07/26 11:46:44 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB)
|
||||||
|
2026/07/26 12:02:42 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB)
|
||||||
|
```
|
||||||
|
|
||||||
|
**4. Writes demonstrably work with this exact token.** The 2026-07-18 snapshot is owned by
|
||||||
|
`felhom@pbs!demo-felhom`, is 9.74 GB, encrypted, and `verification.state == "ok"`.
|
||||||
|
|
||||||
|
**Conclusion: cosmetic.** Scheduling recurring writes here is safe on this ground. Two follow-ups
|
||||||
|
recorded, neither blocking: PVE's storage view will keep showing 0% for this storage (the operator
|
||||||
|
must read fill from the hub's PBS-DR gauge, not `pvesm status`), and granting `Datastore.Audit` at
|
||||||
|
the namespace would be the fix if the PVE-side number is ever wanted — a tenantsync `provision`
|
||||||
|
change, deliberately **not** made here.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P0.3 — Capacity: 🛑 STOP
|
||||||
|
|
||||||
|
### Measured
|
||||||
|
|
||||||
|
| Fact | Value | Source |
|
||||||
|
|---|---|---|
|
||||||
|
| `felhom-offsite` datastore total | **37.2 GB** | ep0 `df`, via the hub's usage op |
|
||||||
|
| Used now | **10.8 GB (28.9%)** | same, 2026-07-26 12:02 |
|
||||||
|
| 80% alert threshold | **29.8 GB** | hub PBS-DR box checker (`fill warn=80% crit=90%`) |
|
||||||
|
| **Headroom to the alert** | **19.0 GB** | |
|
||||||
|
| demo-felhom snapshot, logical | **9.74 GB** (`root.pxar.didx` 9.739 GB) | PBS API |
|
||||||
|
| demo-hp local vzdump, compressed | 1.48 GB | host-report |
|
||||||
|
| demo-felhom local vzdump, compressed | 5.69 GB | host-report |
|
||||||
|
| Snapshots in the datastore today | **1** (demo-felhom); demo-hp namespace exists, **0 snapshots** | PBS API |
|
||||||
|
|
||||||
|
The single 9.74 GB logical snapshot accounts for essentially all of the 10.8 GB used — first-snapshot
|
||||||
|
compression is ≈1:1 here, which is expected when the source is already-compressed Docker layers.
|
||||||
|
|
||||||
|
### The dedup fact that makes this worse
|
||||||
|
|
||||||
|
**Backups are client-side encrypted per tenant** (`encryption-key` in `storage.cfg`, one key per
|
||||||
|
customer). PBS derives chunk digests under the crypt key, so **chunks do not dedup across
|
||||||
|
customers.** Every customer's snapshots cost their full independent size. There is no fleet-scale
|
||||||
|
dedup dividend to lean on.
|
||||||
|
|
||||||
|
### Projection (weekly, `keep-last=3` = three weeks)
|
||||||
|
|
||||||
|
Weekly *incremental* cost is **not measured** — there has never been a second PBS snapshot of any
|
||||||
|
guest to measure it from. Bracketed at 5–20% of full per week and stated as a range rather than a
|
||||||
|
point estimate:
|
||||||
|
|
||||||
|
| Box | First snapshot | 3 retained weekly | Note |
|
||||||
|
|---|---|---|---|
|
||||||
|
| demo-felhom | 9.7 GB | **10.7 – 13.7 GB** | already on disk |
|
||||||
|
| demo-hp | ~2.5 GB | **2.8 – 3.5 GB** | scaled from local-vzdump ratio 5.69→9.74 (×1.71) |
|
||||||
|
| drill-r50 | ~1.5–3 GB | **1.7 – 4.0 GB** | never backed up; estimate only |
|
||||||
|
| **Fleet total** | | **≈ 15 – 21 GB (40–57%)** | comfortably under 80% |
|
||||||
|
|
||||||
|
The current three boxes fit. The problem is the next ones:
|
||||||
|
|
||||||
|
```
|
||||||
|
Steady state, 3 boxes: ~18 GB (48%)
|
||||||
|
80% alert: 29.8 GB
|
||||||
|
Headroom: ~12 GB
|
||||||
|
Per additional customer: ~5–10 GB (first snapshot + 2 retained weekly increments,
|
||||||
|
no cross-tenant dedup)
|
||||||
|
→ the 80% alert fires at roughly the SECOND additional customer.
|
||||||
|
```
|
||||||
|
|
||||||
|
### Why this is a STOP
|
||||||
|
|
||||||
|
The alpha horizon is *"the first remote tester"* (ROADMAP pre-invite checklist) and beyond — that is
|
||||||
|
squarely within one-to-two additional customers. The gate condition is met: **projected fill crosses
|
||||||
|
the 80% alert threshold within the alpha horizon.**
|
||||||
|
|
||||||
|
This is not a reason to abandon R-82 — it is a reason not to schedule recurring writes into a 37.2 GB
|
||||||
|
datastore without first ruling on which lever moves:
|
||||||
|
|
||||||
|
| Lever | Effect | Cost |
|
||||||
|
|---|---|---|
|
||||||
|
| **Grow the ep0 datastore** | the clean fix; 37.2 GB is small for fleet DR | Hetzner resize — money + an operator action |
|
||||||
|
| **Retention `keep-last=2`** weekly | ~1 snapshot/guest saved, ≈15% | two weeks of recovery depth instead of three |
|
||||||
|
| **Fortnightly PBS** | halves increment accrual | doubles the P0.1 exposure window to 14 days — would reopen P0.1 |
|
||||||
|
| **Exclude mp1 `/mnt/sys_drive`** from the snapshot | 1.8 GB used of 50 GB today | needs a ruling on what is on it |
|
||||||
|
| **Stage the rollout** | drill + demo-hp now, defer real customers | buys time, does not solve it |
|
||||||
|
|
||||||
|
**Recommendation (operator's call):** grow the datastore, and set PBS retention to `keep-last=2`
|
||||||
|
weekly for now. Together they roughly double the runway. But the sizing question deserves an explicit
|
||||||
|
answer before recurring writes start, because the failure mode — a DR datastore that silently
|
||||||
|
refuses writes at full — is exactly the class of "applied and empty" fault R-82 exists to fix.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Not collected
|
||||||
|
|
||||||
|
| Item | Why | Needed access |
|
||||||
|
|---|---|---|
|
||||||
|
| Whether the ep0 datastore filesystem is **dedicated** or shares the root disk | no direct ep0 SSH; the hub's `usage` op returns only total/used/avail for the datastore path | root on ep0 (167.233.158.164) |
|
||||||
|
| Actual PBS **weekly incremental** size for these guests | no second snapshot has ever existed to diff against | resolved by Slice D's first weekly cycle — the projection should be re-measured then, not trusted |
|
||||||
|
| Other namespaces' usage in `felhom-offsite` | the PVE token is namespace-scoped (sees only `demo-felhom`) | PBS admin token |
|
||||||
|
| demo-hp / drill-r50 host-level state | no SSH key; break-glass not used, per standing scope | baked key or explicit break-glass authorisation |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Observations
|
||||||
|
|
||||||
|
1. **drill-r50 has no offsite tier at all** (`offsite: null`) and no backups of any kind. It is the
|
||||||
|
box the rollout plan puts *first*, which is right — but it is also the box where the P0.1 verdict
|
||||||
|
does not hold, so its PBS cadence cannot simply inherit the fleet ruling.
|
||||||
|
2. **`last_success` is not a field** on the controller's offsite object — the shape is
|
||||||
|
`last_run` + `last_status`. An early probe here asked for `last_success`, got `null`, and briefly
|
||||||
|
looked like a failing daily tier. Recorded because it is the same shape as the R-80 false alarm:
|
||||||
|
*absence of a field read as evidence of failure.* The real values are `last_status: "ok"` on both
|
||||||
|
production boxes.
|
||||||
|
3. **The PVE storage view will permanently under-report this datastore at 0%.** Any future operator
|
||||||
|
check of PBS fill must use the hub's PBS-DR gauge. Worth a line in the runbook when R-82 lands.
|
||||||
|
4. **`/mnt/sys_drive` (mp1, `backup=1`, 50 GB allocated, 1.8 GB used)** is inside the snapshot. Nobody
|
||||||
|
asked for it to be; it is worth confirming during Slice D that its contents justify the space at
|
||||||
|
weekly cadence.
|
||||||
|
5. `07-backup-architecture.md`'s tier table already carries a **PBS whole-guest** row saying
|
||||||
|
*"bind-mounted drives out of reach"* — which the live `uncovered_volumes` confirms exactly. The
|
||||||
|
doc is stale on versions but correct on this mechanism.
|
||||||
Reference in New Issue
Block a user