docs(arch): 07 §9.1 — record measured PBS whole-guest capacity state (R-82 Phase 0)

Per the operator's 2026-07-26 ruling: datastore growth deferred, R-82 proceeds.
Records the measurements so the constraint is written down rather than carried
in a session: 37.2 GB total / 10.8 GB used, no cross-customer dedup (per-tenant
encryption keys), 80% warn reached at roughly the second additional customer,
and the pvesm 0/0/0 reporting artifact that means operators must read fill from
the hub gauge. Also records what the tier does and does not carry, and the
conditional on the P0.1 weekly verdict (Tier-3 offsite must be healthy).

Doc NOT marked ratified — that stays Viktor's review (R-83).
This commit is contained in:
Claude Code
2026-07-26 12:11:42 +02:00
parent 65e30f2e42
commit 945b7818b5
@@ -322,6 +322,54 @@ backups/secondary/<stack>/
class member that grows like a media library. The pre-push gate + notification is what makes
that growth a conversation instead of a surprise.
### 9.1 PBS whole-guest tier — measured capacity state (R-82 Phase 0, 2026-07-26)
Recorded per the operator's 2026-07-26 ruling: **the datastore will be grown later; R-82 proceeds
meanwhile.** These are measurements, not projections-of-record — re-measure before relying on them.
Full method + evidence: `audits/SPIKE-r82-phase0-2026-07-26.md`.
| Fact | Value | Source |
|---|---|---|
| `felhom-offsite` datastore total (ep0) | **37.2 GB** | ep0 `df` via the hub `usage` op (`scripts/felhom-tenantsync.sh` v1.2.0) |
| Used | **10.8 GB (28.9%)** | same, 2026-07-26 12:02 UTC |
| Hub alert thresholds | warn **80%** (29.8 GB) / crit **90%** | `monitor.PBSDRBoxChecker` |
| Snapshots present | **1** — demo-felhom `ct/9201/2026-07-18T18:31:06Z`, 9.74 GB logical, `verify ok`. demo-hp namespace exists with **zero** | PBS API |
| First-snapshot compression | ≈ **1:1** (9.74 GB logical ≈ 10.8 GB on disk) | source is already-compressed Docker layers |
| Weekly incremental size | **UNMEASURED** — no guest has ever had a second PBS snapshot | measure at the first Slice-D weekly cycle |
**Three facts that govern the cost model here — all differ from the restic tier above:**
1. **No cross-customer dedup.** Backups are client-side encrypted with a per-customer key
(`encryption-key` in `storage.cfg`), so PBS derives chunk digests under that key and chunks
never dedup between tenants. Every customer's snapshots cost their full independent size. There
is no fleet-scale dividend — the opposite of the shared-model assumption in §9 above.
2. **The datastore is small relative to the fleet.** At weekly `keep-last` retention the current
three boxes project to ≈1521 GB (4057%); each additional customer costs ≈510 GB, so the 80%
warn is reached at roughly the **second additional customer**. This is the open constraint the
operator deferred, not a solved problem.
3. **PVE cannot see this datastore's fill.** `pvesm status felhom-pbs` reports **0/0/0 KiB while
active** — PBS answers the status call with HTTP 200 and zeroed usage because the per-customer
token holds `DatastoreBackup` on its namespace, not `Datastore.Audit` on the datastore root.
**Operators must read PBS fill from the hub's PBS-DR gauge, never from `pvesm`/the PVE UI.**
Cosmetic, not a fault: writes work (the 07-18 snapshot is owned by that exact token).
**What the PBS tier does and does not carry** (confirmed live against `pct config 9201` and the
agent's own `uncovered_volumes`): rootfs + `/var/lib/docker` (mp0, `backup=1`) + `/mnt/sys_drive`
(mp1, `backup=1`) are **in**; `/mnt/felhom-drives` (mp8) and `/etc/felhom-bootstrap` (mp9) are bind
mounts and are **out**. The §2 tier table's "bind-mounted drives out of reach" is correct.
**Why weekly is sufficient (R-82 P0.1 verdict).** The only state exposed by a 7-day-old snapshot is
the non-SMB half of `settings.json``storage_paths`, `app_backup` toggles, notification prefs,
`password_hash`, `launcher_share_token` — all recoverable (drives are physical and re-enrollable;
the password has a hub claim-reset path). Critically, `encryption.key` and the offbox credentials
are **stable files unchanged since first boot**, so a week-old copy is byte-identical. Everything
referentially coupled to app state rides the DAILY tiers.
> ⚠️ **The verdict is conditional on Tier-3 offsite being enabled and healthy.** On a box without it
> (drill-r50 reports `offsite: null`), PBS-weekly is the ONLY DR tier and the 7-day window would
> then cover app data and definitions too. Such a box needs offsite enabled first, or a shorter PBS
> cadence.
---
## 10. Decisions requiring Viktor confirmation (one-line vetoes)