Supervised runbook execution. No code change, no version bump. felhom-offsite moved from ep0's 40 GB root disk (/srv/pbs-felhom) to a dedicated 100 GB Hetzner Cloud Volume (/mnt/pbs-datastore, ext4 -m 0, by-id fstab, relatime). Datastore NAME unchanged, so the PBS-DR descriptors, per-box storage ids, ACLs and namespaces are untouched. Capacity: 37.2 GB -> 98 GB total, 28.9% -> 13% used, headroom to the 80% warn 19 GB -> ~65 GB. This CLEARS the R-82 Phase 0 P0.3 STOP. Per-tenant encryption still precludes cross-customer dedup, so the slope is unchanged - the volume buys runway, not a better cost model. Verified: byte totals and chunk counts identical (9748), 7/7 snapshots across all three namespaces, backup:backup ownership, clean itemised dry-run, full verify job TASK OK with 0 errors, and a restore round-trip (source_tier pbs, pass true, mount_parity ok, clean teardown). Nothing deleted - the original 13 GB stays at /srv/pbs-felhom as the rollback until a new weekly backup lands. GC deliberately not run. Three findings recorded: - the `scratch` datastore points at a non-existent path (pre-existing; now logs ENOENT every start) - operator decision - the runbook's S6 guard test proves the wrong proposition: RequiresMountsFor re-mounts rather than refusing, so the test only bites when the device is genuinely unavailable (re-run that way, and the refusal was observed) - amendment recommended - S11: storage box u629193 has no live backup path, BUT ep0 carries an enabled sshfs mount unit against it that must be removed before the box is deleted Deviations: the volume arrived pre-formatted and mounted; S8 ran on demo-felhom rather than demo-hp (no SSH key for demo-hp); the window was contended by a stale in-memory 10-minute restore-test cadence whose config had already been reverted on disk. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ARoadHBf8rHoscfiqeVZn
15 KiB
SPIKE — R-82 Phase 0: the three gates before the backup target split (2026-07-26)
Class: read-only Phase-0 gate for R-82. No code written. No backup triggered. No config changed.
Baselines: felhom-agent dfd5d73 v0.96.0, felhom-controller 47fda06 v0.173.0, felhom.eu c73800c
hub v0.75.0. All three trees clean at origin/main.
Verdicts — one gate returns STOP
| Gate | Verdict |
|---|---|
| P0.1 — what is exposed for seven days | weekly CONFIRMED, with the exposed set named and one conditional |
P0.2 — the pvesm status 0/0/0 anomaly |
RESOLVED — benign PVE-side reporting artifact. Not a blocker |
| P0.3 — capacity headroom | 🛑 STOP. The 80% alert is reached at roughly the second additional customer, inside the alpha horizon |
R-82 stops here pending an operator ruling on P0.3. P0.1 and P0.2 clear; the blocker is capacity, not correctness.
P0.1 — What is exposed for seven days?
Method
The whole-guest vzdump covers what PVE will actually snapshot. Live guest config (pct config 9201,
demo-felhom):
rootfs: local-lvm:vm-9201-disk-0,size=32G → IN the snapshot
mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/docker,backup=1,size=200G → IN
mp1: local-lvm:vm-9201-disk-2,mp=/mnt/sys_drive,backup=1,size=50G → IN
mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives → BIND — NOT in the snapshot
mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1 → BIND — NOT in
Corroborated by the agent's own report: "uncovered_volumes": ["/etc/felhom-bootstrap", "/mnt/felhom-drives"].
The daily tiers are alive — this is what makes the answer come out the way it does, so it was verified rather than assumed:
| Customer | Tier-3 offsite (restic, daily) | Evidence |
|---|---|---|
| demo-felhom | running | last_run 2026-07-26T02:17:56Z, last_status "ok", 28 snapshots, 1.08 GB / 50 GB |
| demo-hp | running | last_run 2026-07-26T02:15:50Z, last_status "ok", 4 snapshots, 29 MB / 50 GB |
| drill-r50 | offsite: null — not configured |
see the conditional below |
The exposed-state table
| State | Where it lives | In the vzdump? | Covered by a DAILY tier? | 7-day staleness cost |
|---|---|---|---|---|
encryption.key (app-secret key) |
controller data dir | ✅ | ❌ none | NONE — the key is STABLE. mtime 2026-07-18 16:32, unchanged since creation. A 7-day-old copy is byte-identical. This was the biggest theoretical risk and it is clean. |
settings.json — SMB shares (smb, smb_shares) |
controller data dir | ✅ | ✅ yes — R-7b payload in tier-2 and offsite (_shares-manifest.json + passdb.tar, staged 2026-07-26 02:16) |
none |
settings.json — everything else |
controller data dir | ✅ | ❌ none | THE EXPOSED SET. mtime 2026-07-26 10:02 — actively changing. Detailed below |
controller.yaml |
controller data dir | ✅ | ❌ | NONE in practice — the hub is the source of truth (config_pull / bootstrap re-fetch overwrites it; it explicitly never touches settings.json) |
offbox credentials (repo_password, ssh_key, known_hosts) |
controller data dir | ✅ | ❌ | NONE — stable (mtime 2026-07-18 16:55) and the restic password is escrowed hub-side |
App definitions (/opt/docker/stacks/*) |
guest rootfs | ✅ | ✅ yes — Tier-1 recovery unit captures compose/ per app, daily; tier-3 lifts it offsite |
none |
| App data (DB dumps, volume tars) | app drives | ✅ (volumes on mp0) | ✅ yes — Tier-1/2/3, daily | none |
| Customer bulk data | /mnt/felhom-drives |
❌ bind | ✅ Tier-2/3 per class | n/a — never was PBS's job |
bootstrap.json |
host, bind at mp9 | ❌ bind | ❌ | none — agent-side state, regenerated at provision |
metrics.db |
controller data dir | ✅ | ❌ | metrics history only — cosmetic |
| Guest rootfs / OS / packages | rootfs | ✅ | ❌ | 7 days of OS drift. Re-derivable from the golden image + host-install |
What is actually lost at 7 days
Only the non-SMB half of settings.json, and none of it is catastrophic:
| Key | Consequence at 7 days | Recovery |
|---|---|---|
storage_paths |
drive enrollments made in the last week are lost from the registry | drives are physical and still present — re-run the drive wizard |
app_backup |
per-app backup toggles revert | customer re-selects |
notifications |
prefs revert | hub seeds prefs at claim (F11/F12); customer_notifications is hub-side |
password_hash |
dashboard password reverts | hub-issued claim-reset code — the supported path |
launcher_share_token |
shared launcher URLs break | regenerate + reshare (v0.165.0) |
claim_*, db_validations, hub_* |
operational state | re-derived on next hub contact |
Verdict: weekly CONFIRMED
Nothing catastrophic is exposed. The two things that would have overturned this — the app-secret
encryption key and the offsite credentials — are stable files that have not changed since first
boot, so a 7-day-old copy is identical to a fresh one. Everything referentially coupled to app
state (definitions, DB dumps, volumes, bulk data, SMB shares) is carried daily by Tier-1/2/3, and
Tier-3 offsite was verified running and ok on both production boxes this morning.
⚠️ The conditional — and it is load-bearing. This verdict depends on the daily offsite tier being
alive. drill-r50 reports offsite: null. On a box with no Tier-3, PBS-weekly would be the only
DR tier, and the 7-day window would then cover app data and app definitions, not just settings —
a completely different risk. Recommended ruling to carry into Slice D: PBS weekly is only
sufficient where Tier-3 offsite is enabled and healthy; a box without it needs either offsite
enabled first, or a shorter PBS cadence. That is a cadence decision and belongs to the operator.
P0.2 — The pvesm status 0/0/0 KiB anomaly: RESOLVED, benign
It is a PVE-side reporting artifact. The datastore is real, writable, and has capacity.
Reproduced, then settled from three independent directions:
1. The anomaly is real and consistent — both the CLI and the API agree:
pvesm status --storage felhom-pbs → active, Total 0, Used 0, Available 0, 0.00% (rc=0, empty stderr)
pvesh get /nodes/demo-felhom/storage/felhom-pbs/status
→ {"active":1,"avail":0,"content":"backup","enabled":1,"shared":1,"total":0,"type":"pbs","used":0}
2. PBS itself returns zeros to this token — with HTTP 200, not 403. Queried directly over the wg
tunnel as felhom@pbs!demo-felhom:
GET /api2/json/admin/datastore/felhom-offsite/status → http=200
{"data":{"avail":0,"backend-type":"filesystem","total":0,"used":0}}
GET /api2/json/status/datastore-usage → http=200
{"data":[{"backend-type":"filesystem","mount-status":"nonremovable","store":"felhom-offsite"}]} ← usage fields OMITTED
GET /api2/json/admin/datastore/felhom-offsite/snapshots?ns=demo-felhom → http=200, real data
The token is namespace-scoped — GET .../namespace returns exactly [{"ns":"demo-felhom"}]. It
holds DatastoreBackup on /datastore/felhom-offsite/demo-felhom (per the tenantsync provision
op's dual-grant), not Datastore.Audit on the datastore root. PBS answers the status call
rather than refusing it, and zeroes/omits the usage figures it will not disclose. PVE has no way to
distinguish "zero" from "not permitted" and prints 0/0/0.
3. Ground truth from an independent path says the datastore is healthy. The hub reads fill over
the ep0 SSH usage op (a literal df on the datastore path — scripts/felhom-tenantsync.sh v1.2.0,
read-only, no admin token):
2026/07/26 11:46:44 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB)
2026/07/26 12:02:42 [INFO] PBS-DR box refreshed: 28.9% full (10.8 GB of 37.2 GB)
4. Writes demonstrably work with this exact token. The 2026-07-18 snapshot is owned by
felhom@pbs!demo-felhom, is 9.74 GB, encrypted, and verification.state == "ok".
Conclusion: cosmetic. Scheduling recurring writes here is safe on this ground. Two follow-ups
recorded, neither blocking: PVE's storage view will keep showing 0% for this storage (the operator
must read fill from the hub's PBS-DR gauge, not pvesm status), and granting Datastore.Audit at
the namespace would be the fix if the PVE-side number is ever wanted — a tenantsync provision
change, deliberately not made here.
P0.3 — Capacity: 🛑 STOP → 🟢 CLEARED 2026-07-27
Resolution (2026-07-27). The datastore was moved off ep0's 40 GB root disk onto a dedicated 100 GB Hetzner Cloud Volume. Total 37.2 GB → 98 GB; used 28.9 % → 13 %; headroom to the 80 % warn 19 GB → ≈65 GB. The "80 % at roughly the second additional customer" projection below becomes roughly the seventh to thirteenth. Method, verification and the restore round-trip that re-cleared the tier:
runbooks/RUNBOOK-ep0-datastore-volume-2026-07-27.md.The dedup fact below is unchanged and still governs the slope — per-tenant encryption still means no cross-customer dedup. The volume bought runway, not a better cost model. The weekly incremental size remains unmeasured.
The analysis below is the original 2026-07-26 record, retained as written.
Measured
| Fact | Value | Source |
|---|---|---|
felhom-offsite datastore total |
37.2 GB | ep0 df, via the hub's usage op |
| Used now | 10.8 GB (28.9%) | same, 2026-07-26 12:02 |
| 80% alert threshold | 29.8 GB | hub PBS-DR box checker (fill warn=80% crit=90%) |
| Headroom to the alert | 19.0 GB | |
| demo-felhom snapshot, logical | 9.74 GB (root.pxar.didx 9.739 GB) |
PBS API |
| demo-hp local vzdump, compressed | 1.48 GB | host-report |
| demo-felhom local vzdump, compressed | 5.69 GB | host-report |
| Snapshots in the datastore today | 1 (demo-felhom); demo-hp namespace exists, 0 snapshots | PBS API |
The single 9.74 GB logical snapshot accounts for essentially all of the 10.8 GB used — first-snapshot compression is ≈1:1 here, which is expected when the source is already-compressed Docker layers.
The dedup fact that makes this worse
Backups are client-side encrypted per tenant (encryption-key in storage.cfg, one key per
customer). PBS derives chunk digests under the crypt key, so chunks do not dedup across
customers. Every customer's snapshots cost their full independent size. There is no fleet-scale
dedup dividend to lean on.
Projection (weekly, keep-last=3 = three weeks)
Weekly incremental cost is not measured — there has never been a second PBS snapshot of any guest to measure it from. Bracketed at 5–20% of full per week and stated as a range rather than a point estimate:
| Box | First snapshot | 3 retained weekly | Note |
|---|---|---|---|
| demo-felhom | 9.7 GB | 10.7 – 13.7 GB | already on disk |
| demo-hp | ~2.5 GB | 2.8 – 3.5 GB | scaled from local-vzdump ratio 5.69→9.74 (×1.71) |
| drill-r50 | ~1.5–3 GB | 1.7 – 4.0 GB | never backed up; estimate only |
| Fleet total | ≈ 15 – 21 GB (40–57%) | comfortably under 80% |
The current three boxes fit. The problem is the next ones:
Steady state, 3 boxes: ~18 GB (48%)
80% alert: 29.8 GB
Headroom: ~12 GB
Per additional customer: ~5–10 GB (first snapshot + 2 retained weekly increments,
no cross-tenant dedup)
→ the 80% alert fires at roughly the SECOND additional customer.
Why this is a STOP
The alpha horizon is "the first remote tester" (ROADMAP pre-invite checklist) and beyond — that is squarely within one-to-two additional customers. The gate condition is met: projected fill crosses the 80% alert threshold within the alpha horizon.
This is not a reason to abandon R-82 — it is a reason not to schedule recurring writes into a 37.2 GB datastore without first ruling on which lever moves:
| Lever | Effect | Cost |
|---|---|---|
| Grow the ep0 datastore | the clean fix; 37.2 GB is small for fleet DR | Hetzner resize — money + an operator action |
Retention keep-last=2 weekly |
~1 snapshot/guest saved, ≈15% | two weeks of recovery depth instead of three |
| Fortnightly PBS | halves increment accrual | doubles the P0.1 exposure window to 14 days — would reopen P0.1 |
Exclude mp1 /mnt/sys_drive from the snapshot |
1.8 GB used of 50 GB today | needs a ruling on what is on it |
| Stage the rollout | drill + demo-hp now, defer real customers | buys time, does not solve it |
Recommendation (operator's call): grow the datastore, and set PBS retention to keep-last=2
weekly for now. Together they roughly double the runway. But the sizing question deserves an explicit
answer before recurring writes start, because the failure mode — a DR datastore that silently
refuses writes at full — is exactly the class of "applied and empty" fault R-82 exists to fix.
Not collected
| Item | Why | Needed access |
|---|---|---|
| Whether the ep0 datastore filesystem is dedicated or shares the root disk | no direct ep0 SSH; the hub's usage op returns only total/used/avail for the datastore path |
root on ep0 (167.233.158.164) |
| Actual PBS weekly incremental size for these guests | no second snapshot has ever existed to diff against | resolved by Slice D's first weekly cycle — the projection should be re-measured then, not trusted |
Other namespaces' usage in felhom-offsite |
the PVE token is namespace-scoped (sees only demo-felhom) |
PBS admin token |
| demo-hp / drill-r50 host-level state | no SSH key; break-glass not used, per standing scope | baked key or explicit break-glass authorisation |
Observations
- drill-r50 has no offsite tier at all (
offsite: null) and no backups of any kind. It is the box the rollout plan puts first, which is right — but it is also the box where the P0.1 verdict does not hold, so its PBS cadence cannot simply inherit the fleet ruling. last_successis not a field on the controller's offsite object — the shape islast_run+last_status. An early probe here asked forlast_success, gotnull, and briefly looked like a failing daily tier. Recorded because it is the same shape as the R-80 false alarm: absence of a field read as evidence of failure. The real values arelast_status: "ok"on both production boxes.- The PVE storage view will permanently under-report this datastore at 0%. Any future operator check of PBS fill must use the hub's PBS-DR gauge. Worth a line in the runbook when R-82 lands.
/mnt/sys_drive(mp1,backup=1, 50 GB allocated, 1.8 GB used) is inside the snapshot. Nobody asked for it to be; it is worth confirming during Slice D that its contents justify the space at weekly cadence.07-backup-architecture.md's tier table already carries a PBS whole-guest row saying "bind-mounted drives out of reach" — which the liveuncovered_volumesconfirms exactly. The doc is stale on versions but correct on this mechanism.