Restore round-trip on demo-hp: pass=true, verified=boot+running, mount_parity=ok, source_tier=pbs (the v0.100.0 fix — the earlier attempt said 'local' and died at 600s), 4m5s restore+boot+verify+teardown, clean teardown with no 403 and no leak. That last point confirms 06's reading that the teardown 403 was a phantom, and corrects my earlier framing of it as a standing privilege gap. Multi-tier quiesce driven through the REAL UI endpoint (authed+CSRF): exactly ONE stop/start pair with BOTH backups inside it, local-first/PBS-last, app quiesced through the non-last tier, early resume on the last tier's snapshot. Total downtime 1m27s for both tiers; app healthy after. Capability map row upgraded IMPLEMENTED -> PROVEN-LIVE, kept distinct from the DR-tier row above which proves ACTIVATION not ARRIVAL. Remaining gaps recorded: the SCHEDULED restore-test still only selects the primary tier (manual path proven, unattended not), and the hub infers cadence from storage type.
12 KiB
REPORT — R-82: the backup target split (2026-07-26)
Local daily + offsite weekly, made expressible at all. Spans four artifacts: agent v0.97.0 → v0.102.0, controller v0.174.0 → v0.175.0, hub v0.76.0, host-install 1.20.0.
Phase-0 gates: documentation/audits/SPIKE-r82-phase0-2026-07-26.md.
1. What was wrong
BackupTarget() returned ONE string and BackupCadence() ONE 24 h window, so "local daily and
PBS weekly" could not be said. The consequence was not theoretical: the DR tier reported applied
since 2026-07-21 while demo-felhom held one snapshot (2026-07-18, a healing artifact) and demo-hp
held zero, ever. R-39 was "applied and dead"; this was applied and empty — the same shape,
quieter, and it would have surfaced first at a real restore.
2. Phase 0 — three gates
| Gate | Verdict |
|---|---|
| P0.1 what is exposed for 7 days | weekly CONFIRMED. The only 7-day-exposed state is the non-SMB half of settings.json. encryption.key and the offbox credentials are stable files unchanged since first boot, so a week-old copy is byte-identical — that was the risk that could have overturned it |
P0.2 the pvesm status 0/0/0 anomaly |
RESOLVED, benign. PBS returns HTTP 200 with zeroed usage to a namespace-scoped token (DatastoreBackup, not Datastore.Audit). Ground truth via the hub's ep0 df: the datastore is real and writable |
| P0.3 capacity | STOP raised; operator ruled to proceed and grow later. 37.2 GB total. Per-tenant encryption means no cross-customer dedup |
Capacity, now measured rather than bracketed: the second weekly snapshot cost +2.7 GB on disk
against 14.46 GB logical (~81 % dedup). Weekly top-ups are cheap; first snapshots are not — one
customer at two retained snapshots costs ~13.5 GB, so the 80 % warn arrives at roughly the first
additional customer, not the second as I first estimated. Recorded in 07-backup-architecture.md
§9.1.
3. What shipped
- Agent —
backup_targets[]: each tier carries its own cadence, retention and wait bound (keep-last=3is three DAYS on a daily tier and three WEEKS on a weekly one; one shared knob guarantees one of them is wrong)./backup/due?target=judges a tier against its own newest successful backup.GET /backup/tiersis the controller's capability probe. One runner per tier. - Controller — every due tier collected up front and run in ONE quiesce window. Two cycles on the weekly night would mean two app outages for one night's work. The app stays quiesced until the last tier snapshots, so every tier is app-consistent.
- Hub — per-tier thresholds (host 26 h, offsite 8 d), preserving R-81's three-valued verdicts,
anchored absence and distinct reason strings. Classification is by target type
(
target_id→storage_targets[].name→.type), never by array membership. - Installer — a fresh box defaults to local-daily + offsite-weekly; an unprovisioned tier defers rather than firing at a storage that does not exist.
The untargeted local-API contract is frozen. No ?target= ⇒ the primary tier, same response
bytes (Target is omitempty and stays empty). An old controller cannot tell the new agent from
the old one; a new controller against an old agent degrades on a 404 probe, logs once, and still
takes the backup.
4. Operator rulings (2026-07-26), all implemented
| Ruling | Implementation |
|---|---|
| Two weeks of offsite backups | keep_last=2; the blanket PBS-prune refusal scoped to additional tiers with an explicit setting — the primary keeps the absolute refusal, because its target and retention both default and could prune the DR by accident |
| Grow the datastore before any real tester | recorded in 07 §9.1; no action taken |
| First backup runs as long as needed; nothing else starts until done | wait bound → 12 h (measured ~5 h for a first full snapshot); one backup at a time per guest — a second tier gets a 409 naming the busy tier, with no job id it could mistake for its own; a tier overrunning the quiesce bound defers the rest |
| Drill box is temporary | dropped from the rollout |
| Restore test, then next slice | done — see §6 |
5. Four defects found by RUNNING it, not reviewing it
- 30-minute wait bound vs a 41-minute backup (v0.98.0). The agent recorded
success:falsewhile the vzdump was still running, and it later completedTASK OK. Not "the backup didn't happen" but worse: the tier stays permanently due and the retry collides with the guest lock. - The restore tier read from the configured target, not the archive (v0.100.0). A
felhom-pbs:archive was classifiedlocaland got the 10-minute bound against a 14.46 GB WAN restore, failing at 600 s. This was a silent regression of the S4.1 fix — the mechanism was never removed, its input changed whenlocal_backup_targetwas retargeted tolocal. The lesson is not "add a timeout" (one was already there) but that a fix keyed on "the configured target" stops holding the moment more than one target exists. Recorded in06-offsite-connectivity.md. - A leaked scratch guest kept
onboot: 1(v0.101.0) — a host reboot would have started a clone of the live guest. Nowonboot=0is set at restore time, because "after" is the path that leaks. - A tier fires at a not-yet-provisioned storage (v0.102.0) — would have quiesced the apps and failed every cadence on a fresh box until DR provisioning.
6. Live validation
demo-felhom — the first real PBS-targeted backup: TASK OK, 41 minutes, 14.46 GB snapshot,
and it restored cleanly (vzrestore: stopped OK, all volumes back). Both tiers armed and
verified over the real local API; the untargeted response confirmed byte-identical.
demo-hp — reached via the documented break-glass path; binary and config backed up first. Its first ever PBS backup landed: 4.25 GB into a namespace that was verifiably empty.
THE RESTORE ROUND-TRIP PASSED — the bar for calling a tier real:
source_archive : felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z
source_tier : pbs <- the v0.100.0 fix; the earlier attempt said "local" and died at 600s
pass : true
verified : boot+running
mount_parity : ok <- mp0=/var/lib/docker 50G, mp1=/mnt/sys_drive 20G, mp8/mp9 stand-ins
duration : 4m5s (restore + boot + verify + teardown)
mount_parity is the non-hollow half: a boot-only verify cannot see a missing data volume. The
scratch guest tore down cleanly — no 403, no leak — which confirms 06-offsite-connectivity.md's
reading that the teardown 403 was a phantom (a consequence of the short timeout, not an ACL gap),
and corrects my earlier framing of it as a standing privilege gap. Afterwards: scratch band empty,
thin pool back to its exact pre-restore figure, live guest running, snapshot intact.
THE MULTI-TIER QUIESCE RAN LIVE, through the real UI endpoint (POST /api/guest-backup/trigger,
session auth + CSRF — the exact call the "Mentés most" button makes):
17:01:39 manual backup requested — quiescing now
17:01:39 backup due on 2 tier(s) — quiescing 1 stack(s): [paperless-ngx] <- ONE stop
17:01:46 tier local: backup job ... started
17:02:56 tier local: ... done — next tier may start (app still quiesced) <- app stays DOWN
17:02:56 tier felhom-pbs: backup job backup-9201-felhom-pbs-... started
17:03:06 tier felhom-pbs: ... snapshotted — resuming app early (8B.2)
17:03:06 unquiescing (snapshotted (early resume, last tier)): restarting 1 stack(s) <- ONE start
Exactly one stop/start pair with both backups inside it — the assertion that matters, since "both backups ran" would also pass against an implementation that quiesces twice. Tier order was local-first/PBS-last as designed, the app stayed quiesced through the non-last tier (app-consistency preserved on the DR tier), and it resumed at the last tier's snapshot rather than waiting for the upload. Total app downtime 1m27s for both tiers; paperless came back healthy.
Hub Slice C replayed against the live DB before deploying:
demo-felhom host=07-26T14:38Z offsite=07-26T12:21Z -> OK
demo-hp host=07-26T07:06Z offsite=none -> UNKNOWN (119h of a 192h grace)
drill-r50 host=none offsite=not expected -> MISSED (no evidence in 29h)
No customer email results from the deploy. demo-hp defers correctly and will alarm in ~3 days if its offsite tier stays empty — the true finding arriving on schedule, not a false alarm.
7. A correction I had to make mid-arc
I reported that the restore-test would boot a scratch guest carrying the live guest's MAC, static
island IP and hostname, and so would break the controller→agent link. That was wrong.
RunRestoreTest step 2 link-downs every interface before the guest is started, and it is
unit-tested. I read a restored config artifact, inferred the boot behaviour from it, and escalated
before reading the code path that consumes it. I also disabled the scheduled restore-test on that
basis, which was an unnecessary reduction in safety coverage; it is re-enabled.
The residual hazard was real but far narrower — it needed the restore to fail before the link-down step, which is what defect 1 caused — and that is what v0.101.0 fixes.
Separately, 06-offsite-connectivity.md records that the teardown 403 I flagged as a standing
privilege gap was already diagnosed in S4.1 as a phantom: it is a consequence of the short
timeout, not an ACL problem. With the timeout fixed the guest is pool-associated by teardown time.
8. Tests
| Repo | Result |
|---|---|
| felhom-agent | build/vet/test rc=0, 29 packages |
| felhom-controller | build/vet/test rc=0, 27 packages |
| felhom.eu (hub) | build/vet/test rc=0, 17 packages |
Red-proofs observed and restored for every mandatory scenario: old-controller compat, new-controller
degrade (the hollow version asserts "no error" while silently skipping the backup), the both-due
night (the COUNT is the assertion — asserting only "both ran" passes against a double-quiesce),
the merged threshold, the per-tier wait bound, the onboot override, and the overrun defer.
A process failure worth recording: I ran the agent suite and committed in the same command, read
packages ok: 28, and pushed without reading rc=1. Five of my own Slice A tests were failing —
a harness artifact, not a product bug, but the commit went out red. Fixed in 13ca2d9. This is the
exact exit-code trap recorded twice earlier in this arc.
9. NOT done — explicitly
- The offsite tier is never AUTOMATICALLY restore-tested. The scheduled restore-test picks candidates from a runner built on the primary target, so it can never select a PBS archive. The manual/selftest path is now proven end-to-end; the unattended one is not. Needs a per-tick spec.
- The hub infers "PBS ⇒ weekly" from storage TYPE.
defaultBackupTargetisfelhom-pbs, so a box that never setslocal_backup_targetwould run PBS as its daily tier and be judged against 8 days — seven days of blindness. No box is in that shape today. The real fix is the agent reporting each tier's actual cadence. - The installer-default fleet flip (Slice D step 4) waits on a full weekly cycle holding — a genuine gate, not an oversight.
07-backup-architecture.mdis NOT ratified — brought current with an honest staleness header; ratification is Viktor's review of the §10 list.
10. Observations
- The
felhom-pbsPVE storage will permanently show 0 % in the PVE UI (namespace-scoped token). Operators must read fill from the hub's PBS-DR gauge. Worth a runbook line. - demo-felhom's guest grew 9.74 → 14.46 GB logical in eight days. Probably one-off from app testing, but if it is a rate the capacity sizing changes quickly.
- demo-felhom was enabled before demo-hp, out of the specified rollout order, because Slice A could not be validated otherwise.
- The customer-facing Hungarian copy still overstates scope ("A mai biztonsági mentés nem készült el a határidőig!" covers only the host/PBS tier). Unchanged; flagged since the R-80 diagnostic.