195 lines
12 KiB
Markdown
195 lines
12 KiB
Markdown
# REPORT — R-82: the backup target split (2026-07-26)
|
|
|
|
Local **daily** + offsite **weekly**, made expressible at all. Spans four artifacts:
|
|
agent **v0.97.0 → v0.103.0**, controller **v0.174.0 → v0.175.0**, hub **v0.76.0**,
|
|
host-install **1.20.0**.
|
|
|
|
Phase-0 gates: `documentation/audits/SPIKE-r82-phase0-2026-07-26.md`.
|
|
|
|
---
|
|
|
|
## 1. What was wrong
|
|
|
|
`BackupTarget()` returned ONE string and `BackupCadence()` ONE 24 h window, so "local daily **and**
|
|
PBS weekly" could not be said. The consequence was not theoretical: the DR tier reported `applied`
|
|
since 2026-07-21 while demo-felhom held **one** snapshot (2026-07-18, a healing artifact) and demo-hp
|
|
held **zero, ever**. R-39 was "applied and dead"; this was **applied and empty** — the same shape,
|
|
quieter, and it would have surfaced first at a real restore.
|
|
|
|
## 2. Phase 0 — three gates
|
|
|
|
| Gate | Verdict |
|
|
|---|---|
|
|
| P0.1 what is exposed for 7 days | **weekly CONFIRMED.** The only 7-day-exposed state is the non-SMB half of `settings.json`. `encryption.key` and the offbox credentials are **stable files unchanged since first boot**, so a week-old copy is byte-identical — that was the risk that could have overturned it |
|
|
| P0.2 the `pvesm status` 0/0/0 anomaly | **RESOLVED, benign.** PBS returns HTTP 200 with zeroed usage to a namespace-scoped token (`DatastoreBackup`, not `Datastore.Audit`). Ground truth via the hub's ep0 `df`: the datastore is real and writable |
|
|
| P0.3 capacity | **STOP raised; operator ruled to proceed and grow later.** 37.2 GB total. Per-tenant encryption means **no cross-customer dedup** |
|
|
|
|
**Capacity, now measured rather than bracketed:** the second weekly snapshot cost **+2.7 GB on disk**
|
|
against 14.46 GB logical (~81 % dedup). Weekly top-ups are cheap; **first** snapshots are not — one
|
|
customer at two retained snapshots costs ~13.5 GB, so the 80 % warn arrives at roughly the **first**
|
|
additional customer, not the second as I first estimated. Recorded in `07-backup-architecture.md`
|
|
§9.1.
|
|
|
|
## 3. What shipped
|
|
|
|
- **Agent** — `backup_targets[]`: each tier carries its **own** cadence, retention and wait bound
|
|
(`keep-last=3` is three DAYS on a daily tier and three WEEKS on a weekly one; one shared knob
|
|
guarantees one of them is wrong). `/backup/due?target=` judges a tier against **its own** newest
|
|
successful backup. `GET /backup/tiers` is the controller's capability probe. One runner per tier.
|
|
- **Controller** — every due tier collected up front and run in **ONE quiesce window**. Two cycles on
|
|
the weekly night would mean two app outages for one night's work. The app stays quiesced until the
|
|
**last** tier snapshots, so every tier is app-consistent.
|
|
- **Hub** — per-tier thresholds (host 26 h, offsite 8 d), preserving R-81's three-valued verdicts,
|
|
anchored absence and distinct reason strings. Classification is by **target type**
|
|
(`target_id` → `storage_targets[].name` → `.type`), never by array membership.
|
|
- **Installer** — a fresh box defaults to local-daily + offsite-weekly; an unprovisioned tier
|
|
**defers** rather than firing at a storage that does not exist.
|
|
|
|
**The untargeted local-API contract is frozen.** No `?target=` ⇒ the primary tier, same response
|
|
**bytes** (`Target` is `omitempty` and stays empty). An old controller cannot tell the new agent from
|
|
the old one; a new controller against an old agent degrades on a 404 probe, logs once, and **still
|
|
takes the backup**.
|
|
|
|
## 4. Operator rulings (2026-07-26), all implemented
|
|
|
|
| Ruling | Implementation |
|
|
|---|---|
|
|
| Two weeks of offsite backups | `keep_last=2`; the blanket PBS-prune refusal scoped to *additional* tiers with an explicit setting — the primary keeps the absolute refusal, because its target **and** retention both default and could prune the DR by accident |
|
|
| Grow the datastore before any real tester | recorded in `07` §9.1; no action taken |
|
|
| First backup runs as long as needed; nothing else starts until done | wait bound → 12 h (measured ~5 h for a first full snapshot); **one backup at a time per guest** — a second tier gets a 409 naming the busy tier, with no job id it could mistake for its own; a tier overrunning the quiesce bound defers the rest |
|
|
| Drill box is temporary | dropped from the rollout |
|
|
| Restore test, then next slice | done — see §6 |
|
|
|
|
## 5. Four defects found by RUNNING it, not reviewing it
|
|
|
|
1. **30-minute wait bound vs a 41-minute backup** (v0.98.0). The agent recorded `success:false`
|
|
**while the vzdump was still running**, and it later completed `TASK OK`. Not "the backup didn't
|
|
happen" but worse: the tier stays permanently due and the retry collides with the guest lock.
|
|
2. **The restore tier read from the configured target, not the archive** (v0.100.0). A `felhom-pbs:`
|
|
archive was classified `local` and got the 10-minute bound against a 14.46 GB WAN restore, failing
|
|
at 600 s. **This was a silent regression of the S4.1 fix** — the mechanism was never removed, its
|
|
*input* changed when `local_backup_target` was retargeted to `local`. The lesson is not "add a
|
|
timeout" (one was already there) but that a fix keyed on *"the configured target"* stops holding
|
|
the moment more than one target exists. Recorded in `06-offsite-connectivity.md`.
|
|
3. **A leaked scratch guest kept `onboot: 1`** (v0.101.0) — a host reboot would have started a clone
|
|
of the live guest. Now `onboot=0` is set **at restore time**, because "after" is the path that
|
|
leaks.
|
|
4. **A tier fires at a not-yet-provisioned storage** (v0.102.0) — would have quiesced the apps and
|
|
failed every cadence on a fresh box until DR provisioning.
|
|
|
|
## 6. Live validation
|
|
|
|
**demo-felhom** — the first real PBS-targeted backup: **`TASK OK`, 41 minutes, 14.46 GB snapshot**,
|
|
and it **restored cleanly** (`vzrestore: stopped OK`, all volumes back). Both tiers armed and
|
|
verified over the real local API; the untargeted response confirmed byte-identical.
|
|
|
|
**demo-hp** — reached via the documented break-glass path; binary and config backed up first.
|
|
Its **first ever** PBS backup **landed**: 4.25 GB into a namespace that was verifiably empty.
|
|
|
|
**THE RESTORE ROUND-TRIP PASSED** — the bar for calling a tier real:
|
|
|
|
```
|
|
source_archive : felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z
|
|
source_tier : pbs <- the v0.100.0 fix; the earlier attempt said "local" and died at 600s
|
|
pass : true
|
|
verified : boot+running
|
|
mount_parity : ok <- mp0=/var/lib/docker 50G, mp1=/mnt/sys_drive 20G, mp8/mp9 stand-ins
|
|
duration : 4m5s (restore + boot + verify + teardown)
|
|
```
|
|
|
|
`mount_parity` is the non-hollow half: a boot-only verify cannot see a missing data volume. The
|
|
scratch guest tore down **cleanly** — no 403, no leak — which confirms `06-offsite-connectivity.md`'s
|
|
reading that the teardown 403 was a **phantom** (a consequence of the short timeout, not an ACL gap),
|
|
and corrects my earlier framing of it as a standing privilege gap. Afterwards: scratch band empty,
|
|
thin pool back to its exact pre-restore figure, live guest running, snapshot intact.
|
|
|
|
**THE MULTI-TIER QUIESCE RAN LIVE**, through the real UI endpoint (`POST /api/guest-backup/trigger`,
|
|
session auth + CSRF — the exact call the "Mentés most" button makes):
|
|
|
|
```
|
|
17:01:39 manual backup requested — quiescing now
|
|
17:01:39 backup due on 2 tier(s) — quiescing 1 stack(s): [paperless-ngx] <- ONE stop
|
|
17:01:46 tier local: backup job ... started
|
|
17:02:56 tier local: ... done — next tier may start (app still quiesced) <- app stays DOWN
|
|
17:02:56 tier felhom-pbs: backup job backup-9201-felhom-pbs-... started
|
|
17:03:06 tier felhom-pbs: ... snapshotted — resuming app early (8B.2)
|
|
17:03:06 unquiescing (snapshotted (early resume, last tier)): restarting 1 stack(s) <- ONE start
|
|
```
|
|
|
|
**Exactly one stop/start pair with both backups inside it** — the assertion that matters, since
|
|
"both backups ran" would also pass against an implementation that quiesces twice. Tier order was
|
|
local-first/PBS-last as designed, the app stayed quiesced through the non-last tier (app-consistency
|
|
preserved on the DR tier), and it resumed at the **last** tier's snapshot rather than waiting for the
|
|
upload. **Total app downtime 1m27s for both tiers**; paperless came back healthy.
|
|
|
|
**Hub Slice C replayed against the live DB before deploying:**
|
|
|
|
```
|
|
demo-felhom host=07-26T14:38Z offsite=07-26T12:21Z -> OK
|
|
demo-hp host=07-26T07:06Z offsite=none -> UNKNOWN (119h of a 192h grace)
|
|
drill-r50 host=none offsite=not expected -> MISSED (no evidence in 29h)
|
|
```
|
|
|
|
**No customer email results from the deploy.** demo-hp defers correctly and will alarm in ~3 days if
|
|
its offsite tier stays empty — the true finding arriving on schedule, not a false alarm.
|
|
|
|
## 7. A correction I had to make mid-arc
|
|
|
|
I reported that the restore-test would boot a scratch guest carrying the live guest's MAC, static
|
|
island IP and hostname, and so would break the controller→agent link. **That was wrong.**
|
|
`RunRestoreTest` step 2 link-downs **every** interface before the guest is started, and it is
|
|
unit-tested. I read a restored config artifact, inferred the boot behaviour from it, and escalated
|
|
before reading the code path that consumes it. I also disabled the scheduled restore-test on that
|
|
basis, which was an unnecessary reduction in safety coverage; it is re-enabled.
|
|
|
|
The residual hazard was real but far narrower — it needed the restore to fail *before* the link-down
|
|
step, which is what defect 1 caused — and that is what v0.101.0 fixes.
|
|
|
|
Separately, `06-offsite-connectivity.md` records that the teardown `403` I flagged as a standing
|
|
privilege gap was **already diagnosed in S4.1 as a phantom**: it is a consequence of the short
|
|
timeout, not an ACL problem. With the timeout fixed the guest is pool-associated by teardown time.
|
|
|
|
## 8. Tests
|
|
|
|
| Repo | Result |
|
|
|---|---|
|
|
| felhom-agent | `build/vet/test` rc=0, **29 packages** |
|
|
| felhom-controller | `build/vet/test` rc=0, **27 packages** |
|
|
| felhom.eu (hub) | `build/vet/test` rc=0, **17 packages** |
|
|
|
|
Red-proofs observed and restored for every mandatory scenario: old-controller compat, new-controller
|
|
degrade (the hollow version asserts "no error" while silently skipping the backup), the both-due
|
|
night (**the COUNT is the assertion** — asserting only "both ran" passes against a double-quiesce),
|
|
the merged threshold, the per-tier wait bound, the `onboot` override, and the overrun defer.
|
|
|
|
**A process failure worth recording:** I ran the agent suite and committed in the same command, read
|
|
`packages ok: 28`, and pushed **without reading `rc=1`**. Five of my own Slice A tests were failing —
|
|
a harness artifact, not a product bug, but the commit went out red. Fixed in `13ca2d9`. This is the
|
|
exact exit-code trap recorded twice earlier in this arc.
|
|
|
|
## 9. NOT done — explicitly
|
|
|
|
1. **The offsite tier is never AUTOMATICALLY restore-tested.** The scheduled restore-test picks
|
|
candidates from a runner built on the primary target, so it can never select a PBS archive. The
|
|
**manual/selftest** path is now proven end-to-end; the **unattended** one is not. Needs a per-tick
|
|
spec.
|
|
2. **The hub infers "PBS ⇒ weekly" from storage TYPE.** `defaultBackupTarget` is `felhom-pbs`, so a
|
|
box that never sets `local_backup_target` would run PBS as its **daily** tier and be judged
|
|
against 8 days — seven days of blindness. No box is in that shape today. The real fix is the agent
|
|
reporting each tier's actual cadence.
|
|
3. **The installer-default fleet flip** (Slice D step 4) waits on a full weekly cycle holding — a
|
|
genuine gate, not an oversight.
|
|
5. **`07-backup-architecture.md` is NOT ratified** — brought current with an honest staleness header;
|
|
ratification is Viktor's review of the §10 list.
|
|
|
|
## 10. Observations
|
|
|
|
- The `felhom-pbs` PVE storage will permanently show **0 %** in the PVE UI (namespace-scoped token).
|
|
Operators must read fill from the hub's PBS-DR gauge. Worth a runbook line.
|
|
- demo-felhom's guest grew **9.74 → 14.46 GB logical in eight days**. Probably one-off from app
|
|
testing, but if it is a rate the capacity sizing changes quickly.
|
|
- demo-felhom was enabled **before** demo-hp, out of the specified rollout order, because Slice A
|
|
could not be validated otherwise.
|
|
- The customer-facing Hungarian copy still overstates scope ("A mai biztonsági mentés nem készült el
|
|
a határidőig!" covers only the host/PBS tier). Unchanged; flagged since the R-80 diagnostic.
|