Files
felhom.eu/REPORT.md
T
Claude Code adf1d1e619 R-82 Slice D/E: installer default 1.20.0 + architecture docs brought current
Slice D.1 — host-install 1.20.0: a FRESH box defaults to local-daily +
offsite-weekly (felhom-pbs, 604800s, keep_last=2). setdefault semantics proven
both ways: fresh gets the tier, an UPGRADE preserves the existing backup block
verbatim — so an in-place upgrade can never silently start writing to an
offsite datastore. Existing boxes are migrated explicitly.

Slice E:
- 07-backup-architecture.md: honest status header per CONTEXT ruling S-2, with
  an explicit STALE-outside-the-PBS-tier verdict (the controller tiers were last
  verified 41 controller versions ago). The PBS row claimed 'PBS on DooPlex'
  (the retired spike store) with no cadence; it now names felhom-pbs ->
  felhom-offsite on ep0 over wg-felhom, weekly, keep_last=2. NOT marked
  ratified — that is Viktor's review of the section 10 list. Discharges R-83.
- 06-offsite-connectivity.md: the target-split remaining-work note collapsed
  (shipped), and records HOW S4.1's tier-aware timeout silently regressed — the
  mechanism was never removed, its INPUT changed when local_backup_target was
  retargeted to 'local'. Also notes S4.1 already diagnosed the teardown 403 as a
  phantom (a timeout consequence, not an ACL gap).
- capability map: new row for recurring offsite backups actually LANDING, as
  distinct from the existing row proving ACTIVATION. IMPLEMENTED, not
  PROVEN-LIVE — the restore round-trip has not completed under the fixed code.
- ROADMAP: R-82 SHIPPED with its remaining gate named, R-83 DISCHARGED, R-84
  left open.
- CONTEXT + REPORT: the arc, including the mid-arc correction I had to make.
2026-07-26 17:54:10 +02:00

163 lines
10 KiB
Markdown

# REPORT — R-82: the backup target split (2026-07-26)
Local **daily** + offsite **weekly**, made expressible at all. Spans four artifacts:
agent **v0.97.0 → v0.102.0**, controller **v0.174.0 → v0.175.0**, hub **v0.76.0**,
host-install **1.20.0**.
Phase-0 gates: `documentation/audits/SPIKE-r82-phase0-2026-07-26.md`.
---
## 1. What was wrong
`BackupTarget()` returned ONE string and `BackupCadence()` ONE 24 h window, so "local daily **and**
PBS weekly" could not be said. The consequence was not theoretical: the DR tier reported `applied`
since 2026-07-21 while demo-felhom held **one** snapshot (2026-07-18, a healing artifact) and demo-hp
held **zero, ever**. R-39 was "applied and dead"; this was **applied and empty** — the same shape,
quieter, and it would have surfaced first at a real restore.
## 2. Phase 0 — three gates
| Gate | Verdict |
|---|---|
| P0.1 what is exposed for 7 days | **weekly CONFIRMED.** The only 7-day-exposed state is the non-SMB half of `settings.json`. `encryption.key` and the offbox credentials are **stable files unchanged since first boot**, so a week-old copy is byte-identical — that was the risk that could have overturned it |
| P0.2 the `pvesm status` 0/0/0 anomaly | **RESOLVED, benign.** PBS returns HTTP 200 with zeroed usage to a namespace-scoped token (`DatastoreBackup`, not `Datastore.Audit`). Ground truth via the hub's ep0 `df`: the datastore is real and writable |
| P0.3 capacity | **STOP raised; operator ruled to proceed and grow later.** 37.2 GB total. Per-tenant encryption means **no cross-customer dedup** |
**Capacity, now measured rather than bracketed:** the second weekly snapshot cost **+2.7 GB on disk**
against 14.46 GB logical (~81 % dedup). Weekly top-ups are cheap; **first** snapshots are not — one
customer at two retained snapshots costs ~13.5 GB, so the 80 % warn arrives at roughly the **first**
additional customer, not the second as I first estimated. Recorded in `07-backup-architecture.md`
§9.1.
## 3. What shipped
- **Agent** — `backup_targets[]`: each tier carries its **own** cadence, retention and wait bound
(`keep-last=3` is three DAYS on a daily tier and three WEEKS on a weekly one; one shared knob
guarantees one of them is wrong). `/backup/due?target=` judges a tier against **its own** newest
successful backup. `GET /backup/tiers` is the controller's capability probe. One runner per tier.
- **Controller** — every due tier collected up front and run in **ONE quiesce window**. Two cycles on
the weekly night would mean two app outages for one night's work. The app stays quiesced until the
**last** tier snapshots, so every tier is app-consistent.
- **Hub** — per-tier thresholds (host 26 h, offsite 8 d), preserving R-81's three-valued verdicts,
anchored absence and distinct reason strings. Classification is by **target type**
(`target_id``storage_targets[].name``.type`), never by array membership.
- **Installer** — a fresh box defaults to local-daily + offsite-weekly; an unprovisioned tier
**defers** rather than firing at a storage that does not exist.
**The untargeted local-API contract is frozen.** No `?target=` ⇒ the primary tier, same response
**bytes** (`Target` is `omitempty` and stays empty). An old controller cannot tell the new agent from
the old one; a new controller against an old agent degrades on a 404 probe, logs once, and **still
takes the backup**.
## 4. Operator rulings (2026-07-26), all implemented
| Ruling | Implementation |
|---|---|
| Two weeks of offsite backups | `keep_last=2`; the blanket PBS-prune refusal scoped to *additional* tiers with an explicit setting — the primary keeps the absolute refusal, because its target **and** retention both default and could prune the DR by accident |
| Grow the datastore before any real tester | recorded in `07` §9.1; no action taken |
| First backup runs as long as needed; nothing else starts until done | wait bound → 12 h (measured ~5 h for a first full snapshot); **one backup at a time per guest** — a second tier gets a 409 naming the busy tier, with no job id it could mistake for its own; a tier overrunning the quiesce bound defers the rest |
| Drill box is temporary | dropped from the rollout |
| Restore test, then next slice | done — see §6 |
## 5. Four defects found by RUNNING it, not reviewing it
1. **30-minute wait bound vs a 41-minute backup** (v0.98.0). The agent recorded `success:false`
**while the vzdump was still running**, and it later completed `TASK OK`. Not "the backup didn't
happen" but worse: the tier stays permanently due and the retry collides with the guest lock.
2. **The restore tier read from the configured target, not the archive** (v0.100.0). A `felhom-pbs:`
archive was classified `local` and got the 10-minute bound against a 14.46 GB WAN restore, failing
at 600 s. **This was a silent regression of the S4.1 fix** — the mechanism was never removed, its
*input* changed when `local_backup_target` was retargeted to `local`. The lesson is not "add a
timeout" (one was already there) but that a fix keyed on *"the configured target"* stops holding
the moment more than one target exists. Recorded in `06-offsite-connectivity.md`.
3. **A leaked scratch guest kept `onboot: 1`** (v0.101.0) — a host reboot would have started a clone
of the live guest. Now `onboot=0` is set **at restore time**, because "after" is the path that
leaks.
4. **A tier fires at a not-yet-provisioned storage** (v0.102.0) — would have quiesced the apps and
failed every cadence on a fresh box until DR provisioning.
## 6. Live validation
**demo-felhom** — the first real PBS-targeted backup: **`TASK OK`, 41 minutes, 14.46 GB snapshot**,
and it **restored cleanly** (`vzrestore: stopped OK`, all volumes back). Both tiers armed and
verified over the real local API; the untargeted response confirmed byte-identical.
**demo-hp** — reached via the documented break-glass path; binary and config backed up first;
v0.102.0 deployed, tier armed, and its **first ever** PBS backup triggered against a namespace that
was verifiably empty. *(In flight at the time of writing.)*
**Hub Slice C replayed against the live DB before deploying:**
```
demo-felhom host=07-26T14:38Z offsite=07-26T12:21Z -> OK
demo-hp host=07-26T07:06Z offsite=none -> UNKNOWN (119h of a 192h grace)
drill-r50 host=none offsite=not expected -> MISSED (no evidence in 29h)
```
**No customer email results from the deploy.** demo-hp defers correctly and will alarm in ~3 days if
its offsite tier stays empty — the true finding arriving on schedule, not a false alarm.
## 7. A correction I had to make mid-arc
I reported that the restore-test would boot a scratch guest carrying the live guest's MAC, static
island IP and hostname, and so would break the controller→agent link. **That was wrong.**
`RunRestoreTest` step 2 link-downs **every** interface before the guest is started, and it is
unit-tested. I read a restored config artifact, inferred the boot behaviour from it, and escalated
before reading the code path that consumes it. I also disabled the scheduled restore-test on that
basis, which was an unnecessary reduction in safety coverage; it is re-enabled.
The residual hazard was real but far narrower — it needed the restore to fail *before* the link-down
step, which is what defect 1 caused — and that is what v0.101.0 fixes.
Separately, `06-offsite-connectivity.md` records that the teardown `403` I flagged as a standing
privilege gap was **already diagnosed in S4.1 as a phantom**: it is a consequence of the short
timeout, not an ACL problem. With the timeout fixed the guest is pool-associated by teardown time.
## 8. Tests
| Repo | Result |
|---|---|
| felhom-agent | `build/vet/test` rc=0, **29 packages** |
| felhom-controller | `build/vet/test` rc=0, **27 packages** |
| felhom.eu (hub) | `build/vet/test` rc=0, **17 packages** |
Red-proofs observed and restored for every mandatory scenario: old-controller compat, new-controller
degrade (the hollow version asserts "no error" while silently skipping the backup), the both-due
night (**the COUNT is the assertion** — asserting only "both ran" passes against a double-quiesce),
the merged threshold, the per-tier wait bound, the `onboot` override, and the overrun defer.
**A process failure worth recording:** I ran the agent suite and committed in the same command, read
`packages ok: 28`, and pushed **without reading `rc=1`**. Five of my own Slice A tests were failing —
a harness artifact, not a product bug, but the commit went out red. Fixed in `13ca2d9`. This is the
exact exit-code trap recorded twice earlier in this arc.
## 9. NOT done — explicitly
1. **The offsite tier is never AUTOMATICALLY restore-tested.** The scheduled restore-test picks
candidates from a runner built on the primary target, so it can never select a PBS archive. This
is arguably the more important half of "is the DR tier real?" and it needs a per-tick spec.
2. **The hub infers "PBS ⇒ weekly" from storage TYPE.** `defaultBackupTarget` is `felhom-pbs`, so a
box that never sets `local_backup_target` would run PBS as its **daily** tier and be judged
against 8 days — seven days of blindness. No box is in that shape today. The real fix is the agent
reporting each tier's actual cadence.
3. **Every agent restart re-triggers an offsite backup** (the in-memory store, R-84). On the local
tier that is wasted minutes; on the offsite tier it is a wasted multi-hour upload after every
deploy. R-84 is now closer to a prerequisite than a tidy-up.
4. **The installer-default fleet flip** (Slice D step 4) waits on a full weekly cycle holding — a
genuine gate, not an oversight.
5. **The boot+verify half of the restore round-trip has not run under the fixed code.**
6. **`07-backup-architecture.md` is NOT ratified** — brought current with an honest staleness header;
ratification is Viktor's review of the §10 list.
## 10. Observations
- The `felhom-pbs` PVE storage will permanently show **0 %** in the PVE UI (namespace-scoped token).
Operators must read fill from the hub's PBS-DR gauge. Worth a runbook line.
- demo-felhom's guest grew **9.74 → 14.46 GB logical in eight days**. Probably one-off from app
testing, but if it is a rate the capacity sizing changes quickly.
- demo-felhom was enabled **before** demo-hp, out of the specified rollout order, because Slice A
could not be validated otherwise.
- The customer-facing Hungarian copy still overstates scope ("A mai biztonsági mentés nem készült el
a határidőig!" covers only the host/PBS tier). Unchanged; flagged since the R-80 diagnostic.