Files
felhom.eu/REPORT.md
T
admin b5a73e050b Move the local whole-guest backup off the guest's own device (demo-hp + demo-felhom)
Supervised operational run. No code, no version bump, nothing deleted.

Primary backup tier on both demo boxes moved from `local` (a dir storage on
/var/lib/vz -- the SAME physical device as the guest) to `felhom-backup`, a dir
storage on each box's secondary drive:

  demo-hp      /mnt/nvme-1tb  uuid:91d2dc2d-...  archive 2,256,044,492 B
  demo-felhom  /mnt/hdd_1     uuid:47a3361a-...  archive 5,957,878,962 B

Both proven end to end via the real UI path: archive lands on the secondary
drive (df delta matches the archive byte-for-byte), restore-test auto-selects it
and passes with mount_parity: ok, and freshness survives an agent restart with
an empty in-memory store -- so the age can only have come from the new storage.

Phase 0: the target is CONFIGURATION, not converged (the sole writer of
agent.json touches only escrow.pbs_storage_id and preserves unknown keys), so
the runbook's STOP did not fire. No consumer hardcodes "local" on the backup path.

Findings:
- F-1  the storage path must BE the mountpoint; a subdirectory fails exactMount
       and the target reports disconnected permanently (observe.go:321)
- F-2  --is_mountpoint 1 is load-bearing; proven live, an unguarded storage on a
       non-mounted path reports active with the ROOT filesystem's free space and
       had already created dump/ on pve-root -- a silent retarget onto the very
       device this change escapes
- F-3  FelhomAgentStore is granted per storage path; without it every backup
       403s. felhom-host-install.sh must issue it for new installs
- R-109 (new) the DR recipe records no backup target, and each box now carries
       two content=backup dir storages, one live and one frozen
- R-105 narrowed and TRACED: dr_recipe drives was [] fleet-wide because the
       enrolled drives were never PVE storages, so isUserDataDrive never saw
       them. Both boxes now populate drives; SMART on the backup drives too

Absent-drive behaviour today is fail-loudly with no silent retarget (PVE half
live-proven; agent half source-traced). That is NOT the intended fall-back-and-
alarm design -- filed as E-2 with the honest single-drive label.

Reported in full in the record: the agent was restarted with a felhom-pbs backup
in flight, producing a spurious tier failure. The backup had in fact succeeded
(PVE task OK, 6,264,034,053 B snapshot) and the spurious failure reached no
channel -- R-84 ground truth superseded it.

Outstanding: full drive-loss recovery (needs physical access) and the agent half
of the absent-drive behaviour.
2026-07-28 21:38:13 +02:00

97 lines
5.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — the local whole-guest backup moved off the guest's own device (2026-07-28)
**Overwritten** per the standing rule (the prior C9-F1/C9-F2 text stays in git history).
**Supervised operational run — no code, no version bump.** Fleet unchanged: hub v0.80.0,
agent v0.110.0, controller v0.183.0. `peti-felhom` untouched. **Nothing was deleted.**
Record: `documentation/runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md`.
---
## What changed
On both demo boxes the primary backup tier moved from `local` — a `dir` storage on `/var/lib/vz`,
i.e. **the same physical device as the guest itself** — to **`felhom-backup`**, a `dir` storage on
the secondary drive.
| | demo-felhom | demo-hp |
|---|---|---|
| Target drive | `/dev/sdb` USB HDD → `/mnt/hdd_1` | `/dev/nvme0n1``/mnt/nvme-1tb` |
| Durable id | `uuid:47a3361a-91e0-4831-a69d-27f540ed3f48` | `uuid:91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae` |
| Archive proven | 5,957,878,962 B | 2,256,044,492 B |
| Restore-test | `pass`, **`mount_parity: ok`**, 1 m 24 s | `pass`, **`mount_parity: ok`**, 1 m 54 s |
Three steps per box: `pvesm add dir … --is_mountpoint 1`, two `pveum acl modify`, and a one-line
`local_backup_target` edit in `agent.json`.
**Drive loss is now locally recoverable in principle on both boxes** — the whole-guest archive lives
on different hardware from the guest, and a restore from it boots and passes mount parity. The proof
that closes matrix row 4 (pull the system drive and restore for real) needs physical access and is
outstanding.
## Phase 0 verdict
**The target is configuration, not converged — the runbook held and the §3 STOP did not fire.**
Exactly one writer of `agent.json` exists (`pbsdr.seedEscrowStorageID`); it touches only
`escrow.pbs_storage_id` and preserves unknown keys verbatim via a `map[string]json.RawMessage`
read-modify-write. Nothing on the hub pushes agent config. No consumer hardcodes `"local"` on the
backup path — runners are one-per-tier, `NewestArchiveTime` reads its own target, and
`restoreTierForArchive` classifies from the archive, not from config.
## What the run found
- **F-1 (before any command).** The storage `path` must **be** the drive's mountpoint. A
subdirectory fails `exactMount` (`internal/storage/observe.go:321`), so the target would report
`disconnected` **permanently** and its durable id would degrade off the filesystem UUID.
- **F-2 (before any command).** `--is_mountpoint 1` is load-bearing. **Proven live:** an unguarded
`dir` storage on a non-mounted path reports `active`, advertises the **root filesystem's** free
space, and had already created `dump/` on `/dev/mapper/pve-root` — a silent retarget onto the
exact device this change exists to escape. The guarded one refuses outright.
- **F-3 (found by the first real backup).** `FelhomAgentStore` is granted **per storage path**; a
new target without its own grant 403s every backup. **`felhom-host-install.sh` must issue it for
new installs**, or a new box ships with a tier that fails on its first run.
- **R-109 (new).** The DR recipe records **no backup target** — and each box now carries two
`content=backup` dir storages, one live and one holding frozen 2026-07-28 archives.
- **R-105 narrowed and traced.** `dr_recipe.host_half.drives` was `[]` fleet-wide with the cause
untraced. Cause: the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them.
Both demo boxes now populate `drives`, and the backup drives gained SMART reporting. R-105's other
two fields are untouched.
## Absent-drive behaviour (Part 3)
Today: **fail loudly, no silent retarget.** The PVE half is live-proven with throwaway storages (no
live drive was unmounted). The agent half is source-traced: `targetStoragePresent` checks name
presence only, never `Reachable`, so the tier stays DUE, the controller quiesces, vzdump is refused,
and the run fails and alarms.
**This is not §6's intended design** (fall back to the system drive and alarm). There is no fallback
at all, so an absent drive means no local backup until a human intervenes. Filed as **E-2**, together
with the honest single-drive label — a one-drive box protects against corruption only, and two
drives is effectively a hardware requirement for drive-loss protection.
## One operational error, reported in full
**The agent was restarted on demo-hp with a `felhom-pbs` backup in flight**, against §4.3. The
in-flight check was done before the first restart and not repeated before the second. It produced a
**spurious tier failure** (`context canceled` while waiting) — the F-A1 class the project already
fixed once.
**The backup had not failed:** the PVE task returned OK and the PBS snapshot `2026-07-28T19:19:45Z`
is 6,264,034,053 B. The spurious failure **reached no channel** — the agent died with its in-memory
record and R-84 ground truth superseded it; the controller's event trail for the window shows only
the real 403 failure and its recovery. The design absorbed it; the mistake was still a mistake, and
on a slower tier the same slip could have aborted a multi-hour WAN upload.
Also recorded: `pgrep -f vzdump` self-matches a polling script's own command line and is not a safe
in-flight check — use the PVE task list.
## State at close
Both boxes healthy. Primary and offsite tiers `due=false` with `age_state=known` on both; breakers
clear; agents `active`; no thrash and no spurious staleness. Target drives at 12 % used with SMART
`PASSED`. The old archives (18 GB demo-felhom, 5.2 GB demo-hp) are **left in place** as the rollback
and as the only evidence of what the previous configuration produced.
**Outstanding:** full drive-loss recovery (physical access), and the agent half of the absent-drive
behaviour (needs a drive unmount that would break the guest bind on a remote box).