b5a73e050b
Supervised operational run. No code, no version bump, nothing deleted.
Primary backup tier on both demo boxes moved from `local` (a dir storage on
/var/lib/vz -- the SAME physical device as the guest) to `felhom-backup`, a dir
storage on each box's secondary drive:
demo-hp /mnt/nvme-1tb uuid:91d2dc2d-... archive 2,256,044,492 B
demo-felhom /mnt/hdd_1 uuid:47a3361a-... archive 5,957,878,962 B
Both proven end to end via the real UI path: archive lands on the secondary
drive (df delta matches the archive byte-for-byte), restore-test auto-selects it
and passes with mount_parity: ok, and freshness survives an agent restart with
an empty in-memory store -- so the age can only have come from the new storage.
Phase 0: the target is CONFIGURATION, not converged (the sole writer of
agent.json touches only escrow.pbs_storage_id and preserves unknown keys), so
the runbook's STOP did not fire. No consumer hardcodes "local" on the backup path.
Findings:
- F-1 the storage path must BE the mountpoint; a subdirectory fails exactMount
and the target reports disconnected permanently (observe.go:321)
- F-2 --is_mountpoint 1 is load-bearing; proven live, an unguarded storage on a
non-mounted path reports active with the ROOT filesystem's free space and
had already created dump/ on pve-root -- a silent retarget onto the very
device this change escapes
- F-3 FelhomAgentStore is granted per storage path; without it every backup
403s. felhom-host-install.sh must issue it for new installs
- R-109 (new) the DR recipe records no backup target, and each box now carries
two content=backup dir storages, one live and one frozen
- R-105 narrowed and TRACED: dr_recipe drives was [] fleet-wide because the
enrolled drives were never PVE storages, so isUserDataDrive never saw
them. Both boxes now populate drives; SMART on the backup drives too
Absent-drive behaviour today is fail-loudly with no silent retarget (PVE half
live-proven; agent half source-traced). That is NOT the intended fall-back-and-
alarm design -- filed as E-2 with the honest single-drive label.
Reported in full in the record: the agent was restarted with a felhom-pbs backup
in flight, producing a spurious tier failure. The backup had in fact succeeded
(PVE task OK, 6,264,034,053 B snapshot) and the spurious failure reached no
channel -- R-84 ground truth superseded it.
Outstanding: full drive-loss recovery (needs physical access) and the agent half
of the absent-drive behaviour.
97 lines
5.7 KiB
Markdown
97 lines
5.7 KiB
Markdown
# REPORT — the local whole-guest backup moved off the guest's own device (2026-07-28)
|
||
|
||
**Overwritten** per the standing rule (the prior C9-F1/C9-F2 text stays in git history).
|
||
**Supervised operational run — no code, no version bump.** Fleet unchanged: hub v0.80.0,
|
||
agent v0.110.0, controller v0.183.0. `peti-felhom` untouched. **Nothing was deleted.**
|
||
|
||
Record: `documentation/runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md`.
|
||
|
||
---
|
||
|
||
## What changed
|
||
|
||
On both demo boxes the primary backup tier moved from `local` — a `dir` storage on `/var/lib/vz`,
|
||
i.e. **the same physical device as the guest itself** — to **`felhom-backup`**, a `dir` storage on
|
||
the secondary drive.
|
||
|
||
| | demo-felhom | demo-hp |
|
||
|---|---|---|
|
||
| Target drive | `/dev/sdb` USB HDD → `/mnt/hdd_1` | `/dev/nvme0n1` → `/mnt/nvme-1tb` |
|
||
| Durable id | `uuid:47a3361a-91e0-4831-a69d-27f540ed3f48` | `uuid:91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae` |
|
||
| Archive proven | 5,957,878,962 B | 2,256,044,492 B |
|
||
| Restore-test | `pass`, **`mount_parity: ok`**, 1 m 24 s | `pass`, **`mount_parity: ok`**, 1 m 54 s |
|
||
|
||
Three steps per box: `pvesm add dir … --is_mountpoint 1`, two `pveum acl modify`, and a one-line
|
||
`local_backup_target` edit in `agent.json`.
|
||
|
||
**Drive loss is now locally recoverable in principle on both boxes** — the whole-guest archive lives
|
||
on different hardware from the guest, and a restore from it boots and passes mount parity. The proof
|
||
that closes matrix row 4 (pull the system drive and restore for real) needs physical access and is
|
||
outstanding.
|
||
|
||
## Phase 0 verdict
|
||
|
||
**The target is configuration, not converged — the runbook held and the §3 STOP did not fire.**
|
||
Exactly one writer of `agent.json` exists (`pbsdr.seedEscrowStorageID`); it touches only
|
||
`escrow.pbs_storage_id` and preserves unknown keys verbatim via a `map[string]json.RawMessage`
|
||
read-modify-write. Nothing on the hub pushes agent config. No consumer hardcodes `"local"` on the
|
||
backup path — runners are one-per-tier, `NewestArchiveTime` reads its own target, and
|
||
`restoreTierForArchive` classifies from the archive, not from config.
|
||
|
||
## What the run found
|
||
|
||
- **F-1 (before any command).** The storage `path` must **be** the drive's mountpoint. A
|
||
subdirectory fails `exactMount` (`internal/storage/observe.go:321`), so the target would report
|
||
`disconnected` **permanently** and its durable id would degrade off the filesystem UUID.
|
||
- **F-2 (before any command).** `--is_mountpoint 1` is load-bearing. **Proven live:** an unguarded
|
||
`dir` storage on a non-mounted path reports `active`, advertises the **root filesystem's** free
|
||
space, and had already created `dump/` on `/dev/mapper/pve-root` — a silent retarget onto the
|
||
exact device this change exists to escape. The guarded one refuses outright.
|
||
- **F-3 (found by the first real backup).** `FelhomAgentStore` is granted **per storage path**; a
|
||
new target without its own grant 403s every backup. **`felhom-host-install.sh` must issue it for
|
||
new installs**, or a new box ships with a tier that fails on its first run.
|
||
- **R-109 (new).** The DR recipe records **no backup target** — and each box now carries two
|
||
`content=backup` dir storages, one live and one holding frozen 2026-07-28 archives.
|
||
- **R-105 narrowed and traced.** `dr_recipe.host_half.drives` was `[]` fleet-wide with the cause
|
||
untraced. Cause: the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them.
|
||
Both demo boxes now populate `drives`, and the backup drives gained SMART reporting. R-105's other
|
||
two fields are untouched.
|
||
|
||
## Absent-drive behaviour (Part 3)
|
||
|
||
Today: **fail loudly, no silent retarget.** The PVE half is live-proven with throwaway storages (no
|
||
live drive was unmounted). The agent half is source-traced: `targetStoragePresent` checks name
|
||
presence only, never `Reachable`, so the tier stays DUE, the controller quiesces, vzdump is refused,
|
||
and the run fails and alarms.
|
||
|
||
**This is not §6's intended design** (fall back to the system drive and alarm). There is no fallback
|
||
at all, so an absent drive means no local backup until a human intervenes. Filed as **E-2**, together
|
||
with the honest single-drive label — a one-drive box protects against corruption only, and two
|
||
drives is effectively a hardware requirement for drive-loss protection.
|
||
|
||
## One operational error, reported in full
|
||
|
||
**The agent was restarted on demo-hp with a `felhom-pbs` backup in flight**, against §4.3. The
|
||
in-flight check was done before the first restart and not repeated before the second. It produced a
|
||
**spurious tier failure** (`context canceled` while waiting) — the F-A1 class the project already
|
||
fixed once.
|
||
|
||
**The backup had not failed:** the PVE task returned OK and the PBS snapshot `2026-07-28T19:19:45Z`
|
||
is 6,264,034,053 B. The spurious failure **reached no channel** — the agent died with its in-memory
|
||
record and R-84 ground truth superseded it; the controller's event trail for the window shows only
|
||
the real 403 failure and its recovery. The design absorbed it; the mistake was still a mistake, and
|
||
on a slower tier the same slip could have aborted a multi-hour WAN upload.
|
||
|
||
Also recorded: `pgrep -f vzdump` self-matches a polling script's own command line and is not a safe
|
||
in-flight check — use the PVE task list.
|
||
|
||
## State at close
|
||
|
||
Both boxes healthy. Primary and offsite tiers `due=false` with `age_state=known` on both; breakers
|
||
clear; agents `active`; no thrash and no spurious staleness. Target drives at 1–2 % used with SMART
|
||
`PASSED`. The old archives (18 GB demo-felhom, 5.2 GB demo-hp) are **left in place** as the rollback
|
||
and as the only evidence of what the previous configuration produced.
|
||
|
||
**Outstanding:** full drive-loss recovery (physical access), and the agent half of the absent-drive
|
||
behaviour (needs a drive unmount that would break the guest bind on a remote box).
|