Supervised operational run. No code, no version bump, nothing deleted.
Primary backup tier on both demo boxes moved from `local` (a dir storage on
/var/lib/vz -- the SAME physical device as the guest) to `felhom-backup`, a dir
storage on each box's secondary drive:
demo-hp /mnt/nvme-1tb uuid:91d2dc2d-... archive 2,256,044,492 B
demo-felhom /mnt/hdd_1 uuid:47a3361a-... archive 5,957,878,962 B
Both proven end to end via the real UI path: archive lands on the secondary
drive (df delta matches the archive byte-for-byte), restore-test auto-selects it
and passes with mount_parity: ok, and freshness survives an agent restart with
an empty in-memory store -- so the age can only have come from the new storage.
Phase 0: the target is CONFIGURATION, not converged (the sole writer of
agent.json touches only escrow.pbs_storage_id and preserves unknown keys), so
the runbook's STOP did not fire. No consumer hardcodes "local" on the backup path.
Findings:
- F-1 the storage path must BE the mountpoint; a subdirectory fails exactMount
and the target reports disconnected permanently (observe.go:321)
- F-2 --is_mountpoint 1 is load-bearing; proven live, an unguarded storage on a
non-mounted path reports active with the ROOT filesystem's free space and
had already created dump/ on pve-root -- a silent retarget onto the very
device this change escapes
- F-3 FelhomAgentStore is granted per storage path; without it every backup
403s. felhom-host-install.sh must issue it for new installs
- R-109 (new) the DR recipe records no backup target, and each box now carries
two content=backup dir storages, one live and one frozen
- R-105 narrowed and TRACED: dr_recipe drives was [] fleet-wide because the
enrolled drives were never PVE storages, so isUserDataDrive never saw
them. Both boxes now populate drives; SMART on the backup drives too
Absent-drive behaviour today is fail-loudly with no silent retarget (PVE half
live-proven; agent half source-traced). That is NOT the intended fall-back-and-
alarm design -- filed as E-2 with the honest single-drive label.
Reported in full in the record: the agent was restarted with a felhom-pbs backup
in flight, producing a spurious tier failure. The backup had in fact succeeded
(PVE task OK, 6,264,034,053 B snapshot) and the spurious failure reached no
channel -- R-84 ground truth superseded it.
Outstanding: full drive-loss recovery (needs physical access) and the agent half
of the absent-drive behaviour.
5.7 KiB
REPORT — the local whole-guest backup moved off the guest's own device (2026-07-28)
Overwritten per the standing rule (the prior C9-F1/C9-F2 text stays in git history).
Supervised operational run — no code, no version bump. Fleet unchanged: hub v0.80.0,
agent v0.110.0, controller v0.183.0. peti-felhom untouched. Nothing was deleted.
Record: documentation/runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md.
What changed
On both demo boxes the primary backup tier moved from local — a dir storage on /var/lib/vz,
i.e. the same physical device as the guest itself — to felhom-backup, a dir storage on
the secondary drive.
| demo-felhom | demo-hp | |
|---|---|---|
| Target drive | /dev/sdb USB HDD → /mnt/hdd_1 |
/dev/nvme0n1 → /mnt/nvme-1tb |
| Durable id | uuid:47a3361a-91e0-4831-a69d-27f540ed3f48 |
uuid:91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae |
| Archive proven | 5,957,878,962 B | 2,256,044,492 B |
| Restore-test | pass, mount_parity: ok, 1 m 24 s |
pass, mount_parity: ok, 1 m 54 s |
Three steps per box: pvesm add dir … --is_mountpoint 1, two pveum acl modify, and a one-line
local_backup_target edit in agent.json.
Drive loss is now locally recoverable in principle on both boxes — the whole-guest archive lives on different hardware from the guest, and a restore from it boots and passes mount parity. The proof that closes matrix row 4 (pull the system drive and restore for real) needs physical access and is outstanding.
Phase 0 verdict
The target is configuration, not converged — the runbook held and the §3 STOP did not fire.
Exactly one writer of agent.json exists (pbsdr.seedEscrowStorageID); it touches only
escrow.pbs_storage_id and preserves unknown keys verbatim via a map[string]json.RawMessage
read-modify-write. Nothing on the hub pushes agent config. No consumer hardcodes "local" on the
backup path — runners are one-per-tier, NewestArchiveTime reads its own target, and
restoreTierForArchive classifies from the archive, not from config.
What the run found
- F-1 (before any command). The storage
pathmust be the drive's mountpoint. A subdirectory failsexactMount(internal/storage/observe.go:321), so the target would reportdisconnectedpermanently and its durable id would degrade off the filesystem UUID. - F-2 (before any command).
--is_mountpoint 1is load-bearing. Proven live: an unguardeddirstorage on a non-mounted path reportsactive, advertises the root filesystem's free space, and had already createddump/on/dev/mapper/pve-root— a silent retarget onto the exact device this change exists to escape. The guarded one refuses outright. - F-3 (found by the first real backup).
FelhomAgentStoreis granted per storage path; a new target without its own grant 403s every backup.felhom-host-install.shmust issue it for new installs, or a new box ships with a tier that fails on its first run. - R-109 (new). The DR recipe records no backup target — and each box now carries two
content=backupdir storages, one live and one holding frozen 2026-07-28 archives. - R-105 narrowed and traced.
dr_recipe.host_half.driveswas[]fleet-wide with the cause untraced. Cause: the enrolled drives were never PVE storages, soisUserDataDrivenever saw them. Both demo boxes now populatedrives, and the backup drives gained SMART reporting. R-105's other two fields are untouched.
Absent-drive behaviour (Part 3)
Today: fail loudly, no silent retarget. The PVE half is live-proven with throwaway storages (no
live drive was unmounted). The agent half is source-traced: targetStoragePresent checks name
presence only, never Reachable, so the tier stays DUE, the controller quiesces, vzdump is refused,
and the run fails and alarms.
This is not §6's intended design (fall back to the system drive and alarm). There is no fallback at all, so an absent drive means no local backup until a human intervenes. Filed as E-2, together with the honest single-drive label — a one-drive box protects against corruption only, and two drives is effectively a hardware requirement for drive-loss protection.
One operational error, reported in full
The agent was restarted on demo-hp with a felhom-pbs backup in flight, against §4.3. The
in-flight check was done before the first restart and not repeated before the second. It produced a
spurious tier failure (context canceled while waiting) — the F-A1 class the project already
fixed once.
The backup had not failed: the PVE task returned OK and the PBS snapshot 2026-07-28T19:19:45Z
is 6,264,034,053 B. The spurious failure reached no channel — the agent died with its in-memory
record and R-84 ground truth superseded it; the controller's event trail for the window shows only
the real 403 failure and its recovery. The design absorbed it; the mistake was still a mistake, and
on a slower tier the same slip could have aborted a multi-hour WAN upload.
Also recorded: pgrep -f vzdump self-matches a polling script's own command line and is not a safe
in-flight check — use the PVE task list.
State at close
Both boxes healthy. Primary and offsite tiers due=false with age_state=known on both; breakers
clear; agents active; no thrash and no spurious staleness. Target drives at 1–2 % used with SMART
PASSED. The old archives (18 GB demo-felhom, 5.2 GB demo-hp) are left in place as the rollback
and as the only evidence of what the previous configuration produced.
Outstanding: full drive-loss recovery (physical access), and the agent half of the absent-drive behaviour (needs a drive unmount that would break the guest bind on a remote box).