E-2 increment 1: report + close E-2b/E-2c as shipped and proven live

hub 0.81.0, agent 0.112.0, controller 0.184.1 live on BOTH demo boxes.

E-2c: eject/decommission of the backup-target drive refused 409 on both boxes,
drives unmoved. E-2b: the never-called disconnect seam is wired, with the target
case raising the specific backup_target_absent.

Records the keying bug caught before deploy (a.Path is the GUEST path, so the
target branch was unreachable -- 0.184.0 superseded, never deployed) and states
plainly that backup_target_absent is NOT proven end-to-end live: proving it needs
a live enrolled drive to go absent.

Parts 2/3/4 and E-2a remain open; Peti risk stays parked.
This commit is contained in:
2026-07-29 08:34:41 +02:00
parent 6869a14015
commit 3696188636
2 changed files with 90 additions and 75 deletions
+88 -73
View File
@@ -1,96 +1,111 @@
# REPORT — the local whole-guest backup moved off the guest's own device (2026-07-28)
# REPORT — E-2 increment 1: the backup target is now defended and audible (2026-07-29)
**Overwritten** per the standing rule (the prior C9-F1/C9-F2 text stays in git history).
**Supervised operational run — no code, no version bump.** Fleet unchanged: hub v0.80.0,
agent v0.110.0, controller v0.183.0. `peti-felhom` untouched. **Nothing was deleted.**
**Overwritten** per the standing rule. **E-2 is PARTIALLY shipped** — this increment is Parts 1 + 5
plus findings E-2b/E-2c. Parts 2, 3, 4 and E-2a remain open (see §6).
Record: `documentation/runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md`.
**Live fleet after this run:** hub **0.81.0**, agent **0.112.0** (both boxes), controller **0.184.1**
(both boxes). `peti-felhom` deliberately untouched.
---
## What changed
## 1. Why this order
On both demo boxes the primary backup tier moved from `local` — a `dir` storage on `/var/lib/vz`,
i.e. **the same physical device as the guest itself** — to **`felhom-backup`**, a `dir` storage on
the secondary drive.
The operator reordered the plan, correctly: **E-2b and E-2c protect a configuration that exists
right now; E-2a builds for boxes that do not.** Both demo boxes are configured, Peti is parked, and
no new installs are happening this week — so Part 3 has no user yet, while E-1's new configuration
had a drive that could be ejected without warning and no event when it went absent.
| | demo-felhom | demo-hp |
|---|---|---|
| Target drive | `/dev/sdb` USB HDD → `/mnt/hdd_1` | `/dev/nvme0n1``/mnt/nvme-1tb` |
| Durable id | `uuid:47a3361a-91e0-4831-a69d-27f540ed3f48` | `uuid:91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae` |
| Archive proven | 5,957,878,962 B | 2,256,044,492 B |
| Restore-test | `pass`, **`mount_parity: ok`**, 1 m 24 s | `pass`, **`mount_parity: ok`**, 1 m 54 s |
## 2. E-2c — the backup drive could be ejected out from under the backup
Three steps per box: `pvesm add dir … --is_mountpoint 1`, two `pveum acl modify`, and a one-line
`local_backup_target` edit in `agent.json`.
`RoleForStorage` types a `local-dir` on a non-system device as **user-data**, so the pre-existing
eject role gate PASSED the drive E-1 had just made the vzdump target. `POST /disks/eject` on
`/mnt/nvme-1tb` or `/mnt/hdd_1` would have **succeeded silently**, taking the only local whole-guest
backup with it while the box kept reporting a configured tier.
**Drive loss is now locally recoverable in principle on both boxes** — the whole-guest archive lives
on different hardware from the guest, and a restore from it boots and passes mount parity. The proof
that closes matrix row 4 (pull the system drive and restore for real) needs physical access and is
outstanding.
Fixed in agent v0.111.0/0.112.0: `handleDiskEject` and `handleDiskDecommission` refuse `409` when the
mount backs a configured tier, naming the storage and the remedy — the op is *ordered*, not
forbidden.
## Phase 0 verdict
**Deliberately NOT a role reclassification.** Making `RoleForStorage` return `RoleBackup` would
refuse every legitimate eject of the customer's own data drive, because on both boxes that drive
**is** the target. `TestEjectStillAllowedOnANonTargetDrive` pins that, and the over-correction is
red-proofed.
**The target is configuration, not converged — the runbook held and the §3 STOP did not fire.**
Exactly one writer of `agent.json` exists (`pbsdr.seedEscrowStorageID`); it touches only
`escrow.pbs_storage_id` and preserves unknown keys verbatim via a `map[string]json.RawMessage`
read-modify-write. Nothing on the hub pushes agent config. No consumer hardcodes `"local"` on the
backup path — runners are one-per-tier, `NewestArchiveTime` reads its own target, and
`restoreTierForArchive` classifies from the archive, not from config.
**Live on both boxes:**
## What the run found
```
demo-hp felhom-backup mount=/mnt/nvme-1tb role=user-data backup_target=True
demo-felhom felhom-backup mount=/mnt/hdd_1 role=user-data backup_target=True
```
- **F-1 (before any command).** The storage `path` must **be** the drive's mountpoint. A
subdirectory fails `exactMount` (`internal/storage/observe.go:321`), so the target would report
`disconnected` **permanently** and its durable id would degrade off the filesystem UUID.
- **F-2 (before any command).** `--is_mountpoint 1` is load-bearing. **Proven live:** an unguarded
`dir` storage on a non-mounted path reports `active`, advertises the **root filesystem's** free
space, and had already created `dump/` on `/dev/mapper/pve-root` — a silent retarget onto the
exact device this change exists to escape. The guarded one refuses outright.
- **F-3 (found by the first real backup).** `FelhomAgentStore` is granted **per storage path**; a
new target without its own grant 403s every backup. **`felhom-host-install.sh` must issue it for
new installs**, or a new box ships with a tier that fails on its first run.
- **R-109 (new).** The DR recipe records **no backup target** — and each box now carries two
`content=backup` dir storages, one live and one holding frozen 2026-07-28 archives.
- **R-105 narrowed and traced.** `dr_recipe.host_half.drives` was `[]` fleet-wide with the cause
untraced. Cause: the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them.
Both demo boxes now populate `drives`, and the backup drives gained SMART reporting. R-105's other
two fields are untouched.
That single line is the exposure: `role=user-data` (which the gate permits) yet `backup_target=True`.
Both ejects returned **409** with the full message; both decommissions **409**; both drives still
mounted with `dump/` intact — the refusal had no side effect. `/var/lib/vz` is still refused by the
**pre-existing** role gate (`role: system`), proving the new gate is narrow.
## Absent-drive behaviour (Part 3)
## 3. E-2b — the alarm that was defined and never called
Today: **fail loudly, no silent retarget.** The PVE half is live-proven with throwaway storages (no
live drive was unmounted). The agent half is source-traced: `targetStoragePresent` checks name
presence only, never `Reachable`, so the tier stays DUE, the controller quiesces, vzdump is refused,
and the run fails and alarms.
`NotifyStorageDisconnected` and `NotifyStorageReconnected` were registered in `allowedEventTypes`, in
`DefaultEnabledEvents`, and given Hungarian hub copy — and **called from nowhere**. A drive going
absent produced stopped apps, a `[WARN]` log and a UI badge, then silence on every channel. A drive
that is *only* a backup target has no apps to stop, so it was silent twice over. Fifth instance of
this class; found by Phase 0, not by a failure.
**This is not §6's intended design** (fall back to the system drive and alarm). There is no fallback
at all, so an absent drive means no local backup until a human intervenes. Filed as **E-2**, together
with the honest single-drive label — a one-drive box protects against corruption only, and two
drives is effectively a hardware requirement for drive-loss protection.
Now wired in `ReconcileDriveGates`, with the target case raising the more specific
`backup_target_absent` (error) and recovering as `backup_target_restored` (info — the pairing-gated
pattern; `severityNotifies` NOT widened). Never both: two mails for one event trains people to
ignore the channel.
## One operational error, reported in full
## 4. The bug I introduced and caught before deploy
**The agent was restarted on demo-hp with a `felhom-pbs` backup in flight**, against §4.3. The
in-flight check was done before the first restart and not repeated before the second. It produced a
**spurious tier failure** (`context canceled` while waiting) — the F-A1 class the project already
fixed once.
`ReconcileDriveGates` resolves the target as `isTarget[a.Path]`, and **`a.Path` is the registered
GUEST path** (`/mnt/felhom-drives/<name>`), not the agent's host `MountPath`. I had keyed the map on
`MountPath` alone, so the backup-target branch was **unreachable** — every absent drive, the target
included, would have fallen through to the generic event.
**The backup had not failed:** the PVE task returned OK and the PBS snapshot `2026-07-28T19:19:45Z`
is 6,264,034,053 B. The spurious failure **reached no channel** — the agent died with its in-memory
record and R-84 ground truth superseded it; the controller's event trail for the window shows only
the real 403 failure and its recovery. The design absorbed it; the mistake was still a mistake, and
on a slower tier the same slip could have aborted a multi-hour WAN upload.
It would have looked wired, passed its own unit tests, and shipped silently wrong on exactly the
drive it exists for. Caught by tracing `a.Path` back to its source instead of assuming. Fixed in
v0.184.1 (both keys, mirroring `planDriveGates`), red-proofed, and the **0.184.0 image is superseded
and was never deployed**.
Also recorded: `pgrep -f vzdump` self-matches a polling script's own command line and is not a safe
in-flight check — use the PVE task list.
## 5. Tests and red-proofs
## State at close
| Red-proof | Observed failure |
|---|---|
| hub allowlist entry removed | `backup_target_absent must be in allowedEventTypes, or POST /event 400s…` |
| hub Hungarian copy shortened | `must state the consequence…, got: "Meghajtó hiányzik."` |
| C — auto-elevation on register | `registering drives assigned the backup target "/mnt/hdd_1"` |
| E-2c — eject guard removed | `eject of the backup-target drive SUCCEEDED (200)` |
| E-2c over-correction | the gate blocked `/mnt/spare` (a non-target) |
| keying reverted to MountPath | `not resolvable by its GUEST path — …the branch would never fire` |
Both boxes healthy. Primary and offsite tiers `due=false` with `age_state=known` on both; breakers
clear; agents `active`; no thrash and no spurious staleness. Target drives at 12 % used with SMART
`PASSED`. The old archives (18 GB demo-felhom, 5.2 GB demo-hp) are **left in place** as the rollback
and as the only evidence of what the previous configuration produced.
Every mutation was **verified to have landed** before the result was trusted. The first attempt
silently no-op'd (gofmt had realigned the map) and the test "passed" — a false proof.
**Outstanding:** full drive-loss recovery (physical access), and the agent half of the absent-drive
behaviour (needs a drive unmount that would break the guest bind on a remote box).
Controller notify tests observe the **wire** (a real `Notifier` against an `httptest` hub), not a
mock, because the failure class is "nothing arrives".
## 6. What is NOT done
- **Part 2** (installer Case A/B), **Part 3** (the offer + the agent-side move), **Part 4** (the
degraded Hungarian banner). Phase 0 inverted their emphasis: the installer has **no drive-enrollment
step**, so at install the common case is system-drive-only and **Part 3 is the primary path**.
- **E-2a** — the guarded wrapper. The agent holds neither `Datastore.Allocate` at `/storage` nor
`Permissions.Modify`, and its sudoers has no `pvesm`/`pveum`. Use the wrapper pattern; **do not
widen the agent's PVE role.**
- **`backup_target_absent` is not proven END-TO-END live.** Its components are: the controller emits
the right type/severity (proven at the wire), the hub allowlists it (unit-pinned, image deployed),
and the keying is pinned. Proving the whole chain needs a live drive to actually go absent, which
means unmounting an enrolled drive carrying customer data and the guest bind on a remote box —
the same thing declined in E-1, for the same reason.
- Full drive-loss recovery (physical access) — carried forward from E-1.
## 7. Observations
- **`peti-felhom` untouched.** Its whole-guest backup still shares a device with its guest, so a
drive failure there is offsite-only recovery. Accepted until the operator's reinstall; re-evaluate
if that slips past ~2026-09-01.
- The fleet was briefly split (demo-hp updated first). Left that way it would have protected one of
two identically-exposed boxes, so demo-felhom was rolled to the same proven artifacts.
- `StorageTarget.Role` carries a documented vocabulary including `vzdump-target` and is always `""`.
Checked — it is a deliberate unpopulated field ("hub-owned, slice 10"), **not** a sixth dead seam.