Files
felhom.eu/REPORT.md
T
admin 3696188636 E-2 increment 1: report + close E-2b/E-2c as shipped and proven live
hub 0.81.0, agent 0.112.0, controller 0.184.1 live on BOTH demo boxes.

E-2c: eject/decommission of the backup-target drive refused 409 on both boxes,
drives unmoved. E-2b: the never-called disconnect seam is wired, with the target
case raising the specific backup_target_absent.

Records the keying bug caught before deploy (a.Path is the GUEST path, so the
target branch was unreachable -- 0.184.0 superseded, never deployed) and states
plainly that backup_target_absent is NOT proven end-to-end live: proving it needs
a live enrolled drive to go absent.

Parts 2/3/4 and E-2a remain open; Peti risk stays parked.
2026-07-29 08:34:41 +02:00

6.4 KiB

REPORT — E-2 increment 1: the backup target is now defended and audible (2026-07-29)

Overwritten per the standing rule. E-2 is PARTIALLY shipped — this increment is Parts 1 + 5 plus findings E-2b/E-2c. Parts 2, 3, 4 and E-2a remain open (see §6).

Live fleet after this run: hub 0.81.0, agent 0.112.0 (both boxes), controller 0.184.1 (both boxes). peti-felhom deliberately untouched.


1. Why this order

The operator reordered the plan, correctly: E-2b and E-2c protect a configuration that exists right now; E-2a builds for boxes that do not. Both demo boxes are configured, Peti is parked, and no new installs are happening this week — so Part 3 has no user yet, while E-1's new configuration had a drive that could be ejected without warning and no event when it went absent.

2. E-2c — the backup drive could be ejected out from under the backup

RoleForStorage types a local-dir on a non-system device as user-data, so the pre-existing eject role gate PASSED the drive E-1 had just made the vzdump target. POST /disks/eject on /mnt/nvme-1tb or /mnt/hdd_1 would have succeeded silently, taking the only local whole-guest backup with it while the box kept reporting a configured tier.

Fixed in agent v0.111.0/0.112.0: handleDiskEject and handleDiskDecommission refuse 409 when the mount backs a configured tier, naming the storage and the remedy — the op is ordered, not forbidden.

Deliberately NOT a role reclassification. Making RoleForStorage return RoleBackup would refuse every legitimate eject of the customer's own data drive, because on both boxes that drive is the target. TestEjectStillAllowedOnANonTargetDrive pins that, and the over-correction is red-proofed.

Live on both boxes:

demo-hp      felhom-backup mount=/mnt/nvme-1tb role=user-data backup_target=True
demo-felhom  felhom-backup mount=/mnt/hdd_1    role=user-data backup_target=True

That single line is the exposure: role=user-data (which the gate permits) yet backup_target=True. Both ejects returned 409 with the full message; both decommissions 409; both drives still mounted with dump/ intact — the refusal had no side effect. /var/lib/vz is still refused by the pre-existing role gate (role: system), proving the new gate is narrow.

3. E-2b — the alarm that was defined and never called

NotifyStorageDisconnected and NotifyStorageReconnected were registered in allowedEventTypes, in DefaultEnabledEvents, and given Hungarian hub copy — and called from nowhere. A drive going absent produced stopped apps, a [WARN] log and a UI badge, then silence on every channel. A drive that is only a backup target has no apps to stop, so it was silent twice over. Fifth instance of this class; found by Phase 0, not by a failure.

Now wired in ReconcileDriveGates, with the target case raising the more specific backup_target_absent (error) and recovering as backup_target_restored (info — the pairing-gated pattern; severityNotifies NOT widened). Never both: two mails for one event trains people to ignore the channel.

4. The bug I introduced and caught before deploy

ReconcileDriveGates resolves the target as isTarget[a.Path], and a.Path is the registered GUEST path (/mnt/felhom-drives/<name>), not the agent's host MountPath. I had keyed the map on MountPath alone, so the backup-target branch was unreachable — every absent drive, the target included, would have fallen through to the generic event.

It would have looked wired, passed its own unit tests, and shipped silently wrong on exactly the drive it exists for. Caught by tracing a.Path back to its source instead of assuming. Fixed in v0.184.1 (both keys, mirroring planDriveGates), red-proofed, and the 0.184.0 image is superseded and was never deployed.

5. Tests and red-proofs

Red-proof Observed failure
hub allowlist entry removed backup_target_absent must be in allowedEventTypes, or POST /event 400s…
hub Hungarian copy shortened must state the consequence…, got: "Meghajtó hiányzik."
C — auto-elevation on register registering drives assigned the backup target "/mnt/hdd_1"
E-2c — eject guard removed eject of the backup-target drive SUCCEEDED (200)
E-2c over-correction the gate blocked /mnt/spare (a non-target)
keying reverted to MountPath not resolvable by its GUEST path — …the branch would never fire

Every mutation was verified to have landed before the result was trusted. The first attempt silently no-op'd (gofmt had realigned the map) and the test "passed" — a false proof.

Controller notify tests observe the wire (a real Notifier against an httptest hub), not a mock, because the failure class is "nothing arrives".

6. What is NOT done

  • Part 2 (installer Case A/B), Part 3 (the offer + the agent-side move), Part 4 (the degraded Hungarian banner). Phase 0 inverted their emphasis: the installer has no drive-enrollment step, so at install the common case is system-drive-only and Part 3 is the primary path.
  • E-2a — the guarded wrapper. The agent holds neither Datastore.Allocate at /storage nor Permissions.Modify, and its sudoers has no pvesm/pveum. Use the wrapper pattern; do not widen the agent's PVE role.
  • backup_target_absent is not proven END-TO-END live. Its components are: the controller emits the right type/severity (proven at the wire), the hub allowlists it (unit-pinned, image deployed), and the keying is pinned. Proving the whole chain needs a live drive to actually go absent, which means unmounting an enrolled drive carrying customer data and the guest bind on a remote box — the same thing declined in E-1, for the same reason.
  • Full drive-loss recovery (physical access) — carried forward from E-1.

7. Observations

  • peti-felhom untouched. Its whole-guest backup still shares a device with its guest, so a drive failure there is offsite-only recovery. Accepted until the operator's reinstall; re-evaluate if that slips past ~2026-09-01.
  • The fleet was briefly split (demo-hp updated first). Left that way it would have protected one of two identically-exposed boxes, so demo-felhom was rolled to the same proven artifacts.
  • StorageTarget.Role carries a documented vocabulary including vzdump-target and is always "". Checked — it is a deliberate unpopulated field ("hub-owned, slice 10"), not a sixth dead seam.