Files
felhom.eu/REPORT.md
T
admin f1df01b5c5 docs(spike): stable intermediary mount — swap drive underneath (C1-immune, no-restart)
Live-validated on demo guest 9201 (felhom-flash throwaway): shared/slave mount
propagation delivers host->guest drive swaps live with no restart and no pct
change; bind source = permanent stable dir => inherently C1-immune (clean boot
with drive absent). nsenter fallback non-viable on unprivileged LXC; ro-tmpfs
placeholder is the CAP_DAC_OVERRIDE-proof fail-close. Report-only, no code change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 15:26:37 +02:00

12 KiB
Raw Blame History

felhom.eu — task reports

Overwrite this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in hub/CHANGELOG.md.


REPORT — SPIKE: stable intermediary mount (swap the backing drive underneath) (2026-06-15)

Type

Architecture spike — report-only + live throwaway experiments on felhom-flash (/dev/sdc1) on demo guest 9201 (felhom-pve, PVE 9.2.2, unprivileged LXC). No production code, no version bump. Full doc: documentation/audits/SPIKE-intermediary-mount-2026-06-15.md. Decides the bigger architecture behind the C1 critical bug surfaced in the B3 run below; the self-heal hook ships separately as the immediate C1 net.

Result

Viable and strictly better than today's per-drive pct set -mpN model. Bind a permanent stable host dir into the guest once (one provision-time reboot), then mount/unmount the real drive underneath it host-side — changes reach the running guest live, no restart, no pct change — via shared(host)/slave(guest) mount propagation that PVE/LXC delivers automatically. Proven both directions (mount appears / unmount disappears), both bind forms.

Inherently C1-immune (headline): rebooted guest 9201 with the drive absent → clean boot in 12 s, because the mp bind source is the always-present stable dir, not <drive>/felhom-data (which vanishes today → guest brick). End-state: 9201 restored to original config (mp0/1/2/9), 25 containers healthy; all experiment mounts/dirs/artifacts removed; felhom-usb/sda/PBS/local(-lvm) never touched.

Findings / decisions

  • Q1 propagation: live no-restart, both forms; guest receives binds as shared,slave (no custom lxc.mount.entry flags needed).
  • Q2 nsenter fallback: not viable on the unprivileged LXC (mount: must be superuser; block device absent from guest /dev). Use propagation.
  • Q3 idmap: guest uid/gid 1000 → host 101000 read/write unchanged; no regression.
  • Q4 fail-closed: naturally fail-closed (bare stable dir host-root-owned/unmapped → guest can't write). chmod 0000 is reliable only on a host-root-owned dir — container-root's CAP_DAC_OVERRIDE bypasses 0000 on a mapped-owned dir (proven). ro tmpfs placeholder = the capability-proof EROFS fail-close, auto-shadowed by the drive on reconnect, auto-re-exposed on eject. App-stop stays primary.
  • Q5 C1 immunity: reboot-with-drive-absent boots clean.
  • Q6 round-trip: eject/reconnect = pure host mount/umount, no restart; reboot mid-eject = no brick.
  • Q7/Q8 fit: parent-bind makes enroll/plug/eject/reconnect all restart-free; only whole-guest decommission touches the bind. Migration/FileBrowser cleaner. Confinement correction: bind only <raw>/felhom-data onto the stable path (proven — customer's dump/images/private/lost+found stay out of the guest), not the raw whole-drive mount.

Fork picks (for the build spec)

parent-bind form · propagation (not nsenter) · ro-placeholder fail-close (+ host-root-owned dir as 2nd layer) · HDD_PATH kept identical via in-guest bind/symlink or repoint (decide) · agent mounts <raw>/felhom-data onto the stable path. See the doc's FORK LIST + current→intermediary migration sketch.


REPORT — B3: live unattended storage-lifecycle exercise + feature sweep (2026-06-15)

Type

Live, unattended, full-UI exercise of migration / decommission / re-enroll / eject + a broad feature sweep on shipped controller v0.66.2 / agent v0.32.0 (demo guest 9201). No production code changed. Full doc: documentation/audits/B3-storage-lifecycle-2026-06-15.md.

Result

Migration engine is solid (every move checksum-verified identical, sources cleaned, HDD_PATH flipped, apps healthy). The decommission/eject/re-enroll lifecycle has serious gaps, incl. one critical guest-bricking bug and three unrouted endpoints the UI calls. End-state is sane: all apps healthy on felhom-flash (now default), both data drives Aktív, sda/PBS/controller untouched.

UI-driven per A4 where the dashboard exposes the action (per-app migrate + migrate-all driven via the real buttons, native confirm() auto-accepted); operations the UI lacks a button for / calls an unrouted endpoint for were driven via the controller's own JSON API in-guest (correct Host header, bypassing the body-mangling proxy) — each gap documented as a finding.

Findings (→ next-slice fixes)

  • 🔴 C1 (critical): decommission/eject never pct set --delete mpN, so the stale LXC bind bricks the guest on next reboot (pre-start hook exit 255, ALL apps down). Recovered live via manual pct set --delete mp1.
  • 🟠 H1: /api/storage/disconnect, /api/storage/reconnect, /api/storage/restart-apps are 404 (unrouted) though settings.js calls them.
  • 🟠 H2: no UI button for /api/storage/decommission (B2b flow unreachable from the dashboard).
  • 🟠 H3: no one-click re-enroll for a decommissioned drive (only "Eltávolítás").
  • 🟡 M1: decommissioning the default drive doesn't reassign the default (orphaned, no failover).
  • 🟡 M2: eject does NOT auto-remount on reconnect — intent-gated-skipped like decommission (premise contradicted); difference is bind-record retention + soft-marker.
  • 🟡 M3: userdata/import/calibre created 755 at enroll (skeleton gap) and migrate-all doesn't re-assert setgid on pre-existing dirs (16/17 dirs preserved 2775).
  • 🟡 M4: felhom-usb's systemd mount unit had a stale UUID (couldn't auto-mount); fixed in restore. M6: test-email fails (no SMTP).

Passes

Per-app migration, migrate-all (7 apps, checksums identical, setgid 16/17), decommission state (intent/unmount/prune/Kiváltva), no-auto-rebind invariant (intent-gated), missing-storage badges (8×), re-enroll marker-clear, encrypted backup→PBS (snapshot landed), FileBrowser umask (PID1 0002 → 2775 group-writable + collab write), fresh-app deploy+lifecycle+logs, monitoring page.


REPORT — SPIKE: self-serve raw-drive plug-and-enroll (GATE) (2026-06-15)

Type

GATE spike — report-only + one live throwaway experiment on /dev/sdc. No production code, no version bump, no service restart. Build spec is gated on the fork picks below. Full doc: documentation/audits/SPIKE-raw-drive-enroll-2026-06-15.md.

One-screen summary

The plug-and-enroll flow is mostly assemblable from primitives that already exist. The genuine net-new is (a) a raw-disk discovery source — the agent's Observe and localapi GET /disks are both built only from PVE storages (ListStorage/NodeStorage), so a disk that isn't already a PVE storage is invisible — and (b) a create-storage call (the agent creates no PVE storages today). The safety classifier already exists (RoleForRawDevice, SystemDisks, RoleForStorage in internal/storage/role.go) and correctly protects the live system disk, but needs two hardening additions before it can drive a discovery list.

Validated live on felhom-pve:

  • sda (ESP + sole LVM PV pve VG holding root, local-lvm, guest 9201's disks) → PROTECTED (system); pinned via /boot/efi=sda2.
  • sdb (felhom-usb, ext4 /mnt/felhom-usb) → PROTECTED (enrolled data).
  • sdc (removable junk iso9660+EFI+hfsplus) → OFFERABLE.
  • Premise correction: felhom-pbs is a remote pbs: storage (server 192.168.0.180) — no local PBS disk exists; it's protected at the storage tier, not a raw-disk tier.

Live FORMAT experiment (serial-fenced to sdc): pvesh create …/disks/directory --name felhom-flash --device /dev/sdc --filesystem ext4 --add_storage 1 did partition+mkfs+mount+register in one call (fs-UUID mount unit, is_mountpoint 1, /mnt/pve/felhom-flash). It immediately appeared in the controller's agent-drive view as user-data / data-bearing / "Nem regisztrált" with a one-click Regisztrálás — no agent restart. sdc left as the clean enrolled felhom-flash B3 migration target.

The two safety-gate hardening gaps (must close)

  1. Whole-disk-granularity enrolled-subtractiondeviceRole matches a registered storage by exact BackingDevice==device; a discovery flow asks about whole disk /dev/sdb while felhom-usb backs /dev/sdb1, so the match fails and sdb would be offered for format. Compare at whole-disk granularity.
  2. LVM-PV / guest-image membershipSystemDisks resolves the system disk from the mount table only; a non-boot disk holding local-lvm PVs / guest images would be missed. Add a disks/lvm PV-membership check.

Fork list (operator decides)

  1. Discovery source: disks/list+disks/lvm+mount-table+Observe (recommended) ± lsblk for display.
  2. Safety signals: mount-table-only (ships, but can offer a non-boot LVM disk) vs hardened (the 2 checks above — recommended).
  3. Mount path: PVE default /mnt/pve/<name> vs /mnt/<name> via EnsureMount+pvesm add dir (recommended, matches felhom-usb).
  4. Format-vs-enroll-as-is: single supported-fs partition → as-is; else format; format on ambiguity.
  5. Endpoints: GET /disks/raw + POST /disks/init + POST /disks/attach-existing, withGuest-scoped, agent re-derives the gate.
  6. v1 scope: format-only first (proven live) + enroll-as-is fast-follow (recommended) vs both in v1.
  7. Normalize content (vs PVE all-types default); typed proxmox-client create method (vs host-side pvesh).

REPORT — Slice 10D (hub half): DR capstone — recovery mode + re-enroll + directive serving (hub v0.11.0) (2026-06-10)

Type

TASK (CC-implemented). The hub half of the slice-10 DR capstone (closes slice 10). Pairs with felhom-agent v0.18.0 (identity escrow + restore-mode consumption).

What changed (hub)

The hub ORCHESTRATES recovery but holds no usable secret and no Cloudflare write-power — a compromised hub can at most hand out opaque blobs (they need R, which the hub never has) + rotate its own per-host credential. It cannot hijack a customer's tunnel (the destructive rotation is the operator's job).

API

  • PUT/DELETE /admin/hosts/{id}/recovery-mode (global key) — arm/disable recovery mode with a bounded TTL (clamped [60s, 4h], default 30m → auto-expires). Directive + re-enroll are served ONLY while active.
  • POST /hosts/{id}/re-enroll — gated ONLY on recovery mode (the lost box has no old key). Rotates the host's API key to the new box's key (old box revoked) + returns the directive + opaque blobs.
  • GET /hosts/{id}/restore-directive (re-enrolled key, recovery-gated) — re-fetch.
  • The slice-7 escrow upload now also accepts the identity blob + non-secret directive (additive).

Store

  • hosts.recovery_mode_until; host_escrow.identity_blob + directive_json. Methods: SetRecoveryMode/ClearRecoveryMode, RotateHostAPIKey, SaveHostDRBundle/GetHostDRBundle.

Tests (green)

  • re-enroll refused without recovery mode (403); recovery-arm is global-key-only; re-enroll rotates + revokes (old key→401, new key→200); directive served only in recovery mode + expires; clear disables re-enroll.

Docs

  • Doc 03 §9 (10D done → SLICE 10 CLOSED) + the host-loss DR flow with the operator-side rotation model (hub orchestrates + read-only verifies; the operator deletes the stale connector + rotates the tunnel/PBS token from a trusted environment).

Deferred (non-blocking, per the locked model)

  • The Config DR/Recovery web UI (functional today via the recovery-mode admin API) + a small operator rotation CLI. No Cloudflare write-credential is in the hub by design.

Pending

  • Build + deploy hub v0.11.0 + agent v0.18.0; run the operator-in-the-loop DR drill (throwaway identity).