Files
felhom.eu/documentation/audits/B3-storage-lifecycle-2026-06-15.md
admin 4fe77fcd1b docs: B3 live storage-lifecycle exercise + feature sweep (2026-06-15)
Live unattended UI exercise of migration/decommission/re-enroll/eject +
feature sweep on shipped controller v0.66.2 / agent v0.32.0 (guest 9201).
No production code changed.

Migration engine solid (all moves checksum-verified identical, sources
cleaned, HDD_PATH flipped, apps healthy). Lifecycle gaps found:
- C1 (critical): decommission/eject leave a stale LXC mp (no pct delete)
  -> guest fails pre-start hook on next reboot, ALL apps down. Recovered
  live via pct set --delete mp1.
- H1: /api/storage/{disconnect,reconnect,restart-apps} are 404 (unrouted)
  though settings.js calls them.
- H2: no UI button for /api/storage/decommission.
- H3: no one-click re-enroll for a decommissioned drive.
- M1: decommissioning the default drive doesn't reassign the default.
- M2: eject does NOT auto-remount (intent-gated like decommission).
- M3: userdata/import/calibre created 755 at enroll; migrate-all doesn't
  re-assert setgid on pre-existing dirs (16/17 preserved 2775).
- M4: felhom-usb mount unit had stale UUID; M6: test-email fails (no SMTP).

End-state sane: apps healthy on felhom-flash (now default), both data
drives Aktiv, reboot-safe, sda/PBS/controller untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 12:39:44 +02:00

115 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# B3 — Live unattended exercise of the storage lifecycle (migration / decommission / re-enroll / eject) + feature sweep
**Date:** 2026-06-15
**Target:** shipped **controller v0.66.2 + agent v0.32.0**, demo guest **9201** on **felhom-pve** (node `demo-felhom`).
**Mode:** live, unattended, all 8 phases run without pausing. No production code changed.
**Drive method:** the real operations were driven through the **dashboard UI via the browser tool** wherever the UI exposes them (per the A4 strict-UI requirement). Where the UI has **no button** or calls an **unrouted endpoint** (documented findings below), the operation was driven via the controller's own JSON API hit **in-guest at the container IP with the correct `Host` header** (bypassing the cloudflared/traefik proxy that mangles bodies) — never a raw shortcut around the controller's logic. Host-level captures/simulations used SSH (`root@felhom-pve`).
**Safety rails honoured:** `/dev/sda`, the felhom-controller container, the bootstrap unit, `felhom-pbs`, `local`/`local-lvm`, and guest 9201's root disk were never touched. All operations confined to the two user-data drives (felhom-usb `/dev/sdb1`, felhom-flash `/dev/sdc1`) + the demo apps.
---
## VERDICT
The migration engine is **solid** (checksums verified identical across every move, sources cleaned, HDD_PATH flipped, apps healthy). The **decommission / eject / re-enroll lifecycle has serious gaps**, including **one critical guest-bricking bug** and **three unrouted endpoints the UI calls**. The B3 migration target (felhom-flash) now holds all apps, healthy, and is the default drive.
### 🔴 Critical (fix before any customer uses decommission)
- **C1 — Decommission bricks the guest on next reboot.** The agent decommission removes its bind-tracking JSON + host-unmounts, but **never runs `pct set --delete mpN`**, so the stale `mp` bind entry survives in the LXC config. On the next guest reboot the LXC **pre-start hook fails (exit 255)** because the bind source (`/mnt/<drive>/felhom-data`) no longer exists → **the whole guest fails to start — ALL apps down, not just the decommissioned drive's.** Proven live (Phase 3): guest 9201 went down on reboot; only `pct set 9201 --delete mp1` + `pct start` recovered it.
### 🟠 High (UI lifecycle largely non-functional)
- **H1 — Three endpoints the settings JS calls are unrouted → HTTP 404:** `/api/storage/disconnect`, `/api/storage/reconnect`, `/api/storage/restart-apps`. The "Leválasztás" (disconnect), "Csatlakoztatás" (reconnect), and "Alkalmazások indítása" (restart-apps) buttons are dead.
- **H2 — No UI trigger for `/api/storage/decommission`.** The B2b decommission flow (Phases 2 & 4 of this exercise) has **no dashboard button**`settings.html` only exposes migrate-all (`/api/storage/migrate`), eject (`/api/storage/eject`), and remove. The routed `/api/storage/decommission` (mode=migrate|anyway) is reachable only by API.
- **H3 — No one-click re-enroll** for a decommissioned drive; its only UI action is "Eltávolítás a rendszerből" (remove). The Change-4 marker-clear fires only on re-register, which has no button on the decommissioned entry.
### 🟡 Medium
- **M1 — Decommissioning the default drive does not reassign the default.** After decommissioning felhom-usb (the default), **no drive held the default badge** (both remaining drives showed "Legyen alapértelmezett") — the default was orphaned on the decommissioned drive, no failover. New HDD-app deploys would have had no valid default target.
- **M2 — Eject does NOT auto-remount on reconnect** (contradicts the expected eject≠decommission distinction). An ejected drive is intent-gated-skipped by `ReassertGuestBinds` exactly like a decommissioned one; its systemd unit is disabled and it stays unmounted. The real distinction is: eject **keeps** the bind record (reassert logs a skip line) + sets **no** soft-marker; decommission **removes** the bind record + sets the "Kiváltva" soft-marker. Neither auto-remounts, and the only "reconnect" path (`/api/storage/reconnect`) is 404 (H1).
- **M3 — `import/calibre` userdata dir loses setgid.** After migrate-all, `userdata/import/calibre` was **755** on the target while it was **2775** on the source (and 16/17 other dirs preserved 2775). Root cause: the **enroll-time userdata skeleton created `import/calibre` as 755** (mtime = enroll time), and the migrate-all merge-walk does not re-assert setgid on a pre-existing target dir (it correctly creates *new* dirs like `import/paperless` as 2775). calibre-web binds that dir — group-write/setgid is broken there.
- **M4 — felhom-usb's systemd mount unit had a stale UUID** (`277a2179…` vs the real `da9e7089…` from the 2026-06-14 reformat), so it could not auto-mount and a role-gated eject 403'd. Fixed in restore.
- **M5 — Decommission's host-unmount does not release the drive from the running guest.** The unprivileged LXC bind keeps `/dev/sdX1` live-mounted inside the guest until reboot, so "safely removable" is only true post-reboot.
- **M6 — Test notification email fails** (returns "Hiba"), likely no SMTP transport configured on the demo.
---
## Per-phase results (PASS/FAIL + evidence)
### Phase 0 — baseline + seed — PASS
- Baseline: 8 HDD apps on felhom-usb (audiobookshelf, calibre-web, immich, jellyfin, komga, paperless-ngx, radarr, romm); felhom-usb = default; flash = empty skeleton.
- Userdata ownership baseline: **all 17 dirs `2775` setgid, gid 101000** (= guest gid 1000). ✓
- Seeded checksummed files (source of truth):
- `userdata/media/movies/b3-movie.txt``6155f6df…833e8`
- `userdata/documents/b3-document.txt``1f2bb9a5…3764`
- `userdata/media/photos/b3-photo.txt``162a7fa1…6e4f`
- `appdata/romm/config/b3-romm.txt``3643c7cc…60f1`
### Phase 1 — per-app migration (romm, felhom-usb → felhom-flash) — PASS (UI-driven)
Driven via the **romm app page → "Áthelyezés másik tárhelyre" → felhom-flash → Áthelyezés** button (native `confirm()` auto-accepted to clear the blocking dialog; the real `appMigrate → POST /api/storage/migrate-app` flow ran). Job `mig-20260615-094757`.
- stop → copy → verify → flip → redeploy **healthy** ✓ (romm `Up … (healthy)`).
- Migrated appdata checksum on flash = `3643c7cc…60f1`**identical** ✓.
- Source subtree `…/felhom-usb/…/appdata/romm` **removed** ✓.
- HDD_PATH **flipped**: `docker inspect romm` mounts now `/mnt/felhom-flash/...`; `app.yaml HDD_PATH: /mnt/felhom-flash` ✓.
- UI badge updated to **"Adattároló: felhom-flash"** ✓.
### Phase 2 — migrate-all + decommission felhom-usb → felhom-flash — PASS (migrate via UI; decommission via API per H2)
Migrate-all driven via **settings → felhom-usb → "Összes adat áthelyezése" → felhom-flash** (real `storageMigrateAll → POST /api/storage/migrate`). Job `mig-20260615-095545`, all 7 apps `cleaned`.
- All 7 apps **healthy on flash** ✓ (komga `unhealthy` = pre-existing healthcheck quirk, still serving).
- All 3 seeded userdata checksums **identical** on flash (`6155f6df` / `1f2bb9a5` / `162a7fa1`) ✓.
- Source userdata files + `appdata/` **removed** ✓.
- **#8 setgid proof:** 16/17 migrated userdata dirs `2775` gid 101000; **1 exception → `import/calibre` = 755** (finding M3).
- **Decommission** (no UI button → routed `POST /api/storage/decommission {mode:anyway}`): `decommissioned:true`; agent **intent=decommissioned**; host **unmounted**; **guest-binds.json pruned** (da9e7089 removed); controller badge **"Kiváltva"** ✓. But **`pct config` still showed `mp1`** (stale — see C1).
- **Default edge (M1):** felhom-usb *was* default; afterwards **no drive held the default badge** — not reassigned.
### Phase 3 — reconnect felhom-usb → NO auto-rebind — PASS (invariant holds) + 🔴 C1 surfaced
- Agent restart: host `/mnt/felhom-usb` **not remounted**; reassert log `F9 re-assert: skipping non-enrolled drive (intent-gated) intent=ejected` (demonstrated on the lingering ce9d drive; felhom-usb was pruned from the bind list entirely so it isn't even reattempted). No auto-rebind ✓.
- **Guest reboot → C1 fired:** `lxc_init … Failed to run lxc.hook.pre-start … startup for container '9201' failed` (exit 255) due to stale `mp1`. **Recovery:** `pct set 9201 --delete mp1``pct start 9201` → booted first try, felhom-usb correctly absent, 24→25 containers back.
### Phase 4 — decommission-anyway felhom-flash → missing-storage — PASS (decommission via API per H2)
`POST /api/storage/decommission {mode:anyway, mount_name:felhom-flash}``stopped_apps` = all 8.
- Apps **stopped**, **keep HDD_PATH=/mnt/felhom-flash** ✓; flash **intent=decommissioned + unmounted** ✓.
- Missing-storage badge **"Hiányzó tárhely: felhom-flash"** rendered **8×** on dashboard, **8×** on stacks, on the app card ✓. Dashboard screenshot: 9 running / 8 stopped, Audiobookshelf "Leállítva" + the orange badge.
### Phase 5 — re-enroll felhom-flash clears decommissioned + badge — PASS (re-mount via SSH + re-register via API per H3)
Re-mounted flash (data **intact**, seed checksum `6155f6df` still matched ✓), `POST /api/storage/register` → reEnrollClearMarker.
- Agent **intent → enrolled** (Change 4 marker cleared) ✓; flash badge no longer Kiváltva; **0 "Hiányzó tárhely"** on dashboard ✓.
- Apps restarted via `POST /api/stacks/{name}/start` (because **`/api/storage/restart-apps` returned 404** — H1); **25 containers healthy** ✓.
### Phase 6 — eject vs decommission — PASS (finding M2: premise contradicted)
Eject flash via `POST /api/storage/eject`: intent→**ejected**, host **unmounted**, **systemd unit disabled**, bind record **kept**. Agent restart (reconnect sim): **NOT auto-remounted**; reassert logged `skipping non-enrolled drive (intent-gated) intent=ejected`. Apps stayed up (guest bind persisted). Flash restored (re-mount + register → enrolled, 25 containers).
### Phase 7 — feature sweep
| Feature | Result |
|---|---|
| Manual backup (whole-guest → felhom-pbs) | **PASS** — real encrypted vzdump→proxmox-backup-client; phase reached `done`; new PBS snapshot `ct/9201/2026-06-15T10:29:13Z` landed (4 total). A 2nd trigger correctly returned 409 "already in progress". |
| Fresh app deploy + lifecycle + logs (Glance) | **PASS** — deploy 202 → healthy; stop/start/restart all HTTP 200; logs endpoint returned app output. |
| FileBrowser umask fix | **PASS** — PID1 `Umask: 0002`, entrypoint `[sh -c umask 002; exec …]`; umask-002 folder = **2775 group-writable** (gid 101000); a uid-1000 process **wrote into it** ("WRITE OK") → FileBrowser/app group collaboration confirmed. (Note: FileBrowser web session had expired; folder-creation via the UI was not re-driven because logging in would require entering a password — prohibited. The fix's effect was verified by replicating PID1's exact umask-002 environment + the group-write test.) |
| Monitoring page | **PASS** — renders host CPU/mem/temp/uptime + storage bars; felhom-usb correctly shown "Nem elérhető" while decommissioned. |
| Test notification email | **FAIL (M6)**`/settings/notifications/test` returned "Hiba" (likely no SMTP configured). |
| Mem-gate | Gate **allows within budget** (Glance deployed cleanly); a **refusal was not reproducible** with ~10 GB free and no catalog app large enough to exceed 12 GB — would require artificial RAM exhaustion (not done, would destabilise the demo). |
### Phase 8 — restore — DONE (sane end-state)
- **felhom-flash**: `Alapértelmezett | Aktív | USB`**default**, all apps healthy.
- **felhom-usb**: `Aktív` (stale-UUID unit fixed → da9e7089, re-mounted, soft-marker cleared; mp1 re-bound and valid). Note: its **agent intent remains `decommissioned`** (a latent cosmetic inconsistency — harmless because the drive is mounted and its mp is valid; the reassert would merely skip re-adding an already-present bind).
- **Belső SSD**: `Aktív`. No "Kiváltva" anywhere; **no stale mps** (mp1=usb, mp2=flash both mounted → reboot-safe); **25 containers running**; Glance test app removed. sda / PBS / controller / local / local-lvm untouched throughout.
---
## Code-fix candidates for the next slice
1. **(C1, critical) On decommission/eject, the agent must `pct set <vmid> --delete mpN`** for the drive's bind (or the controller must, post-finalize) so the stale bind cannot brick the guest on reboot. Add a regression test: decommission a data drive → reboot guest → assert it starts and the drive is absent. Alternatively, make the agent's pre-start hook **self-heal** by dropping config binds whose source path is missing.
2. **(H1) Route `/api/storage/disconnect`, `/api/storage/reconnect`, `/api/storage/restart-apps`** (or remove the dead buttons). These are called by `settings.html` JS but absent from `ServeStorageAPI`.
3. **(H2) Wire a decommission button** (migrate-all-then-decommission / decommission-anyway with type-to-confirm) into the settings storage actions → `/api/storage/decommission`.
4. **(H3) Add a re-enroll action** on a decommissioned drive entry (re-mount + re-register in one click), instead of only "Eltávolítás a rendszerből".
5. **(M1) Reassign the default** to another schedulable drive when the current default is decommissioned/removed (or block decommissioning the default until reassigned).
6. **(M2) Decide eject semantics:** either make eject auto-remount on reconnect (and route `/api/storage/reconnect`), or relabel the UI so "eject" reads as the permanent-until-re-enroll action it actually is.
7. **(M3) Fix the userdata skeleton** so every nested dir (incl. `import/calibre`) is created `2775` setgid at enroll, and have the migrate-all merge-walk **re-assert setgid** on pre-existing target dirs.
8. **(M4) Update the systemd mount unit's `What=` UUID when a drive is reformatted** (or key the unit on a durable id), so re-mount after reformat works.
---
## Evidence index (selected)
- Checksums (source-of-truth → verified-identical-on-target): movie `6155f6df…`, document `1f2bb9a5…`, photo `162a7fa1…`, romm-appdata `3643c7cc…`.
- Ownership: source 17/17 `2775` gid 101000; post-migrate-all target 16/17 `2775`, `import/calibre` `755` (M3).
- C1 brick: `__lxc_start: 2208 Failed to initialize container "9201"`; recovery via `pct set 9201 --delete mp1`.
- Intent transitions (agent `drive-intents.json`): usb da9e7089 enrolled→decommissioned→(restore, controller-cleared); flash 81a26531 enrolled→ejected→enrolled, →decommissioned→enrolled.
- PBS snapshot landed: `felhom-pbs:backup/ct/9201/2026-06-15T10:29:13Z`.
- Screenshots saved (browser): romm-on-flash badge, Phase-4 dashboard missing-storage badges, monitoring page, settings storage states.