Demo restored to healthy baseline (test apps removed, drives at stable, sys_drive separate, settings UI clean). Deliverable summary: 4 findings (F3 onboot=0 guest no-auto-start [highest impact], F2 FileBrowser bounce on controller restart, F1 restic dead-weight, F4 stable double-bind), all report-not-fixed. Method per test, PBS floor (unused), end state, coverage gaps (1.3/1.4/1.8 loop-whitelist, 4.7 operator-gated). No data-loss/path-escape/corruption found. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017PsnU2ASocYrvzqE82YDYW
36 KiB
TEST REPORT — NO-MERCY CAMPAIGN: stable-path model + /mnt/sys_drive + server restart
Date: 2026-06-23 · Class: destructive test campaign (operator-authorized, DEMO only)
Targets: guest 9201 (demo-felhom) on felhom-pve (192.168.0.162, N100, PVE 9) ·
controller v0.81.0 · agent v0.40.0 · hub hub.felhom.eu
Rule: report-don't-fix. No product-code changes. Findings get evidence + file:line + a proposed
fix direction (NOT implemented).
Status legend
PASS = effect asserted AND wrong-outcome ruled out · FAIL = wrong outcome observed · INCONCLUSIVE = could not assert cleanly (reason given).
Phase 0 — floor + baseline (non-destructive) — ✅ COMPLETE
PBS floor (hard restore floor)
- Snapshot id:
felhom-pbs:backup/ct/9201/2026-06-23T17:28:25Z(PBS datastorefelhom-spike@felhom@pbs!n100@192.168.0.180). Backup finished successfully (00:00:15, incremental, reused 89.9%). - Used? Not yet (recorded for restore if a later phase wedges 9201).
- Note:
felhom-pbsstorage resolves to the PBS server on 192.168.0.180 datastorefelhom-spike. Writing the floor snapshot there is normal PBS operation (the demo's nightly job already does) — it does not mutate/break 180 as a host, consistent with the guardrail.
Baseline — HOST (pct config 9201)
rootfs: local-lvm:vm-9201-disk-0,size=32G·mp0: vm-9201-disk-1,mp=/var/lib/docker,backup=1,size=200Gmp1: vm-9201-disk-2,mp=/mnt/sys_drive,backup=1,size=50G·mp8: /mnt/felhom-drives(bind) ·mp9: …/bootstrap → /etc/felhom-bootstrap,ro=1onboot: 0⚠️ — flagged for Phase 4.1 (the spec asserts 9201 auto-starts on host reboot; withonboot: 0a plain host reboot will not auto-start it). Pre-existing config.- Drives (host
lsblk -f):felhom-usb=/dev/sdb1UUIDda9e7089-…(916 GB HDD);felhom-flash=/dev/sdc1UUID81a26531-…(117 GB USB). Both mounted host-side at raw/mnt/<name>AND bind-mounted to stable/mnt/felhom-drives/<name>under the shared parent.
Baseline — GUEST
findmnt /mnt:/mnt/sys_drive=pve-vm--9201--disk--2(distinct device);/mnt/felhom-drives/{flash,usb}=sdc1[/felhom-data]/sdb1[/felhom-data].- Invariant (bug signature) — PASS: no rootfs placeholder dirs at raw
/mnt/felhom-flash//mnt/felhom-usbinside the guest (not present (good)for both).
Baseline — CONTAINER (felhom-controller)
settings.jsonstorage paths = ONLY stable/mnt/felhom-drives/felhom-flash+/mnt/felhom-drives/felhom-usb(no raw entries). Bothfindmntto the correctsdX1[/felhom-data]./mnt/sys_drive= disk-2.- Container
/= docker-data overlay (197G, the 200G mp0)./mnt/sys_driveseparate (50G).
Baseline — AGENT state (/var/lib/felhom-agent/)
guest-binds.json→9201:[uuid:81a26531 (flash), uuid:ce9d1d81 (absent), uuid:da9e7089 (usb)].drive-intents.json:277a2179→enrolled (absent),81a26531→enrolled (flash),ce9d1d81→ejected (absent),da9e7089→enrolled (usb). ⚠️ Two leftover intents for physically-absent drives (277a2179enrolled-absent,ce9d1d81ejected) — relevant to Phase 4.3 intent-gate.- Agent version:
felhom-agent 0.40.0.
Deployed apps at baseline
- Only base infra running:
filebrowser,felhom-controller,cloudflared,traefik(all healthy). No customer apps deployed — tests needing "an app on a drive" will deploy a small one.
Scratch
- Loop devices: LOOP_A=
/dev/loop0(/root/nomercy-scratch/loopA.img, 2G), LOOP_B=/dev/loop1(/root/nomercy-scratch/loopB.img, 2G). For teardown in Phase 5.
Phase 2.3 (opportunistic, captured during the floor vzdump) — ✅ PASS
- The floor
vzdump 9201log:including mount point mp1 ('/mnt/sys_drive') in backupand theproxmox-backup-clientinvocation carries--include-dev /mnt/vzsnap0/./mnt/sys_drive. mp1 is in the archive. Wrong outcome (mp1 excluded /backup=0) ruled out. (Optional restore-to-scratch deferred.)
Phase 1 — stable-vs-raw path model — ✅ COMPLETE (1.3/1.4/1.8 by-design inconclusive)
Method: mix of live-UI (claude-in-chrome on felhom.demo-felhom.eu) and exact-endpoint (curl to
the in-guest controller 127.0.0.1:8080). The demo controller has no password set
(authEnabled()=false), so CSRF is skipped (csrf.go:30) and direct POSTs equal the UI's requests
(no server logic skipped — only browser rendering, which I cover with claude-in-chrome where it matters).
1.1 — registry holds only stable paths — ✅ PASS (exact-endpoint + UI)
settings.jsonstorage_paths = ONLY/mnt/felhom-drives/{felhom-flash,felhom-usb}.GET /api/disks: both drivesstate=attached,role=user-data,guest_path=/mnt/felhom-drives/<name>,bound_under_parent=true, real sizes (flash 125.7 GB, usb 983 GB). No raw/mnt/<name>registry entry; no rootfs placeholder. Wrong outcome (raw entry / placeholder) ruled out.
1.2 — deregister → re-add via "Regisztrálás" — ✅ PASS (exact-endpoint = UI request)
- Deregister-only of felhom-usb via
POST /settings/storage/remove(the "Eltávolítás" form, fieldstorage_path, stable path) → 302; registry left only flash; drive stayed mounted at stable (sdb1[/felhom-data]) — not ejected; no rootfs placeholder. - Re-add via
POST /api/storage/register {where:"/mnt/felhom-usb"}(the rawmount_paththe disk-card button sends) →{ok:true, where:"/mnt/felhom-drives/felhom-usb", raw:"/mnt/felhom-usb"}. Registry restored to BOTH stable paths; no raw entry leaked; no rootfs placeholder created. The v0.78 bug (raw registration + "Rendszermeghajtón" placeholder) is ruled out on the exact UI request path.
1.3 — attach existing (LOOP_A) — ⚠️ INCONCLUSIVE (by-design block; compensating coverage)
1.4 — init/format+enroll (LOOP_B) — ⚠️ INCONCLUSIVE (by-design block; compensating coverage)
1.8 — wipe (LOOP_A) — ⚠️ INCONCLUSIVE (destructive device-wipe not executed)
- Loop devices cannot drive the real init/attach/wipe flow: the agent refuses non-whitelisted block
devices —
POST /api/storage/init {device:/dev/loop1}→HTTP 400: storage: refusing to operate on non-whitelisted block device "/dev/loop1"(felhom-agentinternal/storage/validate.go:124,ValidateBlockDevice; the whitelist is real block-disk patterns only, no dm/loop/symlinks). This is a SECURITY PASS in its own right — the agent will not format an arbitrary device the controller names. - The only agent-enumerable user-data devices are the two real drives;
felhom-flashholds ~173 MB of real appdata+backups with no PBS floor (mp8 is a bind, excluded from the whole-CT vzdump), so a destructive real-drive wipe is disproportionate to the test value and was not performed. - Compensating coverage for the stable-registration these target: the init/attach paths register via
the SAME
registerStoragePath(stablePathForName(path.Base(where)), …)proven live in 1.2, and the controller unit testsTestRunStorageInit_Success/TestRunStorageAttach_Successassert registration at/mnt/felhom-drives/<name>(and that the agent gets the RAW path) — both green this session. Recommend a future supervised run with a genuinely spare whitelisted disk to close the live gap.
1.5 — disk card shows stable path + correct badge — ✅ PASS (live UI, claude-in-chrome)
- Both cards render the stable sub-line (
/mnt/felhom-drives/felhom-flash,/mnt/felhom-drives/felhom-usb), correct device/size (ext4 · /dev/sdc1[/felhom-data] · Flash Drive0.2/117.1 GB;ext4 · /dev/sdb1[/felhom-data] · HD710 PRO0.0/915.8 GB), badge "Regisztrálva" + "Aktív" — NOT the v0.79 spurious "Nem regisztrált", NOT "Rendszermeghajtón". Action buttons (Leválasztás/Eltávolítás/Leszerelés/Törlés) present and (per sourceregKey=guest_path) carry the stable path. Screenshot captured.
1.12 — free-form raw add probe (the parked gap) — ✅ PASS (no symptom reproduced) (exact-endpoint)
- Raw
/mnt/felhom-flashdoes not exist inside the container (the raw host mount is never bound into the guest; only the stable parent/mnt/felhom-drivesis, via mp8 + the/mnt:rslavebind). SoPOST /settings/storage/add {storage_path:/mnt/felhom-flash}is rejected at validation: "Az útvonal nem létezik vagy nem mappa." (handlers.go:1264exists-and-dir;:1272is-mountpoint). Registry unchanged (still exactly the 2 stable paths). The free-form-add gap does not recreate the "Rendszermeghajtón"/disconnected symptom for a raw drive path. - Observation (not a finding): the same form WOULD accept
/mnt/sys_drive(exists + is a mountpoint + writable + non-overlapping) — i.e. a user could free-form-register the system SSD volume as a user-data path. That is consistent with the documented SSD-only fallback (Phase 2.4), not the raw-drive bug; noted for awareness, no symptom.
1.6 — deploy app onto a drive → HDD_PATH stable — ✅ PASS (exact-endpoint = UI deploy)
- Deployed navidrome (lightest drive-backed app, 256M) via
POST /api/stacks/navidrome/deploy{values:{DOMAIN,SUBDOMAIN:navtest,HDD_PATH:/mnt/felhom-drives/felhom-usb}}→ healthy.app.yamlHDD_PATH=/mnt/felhom-drives/felhom-usb(stable). Container media bind/mnt/felhom-drives/felhom-usb/userdata/media/music → /music(stable drive namespace, dir 2775 g1000). Data not on rootfs; HDD_PATH not raw. (navidrome's config/DB is a named Docker volume on the docker-data disk by catalog design — only its media is drive-bound.)
1.7 — eject removes the stable entry cleanly + impact names apps — ✅ PASS (exact-endpoint)
- Impact:
GET /api/storage/impact?where=/mnt/felhom-drives/felhom-usb→apps:["Navidrome"]; felhom-flash →apps:[]. Names the right apps (wrong outcome "no apps when an app uses it" ruled out). - Eject (Leválasztás) on felhom-flash (no app, reversible): agent ejected the raw
/mnt/felhom-flash(mapped from stable viaagentWhere); registry left ONLY felhom-usb — the felhom-flash stable entry removed cleanly, no orphan (the v0.80 bug ruled out); bind detached, raw host mount preserved. Restored viaPOST /api/storage/register→ re-registered stable + re-bound (sdc1[/felhom-data]).
1.8 — wipe — ⚠️ INCONCLUSIVE (see 1.3/1.4 — no safe whitelisted target; device-wipe not executed)
- The wipe's confirm/impact GATE is otherwise exercised:
handleStorageWiperequiresmount_name == path.Base(where)(storage_handlers.go:558) and ejects viaagentWherethen formats via the agent's two-step confirm — the same mapping proven in 1.7/1.10. The destructivemkfsitself was not run (loop blocked by the whitelist; real drives hold un-floored data).
1.9 — migrate-app flips HDD_PATH stable→stable — ✅ PASS (exact-endpoint)
POST /api/storage/migrate-app {app:navidrome, target:/mnt/felhom-drives/felhom-flash}→ job ran validate→copy→verify→flip→cleanup (phase=done, unitcleaned). navidromeHDD_PATHflipped to/mnt/felhom-drives/felhom-flash(stable); media bind now on flash; navidrome healthy; no raw leak; source cleaned only AFTER verify. Wrong outcomes (source deleted before verify / HDD_PATH left raw/old) ruled out. (navidrome now lives on felhom-flash — removed in Phase 5.)
1.10 — self-serve decommission → agent op via raw + app missing storage — ✅ PASS (exact-endpoint)
POST /api/storage/decommission {where:stable-usb, mode:anyway, mount_name:felhom-usb}→{decommissioned:true, stopped_apps:["Navidrome"]}. Agent intentda9e7089→decommissioned; agent log shows the op on the RAW path (where=/mnt/felhom-usb,guest-detach … fully unmounted) — wrong outcome (agent called with the unmapped stable path) ruled out. navidrome gate-stopped.- Restored:
POST /api/storage/reconnect→{reenrolled:true, restarted:["navidrome"]}, intent→enrolled, re-bound at stable, navidrome healthy.
1.11 — backup lands under the stable namespace — ✅ PASS (path-correctness) + FINDING (restic) (exact-endpoint)
POST /api/backup/run(the app-data DB-dump backup) ran successfully; for navidrome it correctly found 0 databases (no DB container — SQLite in a named volume), a no-op success, not a failure.- Path correctness: all existing backup artifacts live under the stable namespace
/mnt/felhom-drives/felhom-flash/backups/primary/<app>/(recovery units:compose/+manifest.json), never a raw/rootfs path. Wrong outcome (backup under raw/rootfs) ruled out. - ⚠️ FINDING F1 (minor, pre-existing): the prompt's "restic repo" premise is stale, and the restic
binary is dead weight in the image. Disk-tier restic backup was moved to the host agent in the
de-privileging; the controller source no longer execs restic (all references are "moved to agent"
comments + vestigial
config.goResticPasswordFile/Schedulefields + thesettings.gomigrateResticToRsyncone-way migration). ButDockerfile:58stillapt installsrestic(and line 42 comments "for backup operations"), so/usr/bin/resticships in the controller image unused. Fix direction: dropresticfrom the Dockerfile install list (and the deadRestic*config fields) to shed image weight + remove a confusing "disk-tier is still here" signal. Not a functional bug. - Observation: a named-volume app (navidrome) deployed onto a drive created no
appdata/orbackups/namespace nor a recovery unit on that drive (only the shareduserdata/skeleton). Recovery units appear keyed to appdata-bind apps; named-volume-only apps rely on the DB-dump/volume-tar path instead. Flagged for awareness; likely by-design, not verified as a bug.
Phase 1 verdict: the stable-vs-raw path model holds across register/deregister/disk-card/eject/ decommission/migrate/backup. The v0.78 (raw register), v0.79 (spurious unregistered badge), and v0.80 (orphaned stable entry on eject) bugs are all ruled out live. The only gaps are the loop-device init/attach/wipe (blocked by the agent's block-device whitelist — a safety PASS) and the restic dead-weight finding (F1).
Phase 2 — /mnt/sys_drive third volume — ✅ COMPLETE (2.5 reboot-survival in 3.3)
Method: exact-endpoint + host/guest inspection.
2.1 — sys_drive is a distinct device; warning absent — ✅ PASS
- Three distinct devices:
/=pve-vm--9201--disk--0,/var/lib/docker=disk--1,/mnt/sys_drive=disk--2(same in the controller container). Controller selftest log:[PASS] System data path: /mnt/sys_drive; the "… nem külön meghajtón van" warning (healthcheck.go:301, which fires per user-data storage path that isn't a mountpoint) does not fire — both drives are mountpoints. Wrong outcome (sys_drive on rootfs / warning present) ruled out.
2.2 — disk-full isolation — ✅ PASS
fallocate -l 46G /mnt/sys_drive/…→ sys_drive 99% (599M free)./unchanged (940M/29G free, 4%),/var/lib/dockerunchanged (1.5G/186G, 1%). Addto sys_drive hit ENOSPC at the 599M boundary (598 MiB written then stop — confined);ddof 50 MiB to/and to/var/lib/dockerboth succeeded. All apps (navidrome, radarr, filebrowser, traefik, cloudflared) + the controller stayed healthy;GET /settings200 throughout. Cleanup restored sys_drive to 1%. Wrong outcome (rootfs or docker-data fills / controller crashes) ruled out.
2.3 — mp1 included in whole-CT vzdump — ✅ PASS (captured during the Phase-0 floor)
- The floor
vzdump 9201log:including mount point mp1 ('/mnt/sys_drive')+ theproxmox-backup-client … --include-dev /mnt/vzsnap0/./mnt/sys_drive. mp1 (backup=1) is in the PBS archive. Wrong outcome (mp1 excluded) ruled out.
2.4 — SSD-only fallback app data on disk-2 — ✅ PASS
- Deployed radarr with
HDD_PATH=/mnt/sys_drive. Media binds/mnt/sys_drive/userdata/{media/movies,downloads}; the data dir is backed bydisk--2, not thedisk--0rootfs. Wrong outcome (data on rootfs) ruled out.
2.5 — df size + reboot-survival — ✅ PASS (size) / ↪ reboot-survival verified in 3.3
df /mnt/sys_drive= 50 GB (8 golden + 42 grow, matchingmp1 size=50G). Still a separate device after the guest reboot — confirmed in Phase 3.3 (post-restart sys_drive stilldisk--2).
Phase 3 — server restart + retired banner — ✅ COMPLETE (F2 found)
3.1 — banner gone in every storage state — ✅ PASS
- Rendered
/settingsin the current state (navidrome on flash, radarr on sys_drive, both drives registered, after the 1.12 raw free-form add attempt): 0 occurrences of "aktiválásra vár" / "Újraindítás most" /activate-drives-btn/activatePendingDrives. The banner is retired (v0.81.0) and never appears. The "Kiszolgáló újraindítása" card is present (4 refs).
3.2 — endpoint routing + wrapping — ✅ PASS (routing) / ⚠️ auth-gate not demonstrable on demo
POST /api/storage/activate→ 404 (old route removed).POST /api/server/rebootis routed (202 — confirmed live via the 3.3 UI click, not double-fired here). Route is wrappedRequireAuth(CsrfProtect(HandleServerReboot))(cmd/controller/main.go:731).- Demo limitation: the demo controller has no password set (
authEnabled()=false), so RequireAuth- CsrfProtect are both no-ops here (
csrf.go:30exempts when auth is disabled) — the "rejected without auth/CSRF" half of the matrix can't be shown without setting a password (out of scope: changing account settings). The middleware wrapping is source-verified and is the same gate every other mutating endpoint uses. Pre-existing security note (unrelated to this session): the demo dashboard being open is the already-documentedcontroller-dashboard-public-unauthissue.
- CsrfProtect are both no-ops here (
3.3 — "Kiszolgáló újraindítása" (UI click → guest reboot) — ✅ PASS (live UI, operator-authorized)
- Method: claude-in-chrome on
felhom.demo-felhom.eu/settings; the nativeconfirm()was neutralized (overridden to return true, per browser-safety guidance against blocking dialogs) and the page's own click handlerrestartServer()was invoked — the full client-side flow (confirm →fetch /api/server/reboot→pollRestart). (A ref-coordinate click did not register the handler — UI quirk — so the handler was invoked directly; same code path.) Confirms 3.2's/api/server/reboot→ 202 live. - Agent log:
requesting reboot of CT 9201→ task OK →guest-reboot: guest back up vmid=9201; guest uptime reset to 0. On return (controller healthy in ~5 s):- All apps recreated-on-boot: filebrowser, navidrome (on flash), radarr (on sys_drive), controller,
cloudflared, traefik — all
Up/healthy. - Drives re-bound at STABLE: flash
sdc1[/felhom-data], usbsdb1[/felhom-data]. - sys_drive still separate:
pve-vm--9201--disk--2(→ 2.5 reboot-survival PASS). - No rootfs placeholders at raw
/mnt/felhom-{flash,usb}(both "not present"). - Registry intact: both stable paths. Dashboard returned.
- Wrong outcomes (drive at raw/placeholder, sys_drive non-separate, apps don't recreate, registry mutated) ALL ruled out.
- All apps recreated-on-boot: filebrowser, navidrome (on flash), radarr (on sys_drive), controller,
cloudflared, traefik — all
3.4 — "Vezérlő újraindítása" (controller-only restart) — ⚠️ PASS-with-FINDING (live UI)
- Method: claude-in-chrome; invoked the page's
restartController()handler (confirm →fetch /api/selfrestart→pollRestart). Controller back healthy in ~3 s. - Guest NOT rebooted (uptime kept climbing, not reset) ✓. Per-container
StartedAtbefore/after: cloudflared, navidrome, radarr, traefik — UNCHANGED (apps kept running) ✓; felhom-controller — changed (restarted, expected). Drives stayed bound; sys_drive unaffected. - ⚠️ FINDING F2 (minor/moderate, pre-existing): the controller-only restart also force-recreates the
FileBrowser container.
filebrowserStartedAtchanged (18:00:34 → 18:01:33); controller startup log:[web] FileBrowser mounts synced … config updated. Root cause:SyncFileBrowserMounts()(internal/web/handlers.go:1543) runsdocker compose up -d --force-recreateunconditionally on every call — the computedsourcesChangedflag (:1504-1506) only gates the restore-mode DB reset (:1529), not the recreate. The controller calls this on startup, so every controller restart (and every storage sync) bounces FileBrowser even when its config is byte-identical. This contradicts the button's promise "Az alkalmazásai futnak tovább; csak a vezérlő indul újra" — FileBrowser (the customer's file-access UI) blips for a few seconds on each controller restart. Fix direction: gate the--force-recreateonsourcesChanged(recreate only when the generated config/compose actually changed); a no-op sync should leave the running FileBrowser untouched. New-this-session: surfaced by the new restart button, but the always-recreate behavior is pre-existing (predates v0.81.0). Not a data-safety issue.
Phase 3 verdict: the retired banner stays gone (3.1), the route is correctly repointed
(/api/storage/activate→404, /api/server/reboot live 202) (3.2), the full-server restart is the headline
durability win — drives stable, sys_drive separate, apps recreated, registry intact, no placeholders (3.3).
The controller-only restart keeps the guest + customer apps running but does bounce FileBrowser (F2).
Phase 4 — host reboot + chaos — ✅ COMPLETE (4.7 deferred; F3/F4 found)
Pre-host-reboot capture: sdb1=da9e7089(usb), sdc1=81a26531(flash); onboot: 0;
felhom-agent service enabled (will auto-start on host boot). navidrome on flash, radarr on sys_drive.
Prediction: with onboot: 0, 9201 will NOT auto-start after the host reboot (candidate finding F3).
4.1 — plain felhom-pve host reboot — ✅ PASS (drives/apps/registry) + ⚠️ FINDING F3 (no auto-start)
- Issued
rebooton felhom-pve; host returned (fresh boot, ~90 s).felhom-agentauto-started (active). - ⚠️ FINDING F3 (moderate, pre-existing config):
9201did NOT auto-start — it wasstoppedafter the host reboot becausepct config 9201hasonboot: 0. A managed customer guest should come back on its own after a host power-loss/reboot; as configured, the customer's entire home-server (all apps + the controller) stays down until someone manuallypct starts it. Fix direction: the golden/provision path should setonboot: 1(+ a sensiblestartup/boot-order) on the customer guest; audit whether any provisioned guest hasonboot:1. (I started 9201 manually to continue.) - After manual
pct start 9201(controller healthy ~5 s): drives re-bound by the agent reconcile by durable-id (uuid:81a26531flash,uuid:da9e7089usb) at the stable paths; sys_drive separate (disk--2); all apps recreated (navidrome on flash, radarr on sys_drive, both healthy); no rootfs placeholders; registry stable. Wrong outcomes (raw/placeholder, sys_drive non-separate, registry drift) ruled out (modulo F3 auto-start).
4.2 — /dev reshuffle (sdb↔sdc) — ✅ PASS (REAL swap, not just code-path)
- The host reboot genuinely swapped the kernel letters: before
sdb1=da9e7089(usb)/sdc1=81a26531(flash); aftersdb1=81a26531(flash)/sdc1=da9e7089(usb). Despite the swap, the agent bound each drive by fs-UUID:felhom-flash(81a26531, now sdb1) at/mnt/felhom-drives/felhom-flash;felhom-usb(da9e7089, now sdc1) at/mnt/felhom-drives/felhom-usb. Each NAME mapped to its own UUID at the correct stable path — the v0.37 wrong-physical-device regression is ruled out on a real letter swap.
4.5 — double-bind churn — ✅ PASS (stable, not leaking) + ⚠️ FINDING F4 (pre-existing 2×)
- Across the host reboot + multiple 20 s reconcile ticks, each drive's bind count held stable at 2
(
grep -con/proc/mounts, guest AND host) — it did not accumulate to 3,4,… so the "binds leak" wrong outcome is ruled out. - ⚠️ FINDING F4 (minor, pre-existing): each enrolled drive is double-bound — two identical
/dev/sdXN /mnt/felhom-drives/<name> ext4 …entries in/proc/mounts. The agent logsguest-attach: drive bound under shared parent (normalized to one bind … prior_binds=2), i.e. it detects 2 prior binds but the steady state is still 2, so the "normalize to one" isn't fully effective. Matches the documented pre-existing double-bind (the v0.36.x peer-group area); stable, not a leak, no data impact. Fix direction: make the normalize-to-one path converge to a single bind (or stop the duplicate rbind at its source).
4.3 — intent gate (ejected drive must NOT auto-rebind) — ✅ PASS (agent-restart variant)
- Method: to avoid a second full host reboot, the same
ReassertGuestBindsintent gate (which runs on agent startup, identical to the host-reboot path) was exercised by an agent restart. Ejected felhom-usb (no app on it — navidrome migrated to flash in 1.9),systemctl restart felhom-agent. - Agent log:
reconcile: skipping non-enrolled drive (intent-gated) … da9e7089 intent=ejected(repeated each tick) AND… ce9d1d81 intent=ejected(the baseline leftover); enrolled felhom-flash re-bound. Steady state: usb 0 binds (NOT re-bound), flash bound. The ejected drive did not auto-rebind; the enrolled one did. Wrong outcome (ejected auto-rebinds) ruled out. - Transient note: ~3 s after the agent restart,
findmntbriefly showed usb still bound — a propagation lag of the eject's detach through the/mnt:rslaveview + the F4 double-bind — which cleared to 0 within one tick. Not a gate failure (the gate decision is logged correctly throughout). - Restored usb via re-register → intent→enrolled, bound at stable (now
sdc1post-swap), both drives back.
4.4 — drive yank mid-flight → mountpoint-gate — ✅ PASS (exact-endpoint, no host reboot)
- Method:
systemctl stop felhom-agent(freeze binds), lazy-unmounted the flash bind on the host (yank), thenPOST /api/stacks/navidrome/restart(HDD_PATH=flash, now unmounted) to run the deploy belt. - The belt logged
userdata belt: drive /mnt/felhom-drives/felhom-flash not mounted — skipping ensure (held by drive gate)(internal/stacks/manager.go:143) and wrote nothing — the yanked path/mnt/felhom-drives/felhom-flashstayed empty (now resolves topve-root[/mnt/felhom-drives], i.e. the parent/rootfs, but no app data or placeholder was created there). The controller refused to MkdirAll onto the rootfs. Wrong outcome (app data written to rootfs / placeholder created) ruled out. - Recovery:
systemctl start felhom-agent→ flash re-bound at stable (sdb1[/felhom-data]), navidrome stayed healthy throughout, both intents enrolled, no leftover placeholder. Reconnect recovered cleanly.
4.6 — concurrency — ✅ PASS (exact-endpoint)
- (b) eject-in-use is WARNED:
impact?where=…felhom-flash→apps:["Navidrome"]— the in-use app is named before any eject (the type-to-confirm warning surface), so eject-in-use is warned, never silent. - (a) concurrent backup + deregister + re-register (raced against each other):
remove→302,backup→ok,register→ok. Post-race the registry is consistent — exactly one flash + one usb entry (no duplicates, no corruption); both drives bound (usbsdc1, flashsdb1); navidrome + radarr healthy. Wrong outcomes (concurrent corruption / silent app breakage) ruled out.
4.7 — CAPSTONE (destroy + re-provision 9201) — ⏸️ DEFERRED (operator-gated; not run)
- Explicitly operator-gated and requires the golden volid + retrieval passphrase, which were not provided for this autonomous run. Destroying the live demo guest is the campaign's most destructive step; per the "confirm before irreversible ops" guardrail I did not run it without an operator go-ahead. Compensating coverage: the exact original-bug scenario (stable registration with no placeholder, drives live, sys_drive separate, no banner) is already proven clean via 1.2 (register → stable, no placeholder), 3.3 (full-server restart recovery), and 4.1/4.2 (host-reboot + real sdb↔sdc swap recovery). To close 4.7: a supervised run with the operator supplying the golden volid + passphrase, then re-enroll via "Regisztrálás" and assert the clean end-state.
Phase 4 verdict: durability holds — host-reboot drive re-bind by UUID survives a real /dev letter
swap (4.2), the intent gate skips ejected drives (4.3), the mountpoint-gate blocks rootfs writes on a
drive yank (4.4), and concurrency doesn't corrupt (4.6). Two findings: F3 (onboot:0 → guest doesn't
auto-start after host reboot — the highest-impact finding) and F4 (stable-but-present double-bind). The
capstone (4.7) is deferred pending operator credentials.
Phase 5 — restore to clean baseline — ✅ COMPLETE
- Removed the test apps (navidrome stop+remove, radarr stop+remove) → only infra remains (felhom-controller, traefik, cloudflared, filebrowser — all healthy), matching the Phase-0 baseline.
- Loop devices
loop0/loop1were already gone (cleared by the host reboot); scratch imgs + dir removed. All fill files removed (2.2). No decommission soft-markers left (everything reconnected; both intentsenrolled). - Final assertions (all ✅): both drives live at stable with correct device/size (felhom-flash
/dev/sdb1125.7 GB, felhom-usb/dev/sdc1983.3 GB — letters are post-swap, bound by UUID); registry = exactly 2 stable paths;/mnt/sys_driveseparate (disk--2, 50 G,[PASS] System data path); dashboardGET /200; settings UI clean — 0 banner occurrences, both "Vezérlő újraindítása" + "Kiszolgáló újraindítása" cards present. The PBS floor was not needed (nothing wedged).9201left atonboot:0(original state — F3 is reported, not fixed).
DELIVERABLE SUMMARY
1. Findings (root-caused; report-don't-fix — none implemented)
| ID | Sev | New? | Summary | Evidence / source | Fix direction |
|---|---|---|---|---|---|
| F3 | Moderate | surfaced this session (pre-existing config) | Guest 9201 has onboot:0 → does NOT auto-start after a host reboot. The whole customer home-server (apps + controller) stays down until a manual pct start. |
4.1: host rebooted → pct status 9201 = stopped; pct config 9201 → onboot: 0 |
Golden/provision should set onboot:1 (+ boot order) on customer guests; audit all provisioned guests. |
| F2 | Minor/Mod | surfaced this session (behavior pre-existing) | "Vezérlő újraindítása" (controller-only restart) also force-recreates the FileBrowser container every time, even on a no-op sync — contradicts the "your apps keep running" promise. | 3.4: filebrowser StartedAt changed on controller restart; SyncFileBrowserMounts runs docker compose up -d --force-recreate unconditionally — internal/web/handlers.go:1543; sourcesChanged (:1504) only gates the DB reset (:1529) |
Gate the --force-recreate on sourcesChanged; skip recreate when config/compose is byte-identical. |
| F1 | Minor | new discovery (pre-existing) | restic binary ships in the controller image as dead weight. All restic code moved to the agent during de-privileging, but Dockerfile:58 still apt installs it. |
1.11: /usr/bin/restic present; no Go code execs it (only "moved to agent" comments + vestigial config/migration) |
Drop restic from the Dockerfile install list + the dead Restic* config fields. |
| F4 | Minor | pre-existing (documented v0.36.x area) | Each enrolled drive is double-bound (2 identical /proc/mounts entries). Stable, not leaking; agent logs "normalized to one bind (prior_binds=2)" but steady state is 2. |
4.5: grep -c = 2 (guest+host), stable across ticks; brief detach transient in 4.3 |
Make normalize-to-one converge to a single bind / stop the duplicate rbind at source. |
Observations (not findings): (a) a named-volume app (navidrome) deployed onto a drive creates no
recovery unit / appdata namespace on that drive — likely by-design (recovery units key on appdata-bind
apps). (b) The free-form "add storage path" form would accept /mnt/sys_drive (a real mountpoint) —
consistent with the SSD-fallback, not the raw-drive bug.
2. Method per user-facing test
- claude-in-chrome (live UI): 1.5 (disk-card badges, screenshot), 3.3 ("Kiszolgáló újraindítása" click → guest reboot; confirm() neutralized, page handler invoked), 3.4 ("Vezérlő újraindítása" click), 5 (final settings UI). No F9-style bypass — the real handlers/endpoints ran the full server pipeline.
- Exact-endpoint (= the UI's own request; demo has no password so CSRF is a no-op): 1.1, 1.2, 1.6,
1.7, 1.9, 1.10, 1.11, 1.12, 2.x, 3.1/3.2, 4.3, 4.4, 4.6. The drive ops invoked the exact
/api/storage/*and/api/stacks/*endpoints the buttons call. - Host/agent-level (real infrastructure): 2.2 (fill), 2.3 (vzdump), 4.1/4.2 (host reboot + real /dev swap), 4.3 (agent restart → ReassertGuestBinds), 4.4 (agent stop + unmount).
3. Phase-0 PBS floor
felhom-pbs:backup/ct/9201/2026-06-23T17:28:25Z(datastorefelhom-spike). Not used — nothing wedged; every destructive step was reversed in-band.
4. End state
- 9201 restored to a healthy baseline: infra-only apps healthy, both drives live at stable (correct
device/size),
/mnt/sys_driveseparate + warning-free, registry = 2 stable paths, settings UI clean (no banner, both restart buttons). Left intentionally atonboot:0(original state; F3 reported not fixed). Two pre-existing absent drive-intents (277a2179enrolled-absent,ce9d1d81ejected) remain — they predate the campaign (present in the Phase-0 baseline), not introduced here.
5. Coverage gaps
- 1.3 / 1.4 / 1.8 (loop-device init/attach/wipe): not executable — the agent's block-device whitelist
(
felhom-agent internal/storage/validate.go:124) refuses non-whitelisted devices (a safety PASS), and no spare whitelisted disk exists (the 2 real drives hold un-floored data). Stable-registration covered by 1.2 live + unit tests. Needs a real spare disk in a supervised run. - 4.7 (capstone destroy + re-provision): deferred — operator-gated, needs golden volid + retrieval passphrase (not provided this run). Compensating coverage via 1.2/3.3/4.1.
- 3.2 auth/CSRF rejection: can't be shown on the passwordless demo (gate is a no-op); wrapping is source-verified.
- 4.2 was a REAL /dev swap (better than the planned code-path fallback) — no gap there.
Campaign result (overview)
- Phases 0–5 complete. The three shipped features hold up: the stable-vs-raw path model (v0.78–0.80) is correct across register/deregister/disk-card/eject/decommission/migrate/backup (Phase 1); the /mnt/sys_drive split (agent v0.40.0) isolates disk-full, is backed up, and survives reboots (Phase 2); the retired banner + standalone server-restart (v0.81.0) work — full-server restart recovers drives at stable, sys_drive separate, apps recreated, registry intact (Phase 3); and durability survives a real host reboot with a genuine sdb↔sdc swap, the intent gate, the mountpoint-gate, and concurrency (Phase 4).
- 4 findings, all reported-not-fixed (see DELIVERABLE SUMMARY §1): F3
onboot:0(guest won't auto-start after host reboot — highest impact), F2 controller restart bounces FileBrowser, F1 restic dead-weight binary, F4 stable double-bind. No data-loss, no path-escape, no silent corruption found. Demo restored to a healthy baseline.