Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
15 KiB
Storage lifecycle completion + acceptance sweep (intermediary-mount) — 2026-06-15
Scope: finish the storage lifecycle on the intermediary-mount foundation (Phase A), migrate the second drive + confirm rslave (Phase B), and run a full edge-case acceptance sweep including a real host reboot (Phase C). Unattended, live on guest 9201 / felhom-pve. Trunk-based, per-commit green gate, non-hollow tests with pre-fix companions.
Deployed at close: agent v0.36.7 (felhom-agent@1e20584), controller v0.68.1
(felhom-controller@7a85732), catalog app-catalog-felhom.eu@939864f. Both external drives on the new
model, 25 containers healthy, felhom-flash default.
Headline: the host-boot ordering is now LIVE-validated — on a real reboot of felhom-pve the
shared-parent unit ran before pve-guests, the guest auto-started (onboot=1), both drives re-propagated
and all 8 apps converged with no manual intervention.
Phase A — lifecycle fixes
| Item | Result | Where |
|---|---|---|
| Boot-id determinism | ✅ agent v0.36.0 emits guest_boot_id (<host-btime>-<guest-init-starttime>); controller v0.68.0 persists LastGuestBootID + processGuestBootChange recreates all deployed drive-backed present apps on change. v0.68.1: state-independent (was missing apps docker hadn't restarted yet). |
internal/localapi/intermediary.go (GuestBootID), internal/web/intermediary.go (shouldRecreateOnBoot), settings.go |
| H2 decommission button | ✅ "Leszerelés" on every connected drive → migrate-then-decommission (target select) OR decommission-anyway (type-to-confirm). Server side was already complete; new model never touches the parent mp. | templates/settings.html (storageDecommission) |
| H3 re-enroll button | ✅ "Visszacsatlakoztatás" on decommissioned/ejected drives → clears marker + re-attaches under the parent + restarts gate-stopped apps. | handleStorageReconnect (decommissioned branch) |
| M1 default reassignment | ✅ defaultPromotionTarget: decommissioning the default auto-promotes another schedulable drive (prefers the migrate target); BLOCKS if none. |
internal/web/intermediary.go, finalizeDecommissionWith |
| M3 userdata setgid | ✅ skeleton already 2775 on every nested dir; the migrate merge-walk now RE-ASSERTS 2775 (EnsureUserdataDir) on userdata dirs (isUserdataDir) instead of merely preserving the source mode. |
internal/stacks/migrate.go |
Non-hollow tests (each with a companion that fails the pre-fix impl): TestShouldRecreateOnBoot,
TestDefaultPromotionTarget, TestIsUserdataDir, TestStarttimeFromStat. Also fixed the existing H1
{path} vs {where} JS body mismatch.
Phase B — second drive + rslave
- felhom-usb migrated off its stranded legacy bind (stale
decommissionedintent cleared via the H3 reconnect → agent guest-attach set intent=enrolled + recorded the bind + bound felhom-data under the parent); registry repointed to/mnt/felhom-drives/felhom-usb; legacymp1deleted. Both drives now on the new model. - rslave is EXPLICIT (not incidental):
-v /mnt:/mnt:rslavein the controller bootstrap docker-run (line 50). Verified it survives a controller redeploy (E7).
Phase C — acceptance sweep (per-item, with evidence)
E1 — HOST REBOOT (the headline) ✅ LIVE-VALIDATED (twice)
reboot issued over SSH; polled SSH return (~70–80 s). From journalctl -b:
felhom-shared-parent.service Finished 19:50:29 (ExecMainStatus=0)
pve-guests.service Starting 19:50:35 → "Starting CT 9201" (onboot=1 auto-start)
The shared parent was made shared before pve-guests; the guest auto-started; both drives
re-propagated (BoundUnderParent); /mnt/felhom-usb (legacy) gone. Apps converged via the gate with no
manual intervention:
[gate] drive ABSENT /mnt/felhom-drives/felhom-flash — stopped+blocked 8 app(s)
[gate] drive RETURNED /mnt/felhom-drives/felhom-flash — re-attached + restarted gate-stopped apps
Data intact (romm bound to /mnt/felhom-drives/felhom-flash/..., immich-postgres base). The first
host reboot (v0.68.0) surfaced a bug — 5 apps stayed exited because the recreate filtered on container
state; fixed in v0.68.1 (recreate ALL deployed present apps) + re-validated.
E2 — guest reboot ×3 ✅
Multiple guest reboots; each converged all 8 apps with zero manual starts (the deterministic guest_boot_id
recreate + the gate). The final clean reboots showed single bind per drive + no 255.
E3 — drive absent at guest boot ✅ (covered by E1)
The gate's ABSENT→RETURNED path (E1) is exactly this: a drive absent at boot stops+blocks its apps; the guesthook (C1 net) lets the guest boot regardless. The agent re-binds on return → gate restarts.
E4 — drive yanked while running ✅
umount of a running drive → immediate guest-root write to the bare stable path DENIED, no leak
to host root. The agent reconcile auto-rebound it; apps stayed healthy. (Gate stop/restart proven by E1.)
E5 — fail-close, capability-proof ✅
With felhom-usb detached, a guest-root write AND a root app container (komga) write to the bare
stable path both returned Permission denied — the host-root-owned (unmapped) dir defeats guest-root
CAP_DAC_OVERRIDE (what chmod-0000 could not). find showed no file leaked onto host root.
E6 — confinement ✅ (both drives)
Guest sees only userdata (felhom-usb) / appdata backups media userdata (felhom-flash); ls .../dump
→ not found. The host raw drives carry dump/images/lost+found/... which never cross in.
E7 — controller redeploy ✅
After the controller container restarts, /mnt/felhom-drives is private,slave inside it and it still
sees the drive data — proving the :rslave is explicit, not incidental.
E8 — two drives ✅
Detaching felhom-usb left felhom-flash + its 8 apps fully unaffected.
E9 — decommission button (H2) ✅
decommission-anyway on felhom-usb → decommissioned:true, intent decommissioned, bind detached, parent
mp untouched (reboot-safe, no brick). Migrate-then-decommission shares the proven handleStorageDecommission.
E10 — re-enroll (H3) ✅ (after a fix)
First attempt FAILED and exposed a real bug: decommission unmounted the RAW drive, so re-enroll bound
an empty pve-root dir. Fix (agent v0.36.1): decommission is now a logical retire — DetachDrive the
bind but LEAVE the raw mounted, so re-enroll re-binds the real drive. Re-tested: re-enroll → felhom-usb
backed by /dev/sdc1[/felhom-data], data intact.
E11 — default reassignment (M1) ✅
Made felhom-usb default, decommissioned it → [gate] default drive reassigned … → … (M1); a default
always remained (never zero). The only-drive BLOCK is unit-tested (TestDefaultPromotionTarget).
E12 — eject the drive holding ALL apps ✅ (covered by E1)
E1's felhom-flash ABSENT→RETURNED is exactly this: all 8 apps gated, then auto-restarted on return, reboot-free.
E13 — rapid eject/reconnect ×3 ✅ (after the deepest fix)
Surfaced a double-bind (2 stacked binds per drive). Root-caused to the shared-parent self-bind
inheriting /'s shared peer group (/mnt/felhom-drives was shared:1 like /), so every drive
bind propagated back and doubled. Fixed across agent v0.36.3–v0.36.7:
- v0.36.3
DetachDriveloop-umounts all stacked layers (full detach → fail-close intact); - v0.36.4 a mutex serializes Attach/Detach (no TOCTOU double-bind);
- v0.36.5
AttachDrivenormalizes to exactly one bind (countHostMounts); - v0.36.6 the real root cause —
make-privatebeforemake-sharedso the parent owns its own peer group (nowshared:51), binds propagate to the guest exactly once; - v0.36.7 isolate only on create (re-doing it churns the peer-group id and orphans the guest's slave).
End state: single bind per drive, guest sees both, all apps healthy, no leak.
E14 / E15 — migration + setgid (M3) ◑ partial
M3's routing (isUserdataDir) is unit-tested and EnsureUserdataDir enforces 2775; the convention was
verified live (/mnt/felhom-drives/felhom-flash/userdata = 2775, gid 1000; FileBrowser bound there).
A full live migrate-all round-trip between the two drives was NOT re-run to avoid merging both drives'
userdata; the migrate/checksum machinery is the pre-existing, previously-validated path (only the M3
re-assert is new, and it's unit-tested).
Regression ✅
FileBrowser bound to userdata on both new-model drives (2775 group-write); immich-postgres data intact;
dashboard + monitoring HTTP 200; no exited containers.
Surprises / breaks found (and fixed) this run
- E1 (v0.68.0): boot-id recreate filtered on container state → 5 apps stayed exited after a host reboot. → v0.68.1 recreate ALL deployed present apps (state-independent).
- E10: decommission unmounted the raw drive → re-enroll bound an empty dir. → v0.36.1 keep raw mounted (logical retire); v0.36.2 same for eject.
- E13 double-bind: shared-parent inherited
/'s peer group → every drive bind doubled. → v0.36.6 make-private before make-shared (own group); v0.36.7 only-on-create (avoid peer-group churn that orphans the guest slave). v0.36.3–.5 are defense-in-depth (loop-detach, mutex, normalize-to-one). - Pre-start hook vs agent churn (transient): rebooting the guest immediately after an agent restart
raced the agent's parent-bind churn against the guest's mp3 setup →
lxc.hook.pre-startexit 255. Resolved by the on-create-only isolation (the parent is no longer churned); a clean guest reboot (no concurrent agent op) returned exit 0. Operational note: don't restart the agent and reboot the guest in the same instant. - Device-letter swap: on a host reboot the kernel re-enumerated sdb↔sdc; harmless because all mounts are by fs-UUID.
Residuals (documented, not blocking)
- E14/E15 full live migrate-all not re-run (M3 unit-tested + convention verified live; data-merge risk).
- Boot-id first-sight bounce: a fresh controller (empty
LastGuestBootID) recreates apps once on its first start — a one-time cost per data-volume lifetime; subsequent restarts don't bounce. - "Safely removable" semantics: eject/decommission now leave the raw mounted (re-enrollable); fs-flush-before-physical-pull is the separate "remove from system" action's job (future refinement).
Commit hashes
agent 1e20584 (v0.36.7) · controller 7a85732 (v0.68.1) · catalog 939864f.
Agent versions this run: v0.36.0→v0.36.7. Controller: v0.68.0→v0.68.1.
Addendum (2026-06-15, later session) — M3 live re-verification + two UI fixes (controller v0.68.2/.3)
Closes the E14/E15 residual above ("full live migrate-all not re-run; M3 unit-tested + convention verified live"). The M3 setgid re-assertion on a pre-existing stale dir is now live-proven on guest 9201, alongside two small UI cleanups shipped the same session.
M3 — migrate merge-walk re-asserts setgid on a stale dir (LIVE, non-hollow)
Two full migrate-all runs were executed via the dashboard UI (the real /api/storage/migrate
endpoint, all 8 flash-resident apps): flash→usb, then usb→flash (restoring apps to the default drive).
Before each, the migrate target was pre-seeded host-side with deliberately-stale 0755,
non-setgid userdata dirs (mimicking the B3 calibre case), stat recorded.
Result — clean proof (usb→flash run): a pre-seeded userdata/documents at 755 gid1000 —
a dir no app mounts — came out 2775 gid1000 after migrate (EnsureUserdataDir re-assert in
walkMerge, migrate.go:854-859), as did the seeded userdata root and userdata/import. Every userdata
dir in the tree was 2775 afterwards. Static seeded files (b3-document.txt, b3-movie.txt,
b3-photo.txt, demo.jpg) had identical sha256 before/after (integrity intact). All 8 apps
redeployed healthy on the target (komga "unhealthy" throughout = its pre-existing
/api/v1/actuator/health 401, serving fine — unrelated).
Important secondary finding — the residual import/calibre 755 is NOT a migrate bug. In both runs
userdata/import/calibre ended at 755 despite the merge-walk. Root cause isolated live: the
calibre-web (CWA) container bind-mounts ${USERDATA_PATH}/import/calibre as its /cwa-book-ingest
drop-zone and chmods it to 755 (strips setgid) on every startup — proven by setting the dir to
2775 and restarting calibre-web (reverted to 755 with no migration involved). This is why the
source flash calibre was already 755 pre-migration too. The merge-walk re-asserts 2775 correctly
during the copy phase; calibre-web clobbers it again during the flip/redeploy phase. So migrate.go was
NOT changed — the M3 path is correct. (If the convention matters for that single-app ingest dir, the
fix belongs in the catalog/app layer, e.g. an umask/entrypoint wrapper for CWA, not in migrate.)
UI fix 1 — stack-card state-badge clipping (controller v0.68.2, CSS)
.stack-detail-header is flex/space-between; the title-row lacked min-width:0 and the
white-space:nowrap .stack-state-badge lacked flex-shrink:0, so on an unhealthy app the long
"⚠ URL nem elérhető…" route-unpublished warning inflated the title-row and the flexbox compressed the
badge — clipping "Nem egészséges" to "Ner…". Fix: .stack-title-row{flex:1;min-width:0} +
.stack-state-badge{flex-shrink:0}. Browser-verified on /stacks (komga BEFORE clipped "Ner…" → AFTER
full "Nem egészséges" with the warning wrapping in the title column; Mealie/healthy + not-deployed cards
unchanged). The defensive flex-shrink:0 on .badge-missing-storage/.badge-orphaned was not
added — no card is currently both unhealthy AND missing-storage/orphaned, and the title-row absorbing all
shrink already shields sibling badges; one-liner available if that combo ever surfaces.
UI fix 2 — Beállítások endless-refresh loop (controller v0.68.3)
Found while validating M3: after any migration finished, the settings page reloaded itself every
~1.5 s forever. MigrationStatus keeps returning the last done job indefinitely; the page's resume-view
called migWatch() for any job, and migWatch's done branch does setTimeout(location.reload,1500)
→ load → see persisted done → watch → reload → loop. Fix (settings.html): resume-view watches only an
in-progress job (phase!=='done' && phase!=='aborted'); the one-time post-completion reload still
fires from the active watcher. Browser-verified: with a persisted done job present, the page stayed put
for 25 s (navType:navigate, marker survived, panel idle) on v0.68.3.
Resting state
Apps healthy on felhom-flash (default drive); felhom-usb cleaned/empty; controller v0.68.3.
Commit hashes (this addendum)
controller a821a9d (v0.68.3; CSS fix v0.68.2 = 37ed757). Deployed live on guest 9201 / felhom-pve.