25e5cb58504b237df5135496774d1c856796f059
Sub-cause: on guest pct reboot, in-guest dockerd auto-starts unless-stopped drive-backed apps ~18s BEFORE the agent re-binds the drive; the create-time volume bind fails (mkdir /mnt/felhom-drives/<drive>/userdata: permission denied) and RestartCount=0 means it's never retried -> stuck Exited. The existing recovery (processGuestBootChange) RAN but raced the rebind: it sampled the agent's BoundUnderParent once during fast startup (not live yet), recreated nothing, and persisted the new boot-id -> burned its one-shot. The periodic gate never recovered them either (first observation after the rebind -> no transition). Fix (harden the existing mechanism, no parallel one): processGuestBootChange now gates on the REAL live in-guest bind. driveBindLive checks whether /mnt/felhom-drives/<drive> is an actual mountpoint in the controller's own /mnt (rslave) /proc/self/mountinfo -- true only once the agent's bind propagated, exactly when docker can recreate the app. pollLiveBinds waits for that (bounded ~120s, poll 2s; rebind lands ~18s) and only then recreates via the normal pipeline, including stuck-Exited create-time-failure apps (shouldRecreateOnBoot is state-independent). Single-flight; absent-after-window drives left to the gate; host-reboot path unaffected; guest-only reboot path now covered. Tests: pollLiveBinds waits through the rebind then reports live (recreate fires); never-live drive stays absent; pre-fix companion (single early sample misses the not-yet-live bind). Red-proofed against a no-wait single-sample. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Description
No description provided
Languages
Go
84.2%
HTML
11.1%
Shell
2.1%
CSS
1.8%
Python
0.6%
Other
0.1%