# DIAG — does the R-166 boot sweep start an app whose data drive is absent? **Date:** 2026-08-02 · **Box:** demo `felhom-pve` guest 9201 · **Controller:** v0.189.0 **Question raised by:** reading the v0.189.0 diff, **not** by an incident. Diagnosed before any fix was written (task §9.2, "diagnose before theorising"). ## Verdict **CONFIRMED — the sweep starts it.** Observed directly, twice in the same run (both attempts). **With a qualification that changes the severity but not the fix:** on this box the resulting `docker compose up -d` **failed** and **nothing was written** to the wrong disk. The mechanism that prevented the write is a **filesystem-permission accident that no code owns**, and the harm that DID occur is a different one: a false dead-app alarm for an app the drive gate is deliberately holding. ## The reasoning under test `bootrecon` imports `stacks` alone and has no storage awareness. Since v0.189.0 `isBootOrphan` returns true for `desired_state: running` + zero containers. The drive-absent gate stops apps with `compose down` (leaving exactly zero containers) and never touches `desired_state`, because it is not the customer. `Manager.StartStack` has no drive gate. Therefore the sweep should start an app whose drive is absent. ## Method The drive is a real USB disk (`/dev/sdb`), bound under the stable parent at `/mnt/felhom-drives/hdd_1`. The gate's absence signal is the agent's `BoundUnderParent` (`planDriveGates`, `intermediary.go:226`), so the drive was made absent **by unmounting it**, not by editing controller state — the state edit would have proven a different thing. **Run 1 was contaminated and is reported because it produced a mechanism.** Unmounting only the parent bind was not enough: **the agent re-binds it within ~60 s** while `/dev/sdb` is still mounted at `/mnt/hdd_1`. In that run the drive gate's startup reconcile re-attached the drive and restarted the apps at `17:18:47`, **one second before** `bootrecon` looked at `17:18:48` — which then logged `no boot-orphaned apps`. That is a race that happened to go the safe way, **not** a disproof. Had it been reported as one, the conclusion would have been wrong. Run 2 therefore unmounted **both** `/mnt/felhom-drives/hdd_1` and `/mnt/hdd_1`, and held them unmounted against the agent's healing for the duration. ## Preconditions, all verified before the observation | # | Precondition | Observed | |---|---|---| | 1 | drive absent | `mount \| grep -c hdd_1` → **0** | | 2 | app has zero containers | `docker ps -a --filter name=calibre` → **0** | | 3 | intent still says running | `desired_state: running` | | 4 | app still deployed | `deployed: true` | The gate stopped the apps itself first, on its own signal: ``` 17:22:07 [WARN] [gate] drive ABSENT /mnt/felhom-drives/hdd_1 — stopped+blocked 2 app(s): [calibre-web immich] ``` `immich` was set to `desired_state: stopped` beforehand to scope the blast radius to one app; the gate is per-drive, so it stops every app on that drive regardless. ## The observation ``` 17:22:29 [INFO] [stacks] desired-state backfill: 0 app(s) recorded as running, 0 left unrecorded 17:22:34 [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [calibre-web] — up to 2 attempt(s) 17:22:34 [INFO] [stacks] Starting stack: calibre-web 17:22:35 [WARN] [bootrecon] attempt 1/2: start "calibre-web" failed after 0.4s: exit code 1 17:23:05 [WARN] [bootrecon] attempt 2/2: start "calibre-web" failed after 0.4s: exit code 1 17:23:05 [WARN] [bootrecon] gave up after 2 attempt(s): recovered=[] still down=[calibre-web] (the dead-app alarm now owns these) ``` **The sweep selected a drive-absent app and called `StartStack` on it.** That is the confirmation. ## What stopped the write, and why it must not be relied on ``` Error response from daemon: error while creating mount source path '/mnt/felhom-drives/hdd_1/userdata/media/books': mkdir /mnt/felhom-drives/hdd_1/userdata: permission denied ``` With the drive unbound, `/mnt/felhom-drives/hdd_1` is an empty directory on the host's `pve-root`, and it is **host-root-owned**: ``` /mnt/felhom-drives uid=0 gid=0 mode=755 /mnt/felhom-drives/hdd_1 uid=0 gid=0 mode=755 ``` Guest 9201 is an **unprivileged** LXC, so its container root is uid 100000 and cannot `mkdir` there. Confirmed no write occurred: `find /mnt/felhom-drives/hdd_1/` returned the directory alone, and `df -h /` was unchanged at `25G used / 28%` before and after. **This protection is accidental.** Nothing in the controller chose it, no test pins it, and it rests on two conditions that are not guaranteed and are not checked anywhere: 1. the guest is unprivileged (a privileged guest maps root→0 and the `mkdir` succeeds); 2. the stable-parent mountpoint directory is root-owned. When the drive **is** bound, that same path shows `uid=100000 gid=100000` — i.e. guest-writable. Any code path or agent version that pre-creates the mountpoint with guest ownership removes the protection silently. It is one `chown` away from being gone, and its removal would be invisible until data landed on the wrong disk. It is therefore **not** a reason to leave the sweep unguarded. ## The harm that DID occur Independent of the write question, and real on every box: - the sweep burns **both** attempts and 30 s of retry delay on an app that cannot start by design; - it then hands the app to the **dead-app alarm** — `still down=[calibre-web] (the dead-app alarm now owns these)` — producing a **false alarm about an app the drive gate is deliberately holding**, which is exactly the noise class R-97/F-A1 exist to prevent; - it leaves a `Created` container behind on each attempt. Before v0.189.0 none of this happened: `isBootOrphan` required `len(Containers) > 0`, and a gate-stopped app has zero. **This is a regression introduced by v0.189.0.** ## A second asymmetry found while restoring the box The **API** start path already refuses this correctly. Restoring `immich` through the endpoint the UI calls, while the path was still flagged disconnected, returned: ``` {"ok":false,"error":"A(z) /mnt/felhom-drives/hdd_1 tárhely jelenleg nem elérhető — az alkalmazás nem indítható, amíg a meghajtó vissza nem csatlakozik."} ``` That is `startGatedByMissingDrive` (`internal/api/router.go`). So the controller already holds the rule "do not start an app whose drive is missing" — it is enforced on the customer's path and **bypassed by the sweep**, which calls `Manager.StartStack` directly. The fix is to give the sweep the same question to ask, not to invent a new rule. ## Consequence for the task **Part 3 applies and is implemented first**, before the sweep's window is widened — a wider window makes both the false alarm and the (currently accident-blocked) write hazard wider. Fail-safe direction per §8.4: **cannot determine drive liveness → do not start.** Not starting is recoverable (the gate's `Return` branch restarts the app when the drive comes back, and the alarm reports it meanwhile); starting on an absent drive is not recoverable by anything automatic. Register row: **R-171**, marked as a regression from v0.189.0.