From 67356c9e5caa59e215b16a51ee210ef94d63e8ed Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 21 Aug 2026 21:25:07 +0200 Subject: [PATCH] register: R-351 CLOSED, R-352 partly, R-353 OPEN (next session's first item) + placement spec R-351 - the restore never read back where the backup said the data lived, and a second press started a second restore. Both shipped in controller v0.217.0. R-352 - four measured untruths about where an app's data goes: 1. 40 of 53 catalogue templates declare no data path (13 declare env_var: HDD_PATH) 2. GetDefaultStoragePath() has three non-test callers and NONE of them places data; its field comment "new apps use this by default" has never been true 3. the first-tier backup follows the data onto the same disk (GetAppDrivePath -> systemDataPath) - the posture Tier 2 refuses outright at tier2.go:329 4. "1 alkalmazas hasznalja" counts only Env["HDD_PATH"] == path, so it can never include the 40-class; it means "1 of the apps that CAN use a drive does" Visibility shipped tonight; PLACEMENT IS OPEN and is the operator's ruling. An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN - it assumed the customer had failed to choose, and they had no choice to make. R-353 - a restore reported success having returned configuration and no data. OpenGist's unit holds manifest.json + compose/ and nothing else (volume_dumps: None, db_dumps: None); the off-site snapshot was 182.3 KB; the outcome said only "completed in 8.666896042s". Ranked as the NEXT SESSION'S FIRST ITEM. Compounding and recorded as UNKNOWN rather than fine: whether the 40-class reaches the off-site tier at all has not been observed - runVolumeDumps covers them on paper, but no nightly dump run had happened on a one-hour-old box. New: documentation/backlog/SPEC-app-data-placement-2026-08-21.md - specification only, nothing implemented, listing the five points a placement ruling must settle. Records that the OpenGist instance meant to be left as evidence was removed by someone between 16:57 and 17:02 UTC (not by this session); its unit and manifest survive, and privatebin is now a live specimen. Ceiling moved R-350 -> R-353. --- documentation/backlog/OPEN-ITEMS.md | 3 + .../SPEC-app-data-placement-2026-08-21.md | 162 ++++++++++++++++++ 2 files changed, 165 insertions(+) create mode 100644 documentation/backlog/SPEC-app-data-placement-2026-08-21.md diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 78502532..4d6bd24c 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -650,6 +650,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | | **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | | **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes | +| **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** Two findings, one session, both shipped. **(a) The blindness.** Every recovery-unit `manifest.json` has carried `drive` and `namespace_root` since schema 1 (`controller/internal/backup/recovery_unit.go:48-49`), written at capture from the app's own placement. `grep -rE '\.Drive\b\|\.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | **CLOSED 2026-08-21** — controller | — | **Shipped:** `backup/offbox_placement.go` (`CheckPlacement`, `PlacementMismatchMessage`, `RecordedUnitForStack`); mismatch **named and refused** before the safety dump, with `ack_placement` as a **separate** field from `confirm=1`; the not-installed refusal names the recorded drive; deploy page **prefills the address and folder from the app's own backup**; `Server.restoreOpBlocked()` reads BOTH flags; `RestoreOpStatus.LastRecent` + `RestoreResultWindow` moved to `internal/backup` as ONE expression for two surfaces. Red-proofs: B with **both** guards removed **was seen starting a restore with no drive attached** (no error, full 3.00 s run into `/tmp/mutant-destination`); C, E and A each returned their wrong outcome; D forced on broke 8 ordinary reconstitute tests, proving reachability both ways. | CC | +| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | +| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. | CC | | **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC | | **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc//fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC | diff --git a/documentation/backlog/SPEC-app-data-placement-2026-08-21.md b/documentation/backlog/SPEC-app-data-placement-2026-08-21.md new file mode 100644 index 00000000..ce6f9e05 --- /dev/null +++ b/documentation/backlog/SPEC-app-data-placement-2026-08-21.md @@ -0,0 +1,162 @@ +# SPEC — where an app's data is placed, and who decides + +**Filed 2026-08-21. Status: SPECIFICATION ONLY — nothing here is implemented, and nothing here may be +implemented without the operator's ruling. It deserves its own session.** + +Measured on `demo-hp` (HP t740, controller 0.216.0, agent 0.130.0) on 2026-08-21, during and after the +box was reinstalled. Every claim below carries the `file:line` or the live observation it came from. + +--- + +## 1. What was found + +An operator walking the deploy screens saw that OpenGist's deploy page offered **only Domain and +Subdomain** — no storage field of any kind — while the Drives page showed the NVMe at +`/mnt/felhom-drives/hdd_1` marked **Alapértelmezett** and **Aktív**, with **"1 alkalmazás használja"**. +OpenGist's data and its first-tier backup were both under `/mnt/sys_drive/`. + +The first hypothesis — that a customer had typed a bad path, or had failed to choose a drive — is +**wrong**. There is nothing to type and nothing to choose. The finding is narrower and worse: + +> **The configured default data store is not consulted on the deploy route at all.** + +## 2. The four measured facts + +### 2.1 The affected class is 40 of 53 catalogue templates + +``` +find . -name '.felhom.yml' | wc -l -> 53 +grep -rl 'env_var: HDD_PATH' --include='.felhom.yml' . | wc -> 13 + do not -> 40 +``` + +The 13 that declare a data path: `audiobookshelf calibre-web emby immich jellyfin komga navidrome +nextcloud paperless-ngx plex radarr romm sonarr`. All media libraries. The same 13 carry a `backup:` +block. + +Negative control: `grep -c HDD_PATH templates/opengist/{.felhom.yml,docker-compose.yml}` returns **0** +for both — such an app cannot receive an `HDD_PATH` even if one were supplied. + +### 2.2 The default store is a preference with no effect on what it names + +`settings.GetDefaultStoragePath()` (`hub`-side equivalent none; controller +`internal/settings/settings.go:1196`) has exactly **three** non-test callers: + +| Caller | What it actually decides | +|---|---| +| `controller/cmd/controller/main.go:410` | which drive the **metrics collector** measures | +| `controller/internal/web/server.go:733` (`primaryHDDPath`) | the dashboard **SystemInfo** panel (`handlers.go:176, 732, 876`) | +| `controller/internal/web/handler_export_upload.go:154` | where an uploaded `.fab` **import** lands | + +`grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' controller/internal/stacks/deploy.go +controller/internal/stacks/manager.go` returns **nothing**. + +The field's own comment at `internal/settings/settings.go:453` reads `// new apps use this by default`. +**No new app has ever used it.** That comment is an invariant with no test pinning it — the class +CLAUDE.md names. + +Where the data actually goes: a **named Docker volume** on the guest root filesystem. Verified live: + +``` +privatebin volume src=/var/lib/docker/volumes/privatebin_privatebin_data/_data +calibre-web bind src=/mnt/felhom-drives/hdd_1/userdata/media/books +``` + +`withPathVars` (`internal/stacks/deploy.go:600-607`) injects `USERDATA_PATH` **only if `hdd != ""`**. + +### 2.3 The first-tier backup follows the data onto the same disk + +`GetAppDrivePath` (`internal/backup/backup.go:324-334`): no `HDD_PATH` → returns `m.systemDataPath`. +Live: + +``` +/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json + drive='/mnt/sys_drive' namespace_root='/mnt/sys_drive/felhom-data' +``` + +So for these 40 apps **the data and its nearest copy sit on the same physical device**, reached by a +customer doing nothing wrong. + +The project already treats that posture as unacceptable — for the *other* tier. Tier 2 refuses it +outright at `internal/backup/tier2.go:329`, recording +`a kiválasztott cél ugyanazon a fizikai lemezen van`. Tier 1 has no such notion, and for a +drive-resident app it is correct that it has none: the unit is meant to live beside the data so a +restore needs the drive and nothing else. The defect is not Tier 1's rule; it is that these apps are +on the system drive in the first place. + +### 2.4 The Drives page count cannot include most apps + +`countAppsUsingPath` (`internal/web/handlers.go:2118-2131`) counts only +`appCfg.Env["HDD_PATH"] == storagePath`. An app with no `HDD_PATH` can never match any drive. + +So **"1 alkalmazás használja" truthfully means "1 of the apps that CAN use a drive does"**, and the +40-of-53 class is invisible on that page. The code already names the class deliberately at +`handlers.go:2140`: `// An app with no HDD_PATH (SSD-resident) is never "missing".` This is an +unstated design, not an accident. + +## 3. Is that class protected? + +**Whole-machine tier: yes.** Verified by `df`, not assumed — `/mnt/sys_drive`, `/var/lib/docker` and +`/var/lib/felhom` are all on `pve-vm-9201-disk-1`, which is `mp0` in `/etc/pve/lxc/9201.conf` with +`backup=1`. Named volumes and sys-drive units land in the guest backup. + +**Off-site tier: covered by code, NOT demonstrated.** `runVolumeDumps` +(`internal/backup/backup.go:607+`) iterates every deployed unprotected stack with named volumes; its +drive-state gates use `GetAppDrivePath`, which returns the system path — neither disconnected nor +decommissioned, so the gates pass. On paper these apps are dumped. + +**It has not been seen happen.** All three units on the box reported `volume_dumps: None, +db_dumps: None` — including `calibre-web`, which is on the data drive. No nightly dump run had +occurred on a box one hour old. **This is recorded as unknown rather than fine.** + +## 4. What is ruled, and what is not + +**Ruled and shipped 2026-08-21 (R-351, controller):** the deploy page now **states where the app's +data will live before the button is pressed** — naming the system drive for the 40-class, and the +selected drive for the 13. Visibility only. **No placement changed. Nothing was migrated.** + +**Explicitly NOT done, and not to be done without a ruling:** + +- **No storage selector was added for apps that do not need one.** The field is not the point; the + placement is. Adding a drive dropdown to 40 apps whose compose never references a path would be a + control that changes nothing. +- **Deployment is NOT refused when no drive is registered.** An earlier draft of this session + recommended that and it was **withdrawn**: it was built on the belief that the customer had failed + to choose. They had no choice to make. Refusing 40 of 53 apps for a drive they cannot use would + break the ordinary path to fix a hazard the customer never touched. +- **No placement change, no migration.** Moving where apps write has consequences for every existing + deployment on every box in the fleet. + +## 5. The open question this document exists to hand over + +**Should a named-volume app's data live on the default data drive rather than the system drive?** + +Points the next session must settle, each of which is a reason this was not decided tonight: + +1. **It is a compose-template question, not only a controller question.** A named volume + (`opengist_data:/opengist`) has no path to redirect. Either the templates gain binds under + `${USERDATA_PATH}` — a 40-template catalogue change — or Docker's data-root moves, which relocates + *every* container's storage including the controller's own. +2. **Existing deployments.** Any change must answer what happens to the apps already running on the + system drive. Leaving them and changing only new installs creates two classes with no visible + difference — which is how this defect became invisible in the first place. +3. **The system drive is not always wrong.** A small config-only app on the SSD is a reasonable + placement; a media library is not. A rule that says "always the data drive" would move things that + were fine. +4. **`IsDefault` must either be consulted or removed.** A setting that names a behaviour it does not + have is worse than no setting. Whichever way the placement question goes, that field's comment at + `settings.go:453` has to become true or go away — **with a test pinning it**, since it has been a + wish since it was written. +5. **The Drives page count needs the same decision.** Whatever the rule becomes, the page must stop + implying that the apps it does not count are not using storage. + +## 6. Evidence deliberately left in place + +The OpenGist instance on `demo-hp` was to be left exactly where it was, as a real example of the +defect on real hardware. **It was removed by someone between 16:57 and 17:02 UTC on 2026-08-21** — +`ScanStacks: found stack "opengist" deployed=false` from 17:02:58 onward — not by this session. + +**Its evidence survives:** the recovery unit and manifest at +`/mnt/sys_drive/felhom-data/backups/primary/opengist/` are intact and carry +`drive='/mnt/sys_drive'`. **`privatebin` is now a live specimen of the same class** on the same box, +with its data at `/var/lib/docker/volumes/privatebin_privatebin_data/_data`.