091a4b7444
gates / gates (push) Successful in 16s
Documentation and survey only. No code, no machine contacted. THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands. R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so reading it as "not installed" misreads a correct configuration. The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy together is REAL and is what the other tiers exist for. A full data volume stopping the OS is NOT real and was the overstated one. THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed. Its positive control convicted the sweep itself twice before it convicted the corpus - markdown bold broke the strongest pattern, and the reporter re-searched a truncated line - both false zeros of the exact class being hunted, and together worth 2 of the 14. THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122, M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH). Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373 (20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go and never the templates. R-370 records the process failure and is closed by the template change. PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and say what it says (with a file->area map and the test "is this something we chose?"), and an enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count. Ceiling R-367 -> R-375.
249 lines
14 KiB
Markdown
249 lines
14 KiB
Markdown
# SPEC — where an app's data is placed, and who decides
|
|
|
|
> ## ⚠ CORRECTED 2026-08-22 — THE MEASUREMENTS STAND, THE FRAMING WAS WRONG
|
|
>
|
|
> **This document called a documented architectural decision a defect.** It is left in place rather
|
|
> than rewritten, because a document that quietly changes its mind teaches nobody. Every measurement
|
|
> below is good and was re-checked on 2026-08-22; the corrections are marked inline as
|
|
> **[CORRECTED 2026-08-22]**.
|
|
>
|
|
> **What it got wrong.** It treated the absence of a storage field on 40 of 53 templates as a choice
|
|
> being denied. There is no choice to deny. The architecture states the placement rule and states it
|
|
> as *enforced*:
|
|
>
|
|
> > *"**App data placement is per-volume, not per-app:** `.felhom.yml` classifies each volume **hot**
|
|
> > (DB/config/cache → fast storage, **enforced**) vs **bulk** (media/files → may be slow). A photo
|
|
> > app's DB stays on SSD while its blobs go to the USB."*
|
|
> > — `documentation/architecture/01-topology-and-trust.md:150-152`
|
|
>
|
|
> The 40 templates are **all-hot apps**: their data *is* app internals — database, config, cache —
|
|
> and hot data belongs on fast storage inside the guest, by design. The 13 with a configurable path
|
|
> are the media and library apps, where **bulk** content legitimately belongs on an attached drive.
|
|
> The split is not 40 apps missing a feature; it is the hot/bulk contract working.
|
|
>
|
|
> **And the guest volume is deliberately its own volume**, so that filling it cannot stop the
|
|
> operating system: the golden ships a small OS rootfs plus a single data volume at `/var/lib/felhom`
|
|
> (`felhom-agent/configs/build-golden.sh:29-40`), and a capture that would exhaust it is refused per
|
|
> app rather than allowed to stop the container runtime
|
|
> (`documentation/architecture/00-capability-map.md:94`, PROVEN-LIVE 2026-08-03).
|
|
>
|
|
> **One real defect survives the correction and is unchanged:** `settings.go:453` says
|
|
> `// new apps use this by default` and no app has ever used it. That is §2.2 below and it is now
|
|
> filed on its own, scoped by measurement rather than by assumption — see the register.
|
|
>
|
|
> **What still needs a ruling: far less than §5 implies.** See the corrected §5.
|
|
|
|
**Filed 2026-08-21. Status: SPECIFICATION ONLY — nothing here is implemented, and nothing here may be
|
|
implemented without the operator's ruling. It deserves its own session.**
|
|
|
|
Measured on `demo-hp` (HP t740, controller 0.216.0, agent 0.130.0) on 2026-08-21, during and after the
|
|
box was reinstalled. Every claim below carries the `file:line` or the live observation it came from.
|
|
|
|
---
|
|
|
|
## 1. What was found
|
|
|
|
An operator walking the deploy screens saw that OpenGist's deploy page offered **only Domain and
|
|
Subdomain** — no storage field of any kind — while the Drives page showed the NVMe at
|
|
`/mnt/felhom-drives/hdd_1` marked **Alapértelmezett** and **Aktív**, with **"1 alkalmazás használja"**.
|
|
OpenGist's data and its first-tier backup were both under `/mnt/sys_drive/`.
|
|
|
|
The first hypothesis — that a customer had typed a bad path, or had failed to choose a drive — is
|
|
**wrong**. There is nothing to type and nothing to choose. The finding is narrower and worse:
|
|
|
|
> **The configured default data store is not consulted on the deploy route at all.**
|
|
|
|
> **[CORRECTED 2026-08-22]** The observation is right; the conclusion drawn from it was not. There is
|
|
> nothing to choose **because these apps have no bulk volume to place** — the hot/bulk contract
|
|
> (`01-topology-and-trust.md:150-152`) puts their data on fast storage inside the guest by design, and
|
|
> OpenGist is a pure hot app. The sentence in the quote above is still true as a statement about the
|
|
> deploy route, and it is still a defect **for the 13 templates that DO take a path** — that is the
|
|
> real remainder, established in Part 4 of the 2026-08-22 session and filed separately. For the other
|
|
> 40 it is not a defect at all: a default drive cannot be "consulted" for an app that has no path to
|
|
> put it in.
|
|
|
|
## 2. The four measured facts
|
|
|
|
### 2.1 The affected class is 40 of 53 catalogue templates
|
|
|
|
```
|
|
find . -name '.felhom.yml' | wc -l -> 53
|
|
grep -rl 'env_var: HDD_PATH' --include='.felhom.yml' . | wc -> 13
|
|
do not -> 40
|
|
```
|
|
|
|
The 13 that declare a data path: `audiobookshelf calibre-web emby immich jellyfin komga navidrome
|
|
nextcloud paperless-ngx plex radarr romm sonarr`. All media libraries. The same 13 carry a `backup:`
|
|
block.
|
|
|
|
Negative control: `grep -c HDD_PATH templates/opengist/{.felhom.yml,docker-compose.yml}` returns **0**
|
|
for both — such an app cannot receive an `HDD_PATH` even if one were supplied.
|
|
|
|
### 2.2 The default store is a preference with no effect on what it names
|
|
|
|
`settings.GetDefaultStoragePath()` (`hub`-side equivalent none; controller
|
|
`internal/settings/settings.go:1196`) has exactly **three** non-test callers:
|
|
|
|
| Caller | What it actually decides |
|
|
|---|---|
|
|
| `controller/cmd/controller/main.go:410` | which drive the **metrics collector** measures |
|
|
| `controller/internal/web/server.go:733` (`primaryHDDPath`) | the dashboard **SystemInfo** panel (`handlers.go:176, 732, 876`) |
|
|
| `controller/internal/web/handler_export_upload.go:154` | where an uploaded `.fab` **import** lands |
|
|
|
|
`grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' controller/internal/stacks/deploy.go
|
|
controller/internal/stacks/manager.go` returns **nothing**.
|
|
|
|
The field's own comment at `internal/settings/settings.go:453` reads `// new apps use this by default`.
|
|
**No new app has ever used it.** That comment is an invariant with no test pinning it — the class
|
|
CLAUDE.md names.
|
|
|
|
Where the data actually goes: a **named Docker volume** on the guest root filesystem. Verified live:
|
|
|
|
```
|
|
privatebin volume src=/var/lib/docker/volumes/privatebin_privatebin_data/_data
|
|
calibre-web bind src=/mnt/felhom-drives/hdd_1/userdata/media/books
|
|
```
|
|
|
|
`withPathVars` (`internal/stacks/deploy.go:600-607`) injects `USERDATA_PATH` **only if `hdd != ""`**.
|
|
|
|
### 2.3 The first-tier backup follows the data onto the same disk
|
|
|
|
`GetAppDrivePath` (`internal/backup/backup.go:324-334`): no `HDD_PATH` → returns `m.systemDataPath`.
|
|
Live:
|
|
|
|
```
|
|
/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json
|
|
drive='/mnt/sys_drive' namespace_root='/mnt/sys_drive/felhom-data'
|
|
```
|
|
|
|
So for these 40 apps **the data and its nearest copy sit on the same physical device**, reached by a
|
|
customer doing nothing wrong.
|
|
|
|
> **[CORRECTED 2026-08-22] — and be precise about which risk is real.**
|
|
>
|
|
> **The device claim is true. The framing "by accident" was not.** Since R-165 (golden `build-golden.sh`
|
|
> v3.0.0, proven live 2026-08-03) the guest carries a small OS rootfs plus **ONE** data volume at
|
|
> `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume**, and
|
|
> the script says so in terms — *"There is deliberately no mp1"* (`build-golden.sh:99`). The
|
|
> `mp0`/`mp1` split this document's era assumed was retired six days before this document was written
|
|
> and is on no current box.
|
|
>
|
|
> **Which risk is real, stated exactly:**
|
|
> - **A failure of the physical disk loses both** the app's data and its first-tier copy. **REAL, and
|
|
> it is what the off-site and whole-machine tiers exist for.** It is not specific to the 40-class:
|
|
> a drive-resident app's unit lives beside its data on the drive too, deliberately, so a restore
|
|
> needs the drive and nothing else.
|
|
> - **A full data volume taking out the operating system. NOT REAL, and it was the overstated one.**
|
|
> The OS rootfs is a separate volume, and the capture floor refuses a capture that would exhaust the
|
|
> data volume rather than letting it stop the container runtime (`00-capability-map.md:94`). Watched
|
|
> working on 2026-08-21: with `/var/lib/felhom` at 99% the reserve refused one app per run and told
|
|
> the hub, while all 15 containers stayed healthy.
|
|
>
|
|
> So the sentence *"The defect is not Tier 1's rule; it is that these apps are on the system drive in
|
|
> the first place"* below is **withdrawn**. Being on the guest's data volume is where hot data belongs.
|
|
|
|
The project already treats that posture as unacceptable — for the *other* tier. Tier 2 refuses it
|
|
outright at `internal/backup/tier2.go:329`, recording
|
|
`a kiválasztott cél ugyanazon a fizikai lemezen van`. Tier 1 has no such notion, and for a
|
|
drive-resident app it is correct that it has none: the unit is meant to live beside the data so a
|
|
restore needs the drive and nothing else. The defect is not Tier 1's rule; it is that these apps are
|
|
on the system drive in the first place.
|
|
|
|
### 2.4 The Drives page count cannot include most apps
|
|
|
|
`countAppsUsingPath` (`internal/web/handlers.go:2118-2131`) counts only
|
|
`appCfg.Env["HDD_PATH"] == storagePath`. An app with no `HDD_PATH` can never match any drive.
|
|
|
|
So **"1 alkalmazás használja" truthfully means "1 of the apps that CAN use a drive does"**, and the
|
|
40-of-53 class is invisible on that page. The code already names the class deliberately at
|
|
`handlers.go:2140`: `// An app with no HDD_PATH (SSD-resident) is never "missing".` This is an
|
|
unstated design, not an accident.
|
|
|
|
## 3. Is that class protected?
|
|
|
|
**Whole-machine tier: yes.** Verified by `df`, not assumed — `/mnt/sys_drive`, `/var/lib/docker` and
|
|
`/var/lib/felhom` are all on `pve-vm-9201-disk-1`, which is `mp0` in `/etc/pve/lxc/9201.conf` with
|
|
`backup=1`. Named volumes and sys-drive units land in the guest backup.
|
|
|
|
**Off-site tier: covered by code, NOT demonstrated.** `runVolumeDumps`
|
|
(`internal/backup/backup.go:607+`) iterates every deployed unprotected stack with named volumes; its
|
|
drive-state gates use `GetAppDrivePath`, which returns the system path — neither disconnected nor
|
|
decommissioned, so the gates pass. On paper these apps are dumped.
|
|
|
|
**It has not been seen happen.** All three units on the box reported `volume_dumps: None,
|
|
db_dumps: None` — including `calibre-web`, which is on the data drive. No nightly dump run had
|
|
occurred on a box one hour old. **This is recorded as unknown rather than fine.**
|
|
|
|
## 4. What is ruled, and what is not
|
|
|
|
**Ruled and shipped 2026-08-21 (R-351, controller):** the deploy page now **states where the app's
|
|
data will live before the button is pressed** — naming the system drive for the 40-class, and the
|
|
selected drive for the 13. Visibility only. **No placement changed. Nothing was migrated.**
|
|
|
|
**Explicitly NOT done, and not to be done without a ruling:**
|
|
|
|
- **No storage selector was added for apps that do not need one.** The field is not the point; the
|
|
placement is. Adding a drive dropdown to 40 apps whose compose never references a path would be a
|
|
control that changes nothing.
|
|
- **Deployment is NOT refused when no drive is registered.** An earlier draft of this session
|
|
recommended that and it was **withdrawn**: it was built on the belief that the customer had failed
|
|
to choose. They had no choice to make. Refusing 40 of 53 apps for a drive they cannot use would
|
|
break the ordinary path to fix a hazard the customer never touched.
|
|
- **No placement change, no migration.** Moving where apps write has consequences for every existing
|
|
deployment on every box in the fleet.
|
|
|
|
## 5. The open question this document exists to hand over
|
|
|
|
> **[CORRECTED 2026-08-22] — THIS QUESTION IS ALREADY ANSWERED, AND THE ANSWER IS NO.**
|
|
>
|
|
> The architecture answers it: hot data (DB/config/cache) belongs on fast storage inside the guest and
|
|
> that placement is **enforced**, not preferred (`01-topology-and-trust.md:150-152`). A named-volume
|
|
> app is all-hot. Moving it to the attached drive would be moving hot data onto storage the same
|
|
> document classes as possibly-slow, and would put it outside the guest vzdump that currently protects
|
|
> it (`01-topology-and-trust.md:154-156`).
|
|
>
|
|
> **So points 1, 2 and 3 below are withdrawn as a live question.** They remain a good record of what a
|
|
> change would have cost, which is why they are not deleted. Point 3 was already arguing against the
|
|
> premise of the question it sat under.
|
|
>
|
|
> **WHAT ACTUALLY STILL NEEDS YOUR RULING FROM THIS DOCUMENT: point 4 only, and it is smaller than it
|
|
> looks** — `IsDefault`'s comment must become true or go away, for the 13 apps it could apply to. That
|
|
> is filed as its own row with its scope measured rather than assumed. Point 5 (the Drives page count)
|
|
> needs **no ruling**: the count is honest, it is only easy to misread, and that is a wording question
|
|
> the design system already owns.
|
|
>
|
|
> **Everything else in this section: nothing to decide.**
|
|
|
|
**Should a named-volume app's data live on the default data drive rather than the system drive?**
|
|
*(superseded — see the correction directly above)*
|
|
|
|
Points the next session must settle, each of which is a reason this was not decided tonight:
|
|
|
|
1. **It is a compose-template question, not only a controller question.** A named volume
|
|
(`opengist_data:/opengist`) has no path to redirect. Either the templates gain binds under
|
|
`${USERDATA_PATH}` — a 40-template catalogue change — or Docker's data-root moves, which relocates
|
|
*every* container's storage including the controller's own.
|
|
2. **Existing deployments.** Any change must answer what happens to the apps already running on the
|
|
system drive. Leaving them and changing only new installs creates two classes with no visible
|
|
difference — which is how this defect became invisible in the first place.
|
|
3. **The system drive is not always wrong.** A small config-only app on the SSD is a reasonable
|
|
placement; a media library is not. A rule that says "always the data drive" would move things that
|
|
were fine.
|
|
4. **`IsDefault` must either be consulted or removed.** A setting that names a behaviour it does not
|
|
have is worse than no setting. Whichever way the placement question goes, that field's comment at
|
|
`settings.go:453` has to become true or go away — **with a test pinning it**, since it has been a
|
|
wish since it was written.
|
|
5. **The Drives page count needs the same decision.** Whatever the rule becomes, the page must stop
|
|
implying that the apps it does not count are not using storage.
|
|
|
|
## 6. Evidence deliberately left in place
|
|
|
|
The OpenGist instance on `demo-hp` was to be left exactly where it was, as a real example of the
|
|
defect on real hardware. **It was removed by someone between 16:57 and 17:02 UTC on 2026-08-21** —
|
|
`ScanStacks: found stack "opengist" deployed=false` from 17:02:58 onward — not by this session.
|
|
|
|
**Its evidence survives:** the recovery unit and manifest at
|
|
`/mnt/sys_drive/felhom-data/backups/primary/opengist/` are intact and carry
|
|
`drive='/mnt/sys_drive'`. **`privatebin` is now a live specimen of the same class** on the same box,
|
|
with its data at `/var/lib/docker/volumes/privatebin_privatebin_data/_data`.
|