Correct the placement mis-framing, and file what we wrote down and never filed (R-368..R-375)
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
Documentation and survey only. No code, no machine contacted. THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands. R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so reading it as "not installed" misreads a correct configuration. The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy together is REAL and is what the other tiers exist for. A full data volume stopping the OS is NOT real and was the overstated one. THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed. Its positive control convicted the sweep itself twice before it convicted the corpus - markdown bold broke the strongest pattern, and the reporter re-searched a truncated line - both false zeros of the exact class being hunted, and together worth 2 of the 14. THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122, M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH). Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373 (20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go and never the templates. R-370 records the process failure and is closed by the template change. PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and say what it says (with a file->area map and the test "is this something we chose?"), and an enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count. Ceiling R-367 -> R-375.
This commit is contained in:
@@ -651,11 +651,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC |
|
||||
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes |
|
||||
| **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** Two findings, one session, both shipped. **(a) The blindness.** Every recovery-unit `manifest.json` has carried `drive` and `namespace_root` since schema 1 (`controller/internal/backup/recovery_unit.go:48-49`), written at capture from the app's own placement. `grep -rE '\.Drive\b\|\.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | **CLOSED 2026-08-21** — controller | — | **Shipped:** `backup/offbox_placement.go` (`CheckPlacement`, `PlacementMismatchMessage`, `RecordedUnitForStack`); mismatch **named and refused** before the safety dump, with `ack_placement` as a **separate** field from `confirm=1`; the not-installed refusal names the recorded drive; deploy page **prefills the address and folder from the app's own backup**; `Server.restoreOpBlocked()` reads BOTH flags; `RestoreOpStatus.LastRecent` + `RestoreResultWindow` moved to `internal/backup` as ONE expression for two surfaces. Red-proofs: B with **both** guards removed **was seen starting a restore with no drive attached** (no error, full 3.00 s run into `/tmp/mutant-destination`); C, E and A each returned their wrong outcome; D forced on broke 8 ordinary reconstitute tests, proving reachability both ways. | CC |
|
||||
| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
|
||||
| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
|
||||
| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC |
|
||||
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | `restoreDockerVolumesFrom` is the LOCAL path's own replay with an explicit directory — ONE implementation, two callers, because a second copy of that loop is what produced the divergence. It reads the SCRATCH unit; the live unit is still never written. Volumes replay BEFORE the database (a logical dump must still win over a volume-tar copy of the same database) and inside the stopped window (Docker will not replace a volume a container holds). `VolumesReplayed` is on the result and in the sentence. **The comment beside the skip was half false and is corrected, not left:** it justified the skip by saying the dump is replayed from the scratch "so nothing is lost" — true of the database, false of the volumes. The half that still holds (the live unit is the local path's source) is named. **PROVEN LIVE on `demo-hp`, negative control first.** On 0.217.0 with the fixture planted, hashed and then deleted from the live volume: „A(z) calibre-web: **0 fájl visszaállítva** … — az alkalmazás újraindult.", `ok=true`, and `ls` reported the directory absent. On 0.218.0, same fixture, same steps: „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** …" and **5/5 files byte-identical**, both Hungarian accented filenames included. **Four red-proofs, each mutation asserted applied:** the volume leg removed returned the silent loss; the count dropped from the message returned the true-but-incomplete sentence verbatim; the unit guard removed was SEEN writing into the live unit; the error swallowed let a partial replay report success. **The unit-guard proof initially PASSED against the mutation** — the fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit root (the R-181 class, in the test rather than the code). Widened, and it convicts. **NOT reached by this fix, and it is the blocker:** the 40 apps that declare no data drive still cannot run this restore at all — **R-356** refuses first. | CC |
|
||||
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | **Neither candidate: a third that removes the guessing.** Every container the controller starts carries `com.docker.compose.project`, and that label IS the stack name BY CONSTRUCTION — compose is run with `cmd.Dir` set to `/opt/docker/stacks/<stack>` and never `-p` (`stacks/manager.go:1218`). `resolveStackName` prefers it whenever it names a deployed stack; `deriveStackName` stays as the fallback for containers not started by compose; an attribution that resolves to NO known stack is now **loud** instead of silently returned. **The fix is in the CONTROLLER, not the catalogue** — renaming the container would have fixed this one app and left the guessing for the next. **The sweep was proven before its answer was trusted:** a second mismatch planted in a scratch copy (`kimai-db`→`timetrack-db`) was convicted by name, removal returned it to 1, and a catalogue with every mismatch removed exits **0** — so "1" is not a stuck value. **1 affected app of 53**, 15 DB containers checked. **PROVEN LIVE on `demo-hp`, negative control first.** 0.217.0 at 09:37: unit `db_dumps = None`, dump refreshed into the phantom `…/primary/paperless/db-dumps/paperless-postgres.sql`. 0.218.0 at 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` **inside the app's own unit**, and present in the off-site snapshot for the first time. Scenario B: the destructive restore said „**0 fájl és 3 adatkötet és az adatbázis visszaállítva**" and wrote `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql` (312 957 B) where **zero** undo copies had existed. Scenario C: with the undo made impossible the restore REFUSED, the marker kept its mutation and the container's `StartedAt` was unchanged — **the app was never stopped**. Scenario D: `romm-mariadb.sql` and `kimai-mariadb.sql` unchanged in name and location. **Three red-proofs, each asserted applied:** the name fix reverted printed both divergent paths; the message predicate reverted returned the false „nincs adatbázisa" sentence verbatim; the refusal removed was seen letting a restore proceed with no undo. **The dumps already written under the wrong name are NOT deleted** — see **R-367**. | CC |
|
||||
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. | CC |
|
||||
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. **RE-CHECKED 2026-08-22 against the corrected placement framing: this row SURVIVES UNCHANGED and is strengthened by it.** The 40-class has no `HDD_PATH` *because the architecture puts its hot data inside the guest by design* (`01-topology-and-trust.md:150-152`) — so an absent HDD path is the NORMAL state for most of the catalogue, and a restore that reads it as "the app is not installed" is misreading a correct configuration, not reporting a misconfiguration. The one sentence in this row that leaned on the old framing — "because they are offered no storage field at deploy time (R-352's own measurement)" — should read: *because these apps have no bulk volume to place, so the field correctly does not exist.* | CC |
|
||||
| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC |
|
||||
| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC |
|
||||
| **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC |
|
||||
@@ -669,6 +669,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-367** | **The database dumps already written under the wrong name are stranded, and nothing will ever collect them.** R-355's fix sends `paperless-ngx`'s dump to the right place from now on; it does not move the ones already written. On `demo-hp` that is `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql` (312 381 B, 2026-08-22 07:38, the last pre-fix cycle). **Nothing deletes them and that is by design, not by luck:** the F5 stale-primary prune (`backup.go:1248`) skips any directory whose name is not a deployed app, under the guard *"an undeployed app's last backup is still its restore point"* — verified still present after the fix. They are equally invisible to the off-site push, which resolves paths from the app's own unit. **They CAN be adopted, by hand:** move the file to `…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point for that app. **It is deliberately not automatic.** The adopted dump would sit beside volume tars taken at a different time, i.e. an INCOHERENT pair — the exact shape R-43/R-44's coherence stamp exists to make visible — and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. **Filed rather than done**, because whether a stale orphan is worth adopting at all is a judgement about one machine's history, not a rule. | **OPEN — LOW** | follows R-355 | Decide per box: adopt (and say the pair is skewed), or delete deliberately. Neither on an upgrade path. | Viktor rules, CC executes |
|
||||
|
||||
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC |
|
||||
| **R-368** | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks/<n>/deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC |
|
||||
| **R-369** | **There are TWO registers, only one calls itself the source of truth, and work filed in the other is invisible to every standing rule that says "grep the register".** `OPEN-ITEMS.md` opens *"the single source of truth for open work"*; `ROADMAP.md` opens *"the prioritized decision log of planned/open work"*. Both hold open work. **Measured 2026-08-22: 72 `R-` ids appear in ROADMAP and not in OPEN-ITEMS; 29 of those are not marked shipped/closed/killed.** Most are feature ideas that arguably belong only in ROADMAP — but some are FINDINGS: R-30/R-31/R-32 (all marked P2-HIGH), R-35, R-40, R-76, R-79, R-25, R-49, R-10 and **R-107**. **The cost is measured, not hypothetical:** R-107 — *"No offsite action unpacks the named-volume tars Tier-3 captures on every run"* — was filed **2026-07-28, READY, severity M**, and re-stated in `07-backup-architecture.md:337,902`. It is absent from OPEN-ITEMS. On 2026-08-21 an overnight drill rediscovered it from scratch by planting files and watching them not come back, and it shipped as R-354 on 2026-08-22. **25 days.** **The dating is exact and makes it a rule violation rather than a gap in the rules:** OPEN-ITEMS.md was created and the template gained its *"single source of truth"* bullet on **2026-07-27** (`655b69f`); R-107 went into ROADMAP alone on **2026-07-28** (`070b0ce`) — the day after. **The risk is NOT double-minting** (ids are shared; OPEN-ITEMS' ceiling 368 exceeds ROADMAP's 331) — it is that a session greps one file, finds nothing, and redoes the work. | **OPEN — HIGH** | R-107 (ROADMAP-only), R-354 | Decide the division and make it mechanical: either one register, or a gate that fails when a ROADMAP row that is not shipped/killed has no OPEN-ITEMS counterpart. Triage the 29 first — most are ideas, a minority are findings. | **Viktor rules**, CC executes |
|
||||
| **R-370** | **PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never `documentation/architecture/`.** The decision: `.felhom.yml` classes each volume **hot** (DB/config/cache → fast storage, **enforced**) or **bulk** (media/files → may be slow) — `01-topology-and-trust.md:150-152`. The 40 templates with no `HDD_PATH` are all-hot apps; the 13 with one are the bulk apps. **The four instances**, all authored 2026-08-21: `SPEC-app-data-placement-2026-08-21.md` §1 (absence of a storage field framed as the finding), §2.3 (*"the defect … is that these apps are on the system drive in the first place"*), §5 (posed a question the architecture already answers), and **R-352 point (3)** (compared it to Tier 2's same-disk refusal). The prompt that commissioned this correction said *three*; the evidenced count is **four**. **Not a personal failure — a missing step:** nothing in `PROMPT-TEMPLATE.md` §4 required naming the architecture document for the area being touched. §4 item 4 named only `02-controller-module-map.md`, a package classification, and the S-1 rule at the end governs UPDATING a design doc, not READING one first. Fixed in the same session — see the template's new §4 rule. | **CLOSED — corrected 2026-08-22** | R-352 (re-framed), R-369 | The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | CC |
|
||||
| **R-371** | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
||||
| **R-372** | **A Tier-2 copy that has NEVER been produced because its source path is missing is not surfaced prominently to the operator.** Written down **2026-07-15**, in `audits/CAMPAIGN-6E-2026-07-15.md:128` (F-6E-1), whose disposition ends *"Optional product idea: surface 'tier-2 has never produced a copy (source missing)' more prominently in the operator UI — not filed."* The finding it sits on was correctly judged demo-data churn rather than a product defect (the code warns loudly and does not silently succeed), but the surfacing idea was never carried anywhere. **Age when filed: 38 days — the oldest gap this sweep recovered.** | **OPEN — LOW** | — | Decide whether "never produced a copy" deserves its own operator surface, distinct from "last copy failed". | CC |
|
||||
| **R-373** | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
|
||||
| **R-374** | **Three C1 refusal cases were judged borderline, left unfiled, and never named — so nobody can re-open the judgement.** `audits/CAMPAIGN-12-class-sweep-2026-08-08.md:123`: *"'Names a route' is a judgement, not a predicate — two readers could disagree on the borderline cases, and three of the 19 were called borderline and left unfiled."* The disclosure is honest and is exactly the right thing to write; **what is missing is WHICH three.** An unnamed borderline case cannot be re-judged by a second reader, which is the only remedy a judgement call has. **Age when filed: 14 days.** | **OPEN — LOW** | — | Name the three in that document, or file them as one row listing them. No code. | CC |
|
||||
| **R-375** | **A PBS datastore signal was noted and explicitly not filed.** `audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:171`: *"Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here."* Recorded so the note has a number and stops depending on someone re-reading that report. **Age when filed: 4 days.** | **OPEN — LOW** | — | Confirm the cause on ep0 the next time it is touched; it is a read-only check. | CC |
|
||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
|
||||
@@ -1,5 +1,38 @@
|
||||
# SPEC — where an app's data is placed, and who decides
|
||||
|
||||
> ## ⚠ CORRECTED 2026-08-22 — THE MEASUREMENTS STAND, THE FRAMING WAS WRONG
|
||||
>
|
||||
> **This document called a documented architectural decision a defect.** It is left in place rather
|
||||
> than rewritten, because a document that quietly changes its mind teaches nobody. Every measurement
|
||||
> below is good and was re-checked on 2026-08-22; the corrections are marked inline as
|
||||
> **[CORRECTED 2026-08-22]**.
|
||||
>
|
||||
> **What it got wrong.** It treated the absence of a storage field on 40 of 53 templates as a choice
|
||||
> being denied. There is no choice to deny. The architecture states the placement rule and states it
|
||||
> as *enforced*:
|
||||
>
|
||||
> > *"**App data placement is per-volume, not per-app:** `.felhom.yml` classifies each volume **hot**
|
||||
> > (DB/config/cache → fast storage, **enforced**) vs **bulk** (media/files → may be slow). A photo
|
||||
> > app's DB stays on SSD while its blobs go to the USB."*
|
||||
> > — `documentation/architecture/01-topology-and-trust.md:150-152`
|
||||
>
|
||||
> The 40 templates are **all-hot apps**: their data *is* app internals — database, config, cache —
|
||||
> and hot data belongs on fast storage inside the guest, by design. The 13 with a configurable path
|
||||
> are the media and library apps, where **bulk** content legitimately belongs on an attached drive.
|
||||
> The split is not 40 apps missing a feature; it is the hot/bulk contract working.
|
||||
>
|
||||
> **And the guest volume is deliberately its own volume**, so that filling it cannot stop the
|
||||
> operating system: the golden ships a small OS rootfs plus a single data volume at `/var/lib/felhom`
|
||||
> (`felhom-agent/configs/build-golden.sh:29-40`), and a capture that would exhaust it is refused per
|
||||
> app rather than allowed to stop the container runtime
|
||||
> (`documentation/architecture/00-capability-map.md:94`, PROVEN-LIVE 2026-08-03).
|
||||
>
|
||||
> **One real defect survives the correction and is unchanged:** `settings.go:453` says
|
||||
> `// new apps use this by default` and no app has ever used it. That is §2.2 below and it is now
|
||||
> filed on its own, scoped by measurement rather than by assumption — see the register.
|
||||
>
|
||||
> **What still needs a ruling: far less than §5 implies.** See the corrected §5.
|
||||
|
||||
**Filed 2026-08-21. Status: SPECIFICATION ONLY — nothing here is implemented, and nothing here may be
|
||||
implemented without the operator's ruling. It deserves its own session.**
|
||||
|
||||
@@ -20,6 +53,15 @@ The first hypothesis — that a customer had typed a bad path, or had failed to
|
||||
|
||||
> **The configured default data store is not consulted on the deploy route at all.**
|
||||
|
||||
> **[CORRECTED 2026-08-22]** The observation is right; the conclusion drawn from it was not. There is
|
||||
> nothing to choose **because these apps have no bulk volume to place** — the hot/bulk contract
|
||||
> (`01-topology-and-trust.md:150-152`) puts their data on fast storage inside the guest by design, and
|
||||
> OpenGist is a pure hot app. The sentence in the quote above is still true as a statement about the
|
||||
> deploy route, and it is still a defect **for the 13 templates that DO take a path** — that is the
|
||||
> real remainder, established in Part 4 of the 2026-08-22 session and filed separately. For the other
|
||||
> 40 it is not a defect at all: a default drive cannot be "consulted" for an app that has no path to
|
||||
> put it in.
|
||||
|
||||
## 2. The four measured facts
|
||||
|
||||
### 2.1 The affected class is 40 of 53 catalogue templates
|
||||
@@ -77,6 +119,29 @@ Live:
|
||||
So for these 40 apps **the data and its nearest copy sit on the same physical device**, reached by a
|
||||
customer doing nothing wrong.
|
||||
|
||||
> **[CORRECTED 2026-08-22] — and be precise about which risk is real.**
|
||||
>
|
||||
> **The device claim is true. The framing "by accident" was not.** Since R-165 (golden `build-golden.sh`
|
||||
> v3.0.0, proven live 2026-08-03) the guest carries a small OS rootfs plus **ONE** data volume at
|
||||
> `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume**, and
|
||||
> the script says so in terms — *"There is deliberately no mp1"* (`build-golden.sh:99`). The
|
||||
> `mp0`/`mp1` split this document's era assumed was retired six days before this document was written
|
||||
> and is on no current box.
|
||||
>
|
||||
> **Which risk is real, stated exactly:**
|
||||
> - **A failure of the physical disk loses both** the app's data and its first-tier copy. **REAL, and
|
||||
> it is what the off-site and whole-machine tiers exist for.** It is not specific to the 40-class:
|
||||
> a drive-resident app's unit lives beside its data on the drive too, deliberately, so a restore
|
||||
> needs the drive and nothing else.
|
||||
> - **A full data volume taking out the operating system. NOT REAL, and it was the overstated one.**
|
||||
> The OS rootfs is a separate volume, and the capture floor refuses a capture that would exhaust the
|
||||
> data volume rather than letting it stop the container runtime (`00-capability-map.md:94`). Watched
|
||||
> working on 2026-08-21: with `/var/lib/felhom` at 99% the reserve refused one app per run and told
|
||||
> the hub, while all 15 containers stayed healthy.
|
||||
>
|
||||
> So the sentence *"The defect is not Tier 1's rule; it is that these apps are on the system drive in
|
||||
> the first place"* below is **withdrawn**. Being on the guest's data volume is where hot data belongs.
|
||||
|
||||
The project already treats that posture as unacceptable — for the *other* tier. Tier 2 refuses it
|
||||
outright at `internal/backup/tier2.go:329`, recording
|
||||
`a kiválasztott cél ugyanazon a fizikai lemezen van`. Tier 1 has no such notion, and for a
|
||||
@@ -129,7 +194,28 @@ selected drive for the 13. Visibility only. **No placement changed. Nothing was
|
||||
|
||||
## 5. The open question this document exists to hand over
|
||||
|
||||
> **[CORRECTED 2026-08-22] — THIS QUESTION IS ALREADY ANSWERED, AND THE ANSWER IS NO.**
|
||||
>
|
||||
> The architecture answers it: hot data (DB/config/cache) belongs on fast storage inside the guest and
|
||||
> that placement is **enforced**, not preferred (`01-topology-and-trust.md:150-152`). A named-volume
|
||||
> app is all-hot. Moving it to the attached drive would be moving hot data onto storage the same
|
||||
> document classes as possibly-slow, and would put it outside the guest vzdump that currently protects
|
||||
> it (`01-topology-and-trust.md:154-156`).
|
||||
>
|
||||
> **So points 1, 2 and 3 below are withdrawn as a live question.** They remain a good record of what a
|
||||
> change would have cost, which is why they are not deleted. Point 3 was already arguing against the
|
||||
> premise of the question it sat under.
|
||||
>
|
||||
> **WHAT ACTUALLY STILL NEEDS YOUR RULING FROM THIS DOCUMENT: point 4 only, and it is smaller than it
|
||||
> looks** — `IsDefault`'s comment must become true or go away, for the 13 apps it could apply to. That
|
||||
> is filed as its own row with its scope measured rather than assumed. Point 5 (the Drives page count)
|
||||
> needs **no ruling**: the count is honest, it is only easy to misread, and that is a wording question
|
||||
> the design system already owns.
|
||||
>
|
||||
> **Everything else in this section: nothing to decide.**
|
||||
|
||||
**Should a named-volume app's data live on the default data drive rather than the system drive?**
|
||||
*(superseded — see the correction directly above)*
|
||||
|
||||
Points the next session must settle, each of which is a reason this was not decided tonight:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user