From 4b2e5608c2270d28c0fb532ee230cee1ded958c1 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 13 Sep 2026 09:05:52 +0200 Subject: [PATCH] R-442 CLOSED (controller v0.236.0): "delete my data too" deletes it or refuses; R-465..467 opened MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - OPEN-ITEMS: R-442 row removed; R-465 (six remaining Paths.HDDPath readers — audit), R-466 (recovery-unit residue after "delete backups"), R-467 (v0.236.0 owes a golden). - CLOSED-ITEMS: R-442 compressed, reasoning kept, fleet shape now established. - 00-capability-map: lifecycle row narrowed (Campaign 3 proved remove removes the APP, not the data) and re-proven from audits/R442-2026-09-13/. - STATUS: item 13 in plain language; item 7 closed (ruled 2026-09-02, 09 §3). - audits/R442-2026-09-13/: live evidence (A, C, D bodies, controls, log window, teardown). Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- STATUS.md | 10 +++- .../architecture/00-capability-map.md | 2 +- .../R442-2026-09-13/A-nextcloud-remove.txt | 37 ++++++++++++ .../R442-2026-09-13/C-nextcloud-refused.txt | 26 +++++++++ .../R442-2026-09-13/D-gokapi-remove.txt | 16 ++++++ .../audits/R442-2026-09-13/controls.txt | 4 ++ .../R442-2026-09-13/teardown-and-log.txt | 56 +++++++++++++++++++ documentation/backlog/CLOSED-ITEMS.md | 1 + documentation/backlog/OPEN-ITEMS.md | 4 +- scripts/closed_register_gate.py | 9 ++- 10 files changed, 159 insertions(+), 6 deletions(-) create mode 100644 documentation/audits/R442-2026-09-13/A-nextcloud-remove.txt create mode 100644 documentation/audits/R442-2026-09-13/C-nextcloud-refused.txt create mode 100644 documentation/audits/R442-2026-09-13/D-gokapi-remove.txt create mode 100644 documentation/audits/R442-2026-09-13/controls.txt create mode 100644 documentation/audits/R442-2026-09-13/teardown-and-log.txt diff --git a/STATUS.md b/STATUS.md index 3efa9360..180f4c47 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,5 +1,7 @@ # STATUS — what works, what's broken, what's next +**Updated 2026-09-13 — "delete my data too" now deletes the data, or tells you it could not. Until today the box said it worked and left everything on the drive. Live on both machines (0.236.0), proven on the HP with a throwaway Nextcloud. Nothing needs you. Item 7 is closed: you decided it on 2026-09-02 and the page still listed it open.** + **Updated 2026-09-06 (third pass) — I chased down the BookStack database problem I found this morning. Good news: it does not get worse, and fixing it costs seven seconds and loses us nothing. The catch I expected — that fixing it would stop us being able to go back — turned out not to be @@ -41,7 +43,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Two things are waiting on you: item 11 (how wide to take the upgrade testing) and item 12 (a small yes/no about the database setting — I have measured both costs).** Item 4 (the Hetzner e-mails) is answered and is being handled in a separate session. Item 7 — the safety-copy decision — is the one open question, and it is **not urgent any more**: the thing that made it urgent was that a restart could upgrade an app behind your back, and as of today it cannot. Item 10 is new and needs nothing from you. Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on +1. **Two things are waiting on you: item 11 (how wide to take the upgrade testing) and item 12 (a small yes/no about the database setting — I have measured both costs).** Item 4 (the Hetzner e-mails) is answered and is being handled in a separate session. Item 7 — the safety-copy decision — was ruled on 2026-09-02 and is now marked closed below (it was **not urgent any more**: the thing that made it urgent was that a restart could upgrade an app behind your back, and as of today it cannot. Item 10 is new and needs nothing from you. Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -93,7 +95,9 @@ nothing.* Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as it is, at the risk you accept by leaving it. I can change it without ever showing you the new one. -7. **Where should the safety go before an app updates? This is the one decision from today's +7. **CLOSED 2026-09-13 — you decided this on 2026-09-02, and the page kept listing it open.** The ruling is recorded in `documentation/architecture/09-update-architecture.md` §3: the safety copy is a *verified recent backup as a precondition*, not a new copy made for the update; the guest-snapshot idea is to be spiked before anything is built on it. The text below is kept as the record of what was asked. + + *Original question:* **Where should the safety go before an app updates? This is the one decision from today's measurement, and it is a design choice, not a bug report.** **What I measured.** The box downloads new app versions by itself every 15 minutes and writes them @@ -269,6 +273,8 @@ nothing.* actually moved a major version so far; the other three will land in the same place when their turn comes. I have not changed any template — this is your call, not a quiet edit. +13. **"Delete my data too" now deletes the data — or tells you it could not.** Until today, when a customer removed an app and ticked the box, the box said it worked and left everything on the drive (128 MB of a Nextcloud on 2026-09-01). The cause: the removal asked one global setting for the drive, and no machine fills that setting in. Every other part of the box already asks the app itself where its data is. Now the removal does too. If the box cannot work out where the data is, it refuses and keeps the app, so you can try again — it never again reports success over data left behind. Proven on the HP with a throwaway Nextcloud: 63 MB the app wrote itself was gone after removal, and the answer listed it; the refusal was shown with the app still in place; an app with no drive data gets a plain "nothing to delete" note. Live on both machines (0.236.0). No standing app was touched. **If you do nothing:** nothing to do. + 8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, as it misled one by an hour. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 6c888811..6d20edff 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -99,7 +99,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Scenario | Components | Status | Evidence | Gap / roadmap | |---|---|---|---|---| | Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only | -| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. What they do to app DATA is now measured too, and it is a separate row-worth of facts (below).** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3`; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | +| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | | **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | diff --git a/documentation/audits/R442-2026-09-13/A-nextcloud-remove.txt b/documentation/audits/R442-2026-09-13/A-nextcloud-remove.txt new file mode 100644 index 00000000..11212fca --- /dev/null +++ b/documentation/audits/R442-2026-09-13/A-nextcloud-remove.txt @@ -0,0 +1,37 @@ +=== 2026-09-13T06:59:26Z SCENARIO A: nextcloud remove_hdd_data=true remove_backups=true === +state=stopped deployed=True deploying=False err=None +--- GET hdd-data --- +{"ok":true,"data":{"stack":"nextcloud","hdd_paths":[{"path":"/mnt/felhom-drives/hdd_1/appdata/nextcloud","size_bytes":65241015,"size_human":"63M","exists":true}],"has_hdd_data":true}} + +HTTP 200 +--- GET backup-data --- +{"ok":true,"data":{"stack":"nextcloud","backup_paths":[{"path":"/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps","size_bytes":4117,"size_human":"8.0K","exists":true},{"path":"/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud/rsync","size_bytes":0,"size_human":"","exists":false}],"has_backups":true}} + +HTTP 200 +--- BEFORE: find --- +69 +/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps/r442-fixture.sql +/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/compose/.felhom.yml +/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/compose/app.yaml +/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/manifest.json +--- POST remove --- +{"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":["/mnt/felhom-drives/hdd_1/appdata/nextcloud (63M)"],"hdd_paths_preserved":[],"backup_paths_removed":["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps (8.0K)"]},"message":"Stack nextcloud removed"} + +HTTP 200 +--- AFTER: find --- +appdata files: 1 +ls: cannot access '/mnt/felhom-drives/hdd_1/appdata/nextcloud': No such file or directory +ls: cannot access '/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps': No such file or directory +paperless +romm +state=not_deployed deployed=False deploying=False err=None +--- controller log --- +2026/09/13 06:59:17 [INFO] [stacks] ParseComposeHDDMounts: found 2 HDD mounts for /opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/13 06:59:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml +2026/09/13 06:59:17 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/romm/docker-compose.yml +2026/09/13 06:59:27 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/13 06:59:27 delete.go:461: [INFO] Removing deployed stack: nextcloud (removeHDDData=true, hddDeclared=true, backupPaths=1) +2026/09/13 06:59:27 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/13 06:59:28 delete.go:523: [INFO] Removed HDD data: /mnt/felhom-drives/hdd_1/appdata/nextcloud (63M) +2026/09/13 06:59:28 delete.go:569: [INFO] Removed backup data: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps (8.0K) diff --git a/documentation/audits/R442-2026-09-13/C-nextcloud-refused.txt b/documentation/audits/R442-2026-09-13/C-nextcloud-refused.txt new file mode 100644 index 00000000..655b1360 --- /dev/null +++ b/documentation/audits/R442-2026-09-13/C-nextcloud-refused.txt @@ -0,0 +1,26 @@ +=== 2026-09-13T07:00:17Z SCENARIO C: nextcloud with HDD_PATH removed from app.yaml (TEST-ONLY EDIT), remove_hdd_data=true === +state=running deployed=True deploying=False err=None +state=stopped deployed=True deploying=False err=None +HDD_PATH lines in app.yaml now: 1 +data files before: 69 +--- POST remove --- +{"ok":false,"error":"Az alkalmazás adatainak helye nem állapítható meg, ezért semmit nem töröltünk. Az alkalmazás nem lett eltávolítva."} + +HTTP 409 +--- after the refusal --- +state=stopped deployed=True deploying=False err=None +-rw------- 1 root root 1302 Sep 13 07:00 /opt/docker/stacks/nextcloud/app.yaml +containers still present: +data files after: 69 +--- controller log --- +2026/09/13 07:00:20 delete.go:130: [ERROR] [stacks] RemoveStack nextcloud refused: data removal requested, the compose binds a drive path, but app.yaml records no HDD_PATH — nothing removed, app kept (R-442) +2026/09/13 07:00:20 router.go:830: [ERROR] [api] Remove failed for nextcloud: Az alkalmazás adatainak helye nem állapítható meg, ezért semmit nem töröltünk. Az alkalmazás nem lett eltávolítva. +--- restore app.yaml, then the SAME call (positive control) --- +HDD_PATH lines restored: 2 +{"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":["/mnt/felhom-drives/hdd_1/appdata/nextcloud (63M)"],"hdd_paths_preserved":[]},"message":"Stack nextcloud removed"} + +HTTP 200 +ls: cannot access '/mnt/felhom-drives/hdd_1/appdata/nextcloud': No such file or directory +paperless +romm +state=not_deployed deployed=False deploying=False err=None diff --git a/documentation/audits/R442-2026-09-13/D-gokapi-remove.txt b/documentation/audits/R442-2026-09-13/D-gokapi-remove.txt new file mode 100644 index 00000000..e34750dc --- /dev/null +++ b/documentation/audits/R442-2026-09-13/D-gokapi-remove.txt @@ -0,0 +1,16 @@ +=== 2026-09-13T06:59:49Z SCENARIO D: gokapi (SSD, no HDD_PATH) remove_hdd_data=true === +0 +app.yaml has NO HDD_PATH line +0 +compose binds NO drive path +state=stopped deployed=True deploying=False err=None +--- GET hdd-data --- +{"ok":true,"data":{"stack":"gokapi","hdd_paths":null,"has_hdd_data":false}} + +HTTP 200 +--- POST remove --- +{"ok":true,"data":{"removed":"gokapi","volumes_removed":null,"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni."},"message":"Stack gokapi removed"} + +HTTP 200 +state=not_deployed deployed=False deploying=False err=None +--- controls on the D body: llap=0 removed_empty= hdd_note=1 diff --git a/documentation/audits/R442-2026-09-13/controls.txt b/documentation/audits/R442-2026-09-13/controls.txt new file mode 100644 index 00000000..e47f25b3 --- /dev/null +++ b/documentation/audits/R442-2026-09-13/controls.txt @@ -0,0 +1,4 @@ +=== ASCII-fragment controls (searched on DooPlex, saved bodies) === +A-nextcloud-remove.txt: llap=0 'nem lett elt'=0 hdd_paths_removed=1 HTTP409=0 HTTP200=3 +C-nextcloud-refused.txt: llap=2 'nem lett elt'=2 hdd_paths_removed=1 HTTP409=1 HTTP200=1 +D-gokapi-remove.txt: llap=1 'nem lett elt'=0 hdd_paths_removed=1 HTTP409=0 HTTP200=2 diff --git a/documentation/audits/R442-2026-09-13/teardown-and-log.txt b/documentation/audits/R442-2026-09-13/teardown-and-log.txt new file mode 100644 index 00000000..ed115e6f --- /dev/null +++ b/documentation/audits/R442-2026-09-13/teardown-and-log.txt @@ -0,0 +1,56 @@ +=== controller log, R-442 window === +2026/09/13 06:54:17 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/13 06:54:17 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/13 06:56:17 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/13 06:56:17 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/13 06:56:57 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/nextcloud/deploy +2026/09/13 06:56:57 router.go:81: [DEBUG] [api] POST /api/stacks/nextcloud/deploy (path=/stacks/nextcloud/deploy) +2026/09/13 06:56:57 router.go:403: [INFO] [api] Deploy requested for stack: nextcloud +2026/09/13 06:56:57 router.go:81: [DEBUG] [api] deployStack: name=nextcloud contentLength=263 +2026/09/13 06:56:58 deploy.go:288: [DEBUG] Deploy nextcloud: received 7 user values +2026/09/13 06:56:58 deploy.go:418: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH, DOMAIN, SUBDOMAIN, DB_PASSWORD, MYSQL_ROOT_PASSWORD] +2026/09/13 06:56:58 manager.go:1515: [INFO] [stacks] Deploying stack nextcloud — checking 3 images... +2026/09/13 06:56:58 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/gokapi/deploy +2026/09/13 06:56:58 router.go:81: [DEBUG] [api] POST /api/stacks/gokapi/deploy (path=/stacks/gokapi/deploy) +2026/09/13 06:56:58 router.go:403: [INFO] [api] Deploy requested for stack: gokapi +2026/09/13 06:56:58 router.go:81: [DEBUG] [api] deployStack: name=gokapi contentLength=103 +2026/09/13 06:56:58 deploy.go:288: [DEBUG] Deploy gokapi: received 3 user values +2026/09/13 06:56:58 deploy.go:418: [INFO] [stacks] Deploying stack gokapi with 3 env vars: [DOMAIN, SUBDOMAIN, GOKAPI_PASSWORD] +2026/09/13 06:56:58 manager.go:1515: [INFO] [stacks] Deploying stack gokapi — checking 1 images... +2026/09/13 06:57:14 deploy.go:466: [INFO] [stacks] Stack gokapi deployed successfully (took 16.2s) +2026/09/13 06:57:56 deploy.go:466: [INFO] [stacks] Stack nextcloud deployed successfully (took 58.7s) +2026/09/13 06:58:17 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=true composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/13 06:58:17 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=true composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/13 06:59:10 manager.go:1173: [DEBUG] [stacks] StopStack nextcloud: current state=running deployed=true containers=3 +2026/09/13 06:59:27 delete.go:432: [DEBUG] [stacks] RemoveStack nextcloud: state=stopped, deployed=true, orphaned=false, deploying=false +2026/09/13 06:59:27 delete.go:461: [INFO] Removing deployed stack: nextcloud (removeHDDData=true, hddDeclared=true, backupPaths=1) +2026/09/13 06:59:28 delete.go:523: [INFO] Removed HDD data: /mnt/felhom-drives/hdd_1/appdata/nextcloud (63M) +2026/09/13 06:59:28 delete.go:569: [INFO] Removed backup data: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps (8.0K) +2026/09/13 06:59:28 delete.go:584: [INFO] Stack nextcloud removed successfully (took 1.0s) +2026/09/13 06:59:28 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=true composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/13 06:59:28 manager.go:527: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/13 06:59:39 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/nextcloud/deploy +2026/09/13 06:59:39 router.go:81: [DEBUG] [api] POST /api/stacks/nextcloud/deploy (path=/stacks/nextcloud/deploy) +2026/09/13 06:59:39 router.go:403: [INFO] [api] Deploy requested for stack: nextcloud +2026/09/13 06:59:39 router.go:81: [DEBUG] [api] deployStack: name=nextcloud contentLength=263 +2026/09/13 06:59:39 deploy.go:288: [DEBUG] Deploy nextcloud: received 7 user values +2026/09/13 06:59:39 deploy.go:418: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH, DOMAIN, SUBDOMAIN, DB_PASSWORD, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER] +2026/09/13 06:59:39 manager.go:1515: [INFO] [stacks] Deploying stack nextcloud — checking 3 images... +2026/09/13 06:59:49 manager.go:1173: [DEBUG] [stacks] StopStack gokapi: current state=running deployed=true containers=1 +2026/09/13 06:59:50 delete.go:432: [DEBUG] [stacks] RemoveStack gokapi: state=stopped, deployed=true, orphaned=false, deploying=false +2026/09/13 06:59:50 delete.go:461: [INFO] Removing deployed stack: gokapi (removeHDDData=true, hddDeclared=false, backupPaths=0) +=== standing apps (deployed) === +bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin romm +=== residue check === +paperless +romm +calibre-web +nextcloud +paperless-ngx +romm +=== teardown: hand-removing the throwaway apps recovery-unit residue (compose/ + manifest.json ��� NOT customer data) === +calibre-web +paperless-ngx +romm +0 +0 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 0b455638..1557e886 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,7 @@ --- +| **R-442** | **`remove_hdd_data: true` was INERT — the customer's data stayed on the drive while the API reported success (HTTP 200, `hdd_paths_removed: null`, 128 MB of Nextcloud left; demo-hp 2026-09-01).** Shipped in controller **v0.236.0** (2026-09-13). Removal resolved the drive from the GLOBAL `cfg.Paths.HDDPath` (no default, set on NO box — demo-hp AND demo-felhom both measured 0 `hdd_path` / 0 `FELHOM_PATHS_*`, so the fleet shares the shape); it now reads the app's OWN `app.yaml` `HDD_PATH` — the `07-backup-architecture.md` ~L437 rule that deploy, the start gate and the backup destination already implemented — and a data removal it cannot resolve, or whose drive is absent, is REFUSED (409, exact Hungarian sentence, typed `stacks.RemoveRefusedError`) BEFORE `compose down`, app kept. SSD app → `hdd_paths_removed: []`, never `null`, plus `hdd_note`. Missing folders stated in `hdd_paths_missing`. The backup-half refusal reaches the response (`backup_paths_refused`) and its base follows the same rule. **Reasoning kept:** *"declares no drive" ≠ "could not resolve the drive" — the first is a fact, the second a refusal*; *an app gone with its data left behind is unrecoverable from the UI — the customer cannot even re-run the removal*; *no fallback to the global — that silent fallback is the exact path this closes*. **Observation carried:** 8 of the 13 `needs_hdd` catalog apps bind ONLY `${USERDATA_PATH}` (the shared library) and no `${HDD_PATH}` folder, so for them "delete my data" correctly removes nothing on the drive and the modal shows no checkbox. | **CLOSED 2026-09-13 — shipped controller v0.236.0, proven live on demo-hp** | `audits/R442-2026-09-13/` — A: 63 MB written by Nextcloud ITSELF, gone after removal and listed with its size; C: 409 + sentence, all 69 files untouched, app still deployed, `[ERROR] … refused` logged; D: gokapi `[]` + note; ASCII controls (`llap`: C=2 A=0 D=0). 15 tests + two red-proofs in `felhom-controller/REPORT.md`. Full original text: `git show d6837d98ee24:documentation/backlog/OPEN-ITEMS.md`. | | **R-449** | **UPDATE ARC SLICE 5 — an upgrade test that runs again. BUILT AND RUN 2026-09-06:** `app-catalog-felhom.eu/scripts/upgrade-test.py` + `upgrade_fixtures.py`, 7 edges across 3 apps, evidence in `audits/upgrade-spike-2026-09-06/`. **Reasoning kept — success is an APPLICATION-LEVEL READBACK, never file identity:** `survive2.py`'s sha256+inode rule is right for a redeploy and WRONG for an upgrade, because a migration is supposed to rewrite files and that rule would fail every correct upgrade. **Reasoning kept — nothing is ever seeded into a volume by hand (R-156);** an app with no non-browser route is recorded `inconclusive`, which is a result and not a licence to plant a file. **Reasoning kept — run the negative control FIRST:** C3's TO image exits immediately and came back `failed`; a harness that cannot fail a known-broken upgrade proves nothing with its greens. **WHAT IT MEASURED:** all five real catalog upgrades kept the customer's data; and **whether an upgrade can be undone is a property of the individual APP, not of upgrades** — docmost REFUSES (*"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*), privatebin does not, which independently reproduces the Nextcloud finding on a second app by a DIFFERENT mechanism and puts two measurements behind §4's ruling that "rollback" is the wrong word. **It also found a defect in our own catalog (R-459).** **What stays open, as its own rows rather than inside this one:** R-459 (the skipped MariaDB datadir upgrade), R-460 (bookstack's file half is unprovable headlessly), R-462 (the widening, costed). Full original text: `git show 417df06f3529:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — harness built, run, and proven by a red negative control** | `audits/SPIKE-upgrade-test-2026-09-06.md`; `audits/upgrade-spike-2026-09-06/evidence/`; catalog `0474ce387e6f` | | **R-438** | **The catalog sync rewrote a DEPLOYED app's `docker-compose.yml` and no architecture document recorded that it did. BOTH HALVES NOW DISCHARGED — the document was written 2026-09-02, the behaviour was changed in controller v0.235.0 (2026-09-06).** `Syncer.copyTemplates` copied into every stack folder on a 15-minute cycle with **no deployed check**, so a deployed app's file and its running containers disagreed from that moment, and the next `compose up -d` from any of thirteen call sites resolved the disagreement by upgrading — measured live: the sync rewrote the file at 17:45:17Z while the container went on running the old image, a restart then upgraded it in 18.3 s **with a pull**, and a boot reconciliation upgraded it **with nobody pressing anything**. **Reasoning kept — the distinction this row existed to protect:** `RestartStack`'s use of `up -d` to pick up template changes was **CHOSEN and written down in its own comment**, so reversing it was an operator DECISION, not a bug fix; that is why the row stayed open through v0.233.0 and v0.234.0 while only the documentation half was done. **Reasoning kept — one fear was measured SMALLER than stated:** a plain power cut does NOT upgrade anything, because Docker's `restart: unless-stopped` restores the containers on the old image and the reconciler logs `no boot-orphaned apps`; the unattended upgrade needs the narrower precondition *"and the app did not come back"*. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — documented 2026-09-02, behaviour changed in controller v0.235.0** | `audits/SPIKE-app-update-2026-09-01.md` §2, §3, §8; `architecture/09-update-architecture.md`; `tests/VALIDATION-update-slice3-2026-09-06.md` | | **R-447** | **UPDATE ARC SLICE 3 — the live compose file is now DERIVED; the syncer renders instead of copying. SHIPPED controller v0.235.0 (2026-09-06).** The pin lives in `app.yaml` (`pinned_images`), the definition it came from is stored beside the app as `applied-compose.yml`, and `Syncer.renderSource` writes the catalog template verbatim while the catalog still offers the pinned version and the stored definition once it moves past it. **Reasoning kept — the ruling, in the operator's own words:** *while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update.* **Reasoning kept — why nothing was added to the thirteen `compose up -d` call sites:** most of them are REPAIRS (the boot reconciler, the drive-return gate, the app-stop guard), and **a repair path that refuses to repair leaves a customer's app down, which is worse than the problem**; they were made safe by removing the reason, not by gating them. **Reasoning kept — why the frozen branch writes a WHOLE file and never a substitution:** `wger 2.6` needs a full DB configuration the older template cannot supply, so an old image under a new template is a third state nobody chose. **Reasoning kept — why this is not "skip deployed apps" (option B, rejected):** that also stops health-check fixes, memory limits and new deploy fields, and destroys the self-healing measured live in the spike §3. **Reasoning kept — `pinned_images` is INTENT and `installed_images` is an OBSERVATION; never feed one from the other** (the R-166 category error, one field over). Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — SHIPPED controller v0.235.0** | `architecture/09-update-architecture.md` §3.4, §5; `tests/VALIDATION-update-slice3-2026-09-06.md`; controller `CHANGELOG.md` v0.235.0 | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index a0f368bc..c05d3fc3 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -677,7 +677,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** | | **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | -| **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** | | **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | @@ -697,6 +696,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 | **READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements** | | **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 | **READY — rank P2-MEDIUM; owner: CC** | | **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** | +| **R-465** | **[P3-LOW] `cfg.Paths.HDDPath` — the global that R-442 proved is set on NO box (demo-hp AND demo-felhom: 0 `hdd_path`, 0 `FELHOM_PATHS_*`) — still has SIX readers, each reading an always-empty value:** `report/builder.go:69`, `monitor/healthcheck.go:35`, `api/router.go` (system-info), `web/server.go:740`, `cmd/controller/main.go` (auto-discovery seed + metrics HDD path). Removal was silently inert for months on the very same read. Whether any of these is inert the same way — a report field that is always empty, a health check that never fires, a metric never collected — is a one-hour audit: for each reader, name the POSITIVE observable that must appear when it works and check it on the box. R-442 §5 said "do not delete it here"; this row is the audit it deferred. Owner: **CC.** `felhom-controller/REPORT.md` (v0.236.0, Observations 1) | **READY — rank P3-LOW; owner: CC** | +| **R-466** | **[P3-LOW] Removing an app with „Mentési adatok törlése" ticked leaves its recovery unit's `compose/` + `manifest.json` on the drive.** The router passes only `backup.AppDBDumpPath(nsRoot, name)` to `RemoveStack`, so `backups/primary//db-dumps/` goes and the unit root keeps `compose/` (the app's `app.yaml` with the portable secret class at 0600) and `manifest.json`. MEASURED 2026-09-13 on demo-hp after the R-442 Scenario A removal — `audits/R442-2026-09-13/teardown-and-log.txt`, residue check: `backups/primary/nextcloud` still listed with the app gone; removed by hand at teardown. A customer who asked for the backups to go is left with the app's definition and a manifest. **Decide:** the button means the WHOLE unit (pass the unit root, under the same `backups/`-prefix guard) or stays db-dumps-only (then the modal must say so). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | +| **R-467** | **[P3-LOW] Controller v0.236.0 owes a golden** (golden-notice gate, 2026-09-13): newest golden baked is **0.232.0**, so a machine installed now receives 0.232.0 and reaches 0.236.0 only by self-update. Bake + vouch per `runbooks/RUNBOOK-manual-build.md` §4.1 (the THREE-field vouch: `golden_version` + `agent_version` + `min_agent`); bake record to `documentation/tests/golden--/`. If the train deliberately skips this version, this row is the waiver the gate asks for. Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |