diff --git a/CONTEXT.md b/CONTEXT.md index e97f7ac0..d9a71687 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -34,6 +34,12 @@ on demo-hp — `documentation/audits/slice4-2026-09-13/`, design `09-update-arch - **Removal keeps the backups AND the Tier-2 record unless the customer asked for the backups to go** (v0.240.0, R-486/R-474). A removal that forgets the record makes an intact mirror unrestorable — that was live until v0.239.0. The first nightly rotation (adventurelog, 2026-09-13) found it. +- **A bind-data app's route back is off-site before its own unit, and the hold names what the copy + holds** (operator ruling 2026-09-13, R-479, controller v0.241.0). The unit of such an app holds the + definition and the database dumps, not the files. Volume-data apps keep the v0.239.0 order. +- **One host, one customer** (schema: `hosts.host_id` PK, one `customer_id`; one agent per box). A + scratch guest "enrolled as its own customer" on a box that already belongs to a customer is not + something the product can do (R-481, 2026-09-13); a second guest of the same customer is. - **A release reaches the fleet by floor between golden bakes when its MinAgent is declared with the floor** (ruling 2026-09-13, hub v0.112.0, R-472 CLOSED). An undeclared floor above the golden is still held, and the forms refuse it. v0.239.0 arrived on both demo boxes this way in ~15 s. Never hand-deploy diff --git a/REPORT.md b/REPORT.md index da93e2ee..210678e5 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,43 +1,45 @@ -# REPORT — the first "be a customer for the night" run: adventurelog on demo-hp (2026-09-13/14) +# REPORT — the four evening items before the second night (2026-09-13) -*Overwritten each session. Every finding below has a register row. The controller report is +*Overwritten each session. The night's own report (adventurelog rotation, v0.240.0) is preserved as +`audits/nightly-2026-09-13-adventurelog/` and the register; the controller report is `felhom-controller/REPORT.md`, the catalog's `app-catalog-felhom.eu/REPORT.md`.* -> One app walked end to end as a household would; seven product defects found; a controller release -> (v0.240.0) shipped the same night by the managed floor; two catalog fixes; two docs fixes. The -> morning note is the top block of `STATUS.md`. +## 1. The rules file -## Step 0 — the rules file +`.claude/rules/unprompted-work.md` copied byte-identical into `felhom-controller`, `felhom.eu` and +`app-catalog-felhom.eu` (documents-only pushes). `instructions_gate.py` refused the first push: a rule +file needs `paths:` or `unconditional: true`; all four copies (root included) now start with the +three-line `unconditional: true` frontmatter — sha256 `d1aa3af83ae3b220…` everywhere. -`.claude/rules/unprompted-work.md` is in none of the three repos and the brief's text block did not -reach this session (searched the workspace, the transcript and memory). Not created — the operator's -rules cannot be invented byte-identically. The brief's own fences were applied instead. +## 2. R-481 — the scratch guest (ruled; NOT built, and why) -## What changed here +The ruling ("a second enrolled LXC … enrolled as a scratch customer") collides with the model: the hub +keys a host to one customer (`hosts.host_id` PK) and the box runs one agent with one host identity; +re-running the installer with another customer-id would rewrite demo-hp's enrolment. Options and a +recommendation are in the row and STATUS; 9201's shape and the disk placement (`pct move-volume` +onto a `dir` storage at `/mnt/hdd_1`) are recorded for whichever is chosen. -- `documentation/runbooks/nightly-rotation.md` — new: the tick list (standing nine first, blocked on - R-481; then the catalog); `actualbudget` recorded (no non-browser front door), `adventurelog` ticked. -- `scripts/observations_gate.py` — R-471: every observations section is read (`scripts/CHANGELOG.md`). -- `documentation/runbooks/target-selection.md` — R-461: `/mnt/hdd_1` is the NVMe; `drill-r50` exists - on neither box (measured), fence kept with the fact beside it; R-93 carries the fact. -- `09-update-architecture.md` §6.1 + §8.2 (R-452 lifted), `07-backup-architecture.md` §6, `CONTEXT.md` — v0.240.0 notes. -- Catalog (`app-catalog-felhom.eu`): the `catalog-since` gate (R-452), see its REPORT. -- Register: R-465 (audited: five readers fall back, one is inert and unreachable → R-490), R-452, R-473, R-474, R-466, R-471, R-453, R-461, R-477, R-478, R-480, R-482, R-484, R-485, R-486 - closed and compressed; R-481, R-482..R-489 opened (R-482 closed the same night). 212 → 206 open, - 177 → 189 closed (byte sizes: see the commit). -- Evidence: `documentation/audits/nightly-2026-09-13-adventurelog/` (01–08), `audits/v0240-2026-09-13/` - (release log, floor, validation, red-proofs incl. R-471). +## 3. R-483 — photos (re-ranked P2, catalog) + +Diffed against `homelab-manifests` `adventurelog-system/adventurelog.yaml`: the k3s ingress sends +`/media`, `/static`, `/admin`, `/accounts` to the backend **service port 80**; the catalog sent +everything to the frontend. Two catalog cuts (`3172258`, `ed62cfd`): a backend traefik router for the +four prefixes, then port 80 instead of 8000 — the backend image runs nginx in front of gunicorn and +Django serves a photo by `X-Accel-Redirect`, so port 8000 returned 200 with an empty body. Applied to +the operator's own instance through the guarded Update; proven without a browser +(`audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt`): `GET /media/images/.webp` with a +session → 200 `image/webp`, RIFF/WEBP magic; anonymous 403 (the app's rule). **Confirmed by the operator in a browser at 21:49 — closed.** The frontend's 500 on scripted multipart uploads is upstream and +recorded in the row; browser uploads work. + +## 4. R-479 — tier order for bind-data apps (ruled; controller v0.241.0) + +Ruling 9 in `09-update-architecture.md` §3. Built, tested, red-proofed; delivered by the floor and +live-checked on a nextcloud throwaway — see `felhom-controller/REPORT.md` and +`audits/v0241-2026-09-13/`. ## Observations -1. **The brief's rules block was not delivered; the file is absent in all repos.** NOT-A-FINDING: an input gap for the operator, recorded in STATUS "needs you", not a product defect. -2. **No scratch guest on demo-hp; the standing nine cannot be throwaways on 9201.** FILED: R-481 -3. **adventurelog ran Django DEBUG=True on the public origin.** FILED: R-482 -4. **Photo upload fails 500 from every non-browser client (frontend proxy body-length bug).** FILED: R-483 -5. **PostGIS not recognised as a database.** FILED: R-484 -6. **The backup card reads dead paths.** FILED: R-485 -7. **Removal with backups kept forgot the Tier-2 record; the restore was refused.** FILED: R-486 -8. **A removed app is on neither backup page.** FILED: R-487 -9. **The backup test package takes 5½ minutes of real waits.** FILED: R-488 -10. **`volumes_removed: null` over removed volumes, five times.** FILED: R-489 -11. **The monitoring page's memory-distribution card never renders — `/api/system/info` is shadowed by `ServeSystemAPI` (found by the R-465 audit).** FILED: R-490 +0. **A removed app keeps its update hold; a reinstall would start held.** FILED: R-491 + +1. **The controller's evening release rides the same session as the night's — two releases in one calendar day, one per session.** NOT-A-FINDING: the rule is per session; both are recorded with their MinAgent lines. +2. **`homelab-manifests` is not in the workspace; its manifests were read from Gitea's API.** NOT-A-FINDING: the workspace CLAUDE.md says non-felhom repos are unrelated; a read-only fetch was enough. diff --git a/STATUS.md b/STATUS.md index d9b122ed..e79687ed 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,5 +1,34 @@ # STATUS — what works, what's broken, what's next +**Updated 2026-09-13 (evening, before the night) — your four items.** + +**Decisions I took.** (1) The rules file now sits in the workspace root and all three repos, +byte-identical; the gate that checks instruction files required three lines at its top saying it +loads in every session on purpose, so all four copies carry them. (2) The photo fix was applied to +your own test instance on the HP through the Update button, not by hand, and I added one test +account with one photo there. + +**What I did.** The rules file: done. The photo problem: the app's backend serves photos itself, +through a small web server inside its own image; our template sent every request to the front page +server, which has no photos. Fixed in two cuts — the second one because the first spoke to the wrong +port and got empty pictures. Proven without a browser: a photo now comes back as a real image +through the public address, and you confirmed it in your browser at 21:49. The tier-order +ruling: built as controller 0.241.0 — an app whose data lives in mounted folders now prefers the +remote copy over its own-drive copy, and the hold message ends with what the chosen copy holds. +Live on both machines (0.241.0, delivered by the floor in 16 and 18 seconds) and proven on the HP +with a throwaway Nextcloud. + +**What broke on the way.** Nothing standing. One new finding: after an app is removed, its +"held after a failed update" mark stays behind, so a reinstall would start blocked. Cleared by hand, +filed (R-491) for the next release. Register: 206 open / 191 closed. + +**Needs you.** The scratch guest: the ruling as written cannot be built — the hub ties one box to +one customer, and a second enrolled customer on the HP would replace the HP's own enrolment. Two +options: a second guest under the HP's own customer (reversible, tonight-ready, but its reports +would clash with the HP's page unless it stays unenrolled), or a nested appliance VM enrolled as its +own box (fully enrolled, hours to build). My pick: the first, unenrolled. **If you do nothing:** +tonight's rotation keeps restoring in place and skips the standing nine. + **Updated 2026-09-14 (the night of 13→14 — "be a customer for the night", first run). Written in the order the rules ask for.** diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 038aec37..72e8e6a2 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -277,7 +277,7 @@ what the unit holds** — for an app whose data is a bind mount that is the defi (R-479). **Removal and the tiers (controller v0.240.0):** removing an app with its backups KEPT keeps the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás" still works afterwards (R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and -never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open). +never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open). **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.** ### 6.1 The four tiers, as configured on the live fleet diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 4ee14ed8..bf70033a 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -178,6 +178,14 @@ These are rulings, not proposals. Anything specced against a different assumptio `audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone), 07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared). +9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy + holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit + that holds the definition and the database dumps and NOT the files (measured: gokapi restored from + „helyi" came back with settings and no data). For such an app the update walks second drive → + off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive → + own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a + customer is never sent to a copy that cannot bring the data back without being told so. + ## 4. The vocabulary ruling — "rollback" is struck **App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually diff --git a/documentation/audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt b/documentation/audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt new file mode 100644 index 00000000..d541a609 --- /dev/null +++ b/documentation/audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt @@ -0,0 +1,79 @@ +=== R-483 repro on demo-hp — 2026-09-13T19:35:02Z === +sync: HTTP 200 {"ok":true,"data":{"ok":true,"message":"Sablonok naprakészek — nincs változás"},"message":"Sablonok naprakészek — nincs változás"} +cache has the backend router: 1 +--- the operator's own throwaway instance (subdomain travel, 253 media files) is used; its containers get the new routers through the guarded Update (same version) --- +labels before: (none) +--- POST /api/stacks/adventurelog/update at 2026-09-13T19:35:05Z --- +HTTP 202 +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} + + 0s phase safety-dump + + 3s phase starting + + 6s phase verifying + + 21s phase done +end (2026-09-13T19:35:26Z, +21s): state=running updating=False phase=done err='' hold='' +labels after: Host(`travel.enkisfelhom.hu`) && (PathPrefix(`/media`) || PathPrefix(`/static`) +healthy: True +signup: 200 +login: (302, '/') +location: 201 +--- upload through the frontend proxy, WebKit-shaped boundary (what a browser sends) --- +proxy upload -> 500 {"id": null, "image": null} +--- proxy still refuses non-browser multipart; upload through the backend's own API inside the guest (the browser's session cookie, same backend) --- +backend upload -> 403 +{"detail":"CSRF Failed: CSRF cookie not set."} +image url: None +--- control: /static and /admin now answer from Django --- +/static/admin/css/base.css -> 200 +/admin/login/ -> 403 +=== done 2026-09-13T19:35:34Z === +=== R-483 repro, second half — 2026-09-13T19:36:33Z: an upload placed through the backend's own API (fresh CSRF cookie from the backend), then the photo fetched through the PUBLIC origin === +login: (302, '/') +have sessionid: True +my place id found: True +backend upload -> 403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."} +image url: None +=== done 2026-09-13T19:36:37Z === +=== R-483 repro, third try (the CSRF token is read with cut -f7, quoting-proof) — 2026-09-13T19:37:35Z === +backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."} +image url: None +=== done 2026-09-13T19:37:39Z === +=== R-483 repro, fourth try (CSRF cookie from /admin/login/, which always sets one) — 2026-09-13T19:38:02Z === +backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."} +image url: None +=== done 2026-09-13T19:38:05Z === +=== R-483 repro, fifth try (curl drops a Secure cookie over http; the token is cut from the raw Set-Cookie header) — 2026-09-13T19:38:30Z === +backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."} +image url: None +=== done 2026-09-13T19:38:33Z === +=== R-483 repro, sixth try (token cut from the config endpoint's Set-Cookie header) — 2026-09-13T19:39:01Z === +backend upload -> tokenlen=32 | http=400 | {"error":"content_type and object_id are required"} +image url: None +=== done 2026-09-13T19:39:05Z === +=== R-483 repro, seventh try (v0.12.1's image API takes content_type=location + object_id) — 2026-09-13T19:39:21Z === +backend upload -> http=201 | {"id":"c0b2b8cd-fe22-45db-b2a2-7c859d771941","image":"https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp","is_primary":false,"user":"51f9e8ee-f303-430e-b8ab-aae231e758c4","immich_id":null} +image url: https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp +GET /media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp through traefik (public origin) -> 200; bytes identical to the upload: False +the app lists the photo on the place: ['https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp'] +=== done 2026-09-13T19:39:25Z === +--- controls --- +real photo : (403, 'text/html; charset=utf-8', 'gunicorn') +bogus /media : (403, 'text/html; charset=utf-8', 'gunicorn') (a Django 404, not the frontend's HTML page) +frontend / : (200, 'text/html', '') +Traceback (most recent call last): + File "", line 4, in +ImportError: cannot import name 'S' from 'app' (/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/0ff9a77d-5916-4eea-91c4-eaea33e5d9a2/scratchpad/night/app.py) +=== R-483 second cut (port 80) applied to the operator's instance — 2026-09-13T19:46:23Z === +sync: HTTP 200 {"ok":true,"data":{"ok":true,"updated":["adventurelog"],"message":"Sablonok frissítve — frissítve: adventurelog"},"messa +cache port 80: 1 +--- POST /api/stacks/adventurelog/update at 2026-09-13T19:46:25Z --- +HTTP 202 +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} + + 0s phase safety-dump + + 3s phase starting + + 6s phase verifying + + 21s phase done +end (2026-09-13T19:46:47Z, +21s): state=running updating=False phase=done err='' hold='' +label after: 80 +with session: 200 image/webp 154 bytes, magic: b'RIFF' b'WEBP' +anonymous: 403 (the app's privacy rule — control) +=== done 2026-09-13T19:46:54Z === diff --git a/documentation/audits/v0241-2026-09-13/12-floor-0.241.0.txt b/documentation/audits/v0241-2026-09-13/12-floor-0.241.0.txt new file mode 100644 index 00000000..2c342b1f --- /dev/null +++ b/documentation/audits/v0241-2026-09-13/12-floor-0.241.0.txt @@ -0,0 +1,54 @@ +=== FLOOR RAISE: 0.241.0 with min_agent 0.129.0 (R-479) === +T0 POST at 2026-09-13T19:48:23Z +HTTP/1.1 303 See Other +Location: /configuration?flash=floor_set +DB floor + declared: id="global-floor-input" name="min_controller_version" value="0.241.0" | name="min_agent" value="0.129.0" placeholder="MinAgent from | +demo-hp on 0.241.0 at 2026-09-13T19:48:40Z (+16s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up 7 seconds (healthy) +demo-felhom on 0.241.0 at 2026-09-13T19:48:41Z (+18s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up Less than a second (health: starting) +--- hub log since T0 (managed floor) --- +2026/09/13 21:48:24 [INFO] Global controller-version floor set to "0.241.0" (declared MinAgent "0.129.0") +2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0) +2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0) + +--- demo-hp: controller log (SetFloor / self-update / version) --- +2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: checking update state in /opt/docker/felhom-controller/data +2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: pending update found — target=0.241.0 previous=0.240.0 +2026/09/13 19:48:33 updater.go:809: [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0) +2026/09/13 19:48:33 main.go:655: [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30) +2026/09/13 19:48:33 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply +2026/09/13 19:48:33 scheduler.go:102: [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s) +2026/09/13 19:48:33 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="selfupdate-check" interval=6h0m0s totalJobs=18 +2026/09/13 19:48:34 client.go:67: [DEBUG] [agentapi] agent version seen: 0.130.0 (was "") +2026/09/13 19:48:38 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.241.0, storagePaths=1 +2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] SetFloor: floor "" → "0.241.0" +2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] maybeAutoUpdate: current 0.241.0 >= floor 0.241.0 — no action +2026/09/13 19:48:43 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running + +--- demo-hp: bootstrap service journal --- +Sep 13 19:48:32 demo-hp systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully. +Sep 13 19:48:32 demo-hp systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). +Sep 13 19:48:32 demo-hp systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 19:48:32 demo-hp systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-hp) +Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305247]: 7d496e05c412d94587ef728dd28b513de5a834e550ca63d6ccae4bb1e7ccc780 +Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] controller started +Sep 13 19:48:32 demo-hp systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). + +--- demo-felhom: controller log (SetFloor / self-update / version) --- +2026/09/13 19:48:42 [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0) +2026/09/13 19:48:42 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30) +2026/09/13 19:48:42 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s) +2026/09/13 19:48:42 [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply +2026/09/13 19:48:52 [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running + +--- demo-felhom: bootstrap service journal --- +Sep 13 19:48:41 demo-felhom systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully. +Sep 13 19:48:41 demo-felhom systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). +Sep 13 19:48:41 demo-felhom systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 19:48:41 demo-felhom systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-felhom) +Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557928]: 52021cc5e06a35d9a7aa8832dba2372335abe431724ca6d2514b3558b8bb02e8 +Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] controller started +Sep 13 19:48:41 demo-felhom systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). + +SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: [] diff --git a/documentation/audits/v0241-2026-09-13/13-R479-live.txt b/documentation/audits/v0241-2026-09-13/13-R479-live.txt new file mode 100644 index 00000000..7808f760 --- /dev/null +++ b/documentation/audits/v0241-2026-09-13/13-R479-live.txt @@ -0,0 +1,60 @@ +=== R-479 live on demo-hp: a bind-data app (nextcloud, HDD_PATH on the registered drive) — 2026-09-13T19:48:56Z === +controller: gitea.dooplex.hu/admin/felhom-controller:0.241.0 +deploy: HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} +healthy: True +classified binds seen by the controller (app has HDD data outside its unit): 2 +--- Tier 2 OFF for the throwaway, then 'backup now' (unit only: definition + db-dump; the files stay on the drive) --- +HTTP/2 303 +location: /stacks/nextcloud/backup?flash=A+2.+ment%C3%A9s+be%C3%A1ll%C3%ADt%C3%A1sa+elmentve. + +HTTP 200 {"ok":true,"message":"Mentés elindítva"} +idle after 121 s +/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT +/mnt/felhom-drives/hdd_1/backups/primary/nextcloud: + 2026-09-13T19:49:12.5566936280 1403 db-dumps/nextcloud-mariadb.sql + 2026-09-13T19:51:13.1432085620 4484 compose/.felhom.yml + 2026-09-13T19:51:13.1432085620 4593 compose/docker-compose.yml + 2026-09-13T19:51:13.1432085620 611 compose/app.yaml + 2026-09-13T19:50:25.3716084070 170598400 volume-dumps/nextcloud_nextcloud_db_data.tar + 2026-09-13T19:50:27.9356406190 6144 volume-dumps/nextcloud_nextcloud_redis_data.tar + 2026-09-13T19:50:27.0496294880 781296640 volume-dumps/nextcloud_nextcloud_html.tar + 2026-09-13T19:51:13.1432085620 1359 manifest.json +/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT +--- a failing update: box-local template edit to an image that exits at once, put back once the job is past the pin --- +16: image: nextcloud:34.0.1-apache +1 +--- POST /api/stacks/nextcloud/update at 2026-09-13T19:51:19Z --- +HTTP 202 +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} + + 0s phase checking + + 9s phase starting + [fixture] template put back: 1 + + 13s phase verifying + +314s phase failed +end (2026-09-13T19:56:34Z, +314s): state=stopped updating=False phase=failed err='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat ' hold='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.' +--- the hold sentence --- +A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. +names 'sajat meghajto': True | says settings+db only, no files ('a fajlokat nem'): True | control 'masodik meghajto': False +--- controller log --- +2026/09/13 19:51:19 update.go:391: [INFO] [stacks] update nextcloud: accepted — guarded update started +2026/09/13 19:51:19 update.go:820: [INFO] [stacks] update nextcloud: phase checking +2026/09/13 19:51:19 update_guard.go:223: [DEBUG] [backup] update precondition for nextcloud: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz) +2026/09/13 19:51:26 update.go:506: [INFO] [stacks] update nextcloud: precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T19:51:13Z (0s old, limit 24h0m0s) +2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase safety-dump +2026/09/13 19:51:26 update.go:538: [INFO] [stacks] update nextcloud: safety dump done (1 file(s)) [/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps/pre-restore-20260913T195126Z-nextcloud-mariadb.sql] +2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pinning +2026/09/13 19:51:26 pin.go:362: [INFO] [stacks] update nextcloud: pin advanced to the catalog's current definition (nextcloud=alpine:3.20, nextcloud-db=mariadb:11.6, nextcloud-redis=redis:7-alpine) +2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pulling +2026/09/13 19:51:28 update.go:820: [INFO] [stacks] update nextcloud: phase starting +2026/09/13 19:51:30 update.go:820: [INFO] [stacks] update nextcloud: phase verifying +2026/09/13 19:56:32 update.go:625: [ERROR] [stacks] update nextcloud FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run) +2026/09/13 19:56:32 update_guard.go:530: [WARN] [backup] nextcloud is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-13T19:51:13Z; holds: "csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem") +--- teardown: restore is not needed for the proof; remove with data + backups --- + +HTTP 200 {"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":["/mnt/felhom-drives/hdd_1/appdata/nextcloud (63M)"],"hdd_paths_preserved":[],"backup_paths_removed":["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud (909M)"]},"message":"Stack nextcloud removed"} +/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT +/mnt/felhom-drives/hdd_1/backups/primary/nextcloud: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT +hold left: {'stack': 'nextcloud', 'at': '2026-09-13T19:56:32Z', 'reason': 'update_failed', 'copy_date': '2026-09-13T19:51:13Z', 'copy_tier': 1, 'copy_holds': 'csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem'} +HDD data left: ls: cannot access '/mnt/felhom-drives/hdd_1/appdata/nextcloud': No such file or directory +=== done 2026-09-13T19:56:44Z === diff --git a/documentation/audits/v0241-2026-09-13/release241.log b/documentation/audits/v0241-2026-09-13/release241.log new file mode 100644 index 00000000..e57946d9 --- /dev/null +++ b/documentation/audits/v0241-2026-09-13/release241.log @@ -0,0 +1,12 @@ +unstaged-left: [] +OK [felhom-controller]: 166 cited paths — exact 151, suffix 10, ambiguous 0, cross-repo 5, FAILED 0 (siblings searched: app-catalog-felhom.eu, felhom-agent, felhom.eu) +all controller gates OK +pre-push [felhom-controller]: gates OK - push proceeding. +HEAD=3e813307cc5688c0cce6ac1f44d284a963cddfd8 origin=3e813307cc5688c0cce6ac1f44d284a963cddfd8 porcelain=[] +build-rc=0 +0.241.0: digest: sha256:a3ef97ba2dbd570922e0876030a497b3476065d5764d712a09c1873b8850eb32 size: 856 +2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0) +2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0) +SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: [] +live-rc=0 +CHAIN-DONE diff --git a/documentation/audits/v0241-2026-09-13/rp-v241-R479-order-layout-blind.txt b/documentation/audits/v0241-2026-09-13/rp-v241-R479-order-layout-blind.txt new file mode 100644 index 00000000..bc2bb75e --- /dev/null +++ b/documentation/audits/v0241-2026-09-13/rp-v241-R479-order-layout-blind.txt @@ -0,0 +1,24 @@ +MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/backup/update_guard.go: + - if m.DataOutsideUnit(stackName) { + return updateTierOrderBindData + } + return updateTierOrder + + _ = updateTierOrderBindData + return updateTierOrder + +=== RUN TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit +[INFO] [settings] No settings.json found, using defaults +[INFO] [settings] Settings saved + r479_tier_order_test.go:40: bind=true: order [2 1 3], want [2 3 1] +[INFO] [settings] No settings.json found, using defaults +[INFO] [settings] Settings saved +[INFO] [settings] No settings.json found, using defaults +[INFO] [settings] Settings saved + r479_tier_order_test.go:60: bind=true: chose tier 1, want 3 +[INFO] [settings] No settings.json found, using defaults +[INFO] [settings] Settings saved +--- FAIL: TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s +FAIL +exit=1 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 3015b2fa..76bf4917 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -279,3 +279,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-461** | **`runbooks/target-selection.md` named a venue that does not exist and fenced a fixture that is gone.** Closed 2026-09-13, both halves checked against both boxes first: (a) no `/mnt/nvme-1tb` on demo-hp or demo-felhom; demo-hp's NVMe is `nvme0n1` at `/mnt/hdd_1` (demo-felhom's `/mnt/hdd_1` is `sdb`) — the runbook now names `/mnt/hdd_1` and says it is the same disk as the data drive; (b) `qm list` is empty on BOTH boxes — `drill-r50` (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | **CLOSED 2026-09-13 — DOCUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` | | **R-452** | **Nothing enforced `catalog_since`, so the badge's one number could silently under-report.** Closed 2026-09-13 (catalog): `scripts/check-catalog-since.py`, the fifth gate in `catalog_gates.py` — a `--range A..B` gate in the engine-major shape: an app whose per-service `image:` lines differ across the range must carry a `catalog_since` on or after the moving commit's day and not in the future; comments, README and CHANGELOG mentions are not the fact. Hook-enforced; the shallow CI clone skips it out loud (the CI-shape half the row named stays as is, by the same reasoning engine-major uses). Five decoy cases; red-proof: dropping the date comparison lets the untouched-date fact through (`audits/v0240-2026-09-13/rp-R452.txt`). | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` | | **R-465** | **`cfg.Paths.HDDPath` — empty on every box — still had six readers; were any inert?** AUDITED 2026-09-13 on demo-hp (registered drive `/mnt/felhom-drives/hdd_1`, `hdd_path` absent, no `FELHOM_PATHS_HDD_PATH`). Five of six fall back before the value matters: `report/builder.go:69` and `monitor/healthcheck.go:35` take `storagePaths[0]`; `web/server.go:740` (`primaryHDDPath`) takes the default storage path; `main.go:511` (metrics) takes the default storage path; `main.go:347` passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. **One is inert AND unreachable:** `api/router.go:981` (`systemInfo`, `GET /api/system/info`) reads the empty value with no fallback (`hdd_configured:false` forever) — and the endpoint itself is shadowed: the web layer's `ServeSystemAPI` claims `/api/system/*` and answers **404 „ismeretlen végpont"** for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as **R-490**. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | **CLOSED 2026-09-13 — AUDITED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` | +| **R-479** | **For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so.** Operator ruling 2026-09-13; closed in controller **v0.241.0** (`3e81330`): an app with classified binds walks second drive → off-site → own unit (`UpdateTierOrderFor`), and the hold sentence ends with what the chosen copy holds (`RestoreHold.CopyHolds`). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). `audits/v0241-2026-09-13/` | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` | +| **R-483** | **adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog).** Closed 2026-09-13: the k3s ingress routes `/media`, `/static`, `/admin`, `/accounts` to the backend service on port **80** — the nginx inside the backend image that serves Django's `X-Accel-Redirect` media; the catalog routed everything to the frontend. Two catalog cuts (`3172258` router, `ed62cfd` port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (`GET /media/…webp` with a session → 200 `image/webp`, RIFF/WEBP) and **confirmed by the operator in a browser at 21:49** (two photos render). Scripted multipart uploads through the frontend's `/api` proxy still 500 (upstream `RequestContentLengthMismatchError`); browser uploads work — not a template matter. `audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt` | **CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 22245cd5..6898bad1 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -695,13 +695,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** | -| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | -| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | -| **R-483** | **[P3-LOW] `adventurelog` v0.12.1: a photo upload through the app's own API proxy fails 500 from every non-browser client — whether a household can upload photos in a browser is UNKNOWN and cannot be measured here.** MEASURED 2026-09-13 on demo-hp (throwaway `travelnight`): `POST /api/images` (multipart: `image`, `location`, `is_primary`) through the public origin returns `{"error":"Internal Server Error"}` from the **frontend** (the backend log shows no request at all); the frontend container logs `RequestContentLengthMismatchError: Request body length does not match content-length header` from its undici forwarder. Same result with an ASCII filename, an accented one, urllib, curl `-F`, and curl with `Transfer-Encoding: chunked`. Everything else on the API — sign-up, the frontend's login form, collections, locations, visits, notes, edit, delete — works. **What is NOT established:** whether the app's own browser page uploads succeed (a browser's FormData body may satisfy the proxy). Per `CLAUDE.md`, that is a manual click-through: open `https://.`, add a photo to a place. **Not ours to fix in the template** (a frontend proxy bug); a newer catalog pin is a version promotion, not tonight's. Evidence: `audits/nightly-2026-09-13-adventurelog/03e-photo-500.txt`. | **WAITING-ON-OPERATOR — needs a browser click-through; rank P3-LOW; owner: VIKTOR checks, CC re-files** | +| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). **RULED 2026-09-13:** a second enrolled LXC on demo-hp, disk on the NVMe path (`/mnt/hdd_1`), sized like 9201, enrolled as a scratch CUSTOMER with its disposition recorded. **MEASURED BEFORE BUILDING — the ruling collides with the product's own model:** the hub's `hosts` table keys a host to exactly ONE customer (`host_id` PK, `customer_id NOT NULL`) and the box runs ONE agent with ONE `host_id`; a second customer on the same Proxmox host would need a second host identity, and the installer (`felhom-host-install.sh --customer-id … --vmid …`) run with another customer-id on demo-hp would rewrite the existing agent's identity — i.e. break the standing demo-hp enrolment. So "enrolled as a scratch customer" is not something the product can do on a box that already belongs to a customer. **What the product CAN do today, two options:** (a) **a second guest of the demo-hp customer** (the `guests` table is per host+vmid; the agent's provision writes a per-vmid bootstrap): supported by the data model, but both controllers report as customer demo-hp and the customer page shows one controller — the standing box's monitoring flips between the two unless the scratch guest's controller is told no hub (unenrolled scratch, disposition recorded on demo-hp's customer page); (b) **a separate scratch HOST** — a nested Proxmox VM on demo-hp (the ISO appliance route the earlier walks used, VM 323/325) enrolled as its own customer: fully enrolled, fully isolated, but an appliance to build (hours) and keep. Disk placement is the same under both: the installer has no rootfs-storage flag (only `--rootfs-grow`, `--archive-storage`), so the guest lands on `local-lvm` and is moved with `pct move-volume` to a `dir` storage created at `/mnt/hdd_1` — a post-provision step, reversible. 9201's shape for "sized like 9201": 7 cores, 25 898 MB, rootfs 32 G + mp0 70 G on local-lvm, unprivileged. **Recommendation: (a) with the hub left out** — it is the reversible one and it gives the rotation its restore target tomorrow; (b) if "enrolled" is the point. Not built tonight: §1 of the rules — a decision that changes what the product promises about hosts is not CC's. | **WAITING-ON-OPERATOR — the ruling cannot be executed as stated; two options below; rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | | **R-487** | **[P2-MEDIUM] A removed app whose backups were kept is listed on NEITHER backup page, so the restore that brings it back has no button — the customer's remove-by-mistake route exists only as an endpoint.** MEASURED 2026-09-13 on demo-hp (nightly rotation, `adventurelog` removed with backups kept, unit + mirror on disk): `GET /backups/apps` and `GET /backups/restore` contain the string `adventurelog` zero times; `POST /backup/restore stack_name=adventurelog snapshot_id=helyi` then restored it in 22 s with the data byte-identical. Cause: `buildAppBackupRows` walks `status.AppDataInfo` = `DiscoverAppData` over DEPLOYED stacks only. The off-site list had exactly this defect and was fixed by keying it on the store (R-237, v0.204.0); the local and Tier-2 lists were not. **Fix shape:** list every app with a recovery unit on a registered drive (`ListRestorePoints` over the primary dirs), marking removed ones „eltávolítva — visszaállítható"; the unit restore already reinstalls (R-253). Not a design reversal — the same rule R-237 set. Evidence: `audits/nightly-2026-09-13-adventurelog/05b-restore-tier1.txt`. | **READY — rank P2-MEDIUM; owner: CC** | | **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** | | **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). | **READY — rank P3-LOW; owner: CC** | | **R-490** | **[P3-LOW] The monitoring page's „Memória-eloszlás" card has never rendered: its `fetch('/api/system/info')` is answered 404 by the web layer's `ServeSystemAPI`, which claims all of `/api/system/*` and knows only the two memory routes.** MEASURED 2026-09-13 on demo-hp: `GET /api/system/info` → `{"error":"ismeretlen végpont","ok":false}`; `monitoring.html` shows the card (`display:none` by default) only when that fetch returns `used_mem_mb`, so it stays hidden on every box. The API router's `systemInfo` handler (`internal/api/router.go:981`) is unreachable, and it is also the one reader of the always-empty `cfg.Paths.HDDPath` with no fallback (R-465). **Fix shape:** let `ServeSystemAPI` fall through to the API router for unknown `/api/system/*` paths (or route `/api/system/info` explicitly), give `systemInfo` the same default-storage-path fallback the other readers have, and pin the card with a render test; then delete the global (R-465's deferred deletion). Next controller release. Evidence: `audits/nightly-2026-09-13-adventurelog/` (audit notes in the R-465 closure). | **READY — rank P3-LOW; owner: CC** | +| **R-491** | **[P2-MEDIUM] Removing an app leaves its update hold in the store, so a reinstall under the same name starts HELD.** MEASURED 2026-09-13 on demo-hp (v0.241.0, R-479 live check): nextcloud was held after a failed update, then removed with data and backups; `settings.json` still carried `restore_holds.nextcloud` (`reason: update_failed`, `copy_tier: 1`). `fillHoldReason` hides the sentence for a not-deployed app (R-480), but every start gate reads the store, so the next deploy of `nextcloud` would be refused as held with a sentence about a backup that no longer exists. Cleared by hand with `-clear-restore-hold`. **Fix shape:** `removeStack` clears an UPDATE hold (never an R-379 restore hold, which stays operator-cleared) — same place the prefs are forgotten; a test that deploys after a held removal. Next controller release. | **READY — rank P2-MEDIUM; owner: CC** |