evening: rules file in place; R-483 CLOSED (catalog, operator-confirmed); R-479 CLOSED (controller v0.241.0, proven live); R-481 blocked on the one-host-one-customer model; R-491 filed
gates / gates (push) Successful in 18s

This commit is contained in:
2026-09-13 21:57:47 +02:00
parent 550fd84754
commit 72ee053a9e
12 changed files with 312 additions and 37 deletions
@@ -277,7 +277,7 @@ what the unit holds** — for an app whose data is a bind mount that is the defi
(R-479). **Removal and the tiers (controller v0.240.0):** removing an app with its backups KEPT keeps
the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás" still works afterwards
(R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and
never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open).
never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open). **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.**
### 6.1 The four tiers, as configured on the live fleet
@@ -178,6 +178,14 @@ These are rulings, not proposals. Anything specced against a different assumptio
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy
holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit
that holds the definition and the database dumps and NOT the files (measured: gokapi restored from
„helyi" came back with settings and no data). For such an app the update walks second drive →
off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive →
own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a
customer is never sent to a copy that cannot bring the data back without being told so.
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
@@ -0,0 +1,79 @@
=== R-483 repro on demo-hp — 2026-09-13T19:35:02Z ===
sync: HTTP 200 {"ok":true,"data":{"ok":true,"message":"Sablonok naprakészek — nincs változás"},"message":"Sablonok naprakészek — nincs változás"}
cache has the backend router: 1
--- the operator's own throwaway instance (subdomain travel, 253 media files) is used; its containers get the new routers through the guarded Update (same version) ---
labels before: (none)
--- POST /api/stacks/adventurelog/update at 2026-09-13T19:35:05Z ---
HTTP 202
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
+ 0s phase safety-dump
+ 3s phase starting
+ 6s phase verifying
+ 21s phase done
end (2026-09-13T19:35:26Z, +21s): state=running updating=False phase=done err='' hold=''
labels after: Host(`travel.enkisfelhom.hu`) && (PathPrefix(`/media`) || PathPrefix(`/static`)
healthy: True
signup: 200
login: (302, '/')
location: 201
--- upload through the frontend proxy, WebKit-shaped boundary (what a browser sends) ---
proxy upload -> 500 {"id": null, "image": null}
--- proxy still refuses non-browser multipart; upload through the backend's own API inside the guest (the browser's session cookie, same backend) ---
backend upload -> 403
{"detail":"CSRF Failed: CSRF cookie not set."}
image url: None
--- control: /static and /admin now answer from Django ---
/static/admin/css/base.css -> 200
/admin/login/ -> 403
=== done 2026-09-13T19:35:34Z ===
=== R-483 repro, second half — 2026-09-13T19:36:33Z: an upload placed through the backend's own API (fresh CSRF cookie from the backend), then the photo fetched through the PUBLIC origin ===
login: (302, '/')
have sessionid: True
my place id found: True
backend upload -> 403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:36:37Z ===
=== R-483 repro, third try (the CSRF token is read with cut -f7, quoting-proof) — 2026-09-13T19:37:35Z ===
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:37:39Z ===
=== R-483 repro, fourth try (CSRF cookie from /admin/login/, which always sets one) — 2026-09-13T19:38:02Z ===
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:38:05Z ===
=== R-483 repro, fifth try (curl drops a Secure cookie over http; the token is cut from the raw Set-Cookie header) — 2026-09-13T19:38:30Z ===
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:38:33Z ===
=== R-483 repro, sixth try (token cut from the config endpoint's Set-Cookie header) — 2026-09-13T19:39:01Z ===
backend upload -> tokenlen=32 | http=400 | {"error":"content_type and object_id are required"}
image url: None
=== done 2026-09-13T19:39:05Z ===
=== R-483 repro, seventh try (v0.12.1's image API takes content_type=location + object_id) — 2026-09-13T19:39:21Z ===
backend upload -> http=201 | {"id":"c0b2b8cd-fe22-45db-b2a2-7c859d771941","image":"https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp","is_primary":false,"user":"51f9e8ee-f303-430e-b8ab-aae231e758c4","immich_id":null}
image url: https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp
GET /media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp through traefik (public origin) -> 200; bytes identical to the upload: False
the app lists the photo on the place: ['https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp']
=== done 2026-09-13T19:39:25Z ===
--- controls ---
real photo : (403, 'text/html; charset=utf-8', 'gunicorn')
bogus /media : (403, 'text/html; charset=utf-8', 'gunicorn') (a Django 404, not the frontend's HTML page)
frontend / : (200, 'text/html', '')
Traceback (most recent call last):
File "<stdin>", line 4, in <module>
ImportError: cannot import name 'S' from 'app' (/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/0ff9a77d-5916-4eea-91c4-eaea33e5d9a2/scratchpad/night/app.py)
=== R-483 second cut (port 80) applied to the operator's instance — 2026-09-13T19:46:23Z ===
sync: HTTP 200 {"ok":true,"data":{"ok":true,"updated":["adventurelog"],"message":"Sablonok frissítve — frissítve: adventurelog"},"messa
cache port 80: 1
--- POST /api/stacks/adventurelog/update at 2026-09-13T19:46:25Z ---
HTTP 202
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
+ 0s phase safety-dump
+ 3s phase starting
+ 6s phase verifying
+ 21s phase done
end (2026-09-13T19:46:47Z, +21s): state=running updating=False phase=done err='' hold=''
label after: 80
with session: 200 image/webp 154 bytes, magic: b'RIFF' b'WEBP'
anonymous: 403 (the app's privacy rule — control)
=== done 2026-09-13T19:46:54Z ===
@@ -0,0 +1,54 @@
=== FLOOR RAISE: 0.241.0 with min_agent 0.129.0 (R-479) ===
T0 POST at 2026-09-13T19:48:23Z
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
DB floor + declared: id="global-floor-input" name="min_controller_version" value="0.241.0" | name="min_agent" value="0.129.0" placeholder="MinAgent from |
demo-hp on 0.241.0 at 2026-09-13T19:48:40Z (+16s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up 7 seconds (healthy)
demo-felhom on 0.241.0 at 2026-09-13T19:48:41Z (+18s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up Less than a second (health: starting)
--- hub log since T0 (managed floor) ---
2026/09/13 21:48:24 [INFO] Global controller-version floor set to "0.241.0" (declared MinAgent "0.129.0")
2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
--- demo-hp: controller log (SetFloor / self-update / version) ---
2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: checking update state in /opt/docker/felhom-controller/data
2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: pending update found — target=0.241.0 previous=0.240.0
2026/09/13 19:48:33 updater.go:809: [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0)
2026/09/13 19:48:33 main.go:655: [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/13 19:48:33 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
2026/09/13 19:48:33 scheduler.go:102: [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/13 19:48:33 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="selfupdate-check" interval=6h0m0s totalJobs=18
2026/09/13 19:48:34 client.go:67: [DEBUG] [agentapi] agent version seen: 0.130.0 (was "")
2026/09/13 19:48:38 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.241.0, storagePaths=1
2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] SetFloor: floor "" → "0.241.0"
2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] maybeAutoUpdate: current 0.241.0 >= floor 0.241.0 — no action
2026/09/13 19:48:43 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running
--- demo-hp: bootstrap service journal ---
Sep 13 19:48:32 demo-hp systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
Sep 13 19:48:32 demo-hp systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
Sep 13 19:48:32 demo-hp systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:32 demo-hp systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-hp)
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305247]: 7d496e05c412d94587ef728dd28b513de5a834e550ca63d6ccae4bb1e7ccc780
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] controller started
Sep 13 19:48:32 demo-hp systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
--- demo-felhom: controller log (SetFloor / self-update / version) ---
2026/09/13 19:48:42 [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0)
2026/09/13 19:48:42 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/13 19:48:42 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/13 19:48:42 [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
2026/09/13 19:48:52 [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running
--- demo-felhom: bootstrap service journal ---
Sep 13 19:48:41 demo-felhom systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
Sep 13 19:48:41 demo-felhom systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
Sep 13 19:48:41 demo-felhom systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:41 demo-felhom systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-felhom)
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557928]: 52021cc5e06a35d9a7aa8832dba2372335abe431724ca6d2514b3558b8bb02e8
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] controller started
Sep 13 19:48:41 demo-felhom systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: []
@@ -0,0 +1,60 @@
=== R-479 live on demo-hp: a bind-data app (nextcloud, HDD_PATH on the registered drive) — 2026-09-13T19:48:56Z ===
controller: gitea.dooplex.hu/admin/felhom-controller:0.241.0
deploy: HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
healthy: True
classified binds seen by the controller (app has HDD data outside its unit): 2
--- Tier 2 OFF for the throwaway, then 'backup now' (unit only: definition + db-dump; the files stay on the drive) ---
HTTP/2 303
location: /stacks/nextcloud/backup?flash=A+2.+ment%C3%A9s+be%C3%A1ll%C3%ADt%C3%A1sa+elmentve.
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
idle after 121 s
/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud:
2026-09-13T19:49:12.5566936280 1403 db-dumps/nextcloud-mariadb.sql
2026-09-13T19:51:13.1432085620 4484 compose/.felhom.yml
2026-09-13T19:51:13.1432085620 4593 compose/docker-compose.yml
2026-09-13T19:51:13.1432085620 611 compose/app.yaml
2026-09-13T19:50:25.3716084070 170598400 volume-dumps/nextcloud_nextcloud_db_data.tar
2026-09-13T19:50:27.9356406190 6144 volume-dumps/nextcloud_nextcloud_redis_data.tar
2026-09-13T19:50:27.0496294880 781296640 volume-dumps/nextcloud_nextcloud_html.tar
2026-09-13T19:51:13.1432085620 1359 manifest.json
/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT
--- a failing update: box-local template edit to an image that exits at once, put back once the job is past the pin ---
16: image: nextcloud:34.0.1-apache
1
--- POST /api/stacks/nextcloud/update at 2026-09-13T19:51:19Z ---
HTTP 202
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
+ 0s phase checking
+ 9s phase starting
[fixture] template put back: 1
+ 13s phase verifying
+314s phase failed
end (2026-09-13T19:56:34Z, +314s): state=stopped updating=False phase=failed err='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat ' hold='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.'
--- the hold sentence ---
A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
names 'sajat meghajto': True | says settings+db only, no files ('a fajlokat nem'): True | control 'masodik meghajto': False
--- controller log ---
2026/09/13 19:51:19 update.go:391: [INFO] [stacks] update nextcloud: accepted — guarded update started
2026/09/13 19:51:19 update.go:820: [INFO] [stacks] update nextcloud: phase checking
2026/09/13 19:51:19 update_guard.go:223: [DEBUG] [backup] update precondition for nextcloud: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz)
2026/09/13 19:51:26 update.go:506: [INFO] [stacks] update nextcloud: precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T19:51:13Z (0s old, limit 24h0m0s)
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase safety-dump
2026/09/13 19:51:26 update.go:538: [INFO] [stacks] update nextcloud: safety dump done (1 file(s)) [/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps/pre-restore-20260913T195126Z-nextcloud-mariadb.sql]
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pinning
2026/09/13 19:51:26 pin.go:362: [INFO] [stacks] update nextcloud: pin advanced to the catalog's current definition (nextcloud=alpine:3.20, nextcloud-db=mariadb:11.6, nextcloud-redis=redis:7-alpine)
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pulling
2026/09/13 19:51:28 update.go:820: [INFO] [stacks] update nextcloud: phase starting
2026/09/13 19:51:30 update.go:820: [INFO] [stacks] update nextcloud: phase verifying
2026/09/13 19:56:32 update.go:625: [ERROR] [stacks] update nextcloud FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
2026/09/13 19:56:32 update_guard.go:530: [WARN] [backup] nextcloud is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-13T19:51:13Z; holds: "csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem")
--- teardown: restore is not needed for the proof; remove with data + backups ---
HTTP 200 {"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":["/mnt/felhom-drives/hdd_1/appdata/nextcloud (63M)"],"hdd_paths_preserved":[],"backup_paths_removed":["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud (909M)"]},"message":"Stack nextcloud removed"}
/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud: ABSENT
/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT
hold left: {'stack': 'nextcloud', 'at': '2026-09-13T19:56:32Z', 'reason': 'update_failed', 'copy_date': '2026-09-13T19:51:13Z', 'copy_tier': 1, 'copy_holds': 'csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem'}
HDD data left: ls: cannot access '/mnt/felhom-drives/hdd_1/appdata/nextcloud': No such file or directory
=== done 2026-09-13T19:56:44Z ===
@@ -0,0 +1,12 @@
unstaged-left: []
OK [felhom-controller]: 166 cited paths — exact 151, suffix 10, ambiguous 0, cross-repo 5, FAILED 0 (siblings searched: app-catalog-felhom.eu, felhom-agent, felhom.eu)
all controller gates OK
pre-push [felhom-controller]: gates OK - push proceeding.
HEAD=3e813307cc5688c0cce6ac1f44d284a963cddfd8 origin=3e813307cc5688c0cce6ac1f44d284a963cddfd8 porcelain=[]
build-rc=0
0.241.0: digest: sha256:a3ef97ba2dbd570922e0876030a497b3476065d5764d712a09c1873b8850eb32 size: 856
2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: []
live-rc=0
CHAIN-DONE
@@ -0,0 +1,24 @@
MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/backup/update_guard.go:
- if m.DataOutsideUnit(stackName) {
return updateTierOrderBindData
}
return updateTierOrder
+ _ = updateTierOrderBindData
return updateTierOrder
=== RUN TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
r479_tier_order_test.go:40: bind=true: order [2 1 3], want [2 3 1]
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
r479_tier_order_test.go:60: bind=true: chose tier 1, want 3
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
--- FAIL: TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s
FAIL
exit=1
+2
View File
@@ -279,3 +279,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-461** | **`runbooks/target-selection.md` named a venue that does not exist and fenced a fixture that is gone.** Closed 2026-09-13, both halves checked against both boxes first: (a) no `/mnt/nvme-1tb` on demo-hp or demo-felhom; demo-hp's NVMe is `nvme0n1` at `/mnt/hdd_1` (demo-felhom's `/mnt/hdd_1` is `sdb`) — the runbook now names `/mnt/hdd_1` and says it is the same disk as the data drive; (b) `qm list` is empty on BOTH boxes — `drill-r50` (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | **CLOSED 2026-09-13 — DOCUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
| **R-452** | **Nothing enforced `catalog_since`, so the badge's one number could silently under-report.** Closed 2026-09-13 (catalog): `scripts/check-catalog-since.py`, the fifth gate in `catalog_gates.py` — a `--range A..B` gate in the engine-major shape: an app whose per-service `image:` lines differ across the range must carry a `catalog_since` on or after the moving commit's day and not in the future; comments, README and CHANGELOG mentions are not the fact. Hook-enforced; the shallow CI clone skips it out loud (the CI-shape half the row named stays as is, by the same reasoning engine-major uses). Five decoy cases; red-proof: dropping the date comparison lets the untouched-date fact through (`audits/v0240-2026-09-13/rp-R452.txt`). | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
| **R-465** | **`cfg.Paths.HDDPath` — empty on every box — still had six readers; were any inert?** AUDITED 2026-09-13 on demo-hp (registered drive `/mnt/felhom-drives/hdd_1`, `hdd_path` absent, no `FELHOM_PATHS_HDD_PATH`). Five of six fall back before the value matters: `report/builder.go:69` and `monitor/healthcheck.go:35` take `storagePaths[0]`; `web/server.go:740` (`primaryHDDPath`) takes the default storage path; `main.go:511` (metrics) takes the default storage path; `main.go:347` passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. **One is inert AND unreachable:** `api/router.go:981` (`systemInfo`, `GET /api/system/info`) reads the empty value with no fallback (`hdd_configured:false` forever) — and the endpoint itself is shadowed: the web layer's `ServeSystemAPI` claims `/api/system/*` and answers **404 „ismeretlen végpont"** for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as **R-490**. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | **CLOSED 2026-09-13 — AUDITED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
| **R-479** | **For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so.** Operator ruling 2026-09-13; closed in controller **v0.241.0** (`3e81330`): an app with classified binds walks second drive → off-site → own unit (`UpdateTierOrderFor`), and the hold sentence ends with what the chosen copy holds (`RestoreHold.CopyHolds`). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). `audits/v0241-2026-09-13/` | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
| **R-483** | **adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog).** Closed 2026-09-13: the k3s ingress routes `/media`, `/static`, `/admin`, `/accounts` to the backend service on port **80** — the nginx inside the backend image that serves Django's `X-Accel-Redirect` media; the catalog routed everything to the frontend. Two catalog cuts (`3172258` router, `ed62cfd` port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (`GET /media/…webp` with a session → 200 `image/webp`, RIFF/WEBP) and **confirmed by the operator in a browser at 21:49** (two photos render). Scripted multipart uploads through the frontend's `/api` proxy still 500 (upstream `RequestContentLengthMismatchError`); browser uploads work — not a template matter. `audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt` | **CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
+2 -3
View File
@@ -695,13 +695,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-483** | **[P3-LOW] `adventurelog` v0.12.1: a photo upload through the app's own API proxy fails 500 from every non-browser client — whether a household can upload photos in a browser is UNKNOWN and cannot be measured here.** MEASURED 2026-09-13 on demo-hp (throwaway `travelnight`): `POST /api/images` (multipart: `image`, `location`, `is_primary`) through the public origin returns `{"error":"Internal Server Error"}` from the **frontend** (the backend log shows no request at all); the frontend container logs `RequestContentLengthMismatchError: Request body length does not match content-length header` from its undici forwarder. Same result with an ASCII filename, an accented one, urllib, curl `-F`, and curl with `Transfer-Encoding: chunked`. Everything else on the API — sign-up, the frontend's login form, collections, locations, visits, notes, edit, delete — works. **What is NOT established:** whether the app's own browser page uploads succeed (a browser's FormData body may satisfy the proxy). Per `CLAUDE.md`, that is a manual click-through: open `https://<sub>.<domain>`, add a photo to a place. **Not ours to fix in the template** (a frontend proxy bug); a newer catalog pin is a version promotion, not tonight's. Evidence: `audits/nightly-2026-09-13-adventurelog/03e-photo-500.txt`. | **WAITING-ON-OPERATOR — needs a browser click-through; rank P3-LOW; owner: VIKTOR checks, CC re-files** |
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). **RULED 2026-09-13:** a second enrolled LXC on demo-hp, disk on the NVMe path (`/mnt/hdd_1`), sized like 9201, enrolled as a scratch CUSTOMER with its disposition recorded. **MEASURED BEFORE BUILDING — the ruling collides with the product's own model:** the hub's `hosts` table keys a host to exactly ONE customer (`host_id` PK, `customer_id NOT NULL`) and the box runs ONE agent with ONE `host_id`; a second customer on the same Proxmox host would need a second host identity, and the installer (`felhom-host-install.sh --customer-id … --vmid …`) run with another customer-id on demo-hp would rewrite the existing agent's identity — i.e. break the standing demo-hp enrolment. So "enrolled as a scratch customer" is not something the product can do on a box that already belongs to a customer. **What the product CAN do today, two options:** (a) **a second guest of the demo-hp customer** (the `guests` table is per host+vmid; the agent's provision writes a per-vmid bootstrap): supported by the data model, but both controllers report as customer demo-hp and the customer page shows one controller — the standing box's monitoring flips between the two unless the scratch guest's controller is told no hub (unenrolled scratch, disposition recorded on demo-hp's customer page); (b) **a separate scratch HOST** — a nested Proxmox VM on demo-hp (the ISO appliance route the earlier walks used, VM 323/325) enrolled as its own customer: fully enrolled, fully isolated, but an appliance to build (hours) and keep. Disk placement is the same under both: the installer has no rootfs-storage flag (only `--rootfs-grow`, `--archive-storage`), so the guest lands on `local-lvm` and is moved with `pct move-volume` to a `dir` storage created at `/mnt/hdd_1` — a post-provision step, reversible. 9201's shape for "sized like 9201": 7 cores, 25 898 MB, rootfs 32 G + mp0 70 G on local-lvm, unprivileged. **Recommendation: (a) with the hub left out** — it is the reversible one and it gives the rotation its restore target tomorrow; (b) if "enrolled" is the point. Not built tonight: §1 of the rules — a decision that changes what the product promises about hosts is not CC's. | **WAITING-ON-OPERATOR — the ruling cannot be executed as stated; two options below; rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-487** | **[P2-MEDIUM] A removed app whose backups were kept is listed on NEITHER backup page, so the restore that brings it back has no button — the customer's remove-by-mistake route exists only as an endpoint.** MEASURED 2026-09-13 on demo-hp (nightly rotation, `adventurelog` removed with backups kept, unit + mirror on disk): `GET /backups/apps` and `GET /backups/restore` contain the string `adventurelog` zero times; `POST /backup/restore stack_name=adventurelog snapshot_id=helyi` then restored it in 22 s with the data byte-identical. Cause: `buildAppBackupRows` walks `status.AppDataInfo` = `DiscoverAppData` over DEPLOYED stacks only. The off-site list had exactly this defect and was fixed by keying it on the store (R-237, v0.204.0); the local and Tier-2 lists were not. **Fix shape:** list every app with a recovery unit on a registered drive (`ListRestorePoints` over the primary dirs), marking removed ones „eltávolítva — visszaállítható"; the unit restore already reinstalls (R-253). Not a design reversal — the same rule R-237 set. Evidence: `audits/nightly-2026-09-13-adventurelog/05b-restore-tier1.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). | **READY — rank P3-LOW; owner: CC** |
| **R-490** | **[P3-LOW] The monitoring page's „Memória-eloszlás" card has never rendered: its `fetch('/api/system/info')` is answered 404 by the web layer's `ServeSystemAPI`, which claims all of `/api/system/*` and knows only the two memory routes.** MEASURED 2026-09-13 on demo-hp: `GET /api/system/info` → `{"error":"ismeretlen végpont","ok":false}`; `monitoring.html` shows the card (`display:none` by default) only when that fetch returns `used_mem_mb`, so it stays hidden on every box. The API router's `systemInfo` handler (`internal/api/router.go:981`) is unreachable, and it is also the one reader of the always-empty `cfg.Paths.HDDPath` with no fallback (R-465). **Fix shape:** let `ServeSystemAPI` fall through to the API router for unknown `/api/system/*` paths (or route `/api/system/info` explicitly), give `systemInfo` the same default-storage-path fallback the other readers have, and pin the card with a render test; then delete the global (R-465's deferred deletion). Next controller release. Evidence: `audits/nightly-2026-09-13-adventurelog/` (audit notes in the R-465 closure). | **READY — rank P3-LOW; owner: CC** |
| **R-491** | **[P2-MEDIUM] Removing an app leaves its update hold in the store, so a reinstall under the same name starts HELD.** MEASURED 2026-09-13 on demo-hp (v0.241.0, R-479 live check): nextcloud was held after a failed update, then removed with data and backups; `settings.json` still carried `restore_holds.nextcloud` (`reason: update_failed`, `copy_tier: 1`). `fillHoldReason` hides the sentence for a not-deployed app (R-480), but every start gate reads the store, so the next deploy of `nextcloud` would be refused as held with a sentence about a backup that no longer exists. Cleared by hand with `-clear-restore-hold`. **Fix shape:** `removeStack` clears an UPDATE hold (never an R-379 restore hold, which stays operator-cleared) — same place the prefs are forgotten; a test that deploys after a held removal. Next controller release. | **READY — rank P2-MEDIUM; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.