evening: rules file in place; R-483 CLOSED (catalog, operator-confirmed); R-479 CLOSED (controller v0.241.0, proven live); R-481 blocked on the one-host-one-customer model; R-491 filed
gates / gates (push) Successful in 18s

This commit is contained in:
2026-09-13 21:57:47 +02:00
parent 550fd84754
commit 72ee053a9e
12 changed files with 312 additions and 37 deletions
+6
View File
@@ -34,6 +34,12 @@ on demo-hp — `documentation/audits/slice4-2026-09-13/`, design `09-update-arch
- **Removal keeps the backups AND the Tier-2 record unless the customer asked for the backups to go**
(v0.240.0, R-486/R-474). A removal that forgets the record makes an intact mirror unrestorable — that
was live until v0.239.0. The first nightly rotation (adventurelog, 2026-09-13) found it.
- **A bind-data app's route back is off-site before its own unit, and the hold names what the copy
holds** (operator ruling 2026-09-13, R-479, controller v0.241.0). The unit of such an app holds the
definition and the database dumps, not the files. Volume-data apps keep the v0.239.0 order.
- **One host, one customer** (schema: `hosts.host_id` PK, one `customer_id`; one agent per box). A
scratch guest "enrolled as its own customer" on a box that already belongs to a customer is not
something the product can do (R-481, 2026-09-13); a second guest of the same customer is.
- **A release reaches the fleet by floor between golden bakes when its MinAgent is declared with the
floor** (ruling 2026-09-13, hub v0.112.0, R-472 CLOSED). An undeclared floor above the golden is still
held, and the forms refuse it. v0.239.0 arrived on both demo boxes this way in ~15 s. Never hand-deploy
+35 -33
View File
@@ -1,43 +1,45 @@
# REPORT — the first "be a customer for the night" run: adventurelog on demo-hp (2026-09-13/14)
# REPORT — the four evening items before the second night (2026-09-13)
*Overwritten each session. Every finding below has a register row. The controller report is
*Overwritten each session. The night's own report (adventurelog rotation, v0.240.0) is preserved as
`audits/nightly-2026-09-13-adventurelog/` and the register; the controller report is
`felhom-controller/REPORT.md`, the catalog's `app-catalog-felhom.eu/REPORT.md`.*
> One app walked end to end as a household would; seven product defects found; a controller release
> (v0.240.0) shipped the same night by the managed floor; two catalog fixes; two docs fixes. The
> morning note is the top block of `STATUS.md`.
## 1. The rules file
## Step 0 — the rules file
`.claude/rules/unprompted-work.md` copied byte-identical into `felhom-controller`, `felhom.eu` and
`app-catalog-felhom.eu` (documents-only pushes). `instructions_gate.py` refused the first push: a rule
file needs `paths:` or `unconditional: true`; all four copies (root included) now start with the
three-line `unconditional: true` frontmatter — sha256 `d1aa3af83ae3b220…` everywhere.
`.claude/rules/unprompted-work.md` is in none of the three repos and the brief's text block did not
reach this session (searched the workspace, the transcript and memory). Not created — the operator's
rules cannot be invented byte-identically. The brief's own fences were applied instead.
## 2. R-481 — the scratch guest (ruled; NOT built, and why)
## What changed here
The ruling ("a second enrolled LXC … enrolled as a scratch customer") collides with the model: the hub
keys a host to one customer (`hosts.host_id` PK) and the box runs one agent with one host identity;
re-running the installer with another customer-id would rewrite demo-hp's enrolment. Options and a
recommendation are in the row and STATUS; 9201's shape and the disk placement (`pct move-volume`
onto a `dir` storage at `/mnt/hdd_1`) are recorded for whichever is chosen.
- `documentation/runbooks/nightly-rotation.md` — new: the tick list (standing nine first, blocked on
R-481; then the catalog); `actualbudget` recorded (no non-browser front door), `adventurelog` ticked.
- `scripts/observations_gate.py` — R-471: every observations section is read (`scripts/CHANGELOG.md`).
- `documentation/runbooks/target-selection.md` — R-461: `/mnt/hdd_1` is the NVMe; `drill-r50` exists
on neither box (measured), fence kept with the fact beside it; R-93 carries the fact.
- `09-update-architecture.md` §6.1 + §8.2 (R-452 lifted), `07-backup-architecture.md` §6, `CONTEXT.md` — v0.240.0 notes.
- Catalog (`app-catalog-felhom.eu`): the `catalog-since` gate (R-452), see its REPORT.
- Register: R-465 (audited: five readers fall back, one is inert and unreachable → R-490), R-452, R-473, R-474, R-466, R-471, R-453, R-461, R-477, R-478, R-480, R-482, R-484, R-485, R-486
closed and compressed; R-481, R-482..R-489 opened (R-482 closed the same night). 212 → 206 open,
177 → 189 closed (byte sizes: see the commit).
- Evidence: `documentation/audits/nightly-2026-09-13-adventurelog/` (01–08), `audits/v0240-2026-09-13/`
(release log, floor, validation, red-proofs incl. R-471).
## 3. R-483 — photos (re-ranked P2, catalog)
Diffed against `homelab-manifests` `adventurelog-system/adventurelog.yaml`: the k3s ingress sends
`/media`, `/static`, `/admin`, `/accounts` to the backend **service port 80**; the catalog sent
everything to the frontend. Two catalog cuts (`3172258`, `ed62cfd`): a backend traefik router for the
four prefixes, then port 80 instead of 8000 — the backend image runs nginx in front of gunicorn and
Django serves a photo by `X-Accel-Redirect`, so port 8000 returned 200 with an empty body. Applied to
the operator's own instance through the guarded Update; proven without a browser
(`audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt`): `GET /media/images/<id>.webp` with a
session → 200 `image/webp`, RIFF/WEBP magic; anonymous 403 (the app's rule). **Confirmed by the operator in a browser at 21:49 — closed.** The frontend's 500 on scripted multipart uploads is upstream and
recorded in the row; browser uploads work.
## 4. R-479 — tier order for bind-data apps (ruled; controller v0.241.0)
Ruling 9 in `09-update-architecture.md` §3. Built, tested, red-proofed; delivered by the floor and
live-checked on a nextcloud throwaway — see `felhom-controller/REPORT.md` and
`audits/v0241-2026-09-13/`.
## Observations
1. **The brief's rules block was not delivered; the file is absent in all repos.** NOT-A-FINDING: an input gap for the operator, recorded in STATUS "needs you", not a product defect.
2. **No scratch guest on demo-hp; the standing nine cannot be throwaways on 9201.** FILED: R-481
3. **adventurelog ran Django DEBUG=True on the public origin.** FILED: R-482
4. **Photo upload fails 500 from every non-browser client (frontend proxy body-length bug).** FILED: R-483
5. **PostGIS not recognised as a database.** FILED: R-484
6. **The backup card reads dead paths.** FILED: R-485
7. **Removal with backups kept forgot the Tier-2 record; the restore was refused.** FILED: R-486
8. **A removed app is on neither backup page.** FILED: R-487
9. **The backup test package takes 5½ minutes of real waits.** FILED: R-488
10. **`volumes_removed: null` over removed volumes, five times.** FILED: R-489
11. **The monitoring page's memory-distribution card never renders — `/api/system/info` is shadowed by `ServeSystemAPI` (found by the R-465 audit).** FILED: R-490
0. **A removed app keeps its update hold; a reinstall would start held.** FILED: R-491
1. **The controller's evening release rides the same session as the night's — two releases in one calendar day, one per session.** NOT-A-FINDING: the rule is per session; both are recorded with their MinAgent lines.
2. **`homelab-manifests` is not in the workspace; its manifests were read from Gitea's API.** NOT-A-FINDING: the workspace CLAUDE.md says non-felhom repos are unrelated; a read-only fetch was enough.
+29
View File
@@ -1,5 +1,34 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-13 (evening, before the night) — your four items.**
**Decisions I took.** (1) The rules file now sits in the workspace root and all three repos,
byte-identical; the gate that checks instruction files required three lines at its top saying it
loads in every session on purpose, so all four copies carry them. (2) The photo fix was applied to
your own test instance on the HP through the Update button, not by hand, and I added one test
account with one photo there.
**What I did.** The rules file: done. The photo problem: the app's backend serves photos itself,
through a small web server inside its own image; our template sent every request to the front page
server, which has no photos. Fixed in two cuts — the second one because the first spoke to the wrong
port and got empty pictures. Proven without a browser: a photo now comes back as a real image
through the public address, and you confirmed it in your browser at 21:49. The tier-order
ruling: built as controller 0.241.0 — an app whose data lives in mounted folders now prefers the
remote copy over its own-drive copy, and the hold message ends with what the chosen copy holds.
Live on both machines (0.241.0, delivered by the floor in 16 and 18 seconds) and proven on the HP
with a throwaway Nextcloud.
**What broke on the way.** Nothing standing. One new finding: after an app is removed, its
"held after a failed update" mark stays behind, so a reinstall would start blocked. Cleared by hand,
filed (R-491) for the next release. Register: 206 open / 191 closed.
**Needs you.** The scratch guest: the ruling as written cannot be built — the hub ties one box to
one customer, and a second enrolled customer on the HP would replace the HP's own enrolment. Two
options: a second guest under the HP's own customer (reversible, tonight-ready, but its reports
would clash with the HP's page unless it stays unenrolled), or a nested appliance VM enrolled as its
own box (fully enrolled, hours to build). My pick: the first, unenrolled. **If you do nothing:**
tonight's rotation keeps restoring in place and skips the standing nine.
**Updated 2026-09-14 (the night of 13→14 — "be a customer for the night", first run). Written in
the order the rules ask for.**
@@ -277,7 +277,7 @@ what the unit holds** — for an app whose data is a bind mount that is the defi
(R-479). **Removal and the tiers (controller v0.240.0):** removing an app with its backups KEPT keeps
the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás" still works afterwards
(R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and
never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open).
never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open). **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.**
### 6.1 The four tiers, as configured on the live fleet
@@ -178,6 +178,14 @@ These are rulings, not proposals. Anything specced against a different assumptio
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy
holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit
that holds the definition and the database dumps and NOT the files (measured: gokapi restored from
„helyi" came back with settings and no data). For such an app the update walks second drive →
off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive →
own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a
customer is never sent to a copy that cannot bring the data back without being told so.
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
@@ -0,0 +1,79 @@
=== R-483 repro on demo-hp — 2026-09-13T19:35:02Z ===
sync: HTTP 200 {"ok":true,"data":{"ok":true,"message":"Sablonok naprakészek — nincs változás"},"message":"Sablonok naprakészek — nincs változás"}
cache has the backend router: 1
--- the operator's own throwaway instance (subdomain travel, 253 media files) is used; its containers get the new routers through the guarded Update (same version) ---
labels before: (none)
--- POST /api/stacks/adventurelog/update at 2026-09-13T19:35:05Z ---
HTTP 202
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
+ 0s phase safety-dump
+ 3s phase starting
+ 6s phase verifying
+ 21s phase done
end (2026-09-13T19:35:26Z, +21s): state=running updating=False phase=done err='' hold=''
labels after: Host(`travel.enkisfelhom.hu`) && (PathPrefix(`/media`) || PathPrefix(`/static`)
healthy: True
signup: 200
login: (302, '/')
location: 201
--- upload through the frontend proxy, WebKit-shaped boundary (what a browser sends) ---
proxy upload -> 500 {"id": null, "image": null}
--- proxy still refuses non-browser multipart; upload through the backend's own API inside the guest (the browser's session cookie, same backend) ---
backend upload -> 403
{"detail":"CSRF Failed: CSRF cookie not set."}
image url: None
--- control: /static and /admin now answer from Django ---
/static/admin/css/base.css -> 200
/admin/login/ -> 403
=== done 2026-09-13T19:35:34Z ===
=== R-483 repro, second half — 2026-09-13T19:36:33Z: an upload placed through the backend's own API (fresh CSRF cookie from the backend), then the photo fetched through the PUBLIC origin ===
login: (302, '/')
have sessionid: True
my place id found: True
backend upload -> 403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:36:37Z ===
=== R-483 repro, third try (the CSRF token is read with cut -f7, quoting-proof) — 2026-09-13T19:37:35Z ===
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:37:39Z ===
=== R-483 repro, fourth try (CSRF cookie from /admin/login/, which always sets one) — 2026-09-13T19:38:02Z ===
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:38:05Z ===
=== R-483 repro, fifth try (curl drops a Secure cookie over http; the token is cut from the raw Set-Cookie header) — 2026-09-13T19:38:30Z ===
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
image url: None
=== done 2026-09-13T19:38:33Z ===
=== R-483 repro, sixth try (token cut from the config endpoint's Set-Cookie header) — 2026-09-13T19:39:01Z ===
backend upload -> tokenlen=32 | http=400 | {"error":"content_type and object_id are required"}
image url: None
=== done 2026-09-13T19:39:05Z ===
=== R-483 repro, seventh try (v0.12.1's image API takes content_type=location + object_id) — 2026-09-13T19:39:21Z ===
backend upload -> http=201 | {"id":"c0b2b8cd-fe22-45db-b2a2-7c859d771941","image":"https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp","is_primary":false,"user":"51f9e8ee-f303-430e-b8ab-aae231e758c4","immich_id":null}
image url: https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp
GET /media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp through traefik (public origin) -> 200; bytes identical to the upload: False
the app lists the photo on the place: ['https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp']
=== done 2026-09-13T19:39:25Z ===
--- controls ---
real photo : (403, 'text/html; charset=utf-8', 'gunicorn')
bogus /media : (403, 'text/html; charset=utf-8', 'gunicorn') (a Django 404, not the frontend's HTML page)
frontend / : (200, 'text/html', '')
Traceback (most recent call last):
File "<stdin>", line 4, in <module>
ImportError: cannot import name 'S' from 'app' (/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/0ff9a77d-5916-4eea-91c4-eaea33e5d9a2/scratchpad/night/app.py)
=== R-483 second cut (port 80) applied to the operator's instance — 2026-09-13T19:46:23Z ===
sync: HTTP 200 {"ok":true,"data":{"ok":true,"updated":["adventurelog"],"message":"Sablonok frissítve — frissítve: adventurelog"},"messa
cache port 80: 1
--- POST /api/stacks/adventurelog/update at 2026-09-13T19:46:25Z ---
HTTP 202
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
+ 0s phase safety-dump
+ 3s phase starting
+ 6s phase verifying
+ 21s phase done
end (2026-09-13T19:46:47Z, +21s): state=running updating=False phase=done err='' hold=''
label after: 80
with session: 200 image/webp 154 bytes, magic: b'RIFF' b'WEBP'
anonymous: 403 (the app's privacy rule — control)
=== done 2026-09-13T19:46:54Z ===
@@ -0,0 +1,54 @@
=== FLOOR RAISE: 0.241.0 with min_agent 0.129.0 (R-479) ===
T0 POST at 2026-09-13T19:48:23Z
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
DB floor + declared: id="global-floor-input" name="min_controller_version" value="0.241.0" | name="min_agent" value="0.129.0" placeholder="MinAgent from |
demo-hp on 0.241.0 at 2026-09-13T19:48:40Z (+16s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up 7 seconds (healthy)
demo-felhom on 0.241.0 at 2026-09-13T19:48:41Z (+18s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up Less than a second (health: starting)
--- hub log since T0 (managed floor) ---
2026/09/13 21:48:24 [INFO] Global controller-version floor set to "0.241.0" (declared MinAgent "0.129.0")
2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
--- demo-hp: controller log (SetFloor / self-update / version) ---
2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: checking update state in /opt/docker/felhom-controller/data
2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: pending update found — target=0.241.0 previous=0.240.0
2026/09/13 19:48:33 updater.go:809: [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0)
2026/09/13 19:48:33 main.go:655: [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/13 19:48:33 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
2026/09/13 19:48:33 scheduler.go:102: [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/13 19:48:33 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="selfupdate-check" interval=6h0m0s totalJobs=18
2026/09/13 19:48:34 client.go:67: [DEBUG] [agentapi] agent version seen: 0.130.0 (was "")
2026/09/13 19:48:38 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.241.0, storagePaths=1
2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] SetFloor: floor "" → "0.241.0"
2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] maybeAutoUpdate: current 0.241.0 >= floor 0.241.0 — no action
2026/09/13 19:48:43 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running
--- demo-hp: bootstrap service journal ---
Sep 13 19:48:32 demo-hp systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
Sep 13 19:48:32 demo-hp systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
Sep 13 19:48:32 demo-hp systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:32 demo-hp systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-hp)
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305247]: 7d496e05c412d94587ef728dd28b513de5a834e550ca63d6ccae4bb1e7ccc780
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] controller started
Sep 13 19:48:32 demo-hp systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
--- demo-felhom: controller log (SetFloor / self-update / version) ---
2026/09/13 19:48:42 [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0)
2026/09/13 19:48:42 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/13 19:48:42 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/13 19:48:42 [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
2026/09/13 19:48:52 [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running
--- demo-felhom: bootstrap service journal ---
Sep 13 19:48:41 demo-felhom systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
Sep 13 19:48:41 demo-felhom systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
Sep 13 19:48:41 demo-felhom systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:41 demo-felhom systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-felhom)
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557928]: 52021cc5e06a35d9a7aa8832dba2372335abe431724ca6d2514b3558b8bb02e8
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] controller started
Sep 13 19:48:41 demo-felhom systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: []
@@ -0,0 +1,60 @@
=== R-479 live on demo-hp: a bind-data app (nextcloud, HDD_PATH on the registered drive) — 2026-09-13T19:48:56Z ===
controller: gitea.dooplex.hu/admin/felhom-controller:0.241.0
deploy: HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
healthy: True
classified binds seen by the controller (app has HDD data outside its unit): 2
--- Tier 2 OFF for the throwaway, then 'backup now' (unit only: definition + db-dump; the files stay on the drive) ---
HTTP/2 303
location: /stacks/nextcloud/backup?flash=A+2.+ment%C3%A9s+be%C3%A1ll%C3%ADt%C3%A1sa+elmentve.
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
idle after 121 s
/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud:
2026-09-13T19:49:12.5566936280 1403 db-dumps/nextcloud-mariadb.sql
2026-09-13T19:51:13.1432085620 4484 compose/.felhom.yml
2026-09-13T19:51:13.1432085620 4593 compose/docker-compose.yml
2026-09-13T19:51:13.1432085620 611 compose/app.yaml
2026-09-13T19:50:25.3716084070 170598400 volume-dumps/nextcloud_nextcloud_db_data.tar
2026-09-13T19:50:27.9356406190 6144 volume-dumps/nextcloud_nextcloud_redis_data.tar
2026-09-13T19:50:27.0496294880 781296640 volume-dumps/nextcloud_nextcloud_html.tar
2026-09-13T19:51:13.1432085620 1359 manifest.json
/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT
--- a failing update: box-local template edit to an image that exits at once, put back once the job is past the pin ---
16: image: nextcloud:34.0.1-apache
1
--- POST /api/stacks/nextcloud/update at 2026-09-13T19:51:19Z ---
HTTP 202
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
+ 0s phase checking
+ 9s phase starting
[fixture] template put back: 1
+ 13s phase verifying
+314s phase failed
end (2026-09-13T19:56:34Z, +314s): state=stopped updating=False phase=failed err='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat ' hold='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.'
--- the hold sentence ---
A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
names 'sajat meghajto': True | says settings+db only, no files ('a fajlokat nem'): True | control 'masodik meghajto': False
--- controller log ---
2026/09/13 19:51:19 update.go:391: [INFO] [stacks] update nextcloud: accepted — guarded update started
2026/09/13 19:51:19 update.go:820: [INFO] [stacks] update nextcloud: phase checking
2026/09/13 19:51:19 update_guard.go:223: [DEBUG] [backup] update precondition for nextcloud: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz)
2026/09/13 19:51:26 update.go:506: [INFO] [stacks] update nextcloud: precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T19:51:13Z (0s old, limit 24h0m0s)
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase safety-dump
2026/09/13 19:51:26 update.go:538: [INFO] [stacks] update nextcloud: safety dump done (1 file(s)) [/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps/pre-restore-20260913T195126Z-nextcloud-mariadb.sql]
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pinning
2026/09/13 19:51:26 pin.go:362: [INFO] [stacks] update nextcloud: pin advanced to the catalog's current definition (nextcloud=alpine:3.20, nextcloud-db=mariadb:11.6, nextcloud-redis=redis:7-alpine)
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pulling
2026/09/13 19:51:28 update.go:820: [INFO] [stacks] update nextcloud: phase starting
2026/09/13 19:51:30 update.go:820: [INFO] [stacks] update nextcloud: phase verifying
2026/09/13 19:56:32 update.go:625: [ERROR] [stacks] update nextcloud FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
2026/09/13 19:56:32 update_guard.go:530: [WARN] [backup] nextcloud is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-13T19:51:13Z; holds: "csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem")
--- teardown: restore is not needed for the proof; remove with data + backups ---
HTTP 200 {"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":["/mnt/felhom-drives/hdd_1/appdata/nextcloud (63M)"],"hdd_paths_preserved":[],"backup_paths_removed":["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud (909M)"]},"message":"Stack nextcloud removed"}
/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud: ABSENT
/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT
hold left: {'stack': 'nextcloud', 'at': '2026-09-13T19:56:32Z', 'reason': 'update_failed', 'copy_date': '2026-09-13T19:51:13Z', 'copy_tier': 1, 'copy_holds': 'csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem'}
HDD data left: ls: cannot access '/mnt/felhom-drives/hdd_1/appdata/nextcloud': No such file or directory
=== done 2026-09-13T19:56:44Z ===
@@ -0,0 +1,12 @@
unstaged-left: []
OK [felhom-controller]: 166 cited paths — exact 151, suffix 10, ambiguous 0, cross-repo 5, FAILED 0 (siblings searched: app-catalog-felhom.eu, felhom-agent, felhom.eu)
all controller gates OK
pre-push [felhom-controller]: gates OK - push proceeding.
HEAD=3e813307cc5688c0cce6ac1f44d284a963cddfd8 origin=3e813307cc5688c0cce6ac1f44d284a963cddfd8 porcelain=[]
build-rc=0
0.241.0: digest: sha256:a3ef97ba2dbd570922e0876030a497b3476065d5764d712a09c1873b8850eb32 size: 856
2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: []
live-rc=0
CHAIN-DONE
@@ -0,0 +1,24 @@
MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/backup/update_guard.go:
- if m.DataOutsideUnit(stackName) {
return updateTierOrderBindData
}
return updateTierOrder
+ _ = updateTierOrderBindData
return updateTierOrder
=== RUN TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
r479_tier_order_test.go:40: bind=true: order [2 1 3], want [2 3 1]
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
r479_tier_order_test.go:60: bind=true: chose tier 1, want 3
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Settings saved
--- FAIL: TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s
FAIL
exit=1
+2
View File
@@ -279,3 +279,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-461** | **`runbooks/target-selection.md` named a venue that does not exist and fenced a fixture that is gone.** Closed 2026-09-13, both halves checked against both boxes first: (a) no `/mnt/nvme-1tb` on demo-hp or demo-felhom; demo-hp's NVMe is `nvme0n1` at `/mnt/hdd_1` (demo-felhom's `/mnt/hdd_1` is `sdb`) — the runbook now names `/mnt/hdd_1` and says it is the same disk as the data drive; (b) `qm list` is empty on BOTH boxes — `drill-r50` (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | **CLOSED 2026-09-13 — DOCUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
| **R-452** | **Nothing enforced `catalog_since`, so the badge's one number could silently under-report.** Closed 2026-09-13 (catalog): `scripts/check-catalog-since.py`, the fifth gate in `catalog_gates.py` — a `--range A..B` gate in the engine-major shape: an app whose per-service `image:` lines differ across the range must carry a `catalog_since` on or after the moving commit's day and not in the future; comments, README and CHANGELOG mentions are not the fact. Hook-enforced; the shallow CI clone skips it out loud (the CI-shape half the row named stays as is, by the same reasoning engine-major uses). Five decoy cases; red-proof: dropping the date comparison lets the untouched-date fact through (`audits/v0240-2026-09-13/rp-R452.txt`). | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
| **R-465** | **`cfg.Paths.HDDPath` — empty on every box — still had six readers; were any inert?** AUDITED 2026-09-13 on demo-hp (registered drive `/mnt/felhom-drives/hdd_1`, `hdd_path` absent, no `FELHOM_PATHS_HDD_PATH`). Five of six fall back before the value matters: `report/builder.go:69` and `monitor/healthcheck.go:35` take `storagePaths[0]`; `web/server.go:740` (`primaryHDDPath`) takes the default storage path; `main.go:511` (metrics) takes the default storage path; `main.go:347` passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. **One is inert AND unreachable:** `api/router.go:981` (`systemInfo`, `GET /api/system/info`) reads the empty value with no fallback (`hdd_configured:false` forever) — and the endpoint itself is shadowed: the web layer's `ServeSystemAPI` claims `/api/system/*` and answers **404 „ismeretlen végpont"** for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as **R-490**. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | **CLOSED 2026-09-13 — AUDITED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
| **R-479** | **For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so.** Operator ruling 2026-09-13; closed in controller **v0.241.0** (`3e81330`): an app with classified binds walks second drive → off-site → own unit (`UpdateTierOrderFor`), and the hold sentence ends with what the chosen copy holds (`RestoreHold.CopyHolds`). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). `audits/v0241-2026-09-13/` | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
| **R-483** | **adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog).** Closed 2026-09-13: the k3s ingress routes `/media`, `/static`, `/admin`, `/accounts` to the backend service on port **80** — the nginx inside the backend image that serves Django's `X-Accel-Redirect` media; the catalog routed everything to the frontend. Two catalog cuts (`3172258` router, `ed62cfd` port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (`GET /media/…webp` with a session → 200 `image/webp`, RIFF/WEBP) and **confirmed by the operator in a browser at 21:49** (two photos render). Scripted multipart uploads through the frontend's `/api` proxy still 500 (upstream `RequestContentLengthMismatchError`); browser uploads work — not a template matter. `audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt` | **CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
+2 -3
View File
@@ -695,13 +695,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-483** | **[P3-LOW] `adventurelog` v0.12.1: a photo upload through the app's own API proxy fails 500 from every non-browser client — whether a household can upload photos in a browser is UNKNOWN and cannot be measured here.** MEASURED 2026-09-13 on demo-hp (throwaway `travelnight`): `POST /api/images` (multipart: `image`, `location`, `is_primary`) through the public origin returns `{"error":"Internal Server Error"}` from the **frontend** (the backend log shows no request at all); the frontend container logs `RequestContentLengthMismatchError: Request body length does not match content-length header` from its undici forwarder. Same result with an ASCII filename, an accented one, urllib, curl `-F`, and curl with `Transfer-Encoding: chunked`. Everything else on the API — sign-up, the frontend's login form, collections, locations, visits, notes, edit, delete — works. **What is NOT established:** whether the app's own browser page uploads succeed (a browser's FormData body may satisfy the proxy). Per `CLAUDE.md`, that is a manual click-through: open `https://<sub>.<domain>`, add a photo to a place. **Not ours to fix in the template** (a frontend proxy bug); a newer catalog pin is a version promotion, not tonight's. Evidence: `audits/nightly-2026-09-13-adventurelog/03e-photo-500.txt`. | **WAITING-ON-OPERATOR — needs a browser click-through; rank P3-LOW; owner: VIKTOR checks, CC re-files** |
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). **RULED 2026-09-13:** a second enrolled LXC on demo-hp, disk on the NVMe path (`/mnt/hdd_1`), sized like 9201, enrolled as a scratch CUSTOMER with its disposition recorded. **MEASURED BEFORE BUILDING — the ruling collides with the product's own model:** the hub's `hosts` table keys a host to exactly ONE customer (`host_id` PK, `customer_id NOT NULL`) and the box runs ONE agent with ONE `host_id`; a second customer on the same Proxmox host would need a second host identity, and the installer (`felhom-host-install.sh --customer-id … --vmid …`) run with another customer-id on demo-hp would rewrite the existing agent's identity — i.e. break the standing demo-hp enrolment. So "enrolled as a scratch customer" is not something the product can do on a box that already belongs to a customer. **What the product CAN do today, two options:** (a) **a second guest of the demo-hp customer** (the `guests` table is per host+vmid; the agent's provision writes a per-vmid bootstrap): supported by the data model, but both controllers report as customer demo-hp and the customer page shows one controller — the standing box's monitoring flips between the two unless the scratch guest's controller is told no hub (unenrolled scratch, disposition recorded on demo-hp's customer page); (b) **a separate scratch HOST** — a nested Proxmox VM on demo-hp (the ISO appliance route the earlier walks used, VM 323/325) enrolled as its own customer: fully enrolled, fully isolated, but an appliance to build (hours) and keep. Disk placement is the same under both: the installer has no rootfs-storage flag (only `--rootfs-grow`, `--archive-storage`), so the guest lands on `local-lvm` and is moved with `pct move-volume` to a `dir` storage created at `/mnt/hdd_1` — a post-provision step, reversible. 9201's shape for "sized like 9201": 7 cores, 25 898 MB, rootfs 32 G + mp0 70 G on local-lvm, unprivileged. **Recommendation: (a) with the hub left out** — it is the reversible one and it gives the rotation its restore target tomorrow; (b) if "enrolled" is the point. Not built tonight: §1 of the rules — a decision that changes what the product promises about hosts is not CC's. | **WAITING-ON-OPERATOR — the ruling cannot be executed as stated; two options below; rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-487** | **[P2-MEDIUM] A removed app whose backups were kept is listed on NEITHER backup page, so the restore that brings it back has no button — the customer's remove-by-mistake route exists only as an endpoint.** MEASURED 2026-09-13 on demo-hp (nightly rotation, `adventurelog` removed with backups kept, unit + mirror on disk): `GET /backups/apps` and `GET /backups/restore` contain the string `adventurelog` zero times; `POST /backup/restore stack_name=adventurelog snapshot_id=helyi` then restored it in 22 s with the data byte-identical. Cause: `buildAppBackupRows` walks `status.AppDataInfo` = `DiscoverAppData` over DEPLOYED stacks only. The off-site list had exactly this defect and was fixed by keying it on the store (R-237, v0.204.0); the local and Tier-2 lists were not. **Fix shape:** list every app with a recovery unit on a registered drive (`ListRestorePoints` over the primary dirs), marking removed ones „eltávolítva — visszaállítható"; the unit restore already reinstalls (R-253). Not a design reversal — the same rule R-237 set. Evidence: `audits/nightly-2026-09-13-adventurelog/05b-restore-tier1.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). | **READY — rank P3-LOW; owner: CC** |
| **R-490** | **[P3-LOW] The monitoring page's „Memória-eloszlás" card has never rendered: its `fetch('/api/system/info')` is answered 404 by the web layer's `ServeSystemAPI`, which claims all of `/api/system/*` and knows only the two memory routes.** MEASURED 2026-09-13 on demo-hp: `GET /api/system/info` → `{"error":"ismeretlen végpont","ok":false}`; `monitoring.html` shows the card (`display:none` by default) only when that fetch returns `used_mem_mb`, so it stays hidden on every box. The API router's `systemInfo` handler (`internal/api/router.go:981`) is unreachable, and it is also the one reader of the always-empty `cfg.Paths.HDDPath` with no fallback (R-465). **Fix shape:** let `ServeSystemAPI` fall through to the API router for unknown `/api/system/*` paths (or route `/api/system/info` explicitly), give `systemInfo` the same default-storage-path fallback the other readers have, and pin the card with a render test; then delete the global (R-465's deferred deletion). Next controller release. Evidence: `audits/nightly-2026-09-13-adventurelog/` (audit notes in the R-465 closure). | **READY — rank P3-LOW; owner: CC** |
| **R-491** | **[P2-MEDIUM] Removing an app leaves its update hold in the store, so a reinstall under the same name starts HELD.** MEASURED 2026-09-13 on demo-hp (v0.241.0, R-479 live check): nextcloud was held after a failed update, then removed with data and backups; `settings.json` still carried `restore_holds.nextcloud` (`reason: update_failed`, `copy_tier: 1`). `fillHoldReason` hides the sentence for a not-deployed app (R-480), but every start gate reads the store, so the next deploy of `nextcloud` would be refused as held with a sentence about a backup that no longer exists. Cleared by hand with `-clear-restore-hold`. **Fix shape:** `removeStack` clears an UPDATE hold (never an R-379 restore hold, which stays operator-cleared) — same place the prefs are forgotten; a test that deploys after a held removal. Next controller release. | **READY — rank P2-MEDIUM; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.