evening: rules file in place; R-483 CLOSED (catalog, operator-confirmed); R-479 CLOSED (controller v0.241.0, proven live); R-481 blocked on the one-host-one-customer model; R-491 filed
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
This commit is contained in:
@@ -34,6 +34,12 @@ on demo-hp — `documentation/audits/slice4-2026-09-13/`, design `09-update-arch
|
||||
- **Removal keeps the backups AND the Tier-2 record unless the customer asked for the backups to go**
|
||||
(v0.240.0, R-486/R-474). A removal that forgets the record makes an intact mirror unrestorable — that
|
||||
was live until v0.239.0. The first nightly rotation (adventurelog, 2026-09-13) found it.
|
||||
- **A bind-data app's route back is off-site before its own unit, and the hold names what the copy
|
||||
holds** (operator ruling 2026-09-13, R-479, controller v0.241.0). The unit of such an app holds the
|
||||
definition and the database dumps, not the files. Volume-data apps keep the v0.239.0 order.
|
||||
- **One host, one customer** (schema: `hosts.host_id` PK, one `customer_id`; one agent per box). A
|
||||
scratch guest "enrolled as its own customer" on a box that already belongs to a customer is not
|
||||
something the product can do (R-481, 2026-09-13); a second guest of the same customer is.
|
||||
- **A release reaches the fleet by floor between golden bakes when its MinAgent is declared with the
|
||||
floor** (ruling 2026-09-13, hub v0.112.0, R-472 CLOSED). An undeclared floor above the golden is still
|
||||
held, and the forms refuse it. v0.239.0 arrived on both demo boxes this way in ~15 s. Never hand-deploy
|
||||
|
||||
@@ -1,43 +1,45 @@
|
||||
# REPORT — the first "be a customer for the night" run: adventurelog on demo-hp (2026-09-13/14)
|
||||
# REPORT — the four evening items before the second night (2026-09-13)
|
||||
|
||||
*Overwritten each session. Every finding below has a register row. The controller report is
|
||||
*Overwritten each session. The night's own report (adventurelog rotation, v0.240.0) is preserved as
|
||||
`audits/nightly-2026-09-13-adventurelog/` and the register; the controller report is
|
||||
`felhom-controller/REPORT.md`, the catalog's `app-catalog-felhom.eu/REPORT.md`.*
|
||||
|
||||
> One app walked end to end as a household would; seven product defects found; a controller release
|
||||
> (v0.240.0) shipped the same night by the managed floor; two catalog fixes; two docs fixes. The
|
||||
> morning note is the top block of `STATUS.md`.
|
||||
## 1. The rules file
|
||||
|
||||
## Step 0 — the rules file
|
||||
`.claude/rules/unprompted-work.md` copied byte-identical into `felhom-controller`, `felhom.eu` and
|
||||
`app-catalog-felhom.eu` (documents-only pushes). `instructions_gate.py` refused the first push: a rule
|
||||
file needs `paths:` or `unconditional: true`; all four copies (root included) now start with the
|
||||
three-line `unconditional: true` frontmatter — sha256 `d1aa3af83ae3b220…` everywhere.
|
||||
|
||||
`.claude/rules/unprompted-work.md` is in none of the three repos and the brief's text block did not
|
||||
reach this session (searched the workspace, the transcript and memory). Not created — the operator's
|
||||
rules cannot be invented byte-identically. The brief's own fences were applied instead.
|
||||
## 2. R-481 — the scratch guest (ruled; NOT built, and why)
|
||||
|
||||
## What changed here
|
||||
The ruling ("a second enrolled LXC … enrolled as a scratch customer") collides with the model: the hub
|
||||
keys a host to one customer (`hosts.host_id` PK) and the box runs one agent with one host identity;
|
||||
re-running the installer with another customer-id would rewrite demo-hp's enrolment. Options and a
|
||||
recommendation are in the row and STATUS; 9201's shape and the disk placement (`pct move-volume`
|
||||
onto a `dir` storage at `/mnt/hdd_1`) are recorded for whichever is chosen.
|
||||
|
||||
- `documentation/runbooks/nightly-rotation.md` — new: the tick list (standing nine first, blocked on
|
||||
R-481; then the catalog); `actualbudget` recorded (no non-browser front door), `adventurelog` ticked.
|
||||
- `scripts/observations_gate.py` — R-471: every observations section is read (`scripts/CHANGELOG.md`).
|
||||
- `documentation/runbooks/target-selection.md` — R-461: `/mnt/hdd_1` is the NVMe; `drill-r50` exists
|
||||
on neither box (measured), fence kept with the fact beside it; R-93 carries the fact.
|
||||
- `09-update-architecture.md` §6.1 + §8.2 (R-452 lifted), `07-backup-architecture.md` §6, `CONTEXT.md` — v0.240.0 notes.
|
||||
- Catalog (`app-catalog-felhom.eu`): the `catalog-since` gate (R-452), see its REPORT.
|
||||
- Register: R-465 (audited: five readers fall back, one is inert and unreachable → R-490), R-452, R-473, R-474, R-466, R-471, R-453, R-461, R-477, R-478, R-480, R-482, R-484, R-485, R-486
|
||||
closed and compressed; R-481, R-482..R-489 opened (R-482 closed the same night). 212 → 206 open,
|
||||
177 → 189 closed (byte sizes: see the commit).
|
||||
- Evidence: `documentation/audits/nightly-2026-09-13-adventurelog/` (01–08), `audits/v0240-2026-09-13/`
|
||||
(release log, floor, validation, red-proofs incl. R-471).
|
||||
## 3. R-483 — photos (re-ranked P2, catalog)
|
||||
|
||||
Diffed against `homelab-manifests` `adventurelog-system/adventurelog.yaml`: the k3s ingress sends
|
||||
`/media`, `/static`, `/admin`, `/accounts` to the backend **service port 80**; the catalog sent
|
||||
everything to the frontend. Two catalog cuts (`3172258`, `ed62cfd`): a backend traefik router for the
|
||||
four prefixes, then port 80 instead of 8000 — the backend image runs nginx in front of gunicorn and
|
||||
Django serves a photo by `X-Accel-Redirect`, so port 8000 returned 200 with an empty body. Applied to
|
||||
the operator's own instance through the guarded Update; proven without a browser
|
||||
(`audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt`): `GET /media/images/<id>.webp` with a
|
||||
session → 200 `image/webp`, RIFF/WEBP magic; anonymous 403 (the app's rule). **Confirmed by the operator in a browser at 21:49 — closed.** The frontend's 500 on scripted multipart uploads is upstream and
|
||||
recorded in the row; browser uploads work.
|
||||
|
||||
## 4. R-479 — tier order for bind-data apps (ruled; controller v0.241.0)
|
||||
|
||||
Ruling 9 in `09-update-architecture.md` §3. Built, tested, red-proofed; delivered by the floor and
|
||||
live-checked on a nextcloud throwaway — see `felhom-controller/REPORT.md` and
|
||||
`audits/v0241-2026-09-13/`.
|
||||
|
||||
## Observations
|
||||
|
||||
1. **The brief's rules block was not delivered; the file is absent in all repos.** NOT-A-FINDING: an input gap for the operator, recorded in STATUS "needs you", not a product defect.
|
||||
2. **No scratch guest on demo-hp; the standing nine cannot be throwaways on 9201.** FILED: R-481
|
||||
3. **adventurelog ran Django DEBUG=True on the public origin.** FILED: R-482
|
||||
4. **Photo upload fails 500 from every non-browser client (frontend proxy body-length bug).** FILED: R-483
|
||||
5. **PostGIS not recognised as a database.** FILED: R-484
|
||||
6. **The backup card reads dead paths.** FILED: R-485
|
||||
7. **Removal with backups kept forgot the Tier-2 record; the restore was refused.** FILED: R-486
|
||||
8. **A removed app is on neither backup page.** FILED: R-487
|
||||
9. **The backup test package takes 5½ minutes of real waits.** FILED: R-488
|
||||
10. **`volumes_removed: null` over removed volumes, five times.** FILED: R-489
|
||||
11. **The monitoring page's memory-distribution card never renders — `/api/system/info` is shadowed by `ServeSystemAPI` (found by the R-465 audit).** FILED: R-490
|
||||
0. **A removed app keeps its update hold; a reinstall would start held.** FILED: R-491
|
||||
|
||||
1. **The controller's evening release rides the same session as the night's — two releases in one calendar day, one per session.** NOT-A-FINDING: the rule is per session; both are recorded with their MinAgent lines.
|
||||
2. **`homelab-manifests` is not in the workspace; its manifests were read from Gitea's API.** NOT-A-FINDING: the workspace CLAUDE.md says non-felhom repos are unrelated; a read-only fetch was enough.
|
||||
|
||||
@@ -1,5 +1,34 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-13 (evening, before the night) — your four items.**
|
||||
|
||||
**Decisions I took.** (1) The rules file now sits in the workspace root and all three repos,
|
||||
byte-identical; the gate that checks instruction files required three lines at its top saying it
|
||||
loads in every session on purpose, so all four copies carry them. (2) The photo fix was applied to
|
||||
your own test instance on the HP through the Update button, not by hand, and I added one test
|
||||
account with one photo there.
|
||||
|
||||
**What I did.** The rules file: done. The photo problem: the app's backend serves photos itself,
|
||||
through a small web server inside its own image; our template sent every request to the front page
|
||||
server, which has no photos. Fixed in two cuts — the second one because the first spoke to the wrong
|
||||
port and got empty pictures. Proven without a browser: a photo now comes back as a real image
|
||||
through the public address, and you confirmed it in your browser at 21:49. The tier-order
|
||||
ruling: built as controller 0.241.0 — an app whose data lives in mounted folders now prefers the
|
||||
remote copy over its own-drive copy, and the hold message ends with what the chosen copy holds.
|
||||
Live on both machines (0.241.0, delivered by the floor in 16 and 18 seconds) and proven on the HP
|
||||
with a throwaway Nextcloud.
|
||||
|
||||
**What broke on the way.** Nothing standing. One new finding: after an app is removed, its
|
||||
"held after a failed update" mark stays behind, so a reinstall would start blocked. Cleared by hand,
|
||||
filed (R-491) for the next release. Register: 206 open / 191 closed.
|
||||
|
||||
**Needs you.** The scratch guest: the ruling as written cannot be built — the hub ties one box to
|
||||
one customer, and a second enrolled customer on the HP would replace the HP's own enrolment. Two
|
||||
options: a second guest under the HP's own customer (reversible, tonight-ready, but its reports
|
||||
would clash with the HP's page unless it stays unenrolled), or a nested appliance VM enrolled as its
|
||||
own box (fully enrolled, hours to build). My pick: the first, unenrolled. **If you do nothing:**
|
||||
tonight's rotation keeps restoring in place and skips the standing nine.
|
||||
|
||||
**Updated 2026-09-14 (the night of 13→14 — "be a customer for the night", first run). Written in
|
||||
the order the rules ask for.**
|
||||
|
||||
|
||||
@@ -277,7 +277,7 @@ what the unit holds** — for an app whose data is a bind mount that is the defi
|
||||
(R-479). **Removal and the tiers (controller v0.240.0):** removing an app with its backups KEPT keeps
|
||||
the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás" still works afterwards
|
||||
(R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and
|
||||
never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open).
|
||||
never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open). **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.**
|
||||
|
||||
### 6.1 The four tiers, as configured on the live fleet
|
||||
|
||||
|
||||
@@ -178,6 +178,14 @@ These are rulings, not proposals. Anything specced against a different assumptio
|
||||
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
|
||||
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
|
||||
|
||||
9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy
|
||||
holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit
|
||||
that holds the definition and the database dumps and NOT the files (measured: gokapi restored from
|
||||
„helyi" came back with settings and no data). For such an app the update walks second drive →
|
||||
off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive →
|
||||
own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a
|
||||
customer is never sent to a copy that cannot bring the data back without being told so.
|
||||
|
||||
## 4. The vocabulary ruling — "rollback" is struck
|
||||
|
||||
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
=== R-483 repro on demo-hp — 2026-09-13T19:35:02Z ===
|
||||
sync: HTTP 200 {"ok":true,"data":{"ok":true,"message":"Sablonok naprakészek — nincs változás"},"message":"Sablonok naprakészek — nincs változás"}
|
||||
cache has the backend router: 1
|
||||
--- the operator's own throwaway instance (subdomain travel, 253 media files) is used; its containers get the new routers through the guarded Update (same version) ---
|
||||
labels before: (none)
|
||||
--- POST /api/stacks/adventurelog/update at 2026-09-13T19:35:05Z ---
|
||||
HTTP 202
|
||||
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
|
||||
+ 0s phase safety-dump
|
||||
+ 3s phase starting
|
||||
+ 6s phase verifying
|
||||
+ 21s phase done
|
||||
end (2026-09-13T19:35:26Z, +21s): state=running updating=False phase=done err='' hold=''
|
||||
labels after: Host(`travel.enkisfelhom.hu`) && (PathPrefix(`/media`) || PathPrefix(`/static`)
|
||||
healthy: True
|
||||
signup: 200
|
||||
login: (302, '/')
|
||||
location: 201
|
||||
--- upload through the frontend proxy, WebKit-shaped boundary (what a browser sends) ---
|
||||
proxy upload -> 500 {"id": null, "image": null}
|
||||
--- proxy still refuses non-browser multipart; upload through the backend's own API inside the guest (the browser's session cookie, same backend) ---
|
||||
backend upload -> 403
|
||||
{"detail":"CSRF Failed: CSRF cookie not set."}
|
||||
image url: None
|
||||
--- control: /static and /admin now answer from Django ---
|
||||
/static/admin/css/base.css -> 200
|
||||
/admin/login/ -> 403
|
||||
=== done 2026-09-13T19:35:34Z ===
|
||||
=== R-483 repro, second half — 2026-09-13T19:36:33Z: an upload placed through the backend's own API (fresh CSRF cookie from the backend), then the photo fetched through the PUBLIC origin ===
|
||||
login: (302, '/')
|
||||
have sessionid: True
|
||||
my place id found: True
|
||||
backend upload -> 403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
|
||||
image url: None
|
||||
=== done 2026-09-13T19:36:37Z ===
|
||||
=== R-483 repro, third try (the CSRF token is read with cut -f7, quoting-proof) — 2026-09-13T19:37:35Z ===
|
||||
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
|
||||
image url: None
|
||||
=== done 2026-09-13T19:37:39Z ===
|
||||
=== R-483 repro, fourth try (CSRF cookie from /admin/login/, which always sets one) — 2026-09-13T19:38:02Z ===
|
||||
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
|
||||
image url: None
|
||||
=== done 2026-09-13T19:38:05Z ===
|
||||
=== R-483 repro, fifth try (curl drops a Secure cookie over http; the token is cut from the raw Set-Cookie header) — 2026-09-13T19:38:30Z ===
|
||||
backend upload -> tokenlen=0 | http=403 | {"detail":"CSRF Failed: CSRF cookie has incorrect length."}
|
||||
image url: None
|
||||
=== done 2026-09-13T19:38:33Z ===
|
||||
=== R-483 repro, sixth try (token cut from the config endpoint's Set-Cookie header) — 2026-09-13T19:39:01Z ===
|
||||
backend upload -> tokenlen=32 | http=400 | {"error":"content_type and object_id are required"}
|
||||
image url: None
|
||||
=== done 2026-09-13T19:39:05Z ===
|
||||
=== R-483 repro, seventh try (v0.12.1's image API takes content_type=location + object_id) — 2026-09-13T19:39:21Z ===
|
||||
backend upload -> http=201 | {"id":"c0b2b8cd-fe22-45db-b2a2-7c859d771941","image":"https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp","is_primary":false,"user":"51f9e8ee-f303-430e-b8ab-aae231e758c4","immich_id":null}
|
||||
image url: https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp
|
||||
GET /media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp through traefik (public origin) -> 200; bytes identical to the upload: False
|
||||
the app lists the photo on the place: ['https://travel.enkisfelhom.hu/media/images/973b9ada-bdab-48f1-85e6-2802a13eb739.webp']
|
||||
=== done 2026-09-13T19:39:25Z ===
|
||||
--- controls ---
|
||||
real photo : (403, 'text/html; charset=utf-8', 'gunicorn')
|
||||
bogus /media : (403, 'text/html; charset=utf-8', 'gunicorn') (a Django 404, not the frontend's HTML page)
|
||||
frontend / : (200, 'text/html', '')
|
||||
Traceback (most recent call last):
|
||||
File "<stdin>", line 4, in <module>
|
||||
ImportError: cannot import name 'S' from 'app' (/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/0ff9a77d-5916-4eea-91c4-eaea33e5d9a2/scratchpad/night/app.py)
|
||||
=== R-483 second cut (port 80) applied to the operator's instance — 2026-09-13T19:46:23Z ===
|
||||
sync: HTTP 200 {"ok":true,"data":{"ok":true,"updated":["adventurelog"],"message":"Sablonok frissítve — frissítve: adventurelog"},"messa
|
||||
cache port 80: 1
|
||||
--- POST /api/stacks/adventurelog/update at 2026-09-13T19:46:25Z ---
|
||||
HTTP 202
|
||||
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
|
||||
+ 0s phase safety-dump
|
||||
+ 3s phase starting
|
||||
+ 6s phase verifying
|
||||
+ 21s phase done
|
||||
end (2026-09-13T19:46:47Z, +21s): state=running updating=False phase=done err='' hold=''
|
||||
label after: 80
|
||||
with session: 200 image/webp 154 bytes, magic: b'RIFF' b'WEBP'
|
||||
anonymous: 403 (the app's privacy rule — control)
|
||||
=== done 2026-09-13T19:46:54Z ===
|
||||
@@ -0,0 +1,54 @@
|
||||
=== FLOOR RAISE: 0.241.0 with min_agent 0.129.0 (R-479) ===
|
||||
T0 POST at 2026-09-13T19:48:23Z
|
||||
HTTP/1.1 303 See Other
|
||||
Location: /configuration?flash=floor_set
|
||||
DB floor + declared: id="global-floor-input" name="min_controller_version" value="0.241.0" | name="min_agent" value="0.129.0" placeholder="MinAgent from |
|
||||
demo-hp on 0.241.0 at 2026-09-13T19:48:40Z (+16s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up 7 seconds (healthy)
|
||||
demo-felhom on 0.241.0 at 2026-09-13T19:48:41Z (+18s): gitea.dooplex.hu/admin/felhom-controller:0.241.0 Up Less than a second (health: starting)
|
||||
--- hub log since T0 (managed floor) ---
|
||||
2026/09/13 21:48:24 [INFO] Global controller-version floor set to "0.241.0" (declared MinAgent "0.129.0")
|
||||
2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
|
||||
2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
|
||||
|
||||
--- demo-hp: controller log (SetFloor / self-update / version) ---
|
||||
2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: checking update state in /opt/docker/felhom-controller/data
|
||||
2026/09/13 19:48:33 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: pending update found — target=0.241.0 previous=0.240.0
|
||||
2026/09/13 19:48:33 updater.go:809: [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0)
|
||||
2026/09/13 19:48:33 main.go:655: [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
|
||||
2026/09/13 19:48:33 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
|
||||
2026/09/13 19:48:33 scheduler.go:102: [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
|
||||
2026/09/13 19:48:33 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="selfupdate-check" interval=6h0m0s totalJobs=18
|
||||
2026/09/13 19:48:34 client.go:67: [DEBUG] [agentapi] agent version seen: 0.130.0 (was "")
|
||||
2026/09/13 19:48:38 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.241.0, storagePaths=1
|
||||
2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] SetFloor: floor "" → "0.241.0"
|
||||
2026/09/13 19:48:39 updater.go:96: [DEBUG] [selfupdate] maybeAutoUpdate: current 0.241.0 >= floor 0.241.0 — no action
|
||||
2026/09/13 19:48:43 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running
|
||||
|
||||
--- demo-hp: bootstrap service journal ---
|
||||
Sep 13 19:48:32 demo-hp systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
|
||||
Sep 13 19:48:32 demo-hp systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
|
||||
Sep 13 19:48:32 demo-hp systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
|
||||
Sep 13 19:48:32 demo-hp systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
|
||||
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-hp)
|
||||
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305247]: 7d496e05c412d94587ef728dd28b513de5a834e550ca63d6ccae4bb1e7ccc780
|
||||
Sep 13 19:48:32 demo-hp felhom-controller-bootstrap.sh[2305171]: [ctrl-bootstrap] controller started
|
||||
Sep 13 19:48:32 demo-hp systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
|
||||
|
||||
--- demo-felhom: controller log (SetFloor / self-update / version) ---
|
||||
2026/09/13 19:48:42 [INFO] [selfupdate] Post-update startup: update successful (0.240.0 → 0.241.0)
|
||||
2026/09/13 19:48:42 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
|
||||
2026/09/13 19:48:42 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
|
||||
2026/09/13 19:48:42 [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
|
||||
2026/09/13 19:48:52 [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.241.0 (we are 0.241.0), no managed update running
|
||||
|
||||
--- demo-felhom: bootstrap service journal ---
|
||||
Sep 13 19:48:41 demo-felhom systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
|
||||
Sep 13 19:48:41 demo-felhom systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
|
||||
Sep 13 19:48:41 demo-felhom systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
|
||||
Sep 13 19:48:41 demo-felhom systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
|
||||
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.241.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-felhom)
|
||||
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557928]: 52021cc5e06a35d9a7aa8832dba2372335abe431724ca6d2514b3558b8bb02e8
|
||||
Sep 13 19:48:41 demo-felhom felhom-controller-bootstrap.sh[3557879]: [ctrl-bootstrap] controller started
|
||||
Sep 13 19:48:41 demo-felhom systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
|
||||
|
||||
SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: []
|
||||
@@ -0,0 +1,60 @@
|
||||
=== R-479 live on demo-hp: a bind-data app (nextcloud, HDD_PATH on the registered drive) — 2026-09-13T19:48:56Z ===
|
||||
controller: gitea.dooplex.hu/admin/felhom-controller:0.241.0
|
||||
deploy: HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
healthy: True
|
||||
classified binds seen by the controller (app has HDD data outside its unit): 2
|
||||
--- Tier 2 OFF for the throwaway, then 'backup now' (unit only: definition + db-dump; the files stay on the drive) ---
|
||||
HTTP/2 303
|
||||
location: /stacks/nextcloud/backup?flash=A+2.+ment%C3%A9s+be%C3%A1ll%C3%ADt%C3%A1sa+elmentve.
|
||||
|
||||
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
|
||||
idle after 121 s
|
||||
/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud:
|
||||
2026-09-13T19:49:12.5566936280 1403 db-dumps/nextcloud-mariadb.sql
|
||||
2026-09-13T19:51:13.1432085620 4484 compose/.felhom.yml
|
||||
2026-09-13T19:51:13.1432085620 4593 compose/docker-compose.yml
|
||||
2026-09-13T19:51:13.1432085620 611 compose/app.yaml
|
||||
2026-09-13T19:50:25.3716084070 170598400 volume-dumps/nextcloud_nextcloud_db_data.tar
|
||||
2026-09-13T19:50:27.9356406190 6144 volume-dumps/nextcloud_nextcloud_redis_data.tar
|
||||
2026-09-13T19:50:27.0496294880 781296640 volume-dumps/nextcloud_nextcloud_html.tar
|
||||
2026-09-13T19:51:13.1432085620 1359 manifest.json
|
||||
/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT
|
||||
--- a failing update: box-local template edit to an image that exits at once, put back once the job is past the pin ---
|
||||
16: image: nextcloud:34.0.1-apache
|
||||
1
|
||||
--- POST /api/stacks/nextcloud/update at 2026-09-13T19:51:19Z ---
|
||||
HTTP 202
|
||||
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
|
||||
+ 0s phase checking
|
||||
+ 9s phase starting
|
||||
[fixture] template put back: 1
|
||||
+ 13s phase verifying
|
||||
+314s phase failed
|
||||
end (2026-09-13T19:56:34Z, +314s): state=stopped updating=False phase=failed err='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat ' hold='A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.'
|
||||
--- the hold sentence ---
|
||||
A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
|
||||
names 'sajat meghajto': True | says settings+db only, no files ('a fajlokat nem'): True | control 'masodik meghajto': False
|
||||
--- controller log ---
|
||||
2026/09/13 19:51:19 update.go:391: [INFO] [stacks] update nextcloud: accepted — guarded update started
|
||||
2026/09/13 19:51:19 update.go:820: [INFO] [stacks] update nextcloud: phase checking
|
||||
2026/09/13 19:51:19 update_guard.go:223: [DEBUG] [backup] update precondition for nextcloud: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz)
|
||||
2026/09/13 19:51:26 update.go:506: [INFO] [stacks] update nextcloud: precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T19:51:13Z (0s old, limit 24h0m0s)
|
||||
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase safety-dump
|
||||
2026/09/13 19:51:26 update.go:538: [INFO] [stacks] update nextcloud: safety dump done (1 file(s)) [/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps/pre-restore-20260913T195126Z-nextcloud-mariadb.sql]
|
||||
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pinning
|
||||
2026/09/13 19:51:26 pin.go:362: [INFO] [stacks] update nextcloud: pin advanced to the catalog's current definition (nextcloud=alpine:3.20, nextcloud-db=mariadb:11.6, nextcloud-redis=redis:7-alpine)
|
||||
2026/09/13 19:51:26 update.go:820: [INFO] [stacks] update nextcloud: phase pulling
|
||||
2026/09/13 19:51:28 update.go:820: [INFO] [stacks] update nextcloud: phase starting
|
||||
2026/09/13 19:51:30 update.go:820: [INFO] [stacks] update nextcloud: phase verifying
|
||||
2026/09/13 19:56:32 update.go:625: [ERROR] [stacks] update nextcloud FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
|
||||
2026/09/13 19:56:32 update_guard.go:530: [WARN] [backup] nextcloud is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-13T19:51:13Z; holds: "csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem")
|
||||
--- teardown: restore is not needed for the proof; remove with data + backups ---
|
||||
|
||||
HTTP 200 {"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":["/mnt/felhom-drives/hdd_1/appdata/nextcloud (63M)"],"hdd_paths_preserved":[],"backup_paths_removed":["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud (909M)"]},"message":"Stack nextcloud removed"}
|
||||
/mnt/sys_drive/felhom-data/backups/primary/nextcloud: ABSENT
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud: ABSENT
|
||||
/mnt/felhom-drives/hdd_1/backups/secondary/nextcloud: ABSENT
|
||||
hold left: {'stack': 'nextcloud', 'at': '2026-09-13T19:56:32Z', 'reason': 'update_failed', 'copy_date': '2026-09-13T19:51:13Z', 'copy_tier': 1, 'copy_holds': 'csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem'}
|
||||
HDD data left: ls: cannot access '/mnt/felhom-drives/hdd_1/appdata/nextcloud': No such file or directory
|
||||
=== done 2026-09-13T19:56:44Z ===
|
||||
@@ -0,0 +1,12 @@
|
||||
unstaged-left: []
|
||||
OK [felhom-controller]: 166 cited paths — exact 151, suffix 10, ambiguous 0, cross-repo 5, FAILED 0 (siblings searched: app-catalog-felhom.eu, felhom-agent, felhom.eu)
|
||||
all controller gates OK
|
||||
pre-push [felhom-controller]: gates OK - push proceeding.
|
||||
HEAD=3e813307cc5688c0cce6ac1f44d284a963cddfd8 origin=3e813307cc5688c0cce6ac1f44d284a963cddfd8 porcelain=[]
|
||||
build-rc=0
|
||||
0.241.0: digest: sha256:a3ef97ba2dbd570922e0876030a497b3476065d5764d712a09c1873b8850eb32 size: 856
|
||||
2026/09/13 21:48:26 [INFO] managed floor SERVED for demo-felhom: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
|
||||
2026/09/13 21:48:27 [INFO] managed floor SERVED for demo-hp: floor 0.241.0, agent requirement "0.129.0" from declared (golden 0.236.0)
|
||||
SUMMARY T0=2026-09-13T19:48:23Z {'demo-hp': ('2026-09-13T19:48:40Z', 16), 'demo-felhom': ('2026-09-13T19:48:41Z', 18)} not done: []
|
||||
live-rc=0
|
||||
CHAIN-DONE
|
||||
@@ -0,0 +1,24 @@
|
||||
MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/backup/update_guard.go:
|
||||
- if m.DataOutsideUnit(stackName) {
|
||||
return updateTierOrderBindData
|
||||
}
|
||||
return updateTierOrder
|
||||
+ _ = updateTierOrderBindData
|
||||
return updateTierOrder
|
||||
|
||||
=== RUN TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit
|
||||
[INFO] [settings] No settings.json found, using defaults
|
||||
[INFO] [settings] Settings saved
|
||||
r479_tier_order_test.go:40: bind=true: order [2 1 3], want [2 3 1]
|
||||
[INFO] [settings] No settings.json found, using defaults
|
||||
[INFO] [settings] Settings saved
|
||||
[INFO] [settings] No settings.json found, using defaults
|
||||
[INFO] [settings] Settings saved
|
||||
r479_tier_order_test.go:60: bind=true: chose tier 1, want 3
|
||||
[INFO] [settings] No settings.json found, using defaults
|
||||
[INFO] [settings] Settings saved
|
||||
--- FAIL: TestR479_BindDataAppWalksSecondDriveOffsiteThenOwnUnit (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s
|
||||
FAIL
|
||||
exit=1
|
||||
@@ -279,3 +279,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-461** | **`runbooks/target-selection.md` named a venue that does not exist and fenced a fixture that is gone.** Closed 2026-09-13, both halves checked against both boxes first: (a) no `/mnt/nvme-1tb` on demo-hp or demo-felhom; demo-hp's NVMe is `nvme0n1` at `/mnt/hdd_1` (demo-felhom's `/mnt/hdd_1` is `sdb`) — the runbook now names `/mnt/hdd_1` and says it is the same disk as the data drive; (b) `qm list` is empty on BOTH boxes — `drill-r50` (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | **CLOSED 2026-09-13 — DOCUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-452** | **Nothing enforced `catalog_since`, so the badge's one number could silently under-report.** Closed 2026-09-13 (catalog): `scripts/check-catalog-since.py`, the fifth gate in `catalog_gates.py` — a `--range A..B` gate in the engine-major shape: an app whose per-service `image:` lines differ across the range must carry a `catalog_since` on or after the moving commit's day and not in the future; comments, README and CHANGELOG mentions are not the fact. Hook-enforced; the shallow CI clone skips it out loud (the CI-shape half the row named stays as is, by the same reasoning engine-major uses). Five decoy cases; red-proof: dropping the date comparison lets the untouched-date fact through (`audits/v0240-2026-09-13/rp-R452.txt`). | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-465** | **`cfg.Paths.HDDPath` — empty on every box — still had six readers; were any inert?** AUDITED 2026-09-13 on demo-hp (registered drive `/mnt/felhom-drives/hdd_1`, `hdd_path` absent, no `FELHOM_PATHS_HDD_PATH`). Five of six fall back before the value matters: `report/builder.go:69` and `monitor/healthcheck.go:35` take `storagePaths[0]`; `web/server.go:740` (`primaryHDDPath`) takes the default storage path; `main.go:511` (metrics) takes the default storage path; `main.go:347` passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. **One is inert AND unreachable:** `api/router.go:981` (`systemInfo`, `GET /api/system/info`) reads the empty value with no fallback (`hdd_configured:false` forever) — and the endpoint itself is shadowed: the web layer's `ServeSystemAPI` claims `/api/system/*` and answers **404 „ismeretlen végpont"** for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as **R-490**. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | **CLOSED 2026-09-13 — AUDITED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-479** | **For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so.** Operator ruling 2026-09-13; closed in controller **v0.241.0** (`3e81330`): an app with classified binds walks second drive → off-site → own unit (`UpdateTierOrderFor`), and the hold sentence ends with what the chosen copy holds (`RestoreHold.CopyHolds`). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). `audits/v0241-2026-09-13/` | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-483** | **adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog).** Closed 2026-09-13: the k3s ingress routes `/media`, `/static`, `/admin`, `/accounts` to the backend service on port **80** — the nginx inside the backend image that serves Django's `X-Accel-Redirect` media; the catalog routed everything to the frontend. Two catalog cuts (`3172258` router, `ed62cfd` port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (`GET /media/…webp` with a session → 200 `image/webp`, RIFF/WEBP) and **confirmed by the operator in a browser at 21:49** (two photos render). Scripted multipart uploads through the frontend's `/api` proxy still 500 (upstream `RequestContentLengthMismatchError`); browser uploads work — not a template matter. `audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt` | **CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -695,13 +695,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
|
||||
|
||||
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-483** | **[P3-LOW] `adventurelog` v0.12.1: a photo upload through the app's own API proxy fails 500 from every non-browser client — whether a household can upload photos in a browser is UNKNOWN and cannot be measured here.** MEASURED 2026-09-13 on demo-hp (throwaway `travelnight`): `POST /api/images` (multipart: `image`, `location`, `is_primary`) through the public origin returns `{"error":"Internal Server Error"}` from the **frontend** (the backend log shows no request at all); the frontend container logs `RequestContentLengthMismatchError: Request body length does not match content-length header` from its undici forwarder. Same result with an ASCII filename, an accented one, urllib, curl `-F`, and curl with `Transfer-Encoding: chunked`. Everything else on the API — sign-up, the frontend's login form, collections, locations, visits, notes, edit, delete — works. **What is NOT established:** whether the app's own browser page uploads succeed (a browser's FormData body may satisfy the proxy). Per `CLAUDE.md`, that is a manual click-through: open `https://<sub>.<domain>`, add a photo to a place. **Not ours to fix in the template** (a frontend proxy bug); a newer catalog pin is a version promotion, not tonight's. Evidence: `audits/nightly-2026-09-13-adventurelog/03e-photo-500.txt`. | **WAITING-ON-OPERATOR — needs a browser click-through; rank P3-LOW; owner: VIKTOR checks, CC re-files** |
|
||||
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). **RULED 2026-09-13:** a second enrolled LXC on demo-hp, disk on the NVMe path (`/mnt/hdd_1`), sized like 9201, enrolled as a scratch CUSTOMER with its disposition recorded. **MEASURED BEFORE BUILDING — the ruling collides with the product's own model:** the hub's `hosts` table keys a host to exactly ONE customer (`host_id` PK, `customer_id NOT NULL`) and the box runs ONE agent with ONE `host_id`; a second customer on the same Proxmox host would need a second host identity, and the installer (`felhom-host-install.sh --customer-id … --vmid …`) run with another customer-id on demo-hp would rewrite the existing agent's identity — i.e. break the standing demo-hp enrolment. So "enrolled as a scratch customer" is not something the product can do on a box that already belongs to a customer. **What the product CAN do today, two options:** (a) **a second guest of the demo-hp customer** (the `guests` table is per host+vmid; the agent's provision writes a per-vmid bootstrap): supported by the data model, but both controllers report as customer demo-hp and the customer page shows one controller — the standing box's monitoring flips between the two unless the scratch guest's controller is told no hub (unenrolled scratch, disposition recorded on demo-hp's customer page); (b) **a separate scratch HOST** — a nested Proxmox VM on demo-hp (the ISO appliance route the earlier walks used, VM 323/325) enrolled as its own customer: fully enrolled, fully isolated, but an appliance to build (hours) and keep. Disk placement is the same under both: the installer has no rootfs-storage flag (only `--rootfs-grow`, `--archive-storage`), so the guest lands on `local-lvm` and is moved with `pct move-volume` to a `dir` storage created at `/mnt/hdd_1` — a post-provision step, reversible. 9201's shape for "sized like 9201": 7 cores, 25 898 MB, rootfs 32 G + mp0 70 G on local-lvm, unprivileged. **Recommendation: (a) with the hub left out** — it is the reversible one and it gives the rotation its restore target tomorrow; (b) if "enrolled" is the point. Not built tonight: §1 of the rules — a decision that changes what the product promises about hosts is not CC's. | **WAITING-ON-OPERATOR — the ruling cannot be executed as stated; two options below; rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-487** | **[P2-MEDIUM] A removed app whose backups were kept is listed on NEITHER backup page, so the restore that brings it back has no button — the customer's remove-by-mistake route exists only as an endpoint.** MEASURED 2026-09-13 on demo-hp (nightly rotation, `adventurelog` removed with backups kept, unit + mirror on disk): `GET /backups/apps` and `GET /backups/restore` contain the string `adventurelog` zero times; `POST /backup/restore stack_name=adventurelog snapshot_id=helyi` then restored it in 22 s with the data byte-identical. Cause: `buildAppBackupRows` walks `status.AppDataInfo` = `DiscoverAppData` over DEPLOYED stacks only. The off-site list had exactly this defect and was fixed by keying it on the store (R-237, v0.204.0); the local and Tier-2 lists were not. **Fix shape:** list every app with a recovery unit on a registered drive (`ListRestorePoints` over the primary dirs), marking removed ones „eltávolítva — visszaállítható"; the unit restore already reinstalls (R-253). Not a design reversal — the same rule R-237 set. Evidence: `audits/nightly-2026-09-13-adventurelog/05b-restore-tier1.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-490** | **[P3-LOW] The monitoring page's „Memória-eloszlás" card has never rendered: its `fetch('/api/system/info')` is answered 404 by the web layer's `ServeSystemAPI`, which claims all of `/api/system/*` and knows only the two memory routes.** MEASURED 2026-09-13 on demo-hp: `GET /api/system/info` → `{"error":"ismeretlen végpont","ok":false}`; `monitoring.html` shows the card (`display:none` by default) only when that fetch returns `used_mem_mb`, so it stays hidden on every box. The API router's `systemInfo` handler (`internal/api/router.go:981`) is unreachable, and it is also the one reader of the always-empty `cfg.Paths.HDDPath` with no fallback (R-465). **Fix shape:** let `ServeSystemAPI` fall through to the API router for unknown `/api/system/*` paths (or route `/api/system/info` explicitly), give `systemInfo` the same default-storage-path fallback the other readers have, and pin the card with a render test; then delete the global (R-465's deferred deletion). Next controller release. Evidence: `audits/nightly-2026-09-13-adventurelog/` (audit notes in the R-465 closure). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-491** | **[P2-MEDIUM] Removing an app leaves its update hold in the store, so a reinstall under the same name starts HELD.** MEASURED 2026-09-13 on demo-hp (v0.241.0, R-479 live check): nextcloud was held after a failed update, then removed with data and backups; `settings.json` still carried `restore_holds.nextcloud` (`reason: update_failed`, `copy_tier: 1`). `fillHoldReason` hides the sentence for a not-deployed app (R-480), but every start gate reads the store, so the next deploy of `nextcloud` would be refused as held with a sentence about a backup that no longer exists. Cleared by hand with `-clear-restore-hold`. **Fix shape:** `removeStack` clears an UPDATE hold (never an R-379 restore hold, which stays operator-cleared) — same place the prefs are forgotten; a test that deploys after a held removal. Next controller release. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user