evening: rules file in place; R-483 CLOSED (catalog, operator-confirmed); R-479 CLOSED (controller v0.241.0, proven live); R-481 blocked on the one-host-one-customer model; R-491 filed
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
This commit is contained in:
@@ -279,3 +279,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-461** | **`runbooks/target-selection.md` named a venue that does not exist and fenced a fixture that is gone.** Closed 2026-09-13, both halves checked against both boxes first: (a) no `/mnt/nvme-1tb` on demo-hp or demo-felhom; demo-hp's NVMe is `nvme0n1` at `/mnt/hdd_1` (demo-felhom's `/mnt/hdd_1` is `sdb`) — the runbook now names `/mnt/hdd_1` and says it is the same disk as the data drive; (b) `qm list` is empty on BOTH boxes — `drill-r50` (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | **CLOSED 2026-09-13 — DOCUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-452** | **Nothing enforced `catalog_since`, so the badge's one number could silently under-report.** Closed 2026-09-13 (catalog): `scripts/check-catalog-since.py`, the fifth gate in `catalog_gates.py` — a `--range A..B` gate in the engine-major shape: an app whose per-service `image:` lines differ across the range must carry a `catalog_since` on or after the moving commit's day and not in the future; comments, README and CHANGELOG mentions are not the fact. Hook-enforced; the shallow CI clone skips it out loud (the CI-shape half the row named stays as is, by the same reasoning engine-major uses). Five decoy cases; red-proof: dropping the date comparison lets the untouched-date fact through (`audits/v0240-2026-09-13/rp-R452.txt`). | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-465** | **`cfg.Paths.HDDPath` — empty on every box — still had six readers; were any inert?** AUDITED 2026-09-13 on demo-hp (registered drive `/mnt/felhom-drives/hdd_1`, `hdd_path` absent, no `FELHOM_PATHS_HDD_PATH`). Five of six fall back before the value matters: `report/builder.go:69` and `monitor/healthcheck.go:35` take `storagePaths[0]`; `web/server.go:740` (`primaryHDDPath`) takes the default storage path; `main.go:511` (metrics) takes the default storage path; `main.go:347` passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. **One is inert AND unreachable:** `api/router.go:981` (`systemInfo`, `GET /api/system/info`) reads the empty value with no fallback (`hdd_configured:false` forever) — and the endpoint itself is shadowed: the web layer's `ServeSystemAPI` claims `/api/system/*` and answers **404 „ismeretlen végpont"** for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as **R-490**. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | **CLOSED 2026-09-13 — AUDITED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-479** | **For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so.** Operator ruling 2026-09-13; closed in controller **v0.241.0** (`3e81330`): an app with classified binds walks second drive → off-site → own unit (`UpdateTierOrderFor`), and the hold sentence ends with what the chosen copy holds (`RestoreHold.CopyHolds`). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). `audits/v0241-2026-09-13/` | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-483** | **adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog).** Closed 2026-09-13: the k3s ingress routes `/media`, `/static`, `/admin`, `/accounts` to the backend service on port **80** — the nginx inside the backend image that serves Django's `X-Accel-Redirect` media; the catalog routed everything to the frontend. Two catalog cuts (`3172258` router, `ed62cfd` port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (`GET /media/…webp` with a session → 200 `image/webp`, RIFF/WEBP) and **confirmed by the operator in a browser at 21:49** (two photos render). Scripted multipart uploads through the frontend's `/api` proxy still 500 (upstream `RequestContentLengthMismatchError`); browser uploads work — not a template matter. `audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt` | **CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -695,13 +695,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
|
||||
|
||||
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-483** | **[P3-LOW] `adventurelog` v0.12.1: a photo upload through the app's own API proxy fails 500 from every non-browser client — whether a household can upload photos in a browser is UNKNOWN and cannot be measured here.** MEASURED 2026-09-13 on demo-hp (throwaway `travelnight`): `POST /api/images` (multipart: `image`, `location`, `is_primary`) through the public origin returns `{"error":"Internal Server Error"}` from the **frontend** (the backend log shows no request at all); the frontend container logs `RequestContentLengthMismatchError: Request body length does not match content-length header` from its undici forwarder. Same result with an ASCII filename, an accented one, urllib, curl `-F`, and curl with `Transfer-Encoding: chunked`. Everything else on the API — sign-up, the frontend's login form, collections, locations, visits, notes, edit, delete — works. **What is NOT established:** whether the app's own browser page uploads succeed (a browser's FormData body may satisfy the proxy). Per `CLAUDE.md`, that is a manual click-through: open `https://<sub>.<domain>`, add a photo to a place. **Not ours to fix in the template** (a frontend proxy bug); a newer catalog pin is a version promotion, not tonight's. Evidence: `audits/nightly-2026-09-13-adventurelog/03e-photo-500.txt`. | **WAITING-ON-OPERATOR — needs a browser click-through; rank P3-LOW; owner: VIKTOR checks, CC re-files** |
|
||||
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). **RULED 2026-09-13:** a second enrolled LXC on demo-hp, disk on the NVMe path (`/mnt/hdd_1`), sized like 9201, enrolled as a scratch CUSTOMER with its disposition recorded. **MEASURED BEFORE BUILDING — the ruling collides with the product's own model:** the hub's `hosts` table keys a host to exactly ONE customer (`host_id` PK, `customer_id NOT NULL`) and the box runs ONE agent with ONE `host_id`; a second customer on the same Proxmox host would need a second host identity, and the installer (`felhom-host-install.sh --customer-id … --vmid …`) run with another customer-id on demo-hp would rewrite the existing agent's identity — i.e. break the standing demo-hp enrolment. So "enrolled as a scratch customer" is not something the product can do on a box that already belongs to a customer. **What the product CAN do today, two options:** (a) **a second guest of the demo-hp customer** (the `guests` table is per host+vmid; the agent's provision writes a per-vmid bootstrap): supported by the data model, but both controllers report as customer demo-hp and the customer page shows one controller — the standing box's monitoring flips between the two unless the scratch guest's controller is told no hub (unenrolled scratch, disposition recorded on demo-hp's customer page); (b) **a separate scratch HOST** — a nested Proxmox VM on demo-hp (the ISO appliance route the earlier walks used, VM 323/325) enrolled as its own customer: fully enrolled, fully isolated, but an appliance to build (hours) and keep. Disk placement is the same under both: the installer has no rootfs-storage flag (only `--rootfs-grow`, `--archive-storage`), so the guest lands on `local-lvm` and is moved with `pct move-volume` to a `dir` storage created at `/mnt/hdd_1` — a post-provision step, reversible. 9201's shape for "sized like 9201": 7 cores, 25 898 MB, rootfs 32 G + mp0 70 G on local-lvm, unprivileged. **Recommendation: (a) with the hub left out** — it is the reversible one and it gives the rotation its restore target tomorrow; (b) if "enrolled" is the point. Not built tonight: §1 of the rules — a decision that changes what the product promises about hosts is not CC's. | **WAITING-ON-OPERATOR — the ruling cannot be executed as stated; two options below; rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-487** | **[P2-MEDIUM] A removed app whose backups were kept is listed on NEITHER backup page, so the restore that brings it back has no button — the customer's remove-by-mistake route exists only as an endpoint.** MEASURED 2026-09-13 on demo-hp (nightly rotation, `adventurelog` removed with backups kept, unit + mirror on disk): `GET /backups/apps` and `GET /backups/restore` contain the string `adventurelog` zero times; `POST /backup/restore stack_name=adventurelog snapshot_id=helyi` then restored it in 22 s with the data byte-identical. Cause: `buildAppBackupRows` walks `status.AppDataInfo` = `DiscoverAppData` over DEPLOYED stacks only. The off-site list had exactly this defect and was fixed by keying it on the store (R-237, v0.204.0); the local and Tier-2 lists were not. **Fix shape:** list every app with a recovery unit on a registered drive (`ListRestorePoints` over the primary dirs), marking removed ones „eltávolítva — visszaállítható"; the unit restore already reinstalls (R-253). Not a design reversal — the same rule R-237 set. Evidence: `audits/nightly-2026-09-13-adventurelog/05b-restore-tier1.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-490** | **[P3-LOW] The monitoring page's „Memória-eloszlás" card has never rendered: its `fetch('/api/system/info')` is answered 404 by the web layer's `ServeSystemAPI`, which claims all of `/api/system/*` and knows only the two memory routes.** MEASURED 2026-09-13 on demo-hp: `GET /api/system/info` → `{"error":"ismeretlen végpont","ok":false}`; `monitoring.html` shows the card (`display:none` by default) only when that fetch returns `used_mem_mb`, so it stays hidden on every box. The API router's `systemInfo` handler (`internal/api/router.go:981`) is unreachable, and it is also the one reader of the always-empty `cfg.Paths.HDDPath` with no fallback (R-465). **Fix shape:** let `ServeSystemAPI` fall through to the API router for unknown `/api/system/*` paths (or route `/api/system/info` explicitly), give `systemInfo` the same default-storage-path fallback the other readers have, and pin the card with a render test; then delete the global (R-465's deferred deletion). Next controller release. Evidence: `audits/nightly-2026-09-13-adventurelog/` (audit notes in the R-465 closure). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-491** | **[P2-MEDIUM] Removing an app leaves its update hold in the store, so a reinstall under the same name starts HELD.** MEASURED 2026-09-13 on demo-hp (v0.241.0, R-479 live check): nextcloud was held after a failed update, then removed with data and backups; `settings.json` still carried `restore_holds.nextcloud` (`reason: update_failed`, `copy_tier: 1`). `fillHoldReason` hides the sentence for a not-deployed app (R-480), but every start gate reads the store, so the next deploy of `nextcloud` would be refused as held with a sentence about a backup that no longer exists. Cleared by hand with `-clear-restore-hold`. **Fix shape:** `removeStack` clears an UPDATE hold (never an R-379 restore hold, which stays operator-cleared) — same place the prefs are forgotten; a test that deploys after a held removal. Next controller release. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user