diff --git a/REPORT.md b/REPORT.md index 77e11bb5..4aeef77d 100644 --- a/REPORT.md +++ b/REPORT.md @@ -29,7 +29,7 @@ ## Register -**194 rows before, 202 after.** Nothing closed, and that is stated rather than implied: R-438 and +**194 rows before, 205 after.** Nothing closed, and that is stated rather than implied: R-438 and R-440 are **amended and stay OPEN** — the mechanism is now documented, not changed — so nothing moved to `CLOSED-ITEMS.md` and that file is untouched. @@ -43,6 +43,9 @@ to `CLOSED-ITEMS.md` and that file is untouched. | R-451 | slice 7 — a fleet sweep (needs a hub change: no image field is reported) | READY, P3-LOW | | R-452 | no gate enforces `catalog_since` (`--depth 1` has no parent to diff) | READY, P3-LOW | | R-453 | **the vaulted dashboard password is stale on BOTH demo boxes** | **WAITING-ON-OPERATOR**, P2-MEDIUM | +| R-454 | five `gofmt`-unclean `internal/web` test files, and no gate notices | READY, P3-LOW | +| R-455 | **DooPlex has no Docker Hub login — the ceiling now blocks BUILDS, not just gates** | **WAITING-ON-OPERATOR**, P2-MEDIUM | +| R-456 | a partly-dead stack is not a boot orphan, and that rule is written down nowhere | READY, P3-LOW | ## Live validation diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 070ba6e2..a84e25e0 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -691,6 +691,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** | | **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** | | **R-453** | **[P2-MEDIUM] The vaulted customer dashboard password is STALE ON BOTH DEMO BOXES, and there is no operator-side route to the real one — so no session can drive a customer page.** MEASURED 2026-09-02 while trying to live-validate the update badge: `PASSWORD` from DooPlex `~/.config/credentials` returns **HTTP 200 with the login page and the body string `Hibás jelszó`** against demo-hp guest 9201 (`https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`) **and** against demo-felhom guest 9201 (`https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu`). **The discriminator is the controller's own log, not the status code** — `auth.go:176: [WARN] [web] Failed login` proves wrong PASSWORD rather than wrong Host header, which is the trap this class always presents (a rejected login renders no flash and looks exactly like a routing problem). Every other key in the credentials file was checked and none is a dashboard password (`HUB_PW`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*`, and the `R_*` keys are escrow recovery codes). **THIS IS THE SECOND TIME:** it drifted on demo-hp on 2026-08-09 and was put back on operator instruction by writing a fresh bcrypt hash into the guest's `data/settings.json`; demo-hp was then reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted as well. **Why it is not merely inconvenient: it silently converts "endpoint-level validation" — this project's STANDARD method, because there is no browser on DooPlex — into "unit tests only" for anything that renders a customer page.** The cost is paid per session and rediscovered each time. **The fix is a decision, not a command:** the claim code is bcrypt-hashed hub-side and only e-mailed (R-119), so either the operator records the current demo passwords out-of-band, or a re-set becomes a standing authorisation for the two Tier-0 demo boxes. **NOT TAKEN UNILATERALLY:** re-setting a dashboard password is a decision about a customer account, and it was done under operator instruction last time. Raised in `STATUS.md` item 9. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4, which lists all five attempts. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** | +| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | +| **R-455** | **[P2-MEDIUM] DooPlex has no Docker Hub login, and the unauthenticated ceiling now blocks BUILDS, not just gates.** MEASURED 2026-09-02: `build.sh 0.233.0 --push` failed at `[internal] load metadata for docker.io/library/debian:bookworm-slim: 429 Too Many Requests`, with `golang:1.24-bookworm` cancelled behind it. Neither base image was in the local store, so there was no fallback. The same window also left the catalog's `image-resolvable` gate INCONCLUSIVE (6 of 65 lookups throttled). **This is R-41's known note — *"the full sweep is still OWED — DooPlex is not logged in to Docker Hub"* — arriving with a bigger bill: it is no longer only a gate that cannot reach a verdict, it is a release that cannot be built.** **WORKAROUND USED AND RECORDED RATHER THAN BURIED:** both base images were pulled from Google's official Docker Hub mirror (`mirror.gcr.io/library/...`) and retagged locally, after which `build.sh` ran unmodified; the two digests are written into `felhom-controller/REPORT.md` §5 so the identity check against Hub is one command when the window clears. **GRADED: that the mirror is byte-identical to Hub is KNOWN, not MEASURED — the throttle was still in force at the end of the session.** **The fix is a credential, and it is a decision:** a Docker Hub account for DooPlex (free tier lifts the ceiling ~6x for authenticated pulls), or an explicit standing ruling that the mirror is the sanctioned source and `build.sh` should name it. **A build that depends on an anonymous third-party quota is not a build you can run when you need to.** Owner: **VIKTOR decides, CC executes.** | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** | +| **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** |