diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 15587392..db9472ed 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -368,7 +368,7 @@ stopping line that lies. | **R-565** | Process & tooling | P3 | **[P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page.** FOUND 2026-09-17 by slice 1 release C (R-556, controller v0.250.0): after the extractor and the tests were green, a by-eye review of the English renders found six Hungarian fragments still in JavaScript strings — „, majd a(z)” and „FIGYELEM:” in the storage decommission dialog, „jelenlegi:” on the drive-init list, the uptime units „mp” and „p” and the count word „ db” on the debug page. All six were converted by hand; **no test failed on any of them**, because `TestI18nEnglishPages` looks for Hungarian letters and the extractor's ASCII word list (`i18n_extract.py` `ASCII_HU`) is used by neither test nor gate. Release B's review had found more of the same kind (Konfig, Megtartva, helyi, pl., Befejezve, automatikus, jelenleg:, kedd/szerda/szombat, szint). **Fix shape:** run the ASCII word list over the English renders in `TestI18nEnglishPages` (after the data mask), with a negative control on an English sentence and a decoy planting „mp” in an English value; extend the list with the words releases B and C found. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-578** | Process & tooling | P3 | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC | -| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** | **OPEN** | — | Operator: list the runners (Site Administration → Runners) and remove any offline duplicate of `felhom-gates-runner`; then push and confirm the new job runs | operator | +| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. | — | Run the 2026-10-12 check; if a job is lost again, read Gitea's log for the assignment within the hour (it rotates) and the runner's fetch timeout | operator | | **R-129** | Process & tooling | P4 | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-208** | Process & tooling | P4 | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** | **READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN** | — | **Measured, not inferred.** All **208** retained `go mod download` records carried **`Usage count: 1`** — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the **same** `--build-arg VERSION` → `RUN go mod download` **CACHED**; the **same** tree with a **new** `VERSION` → **executed**. `COPY go.mod ./` stays CACHED either way, which is the tell: a `COPY`'s key is content-based, while a `RUN`'s key includes the stage **environment**, and `ARG VERSION`/`ARG GIT_COMMIT` are declared *before* the download in `felhom-controller/controller/Dockerfile`. Since every real build passes a fresh version, the layer is invalidated **every single time**. **`felhom.eu/hub/Dockerfile` has the identical defect** (`ARG VERSION`/`ARG BUILD_TIME` above `COPY go.mod go.sum*` → `RUN go mod download`) — and because both Dockerfiles produce byte-identical `buildx du` description strings, the 208 records are a COMBINED count and must not be attributed to one project. **Fix shape (one line each, not applied here):** move the `ARG VERSION`/`ARG GIT_COMMIT`/`ARG BUILD_TIME` declarations down to just above the final `go build`. **Worth more than the cap and the move combined** — the cap bounds the symptom, this removes the source. `build.sh`'s `rm -rf` + `cp -a` and its host-side `go mod tidy` were **ruled out by fingerprinting**: the tree is byte-identical across runs and `tidy` is a no-op | CC | @@ -426,5 +426,6 @@ stopping line that lies. the R-row. Duplicating them here would create the second source this design avoids. --> | item | due (UTC) | what to measure | |---|---|---| +| R-887 | 2026-10-12 | no CI job lost since 2026-10-05 16:00Z (every completed job has a runner `task` line and a log); detail in the R-887 row | | R-872 | 2026-10-06 | the first live 05:00 deadline run judges a down box on the longer lines (Tester 2, if still off): hub log + the two events (detail in the R-872 row) |