From e8c56c440a9e46cd16249b33a4e6393c67bbd949 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 5 Oct 2026 17:59:51 +0200 Subject: [PATCH] R-887 opened: CI jobs never picked up by the runner, failed after ~10 min with no log (register 291 -> 292; 1 opened, 45 closed this session) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 2 +- REPORT-burndown-2026-10-05.md | 13 ++++++++++--- STATUS.md | 9 ++++++--- documentation/backlog/OPEN-ITEMS.md | 3 ++- 4 files changed, 19 insertions(+), 8 deletions(-) diff --git a/CONTEXT.md b/CONTEXT.md index c01c54a5..079a8c6e 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,7 +16,7 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. -> **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 291 (0 opened, 45 +> **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 292 (1 opened — R-887 CI runner fault — 45 > closed): 24 fixed by later work + 2 duplicates (each re-checked; `audits/burndown-2026-10-05/partA-table.md` holds all > 317 P3/P4 verdicts), 19 small fixes with tests/red-proofs (catalog `29ac711`, agent `d833163`, controller `114ff27`, > felhom.eu `ab2b304`…). New gate `script-tests` (every `scripts/**/test_*.py` per push, R-885); `closed_register_gate` diff --git a/REPORT-burndown-2026-10-05.md b/REPORT-burndown-2026-10-05.md index 2095f7da..2b1f8a67 100644 --- a/REPORT-burndown-2026-10-05.md +++ b/REPORT-burndown-2026-10-05.md @@ -9,9 +9,9 @@ | Rows before | Rows after | Opened | Closed | |---|---|---|---| -| **336** | **291** | **0** | **45** | +| **336** | **292** | **1** | **45** | -Counted by `register_shape_gate.py`'s method (`| **R-n** |` lines in `OPEN-ITEMS.md`). Target was ≥ 40 fewer: 45. +Counted by `register_shape_gate.py`'s method (`| **R-n** |` lines in `OPEN-ITEMS.md`). Target was ≥ 40 fewer: 44 net (45 closed, 1 opened — R-887, below). ## Baselines (re-verified at the start) @@ -77,7 +77,14 @@ saved to a file** — re-run them by mutating as described in each CHANGELOG ent the pre-push hook everything was green, so the push went through and CI caught it: jobs 1351, 1352, 1353, 1358, 1359 = failure, each mailed to the operator (`RESEND-ACCEPTED` in the log). Fixed in `40d34c5` (a form both `date`s read; red-proved with a BusyBox-only PATH: the old form gives the same 9 failures) → **job 1360 = success**. The gate did what -it was built for, on its first day. The copy installed on DooPlex is the previous revision (same behaviour on GNU +it was built for, on its first day. + +**A second CI fault, not mine to fix (R-887, opened):** felhom.eu job 1361 (`1122b5c`) and controller job 1357 (`114ff27`) were +never picked up by the runner (no `task` line in its log; task id = job id + 1), and Gitea failed them after ~10–13 min with +no log and no alarm. Re-run through the API: **controller 1357 → success (32 s)**; felhom.eu 1361 → never picked up again, +failure. Suspected stale runner registration; my token cannot list runners. Final CI per repo: agent `d833163` → job 1356 +success; catalog `29ac711` → job 1355 success; controller `114ff27` → job 1357 (re-run) success; felhom.eu `40d34c5` → job +1360 success (the later docs-only commits: see the final check below). The copy installed on DooPlex is the previous revision (same behaviour on GNU `date`); not reinstalled because DooPlex was out of scope. **One slip, said plainly:** R-885's gate commit (`ab2b304`) went out WITHOUT closing the row — my closing file had a diff --git a/STATUS.md b/STATUS.md index f277015d..81d40a8d 100644 --- a/STATUS.md +++ b/STATUS.md @@ -4,7 +4,7 @@ sent to it.** **Updated 2026-10-05 (night, the burn-down session): every box of ours healthy (nothing was changed on any box). The -open-items list went from 336 to 291. Report: `REPORT-burndown-2026-10-05.md`.** +open-items list went from 336 to 292. Report: `REPORT-burndown-2026-10-05.md`.** ## Tonight, later (2026-10-05, night): the list got shorter @@ -19,17 +19,20 @@ open-items list went from 336 to 291. Report: `REPORT-burndown-2026-10-05.md`.** it; it is gone, and a test keeps it gone. - **You got 5 „CI failed" mails tonight — my fault, fixed.** The new test gate ran one of today's test suites on the CI machine for the first time; it uses a different `date` tool, and 9 tests failed. Fixed; CI is green again (job 1360). +- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes, with no log and no mail. One re-run passed; the other was again never picked up. It looks like a leftover runner registration, but I cannot see the runner list. If you do nothing: now and then a CI run reads „failed" without testing anything. Fix: Site Administration → Runners, remove any offline duplicate of `felhom-gates-runner`. - **A new rule, so the list stops growing:** a small problem found during work (about 30 minutes) is fixed in that session and never added to the list. Every report now states four numbers: rows before, after, opened, closed. -**The numbers:** 336 before → **291 after**; 0 opened; 45 closed. +**The numbers:** 336 before → **292 after**; 1 opened (the CI fault below); 45 closed. **Needs you:** 1. **The list below: 45 rows I would close as „accepted".** Answer per row, or „all as picked". If you do nothing, - they stay open and the list stays at 291. + they stay open and the list stays at 292. 2. **Two leaked tokens to rotate** (R-831, R-870, at the end of the list). If you do nothing, whoever has those transcripts keeps that access. 3. Earlier tonight's items are unchanged (pin the other „latest" apps on DooPlex; Alertmanager's file permissions). +4. **Look at the CI runners** (Gitea → Site Administration → Runners; R-887): remove any offline duplicate of + `felhom-gates-runner`. If you do nothing, some CI runs fail without running. ## Your list: rows I would close as 'accepted' (burn-down 2026-10-05) - answer per row, or 'all as picked' diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index c1998eea..2a3dea9e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -391,7 +391,7 @@ stopping line that lies. | **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 64 rows (P3 4, P4 60) +## Process & tooling — 65 rows (P3 5, P4 60) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -399,6 +399,7 @@ stopping line that lies. | **R-578** | Process & tooling | P3 | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-586** | Process & tooling | P3 | **[P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw.** FOUND 2026-09-18 starting slice 4 (R-559): running `scripts/iso/test/bootstrap-modes.sh` unchanged at `183727db9c44` reported `FAIL: R-496: banner painted to the console seam` and `FAIL: R-496: banner names the Tulajdonosi jelmondat`. **Cause:** the script paints each banner with `> "$CONSOLE_DEV"`. On a real console that is a character device and truncation is a no-op, so every paint appears; pointed at a plain file, as the harness did, each paint TRUNCATES. Commit `c033b3b` (ISO 1.28.0, R-535, 2026-09-16) added `print_bound_banner`, which paints immediately after the pairing banner in the same invocation — from that commit the pairing banner was wiped before the check read it. `c033b3b` did not touch the harness. **Why it survived: the harness is in NO gate and NO CI run** — not in `repo_gates.py`, not in `.gitea/`; it runs only when a person runs it, and between 09-16 and 09-18 nobody did. FIXED in the same session: the harness points `FELHOM_CONSOLE_DEV` at a FIFO with a background reader, which restores device semantics (opening a FIFO with `>` truncates nothing) and lets a test see EVERY paint — which slice 4's golden checks then needed anyway. Production code untouched. **What is still open: the harness remains outside every gate.** It needs a container, so it cannot join `repo_gates.py --fast`, which is what both the pre-push hook and CI run — meaning a non-fast entry would still never execute. **Fix shape:** either give CI a container-capable job that runs it, or make the ISO release gate's G16 the place it is required (done for G16 this session — so it now runs at least once per ISO release, which is better than never but later than a push). | **READY - rank P2-MED; owner: CC** **Re-ranked 2026-10-03: P2->P3: test harness gap; now required at each ISO release gate.** | — | — | CC | | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC | +| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. | **OPEN** | — | Operator: list the runners (Site Administration → Runners) and remove any offline duplicate of `felhom-gates-runner`; then push and confirm the new job runs | operator | | **R-93** | Process & tooling | P4 | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture".** | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC | | **R-129** | Process & tooling | P4 | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC | | **R-161** | Process & tooling | P4 | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **OPEN** — **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator |