R-887 opened: CI jobs never picked up by the runner, failed after ~10 min with no log (register 291 -> 292; 1 opened, 45 closed this session)
gates / gates (push) Successful in 1m45s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 17:59:51 +02:00
parent 1122b5c65b
commit e8c56c440a
4 changed files with 19 additions and 8 deletions
+1 -1
View File
@@ -16,7 +16,7 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 291 (0 opened, 45
> **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 292 (1 opened — R-887 CI runner fault — 45
> closed): 24 fixed by later work + 2 duplicates (each re-checked; `audits/burndown-2026-10-05/partA-table.md` holds all
> 317 P3/P4 verdicts), 19 small fixes with tests/red-proofs (catalog `29ac711`, agent `d833163`, controller `114ff27`,
> felhom.eu `ab2b304`…). New gate `script-tests` (every `scripts/**/test_*.py` per push, R-885); `closed_register_gate`
+10 -3
View File
@@ -9,9 +9,9 @@
| Rows before | Rows after | Opened | Closed |
|---|---|---|---|
| **336** | **291** | **0** | **45** |
| **336** | **292** | **1** | **45** |
Counted by `register_shape_gate.py`'s method (`| **R-n** |` lines in `OPEN-ITEMS.md`). Target was ≥ 40 fewer: 45.
Counted by `register_shape_gate.py`'s method (`| **R-n** |` lines in `OPEN-ITEMS.md`). Target was ≥ 40 fewer: 44 net (45 closed, 1 opened — R-887, below).
## Baselines (re-verified at the start)
@@ -77,7 +77,14 @@ saved to a file** — re-run them by mutating as described in each CHANGELOG ent
the pre-push hook everything was green, so the push went through and CI caught it: jobs 1351, 1352, 1353, 1358, 1359 =
failure, each mailed to the operator (`RESEND-ACCEPTED` in the log). Fixed in `40d34c5` (a form both `date`s read;
red-proved with a BusyBox-only PATH: the old form gives the same 9 failures) → **job 1360 = success**. The gate did what
it was built for, on its first day. The copy installed on DooPlex is the previous revision (same behaviour on GNU
it was built for, on its first day.
**A second CI fault, not mine to fix (R-887, opened):** felhom.eu job 1361 (`1122b5c`) and controller job 1357 (`114ff27`) were
never picked up by the runner (no `task` line in its log; task id = job id + 1), and Gitea failed them after ~10–13 min with
no log and no alarm. Re-run through the API: **controller 1357 → success (32 s)**; felhom.eu 1361 → never picked up again,
failure. Suspected stale runner registration; my token cannot list runners. Final CI per repo: agent `d833163` → job 1356
success; catalog `29ac711` → job 1355 success; controller `114ff27` → job 1357 (re-run) success; felhom.eu `40d34c5` → job
1360 success (the later docs-only commits: see the final check below). The copy installed on DooPlex is the previous revision (same behaviour on GNU
`date`); not reinstalled because DooPlex was out of scope.
**One slip, said plainly:** R-885's gate commit (`ab2b304`) went out WITHOUT closing the row — my closing file had a
+6 -3
View File
@@ -4,7 +4,7 @@
sent to it.**
**Updated 2026-10-05 (night, the burn-down session): every box of ours healthy (nothing was changed on any box). The
open-items list went from 336 to 291. Report: `REPORT-burndown-2026-10-05.md`.**
open-items list went from 336 to 292. Report: `REPORT-burndown-2026-10-05.md`.**
## Tonight, later (2026-10-05, night): the list got shorter
@@ -19,17 +19,20 @@ open-items list went from 336 to 291. Report: `REPORT-burndown-2026-10-05.md`.**
it; it is gone, and a test keeps it gone.
- **You got 5 „CI failed" mails tonight — my fault, fixed.** The new test gate ran one of today's test suites on the CI
machine for the first time; it uses a different `date` tool, and 9 tests failed. Fixed; CI is green again (job 1360).
- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes, with no log and no mail. One re-run passed; the other was again never picked up. It looks like a leftover runner registration, but I cannot see the runner list. If you do nothing: now and then a CI run reads „failed" without testing anything. Fix: Site Administration → Runners, remove any offline duplicate of `felhom-gates-runner`.
- **A new rule, so the list stops growing:** a small problem found during work (about 30 minutes) is fixed in that
session and never added to the list. Every report now states four numbers: rows before, after, opened, closed.
**The numbers:** 336 before → **291 after**; 0 opened; 45 closed.
**The numbers:** 336 before → **292 after**; 1 opened (the CI fault below); 45 closed.
**Needs you:**
1. **The list below: 45 rows I would close as „accepted".** Answer per row, or „all as picked". If you do nothing,
they stay open and the list stays at 291.
they stay open and the list stays at 292.
2. **Two leaked tokens to rotate** (R-831, R-870, at the end of the list). If you do nothing, whoever has those
transcripts keeps that access.
3. Earlier tonight's items are unchanged (pin the other „latest" apps on DooPlex; Alertmanager's file permissions).
4. **Look at the CI runners** (Gitea → Site Administration → Runners; R-887): remove any offline duplicate of
`felhom-gates-runner`. If you do nothing, some CI runs fail without running.
## Your list: rows I would close as 'accepted' (burn-down 2026-10-05) - answer per row, or 'all as picked'
+2 -1
View File
@@ -391,7 +391,7 @@ stopping line that lies.
| **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC |
| **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC |
## Process & tooling — 64 rows (P3 4, P4 60)
## Process & tooling — 65 rows (P3 5, P4 60)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -399,6 +399,7 @@ stopping line that lies.
| **R-578** | Process & tooling | P3 | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | — | — | CC |
| **R-586** | Process & tooling | P3 | **[P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw.** FOUND 2026-09-18 starting slice 4 (R-559): running `scripts/iso/test/bootstrap-modes.sh` unchanged at `183727db9c44` reported `FAIL: R-496: banner painted to the console seam` and `FAIL: R-496: banner names the Tulajdonosi jelmondat`. **Cause:** the script paints each banner with `> "$CONSOLE_DEV"`. On a real console that is a character device and truncation is a no-op, so every paint appears; pointed at a plain file, as the harness did, each paint TRUNCATES. Commit `c033b3b` (ISO 1.28.0, R-535, 2026-09-16) added `print_bound_banner`, which paints immediately after the pairing banner in the same invocation — from that commit the pairing banner was wiped before the check read it. `c033b3b` did not touch the harness. **Why it survived: the harness is in NO gate and NO CI run** — not in `repo_gates.py`, not in `.gitea/`; it runs only when a person runs it, and between 09-16 and 09-18 nobody did. FIXED in the same session: the harness points `FELHOM_CONSOLE_DEV` at a FIFO with a background reader, which restores device semantics (opening a FIFO with `>` truncates nothing) and lets a test see EVERY paint — which slice 4's golden checks then needed anyway. Production code untouched. **What is still open: the harness remains outside every gate.** It needs a container, so it cannot join `repo_gates.py --fast`, which is what both the pre-push hook and CI run — meaning a non-fast entry would still never execute. **Fix shape:** either give CI a container-capable job that runs it, or make the ISO release gate's G16 the place it is required (done for G16 this session — so it now runs at least once per ISO release, which is better than never but later than a push). | **READY - rank P2-MED; owner: CC** **Re-ranked 2026-10-03: P2->P3: test harness gap; now required at each ISO release gate.** | — | — | CC |
| **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC |
| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. | **OPEN** | — | Operator: list the runners (Site Administration → Runners) and remove any offline duplicate of `felhom-gates-runner`; then push and confirm the new job runs | operator |
| **R-93** | Process & tooling | P4 | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture".** | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
| **R-129** | Process & tooling | P4 | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC |
| **R-161** | Process & tooling | P4 | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **OPEN** — **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator |