From 7674947ae91caa37b60b95b30298d81f8db765d1 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 7 Oct 2026 05:50:25 +0200 Subject: [PATCH] R-872 closed (proven at 05:00, two channels); R-887 night count 0; numbers 138 -> 137 Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT-night-burndown-2026-10-06.md | 2 +- STATUS.md | 2 +- .../audits/night-burndown-2026-10-06/MORNING-NOTE.md | 3 ++- .../audits/night-burndown-2026-10-06/NIGHT-LOG.md | 1 + .../audits/night-burndown-2026-10-06/r872-check.txt | 11 +++++++++++ documentation/backlog/CLOSED-ITEMS.md | 1 + documentation/backlog/OPEN-ITEMS.md | 6 ++---- 7 files changed, 19 insertions(+), 7 deletions(-) create mode 100644 documentation/audits/night-burndown-2026-10-06/r872-check.txt diff --git a/REPORT-night-burndown-2026-10-06.md b/REPORT-night-burndown-2026-10-06.md index 91ab5744..0f298c09 100644 --- a/REPORT-night-burndown-2026-10-06.md +++ b/REPORT-night-burndown-2026-10-06.md @@ -6,7 +6,7 @@ Row by row, with minutes and commits: `documentation/audits/night-burndown-2026- | Rows before | Rows after | Opened | Closed | |---|---|---|---| -| 138 | 138 | 2 (R-895, R-896) | 2 (R-377, R-392) | +| 138 | 137 | 2 (R-895, R-896) | 3 (R-377, R-392, R-872) | - **§3 R-892:** stopped — no SSH agent here; DooPlex's own keys refused (`s3/`). - **§4 R-894:** fixed on agent main `74b5eae` (4 red-proofs, `07` §6.1). diff --git a/STATUS.md b/STATUS.md index 0df590e1..21b0cdae 100644 --- a/STATUS.md +++ b/STATUS.md @@ -3,7 +3,7 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** **Updated 2026-10-07 05:50: hub 0.140.0; demo-hp, demo-felhom and Tester 1 run controller 0.301.0, agent 0.149.0 — -nothing was delivered tonight. The open-items list is at 138. Night: `documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md`. +nothing was delivered tonight. The open-items list is at 137. Night: `documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md`. Day: Parts B and D of the read-back brief after 08:30, then the releases (hub 0.141.0, agent 0.150.0, controller, the held catalog branch).** diff --git a/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md b/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md index 5b3db907..3ab84feb 100644 --- a/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md +++ b/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md @@ -17,7 +17,7 @@ folder (photos, documents) is empty on every fresh install, so failing it would ## 3. The four numbers -Rows before: **138**. Rows after: **138**. Opened: **2**. Closed: **2**. +Rows before: **138**. Rows after: **137**. Opened: **2**. Closed: **3**. ## 4. What was fixed @@ -30,6 +30,7 @@ Rows before: **138**. Rows after: **138**. Opened: **2**. Closed: **2**. - **Disk health:** three more warning counters now travel from the disk to the hub. Nothing alarms on them yet. - **Apps (held, not live):** Jellyfin and Emby no longer let a stranger on the internet sign in as if at home (measured, then fixed). Also wishlist, homepage, gramps-web and two test tools. +- **Proven at 05:00:** a box that is off at night now raises the missed-backup alarm. Tester 2 got it, and the mail reached your inbox. - **Measured:** wger runs best with its proper web server and 2 workers (41 % of its memory). **Waits for today's releases:** diff --git a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md index 9e36cb04..0fda40ad 100644 --- a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md @@ -53,3 +53,4 @@ no reboot. | R-516 | **22 Go-literal messages moved into the bundle** (14 formal → te-form; English added); a detached NAS attach failure follows the reader language; 2 red-proofs | 60 | controller `3d6ba28` | | R-330 | **wire half built** in three repos (agent sends 187/188/199, controller decodes, hub models; carried only — no verdict reads them); 5 red-proofs | 15 | hub/agent/controller (this batch) | | (hub release) | **NOT DEPLOYED.** 00:10 the release commit (`hub/CHANGELOG.md` v0.141.0: R-366, R-31, R-330 wire) was pushed, green; at 00:12 the image build (`build/felhom-hub/build.sh 0.141.0 --push`) was **refused by the session's permission check** („Production Deploy"). Stopped there (brief Rules 3), no workaround; the manifest still names 0.140.0, nothing changed on the hub. The CHANGELOG says so. The day session builds and deploys 0.141.0 | 5 | (this batch) | +| R-872 | **closed** — the 05:00 run judged Tester 2 on the longer lines (dump missed=1, backup missed=0 correctly) and the mail reached the operator inbox (second channel) | 10 | (this batch) | diff --git a/documentation/audits/night-burndown-2026-10-06/r872-check.txt b/documentation/audits/night-burndown-2026-10-06/r872-check.txt new file mode 100644 index 00000000..94b1c28e --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/r872-check.txt @@ -0,0 +1,11 @@ +R-872 dated check 2026-10-07 — read 05:50 CEST + +Channel 1 — hub pod log (sudo kubectl -n felhom-system logs deploy/hub --since=3h | grep deadline): +2026/10/07 05:00:00 [INFO] Deadline check: Tester-2 is DOWN — judged on the longer lines (dump 48h0m0s, whole-guest 72h0m0s): dump missed=1 backup missed=0 +2026/10/07 05:00:00 [INFO] Deadline check: 4 customers, 0 backup missed, 0 backup unknown (deferred), 1 dbdump missed, 1 skipped (down), 0 unknown (no host ever bound) +2026/10/07 05:00:00 [INFO] deadline-check: next run at 2026-10-08 05:00 CEST (in 23h59m59s) + +Channel 2 — operator mailbox (Gmail connector, newer_than:1d): +2026-10-07T03:00:01Z monitoring@felhom.eu -> admin@felhom.eu [Felhom] Tester-2: expected_dbdump_missed (Severity error, Time 2026-10-07 05:00 CEST, 'No DB dump for over 48h0m0s while the box is down at its deadline (last: none in the last 7 days)') + +backup missed=0 is correct: the whole-guest line is 72 h and Tester 2's last whole-guest copy (ep0 listing, readback C/C2) is 2026-10-04 16:31Z — 60.5 h old at 05:00. diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index ef883d96..067b4b64 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -32,6 +32,7 @@ The full text of every row below: `git show e866a66b56:documentation/backlog/OPE | Row | What | Closed | Evidence | |---|---|---|---| +| **R-872** | **A box that is off every night never raised a missed-backup alarm.** (P2) — full text `git show 3aba200293:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — PROVEN LIVE (hub v0.134.0 + v0.140.0) | 05:00 run: `Deadline check: Tester-2 is DOWN — judged on the longer lines (dump 48h0m0s, whole-guest 72h0m0s): dump missed=1 backup missed=0`; the operator mailbox holds `[Felhom] Tester-2: expected_dbdump_missed` sent 05:00 (a second channel). `backup missed=0` is correct: the last whole-guest copy was 60.5 h old against the 72 h line. `audits/night-burndown-2026-10-06/r872-check.txt` | | **R-392** | **No architecture document covers the two-AI workflow.** (P4) — full text `git show 93b942f136:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-06 night — WRITTEN | `documentation/architecture/12-agent-tooling.md`: the two roles and what each owns, where each kind of instruction lives, who may change an instruction file (decision 150), and why the boundaries sit where they do; it points at the doc-authoring skill and does not restate it. Routed from `.claude/rules/docs.md`. | | **R-377** | **`CONTEXT.md`'s standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them.** (P4) | CLOSED 2026-10-06 night — DONE AS THE ROW SAID | 44 `### S-n (date)` sub-headings added above the rulings in `CONTEXT.md` § Standing rulings; no ruling's text edited, reordered or compressed (the diff is 88 added lines, 0 removed). Five ids are used twice in the log (S-11..S-15) — kept as written, the log is never edited. | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 555b58c0..516a0335 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -229,7 +229,7 @@ stopping line that lies. | **R-468** | Box system & updates | P3 | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | — | — | CC | | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | -## Monitoring & notifications — 16 rows (P2 3, P3 9, P4 4) +## Monitoring & notifications — 15 rows (P2 2, P3 9, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -246,7 +246,6 @@ stopping line that lies. | **R-177** | Monitoring & notifications | P4 | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the same missing operator door as R-314 and R-279. | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | | **R-266** | Monitoring & notifications | P4 | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator | | **R-285** | Monitoring & notifications | P4 | **A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere.** During the 2026-08-09 rehearsal the hub sent, all `status: sent` to the operator channel: `host_stale` 08:58 UTC, `node_stale` 09:00, **`host_down` 09:28 (error)**, **`node_down` 09:30 (error)**, `host_leaf_changed` 09:31, `host_recovered` 09:31, `node_recovered` 09:34, `offsite_delivery_stuck` 09:34 — eight operator mails for work that was deliberate, attended and announced. **This is the OPPOSITE gap from the one R-281 filed:** the alarms are not missing, they are indiscriminate. `host_stale` at 30 min and `host_down` at 60 min (`monitor/host_staleness.go:22-23`, `downAfter = 2 * threshold`) cannot distinguish a wiped-on-purpose box from a dead one, and `host_leaf_changed` firing on a reinstall is correct-but-expected. **Note the interaction with the mute used on 2026-08-09 evening:** blocking a customer silences everything, so today the only two settings are *page me for planned work* and *tell me nothing at all*. **What is owed is a middle:** a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible | **READY (M) — NEW 2026-08-09** | — | The evidence is the operator's mailbox plus `events`/`notification_log` for 2026-08-09 | CC | -| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** **NIGHT WATCH 2026-10-06 05:00 (burn-down night): the run WORKED AS DESIGNED but could not be seen to.** Hub log: `Deadline check: 4 customers, 0 backup missed … 1 skipped (down)` and NO per-box line. Read from a hub.db copy with -wal (a different channel): Tester 2's first host report was 2026-10-04 16:13:44Z, its last 2026-10-04 18:05:48Z — at 05:00 it was 34.8 h old, under the 48 h line, so `judgeDownCustomer` returned early („nothing expected yet") — silently. No alarm was owed, none fired. **Fixed without a row (hub, on `main`, ships with the next hub release):** each early return now logs why (`… is DOWN — not judged yet: first host report … ago`); `TestR872_DownButTooYoungIsLogged`, red-proved. **RE-DATED 2026-10-06 → 2026-10-07:** from 16:14Z today Tester 2 is old enough; if it is still off at 05:00 tomorrow the existing `… is DOWN — judged on the longer lines …` line and its events are the proof. `audits/night-burndown-2026-10-05/r872-watch.txt`. | R-871 | — | CC | | **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. **Checked from source 2026-10-05 (burn-down round 2):** homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager re | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | | **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. **Checked from source 2026-10-05 (burn-down round 2):** Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | @@ -283,7 +282,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** **2026-10-06 night: NARROWED** — the harness records each container's `swap_peak` and the venue's swap (catalog branch `night-held-2026-10-06`, `3f4611c`; SwapRecorded red-proved; bench SwapTotal 0 kB, 9202 524288 kB). LEFT: the golden's swap in its bake evidence, the box walk's record, and whether proofs run with swap off. | — | — | CC | -| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | +| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. **NIGHT WATCH 2026-10-06/07 (second burn-down night): 0 jobs lost** — every push checked by its commit (felhom.eu, controller, agent, catalog main and the held branch), each completed `success`, one filtered call per check. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | | **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config` and no entry in `operations/nodes.md`. **Corrected 2026-10-06 18:24 (`09` §3 decision 158):** its Proxmox host IS known — it runs as **VM 341 on the HP box** (`ssh hp`), recorded in `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`; the session that filed this did not find that file. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** **2026-10-06 18:24: operator ruling — yes, CC may reach it by SSH (decision 158).** **2026-10-06 (night): the route is BUILT in the box walk** (catalog `d63ea35`: `TARGET=tester-1`, `ssh -J demo-hp root@192.168.0.154`, guest 9201, `felhom.enkicsifelhom.hu`; `BOX_ADMIN_SEED_GUESTS` has it) and the identity is matched (the agent's report `host.node=felhom` = VM 341's certificate; the guest answers its domain 200, demo-hp's 404). **Blocked:** DooPlex's key is not authorized on VM 341 (`Permission denied (publickey,password)`); the VM has no guest agent and a disk edit needs a VM stop (a reboot, not allowed); fetching its vaulted password from the hub was refused by the session's permission check. `audits/readback-2026-10-07/` **2026-10-06 20:27 (night brief §3): still blocked** — the operator ran `ssh-copy-id` at 20:20, but it copied the key of the operator's own SSH agent; this shell has no agent, and DooPlex's two keys (`id_rsa`, `id_ed25519`) are still refused, direct and through demo-hp (`audits/night-burndown-2026-10-06/s3/`). The fix: run `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` from DooPlex as kisfenyo. | DooPlex's key on VM 341 | Operator: authorize DooPlex's public key on VM 341 (one line in `/root/.ssh/authorized_keys`, from its console), or allow CC to read the vaulted password; then prove one step there | operator | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | @@ -314,5 +313,4 @@ stopping line that lies. | item | due (UTC) | what to measure | |---|---|---| | R-887 | 2026-10-12 | no CI job lost since 2026-10-05 16:00Z (every completed job has a runner `task` line and a log); detail in the R-887 row | -| R-872 | 2026-10-07 | the 05:00 deadline run judges a down box on the longer lines (Tester 2, if still off; it is old enough from 2026-10-06 16:14Z): hub log `… is DOWN — judged on the longer lines …` + the events (detail in the R-872 row) |