R-265: a CI run failed with NO LOG, and the alarm points at a log that does not exist
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
Run 264 (650cc8a, a DOCUMENTATION-ONLY commit) failed between two greens of byte-identical gate code.
Not waved away as a flake, because this project's own record is that a "known flake" can be a true
positive.
MEASURED. Every other run this session: 18-34s, log present (HTTP 200). Run 264: 834s (07:12:40 ->
07:26:34 UTC) and GET /actions/jobs/264/logs returns HTTP 500 - "264.log.zst: file does not exist".
The act-runner pod never restarted (0 restarts, 5d17h), so the job hung and was reaped; the runner
did not die.
NOT A GATE FINDING, on four independent facts: the diff from the green before it is Markdown only;
the same content is green two commits later (265, 33s); the gate code is identical across 263/264/265;
and 260-262, which WERE real gate failures, each failed in under 35s WITH a log.
THE CAUSE OF THE HANG IS UNDETERMINED and is deliberately recorded as such. DooPlex was doing heavy
work in that window (139 MB kubectl cp, a go run compiling the whole hub module), which is a
plausible contention story - but 40 cores at load ~5 does not establish it, so it is filed as a
hypothesis rather than asserted as a cause.
THE FINDING THAT MATTERS IS SECOND-ORDER, and it is gates.yml's own purpose turned against it. The
workflow exists because "a detector nobody hears is the defect R-29 filed, rebuilt one layer up", and
its alarm mail says "The failing gate names itself in the run log." There is no run log. An operator
following that sentence finds nothing and cannot tell a reap from a conviction. Unverified and worse:
the alarm step is `if: failure()` and whether it ran at all for a reaped job is unknown - if it did
not, this was a red CI that alarmed nobody.
Fix shapes recorded, none built: surface duration + log-presence in the alarm; an explicit
timeout-minutes under the reap so it fails fast and loudly WITH a log; and one deliberate test of
whether the alarm fires on a reaped job, because until that runs, "CI alarms on failure" is an
assumption.
This commit is contained in:
@@ -107,6 +107,18 @@ that is the fail-open shape and would leave it running in neither home (R-29).
|
||||
mail before run **263** went green. The alarm working is the system behaving correctly; the noise was
|
||||
mine.
|
||||
|
||||
**A FOURTH red run, 264, was NOT one of mine and is filed as R-265.** It sat between two greens on a
|
||||
**documentation-only** commit, ran **834 s** against 18–34 s for every other run in the session, and
|
||||
**persisted no log at all** (`jobs/264/logs` → HTTP 500, *file does not exist*). The runner pod never
|
||||
restarted, so the job hung and was reaped rather than the runner dying. Not a gate finding — the diff
|
||||
was Markdown, the gate code was byte-identical to the two greens around it, and the same content is
|
||||
green at run 265. **The cause of the hang is undetermined and is not guessed at**; DooPlex was busy in
|
||||
that window with this session's own live-validation work, but the box has 40 cores at load ~5, so
|
||||
that is a hypothesis, not a cause. **The reusable finding is second-order:** the alarm mail tells the
|
||||
operator "the failing gate names itself in the run log", and here there is no run log — so a reap is
|
||||
indistinguishable from a conviction, and whether the alarm fired at all for a reaped job is
|
||||
**unverified**. R-265 carries it.
|
||||
|
||||
## 6. What `oobDegraded` says when it fails
|
||||
|
||||
```
|
||||
|
||||
@@ -441,6 +441,8 @@ builds the receiving struct by hand cannot see a field that never decodes, which
|
||||
|---|---|---|
|
||||
| **R-264** | **Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed".** Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (`operator_key_configured`) could not be mistaken for having decided what the hub should do with the rest. **The list, grouped by what a consumer would be for.** **(a) Guest-network health — `guest_net` and its seven children** (`checked_at`, `has_route`, `dhclient_alive`, `heal_succeeded`, `heals_last_hour`, `last_heal_at`, `damped`). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed `dhclient` took a tunnel down for 1 h 15 m (`audits/INCIDENT-guest-dhclient-killed-2026-07-20.md`); a recurring-heal signal is exactly what would have surfaced it. **This is the strongest candidate of the twenty-one.** **(b) `selfupdate_pending` + `selfupdate_pending_version`** — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. **(c) `mgmt_plane.healed_recently`** — bounded: the hub DOES alarm on the `privsep_healed_at` timestamp beside it, so the recurring-clobber signal is not lost, only this flag. **(d) `restore_tests.mount_parity` + `mount_inventory`** — R-262's subject; the verdict is not lost (a mismatch fails before `Pass` is set) but the hub cannot tell a full-fidelity pass from a boot-only one. **(e) `pbs_dr.applied_at`.** **(f) Controller-side: `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check`** — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. **For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it** — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. **Deliberately not decided in the G-1 session**, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change | **READY** — owner Viktor |
|
||||
|
||||
| **R-265** | **A CI run can fail with NO LOG PERSISTED, and the alarm mail then points the operator at a log that does not exist.** Observed 2026-08-08 as run **264** (`650cc8a`, a **documentation-only** commit) sat between two green runs of identical gate code. **Measured rather than assumed — the shape is unmistakable:** every other run in the session took **18–34 s and has a log (HTTP 200)**; 264 took **834 s** (07:12:40 → 07:26:34 UTC) and `GET /actions/jobs/264/logs` returns **HTTP 500 — `actions_log/…/264.log.zst: file does not exist`**. The runner pod never restarted (`act-runner`, 0 restarts, 5 d 17 h uptime), so the runner did not die; the JOB hung and was reaped. **It is NOT a gate finding, and four independent facts say so:** the diff from the green run before it is Markdown only; the same content is green two commits later (run 265, `dd55a3f`, 33 s); the gate code is byte-identical across 263/264/265; and 260–262, which WERE real gate failures, all failed in under 35 s **with** logs. **THE CAUSE OF THE HANG IS UNDETERMINED and is deliberately not guessed at.** DooPlex was doing heavy work in that window (a 139 MB `kubectl cp`, and a `go run` compiling the whole hub module for the live-validation harness), which is a plausible contention story — but the box has 40 cores and sat at load ~5, so it is **not established** and is recorded as a hypothesis, not a cause. **THE FINDING THAT MATTERS IS THE SECOND-ORDER ONE, and it is this workflow's own stated purpose turned against it.** `gates.yml` exists because "a detector nobody hears is the defect R-29 filed, rebuilt one layer up", and its alarm mail says *"The failing gate names itself in the run log."* **Here there is no run log**, so an operator following that sentence finds nothing and cannot tell an infrastructure reap from a real conviction. Worse and **unverified**: the alarm step is `if: failure()`, and whether it even ran for a reaped job is unknown — if it did not, this was a red CI that alarmed nobody, which is exactly the shape the workflow was built to prevent. **Fix shape, not a decision:** (a) make the alarm mail state the run's DURATION and whether a log exists, so a log-less reap is self-identifying; (b) give the job an explicit `timeout-minutes` well under the reap so it fails fast, loudly and with a log; (c) establish whether the alarm fires at all on a reaped job — that is one deliberate test, and until it is run, "CI alarms on failure" is an assumption | **READY** — owner Viktor |
|
||||
|
||||
**Explicitly still open, untouched by this session:** R-246 (the wrong stale flag on `demo-hp` —
|
||||
clearing it is an operator act hub-side), R-255, R-256, R-257, R-258, R-259, R-261, R-262, R-263, and
|
||||
**C7's test-comment half**, which Campaign 12 recorded as *owed, not done* (60 of 2652 production
|
||||
|
||||
Reference in New Issue
Block a user