drill 0.243.0: F9'' answers the supervisor budget question; machine torn down
gates / gates (push) Successful in 20s

Three more kills 20 minutes apart: recovered in 61 s / 41 s / 61 s, and none of
them accumulated, because the window is 15 minutes. Four restarts, zero pauses.
So the brake catches a FAST crash loop and is blind to a SLOW one — a controller
dying every 20 minutes is restarted forever, and the only trace is an info event
that mails nobody. Measured, not changed: the options are written into R-531 for
the operator to rule on.

Machine layer torn down: VM 334 purged with its disks, demo-hp's own containers
9201 and 9202 untouched. Evidence copied off the box first, token-leak control 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 13:59:27 +02:00
parent db58af80a2
commit 0b63574293
8 changed files with 2308 additions and 1 deletions
+1 -1
View File
@@ -728,7 +728,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |
| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** |
| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** |
| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** |
| **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.