From bca013edecae8451fb688c4e1da5c63750fd6833 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 23:31:13 +0200 Subject: [PATCH] CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it Round 2 (restore gokapi + power cut, 20s into the restore): the box came back BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being restored when the plug came out - returned healthy. The only alarm was controller_started, which is what the ladder expects for a 60-second outage: no node_stale (30 min threshold), no app_start_failed (90s boot grace). No false alarm, none missed. Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the loop died with the box and zero lines is not zero failures. The round also handed over immich's whole diagnosis. app_oom fired - "immich (immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the app and the exact container. That is why immich saw CONNECTION_CLOSED and crash-looped twelve times. It is added as tonight's line on the EXISTING OOM row rather than filed as a new one, because this project's standing finding is that those signals are invisible inside LXC guests and on this box the scan caught one. The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm feed, correctly labelled, the whole time. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/DRILL-chaos-night-2026-09-17.md | 25 ++++++++++++- .../round-2.txt | 35 +++++++++++++++++++ .../round-3-accident.txt | 7 ++++ .../round-3.txt | 8 +++++ documentation/backlog/OPEN-ITEMS.md | 2 +- 5 files changed, 75 insertions(+), 2 deletions(-) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-3-accident.txt create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index dbc414cb..73437729 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -206,7 +206,30 @@ correct and legible: it refused to route to unhealthy containers, `app_start_fai that was down, `backup_run_failures` said „1 of 12 … nextcloud", the memory guard refused the eleventh and twelfth installs with both numbers quoted, and `storage_fill_critical` fired at 100 %. -### Rounds 2-12 +### Round 2 — `restore` gokapi / accident: **power cut**, 20 s into the restore + +| the five things | | +|---|---| +| what the customer saw | „Visszaállítás elindult — az állapot itt frissül." then the box went dark mid-restore; on return the dashboard and every app were back | +| what the box did by itself | everything — containers 0 → **25 at t+131s** → **26 at t+148s**, nothing stuck, no shell used. **gokapi, the app being restored when the plug came out, returned `Up 30 seconds (healthy)`** | +| time to steady | **148 s**, measured against the container count this round took itself before the accident | +| alarm fired / true? | `controller_started` (info) — true and correct. **No false alarm.** | +| should have fired, did not | **none** — per the ladder a 60-second outage yields no `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace), and neither appeared | + +**Household loop: NOT COLLECTED.** The loop was a transient unit on the VM and died with the power +cut — the first accident that could have produced household failures instead produced no lines at +all. Zero lines is not zero failures, so it is recorded as not collected, and the loop is now a +persistent systemd unit that returns with the box. + +**A finding this round handed over:** `app_oom` (warning) — „Alkalmazás memóriája elfogyott: immich +(immich-postgres) — egy folyamatát a memóriakorlát leállította". That is immich's whole mystery +solved: its Postgres was OOM-killed during the reverse-geocoding import, which is why the server saw +`CONNECTION_CLOSED` and crash-looped twelve times. **The controller caught an OOM inside an LXC guest +and named the exact container** — worth recording against this project's standing finding that those +signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in +the alarm feed, correctly labelled, the whole time. + +### Rounds 3-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-2.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-2.txt index 226b8bd6..38fd395c 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-2.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-2.txt @@ -61,3 +61,38 @@ gokapi Up 30 seconds (healthy) | Sep 16 21:19 | info | app_deploy_started | Alkalmazás telepítése elindult: Immich | controller | Sep 16 21:19 | info | app_removed | Alkalmazás eltávolítva: immich | controller 2026-09-16T21:28:32Z ================ END ROUND 2 ================ + +## ROUND 2 — the five things +**action:** `restore` gokapi from its local backup · **accident:** power cut, 20 s into the restore + +1. **What the customer saw.** „Visszaállítás elindult — az állapot itt frissül." and then the box went + dark mid-restore. On return, the dashboard and every app came back on their own. +2. **What the box did by itself.** Everything. Power returned 21:26:01Z; containers went 0 → **25 at + t+131s** → **26 at t+148s**, with no shell, no help and nothing stuck. **gokapi — the app that was + being restored when the plug came out — came back `Up 30 seconds (healthy)`.** +3. **Time to steady.** **148 seconds** (back to 26 of 26 containers, measured against the count this + round took itself before the accident). +4. **Alarm fired / true?** `controller_started` (info) at 21:28 — true, and correct: a power cut + inside the 30-minute staleness window is not a liveness alarm, so the ladder expects exactly this + one line and nothing else. **No false alarm fired.** +5. **Should have fired and did not.** None. Per `08-alarm-ladder.md` a 60-second outage produces no + `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace covers the restart) — + and neither appeared. + +**Household loop:** NOT COLLECTED for this round — the loop died with the box (see above). Fixed for +every round after this one. + +## A real finding this round handed me: the OOM was DETECTED + Sep 16 21:20 warning **app_oom** + „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította" + +That is immich's `CONNECTION_CLOSED immich-postgres:5432` explained: its Postgres was killed by the +memory limit during the reverse-geocoding import, the connection died mid-query, the metadata service +failed, and the worker exited — twelve times. + +**Why it is worth writing down:** this project carries a standing finding that OOM signals are +invisible inside LXC guests (`OOMKilled` false, no docker oom events). Here the controller's own +scan DID catch it, named the app AND the exact container, and said in plain Hungarian that a process +was stopped by the memory limit. It is an operator-only warning by register, so no customer mail — +also correct. **The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm +feed, correctly labelled, the whole time.** diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-3-accident.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-3-accident.txt new file mode 100644 index 00000000..c1030955 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-3-accident.txt @@ -0,0 +1,7 @@ +2026-09-16T21:29:49Z ACCIDENT=disk-95-full round=3 +2026-09-16T21:29:49Z filling the customer guest's SYSTEM disk to ~95 % — with a POOL GUARD +2026-09-16T21:29:51Z pool now: 39% used; guest / has 29352 MB free +/dev/mapper/pve-vm--9201--disk--0 32G 944M 29G 4% / +/dev/mapper/pve-vm--9201--disk--0 32G 29G 1.5G 96% / +2026-09-16T21:29:55Z pool after the fill: 39.69 +2026-09-16T21:29:55Z full — holding 10 minutes diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt new file mode 100644 index 00000000..02863ab6 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-3.txt @@ -0,0 +1,8 @@ +2026-09-16T21:29:45Z ================ ROUND 3 : use bookstack, while: disk-95-full ================ +2026-09-16T21:29:48Z --- BEFORE --- containers=26 wiki=200 status=200 paste=200 +2026-09-16T21:29:48Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) +2026-09-16T21:29:48Z --- ACTION: use on bookstack --- +2026-09-16T21:29:48Z wiki read 1 -> 200 +2026-09-16T21:29:49Z wiki read 2 -> 200 +2026-09-16T21:29:49Z wiki read 3 -> 200 +2026-09-16T21:29:49Z --- ACCIDENT: disk-95-full (injected after the action started) --- diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index b64d7633..67c9a220 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -734,7 +734,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** | | **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** | | **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** | -| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** | +| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. | | **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | | **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** | | **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** |