CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it
gates / gates (push) Successful in 23s

Round 2 (restore gokapi + power cut, 20s into the restore): the box came back
BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being
restored when the plug came out - returned healthy. The only alarm was
controller_started, which is what the ladder expects for a 60-second outage:
no node_stale (30 min threshold), no app_start_failed (90s boot grace). No
false alarm, none missed.

Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the
loop died with the box and zero lines is not zero failures.

The round also handed over immich's whole diagnosis. app_oom fired - "immich
(immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the
app and the exact container. That is why immich saw CONNECTION_CLOSED and
crash-looped twelve times. It is added as tonight's line on the EXISTING OOM
row rather than filed as a new one, because this project's standing finding is
that those signals are invisible inside LXC guests and on this box the scan
caught one. The diagnosis I spent twenty minutes reaching from logs was sitting
in the alarm feed, correctly labelled, the whole time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 23:31:13 +02:00
parent 5b6e4b5c30
commit bca013edec
5 changed files with 75 additions and 2 deletions
@@ -206,7 +206,30 @@ correct and legible: it refused to route to unhealthy containers, `app_start_fai
that was down, `backup_run_failures` said „1 of 12 … nextcloud", the memory guard refused the
eleventh and twelfth installs with both numbers quoted, and `storage_fill_critical` fired at 100 %.
### Rounds 2-12
### Round 2 — `restore` gokapi / accident: **power cut**, 20 s into the restore
| the five things | |
|---|---|
| what the customer saw | „Visszaállítás elindult — az állapot itt frissül." then the box went dark mid-restore; on return the dashboard and every app were back |
| what the box did by itself | everything — containers 0 → **25 at t+131s** → **26 at t+148s**, nothing stuck, no shell used. **gokapi, the app being restored when the plug came out, returned `Up 30 seconds (healthy)`** |
| time to steady | **148 s**, measured against the container count this round took itself before the accident |
| alarm fired / true? | `controller_started` (info) — true and correct. **No false alarm.** |
| should have fired, did not | **none** — per the ladder a 60-second outage yields no `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace), and neither appeared |
**Household loop: NOT COLLECTED.** The loop was a transient unit on the VM and died with the power
cut — the first accident that could have produced household failures instead produced no lines at
all. Zero lines is not zero failures, so it is recorded as not collected, and the loop is now a
persistent systemd unit that returns with the box.
**A finding this round handed over:** `app_oom` (warning) — „Alkalmazás memóriája elfogyott: immich
(immich-postgres) — egy folyamatát a memóriakorlát leállította". That is immich's whole mystery
solved: its Postgres was OOM-killed during the reverse-geocoding import, which is why the server saw
`CONNECTION_CLOSED` and crash-looped twelve times. **The controller caught an OOM inside an LXC guest
and named the exact container** — worth recording against this project's standing finding that those
signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in
the alarm feed, correctly labelled, the whole time.
### Rounds 3-12
PENDING
@@ -61,3 +61,38 @@ gokapi Up 30 seconds (healthy)
| Sep 16 21:19 | info | app_deploy_started | Alkalmazás telepítése elindult: Immich | controller
| Sep 16 21:19 | info | app_removed | Alkalmazás eltávolítva: immich | controller
2026-09-16T21:28:32Z ================ END ROUND 2 ================
## ROUND 2 — the five things
**action:** `restore` gokapi from its local backup · **accident:** power cut, 20 s into the restore
1. **What the customer saw.** „Visszaállítás elindult — az állapot itt frissül." and then the box went
dark mid-restore. On return, the dashboard and every app came back on their own.
2. **What the box did by itself.** Everything. Power returned 21:26:01Z; containers went 0 → **25 at
t+131s** → **26 at t+148s**, with no shell, no help and nothing stuck. **gokapi — the app that was
being restored when the plug came out — came back `Up 30 seconds (healthy)`.**
3. **Time to steady.** **148 seconds** (back to 26 of 26 containers, measured against the count this
round took itself before the accident).
4. **Alarm fired / true?** `controller_started` (info) at 21:28 — true, and correct: a power cut
inside the 30-minute staleness window is not a liveness alarm, so the ladder expects exactly this
one line and nothing else. **No false alarm fired.**
5. **Should have fired and did not.** None. Per `08-alarm-ladder.md` a 60-second outage produces no
`node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace covers the restart) —
and neither appeared.
**Household loop:** NOT COLLECTED for this round — the loop died with the box (see above). Fixed for
every round after this one.
## A real finding this round handed me: the OOM was DETECTED
Sep 16 21:20 warning **app_oom**
„Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította"
That is immich's `CONNECTION_CLOSED immich-postgres:5432` explained: its Postgres was killed by the
memory limit during the reverse-geocoding import, the connection died mid-query, the metadata service
failed, and the worker exited — twelve times.
**Why it is worth writing down:** this project carries a standing finding that OOM signals are
invisible inside LXC guests (`OOMKilled` false, no docker oom events). Here the controller's own
scan DID catch it, named the app AND the exact container, and said in plain Hungarian that a process
was stopped by the memory limit. It is an operator-only warning by register, so no customer mail —
also correct. **The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm
feed, correctly labelled, the whole time.**
@@ -0,0 +1,7 @@
2026-09-16T21:29:49Z ACCIDENT=disk-95-full round=3
2026-09-16T21:29:49Z filling the customer guest's SYSTEM disk to ~95 % — with a POOL GUARD
2026-09-16T21:29:51Z pool now: 39% used; guest / has 29352 MB free
/dev/mapper/pve-vm--9201--disk--0 32G 944M 29G 4% /
/dev/mapper/pve-vm--9201--disk--0 32G 29G 1.5G 96% /
2026-09-16T21:29:55Z pool after the fill: 39.69
2026-09-16T21:29:55Z full — holding 10 minutes
@@ -0,0 +1,8 @@
2026-09-16T21:29:45Z ================ ROUND 3 : use bookstack, while: disk-95-full ================
2026-09-16T21:29:48Z --- BEFORE --- containers=26 wiki=200 status=200 paste=200
2026-09-16T21:29:48Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:29:48Z --- ACTION: use on bookstack ---
2026-09-16T21:29:48Z wiki read 1 -> 200
2026-09-16T21:29:49Z wiki read 2 -> 200
2026-09-16T21:29:49Z wiki read 3 -> 200
2026-09-16T21:29:49Z --- ACCIDENT: disk-95-full (injected after the action started) ---
+1 -1
View File
@@ -734,7 +734,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
| **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** |
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |
| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. |
| **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** |
| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** |
| **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** |