CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it
gates / gates (push) Successful in 23s

Round 2 (restore gokapi + power cut, 20s into the restore): the box came back
BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being
restored when the plug came out - returned healthy. The only alarm was
controller_started, which is what the ladder expects for a 60-second outage:
no node_stale (30 min threshold), no app_start_failed (90s boot grace). No
false alarm, none missed.

Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the
loop died with the box and zero lines is not zero failures.

The round also handed over immich's whole diagnosis. app_oom fired - "immich
(immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the
app and the exact container. That is why immich saw CONNECTION_CLOSED and
crash-looped twelve times. It is added as tonight's line on the EXISTING OOM
row rather than filed as a new one, because this project's standing finding is
that those signals are invisible inside LXC guests and on this box the scan
caught one. The diagnosis I spent twenty minutes reaching from logs was sitting
in the alarm feed, correctly labelled, the whole time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 23:31:13 +02:00
parent 5b6e4b5c30
commit bca013edec
5 changed files with 75 additions and 2 deletions
@@ -206,7 +206,30 @@ correct and legible: it refused to route to unhealthy containers, `app_start_fai
that was down, `backup_run_failures` said „1 of 12 … nextcloud", the memory guard refused the
eleventh and twelfth installs with both numbers quoted, and `storage_fill_critical` fired at 100 %.
### Rounds 2-12
### Round 2 — `restore` gokapi / accident: **power cut**, 20 s into the restore
| the five things | |
|---|---|
| what the customer saw | „Visszaállítás elindult — az állapot itt frissül." then the box went dark mid-restore; on return the dashboard and every app were back |
| what the box did by itself | everything — containers 0 → **25 at t+131s** → **26 at t+148s**, nothing stuck, no shell used. **gokapi, the app being restored when the plug came out, returned `Up 30 seconds (healthy)`** |
| time to steady | **148 s**, measured against the container count this round took itself before the accident |
| alarm fired / true? | `controller_started` (info) — true and correct. **No false alarm.** |
| should have fired, did not | **none** — per the ladder a 60-second outage yields no `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace), and neither appeared |
**Household loop: NOT COLLECTED.** The loop was a transient unit on the VM and died with the power
cut — the first accident that could have produced household failures instead produced no lines at
all. Zero lines is not zero failures, so it is recorded as not collected, and the loop is now a
persistent systemd unit that returns with the box.
**A finding this round handed over:** `app_oom` (warning) — „Alkalmazás memóriája elfogyott: immich
(immich-postgres) — egy folyamatát a memóriakorlát leállította". That is immich's whole mystery
solved: its Postgres was OOM-killed during the reverse-geocoding import, which is why the server saw
`CONNECTION_CLOSED` and crash-looped twelve times. **The controller caught an OOM inside an LXC guest
and named the exact container** — worth recording against this project's standing finding that those
signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in
the alarm feed, correctly labelled, the whole time.
### Rounds 3-12
PENDING
@@ -61,3 +61,38 @@ gokapi Up 30 seconds (healthy)
| Sep 16 21:19 | info | app_deploy_started | Alkalmazás telepítése elindult: Immich | controller
| Sep 16 21:19 | info | app_removed | Alkalmazás eltávolítva: immich | controller
2026-09-16T21:28:32Z ================ END ROUND 2 ================
## ROUND 2 — the five things
**action:** `restore` gokapi from its local backup · **accident:** power cut, 20 s into the restore
1. **What the customer saw.** „Visszaállítás elindult — az állapot itt frissül." and then the box went
dark mid-restore. On return, the dashboard and every app came back on their own.
2. **What the box did by itself.** Everything. Power returned 21:26:01Z; containers went 0 → **25 at
t+131s** → **26 at t+148s**, with no shell, no help and nothing stuck. **gokapi — the app that was
being restored when the plug came out — came back `Up 30 seconds (healthy)`.**
3. **Time to steady.** **148 seconds** (back to 26 of 26 containers, measured against the count this
round took itself before the accident).
4. **Alarm fired / true?** `controller_started` (info) at 21:28 — true, and correct: a power cut
inside the 30-minute staleness window is not a liveness alarm, so the ladder expects exactly this
one line and nothing else. **No false alarm fired.**
5. **Should have fired and did not.** None. Per `08-alarm-ladder.md` a 60-second outage produces no
`node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace covers the restart) —
and neither appeared.
**Household loop:** NOT COLLECTED for this round — the loop died with the box (see above). Fixed for
every round after this one.
## A real finding this round handed me: the OOM was DETECTED
Sep 16 21:20 warning **app_oom**
„Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította"
That is immich's `CONNECTION_CLOSED immich-postgres:5432` explained: its Postgres was killed by the
memory limit during the reverse-geocoding import, the connection died mid-query, the metadata service
failed, and the worker exited — twelve times.
**Why it is worth writing down:** this project carries a standing finding that OOM signals are
invisible inside LXC guests (`OOMKilled` false, no docker oom events). Here the controller's own
scan DID catch it, named the app AND the exact container, and said in plain Hungarian that a process
was stopped by the memory limit. It is an operator-only warning by register, so no customer mail —
also correct. **The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm
feed, correctly labelled, the whole time.**
@@ -0,0 +1,7 @@
2026-09-16T21:29:49Z ACCIDENT=disk-95-full round=3
2026-09-16T21:29:49Z filling the customer guest's SYSTEM disk to ~95 % — with a POOL GUARD
2026-09-16T21:29:51Z pool now: 39% used; guest / has 29352 MB free
/dev/mapper/pve-vm--9201--disk--0 32G 944M 29G 4% /
/dev/mapper/pve-vm--9201--disk--0 32G 29G 1.5G 96% /
2026-09-16T21:29:55Z pool after the fill: 39.69
2026-09-16T21:29:55Z full — holding 10 minutes
@@ -0,0 +1,8 @@
2026-09-16T21:29:45Z ================ ROUND 3 : use bookstack, while: disk-95-full ================
2026-09-16T21:29:48Z --- BEFORE --- containers=26 wiki=200 status=200 paste=200
2026-09-16T21:29:48Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:29:48Z --- ACTION: use on bookstack ---
2026-09-16T21:29:48Z wiki read 1 -> 200
2026-09-16T21:29:49Z wiki read 2 -> 200
2026-09-16T21:29:49Z wiki read 3 -> 200
2026-09-16T21:29:49Z --- ACCIDENT: disk-95-full (injected after the action started) ---