CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it
gates / gates (push) Successful in 23s
gates / gates (push) Successful in 23s
Round 2 (restore gokapi + power cut, 20s into the restore): the box came back BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being restored when the plug came out - returned healthy. The only alarm was controller_started, which is what the ladder expects for a 60-second outage: no node_stale (30 min threshold), no app_start_failed (90s boot grace). No false alarm, none missed. Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the loop died with the box and zero lines is not zero failures. The round also handed over immich's whole diagnosis. app_oom fired - "immich (immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the app and the exact container. That is why immich saw CONNECTION_CLOSED and crash-looped twelve times. It is added as tonight's line on the EXISTING OOM row rather than filed as a new one, because this project's standing finding is that those signals are invisible inside LXC guests and on this box the scan caught one. The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm feed, correctly labelled, the whole time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -206,7 +206,30 @@ correct and legible: it refused to route to unhealthy containers, `app_start_fai
|
||||
that was down, `backup_run_failures` said „1 of 12 … nextcloud", the memory guard refused the
|
||||
eleventh and twelfth installs with both numbers quoted, and `storage_fill_critical` fired at 100 %.
|
||||
|
||||
### Rounds 2-12
|
||||
### Round 2 — `restore` gokapi / accident: **power cut**, 20 s into the restore
|
||||
|
||||
| the five things | |
|
||||
|---|---|
|
||||
| what the customer saw | „Visszaállítás elindult — az állapot itt frissül." then the box went dark mid-restore; on return the dashboard and every app were back |
|
||||
| what the box did by itself | everything — containers 0 → **25 at t+131s** → **26 at t+148s**, nothing stuck, no shell used. **gokapi, the app being restored when the plug came out, returned `Up 30 seconds (healthy)`** |
|
||||
| time to steady | **148 s**, measured against the container count this round took itself before the accident |
|
||||
| alarm fired / true? | `controller_started` (info) — true and correct. **No false alarm.** |
|
||||
| should have fired, did not | **none** — per the ladder a 60-second outage yields no `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace), and neither appeared |
|
||||
|
||||
**Household loop: NOT COLLECTED.** The loop was a transient unit on the VM and died with the power
|
||||
cut — the first accident that could have produced household failures instead produced no lines at
|
||||
all. Zero lines is not zero failures, so it is recorded as not collected, and the loop is now a
|
||||
persistent systemd unit that returns with the box.
|
||||
|
||||
**A finding this round handed over:** `app_oom` (warning) — „Alkalmazás memóriája elfogyott: immich
|
||||
(immich-postgres) — egy folyamatát a memóriakorlát leállította". That is immich's whole mystery
|
||||
solved: its Postgres was OOM-killed during the reverse-geocoding import, which is why the server saw
|
||||
`CONNECTION_CLOSED` and crash-looped twelve times. **The controller caught an OOM inside an LXC guest
|
||||
and named the exact container** — worth recording against this project's standing finding that those
|
||||
signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in
|
||||
the alarm feed, correctly labelled, the whole time.
|
||||
|
||||
### Rounds 3-12
|
||||
|
||||
PENDING
|
||||
|
||||
|
||||
@@ -61,3 +61,38 @@ gokapi Up 30 seconds (healthy)
|
||||
| Sep 16 21:19 | info | app_deploy_started | Alkalmazás telepítése elindult: Immich | controller
|
||||
| Sep 16 21:19 | info | app_removed | Alkalmazás eltávolítva: immich | controller
|
||||
2026-09-16T21:28:32Z ================ END ROUND 2 ================
|
||||
|
||||
## ROUND 2 — the five things
|
||||
**action:** `restore` gokapi from its local backup · **accident:** power cut, 20 s into the restore
|
||||
|
||||
1. **What the customer saw.** „Visszaállítás elindult — az állapot itt frissül." and then the box went
|
||||
dark mid-restore. On return, the dashboard and every app came back on their own.
|
||||
2. **What the box did by itself.** Everything. Power returned 21:26:01Z; containers went 0 → **25 at
|
||||
t+131s** → **26 at t+148s**, with no shell, no help and nothing stuck. **gokapi — the app that was
|
||||
being restored when the plug came out — came back `Up 30 seconds (healthy)`.**
|
||||
3. **Time to steady.** **148 seconds** (back to 26 of 26 containers, measured against the count this
|
||||
round took itself before the accident).
|
||||
4. **Alarm fired / true?** `controller_started` (info) at 21:28 — true, and correct: a power cut
|
||||
inside the 30-minute staleness window is not a liveness alarm, so the ladder expects exactly this
|
||||
one line and nothing else. **No false alarm fired.**
|
||||
5. **Should have fired and did not.** None. Per `08-alarm-ladder.md` a 60-second outage produces no
|
||||
`node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace covers the restart) —
|
||||
and neither appeared.
|
||||
|
||||
**Household loop:** NOT COLLECTED for this round — the loop died with the box (see above). Fixed for
|
||||
every round after this one.
|
||||
|
||||
## A real finding this round handed me: the OOM was DETECTED
|
||||
Sep 16 21:20 warning **app_oom**
|
||||
„Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította"
|
||||
|
||||
That is immich's `CONNECTION_CLOSED immich-postgres:5432` explained: its Postgres was killed by the
|
||||
memory limit during the reverse-geocoding import, the connection died mid-query, the metadata service
|
||||
failed, and the worker exited — twelve times.
|
||||
|
||||
**Why it is worth writing down:** this project carries a standing finding that OOM signals are
|
||||
invisible inside LXC guests (`OOMKilled` false, no docker oom events). Here the controller's own
|
||||
scan DID catch it, named the app AND the exact container, and said in plain Hungarian that a process
|
||||
was stopped by the memory limit. It is an operator-only warning by register, so no customer mail —
|
||||
also correct. **The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm
|
||||
feed, correctly labelled, the whole time.**
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
2026-09-16T21:29:49Z ACCIDENT=disk-95-full round=3
|
||||
2026-09-16T21:29:49Z filling the customer guest's SYSTEM disk to ~95 % — with a POOL GUARD
|
||||
2026-09-16T21:29:51Z pool now: 39% used; guest / has 29352 MB free
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 944M 29G 4% /
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 29G 1.5G 96% /
|
||||
2026-09-16T21:29:55Z pool after the fill: 39.69
|
||||
2026-09-16T21:29:55Z full — holding 10 minutes
|
||||
@@ -0,0 +1,8 @@
|
||||
2026-09-16T21:29:45Z ================ ROUND 3 : use bookstack, while: disk-95-full ================
|
||||
2026-09-16T21:29:48Z --- BEFORE --- containers=26 wiki=200 status=200 paste=200
|
||||
2026-09-16T21:29:48Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T21:29:48Z --- ACTION: use on bookstack ---
|
||||
2026-09-16T21:29:48Z wiki read 1 -> 200
|
||||
2026-09-16T21:29:49Z wiki read 2 -> 200
|
||||
2026-09-16T21:29:49Z wiki read 3 -> 200
|
||||
2026-09-16T21:29:49Z --- ACCIDENT: disk-95-full (injected after the action started) ---
|
||||
Reference in New Issue
Block a user