2026-09-16T21:24:31Z ================ ROUND 2 : restore gokapi, while: power cut ================
2026-09-16T21:24:31Z --- BEFORE: steady? ---
2026-09-16T21:24:32Z     containers before: 26
2026-09-16T21:24:32Z --- ACTION: restore gokapi from its local backup ---
  gokapi BEFORE: gokapi Up 8 minutes (healthy)
  --- the restore, through the page's own form ---
  POST /backup/restore -> HTTP/1.1 302 Found
  flash: Visszaállítás elindult — az állapot itt frissül.
2026-09-16T21:24:36Z --- ACCIDENT in 20 s: the plug comes out mid-restore ---
## PRE-ROUND-2 STATE — the household the accidents will actually hit (2026-09-16T21:24:16Z)
    wiki 200 · paste 200 · share 200 · recipes 200 · inventory 200 · status 200
    travel 200 · paperless 200 · cloud 200 · media 200 · vault 200
    **photos 404 — immich, crash-looping on `CONNECTION_CLOSED immich-postgres:5432`**
    control: nosuchapp -> 404 (so a 200 above means a real route to a real app, not a stray default)

**11 of 12 serving.** Immich is a KNOWN PRE-EXISTING CONDITION from here on: it was already broken
before this accident landed, no round in the drawn schedule acts on it, and no later failure of it
may be attributed to an accident.

Everything else was rebuilt from my own damage and is verified at the front door rather than by a
health badge — including the two that were dead an hour ago (`cloud`/nextcloud and `share`/gokapi)
and `wiki`/bookstack, which needed a properly formed APP_KEY.
2026-09-16T21:24:56Z     qm stop 336 (no clean shutdown)
2026-09-16T21:24:58Z     dark for 60 s
2026-09-16T21:26:01Z     power back on
2026-09-16T21:26:01Z --- how long until the box is steady again, by itself ---
2026-09-16T21:26:18Z     t+17s containers=0
2026-09-16T21:27:55Z     t+114s containers=0

## THE BACKGROUND HOUSEHOLD DIED WITH THE BOX — found at round 2, fixed for the rest of the night
The loop's last line is **21:23:25Z**; the power cut was at **21:24:56Z**; afterwards
`systemctl is-active household` = **inactive**.

The reason is mine: the loop ran on the VM as a `systemd-run --collect` **transient** unit, and a
transient unit does not survive the machine being switched off. Round 2's accident is a power cut, so
the very first accident that could produce household failures instead produced **no lines at all**.

**Zero lines is not the same as zero failures, and it must not be scored as one.** For round 2 the
household measure is: *not collected* — the loop was dead from 21:24:56Z until it was reinstalled.

**Fixed for every round from here:** `/etc/systemd/system/household.service` with `Restart=always`
and `WantedBy=multi-user.target`, enabled — so it comes back with the box after a power cut, a hard
reset or a docker restart, which rounds 10 and 5 will also deliver. The log carries a marker line at
the changeover so the two regimes are not confused.

This is the eighth thing tonight I have had to fix in my own measuring apparatus rather than in the
product, and the same shape as the rest: **the instrument was silent and silence looked like a pass.**
2026-09-16T21:28:12Z     t+131s containers=25
2026-09-16T21:28:29Z     t+148s containers=26
2026-09-16T21:28:29Z     STEADY after 148s (back to 26 of 26 containers)
2026-09-16T21:28:29Z --- what the box says about gokapi and the restore ---
gokapi Up 30 seconds (healthy)
2026-09-16T21:28:31Z --- alarms in this round ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
  | Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
  | Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
  | Sep 16 21:19 | info | app_deploy_started | Alkalmazás telepítése elindult: Immich | controller
  | Sep 16 21:19 | info | app_removed | Alkalmazás eltávolítva: immich | controller
2026-09-16T21:28:32Z ================ END ROUND 2 ================

## ROUND 2 — the five things
**action:** `restore` gokapi from its local backup · **accident:** power cut, 20 s into the restore

1. **What the customer saw.** „Visszaállítás elindult — az állapot itt frissül." and then the box went
   dark mid-restore. On return, the dashboard and every app came back on their own.
2. **What the box did by itself.** Everything. Power returned 21:26:01Z; containers went 0 → **25 at
   t+131s** → **26 at t+148s**, with no shell, no help and nothing stuck. **gokapi — the app that was
   being restored when the plug came out — came back `Up 30 seconds (healthy)`.**
3. **Time to steady.** **148 seconds** (back to 26 of 26 containers, measured against the count this
   round took itself before the accident).
4. **Alarm fired / true?** `controller_started` (info) at 21:28 — true, and correct: a power cut
   inside the 30-minute staleness window is not a liveness alarm, so the ladder expects exactly this
   one line and nothing else. **No false alarm fired.**
5. **Should have fired and did not.** None. Per `08-alarm-ladder.md` a 60-second outage produces no
   `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace covers the restart) —
   and neither appeared.

**Household loop:** NOT COLLECTED for this round — the loop died with the box (see above). Fixed for
every round after this one.

## A real finding this round handed me: the OOM was DETECTED
    Sep 16 21:20  warning  **app_oom**
      „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította"

That is immich's `CONNECTION_CLOSED immich-postgres:5432` explained: its Postgres was killed by the
memory limit during the reverse-geocoding import, the connection died mid-query, the metadata service
failed, and the worker exited — twelve times.

**Why it is worth writing down:** this project carries a standing finding that OOM signals are
invisible inside LXC guests (`OOMKilled` false, no docker oom events). Here the controller's own
scan DID catch it, named the app AND the exact container, and said in plain Hungarian that a process
was stopped by the memory limit. It is an operator-only warning by register, so no customer mail —
also correct. **The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm
feed, correctly labelled, the whole time.**
