2026-09-16T21:53:25Z ================ ROUND 5 : use privatebin, while: docker-restart ================
2026-09-16T21:53:27Z --- BEFORE --- containers=26  paste=200  status=200  paste=200
2026-09-16T21:53:27Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:53:27Z --- ACTION: use on privatebin ---
2026-09-16T21:53:27Z     paste read 1 -> 200
2026-09-16T21:53:27Z     paste read 2 -> 200
2026-09-16T21:53:27Z     paste read 3 -> 200
2026-09-16T21:53:27Z --- ACCIDENT: docker-restart (injected after the action started) ---
    2026-09-16T21:53:27Z ACCIDENT=docker-restart round=5
    2026-09-16T21:53:27Z systemctl restart docker inside the customer guest
    2026-09-16T21:53:41Z docker restarted
    2026-09-16T21:53:41Z accident docker-restart complete
2026-09-16T21:53:41Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T21:53:43Z     t+16s containers=26 (before 26)
2026-09-16T21:53:43Z     STEADY after 16s
2026-09-16T21:53:43Z     front doors: paste=404  status=404  paste=404  wiki=404
2026-09-16T21:53:44Z     household lines this round: 0  failures: 0
2026-09-16T21:53:44Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
  | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
  | Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
  | Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
2026-09-16T21:53:45Z ================ END ROUND 5 ================

## The controller DID restart with docker — established before judging the alarm
    docker restart ran          21:53:27Z -> 21:53:41Z
    felhom-controller StartedAt **2026-09-16T21:53:38.397Z**  RestartCount=0  Status=running
    traefik           StartedAt 2026-09-16T21:53:38.634Z
    controller log:
        21:54:37 [bootrecon] boot window: fleet settled after 51s (3 identical samples 5s apart) — sweeping
        21:54:37 [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
        21:54:42 [stacks] Status refresh: 25 containers across 56 stacks

So the controller went down and came back with the docker daemon, took 51 seconds to decide the
fleet had settled, and found **nothing boot-orphaned to repair** — which is the right answer, because
every container had already come back on its own restart policy.

**A trap I nearly walked into, recorded because it would have been the eleventh tonight.** Round 5's
own alarm snapshot was taken at **21:53:44Z — six seconds after the controller started**. Concluding
„`controller_started` did not fire" from that snapshot would have been a measurement taken before the
thing it was measuring could have happened. The verdict is therefore taken from a later reading, not
from the round's own dump.

## ROUND 5 — the five things
**action:** `use` privatebin (three reads through its own front door, all 200)
**accident:** `systemctl restart docker` inside the guest — every container down at once

1. **What the customer saw.** A gap of well under a minute. The apps answered 200 before the restart;
   two seconds after docker returned they were **404** (traefik had not re-registered routes yet), and
   by 21:54:19Z they were serving again — LAN **301**, public **200**. Call it ~40 seconds of doors
   being shut, with no error page beyond a plain 404.
2. **What the box did by itself.** Everything. `systemctl restart docker` ran 21:53:27Z → 21:53:41Z;
   **all 26 containers were back by t+16s**, none in a non-Up state. The controller came back with
   them (StartedAt 21:53:38Z), waited 51 s for the fleet to settle, and found **nothing
   boot-orphaned to repair** — correct, because every container had already returned on its own
   restart policy.
3. **Time to steady.** **16 seconds** to 26 of 26 containers; ~40 seconds until the front doors served
   again. The slower of the two numbers is the one a household would feel.
4. **Alarm fired / true?** **`controller_started` (info) at 21:53 — true and correct.** The controller
   really did restart, and the ladder expects exactly this one line: no `app_start_failed` (the 90 s
   boot grace covers a restart) and no liveness alarm (far inside the 30-minute window).
5. **Should have fired and did not.** **None.**

**Household loop: NOT SAMPLED.** It logged 0 lines in this round — not because nothing failed, but
because the round lasted ~18 seconds and the loop samples every 2 minutes. Recorded as not sampled,
never as a pass, exactly as round 2's was.

**A trap avoided, and it would have been my eleventh:** the round's own alarm snapshot was taken at
21:53:44Z, **six seconds after the controller started**. From that snapshot `controller_started`
looked missing. A later reading at 21:55:18Z shows it present at 21:53. The verdict comes from the
later reading; an alarm cannot be called missing by a measurement taken before it could have fired.
