  round 9 armed for 2026-09-16T23:00:48Z (use uptime-kuma + internet cut, hub now blocked too)
=== round 9 launched 2026-09-16T23:00:50Z (due 23:00:48Z) ===
2026-09-16T23:00:50Z ================ ROUND 9 : use uptime-kuma, while: internet-gone-10min ================
2026-09-16T23:00:52Z --- BEFORE --- containers=26  status=200  status=200  paste=200
2026-09-16T23:00:52Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T23:00:52Z --- ACTION: use on uptime-kuma ---
2026-09-16T23:00:52Z     status read 1 -> 200
2026-09-16T23:00:53Z     status read 2 -> 200
2026-09-16T23:00:53Z     status read 3 -> 200
2026-09-16T23:00:53Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---
    2026-09-16T23:00:53Z ACCIDENT=internet-gone-10min round=9
    2026-09-16T23:00:54Z blocking the box traffic off-LAN at the HOST, on tap336i0; the LAN stays up EXCEPT the hub (192.168.0.192)
    2026-09-16T23:00:54Z blocked (LAN allowed EXCEPT the hub at 192.168.0.192, everything else dropped) - 10 minutes
    2026-09-16T23:10:54Z unblocked; host sysctl restored to 0 and both rules removed
    -P FORWARD ACCEPT
    2026-09-16T23:10:55Z accident internet-gone-10min complete
2026-09-16T23:10:55Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T23:10:56Z     t+603s containers=26 (before 26)
2026-09-16T23:10:56Z     STEADY after 603s
2026-09-16T23:10:56Z     front doors, FIRST reading at 2026-09-16T23:10:56Z - TOO EARLY to trust if the accident just ended:
2026-09-16T23:10:57Z       status=200  status=200  paste=200  wiki=200
2026-09-16T23:11:57Z     front doors, SECOND reading at 2026-09-16T23:11:57Z, 60 s later - THIS is the one to trust:
2026-09-16T23:11:58Z       status=200  status=200  paste=200  wiki=200
2026-09-16T23:11:59Z     household lines this round: 12  failures: 0
2026-09-16T23:11:59Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
  | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
  | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
2026-09-16T23:12:00Z ================ END ROUND 9 ================

[exited with code 0]

## THE MEASUREMENT THE NIGHT WAS MISSING (2026-09-16T23:12Z)

With the injector fixed, the hub really was unreachable this time. The controller's own log:

  23:08:42  [INFO]  [scheduler] Running job: hub-report
  23:08:42  [INFO]  [report] Building system report
  23:10:23  [WARN]  [report] Push failed: Post "https://hub.felhom.eu/api/v1/report":
                    context deadline exceeded (Client.Timeout exceeded while awaiting headers)
  23:10:23  [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts
                    (took 1m40.813s)

So the behaviour is now measured, not assumed:
  * the report was built, attempted THREE times, and then GIVEN UP after 1 m 40.8 s;
  * it gave up at 23:10:23, THIRTY-ONE SECONDS before the link came back at 23:10:54;
  * nothing was queued. The next report is simply the next scheduled one.
The earlier WARN says so in plain words: "backing off (the 15-min cycle still reconciles)".
That is the design: the report is a SNAPSHOT, so a lost one costs nothing - the next snapshot
carries the same truth. It is NOT the same as the event path, where a dropped event is a lost
FACT. No event happened to be raised during this cut, so the event-drop path is STILL unmeasured.

### A SECOND fidelity fault in my accident, recorded like the first.
The cut also severed the controller from its own HOST AGENT:
  23:03:56 / 23:08:56 [ERROR] [quiesce] cycle error: check due: agentapi:
           GET /backup/tiers: Get "https://169.254.253.1:8443/backup/tiers": context deadline exceeded
169.254.253.1 is the link-local address of the agent on the host side of the same tap. My blanket
DROP is "everything not in 192.168.0.0/24", so it took the agent link with it.
In a real house an ISP outage does NOT cut the controller from the agent - they sit on one machine.
So "internet gone" as injected is BROADER than its name: it removes the internet, the hub, AND the
local host agent. Stated so nobody reads more into rounds 7-9 than was actually tested.
No further internet cuts are drawn (rounds 10-12 are hard reset, drive pulled, nothing), so the
injector is left as it is and this caveat travels with the three rounds that used it.

### How close the staleness alarm came, by luck, not by design.
Last good report 22:53:43Z. Next scheduled 23:23:42Z. node_stale trips at 30 minutes.
The gap is 29 m 59 s. One second more and the operator would have been paged for a box that was
perfectly healthy and had already repaired itself. That is worth a register row.

### THE RECOVERY, confirmed by a reading that waited for its own precondition (23:24:21Z)
  2026-09-16T23:23:43.762Z [INFO] [report] Hub report pushed successfully (15354 bytes)
The very next scheduled report went through - exactly 15 minutes after the cycle that failed, and
about 13 minutes after the link returned. Nothing was done to the box to achieve this.
The reading was taken at 23:24:21Z, deliberately AFTER the 23:23:42Z due time, so it could not be
premature. That is the sixth-mistake fix working as intended: the measurement states its own
precondition instead of me judging the clock by eye.
So the full shape of a hub outage on this product is now measured end to end:
  build -> 3 attempts -> give up -> keep serving -> next cycle succeeds -> no alarm, no loss.
