chaos night round 9: what a lost hub report actually costs, measured
gates / gates (push) Successful in 21s

With the injector corrected, the hub really was unreachable. The controller
built its 23:08:42Z report, retried the push three times over 1m40.8s and
gave up at 23:10:23Z - 31 seconds before the link returned. Nothing queued,
which is correct: a report is a snapshot, not a fact.

The box passed. 26 containers throughout, every front door serving, both the
hub link and the host-agent link repaired unaided the moment the block lifted,
no alarm fired and none should have.

R-549 filed (P2): the staleness threshold (30 min) is exactly twice the report
cadence (15 min), so ONE failed push spends the entire budget. The measured gap
was 29m59s - one second inside the alarm. A healthy, self-repaired box came that
close to paging the operator.

Also recorded: the injected cut is broader than its name - it severed the
controller from its own host agent too, which a real ISP outage would not do.
The caveat travels with rounds 7, 8 and 9. The event-drop path remains
unmeasured, because no event was raised during any cut.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 01:15:19 +02:00
parent 418f3a2c20
commit 889310ec17
3 changed files with 116 additions and 0 deletions
@@ -379,6 +379,47 @@ public doors **three seconds** after the unblock and reported 530 on all four. T
never have been fair. The re-measurement 43 s later read 200 on all four. The runner measured the
recovery before the recovery could begin — so the runner is the thing that gets fixed, not the note.
### Round 9 — `use` uptime-kuma / accident: **internet cut for ten minutes, hub included**
**23:00:50Z–23:12:00Z.** The first cut of the night that really removed the hub. The injector was
corrected between rounds 8 and 9; the drawn action, app and accident were **not** changed.
| the five things | |
|---|---|
| what the customer saw | **Nothing at home.** uptime-kuma answered 200 on all three reads during the action, and all four front doors read 200 at both post-accident readings. 26 apps before, 26 after, never fewer. |
| what the box did by itself | Kept every container running while it was cut off from the hub **and** from its own host agent. Built its report on schedule, tried to push it **three times over 1 m 40.8 s**, gave up, and carried on serving. Both links repaired themselves the instant the block lifted — no restart, no repair action, nothing from me. |
| time to steady | The apps never left steady. The doors read 200 at the first reading, **2 s** after the unblock, and again 63 s later. |
| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold. Nothing false was raised. |
| should have fired, did not | **none** — but see the near-miss below, which is a finding in its own right. |
**What a lost report costs: measured, not assumed.**
```
23:08:42 [INFO] [scheduler] Running job: hub-report
23:08:42 [INFO] [report] Building system report
23:10:23 [WARN] [report] Push failed: … context deadline exceeded
23:10:23 [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts (took 1m40.813s)
```
Three attempts, then it stops. Nothing is queued. It gave up **31 seconds before** the link returned.
And that is correct: a report is a **snapshot**, so a lost one costs nothing — the next snapshot
carries the same truth, and the controller says so itself („backing off (the 15-min cycle still
reconciles)"). **This is not the event path.** A dropped event is a lost *fact*, not a stale copy of
a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop
behaviour is **still unmeasured** after three rounds of internet cuts.
**The near-miss, which is luck and not design.** Last good report 22:53:43Z; next scheduled
23:23:42Z; `node_stale` trips at 30 minutes. The gap is **29 m 59 s**. The staleness threshold is
exactly twice the report cadence, so **a single failed push spends the entire budget** — one second
of ordinary jitter and the operator is paged about a box that was healthy throughout and had already
repaired itself. Filed as a register row.
**A second fidelity fault in my accident, named like the first.** The cut also severed the controller
from its host agent (`GET /backup/tiers` to the link-local `169.254.253.1` timed out twice). In a
real house an ISP outage does not do that — controller and agent share one machine. So „internet
gone" as injected is **broader than its name**: internet, hub *and* local agent. No further internet
cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9.
### Rounds 8-12
PENDING
@@ -0,0 +1,74 @@
round 9 armed for 2026-09-16T23:00:48Z (use uptime-kuma + internet cut, hub now blocked too)
=== round 9 launched 2026-09-16T23:00:50Z (due 23:00:48Z) ===
2026-09-16T23:00:50Z ================ ROUND 9 : use uptime-kuma, while: internet-gone-10min ================
2026-09-16T23:00:52Z --- BEFORE --- containers=26 status=200 status=200 paste=200
2026-09-16T23:00:52Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T23:00:52Z --- ACTION: use on uptime-kuma ---
2026-09-16T23:00:52Z status read 1 -> 200
2026-09-16T23:00:53Z status read 2 -> 200
2026-09-16T23:00:53Z status read 3 -> 200
2026-09-16T23:00:53Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---
2026-09-16T23:00:53Z ACCIDENT=internet-gone-10min round=9
2026-09-16T23:00:54Z blocking the box traffic off-LAN at the HOST, on tap336i0; the LAN stays up EXCEPT the hub (192.168.0.192)
2026-09-16T23:00:54Z blocked (LAN allowed EXCEPT the hub at 192.168.0.192, everything else dropped) - 10 minutes
2026-09-16T23:10:54Z unblocked; host sysctl restored to 0 and both rules removed
-P FORWARD ACCEPT
2026-09-16T23:10:55Z accident internet-gone-10min complete
2026-09-16T23:10:55Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T23:10:56Z t+603s containers=26 (before 26)
2026-09-16T23:10:56Z STEADY after 603s
2026-09-16T23:10:56Z front doors, FIRST reading at 2026-09-16T23:10:56Z - TOO EARLY to trust if the accident just ended:
2026-09-16T23:10:57Z status=200 status=200 paste=200 wiki=200
2026-09-16T23:11:57Z front doors, SECOND reading at 2026-09-16T23:11:57Z, 60 s later - THIS is the one to trust:
2026-09-16T23:11:58Z status=200 status=200 paste=200 wiki=200
2026-09-16T23:11:59Z household lines this round: 12 failures: 0
2026-09-16T23:11:59Z --- alarms ---
| Time | Severity | Type | Message | Source
| Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
| Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
2026-09-16T23:12:00Z ================ END ROUND 9 ================
[exited with code 0]
## THE MEASUREMENT THE NIGHT WAS MISSING (2026-09-16T23:12Z)
With the injector fixed, the hub really was unreachable this time. The controller's own log:
23:08:42 [INFO] [scheduler] Running job: hub-report
23:08:42 [INFO] [report] Building system report
23:10:23 [WARN] [report] Push failed: Post "https://hub.felhom.eu/api/v1/report":
context deadline exceeded (Client.Timeout exceeded while awaiting headers)
23:10:23 [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts
(took 1m40.813s)
So the behaviour is now measured, not assumed:
* the report was built, attempted THREE times, and then GIVEN UP after 1 m 40.8 s;
* it gave up at 23:10:23, THIRTY-ONE SECONDS before the link came back at 23:10:54;
* nothing was queued. The next report is simply the next scheduled one.
The earlier WARN says so in plain words: "backing off (the 15-min cycle still reconciles)".
That is the design: the report is a SNAPSHOT, so a lost one costs nothing - the next snapshot
carries the same truth. It is NOT the same as the event path, where a dropped event is a lost
FACT. No event happened to be raised during this cut, so the event-drop path is STILL unmeasured.
### A SECOND fidelity fault in my accident, recorded like the first.
The cut also severed the controller from its own HOST AGENT:
23:03:56 / 23:08:56 [ERROR] [quiesce] cycle error: check due: agentapi:
GET /backup/tiers: Get "https://169.254.253.1:8443/backup/tiers": context deadline exceeded
169.254.253.1 is the link-local address of the agent on the host side of the same tap. My blanket
DROP is "everything not in 192.168.0.0/24", so it took the agent link with it.
In a real house an ISP outage does NOT cut the controller from the agent - they sit on one machine.
So "internet gone" as injected is BROADER than its name: it removes the internet, the hub, AND the
local host agent. Stated so nobody reads more into rounds 7-9 than was actually tested.
No further internet cuts are drawn (rounds 10-12 are hard reset, drive pulled, nothing), so the
injector is left as it is and this caveat travels with the three rounds that used it.
### How close the staleness alarm came, by luck, not by design.
Last good report 22:53:43Z. Next scheduled 23:23:42Z. node_stale trips at 30 minutes.
The gap is 29 m 59 s. One second more and the operator would have been paged for a box that was
perfectly healthy and had already repaired itself. That is worth a register row.