chaos night round 9: what a lost hub report actually costs, measured
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
With the injector corrected, the hub really was unreachable. The controller built its 23:08:42Z report, retried the push three times over 1m40.8s and gave up at 23:10:23Z - 31 seconds before the link returned. Nothing queued, which is correct: a report is a snapshot, not a fact. The box passed. 26 containers throughout, every front door serving, both the hub link and the host-agent link repaired unaided the moment the block lifted, no alarm fired and none should have. R-549 filed (P2): the staleness threshold (30 min) is exactly twice the report cadence (15 min), so ONE failed push spends the entire budget. The measured gap was 29m59s - one second inside the alarm. A healthy, self-repaired box came that close to paging the operator. Also recorded: the injected cut is broader than its name - it severed the controller from its own host agent too, which a real ISP outage would not do. The caveat travels with rounds 7, 8 and 9. The event-drop path remains unmeasured, because no event was raised during any cut. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -379,6 +379,47 @@ public doors **three seconds** after the unblock and reported 530 on all four. T
|
||||
never have been fair. The re-measurement 43 s later read 200 on all four. The runner measured the
|
||||
recovery before the recovery could begin — so the runner is the thing that gets fixed, not the note.
|
||||
|
||||
### Round 9 — `use` uptime-kuma / accident: **internet cut for ten minutes, hub included**
|
||||
|
||||
**23:00:50Z–23:12:00Z.** The first cut of the night that really removed the hub. The injector was
|
||||
corrected between rounds 8 and 9; the drawn action, app and accident were **not** changed.
|
||||
|
||||
| the five things | |
|
||||
|---|---|
|
||||
| what the customer saw | **Nothing at home.** uptime-kuma answered 200 on all three reads during the action, and all four front doors read 200 at both post-accident readings. 26 apps before, 26 after, never fewer. |
|
||||
| what the box did by itself | Kept every container running while it was cut off from the hub **and** from its own host agent. Built its report on schedule, tried to push it **three times over 1 m 40.8 s**, gave up, and carried on serving. Both links repaired themselves the instant the block lifted — no restart, no repair action, nothing from me. |
|
||||
| time to steady | The apps never left steady. The doors read 200 at the first reading, **2 s** after the unblock, and again 63 s later. |
|
||||
| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold. Nothing false was raised. |
|
||||
| should have fired, did not | **none** — but see the near-miss below, which is a finding in its own right. |
|
||||
|
||||
**What a lost report costs: measured, not assumed.**
|
||||
|
||||
```
|
||||
23:08:42 [INFO] [scheduler] Running job: hub-report
|
||||
23:08:42 [INFO] [report] Building system report
|
||||
23:10:23 [WARN] [report] Push failed: … context deadline exceeded
|
||||
23:10:23 [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts (took 1m40.813s)
|
||||
```
|
||||
|
||||
Three attempts, then it stops. Nothing is queued. It gave up **31 seconds before** the link returned.
|
||||
And that is correct: a report is a **snapshot**, so a lost one costs nothing — the next snapshot
|
||||
carries the same truth, and the controller says so itself („backing off (the 15-min cycle still
|
||||
reconciles)"). **This is not the event path.** A dropped event is a lost *fact*, not a stale copy of
|
||||
a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop
|
||||
behaviour is **still unmeasured** after three rounds of internet cuts.
|
||||
|
||||
**The near-miss, which is luck and not design.** Last good report 22:53:43Z; next scheduled
|
||||
23:23:42Z; `node_stale` trips at 30 minutes. The gap is **29 m 59 s**. The staleness threshold is
|
||||
exactly twice the report cadence, so **a single failed push spends the entire budget** — one second
|
||||
of ordinary jitter and the operator is paged about a box that was healthy throughout and had already
|
||||
repaired itself. Filed as a register row.
|
||||
|
||||
**A second fidelity fault in my accident, named like the first.** The cut also severed the controller
|
||||
from its host agent (`GET /backup/tiers` to the link-local `169.254.253.1` timed out twice). In a
|
||||
real house an ISP outage does not do that — controller and agent share one machine. So „internet
|
||||
gone" as injected is **broader than its name**: internet, hub *and* local agent. No further internet
|
||||
cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9.
|
||||
|
||||
### Rounds 8-12
|
||||
|
||||
PENDING
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
round 9 armed for 2026-09-16T23:00:48Z (use uptime-kuma + internet cut, hub now blocked too)
|
||||
=== round 9 launched 2026-09-16T23:00:50Z (due 23:00:48Z) ===
|
||||
2026-09-16T23:00:50Z ================ ROUND 9 : use uptime-kuma, while: internet-gone-10min ================
|
||||
2026-09-16T23:00:52Z --- BEFORE --- containers=26 status=200 status=200 paste=200
|
||||
2026-09-16T23:00:52Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T23:00:52Z --- ACTION: use on uptime-kuma ---
|
||||
2026-09-16T23:00:52Z status read 1 -> 200
|
||||
2026-09-16T23:00:53Z status read 2 -> 200
|
||||
2026-09-16T23:00:53Z status read 3 -> 200
|
||||
2026-09-16T23:00:53Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---
|
||||
2026-09-16T23:00:53Z ACCIDENT=internet-gone-10min round=9
|
||||
2026-09-16T23:00:54Z blocking the box traffic off-LAN at the HOST, on tap336i0; the LAN stays up EXCEPT the hub (192.168.0.192)
|
||||
2026-09-16T23:00:54Z blocked (LAN allowed EXCEPT the hub at 192.168.0.192, everything else dropped) - 10 minutes
|
||||
2026-09-16T23:10:54Z unblocked; host sysctl restored to 0 and both rules removed
|
||||
-P FORWARD ACCEPT
|
||||
2026-09-16T23:10:55Z accident internet-gone-10min complete
|
||||
2026-09-16T23:10:55Z --- AFTER: what the box did BY ITSELF ---
|
||||
2026-09-16T23:10:56Z t+603s containers=26 (before 26)
|
||||
2026-09-16T23:10:56Z STEADY after 603s
|
||||
2026-09-16T23:10:56Z front doors, FIRST reading at 2026-09-16T23:10:56Z - TOO EARLY to trust if the accident just ended:
|
||||
2026-09-16T23:10:57Z status=200 status=200 paste=200 wiki=200
|
||||
2026-09-16T23:11:57Z front doors, SECOND reading at 2026-09-16T23:11:57Z, 60 s later - THIS is the one to trust:
|
||||
2026-09-16T23:11:58Z status=200 status=200 paste=200 wiki=200
|
||||
2026-09-16T23:11:59Z household lines this round: 12 failures: 0
|
||||
2026-09-16T23:11:59Z --- alarms ---
|
||||
| Time | Severity | Type | Message | Source
|
||||
| Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
|
||||
| Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
|
||||
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
|
||||
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
|
||||
2026-09-16T23:12:00Z ================ END ROUND 9 ================
|
||||
|
||||
[exited with code 0]
|
||||
|
||||
## THE MEASUREMENT THE NIGHT WAS MISSING (2026-09-16T23:12Z)
|
||||
|
||||
With the injector fixed, the hub really was unreachable this time. The controller's own log:
|
||||
|
||||
23:08:42 [INFO] [scheduler] Running job: hub-report
|
||||
23:08:42 [INFO] [report] Building system report
|
||||
23:10:23 [WARN] [report] Push failed: Post "https://hub.felhom.eu/api/v1/report":
|
||||
context deadline exceeded (Client.Timeout exceeded while awaiting headers)
|
||||
23:10:23 [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts
|
||||
(took 1m40.813s)
|
||||
|
||||
So the behaviour is now measured, not assumed:
|
||||
* the report was built, attempted THREE times, and then GIVEN UP after 1 m 40.8 s;
|
||||
* it gave up at 23:10:23, THIRTY-ONE SECONDS before the link came back at 23:10:54;
|
||||
* nothing was queued. The next report is simply the next scheduled one.
|
||||
The earlier WARN says so in plain words: "backing off (the 15-min cycle still reconciles)".
|
||||
That is the design: the report is a SNAPSHOT, so a lost one costs nothing - the next snapshot
|
||||
carries the same truth. It is NOT the same as the event path, where a dropped event is a lost
|
||||
FACT. No event happened to be raised during this cut, so the event-drop path is STILL unmeasured.
|
||||
|
||||
### A SECOND fidelity fault in my accident, recorded like the first.
|
||||
The cut also severed the controller from its own HOST AGENT:
|
||||
23:03:56 / 23:08:56 [ERROR] [quiesce] cycle error: check due: agentapi:
|
||||
GET /backup/tiers: Get "https://169.254.253.1:8443/backup/tiers": context deadline exceeded
|
||||
169.254.253.1 is the link-local address of the agent on the host side of the same tap. My blanket
|
||||
DROP is "everything not in 192.168.0.0/24", so it took the agent link with it.
|
||||
In a real house an ISP outage does NOT cut the controller from the agent - they sit on one machine.
|
||||
So "internet gone" as injected is BROADER than its name: it removes the internet, the hub, AND the
|
||||
local host agent. Stated so nobody reads more into rounds 7-9 than was actually tested.
|
||||
No further internet cuts are drawn (rounds 10-12 are hard reset, drive pulled, nothing), so the
|
||||
injector is left as it is and this caveat travels with the three rounds that used it.
|
||||
|
||||
### How close the staleness alarm came, by luck, not by design.
|
||||
Last good report 22:53:43Z. Next scheduled 23:23:42Z. node_stale trips at 30 minutes.
|
||||
The gap is 29 m 59 s. One second more and the operator would have been paged for a box that was
|
||||
perfectly healthy and had already repaired itself. That is worth a register row.
|
||||
Reference in New Issue
Block a user