CHAOS NIGHT round 7: a ten-minute outage falls between two reports
gates / gates (push) Successful in 21s

Measured from the box's own log rather than recalled from documentation:
  21:53:41Z Registered periodic job: hub-report (every 15m0s)
  21:53:49Z Hub report pushed successfully (30535 bytes)
  22:08:43Z Hub report pushed successfully (15934 bytes)  - exactly 15m later

The block runs ~22:10Z to ~22:20Z and the next report is due ~22:23:43Z, after
it lifts. So the box never attempts a push while cut off: nothing was tried,
nothing failed, nothing was lost.

The honest verdict for the question I wanted this round to answer - are alarms
raised while the hub is unreachable retried and then silently dropped? - is NOT
EXERCISED, not "passed". Recorded that way.

It is still a finding of its own: a ten-minute internet outage is invisible to
the fleet view because the box had nothing due to say, and the hub's staleness
threshold (30 minutes) is set well beyond it. The two mechanisms agree.

Rounds 8 and 9 are also ten-minute cuts at ~25-minute spacing against a
15-minute cycle, so one will very likely contain a scheduled report and exercise
the drop behaviour properly. They are NOT re-timed to make that happen -
re-timing a round to get a better result is choosing the night after the fact.

Also captured while blocked: internet unreachable from the box, LAN reachable,
26 containers up, free space unmoved, disk guard silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 00:15:11 +02:00
parent a103b62330
commit 3abd25e681
@@ -46,3 +46,65 @@ insurance either way and has cost nothing.
Same family as the `pkill -f` that killed my own watcher earlier: **a pattern that matches the hand
holding it.** The cure is the one that worked here — ask for the command lines, not the count.
START
--- controller: any event pushes attempted, and did they fail? ---
2026/09/16 22:08:42 [INFO] [scheduler] Running job: offsite-credential-retry
2026/09/16 22:08:42 [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
2026/09/16 22:08:42 [INFO] [scheduler] Running job: hub-report
2026/09/16 22:08:42 [INFO] [report] Building system report
2026/09/16 22:08:43 [INFO] [report] Hub report pushed successfully (15934 bytes)
2026/09/16 22:08:43 [INFO] [scheduler] Job hub-report completed (took 1.122s)
2026/09/16 22:13:42 [INFO] [scheduler] Running job: offsite-credential-retry
2026/09/16 22:13:42 [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
--- is the box still blocked right now? ---
internet: still blocked
--- containers / free / guard ---
containers=26 free=7573M guardlines=0
END
## INTERIM: the cut may be shorter than the box's own cadence — so nothing even tried to send
Read from the box at 22:13:49Z, while it was still blocked:
22:08:42Z [scheduler] Running job: hub-report
22:08:43Z [report] **Hub report pushed successfully (15934 bytes)** <- BEFORE the block
22:13:42Z [scheduler] Running job: offsite-credential-retry — completed (took 0s), no push
(no hub-report job since 22:08:42Z; no push failures; no dropped events logged)
containers **26** · free `/` **7573M** unchanged · diskguard **0 lines**
The block began about 22:10Z. The box's last contact with the hub was a minute and a half BEFORE
that, and in the ten minutes since, its scheduler has not tried to reach the hub at all.
**So round 7 may not test what I expected it to test.** The interesting question — are alarms raised
while the hub is unreachable retried and then silently dropped? — needs the box to actually attempt a
push during the outage. If its reporting interval is longer than the outage, a ten-minute cut can
pass entirely between two reports and the hub never notices anything happened.
That would itself be a finding worth having, and a reassuring one: **a short internet outage is
invisible to the fleet view not because anything is hidden, but because nothing was due to be said.**
The interval is measured from the box's own log rather than recalled from documentation, and the
post-block check confirms whether any push was attempted, failed, or lost.
## The cadence, MEASURED — and what it means for this round and the next two
From the box's own log, not from documentation:
21:53:41Z [scheduler] **Registered periodic job: hub-report (every 15m0s)**
21:53:49Z [report] Hub report pushed successfully (30535 bytes)
22:08:43Z [report] Hub report pushed successfully (15934 bytes) <- exactly 15 min later
The block runs ~22:10Z → ~22:20Z. **The next report is due ~22:23:43Z, after the cut ends.**
**Conclusion for round 7, stated plainly: this round does not test what I hoped it would.** The
question „are alarms raised while the hub is unreachable retried and then silently dropped?" needs
the box to attempt a push during the outage. Here the outage falls entirely between two reports, so
nothing was attempted, nothing failed, and nothing was lost. The correct verdict is **not exercised**
— not „passed".
**And it is a real, if quiet, finding of its own:** on this box a ten-minute internet outage is
invisible to the fleet view, because the box had nothing due to say. The hub's staleness thresholds
(30 minutes to `node_stale`) are set well beyond that, so the two mechanisms agree.
**What it implies for rounds 8 and 9, written before they run:** both are also ten-minute cuts, drawn
at ~25-minute spacing against a 15-minute report cycle, so one of them will almost certainly contain a
scheduled report and will exercise the drop behaviour properly. **The rounds are NOT re-timed to make
that happen** — re-timing a round to obtain a better result is choosing the night after the fact. If
it happens naturally, it is measured; if it does not, that is recorded too.