chaos night: the hub link repaired itself on the next cycle, and a late ghost task
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
Round 9's recovery confirmed end to end: 23:23:43Z 'Hub report pushed successfully (15354 bytes)', exactly 15 minutes after the cycle that failed, with nothing done to the box. The reading was deliberately taken after the report was due so it could not be premature. Also recorded: round 1's poll loop reported 'completed, exit 0' two hours after it stopped working. It went silent during round 2's power cut and its SSH hung until TCP reset it. Three familiar classes in one: silence is not completion, the exit code described the local shell not the remote work, and a very late completion notice can be mistaken for a fresh result. It touched nothing after 21:24:56Z, so no round is contaminated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -408,6 +408,13 @@ reconciles)"). **This is not the event path.** A dropped event is a lost *fact*,
|
||||
a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop
|
||||
behaviour is **still unmeasured** after three rounds of internet cuts.
|
||||
|
||||
**And it came back by itself, on schedule.** The very next scheduled report went through —
|
||||
**23:23:43Z, „Hub report pushed successfully (15354 bytes)”** — exactly 15 minutes after the
|
||||
cycle that failed, and nothing was done to the box to achieve it. The reading was taken at
|
||||
23:24:21Z, deliberately *after* the report was due, so it could not be premature. So the whole
|
||||
shape of a hub outage is now measured end to end: build → three attempts → give up →
|
||||
keep serving → next cycle succeeds → no alarm, no loss.
|
||||
|
||||
**The near-miss, which is luck and not design.** Last good report 22:53:43Z; next scheduled
|
||||
23:23:42Z; `node_stale` trips at 30 minutes. The gap is **29 m 59 s**. The staleness threshold is
|
||||
exactly twice the report cadence, so **a single failed push spends the entire budget** — one second
|
||||
|
||||
@@ -196,3 +196,22 @@ mine, with different cures, and it took separating them to fix either.
|
||||
|
||||
Household loop through all of this: unit active, **4 sampled apps since the classifier fix, 0
|
||||
failures** — the box kept serving its other apps while five were being torn down and rebuilt.
|
||||
|
||||
## A LATE GHOST FROM THIS ROUND, recorded 2026-09-17T01:26 CEST (23:26Z)
|
||||
Round 1's off-site poll loop reported "completed, exit code 0" TWO HOURS after it stopped working.
|
||||
last real output line: 2026-09-16T21:24:56Z
|
||||
task actually ended: 2026-09-16T23:24:56Z, with
|
||||
"Read from remote host 192.168.0.115: Connection reset by peer"
|
||||
"client_loop: send disconnect: Broken pipe"
|
||||
It went silent during round 2's power cut (the box it was polling was abruptly stopped), and the
|
||||
SSH session then hung, producing nothing, until TCP reset it two hours later.
|
||||
|
||||
Three things worth keeping, all of them this project's recurring classes:
|
||||
1. A HUNG connection and a FINISHED one look identical from outside: both produce no new output.
|
||||
Silence is not completion. The loop's own timestamps are what distinguish them, which is why
|
||||
every poll line carries one.
|
||||
2. The task EXITED 0 while its connection had been reset. Another instance of "exit codes that
|
||||
lie" - the exit status described the local shell, not the remote work.
|
||||
3. A completion notice arriving two hours late could easily be read as a FRESH result for
|
||||
whatever round is running now. It was checked against its own content before being believed.
|
||||
It did nothing to the box after 21:24:56Z, so no round is contaminated. Recorded, not hidden.
|
||||
|
||||
@@ -72,3 +72,13 @@ injector is left as it is and this caveat travels with the three rounds that use
|
||||
Last good report 22:53:43Z. Next scheduled 23:23:42Z. node_stale trips at 30 minutes.
|
||||
The gap is 29 m 59 s. One second more and the operator would have been paged for a box that was
|
||||
perfectly healthy and had already repaired itself. That is worth a register row.
|
||||
|
||||
### THE RECOVERY, confirmed by a reading that waited for its own precondition (23:24:21Z)
|
||||
2026-09-16T23:23:43.762Z [INFO] [report] Hub report pushed successfully (15354 bytes)
|
||||
The very next scheduled report went through - exactly 15 minutes after the cycle that failed, and
|
||||
about 13 minutes after the link returned. Nothing was done to the box to achieve this.
|
||||
The reading was taken at 23:24:21Z, deliberately AFTER the 23:23:42Z due time, so it could not be
|
||||
premature. That is the sixth-mistake fix working as intended: the measurement states its own
|
||||
precondition instead of me judging the clock by eye.
|
||||
So the full shape of a hub outage on this product is now measured end to end:
|
||||
build -> 3 attempts -> give up -> keep serving -> next cycle succeeds -> no alarm, no loss.
|
||||
|
||||
Reference in New Issue
Block a user