chaos night: the hub link repaired itself on the next cycle, and a late ghost task
gates / gates (push) Successful in 21s

Round 9's recovery confirmed end to end: 23:23:43Z 'Hub report pushed
successfully (15354 bytes)', exactly 15 minutes after the cycle that failed,
with nothing done to the box. The reading was deliberately taken after the
report was due so it could not be premature.

Also recorded: round 1's poll loop reported 'completed, exit 0' two hours after
it stopped working. It went silent during round 2's power cut and its SSH hung
until TCP reset it. Three familiar classes in one: silence is not completion,
the exit code described the local shell not the remote work, and a very late
completion notice can be mistaken for a fresh result. It touched nothing after
21:24:56Z, so no round is contaminated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 01:26:09 +02:00
parent 70f1e01736
commit f973fd7151
3 changed files with 36 additions and 0 deletions
@@ -408,6 +408,13 @@ reconciles)"). **This is not the event path.** A dropped event is a lost *fact*,
a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop
behaviour is **still unmeasured** after three rounds of internet cuts.
**And it came back by itself, on schedule.** The very next scheduled report went through —
**23:23:43Z, „Hub report pushed successfully (15354 bytes)”** — exactly 15 minutes after the
cycle that failed, and nothing was done to the box to achieve it. The reading was taken at
23:24:21Z, deliberately *after* the report was due, so it could not be premature. So the whole
shape of a hub outage is now measured end to end: build → three attempts → give up →
keep serving → next cycle succeeds → no alarm, no loss.
**The near-miss, which is luck and not design.** Last good report 22:53:43Z; next scheduled
23:23:42Z; `node_stale` trips at 30 minutes. The gap is **29 m 59 s**. The staleness threshold is
exactly twice the report cadence, so **a single failed push spends the entire budget** — one second
@@ -196,3 +196,22 @@ mine, with different cures, and it took separating them to fix either.
Household loop through all of this: unit active, **4 sampled apps since the classifier fix, 0
failures** — the box kept serving its other apps while five were being torn down and rebuilt.
## A LATE GHOST FROM THIS ROUND, recorded 2026-09-17T01:26 CEST (23:26Z)
Round 1's off-site poll loop reported "completed, exit code 0" TWO HOURS after it stopped working.
last real output line: 2026-09-16T21:24:56Z
task actually ended: 2026-09-16T23:24:56Z, with
"Read from remote host 192.168.0.115: Connection reset by peer"
"client_loop: send disconnect: Broken pipe"
It went silent during round 2's power cut (the box it was polling was abruptly stopped), and the
SSH session then hung, producing nothing, until TCP reset it two hours later.
Three things worth keeping, all of them this project's recurring classes:
1. A HUNG connection and a FINISHED one look identical from outside: both produce no new output.
Silence is not completion. The loop's own timestamps are what distinguish them, which is why
every poll line carries one.
2. The task EXITED 0 while its connection had been reset. Another instance of "exit codes that
lie" - the exit status described the local shell, not the remote work.
3. A completion notice arriving two hours late could easily be read as a FRESH result for
whatever round is running now. It was checked against its own content before being believed.
It did nothing to the box after 21:24:56Z, so no round is contaminated. Recorded, not hidden.
@@ -72,3 +72,13 @@ injector is left as it is and this caveat travels with the three rounds that use
Last good report 22:53:43Z. Next scheduled 23:23:42Z. node_stale trips at 30 minutes.
The gap is 29 m 59 s. One second more and the operator would have been paged for a box that was
perfectly healthy and had already repaired itself. That is worth a register row.
### THE RECOVERY, confirmed by a reading that waited for its own precondition (23:24:21Z)
2026-09-16T23:23:43.762Z [INFO] [report] Hub report pushed successfully (15354 bytes)
The very next scheduled report went through - exactly 15 minutes after the cycle that failed, and
about 13 minutes after the link returned. Nothing was done to the box to achieve this.
The reading was taken at 23:24:21Z, deliberately AFTER the 23:23:42Z due time, so it could not be
premature. That is the sixth-mistake fix working as intended: the measurement states its own
precondition instead of me judging the clock by eye.
So the full shape of a hub outage on this product is now measured end to end:
build -> 3 attempts -> give up -> keep serving -> next cycle succeeds -> no alarm, no loss.