diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index d361e3af..8d9bf0a1 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -408,6 +408,13 @@ reconciles)"). **This is not the event path.** A dropped event is a lost *fact*, a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop behaviour is **still unmeasured** after three rounds of internet cuts. +**And it came back by itself, on schedule.** The very next scheduled report went through — +**23:23:43Z, „Hub report pushed successfully (15354 bytes)”** — exactly 15 minutes after the +cycle that failed, and nothing was done to the box to achieve it. The reading was taken at +23:24:21Z, deliberately *after* the report was due, so it could not be premature. So the whole +shape of a hub outage is now measured end to end: build → three attempts → give up → +keep serving → next cycle succeeds → no alarm, no loss. + **The near-miss, which is luck and not design.** Last good report 22:53:43Z; next scheduled 23:23:42Z; `node_stale` trips at 30 minutes. The gap is **29 m 59 s**. The staleness threshold is exactly twice the report cadence, so **a single failed push spends the entire budget** — one second diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-1-notes.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-1-notes.txt index 94d1bb9a..52328e4a 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-1-notes.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-1-notes.txt @@ -196,3 +196,22 @@ mine, with different cures, and it took separating them to fix either. Household loop through all of this: unit active, **4 sampled apps since the classifier fix, 0 failures** — the box kept serving its other apps while five were being torn down and rebuilt. + +## A LATE GHOST FROM THIS ROUND, recorded 2026-09-17T01:26 CEST (23:26Z) +Round 1's off-site poll loop reported "completed, exit code 0" TWO HOURS after it stopped working. + last real output line: 2026-09-16T21:24:56Z + task actually ended: 2026-09-16T23:24:56Z, with + "Read from remote host 192.168.0.115: Connection reset by peer" + "client_loop: send disconnect: Broken pipe" +It went silent during round 2's power cut (the box it was polling was abruptly stopped), and the +SSH session then hung, producing nothing, until TCP reset it two hours later. + +Three things worth keeping, all of them this project's recurring classes: + 1. A HUNG connection and a FINISHED one look identical from outside: both produce no new output. + Silence is not completion. The loop's own timestamps are what distinguish them, which is why + every poll line carries one. + 2. The task EXITED 0 while its connection had been reset. Another instance of "exit codes that + lie" - the exit status described the local shell, not the remote work. + 3. A completion notice arriving two hours late could easily be read as a FRESH result for + whatever round is running now. It was checked against its own content before being believed. +It did nothing to the box after 21:24:56Z, so no round is contaminated. Recorded, not hidden. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-9.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-9.txt index 380adaac..699a7c6b 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-9.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-9.txt @@ -72,3 +72,13 @@ injector is left as it is and this caveat travels with the three rounds that use Last good report 22:53:43Z. Next scheduled 23:23:42Z. node_stale trips at 30 minutes. The gap is 29 m 59 s. One second more and the operator would have been paged for a box that was perfectly healthy and had already repaired itself. That is worth a register row. + +### THE RECOVERY, confirmed by a reading that waited for its own precondition (23:24:21Z) + 2026-09-16T23:23:43.762Z [INFO] [report] Hub report pushed successfully (15354 bytes) +The very next scheduled report went through - exactly 15 minutes after the cycle that failed, and +about 13 minutes after the link returned. Nothing was done to the box to achieve this. +The reading was taken at 23:24:21Z, deliberately AFTER the 23:23:42Z due time, so it could not be +premature. That is the sixth-mistake fix working as intended: the measurement states its own +precondition instead of me judging the clock by eye. +So the full shape of a hub outage on this product is now measured end to end: + build -> 3 attempts -> give up -> keep serving -> next cycle succeeds -> no alarm, no loss.