From eb638d303bce77c22f19780cfb6ee2ebda3a6fd4 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 23:56:26 +0200 Subject: [PATCH] CHAOS NIGHT rounds 4-5: the box repairs its own tunnel, and survives docker dying Round 4 (tunnel killed 10 min): cloudflared came back in ~97 seconds and the CONTROLLER did it, not Docker - RestartCount=0 proves the unless-stopped policy never acted, and the log shows the protected-infra recovery redeploying it. health_critical fired at 21:43 and health_recovered closed it at 21:48, so the alarm was not a dead end. Round 4's drawn ACTION never ran, and that is recorded rather than re-run: the round-2 power cut rebooted the guest, /tmp is cleared on boot, and the dashboard password lived there. The off-site run was never triggered. Re-running a round after watching it fail is how a drill starts choosing its own results. The file now lives in /root, so round 10's hard reset cannot disarm rounds 6, 8 and 10. Round 5 (docker restarted): all 26 containers back in 16 seconds, doors serving again in ~40, controller_started true and correct, nothing missed. The household loop is recorded as NOT SAMPLED - 0 lines because the round was shorter than its 2-minute sampling interval, which is not the same as 0 failures. Two traps avoided and written down: round 5's own alarm snapshot was taken six seconds after the controller started, from which controller_started looked missing (it fired); and my check for the password file put its redirect on the wrong host, reporting "not there" for a file that was present all along. Tenth and eleventh of the same family tonight - a check whose own precondition was wrong, answering confidently. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/DRILL-chaos-night-2026-09-17.md | 40 +++++++- .../round-4-accident.txt | 1 + .../round-4.txt | 98 +++++++++++++++++++ .../round-5-accident.txt | 4 + .../round-5.txt | 76 ++++++++++++++ .../round-6.txt | 4 + .../evidence-chaos-night-2026-09-17/round4.sh | 2 +- .../evidence-chaos-night-2026-09-17/round6.sh | 2 +- .../run_round.sh | 2 +- 9 files changed, 225 insertions(+), 4 deletions(-) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-5-accident.txt create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-5.txt create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 6f23edb3..9c7af08e 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -251,7 +251,45 @@ estimating the clock instead of reading it. The readings were unchanged, but „ minutes in" and „nothing in the whole window" are different findings. From here the end-of-window check is taken when the round's own runner reports completion. -### Rounds 4-12 +### Round 4 — `offsite-run` (mealie) / accident: **tunnel killed for ten minutes** + +**The drawn ACTION never ran.** The runner aborted with „no session — mine, not the product's": +round 2's power cut had rebooted the guest, `/tmp` is cleared on boot, and the dashboard password +file lived there. Round 3 was a `use` round and never needed it, so round 4 was the first to find it +gone. **The accident was measured; the off-site run was not.** Recorded as half-measured rather than +re-run and presented as whole — re-running a round after seeing it fail is how a drill starts +choosing its own results. The file now lives in `/root`, which survives a reboot. + +| the five things | | +|---|---| +| what the customer saw | from outside, the apps vanished for ~90 s (public route **530**) and came back on their own (**200**); from inside the house, nothing — traefik answered **301** throughout | +| what the box did by itself | **repaired its own tunnel.** cloudflared killed 21:41:33Z, running again **21:43:07.478Z (~97 s)**, with `RestartCount=0` — so Docker's `unless-stopped` policy did *not* do it; the controller's protected-infra recovery redeployed it („[infra] deploying cloudflared →…") | +| time to steady | the apps never stopped; the way IN was restored in **~97 s** | +| alarm fired / true? | **two, correctly paired** — `health_critical` (error) 21:43, `health_recovered` (info) 21:48. Exactly what the ladder predicts for a missing protected container, and **the alarm was not a dead end** | +| should have fired, did not | none for the accident | + +**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop +does not follow redirects, so it measures „is the app serving on the box", never „can the household +reach it from outside". It saw nothing while the public route was returning 530. + +### Round 5 — `use` privatebin / accident: **docker restarted inside the guest** + +| the five things | | +|---|---| +| what the customer saw | a gap well under a minute: 200 before, **404** two seconds after docker returned (traefik had not re-registered routes), serving again by 21:54:19Z — about **40 s** of shut doors | +| what the box did by itself | everything. Restart ran 21:53:27Z→21:53:41Z; **all 26 containers back at t+16s**; the controller returned with them, waited 51 s for the fleet to settle and found **nothing boot-orphaned to repair** — correct, since every container had already come back on its own policy | +| time to steady | **16 s** to 26 of 26 containers; **~40 s** until the doors served. The slower number is the one a household feels | +| alarm fired / true? | **`controller_started` (info) — true and correct**, and the only line the ladder expects: no `app_start_failed` (90 s boot grace), no liveness alarm | +| should have fired, did not | **none** | + +**Household loop: NOT SAMPLED** — 0 lines, because the round lasted ~18 s and the loop samples every +2 minutes. Recorded as not sampled, never as a pass. + +**A trap avoided.** The round's own alarm snapshot was taken **six seconds** after the controller +started, and from it `controller_started` looked missing. A later reading shows it present at 21:53. +An alarm cannot be called missing by a measurement taken before it could have fired. + +### Rounds 6-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-4-accident.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-4-accident.txt index 4d8a3923..396f7944 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-4-accident.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-4-accident.txt @@ -1,3 +1,4 @@ 2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4 2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest 2026-09-16T21:41:33Z tunnel killed — 10 minutes +2026-09-16T21:51:33Z 10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt index 7f63de11..f6bf2331 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-4.txt @@ -72,3 +72,101 @@ true and answers a narrower question than it appears to. This is the same distinction as the front-door method correction above, and it is recorded in both places because the number and the label live in different files. + 2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4 + 2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest + cloudflared + 2026-09-16T21:41:33Z tunnel killed — 10 minutes + 2026-09-16T21:51:33Z 10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement + /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh: line 109: unexpected EOF while looking for matching `"' +2026-09-16T21:51:33Z --- AFTER: what the box did BY ITSELF --- +2026-09-16T21:51:35Z t+603s containers=26 (before 26) +2026-09-16T21:51:35Z STEADY after 603s +2026-09-16T21:51:36Z front doors: recipes=200 status=200 paste=200 wiki=200 +2026-09-16T21:51:36Z household lines this round: 10 failures: 0 +2026-09-16T21:51:36Z --- alarms --- + | Time | Severity | Type | Message | Source + | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller + | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller + | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller + | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller + | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller + | Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller + | Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller +2026-09-16T21:51:37Z ================ END ROUND 4 ================ + +## ROUND 4 — the five things (and the action that did NOT run) +**action drawn:** `offsite-run` (app: mealie) · **accident:** tunnel killed for ten minutes + +**THE ACTION NEVER RAN.** The runner reported: + ABORT: no session — mine, not the product's +Cause, traced rather than guessed: **round 2's power cut wiped `/tmp` inside the guest**, and the +dashboard password file lived there. Round 3 was a `use` round and needed only the front door, so +round 4 was the first to touch it and the first to find it missing. The off-site run was therefore +never triggered, and **round 4 measured its accident but not its action.** Recorded as such; it is +not quietly re-run and counted as if it had gone to plan. + +1. **What the customer saw.** From outside: the apps went away for about ninety seconds (public route + **530**) and came back on their own (**200**). From the house: nothing — traefik answered **301** + on the box throughout. +2. **What the box did by itself.** Repaired its own tunnel. cloudflared was killed 21:41:33Z and was + running again at **21:43:07.478Z (~97 s)** with `RestartCount=0` — so Docker's `unless-stopped` + policy did NOT do it; the controller's protected-infra recovery redeployed it + („[infra] deploying cloudflared → /opt/docker/stacks/cloudflared"). 26 containers before and after. +3. **Time to steady.** The apps never stopped; the way IN was restored in **~97 seconds**. + (The runner's „STEADY after 603s" is the elapsed ten-minute window, not a recovery time.) +4. **Alarm fired / true?** **Two, and they pair correctly:** + 21:43 error `health_critical` „Rendszer állapot kritikus (volt: ok)" + 21:48 info `health_recovered` „Rendszer állapot helyreállt: ok (volt: fail)" + The ladder predicts exactly `health_critical` for a missing protected container, and the recovery + closes it — **the alarm was not a dead end.** +5. **Should have fired and did not.** None for the accident. The missing off-site run produced no + alarm because it was never started — my fault, not a silence of the product's. + +**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop +does not follow redirects, so it measures the box, not the way in from outside; see the limit +recorded above. Someone away from home would have met 530 for about ninety seconds. + +## Why round 4's action never ran — the root cause, and what it costs the night +`/tmp` inside the customer guest is cleared on boot. **Round 2's power cut rebooted the guest at +21:26**, so the dashboard password file I had pushed to `/tmp/.pw` during Phase 0 was gone from that +moment. Round 3 was a `use` round and needed only the front door, so nothing noticed. Round 4 was the +first round to need a dashboard session, and it aborted before its action: + ABORT: no session — mine, not the product's + +**What it costs:** round 4's drawn action (`offsite-run`) was not exercised. Its accident was. The +round is recorded as half-measured rather than re-run and presented as whole — the schedule was drawn +before the night and re-running a round after seeing it fail is how a drill starts choosing its own +results. + +**What it changes for the rounds still to come:** rounds 6 (`backup-system`), 8 (`backup-app`) and +10 (`restore`) all need that session. The file is being re-placed in **`/root`**, which is on the +guest's own filesystem and survives a reboot, and the runners now read it from there — so the next +power cut (round 10's hard reset) cannot silently disarm the same three rounds. + +## The injector's second EOF, judged and left alone +`inject.sh: line 109: unexpected EOF` appeared again in this round. `bash -n` passes, and both +branches still to run are plain single commands with no nested quoting: + tunnel-down-10min -> G "docker kill cloudflared" + docker-restart -> G "systemctl restart docker" +The accident itself completed correctly — cloudflared was killed at 21:41:33Z and the recovery was +measured — so the message is noise from the remote shell's re-parse, not a failed injection. Recorded +and not chased further; it cannot affect the remaining rounds, whose branches are clean. + +## The password file IS in place — and my check was the thing that was broken + -rw------- 1 root root 21 Sep 16 21:52 /root/.pw + -rw------- 1 root root 21 Sep 16 21:52 /tmp/.pw +Both present, 21 bytes, mode 600, on the guest's own filesystem. + +My first two verifications reported it missing. Both were wrong in the same way: one wrote +`wc -c < /root/.pw` inside an `ssh` command, so the **redirect was evaluated on the VM** rather than +inside the guest and looked for the file on the wrong machine; the other had its output eaten by the +noise filter I pipe everything through. The push had worked the whole time. + +That is the **tenth** instrument error of tonight and the same family as all the others: a check whose +own precondition was wrong, reporting a confident „not there". The cure that finally worked is the +one this project keeps re-learning — **run the check inside a marker block with no filtering and no +host-side redirection**, so the answer cannot be silently dropped: + pct exec 9201 -- sh -c "echo MARKER_START; ls -l /root/.pw /tmp/.pw; echo MARKER_END" +Rounds 6, 8 and 10 (backup-system, backup-app, restore) are therefore armed, and the copy in `/root` +survives the guest reboot that round 10's hard reset will cause. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-5-accident.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-5-accident.txt new file mode 100644 index 00000000..454c0001 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-5-accident.txt @@ -0,0 +1,4 @@ +2026-09-16T21:53:27Z ACCIDENT=docker-restart round=5 +2026-09-16T21:53:27Z systemctl restart docker inside the customer guest +2026-09-16T21:53:41Z docker restarted +2026-09-16T21:53:41Z accident docker-restart complete diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-5.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-5.txt new file mode 100644 index 00000000..09a7a668 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-5.txt @@ -0,0 +1,76 @@ +2026-09-16T21:53:25Z ================ ROUND 5 : use privatebin, while: docker-restart ================ +2026-09-16T21:53:27Z --- BEFORE --- containers=26 paste=200 status=200 paste=200 +2026-09-16T21:53:27Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) +2026-09-16T21:53:27Z --- ACTION: use on privatebin --- +2026-09-16T21:53:27Z paste read 1 -> 200 +2026-09-16T21:53:27Z paste read 2 -> 200 +2026-09-16T21:53:27Z paste read 3 -> 200 +2026-09-16T21:53:27Z --- ACCIDENT: docker-restart (injected after the action started) --- + 2026-09-16T21:53:27Z ACCIDENT=docker-restart round=5 + 2026-09-16T21:53:27Z systemctl restart docker inside the customer guest + 2026-09-16T21:53:41Z docker restarted + 2026-09-16T21:53:41Z accident docker-restart complete +2026-09-16T21:53:41Z --- AFTER: what the box did BY ITSELF --- +2026-09-16T21:53:43Z t+16s containers=26 (before 26) +2026-09-16T21:53:43Z STEADY after 16s +2026-09-16T21:53:43Z front doors: paste=404 status=404 paste=404 wiki=404 +2026-09-16T21:53:44Z household lines this round: 0 failures: 0 +2026-09-16T21:53:44Z --- alarms --- + | Time | Severity | Type | Message | Source + | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller + | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller + | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller + | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller + | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller + | Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller + | Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller +2026-09-16T21:53:45Z ================ END ROUND 5 ================ + +## The controller DID restart with docker — established before judging the alarm + docker restart ran 21:53:27Z -> 21:53:41Z + felhom-controller StartedAt **2026-09-16T21:53:38.397Z** RestartCount=0 Status=running + traefik StartedAt 2026-09-16T21:53:38.634Z + controller log: + 21:54:37 [bootrecon] boot window: fleet settled after 51s (3 identical samples 5s apart) — sweeping + 21:54:37 [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start) + 21:54:42 [stacks] Status refresh: 25 containers across 56 stacks + +So the controller went down and came back with the docker daemon, took 51 seconds to decide the +fleet had settled, and found **nothing boot-orphaned to repair** — which is the right answer, because +every container had already come back on its own restart policy. + +**A trap I nearly walked into, recorded because it would have been the eleventh tonight.** Round 5's +own alarm snapshot was taken at **21:53:44Z — six seconds after the controller started**. Concluding +„`controller_started` did not fire" from that snapshot would have been a measurement taken before the +thing it was measuring could have happened. The verdict is therefore taken from a later reading, not +from the round's own dump. + +## ROUND 5 — the five things +**action:** `use` privatebin (three reads through its own front door, all 200) +**accident:** `systemctl restart docker` inside the guest — every container down at once + +1. **What the customer saw.** A gap of well under a minute. The apps answered 200 before the restart; + two seconds after docker returned they were **404** (traefik had not re-registered routes yet), and + by 21:54:19Z they were serving again — LAN **301**, public **200**. Call it ~40 seconds of doors + being shut, with no error page beyond a plain 404. +2. **What the box did by itself.** Everything. `systemctl restart docker` ran 21:53:27Z → 21:53:41Z; + **all 26 containers were back by t+16s**, none in a non-Up state. The controller came back with + them (StartedAt 21:53:38Z), waited 51 s for the fleet to settle, and found **nothing + boot-orphaned to repair** — correct, because every container had already returned on its own + restart policy. +3. **Time to steady.** **16 seconds** to 26 of 26 containers; ~40 seconds until the front doors served + again. The slower of the two numbers is the one a household would feel. +4. **Alarm fired / true?** **`controller_started` (info) at 21:53 — true and correct.** The controller + really did restart, and the ladder expects exactly this one line: no `app_start_failed` (the 90 s + boot grace covers a restart) and no liveness alarm (far inside the 30-minute window). +5. **Should have fired and did not.** **None.** + +**Household loop: NOT SAMPLED.** It logged 0 lines in this round — not because nothing failed, but +because the round lasted ~18 seconds and the loop samples every 2 minutes. Recorded as not sampled, +never as a pass, exactly as round 2's was. + +**A trap avoided, and it would have been my eleventh:** the round's own alarm snapshot was taken at +21:53:44Z, **six seconds after the controller started**. From that snapshot `controller_started` +looked missing. A later reading at 21:55:18Z shows it present at 21:53. The verdict comes from the +later reading; an alarm cannot be called missing by a measurement taken before it could have fired. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt new file mode 100644 index 00000000..5ac8d8a0 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt @@ -0,0 +1,4 @@ +2026-09-16T21:54:59Z ================ ROUND 6 : backup-system adventurelog, while: nothing ================ +2026-09-16T21:55:00Z --- BEFORE --- containers=26 travel=200 status=200 paste=200 +2026-09-16T21:55:00Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) +2026-09-16T21:55:01Z --- ACTION: backup-system on adventurelog --- diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round4.sh b/documentation/audits/evidence-chaos-night-2026-09-17/round4.sh index 8beb7794..b3c8806c 100755 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round4.sh +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round4.sh @@ -15,7 +15,7 @@ door(){ curl -sL -o /dev/null -w '%{http_code}' --max-time 12 -k -H "Host: $1.en cat > $SCR/r4a.sh <<'PEOF' #!/bin/bash H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080 -curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/tmp/.pw" $B/login +curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/root/.pw" $B/login S=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/') [ -z "$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; } C="Cookie: felhom_session=$S" diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round6.sh b/documentation/audits/evidence-chaos-night-2026-09-17/round6.sh index 8d1d59b2..2a3675d2 100755 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round6.sh +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round6.sh @@ -14,7 +14,7 @@ door(){ curl -sL -o /dev/null -w '%{http_code}' --max-time 12 -k -H "Host: $1.en cat > $SCR/r6a.sh <<'PEOF' #!/bin/bash H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080 -curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/tmp/.pw" $B/login +curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/root/.pw" $B/login S=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/') [ -z "$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; } C="Cookie: felhom_session=$S" diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh b/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh index 6ae25032..bf4620bb 100755 --- a/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh +++ b/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh @@ -22,7 +22,7 @@ SUB=$(sub "$Y") cat > $SCR/act.sh <