CHAOS NIGHT rounds 4-5: the box repairs its own tunnel, and survives docker dying
gates / gates (push) Successful in 22s

Round 4 (tunnel killed 10 min): cloudflared came back in ~97 seconds and the
CONTROLLER did it, not Docker - RestartCount=0 proves the unless-stopped policy
never acted, and the log shows the protected-infra recovery redeploying it.
health_critical fired at 21:43 and health_recovered closed it at 21:48, so the
alarm was not a dead end.

Round 4's drawn ACTION never ran, and that is recorded rather than re-run: the
round-2 power cut rebooted the guest, /tmp is cleared on boot, and the dashboard
password lived there. The off-site run was never triggered. Re-running a round
after watching it fail is how a drill starts choosing its own results. The file
now lives in /root, so round 10's hard reset cannot disarm rounds 6, 8 and 10.

Round 5 (docker restarted): all 26 containers back in 16 seconds, doors serving
again in ~40, controller_started true and correct, nothing missed. The household
loop is recorded as NOT SAMPLED - 0 lines because the round was shorter than its
2-minute sampling interval, which is not the same as 0 failures.

Two traps avoided and written down: round 5's own alarm snapshot was taken six
seconds after the controller started, from which controller_started looked
missing (it fired); and my check for the password file put its redirect on the
wrong host, reporting "not there" for a file that was present all along. Tenth
and eleventh of the same family tonight - a check whose own precondition was
wrong, answering confidently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 23:56:26 +02:00
parent 3e6645413e
commit eb638d303b
9 changed files with 225 additions and 4 deletions
@@ -251,7 +251,45 @@ estimating the clock instead of reading it. The readings were unchanged, but „
minutes in" and „nothing in the whole window" are different findings. From here the end-of-window
check is taken when the round's own runner reports completion.
### Rounds 4-12
### Round 4 — `offsite-run` (mealie) / accident: **tunnel killed for ten minutes**
**The drawn ACTION never ran.** The runner aborted with „no session — mine, not the product's":
round 2's power cut had rebooted the guest, `/tmp` is cleared on boot, and the dashboard password
file lived there. Round 3 was a `use` round and never needed it, so round 4 was the first to find it
gone. **The accident was measured; the off-site run was not.** Recorded as half-measured rather than
re-run and presented as whole — re-running a round after seeing it fail is how a drill starts
choosing its own results. The file now lives in `/root`, which survives a reboot.
| the five things | |
|---|---|
| what the customer saw | from outside, the apps vanished for ~90 s (public route **530**) and came back on their own (**200**); from inside the house, nothing — traefik answered **301** throughout |
| what the box did by itself | **repaired its own tunnel.** cloudflared killed 21:41:33Z, running again **21:43:07.478Z (~97 s)**, with `RestartCount=0` — so Docker's `unless-stopped` policy did *not* do it; the controller's protected-infra recovery redeployed it („[infra] deploying cloudflared →…") |
| time to steady | the apps never stopped; the way IN was restored in **~97 s** |
| alarm fired / true? | **two, correctly paired** — `health_critical` (error) 21:43, `health_recovered` (info) 21:48. Exactly what the ladder predicts for a missing protected container, and **the alarm was not a dead end** |
| should have fired, did not | none for the accident |
**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop
does not follow redirects, so it measures „is the app serving on the box", never „can the household
reach it from outside". It saw nothing while the public route was returning 530.
### Round 5 — `use` privatebin / accident: **docker restarted inside the guest**
| the five things | |
|---|---|
| what the customer saw | a gap well under a minute: 200 before, **404** two seconds after docker returned (traefik had not re-registered routes), serving again by 21:54:19Z — about **40 s** of shut doors |
| what the box did by itself | everything. Restart ran 21:53:27Z→21:53:41Z; **all 26 containers back at t+16s**; the controller returned with them, waited 51 s for the fleet to settle and found **nothing boot-orphaned to repair** — correct, since every container had already come back on its own policy |
| time to steady | **16 s** to 26 of 26 containers; **~40 s** until the doors served. The slower number is the one a household feels |
| alarm fired / true? | **`controller_started` (info) — true and correct**, and the only line the ladder expects: no `app_start_failed` (90 s boot grace), no liveness alarm |
| should have fired, did not | **none** |
**Household loop: NOT SAMPLED** — 0 lines, because the round lasted ~18 s and the loop samples every
2 minutes. Recorded as not sampled, never as a pass.
**A trap avoided.** The round's own alarm snapshot was taken **six seconds** after the controller
started, and from it `controller_started` looked missing. A later reading shows it present at 21:53.
An alarm cannot be called missing by a measurement taken before it could have fired.
### Rounds 6-12
PENDING
@@ -1,3 +1,4 @@
2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4
2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest
2026-09-16T21:41:33Z tunnel killed — 10 minutes
2026-09-16T21:51:33Z 10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement
@@ -72,3 +72,101 @@ true and answers a narrower question than it appears to.
This is the same distinction as the front-door method correction above, and it is recorded in both
places because the number and the label live in different files.
2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4
2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest
cloudflared
2026-09-16T21:41:33Z tunnel killed — 10 minutes
2026-09-16T21:51:33Z 10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement
/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh: line 109: unexpected EOF while looking for matching `"'
2026-09-16T21:51:33Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T21:51:35Z t+603s containers=26 (before 26)
2026-09-16T21:51:35Z STEADY after 603s
2026-09-16T21:51:36Z front doors: recipes=200 status=200 paste=200 wiki=200
2026-09-16T21:51:36Z household lines this round: 10 failures: 0
2026-09-16T21:51:36Z --- alarms ---
| Time | Severity | Type | Message | Source
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
| Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
| Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
2026-09-16T21:51:37Z ================ END ROUND 4 ================
## ROUND 4 — the five things (and the action that did NOT run)
**action drawn:** `offsite-run` (app: mealie) · **accident:** tunnel killed for ten minutes
**THE ACTION NEVER RAN.** The runner reported:
ABORT: no session — mine, not the product's
Cause, traced rather than guessed: **round 2's power cut wiped `/tmp` inside the guest**, and the
dashboard password file lived there. Round 3 was a `use` round and needed only the front door, so
round 4 was the first to touch it and the first to find it missing. The off-site run was therefore
never triggered, and **round 4 measured its accident but not its action.** Recorded as such; it is
not quietly re-run and counted as if it had gone to plan.
1. **What the customer saw.** From outside: the apps went away for about ninety seconds (public route
**530**) and came back on their own (**200**). From the house: nothing — traefik answered **301**
on the box throughout.
2. **What the box did by itself.** Repaired its own tunnel. cloudflared was killed 21:41:33Z and was
running again at **21:43:07.478Z (~97 s)** with `RestartCount=0` — so Docker's `unless-stopped`
policy did NOT do it; the controller's protected-infra recovery redeployed it
(„[infra] deploying cloudflared → /opt/docker/stacks/cloudflared"). 26 containers before and after.
3. **Time to steady.** The apps never stopped; the way IN was restored in **~97 seconds**.
(The runner's „STEADY after 603s" is the elapsed ten-minute window, not a recovery time.)
4. **Alarm fired / true?** **Two, and they pair correctly:**
21:43 error `health_critical` „Rendszer állapot kritikus (volt: ok)"
21:48 info `health_recovered` „Rendszer állapot helyreállt: ok (volt: fail)"
The ladder predicts exactly `health_critical` for a missing protected container, and the recovery
closes it — **the alarm was not a dead end.**
5. **Should have fired and did not.** None for the accident. The missing off-site run produced no
alarm because it was never started — my fault, not a silence of the product's.
**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop
does not follow redirects, so it measures the box, not the way in from outside; see the limit
recorded above. Someone away from home would have met 530 for about ninety seconds.
## Why round 4's action never ran — the root cause, and what it costs the night
`/tmp` inside the customer guest is cleared on boot. **Round 2's power cut rebooted the guest at
21:26**, so the dashboard password file I had pushed to `/tmp/.pw` during Phase 0 was gone from that
moment. Round 3 was a `use` round and needed only the front door, so nothing noticed. Round 4 was the
first round to need a dashboard session, and it aborted before its action:
ABORT: no session — mine, not the product's
**What it costs:** round 4's drawn action (`offsite-run`) was not exercised. Its accident was. The
round is recorded as half-measured rather than re-run and presented as whole — the schedule was drawn
before the night and re-running a round after seeing it fail is how a drill starts choosing its own
results.
**What it changes for the rounds still to come:** rounds 6 (`backup-system`), 8 (`backup-app`) and
10 (`restore`) all need that session. The file is being re-placed in **`/root`**, which is on the
guest's own filesystem and survives a reboot, and the runners now read it from there — so the next
power cut (round 10's hard reset) cannot silently disarm the same three rounds.
## The injector's second EOF, judged and left alone
`inject.sh: line 109: unexpected EOF` appeared again in this round. `bash -n` passes, and both
branches still to run are plain single commands with no nested quoting:
tunnel-down-10min -> G "docker kill cloudflared"
docker-restart -> G "systemctl restart docker"
The accident itself completed correctly — cloudflared was killed at 21:41:33Z and the recovery was
measured — so the message is noise from the remote shell's re-parse, not a failed injection. Recorded
and not chased further; it cannot affect the remaining rounds, whose branches are clean.
## The password file IS in place — and my check was the thing that was broken
-rw------- 1 root root 21 Sep 16 21:52 /root/.pw
-rw------- 1 root root 21 Sep 16 21:52 /tmp/.pw
Both present, 21 bytes, mode 600, on the guest's own filesystem.
My first two verifications reported it missing. Both were wrong in the same way: one wrote
`wc -c < /root/.pw` inside an `ssh` command, so the **redirect was evaluated on the VM** rather than
inside the guest and looked for the file on the wrong machine; the other had its output eaten by the
noise filter I pipe everything through. The push had worked the whole time.
That is the **tenth** instrument error of tonight and the same family as all the others: a check whose
own precondition was wrong, reporting a confident „not there". The cure that finally worked is the
one this project keeps re-learning — **run the check inside a marker block with no filtering and no
host-side redirection**, so the answer cannot be silently dropped:
pct exec 9201 -- sh -c "echo MARKER_START; ls -l /root/.pw /tmp/.pw; echo MARKER_END"
Rounds 6, 8 and 10 (backup-system, backup-app, restore) are therefore armed, and the copy in `/root`
survives the guest reboot that round 10's hard reset will cause.
@@ -0,0 +1,4 @@
2026-09-16T21:53:27Z ACCIDENT=docker-restart round=5
2026-09-16T21:53:27Z systemctl restart docker inside the customer guest
2026-09-16T21:53:41Z docker restarted
2026-09-16T21:53:41Z accident docker-restart complete
@@ -0,0 +1,76 @@
2026-09-16T21:53:25Z ================ ROUND 5 : use privatebin, while: docker-restart ================
2026-09-16T21:53:27Z --- BEFORE --- containers=26 paste=200 status=200 paste=200
2026-09-16T21:53:27Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:53:27Z --- ACTION: use on privatebin ---
2026-09-16T21:53:27Z paste read 1 -> 200
2026-09-16T21:53:27Z paste read 2 -> 200
2026-09-16T21:53:27Z paste read 3 -> 200
2026-09-16T21:53:27Z --- ACCIDENT: docker-restart (injected after the action started) ---
2026-09-16T21:53:27Z ACCIDENT=docker-restart round=5
2026-09-16T21:53:27Z systemctl restart docker inside the customer guest
2026-09-16T21:53:41Z docker restarted
2026-09-16T21:53:41Z accident docker-restart complete
2026-09-16T21:53:41Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T21:53:43Z t+16s containers=26 (before 26)
2026-09-16T21:53:43Z STEADY after 16s
2026-09-16T21:53:43Z front doors: paste=404 status=404 paste=404 wiki=404
2026-09-16T21:53:44Z household lines this round: 0 failures: 0
2026-09-16T21:53:44Z --- alarms ---
| Time | Severity | Type | Message | Source
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
| Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
| Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
2026-09-16T21:53:45Z ================ END ROUND 5 ================
## The controller DID restart with docker — established before judging the alarm
docker restart ran 21:53:27Z -> 21:53:41Z
felhom-controller StartedAt **2026-09-16T21:53:38.397Z** RestartCount=0 Status=running
traefik StartedAt 2026-09-16T21:53:38.634Z
controller log:
21:54:37 [bootrecon] boot window: fleet settled after 51s (3 identical samples 5s apart) — sweeping
21:54:37 [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
21:54:42 [stacks] Status refresh: 25 containers across 56 stacks
So the controller went down and came back with the docker daemon, took 51 seconds to decide the
fleet had settled, and found **nothing boot-orphaned to repair** — which is the right answer, because
every container had already come back on its own restart policy.
**A trap I nearly walked into, recorded because it would have been the eleventh tonight.** Round 5's
own alarm snapshot was taken at **21:53:44Z — six seconds after the controller started**. Concluding
„`controller_started` did not fire" from that snapshot would have been a measurement taken before the
thing it was measuring could have happened. The verdict is therefore taken from a later reading, not
from the round's own dump.
## ROUND 5 — the five things
**action:** `use` privatebin (three reads through its own front door, all 200)
**accident:** `systemctl restart docker` inside the guest — every container down at once
1. **What the customer saw.** A gap of well under a minute. The apps answered 200 before the restart;
two seconds after docker returned they were **404** (traefik had not re-registered routes yet), and
by 21:54:19Z they were serving again — LAN **301**, public **200**. Call it ~40 seconds of doors
being shut, with no error page beyond a plain 404.
2. **What the box did by itself.** Everything. `systemctl restart docker` ran 21:53:27Z → 21:53:41Z;
**all 26 containers were back by t+16s**, none in a non-Up state. The controller came back with
them (StartedAt 21:53:38Z), waited 51 s for the fleet to settle, and found **nothing
boot-orphaned to repair** — correct, because every container had already returned on its own
restart policy.
3. **Time to steady.** **16 seconds** to 26 of 26 containers; ~40 seconds until the front doors served
again. The slower of the two numbers is the one a household would feel.
4. **Alarm fired / true?** **`controller_started` (info) at 21:53 — true and correct.** The controller
really did restart, and the ladder expects exactly this one line: no `app_start_failed` (the 90 s
boot grace covers a restart) and no liveness alarm (far inside the 30-minute window).
5. **Should have fired and did not.** **None.**
**Household loop: NOT SAMPLED.** It logged 0 lines in this round — not because nothing failed, but
because the round lasted ~18 seconds and the loop samples every 2 minutes. Recorded as not sampled,
never as a pass, exactly as round 2's was.
**A trap avoided, and it would have been my eleventh:** the round's own alarm snapshot was taken at
21:53:44Z, **six seconds after the controller started**. From that snapshot `controller_started`
looked missing. A later reading at 21:55:18Z shows it present at 21:53. The verdict comes from the
later reading; an alarm cannot be called missing by a measurement taken before it could have fired.
@@ -0,0 +1,4 @@
2026-09-16T21:54:59Z ================ ROUND 6 : backup-system adventurelog, while: nothing ================
2026-09-16T21:55:00Z --- BEFORE --- containers=26 travel=200 status=200 paste=200
2026-09-16T21:55:00Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:55:01Z --- ACTION: backup-system on adventurelog ---
@@ -15,7 +15,7 @@ door(){ curl -sL -o /dev/null -w '%{http_code}' --max-time 12 -k -H "Host: $1.en
cat > $SCR/r4a.sh <<'PEOF'
#!/bin/bash
H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/tmp/.pw" $B/login
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/root/.pw" $B/login
S=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
[ -z "$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; }
C="Cookie: felhom_session=$S"
@@ -14,7 +14,7 @@ door(){ curl -sL -o /dev/null -w '%{http_code}' --max-time 12 -k -H "Host: $1.en
cat > $SCR/r6a.sh <<'PEOF'
#!/bin/bash
H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/tmp/.pw" $B/login
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/root/.pw" $B/login
S=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
[ -z "$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; }
C="Cookie: felhom_session=$S"
@@ -22,7 +22,7 @@ SUB=$(sub "$Y")
cat > $SCR/act.sh <<PEOF
#!/bin/bash
H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080
curl -s -o /dev/null -D /tmp/.l -H "\$H" -X POST --data-urlencode "password@/tmp/.pw" \$B/login
curl -s -o /dev/null -D /tmp/.l -H "\$H" -X POST --data-urlencode "password@/root/.pw" \$B/login
S=\$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
[ -z "\$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; }
C="Cookie: felhom_session=\$S"