CHAOS NIGHT rounds 4-5: the box repairs its own tunnel, and survives docker dying
gates / gates (push) Successful in 22s
gates / gates (push) Successful in 22s
Round 4 (tunnel killed 10 min): cloudflared came back in ~97 seconds and the CONTROLLER did it, not Docker - RestartCount=0 proves the unless-stopped policy never acted, and the log shows the protected-infra recovery redeploying it. health_critical fired at 21:43 and health_recovered closed it at 21:48, so the alarm was not a dead end. Round 4's drawn ACTION never ran, and that is recorded rather than re-run: the round-2 power cut rebooted the guest, /tmp is cleared on boot, and the dashboard password lived there. The off-site run was never triggered. Re-running a round after watching it fail is how a drill starts choosing its own results. The file now lives in /root, so round 10's hard reset cannot disarm rounds 6, 8 and 10. Round 5 (docker restarted): all 26 containers back in 16 seconds, doors serving again in ~40, controller_started true and correct, nothing missed. The household loop is recorded as NOT SAMPLED - 0 lines because the round was shorter than its 2-minute sampling interval, which is not the same as 0 failures. Two traps avoided and written down: round 5's own alarm snapshot was taken six seconds after the controller started, from which controller_started looked missing (it fired); and my check for the password file put its redirect on the wrong host, reporting "not there" for a file that was present all along. Tenth and eleventh of the same family tonight - a check whose own precondition was wrong, answering confidently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -251,7 +251,45 @@ estimating the clock instead of reading it. The readings were unchanged, but „
|
||||
minutes in" and „nothing in the whole window" are different findings. From here the end-of-window
|
||||
check is taken when the round's own runner reports completion.
|
||||
|
||||
### Rounds 4-12
|
||||
### Round 4 — `offsite-run` (mealie) / accident: **tunnel killed for ten minutes**
|
||||
|
||||
**The drawn ACTION never ran.** The runner aborted with „no session — mine, not the product's":
|
||||
round 2's power cut had rebooted the guest, `/tmp` is cleared on boot, and the dashboard password
|
||||
file lived there. Round 3 was a `use` round and never needed it, so round 4 was the first to find it
|
||||
gone. **The accident was measured; the off-site run was not.** Recorded as half-measured rather than
|
||||
re-run and presented as whole — re-running a round after seeing it fail is how a drill starts
|
||||
choosing its own results. The file now lives in `/root`, which survives a reboot.
|
||||
|
||||
| the five things | |
|
||||
|---|---|
|
||||
| what the customer saw | from outside, the apps vanished for ~90 s (public route **530**) and came back on their own (**200**); from inside the house, nothing — traefik answered **301** throughout |
|
||||
| what the box did by itself | **repaired its own tunnel.** cloudflared killed 21:41:33Z, running again **21:43:07.478Z (~97 s)**, with `RestartCount=0` — so Docker's `unless-stopped` policy did *not* do it; the controller's protected-infra recovery redeployed it („[infra] deploying cloudflared →…") |
|
||||
| time to steady | the apps never stopped; the way IN was restored in **~97 s** |
|
||||
| alarm fired / true? | **two, correctly paired** — `health_critical` (error) 21:43, `health_recovered` (info) 21:48. Exactly what the ladder predicts for a missing protected container, and **the alarm was not a dead end** |
|
||||
| should have fired, did not | none for the accident |
|
||||
|
||||
**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop
|
||||
does not follow redirects, so it measures „is the app serving on the box", never „can the household
|
||||
reach it from outside". It saw nothing while the public route was returning 530.
|
||||
|
||||
### Round 5 — `use` privatebin / accident: **docker restarted inside the guest**
|
||||
|
||||
| the five things | |
|
||||
|---|---|
|
||||
| what the customer saw | a gap well under a minute: 200 before, **404** two seconds after docker returned (traefik had not re-registered routes), serving again by 21:54:19Z — about **40 s** of shut doors |
|
||||
| what the box did by itself | everything. Restart ran 21:53:27Z→21:53:41Z; **all 26 containers back at t+16s**; the controller returned with them, waited 51 s for the fleet to settle and found **nothing boot-orphaned to repair** — correct, since every container had already come back on its own policy |
|
||||
| time to steady | **16 s** to 26 of 26 containers; **~40 s** until the doors served. The slower number is the one a household feels |
|
||||
| alarm fired / true? | **`controller_started` (info) — true and correct**, and the only line the ladder expects: no `app_start_failed` (90 s boot grace), no liveness alarm |
|
||||
| should have fired, did not | **none** |
|
||||
|
||||
**Household loop: NOT SAMPLED** — 0 lines, because the round lasted ~18 s and the loop samples every
|
||||
2 minutes. Recorded as not sampled, never as a pass.
|
||||
|
||||
**A trap avoided.** The round's own alarm snapshot was taken **six seconds** after the controller
|
||||
started, and from it `controller_started` looked missing. A later reading shows it present at 21:53.
|
||||
An alarm cannot be called missing by a measurement taken before it could have fired.
|
||||
|
||||
### Rounds 6-12
|
||||
|
||||
PENDING
|
||||
|
||||
|
||||
@@ -1,3 +1,4 @@
|
||||
2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4
|
||||
2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest
|
||||
2026-09-16T21:41:33Z tunnel killed — 10 minutes
|
||||
2026-09-16T21:51:33Z 10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement
|
||||
|
||||
@@ -72,3 +72,101 @@ true and answers a narrower question than it appears to.
|
||||
|
||||
This is the same distinction as the front-door method correction above, and it is recorded in both
|
||||
places because the number and the label live in different files.
|
||||
2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4
|
||||
2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest
|
||||
cloudflared
|
||||
2026-09-16T21:41:33Z tunnel killed — 10 minutes
|
||||
2026-09-16T21:51:33Z 10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement
|
||||
/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh: line 109: unexpected EOF while looking for matching `"'
|
||||
2026-09-16T21:51:33Z --- AFTER: what the box did BY ITSELF ---
|
||||
2026-09-16T21:51:35Z t+603s containers=26 (before 26)
|
||||
2026-09-16T21:51:35Z STEADY after 603s
|
||||
2026-09-16T21:51:36Z front doors: recipes=200 status=200 paste=200 wiki=200
|
||||
2026-09-16T21:51:36Z household lines this round: 10 failures: 0
|
||||
2026-09-16T21:51:36Z --- alarms ---
|
||||
| Time | Severity | Type | Message | Source
|
||||
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
|
||||
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
|
||||
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
|
||||
| Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
|
||||
| Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
|
||||
2026-09-16T21:51:37Z ================ END ROUND 4 ================
|
||||
|
||||
## ROUND 4 — the five things (and the action that did NOT run)
|
||||
**action drawn:** `offsite-run` (app: mealie) · **accident:** tunnel killed for ten minutes
|
||||
|
||||
**THE ACTION NEVER RAN.** The runner reported:
|
||||
ABORT: no session — mine, not the product's
|
||||
Cause, traced rather than guessed: **round 2's power cut wiped `/tmp` inside the guest**, and the
|
||||
dashboard password file lived there. Round 3 was a `use` round and needed only the front door, so
|
||||
round 4 was the first to touch it and the first to find it missing. The off-site run was therefore
|
||||
never triggered, and **round 4 measured its accident but not its action.** Recorded as such; it is
|
||||
not quietly re-run and counted as if it had gone to plan.
|
||||
|
||||
1. **What the customer saw.** From outside: the apps went away for about ninety seconds (public route
|
||||
**530**) and came back on their own (**200**). From the house: nothing — traefik answered **301**
|
||||
on the box throughout.
|
||||
2. **What the box did by itself.** Repaired its own tunnel. cloudflared was killed 21:41:33Z and was
|
||||
running again at **21:43:07.478Z (~97 s)** with `RestartCount=0` — so Docker's `unless-stopped`
|
||||
policy did NOT do it; the controller's protected-infra recovery redeployed it
|
||||
(„[infra] deploying cloudflared → /opt/docker/stacks/cloudflared"). 26 containers before and after.
|
||||
3. **Time to steady.** The apps never stopped; the way IN was restored in **~97 seconds**.
|
||||
(The runner's „STEADY after 603s" is the elapsed ten-minute window, not a recovery time.)
|
||||
4. **Alarm fired / true?** **Two, and they pair correctly:**
|
||||
21:43 error `health_critical` „Rendszer állapot kritikus (volt: ok)"
|
||||
21:48 info `health_recovered` „Rendszer állapot helyreállt: ok (volt: fail)"
|
||||
The ladder predicts exactly `health_critical` for a missing protected container, and the recovery
|
||||
closes it — **the alarm was not a dead end.**
|
||||
5. **Should have fired and did not.** None for the accident. The missing off-site run produced no
|
||||
alarm because it was never started — my fault, not a silence of the product's.
|
||||
|
||||
**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop
|
||||
does not follow redirects, so it measures the box, not the way in from outside; see the limit
|
||||
recorded above. Someone away from home would have met 530 for about ninety seconds.
|
||||
|
||||
## Why round 4's action never ran — the root cause, and what it costs the night
|
||||
`/tmp` inside the customer guest is cleared on boot. **Round 2's power cut rebooted the guest at
|
||||
21:26**, so the dashboard password file I had pushed to `/tmp/.pw` during Phase 0 was gone from that
|
||||
moment. Round 3 was a `use` round and needed only the front door, so nothing noticed. Round 4 was the
|
||||
first round to need a dashboard session, and it aborted before its action:
|
||||
ABORT: no session — mine, not the product's
|
||||
|
||||
**What it costs:** round 4's drawn action (`offsite-run`) was not exercised. Its accident was. The
|
||||
round is recorded as half-measured rather than re-run and presented as whole — the schedule was drawn
|
||||
before the night and re-running a round after seeing it fail is how a drill starts choosing its own
|
||||
results.
|
||||
|
||||
**What it changes for the rounds still to come:** rounds 6 (`backup-system`), 8 (`backup-app`) and
|
||||
10 (`restore`) all need that session. The file is being re-placed in **`/root`**, which is on the
|
||||
guest's own filesystem and survives a reboot, and the runners now read it from there — so the next
|
||||
power cut (round 10's hard reset) cannot silently disarm the same three rounds.
|
||||
|
||||
## The injector's second EOF, judged and left alone
|
||||
`inject.sh: line 109: unexpected EOF` appeared again in this round. `bash -n` passes, and both
|
||||
branches still to run are plain single commands with no nested quoting:
|
||||
tunnel-down-10min -> G "docker kill cloudflared"
|
||||
docker-restart -> G "systemctl restart docker"
|
||||
The accident itself completed correctly — cloudflared was killed at 21:41:33Z and the recovery was
|
||||
measured — so the message is noise from the remote shell's re-parse, not a failed injection. Recorded
|
||||
and not chased further; it cannot affect the remaining rounds, whose branches are clean.
|
||||
|
||||
## The password file IS in place — and my check was the thing that was broken
|
||||
-rw------- 1 root root 21 Sep 16 21:52 /root/.pw
|
||||
-rw------- 1 root root 21 Sep 16 21:52 /tmp/.pw
|
||||
Both present, 21 bytes, mode 600, on the guest's own filesystem.
|
||||
|
||||
My first two verifications reported it missing. Both were wrong in the same way: one wrote
|
||||
`wc -c < /root/.pw` inside an `ssh` command, so the **redirect was evaluated on the VM** rather than
|
||||
inside the guest and looked for the file on the wrong machine; the other had its output eaten by the
|
||||
noise filter I pipe everything through. The push had worked the whole time.
|
||||
|
||||
That is the **tenth** instrument error of tonight and the same family as all the others: a check whose
|
||||
own precondition was wrong, reporting a confident „not there". The cure that finally worked is the
|
||||
one this project keeps re-learning — **run the check inside a marker block with no filtering and no
|
||||
host-side redirection**, so the answer cannot be silently dropped:
|
||||
pct exec 9201 -- sh -c "echo MARKER_START; ls -l /root/.pw /tmp/.pw; echo MARKER_END"
|
||||
Rounds 6, 8 and 10 (backup-system, backup-app, restore) are therefore armed, and the copy in `/root`
|
||||
survives the guest reboot that round 10's hard reset will cause.
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
2026-09-16T21:53:27Z ACCIDENT=docker-restart round=5
|
||||
2026-09-16T21:53:27Z systemctl restart docker inside the customer guest
|
||||
2026-09-16T21:53:41Z docker restarted
|
||||
2026-09-16T21:53:41Z accident docker-restart complete
|
||||
@@ -0,0 +1,76 @@
|
||||
2026-09-16T21:53:25Z ================ ROUND 5 : use privatebin, while: docker-restart ================
|
||||
2026-09-16T21:53:27Z --- BEFORE --- containers=26 paste=200 status=200 paste=200
|
||||
2026-09-16T21:53:27Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T21:53:27Z --- ACTION: use on privatebin ---
|
||||
2026-09-16T21:53:27Z paste read 1 -> 200
|
||||
2026-09-16T21:53:27Z paste read 2 -> 200
|
||||
2026-09-16T21:53:27Z paste read 3 -> 200
|
||||
2026-09-16T21:53:27Z --- ACCIDENT: docker-restart (injected after the action started) ---
|
||||
2026-09-16T21:53:27Z ACCIDENT=docker-restart round=5
|
||||
2026-09-16T21:53:27Z systemctl restart docker inside the customer guest
|
||||
2026-09-16T21:53:41Z docker restarted
|
||||
2026-09-16T21:53:41Z accident docker-restart complete
|
||||
2026-09-16T21:53:41Z --- AFTER: what the box did BY ITSELF ---
|
||||
2026-09-16T21:53:43Z t+16s containers=26 (before 26)
|
||||
2026-09-16T21:53:43Z STEADY after 16s
|
||||
2026-09-16T21:53:43Z front doors: paste=404 status=404 paste=404 wiki=404
|
||||
2026-09-16T21:53:44Z household lines this round: 0 failures: 0
|
||||
2026-09-16T21:53:44Z --- alarms ---
|
||||
| Time | Severity | Type | Message | Source
|
||||
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
|
||||
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
|
||||
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
|
||||
| Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
|
||||
| Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
|
||||
2026-09-16T21:53:45Z ================ END ROUND 5 ================
|
||||
|
||||
## The controller DID restart with docker — established before judging the alarm
|
||||
docker restart ran 21:53:27Z -> 21:53:41Z
|
||||
felhom-controller StartedAt **2026-09-16T21:53:38.397Z** RestartCount=0 Status=running
|
||||
traefik StartedAt 2026-09-16T21:53:38.634Z
|
||||
controller log:
|
||||
21:54:37 [bootrecon] boot window: fleet settled after 51s (3 identical samples 5s apart) — sweeping
|
||||
21:54:37 [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
|
||||
21:54:42 [stacks] Status refresh: 25 containers across 56 stacks
|
||||
|
||||
So the controller went down and came back with the docker daemon, took 51 seconds to decide the
|
||||
fleet had settled, and found **nothing boot-orphaned to repair** — which is the right answer, because
|
||||
every container had already come back on its own restart policy.
|
||||
|
||||
**A trap I nearly walked into, recorded because it would have been the eleventh tonight.** Round 5's
|
||||
own alarm snapshot was taken at **21:53:44Z — six seconds after the controller started**. Concluding
|
||||
„`controller_started` did not fire" from that snapshot would have been a measurement taken before the
|
||||
thing it was measuring could have happened. The verdict is therefore taken from a later reading, not
|
||||
from the round's own dump.
|
||||
|
||||
## ROUND 5 — the five things
|
||||
**action:** `use` privatebin (three reads through its own front door, all 200)
|
||||
**accident:** `systemctl restart docker` inside the guest — every container down at once
|
||||
|
||||
1. **What the customer saw.** A gap of well under a minute. The apps answered 200 before the restart;
|
||||
two seconds after docker returned they were **404** (traefik had not re-registered routes yet), and
|
||||
by 21:54:19Z they were serving again — LAN **301**, public **200**. Call it ~40 seconds of doors
|
||||
being shut, with no error page beyond a plain 404.
|
||||
2. **What the box did by itself.** Everything. `systemctl restart docker` ran 21:53:27Z → 21:53:41Z;
|
||||
**all 26 containers were back by t+16s**, none in a non-Up state. The controller came back with
|
||||
them (StartedAt 21:53:38Z), waited 51 s for the fleet to settle, and found **nothing
|
||||
boot-orphaned to repair** — correct, because every container had already returned on its own
|
||||
restart policy.
|
||||
3. **Time to steady.** **16 seconds** to 26 of 26 containers; ~40 seconds until the front doors served
|
||||
again. The slower of the two numbers is the one a household would feel.
|
||||
4. **Alarm fired / true?** **`controller_started` (info) at 21:53 — true and correct.** The controller
|
||||
really did restart, and the ladder expects exactly this one line: no `app_start_failed` (the 90 s
|
||||
boot grace covers a restart) and no liveness alarm (far inside the 30-minute window).
|
||||
5. **Should have fired and did not.** **None.**
|
||||
|
||||
**Household loop: NOT SAMPLED.** It logged 0 lines in this round — not because nothing failed, but
|
||||
because the round lasted ~18 seconds and the loop samples every 2 minutes. Recorded as not sampled,
|
||||
never as a pass, exactly as round 2's was.
|
||||
|
||||
**A trap avoided, and it would have been my eleventh:** the round's own alarm snapshot was taken at
|
||||
21:53:44Z, **six seconds after the controller started**. From that snapshot `controller_started`
|
||||
looked missing. A later reading at 21:55:18Z shows it present at 21:53. The verdict comes from the
|
||||
later reading; an alarm cannot be called missing by a measurement taken before it could have fired.
|
||||
@@ -0,0 +1,4 @@
|
||||
2026-09-16T21:54:59Z ================ ROUND 6 : backup-system adventurelog, while: nothing ================
|
||||
2026-09-16T21:55:00Z --- BEFORE --- containers=26 travel=200 status=200 paste=200
|
||||
2026-09-16T21:55:00Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T21:55:01Z --- ACTION: backup-system on adventurelog ---
|
||||
@@ -15,7 +15,7 @@ door(){ curl -sL -o /dev/null -w '%{http_code}' --max-time 12 -k -H "Host: $1.en
|
||||
cat > $SCR/r4a.sh <<'PEOF'
|
||||
#!/bin/bash
|
||||
H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080
|
||||
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/tmp/.pw" $B/login
|
||||
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/root/.pw" $B/login
|
||||
S=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
|
||||
[ -z "$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; }
|
||||
C="Cookie: felhom_session=$S"
|
||||
|
||||
@@ -14,7 +14,7 @@ door(){ curl -sL -o /dev/null -w '%{http_code}' --max-time 12 -k -H "Host: $1.en
|
||||
cat > $SCR/r6a.sh <<'PEOF'
|
||||
#!/bin/bash
|
||||
H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080
|
||||
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/tmp/.pw" $B/login
|
||||
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@/root/.pw" $B/login
|
||||
S=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
|
||||
[ -z "$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; }
|
||||
C="Cookie: felhom_session=$S"
|
||||
|
||||
@@ -22,7 +22,7 @@ SUB=$(sub "$Y")
|
||||
cat > $SCR/act.sh <<PEOF
|
||||
#!/bin/bash
|
||||
H="Host: felhom.enkicsifelhom.hu"; B=http://172.17.0.2:8080
|
||||
curl -s -o /dev/null -D /tmp/.l -H "\$H" -X POST --data-urlencode "password@/tmp/.pw" \$B/login
|
||||
curl -s -o /dev/null -D /tmp/.l -H "\$H" -X POST --data-urlencode "password@/root/.pw" \$B/login
|
||||
S=\$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
|
||||
[ -z "\$S" ] && { echo "ABORT: no session — mine, not the product's"; exit 1; }
|
||||
C="Cookie: felhom_session=\$S"
|
||||
|
||||
Reference in New Issue
Block a user