2026-09-16T21:41:25Z ================ ROUND 4 : offsite-run mealie, while: tunnel-down-10min ================
2026-09-16T21:41:28Z --- BEFORE --- containers=26  recipes=200  status=200  paste=200
2026-09-16T21:41:28Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:41:29Z --- ACTION: offsite-run on mealie ---
ABORT: no session — mine, not the product's
2026-09-16T21:41:32Z --- ACCIDENT: tunnel-down-10min (injected after the action started) ---

## A METHOD CORRECTION that retro-labels every front-door reading tonight
During the tunnel outage the front doors returned **530**. That is a **Cloudflare** status, not one
traefik would produce — which exposed what my sweep was actually measuring.

`curl -sL -H "Host: x.enkicsifelhom.hu" http://192.168.0.116/` gets a **301** from traefik and then
**follows it** to `https://x.enkicsifelhom.hu/`, which resolves through public DNS to Cloudflare and
comes back down the tunnel. So every „front door" reading I have taken tonight — the 200s in Phase 0,
the 10-of-12 and 11-of-12 sweeps, the per-round `door()` calls — has been measuring the **PUBLIC
path**: box -> tunnel -> Cloudflare -> back. Not the LAN.

**What that does and does not invalidate.** It does NOT invalidate the readings: a 200 by that route
is a stronger statement than a LAN 200, because it proves traefik, the app, the tunnel and Cloudflare
all worked. What it invalidates is the LABEL „from DooPlex over the LAN", and with it any conclusion
of the form „the app is fine, only the tunnel is down" — that distinction was never being drawn.

**Corrected method, used from here on:** the two paths are measured separately —
  * LAN: no redirect following, straight to the box
  * public: the hostname pinned to the box, or the ordinary public route
so „the app serves" and „the tunnel serves" are two answers, not one.

This is the eighth instrument error of mine tonight and the most consequential, because it silently
changed what a green reading MEANT rather than producing an obviously wrong one.

## The tunnel came back BY ITSELF in ~97 seconds — and the CONTROLLER did it, not Docker
    killed          ~21:41:30Z   (`docker kill cloudflared`, deliberately never restarted by hand)
    cloudflared     Exited (137) at 21:42:44Z, 25 of 26 containers running
    back up         **StartedAt 2026-09-16T21:43:07.478Z**, Status=running, 26 containers
    elapsed         **~97 seconds**

**Who restarted it is answerable, and it matters.** `RestartCount=**0**` — Docker's own
`unless-stopped` policy did NOT act, because that would have incremented the counter. The controller's
log says what did:
    21:43:07 [INFO] [infra] deploying cloudflared → /opt/docker/stacks/cloudflared
i.e. the product's **protected-infra recovery** noticed a protected stack was gone and redeployed it.
That is a stronger result than „docker restarted it": the box repaired one of its own infrastructure
stacks without anyone asking.

## The alarm the ladder predicted, and it fired TRUE
    Sep 16 21:43  **error  health_critical**  „Rendszer állapot kritikus (volt: ok)"
A missing protected container forces the health verdict to fail, and `NotifyHealthChange` is
edge-triggered on the rank increase — so exactly one alarm, at error severity, customer-facing and ON
by default. It fired within the outage rather than after it, and it is TRUE: the tunnel really was
down and the box really was not fully healthy.

**Front doors during the outage, measured on both paths (see the method correction above):**
    public route (box -> tunnel -> Cloudflare)   **530** at 21:42:44Z, **200** again at 21:43:15Z
    LAN route straight to the box                traefik answered **301** throughout
So the household's apps never stopped serving locally; what failed, and recovered, was the way in
from outside.

## A LIMIT of the household loop, stated before its numbers are read
Through the tunnel outage the loop logged:
    21:41:57Z paperless read ok http=301   21:41:57Z paperless dash ok http=301
    21:43:57Z paste     read ok http=301   21:43:57Z paste     dash ok http=301
i.e. **no failures at all**, while the public way in was returning 530.

That is correct behaviour for the loop and a real blind spot at the same time. The loop does **not**
follow redirects: it takes traefik's 301 on the box as a success, so it measures *„is the app serving
on the box?"* and never *„can the household reach it from outside?"*. During round 4 the first answer
stayed yes throughout, which is why it saw nothing.

**Therefore: „0 household failures in round 4" must NOT be read as „the household was unaffected".**
A household member away from home would have met 530 for about ninety seconds. The loop's number is
true and answers a narrower question than it appears to.

This is the same distinction as the front-door method correction above, and it is recorded in both
places because the number and the label live in different files.
    2026-09-16T21:41:32Z ACCIDENT=tunnel-down-10min round=4
    2026-09-16T21:41:32Z docker kill cloudflared inside the customer guest
    cloudflared
    2026-09-16T21:41:33Z tunnel killed — 10 minutes
    2026-09-16T21:51:33Z 10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement
    /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh: line 109: unexpected EOF while looking for matching `"'
2026-09-16T21:51:33Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T21:51:35Z     t+603s containers=26 (before 26)
2026-09-16T21:51:35Z     STEADY after 603s
2026-09-16T21:51:36Z     front doors: recipes=200  status=200  paste=200  wiki=200
2026-09-16T21:51:36Z     household lines this round: 10  failures: 0
2026-09-16T21:51:36Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
  | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
  | Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
  | Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
2026-09-16T21:51:37Z ================ END ROUND 4 ================

## ROUND 4 — the five things (and the action that did NOT run)
**action drawn:** `offsite-run` (app: mealie) · **accident:** tunnel killed for ten minutes

**THE ACTION NEVER RAN.** The runner reported:
    ABORT: no session — mine, not the product's
Cause, traced rather than guessed: **round 2's power cut wiped `/tmp` inside the guest**, and the
dashboard password file lived there. Round 3 was a `use` round and needed only the front door, so
round 4 was the first to touch it and the first to find it missing. The off-site run was therefore
never triggered, and **round 4 measured its accident but not its action.** Recorded as such; it is
not quietly re-run and counted as if it had gone to plan.

1. **What the customer saw.** From outside: the apps went away for about ninety seconds (public route
   **530**) and came back on their own (**200**). From the house: nothing — traefik answered **301**
   on the box throughout.
2. **What the box did by itself.** Repaired its own tunnel. cloudflared was killed 21:41:33Z and was
   running again at **21:43:07.478Z (~97 s)** with `RestartCount=0` — so Docker's `unless-stopped`
   policy did NOT do it; the controller's protected-infra recovery redeployed it
   („[infra] deploying cloudflared → /opt/docker/stacks/cloudflared"). 26 containers before and after.
3. **Time to steady.** The apps never stopped; the way IN was restored in **~97 seconds**.
   (The runner's „STEADY after 603s" is the elapsed ten-minute window, not a recovery time.)
4. **Alarm fired / true?** **Two, and they pair correctly:**
       21:43  error  `health_critical`   „Rendszer állapot kritikus (volt: ok)"
       21:48  info   `health_recovered`  „Rendszer állapot helyreállt: ok (volt: fail)"
   The ladder predicts exactly `health_critical` for a missing protected container, and the recovery
   closes it — **the alarm was not a dead end.**
5. **Should have fired and did not.** None for the accident. The missing off-site run produced no
   alarm because it was never started — my fault, not a silence of the product's.

**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop
does not follow redirects, so it measures the box, not the way in from outside; see the limit
recorded above. Someone away from home would have met 530 for about ninety seconds.

## Why round 4's action never ran — the root cause, and what it costs the night
`/tmp` inside the customer guest is cleared on boot. **Round 2's power cut rebooted the guest at
21:26**, so the dashboard password file I had pushed to `/tmp/.pw` during Phase 0 was gone from that
moment. Round 3 was a `use` round and needed only the front door, so nothing noticed. Round 4 was the
first round to need a dashboard session, and it aborted before its action:
    ABORT: no session — mine, not the product's

**What it costs:** round 4's drawn action (`offsite-run`) was not exercised. Its accident was. The
round is recorded as half-measured rather than re-run and presented as whole — the schedule was drawn
before the night and re-running a round after seeing it fail is how a drill starts choosing its own
results.

**What it changes for the rounds still to come:** rounds 6 (`backup-system`), 8 (`backup-app`) and
10 (`restore`) all need that session. The file is being re-placed in **`/root`**, which is on the
guest's own filesystem and survives a reboot, and the runners now read it from there — so the next
power cut (round 10's hard reset) cannot silently disarm the same three rounds.

## The injector's second EOF, judged and left alone
`inject.sh: line 109: unexpected EOF` appeared again in this round. `bash -n` passes, and both
branches still to run are plain single commands with no nested quoting:
    tunnel-down-10min ->  G "docker kill cloudflared"
    docker-restart    ->  G "systemctl restart docker"
The accident itself completed correctly — cloudflared was killed at 21:41:33Z and the recovery was
measured — so the message is noise from the remote shell's re-parse, not a failed injection. Recorded
and not chased further; it cannot affect the remaining rounds, whose branches are clean.

## The password file IS in place — and my check was the thing that was broken
    -rw------- 1 root root 21 Sep 16 21:52 /root/.pw
    -rw------- 1 root root 21 Sep 16 21:52 /tmp/.pw
Both present, 21 bytes, mode 600, on the guest's own filesystem.

My first two verifications reported it missing. Both were wrong in the same way: one wrote
`wc -c < /root/.pw` inside an `ssh` command, so the **redirect was evaluated on the VM** rather than
inside the guest and looked for the file on the wrong machine; the other had its output eaten by the
noise filter I pipe everything through. The push had worked the whole time.

That is the **tenth** instrument error of tonight and the same family as all the others: a check whose
own precondition was wrong, reporting a confident „not there". The cure that finally worked is the
one this project keeps re-learning — **run the check inside a marker block with no filtering and no
host-side redirection**, so the answer cannot be silently dropped:
    pct exec 9201 -- sh -c "echo MARKER_START; ls -l /root/.pw /tmp/.pw; echo MARKER_END"
Rounds 6, 8 and 10 (backup-system, backup-app, restore) are therefore armed, and the copy in `/root`
survives the guest reboot that round 10's hard reset will cause.
