  -rwxrwxr-x 1 kisfenyo kisfenyo 5650 Sep 16 23:42 /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh
  -rwxrwxr-x 1 kisfenyo kisfenyo 5712 Sep 16 23:52 /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh
  waiter armed; will launch round 8 at 2026-09-16T22:35:48Z
=== round 8 launched 2026-09-16T22:35:48Z (due 22:35:48Z) ===
2026-09-16T22:35:48Z ================ ROUND 8 : backup-app nextcloud, while: internet-gone-10min ================
2026-09-16T22:35:50Z --- BEFORE --- containers=26  cloud=200  status=200  paste=200
2026-09-16T22:35:50Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T22:35:50Z --- ACTION: backup-app on nextcloud ---
  POST /api/backup/run -> 200
{"ok":true,"message":"Mentés elindítva"}

  {"ok":true,"data":{"enabled":true,"running":true}}
  {"ok":true,"data":{"enabled":true,"running":true}}
  {"ok":true,"data":{"enabled":true,"running":true}}
  {"ok":true,"data":{"enabled":true,"running":true}}
  {"ok":true,"data":{"enabled":true,"running":true}}
  {"ok":true,"data":{"db_dump":{"count":5,"duration":"1m55.578787777s","last_run":"2026-09-16T22:37:48.900952183Z","success":true},"enabled":true,"running":false}}
2026-09-16T22:37:54Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---
    2026-09-16T22:37:54Z ACCIDENT=internet-gone-10min round=8
    2026-09-16T22:37:55Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up
    2026-09-16T22:37:55Z blocked (LAN allowed, everything else dropped) — 10 minutes
    2026-09-16T22:47:55Z unblocked; host sysctl restored to 0 and both rules removed
    -P FORWARD ACCEPT
    2026-09-16T22:47:55Z accident internet-gone-10min complete
2026-09-16T22:47:55Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T22:47:57Z     t+603s containers=26 (before 26)
2026-09-16T22:47:57Z     STEADY after 603s
2026-09-16T22:47:58Z     front doors: cloud=530  status=530  paste=530  wiki=530
2026-09-16T22:47:58Z     household lines this round: 12  failures: 0
2026-09-16T22:47:58Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
  | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
  | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
2026-09-16T22:47:59Z ================ END ROUND 8 ================

[exited with code 0]

## POST-ROUND CORRECTIONS, written immediately (2026-09-16T22:48Z)

### 1. The "530 on every door" line above is a MISTIMED READING, not a fault.
It was taken 3 seconds after the unblock (22:47:58Z; unblock 22:47:55Z). Re-measured at
22:48:38Z, 43 s after the unblock:
    cloud  LAN=301 public=200
    status LAN=301 public=200
    paste  LAN=301 public=200
    wiki   LAN=301 public=200
So the public path came back by itself in <=43 s. This is the SIXTH mistimed reading tonight and
the runner's own AFTER step is the one still doing it -- it measures the public door 3 s after the
network returns, which can never be a fair reading. Fix applied below.

### 2. THE ACCIDENT DID NOT DO WHAT ITS NAME SAYS. This is my fault, and it is the bigger finding.
The hub report due at 22:38:43Z fell INSIDE the ten-minute cut (22:37:55 -> 22:47:55).
It SUCCEEDED: "[report] Hub report pushed successfully (15526 bytes)".
It succeeded because hub.felhom.eu resolves to a LAN address in this lab, and my injector blocks
everything EXCEPT the LAN ("LAN allowed, everything else dropped"). So:
  * the box never lost contact with the hub in round 7 OR round 8;
  * "internet gone" as injected means only "the PUBLIC path is gone";
  * the dropped-event question (events pushed while the hub is unreachable are retried 3x then
    dropped permanently, no queue) is STILL unmeasured after two rounds that looked like they
    measured it.
My own memory file carries this exact warning -- "hub.felhom.eu resolves to the LAN here; an
internet-cut drill must block it too" -- and I did not apply it. Recorded as an instrument fault,
not a product defect. Nothing the box did was wrong.

### 3. The action and the accident barely overlapped, and would not have interacted anyway.
The backup finished at 22:37:48Z (db_dump count=5, 1m55s, success=true). The cut began 22:37:54Z,
six seconds LATER. So "backup-app during an internet cut" was not really exercised. It would not
have mattered: backup-app is the LOCAL app-data tier (DB dump + volumes) and needs no internet.
The off-site tier is a different action (offsite-run), drawn in rounds 1 and 4.

### 4. What the box actually did, which is all true and all good.
  containers 26 -> 26, never dropped, at no point during the ten minutes
  household loop: 12 lines this round, 0 failures
  front doors on the LAN: served throughout
  public path: lost during the cut, restored by itself in <=43 s, unaided
  alarms: NONE fired, and none should have (node_stale threshold is 30 min; and the box was
          never actually stale, because it was reporting to the hub the whole time)
