From d4318529b4b80e515f2e790813ec4e2af792ce2d Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 23:58:00 +0200 Subject: [PATCH] CHAOS NIGHT round 6: what a whole-system backup costs the household Measured while the vzdump was in flight (started 21:55:24Z, 2.1GB and growing): - containers running: 4 of 26. The whole-guest backup stops twenty-two apps. - front doors: LAN 301 (traefik is one of the four still up) but public 404 - nothing behind the proxy to serve. Every app unavailable for the duration. - alarms: NONE. Twenty-two apps went down at once and not one alarm fired. Those two findings point opposite ways and both matter. The downtime is real and total, not a brief pause - this is the standing whole-system-backup downtime row, seen on a fresh box with twelve apps. And the suppression is correct: the backup's own stack stops belong to a suppression set, so the box does not alarm about downtime it caused deliberately. The duration itself is taken from the round's runner when it reports, not estimated from a single sample. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../round-6.txt | 25 +++++++++++++++++++ 1 file changed, 25 insertions(+) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt index 5ac8d8a0..27a0a826 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt @@ -2,3 +2,28 @@ 2026-09-16T21:55:00Z --- BEFORE --- containers=26 travel=200 status=200 paste=200 2026-09-16T21:55:00Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) 2026-09-16T21:55:01Z --- ACTION: backup-system on adventurelog --- + +## MID-BACKUP measurements — what a household actually experiences (21:56:55Z) +The whole-system backup started **21:55:24Z** (`vzdump-lxc-9201-2026_09_16-23_55_24.tar.dat`, local +time 23:55:24 CEST, growing through 2.1 GB) and while it ran: + + containers running **4** (was 26) — the apps are STOPPED for the whole-guest backup + vzdump processes running, plus its .tmp working directory + front doors, LAN **301** — traefik is one of the four still up, so the box still answers + front doors, public **404** — nothing behind the proxy to serve + alarms **none** — newest event still `controller_started` 21:53 + +**Two findings in that, and they point opposite ways.** + +**1. The downtime is real and total.** A whole-system backup stops twenty-two of the twenty-six +containers. For as long as it runs, every app is unavailable — `status`, `paste` and `travel` all +answered **404** through the public route. This is the subject of the standing whole-system-backup +downtime row, seen here on a fresh box with twelve apps: not a brief pause, but the apps down for the +duration of a multi-gigabyte dump. + +**2. The suppression works.** Twenty-two apps went down at once and **not one alarm fired**. That is +correct and deliberate: the backup's own stack stops belong to a suppression set, so the box does not +alarm about downtime it caused on purpose. The distinction this project keeps re-learning — a third +way to stop an app needs a third suppression set — held here. + +The duration is taken from the round's own runner when it reports, not estimated from this sample.