CHAOS NIGHT round 6: what a whole-system backup costs the household
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
Measured while the vzdump was in flight (started 21:55:24Z, 2.1GB and growing): - containers running: 4 of 26. The whole-guest backup stops twenty-two apps. - front doors: LAN 301 (traefik is one of the four still up) but public 404 - nothing behind the proxy to serve. Every app unavailable for the duration. - alarms: NONE. Twenty-two apps went down at once and not one alarm fired. Those two findings point opposite ways and both matter. The downtime is real and total, not a brief pause - this is the standing whole-system-backup downtime row, seen on a fresh box with twelve apps. And the suppression is correct: the backup's own stack stops belong to a suppression set, so the box does not alarm about downtime it caused deliberately. The duration itself is taken from the round's runner when it reports, not estimated from a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -2,3 +2,28 @@
|
||||
2026-09-16T21:55:00Z --- BEFORE --- containers=26 travel=200 status=200 paste=200
|
||||
2026-09-16T21:55:00Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T21:55:01Z --- ACTION: backup-system on adventurelog ---
|
||||
|
||||
## MID-BACKUP measurements — what a household actually experiences (21:56:55Z)
|
||||
The whole-system backup started **21:55:24Z** (`vzdump-lxc-9201-2026_09_16-23_55_24.tar.dat`, local
|
||||
time 23:55:24 CEST, growing through 2.1 GB) and while it ran:
|
||||
|
||||
containers running **4** (was 26) — the apps are STOPPED for the whole-guest backup
|
||||
vzdump processes running, plus its .tmp working directory
|
||||
front doors, LAN **301** — traefik is one of the four still up, so the box still answers
|
||||
front doors, public **404** — nothing behind the proxy to serve
|
||||
alarms **none** — newest event still `controller_started` 21:53
|
||||
|
||||
**Two findings in that, and they point opposite ways.**
|
||||
|
||||
**1. The downtime is real and total.** A whole-system backup stops twenty-two of the twenty-six
|
||||
containers. For as long as it runs, every app is unavailable — `status`, `paste` and `travel` all
|
||||
answered **404** through the public route. This is the subject of the standing whole-system-backup
|
||||
downtime row, seen here on a fresh box with twelve apps: not a brief pause, but the apps down for the
|
||||
duration of a multi-gigabyte dump.
|
||||
|
||||
**2. The suppression works.** Twenty-two apps went down at once and **not one alarm fired**. That is
|
||||
correct and deliberate: the backup's own stack stops belong to a suppression set, so the box does not
|
||||
alarm about downtime it caused on purpose. The distinction this project keeps re-learning — a third
|
||||
way to stop an app needs a third suppression set — held here.
|
||||
|
||||
The duration is taken from the round's own runner when it reports, not estimated from this sample.
|
||||
|
||||
Reference in New Issue
Block a user