CHAOS NIGHT round 6: the local backup tier cannot fit, and the box says so properly
gates / gates (push) Successful in 21s

Round 6 (whole-guest backup, control round). The local tier failed - a ~29GB
source into a 14GB root filesystem - and the product handled it well:
  whole_guest_backup_failed (error): "Whole-guest backup FAILED on the LOCAL
  TIER - retrying with backoff (next attempt in 15m0s)"
It names the tier rather than "the backup", says what it will do next, and its
status surface agrees (target_id local, success false, size_bytes 0). Then the
OFF-SITE tier ran from the same snapshot with the apps already back up,
finishing in ~8.5 minutes, encrypted to ep0, consuming no local disk at all.

During the backup 4 of 26 containers were up and every app answered 404
publicly - and NO alarm fired for those stops, which is correct: the backup's
own stack stops are suppressed, so the box does not alarm about downtime it
caused deliberately.

The downtime number (~5m43s) is recorded as CONTAMINATED by my own intervention
rather than presented as clean: I killed the local leg partway through.

A safety guard now runs on the box before round 7, declared and not a
measurement: the failed local tier retries every ~15 minutes and would consume /
at ~16MB/s while round 7 has the network cut. The guard kills only a local-tier
dump, only below a 2500M floor, and logs every action. If it fires it is an
intervention and will be counted as one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 00:11:37 +02:00
parent aaf0537665
commit c3722e06d2
5 changed files with 113 additions and 2 deletions
@@ -289,7 +289,27 @@ reach it from outside". It saw nothing while the public route was returning 530.
started, and from it `controller_started` looked missing. A later reading shows it present at 21:53.
An alarm cannot be called missing by a measurement taken before it could have fired.
### Rounds 6-12
### Round 6 — `backup-system` (whole-guest backup) / accident: **none** (control round)
| the five things | |
|---|---|
| what the customer saw | their apps went away and came back: during the backup **4 of 26** containers were up and every app answered **404** publicly; all 26 were serving again by 22:01:13Z. No banner, no mail — correctly |
| what the box did by itself | quiesced the apps, snapshotted, ran the local tier, **failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z (~8½ min)**, encrypted to ep0, taking **no local disk at all** |
| time to steady | apps down ~21:55:30Z → 22:01:13Z ≈ **5m43s** — **contaminated by my own intervention** (I killed the local leg at 21:59:45), so it is an upper bound on the quiesce window, not a clean measurement |
| alarm fired / true? | **`whole_guest_backup_failed` (error) — true, and better than true:** „Whole-guest backup FAILED on the **local tier** — retrying with backoff (next attempt in 15m0s)". It names the tier, not „the backup", and says what it will do next. The status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`). **No alarm for the 22 apps it stopped** — correct, those stops are suppressed |
| should have fired, did not | **none** |
**The finding to carry forward:** a whole-guest backup on this box **cannot use its local tier** — a
~29 GB source into a 14 GB root filesystem — and the product handles that honestly: it fails the tier,
says which tier, schedules a retry, and still gets the data out of the house on the off-site tier.
The local tier will keep retrying and keep failing on a box shaped like this one.
**My intervention, and the correction it needed.** I killed the local leg when two samples showed `/`
falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nested PVE. The arithmetic
stands, but I first wrote „I stopped the backup", which was wrong: I stopped **one leg**, and the
off-site leg started by itself seconds later and succeeded. Counted as an intervention either way.
### Rounds 7-12
PENDING
@@ -3,3 +3,17 @@
2026-09-16T22:02:40Z procs=2 apps=26 free=7573M
2026-09-16T22:03:12Z procs=2 apps=26 free=7573M
2026-09-16T22:03:43Z procs=2 apps=26 free=7573M
2026-09-16T22:04:15Z procs=2 apps=26 free=7573M
2026-09-16T22:04:46Z procs=4 apps=26 free=7573M
2026-09-16T22:05:18Z procs=2 apps=26 free=7573M
2026-09-16T22:05:49Z procs=2 apps=26 free=7573M
2026-09-16T22:06:21Z procs=2 apps=26 free=7573M
2026-09-16T22:06:52Z procs=2 apps=26 free=7573M
2026-09-16T22:07:24Z procs=2 apps=26 free=7573M
2026-09-16T22:07:55Z procs=2 apps=26 free=7573M
2026-09-16T22:08:27Z procs=0 apps=26 free=7573M
OFF-SITE LEG FINISHED at 2026-09-16T22:08:27Z
--- what the task list says about it ---
├──────┼────────────┼───────┼────────┼─────────────────────┼────────┼────────────────────────────────────────────────────────────────────────────────┼──────────────────┼─────────────────────┼────────────┤
│ 9201 │ chaosnight │ 82472 │ 175208 │ 2026-09-16 23:55:24 │ vzdump │ UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felhom-agent@pve!agent: │ felhom-agent@pve │ 2026-09-16 23:59:45 │ job errors │
└──────┴────────────┴───────┴────────┴─────────────────────┴────────┴────────────────────────────────────────────────────────────────────────────────┴──────────────────┴─────────────────────┴────────────┘
@@ -154,4 +154,70 @@ product's own path and read only when nothing else can answer.
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent
## ROUND 6 — the five things
**action:** `backup-system` (whole-guest backup; drawn app adventurelog is which row is read after)
**accident:** none — control round
1. **What the customer saw.** Their apps went away and came back. During the backup **4 of 26**
containers were up and every app answered **404** through the public route; by 22:01:13Z all 26 were
back and serving. No banner, no mail — and correctly so.
2. **What the box did by itself.** Quiesced the apps, took an LVM snapshot, ran the local tier,
**failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the
snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z, about 8½ minutes**, encrypted to
ep0, consuming **no local disk at all** (free stayed at 7573M throughout).
3. **Time to steady.** Apps down from ~21:55:30Z to **22:01:13Z ≈ 5m43s**. **This number is
contaminated by my own intervention** — I killed the local leg at 21:59:45, and the apps returned
shortly after — so it is an upper bound on the quiesce window, not a clean measurement of it.
4. **Alarm fired / true?** **`whole_guest_backup_failed` (error, operator-only) at 21:59:**
„Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)".
**True, and better than true: it names the TIER that failed rather than 'the backup', and it says
what it will do next.** The controller's own status surface agreed — `target_id:"local"`,
`success:false`, `size_bytes:0`, quoting the failing task's UPID. Nothing claimed success.
**No alarm fired for the twenty-two apps it stopped** — correct: the backup's own stack stops are
suppressed, so the box does not alarm about downtime it caused deliberately.
5. **Should have fired and did not.** **None.** The off-site tier succeeded and needed no alarm.
**Household loop:** kept sampling throughout, `ok http=301` every two minutes — but see its recorded
blind spot: it does not follow redirects, so it cannot see that the PUBLIC route was 404 while the
apps were stopped. Its silence here is not evidence the household was unaffected.
**The finding worth carrying forward:** a whole-guest backup on this box **cannot use its local tier**
(a ~29 GB source into a 14 GB root filesystem), and the product handles that honestly — it fails the
tier, says which tier, schedules a retry, and still gets the data out of the house on the off-site
tier. The local tier will keep retrying and keep failing on a box shaped like this one.
":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125
## A SAFETY GUARD added before round 7 — declared, and not a measurement
The failed local tier announced its own retry: „next attempt in **15m0s**" from 21:59, i.e. about
**22:14Z**, which falls inside round 7. That retry will fail exactly as before — a ~29 GB source into
a 14 GB root filesystem — but on the way it consumes `/` at ~16 MB/s, and round 7 has the box's
network cut for ten minutes, so nobody would be watching.
Rather than reach in a second time mid-round, a guard now runs on the box:
unit `diskguard`, active; floor **2500 MB**; free at start **7573 MB**
it kills a dump ONLY if free space falls under the floor AND a local `.tar.dat` is being written,
then removes the partial archive, and logs every action to /root/diskguard.log
**What it is and is not.** It is insurance against my fixture's known defect wrecking the remaining
rounds; it is **not** part of any round's result. If it ever fires, that is an intervention and will
be counted as one, with its log line quoted. If it never fires, it changed nothing. Either way the
product's own behaviour — fail the tier, name it, retry with backoff — is what is being measured, and
the guard does not touch the off-site tier at all.
@@ -0,0 +1,3 @@
2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7
2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up
2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes
@@ -0,0 +1,8 @@
2026-09-16T22:10:43Z ================ ROUND 7 : use nextcloud, while: internet-gone-10min ================
2026-09-16T22:10:46Z --- BEFORE --- containers=26 cloud=200 status=200 paste=200
2026-09-16T22:10:46Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T22:10:46Z --- ACTION: use on nextcloud ---
2026-09-16T22:10:46Z cloud read 1 -> 200
2026-09-16T22:10:47Z cloud read 2 -> 200
2026-09-16T22:10:47Z cloud read 3 -> 200
2026-09-16T22:10:47Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---