diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 9c7af08e..432bf3b7 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -289,7 +289,27 @@ reach it from outside". It saw nothing while the public route was returning 530. started, and from it `controller_started` looked missing. A later reading shows it present at 21:53. An alarm cannot be called missing by a measurement taken before it could have fired. -### Rounds 6-12 +### Round 6 — `backup-system` (whole-guest backup) / accident: **none** (control round) + +| the five things | | +|---|---| +| what the customer saw | their apps went away and came back: during the backup **4 of 26** containers were up and every app answered **404** publicly; all 26 were serving again by 22:01:13Z. No banner, no mail — correctly | +| what the box did by itself | quiesced the apps, snapshotted, ran the local tier, **failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z (~8½ min)**, encrypted to ep0, taking **no local disk at all** | +| time to steady | apps down ~21:55:30Z → 22:01:13Z ≈ **5m43s** — **contaminated by my own intervention** (I killed the local leg at 21:59:45), so it is an upper bound on the quiesce window, not a clean measurement | +| alarm fired / true? | **`whole_guest_backup_failed` (error) — true, and better than true:** „Whole-guest backup FAILED on the **local tier** — retrying with backoff (next attempt in 15m0s)". It names the tier, not „the backup", and says what it will do next. The status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`). **No alarm for the 22 apps it stopped** — correct, those stops are suppressed | +| should have fired, did not | **none** | + +**The finding to carry forward:** a whole-guest backup on this box **cannot use its local tier** — a +~29 GB source into a 14 GB root filesystem — and the product handles that honestly: it fails the tier, +says which tier, schedules a retry, and still gets the data out of the house on the off-site tier. +The local tier will keep retrying and keep failing on a box shaped like this one. + +**My intervention, and the correction it needed.** I killed the local leg when two samples showed `/` +falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nested PVE. The arithmetic +stands, but I first wrote „I stopped the backup", which was wrong: I stopped **one leg**, and the +off-site leg started by itself seconds later and succeeded. Counted as an intervention either way. + +### Rounds 7-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt index 959dd58c..389b3d0b 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt @@ -3,3 +3,17 @@ 2026-09-16T22:02:40Z procs=2 apps=26 free=7573M 2026-09-16T22:03:12Z procs=2 apps=26 free=7573M 2026-09-16T22:03:43Z procs=2 apps=26 free=7573M +2026-09-16T22:04:15Z procs=2 apps=26 free=7573M +2026-09-16T22:04:46Z procs=4 apps=26 free=7573M +2026-09-16T22:05:18Z procs=2 apps=26 free=7573M +2026-09-16T22:05:49Z procs=2 apps=26 free=7573M +2026-09-16T22:06:21Z procs=2 apps=26 free=7573M +2026-09-16T22:06:52Z procs=2 apps=26 free=7573M +2026-09-16T22:07:24Z procs=2 apps=26 free=7573M +2026-09-16T22:07:55Z procs=2 apps=26 free=7573M +2026-09-16T22:08:27Z procs=0 apps=26 free=7573M +OFF-SITE LEG FINISHED at 2026-09-16T22:08:27Z +--- what the task list says about it --- +├──────┼────────────┼───────┼────────┼─────────────────────┼────────┼────────────────────────────────────────────────────────────────────────────────┼──────────────────┼─────────────────────┼────────────┤ +│ 9201 │ chaosnight │ 82472 │ 175208 │ 2026-09-16 23:55:24 │ vzdump │ UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felhom-agent@pve!agent: │ felhom-agent@pve │ 2026-09-16 23:59:45 │ job errors │ +└──────┴────────────┴───────┴────────┴─────────────────────┴────────┴────────────────────────────────────────────────────────────────────────────────┴──────────────────┴─────────────────────┴────────────┘ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt index 1d37e0d5..8ddeed15 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt @@ -154,4 +154,70 @@ product's own path and read only when nothing else can answer. {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho - {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent \ No newline at end of file + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent +## ROUND 6 — the five things +**action:** `backup-system` (whole-guest backup; drawn app adventurelog is which row is read after) +**accident:** none — control round + +1. **What the customer saw.** Their apps went away and came back. During the backup **4 of 26** + containers were up and every app answered **404** through the public route; by 22:01:13Z all 26 were + back and serving. No banner, no mail — and correctly so. +2. **What the box did by itself.** Quiesced the apps, took an LVM snapshot, ran the local tier, + **failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the + snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z, about 8½ minutes**, encrypted to + ep0, consuming **no local disk at all** (free stayed at 7573M throughout). +3. **Time to steady.** Apps down from ~21:55:30Z to **22:01:13Z ≈ 5m43s**. **This number is + contaminated by my own intervention** — I killed the local leg at 21:59:45, and the apps returned + shortly after — so it is an upper bound on the quiesce window, not a clean measurement of it. +4. **Alarm fired / true?** **`whole_guest_backup_failed` (error, operator-only) at 21:59:** + „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)". + **True, and better than true: it names the TIER that failed rather than 'the backup', and it says + what it will do next.** The controller's own status surface agreed — `target_id:"local"`, + `success:false`, `size_bytes:0`, quoting the failing task's UPID. Nothing claimed success. + **No alarm fired for the twenty-two apps it stopped** — correct: the backup's own stack stops are + suppressed, so the box does not alarm about downtime it caused deliberately. +5. **Should have fired and did not.** **None.** The off-site tier succeeded and needed no alarm. + +**Household loop:** kept sampling throughout, `ok http=301` every two minutes — but see its recorded +blind spot: it does not follow redirects, so it cannot see that the PUBLIC route was 404 while the +apps were stopped. Its silence here is not evidence the household was unaffected. + +**The finding worth carrying forward:** a whole-guest backup on this box **cannot use its local tier** +(a ~29 GB source into a 14 GB root filesystem), and the product handles that honestly — it fails the +tier, says which tier, schedules a retry, and still gets the data out of the house on the off-site +tier. The local tier will keep retrying and keep failing on a box shaped like this one. +":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho + {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1 + {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1 + {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1 + {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1 + {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1 + {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125 +## A SAFETY GUARD added before round 7 — declared, and not a measurement +The failed local tier announced its own retry: „next attempt in **15m0s**" from 21:59, i.e. about +**22:14Z**, which falls inside round 7. That retry will fail exactly as before — a ~29 GB source into +a 14 GB root filesystem — but on the way it consumes `/` at ~16 MB/s, and round 7 has the box's +network cut for ten minutes, so nobody would be watching. + +Rather than reach in a second time mid-round, a guard now runs on the box: + unit `diskguard`, active; floor **2500 MB**; free at start **7573 MB** + it kills a dump ONLY if free space falls under the floor AND a local `.tar.dat` is being written, + then removes the partial archive, and logs every action to /root/diskguard.log + +**What it is and is not.** It is insurance against my fixture's known defect wrecking the remaining +rounds; it is **not** part of any round's result. If it ever fires, that is an intervention and will +be counted as one, with its log line quoted. If it never fires, it changed nothing. Either way the +product's own behaviour — fail the tier, name it, retry with backoff — is what is being measured, and +the guard does not touch the off-site tier at all. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt new file mode 100644 index 00000000..75796e29 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-7-accident.txt @@ -0,0 +1,3 @@ +2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7 +2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up +2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt new file mode 100644 index 00000000..c4f40743 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-7.txt @@ -0,0 +1,8 @@ +2026-09-16T22:10:43Z ================ ROUND 7 : use nextcloud, while: internet-gone-10min ================ +2026-09-16T22:10:46Z --- BEFORE --- containers=26 cloud=200 status=200 paste=200 +2026-09-16T22:10:46Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) +2026-09-16T22:10:46Z --- ACTION: use on nextcloud --- +2026-09-16T22:10:46Z cloud read 1 -> 200 +2026-09-16T22:10:47Z cloud read 2 -> 200 +2026-09-16T22:10:47Z cloud read 3 -> 200 +2026-09-16T22:10:47Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---