From 36ae3b3191df03ea088793eefa4f148c8fac9aba Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 00:02:46 +0200 Subject: [PATCH] CHAOS NIGHT round 6: I stopped a LEG, and the box carried on by itself Correction to my own account, recorded where the wrong version stood. I wrote "I stopped the backup". What I stopped was its LOCAL leg: the task list shows that job ending 23:59:45 with status "job errors" - my pkill - while a second vzdump was already running, streaming encrypted to ep0 (--repository felhom@pbs!tester-1@...:felhom-offsite --ns tester-1). The arithmetic that forced the intervention still stands: the local leg was writing a ~29GB source into a filesystem with 3.6GB free, falling at ~16MB/s, which gave under four minutes before / filled and the nested PVE wedged - the Phase 0 failure one level up. But the off-site leg needs no local space at all, so the box's design copes with exactly the problem I thought I was rescuing it from. It is still counted as an intervention: I reached in and killed a job. The box then recovered unaided: 22 containers at 22:00:52Z, 26 at 22:01:13Z, and / went from 2539MB free back to 7573MB once the partial archive was removed. The older completed backup was left untouched. Round 6's valuable half stands: during the backup 4 of 26 containers were up, every app returned 404 through the public route, and NO alarm fired - the suppression of the backup's own stack stops held. And the waiting rule for the off-site leg is fixed in advance: round 7 does not start while it runs, unless it is still running at 00:45, in which case round 7 proceeds and records that it cut an in-flight backup - labelled as that, not as a clean internet-cut round. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../round-6-pbs-leg.txt | 3 + .../round-6.txt | 80 +++++++++++++++++++ 2 files changed, 83 insertions(+) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt new file mode 100644 index 00000000..521c190a --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt @@ -0,0 +1,3 @@ +## the OFF-SITE leg of the whole-system backup, watched to its end +2026-09-16T22:02:09Z procs=2 apps=26 free=7573M +2026-09-16T22:02:40Z procs=2 apps=26 free=7573M diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt index 27a0a826..3d79156d 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt @@ -27,3 +27,83 @@ alarm about downtime it caused on purpose. The distinction this project keeps re way to stop an app needs a third suppression set — held here. The duration is taken from the round's own runner when it reports, not estimated from this sample. + +## I STOPPED THE BACKUP — a declared intervention, with the arithmetic that forced it +Two samples thirty seconds apart, taken because the dump was being written to the VM's own root +filesystem: + + 21:58:33Z free **4138 MB** dump 3 602 939 904 B (3.36 GiB) + 21:59:03Z free **3663 MB** dump 4 100 403 200 B (3.82 GiB) + -> the archive grew ~497 MB and free space fell ~475 MB in 30 s = **~16 MB/s** + -> at that rate `/` (14 G total) would be **FULL in under four minutes** + +And the guest's config says what it was trying to fit: + rootfs local-lvm:vm-9201-disk-0, 32G (944 M used) + **mp0 local-lvm:vm-9201-disk-1, mp=/var/lib/felhom, backup=1, size=70G** <- 40.58 % used ≈ **28 GB** + mp8 /mnt/felhom-drives — no `backup=1`, correctly EXCLUDED (the household's data drive) + mp9 bootstrap, read-only + +So a whole-guest dump of ~29 GB of real data was being written into a filesystem with 3.6 GB left. +**It could not fit, and finishing was never possible** — the only question was whether it would fill +`/` on the nested PVE and wedge the box, which is the Phase 0 failure one level up and would have +cost every remaining round of the night. + +**Decision: stopped deliberately, and counted as an intervention.** Not a product failure and not +dressed up as one — the box was doing exactly what it was asked to do. What was wrong is the shape of +my fixture: a 32 GB system disk gives `pve-root` ~14 GB, while the guest's backed-up volume is 70 GB +provisioned with 28 GB used. + +**What round 6 still legitimately measured, and it is the valuable half:** with the backup running, +**4 of 26 containers** were up, the public route returned **404** for every app, and **no alarm +fired** — the suppression of the backup's own stack stops held. The downtime is real, total, and +correctly silent. What is NOT measured is the backup's duration or its completion, and the round says +so rather than implying a finished backup. + +## CORRECTION — I stopped a LEG of the backup, not the backup, and the box carried on by itself +The task list settles what actually happened: + + UPID …6AAB104C vzdump 9201 started 23:55:24 (21:55:24Z) **ended 23:59:45 — status `job errors`** + (that is the local-storage leg, and 23:59:45 is when my `pkill` landed) + +and then, seconds later, a **different** vzdump was already running: + + 86528 task UPID:…:vzdump:9201:felhom-agent@pve!agent: + 86604 /usr/bin/proxmox-backup-client backup … root.pxar:/mnt/vzsnap0 + --include-dev /mnt/vzsnap0/./var/lib/felhom + **--repository felhom@pbs!tester-1@10.77.0.1:felhom-offsite --ns tester-1** --crypt-mode=encrypt + +**So the whole-guest backup has two legs, and only one of them needed local space.** The local leg +was writing a ~29 GB source into a filesystem with 3.6 GB left and could never have finished — the +arithmetic that made me stop it stands. But the off-site leg streams to ep0 over WireGuard, encrypted, +and needs no room on `/` at all. **The box's design copes with exactly the problem I thought I was +rescuing it from**, and my note above said „I stopped the backup" when what I stopped was its local +leg. Corrected here rather than left standing. + +**What my intervention did and did not cost:** it ended a leg that could not have succeeded, and it +did not stop the off-site leg, which began by itself and is the one that carries this box's data out +of the house. It is still counted as an intervention — I reached in and killed a running job. + +## The box came back on its own + 22:00:52Z containers = 22 + 22:01:13Z containers = **26** — back to the full household, unaided + free on `/` recovered 2539 MB -> **7573 MB** once the partial archive was removed + the older completed backup (623 MB, 22:26) was left untouched + +## The waiting rule for the off-site leg — decided NOW, before the clock forces it +The off-site leg streams roughly 28 GB to ep0 over WireGuard. It may finish in minutes or run for +an hour; I cannot tell yet, and I do not want to be choosing between „wait" and „press on" at 00:30 +with rounds queuing behind me. + +**The rule, fixed in advance:** +1. Round 7 does NOT start while the off-site leg is running. Its drawn accident is *internet gone for + ten minutes*, which would cut the very transfer being measured — two experiments spoiling each + other, and neither answerable afterwards. +2. **If the leg is still running at 00:45 CEST**, round 7 starts anyway, and the round records that + its accident interrupted an in-flight off-site backup. That is then a REAL finding about what a + network cut does to a running whole-guest backup — worth having, provided it is labelled as such + and not mistaken for a clean internet-cut round. +3. Either way the remaining rounds keep their ~25-minute spacing from whenever round 7 starts, and + the 05:00 stop is unchanged. If that means fewer than twelve rounds run, the morning verdict says + how many actually ran rather than implying all twelve did. + +Written before the situation arises, because a rule invented at the moment it binds is not a rule.