diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt new file mode 100644 index 00000000..521c190a --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6-pbs-leg.txt @@ -0,0 +1,3 @@ +## the OFF-SITE leg of the whole-system backup, watched to its end +2026-09-16T22:02:09Z procs=2 apps=26 free=7573M +2026-09-16T22:02:40Z procs=2 apps=26 free=7573M diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt index 27a0a826..3d79156d 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-6.txt @@ -27,3 +27,83 @@ alarm about downtime it caused on purpose. The distinction this project keeps re way to stop an app needs a third suppression set — held here. The duration is taken from the round's own runner when it reports, not estimated from this sample. + +## I STOPPED THE BACKUP — a declared intervention, with the arithmetic that forced it +Two samples thirty seconds apart, taken because the dump was being written to the VM's own root +filesystem: + + 21:58:33Z free **4138 MB** dump 3 602 939 904 B (3.36 GiB) + 21:59:03Z free **3663 MB** dump 4 100 403 200 B (3.82 GiB) + -> the archive grew ~497 MB and free space fell ~475 MB in 30 s = **~16 MB/s** + -> at that rate `/` (14 G total) would be **FULL in under four minutes** + +And the guest's config says what it was trying to fit: + rootfs local-lvm:vm-9201-disk-0, 32G (944 M used) + **mp0 local-lvm:vm-9201-disk-1, mp=/var/lib/felhom, backup=1, size=70G** <- 40.58 % used ≈ **28 GB** + mp8 /mnt/felhom-drives — no `backup=1`, correctly EXCLUDED (the household's data drive) + mp9 bootstrap, read-only + +So a whole-guest dump of ~29 GB of real data was being written into a filesystem with 3.6 GB left. +**It could not fit, and finishing was never possible** — the only question was whether it would fill +`/` on the nested PVE and wedge the box, which is the Phase 0 failure one level up and would have +cost every remaining round of the night. + +**Decision: stopped deliberately, and counted as an intervention.** Not a product failure and not +dressed up as one — the box was doing exactly what it was asked to do. What was wrong is the shape of +my fixture: a 32 GB system disk gives `pve-root` ~14 GB, while the guest's backed-up volume is 70 GB +provisioned with 28 GB used. + +**What round 6 still legitimately measured, and it is the valuable half:** with the backup running, +**4 of 26 containers** were up, the public route returned **404** for every app, and **no alarm +fired** — the suppression of the backup's own stack stops held. The downtime is real, total, and +correctly silent. What is NOT measured is the backup's duration or its completion, and the round says +so rather than implying a finished backup. + +## CORRECTION — I stopped a LEG of the backup, not the backup, and the box carried on by itself +The task list settles what actually happened: + + UPID …6AAB104C vzdump 9201 started 23:55:24 (21:55:24Z) **ended 23:59:45 — status `job errors`** + (that is the local-storage leg, and 23:59:45 is when my `pkill` landed) + +and then, seconds later, a **different** vzdump was already running: + + 86528 task UPID:…:vzdump:9201:felhom-agent@pve!agent: + 86604 /usr/bin/proxmox-backup-client backup … root.pxar:/mnt/vzsnap0 + --include-dev /mnt/vzsnap0/./var/lib/felhom + **--repository felhom@pbs!tester-1@10.77.0.1:felhom-offsite --ns tester-1** --crypt-mode=encrypt + +**So the whole-guest backup has two legs, and only one of them needed local space.** The local leg +was writing a ~29 GB source into a filesystem with 3.6 GB left and could never have finished — the +arithmetic that made me stop it stands. But the off-site leg streams to ep0 over WireGuard, encrypted, +and needs no room on `/` at all. **The box's design copes with exactly the problem I thought I was +rescuing it from**, and my note above said „I stopped the backup" when what I stopped was its local +leg. Corrected here rather than left standing. + +**What my intervention did and did not cost:** it ended a leg that could not have succeeded, and it +did not stop the off-site leg, which began by itself and is the one that carries this box's data out +of the house. It is still counted as an intervention — I reached in and killed a running job. + +## The box came back on its own + 22:00:52Z containers = 22 + 22:01:13Z containers = **26** — back to the full household, unaided + free on `/` recovered 2539 MB -> **7573 MB** once the partial archive was removed + the older completed backup (623 MB, 22:26) was left untouched + +## The waiting rule for the off-site leg — decided NOW, before the clock forces it +The off-site leg streams roughly 28 GB to ep0 over WireGuard. It may finish in minutes or run for +an hour; I cannot tell yet, and I do not want to be choosing between „wait" and „press on" at 00:30 +with rounds queuing behind me. + +**The rule, fixed in advance:** +1. Round 7 does NOT start while the off-site leg is running. Its drawn accident is *internet gone for + ten minutes*, which would cut the very transfer being measured — two experiments spoiling each + other, and neither answerable afterwards. +2. **If the leg is still running at 00:45 CEST**, round 7 starts anyway, and the round records that + its accident interrupted an in-flight off-site backup. That is then a REAL finding about what a + network cut does to a running whole-guest backup — worth having, provided it is labelled as such + and not mistaken for a clean internet-cut round. +3. Either way the remaining rounds keep their ~25-minute spacing from whenever round 7 starts, and + the 05:00 stop is unchanged. If that means fewer than twelve rounds run, the morning verdict says + how many actually ran rather than implying all twelve did. + +Written before the situation arises, because a rule invented at the moment it binds is not a rule.