CHAOS NIGHT round 6: I stopped a LEG, and the box carried on by itself
gates / gates (push) Successful in 21s

Correction to my own account, recorded where the wrong version stood. I wrote
"I stopped the backup". What I stopped was its LOCAL leg: the task list shows
that job ending 23:59:45 with status "job errors" - my pkill - while a second
vzdump was already running, streaming encrypted to ep0
(--repository felhom@pbs!tester-1@...:felhom-offsite --ns tester-1).

The arithmetic that forced the intervention still stands: the local leg was
writing a ~29GB source into a filesystem with 3.6GB free, falling at ~16MB/s,
which gave under four minutes before / filled and the nested PVE wedged - the
Phase 0 failure one level up. But the off-site leg needs no local space at all,
so the box's design copes with exactly the problem I thought I was rescuing it
from. It is still counted as an intervention: I reached in and killed a job.

The box then recovered unaided: 22 containers at 22:00:52Z, 26 at 22:01:13Z,
and / went from 2539MB free back to 7573MB once the partial archive was removed.
The older completed backup was left untouched.

Round 6's valuable half stands: during the backup 4 of 26 containers were up,
every app returned 404 through the public route, and NO alarm fired - the
suppression of the backup's own stack stops held.

And the waiting rule for the off-site leg is fixed in advance: round 7 does not
start while it runs, unless it is still running at 00:45, in which case round 7
proceeds and records that it cut an in-flight backup - labelled as that, not as
a clean internet-cut round.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 00:02:46 +02:00
parent d4318529b4
commit 36ae3b3191
2 changed files with 83 additions and 0 deletions
@@ -0,0 +1,3 @@
## the OFF-SITE leg of the whole-system backup, watched to its end
2026-09-16T22:02:09Z procs=2 apps=26 free=7573M
2026-09-16T22:02:40Z procs=2 apps=26 free=7573M
@@ -27,3 +27,83 @@ alarm about downtime it caused on purpose. The distinction this project keeps re
way to stop an app needs a third suppression set — held here.
The duration is taken from the round's own runner when it reports, not estimated from this sample.
## I STOPPED THE BACKUP — a declared intervention, with the arithmetic that forced it
Two samples thirty seconds apart, taken because the dump was being written to the VM's own root
filesystem:
21:58:33Z free **4138 MB** dump 3 602 939 904 B (3.36 GiB)
21:59:03Z free **3663 MB** dump 4 100 403 200 B (3.82 GiB)
-> the archive grew ~497 MB and free space fell ~475 MB in 30 s = **~16 MB/s**
-> at that rate `/` (14 G total) would be **FULL in under four minutes**
And the guest's config says what it was trying to fit:
rootfs local-lvm:vm-9201-disk-0, 32G (944 M used)
**mp0 local-lvm:vm-9201-disk-1, mp=/var/lib/felhom, backup=1, size=70G** <- 40.58 % used ≈ **28 GB**
mp8 /mnt/felhom-drives — no `backup=1`, correctly EXCLUDED (the household's data drive)
mp9 bootstrap, read-only
So a whole-guest dump of ~29 GB of real data was being written into a filesystem with 3.6 GB left.
**It could not fit, and finishing was never possible** — the only question was whether it would fill
`/` on the nested PVE and wedge the box, which is the Phase 0 failure one level up and would have
cost every remaining round of the night.
**Decision: stopped deliberately, and counted as an intervention.** Not a product failure and not
dressed up as one — the box was doing exactly what it was asked to do. What was wrong is the shape of
my fixture: a 32 GB system disk gives `pve-root` ~14 GB, while the guest's backed-up volume is 70 GB
provisioned with 28 GB used.
**What round 6 still legitimately measured, and it is the valuable half:** with the backup running,
**4 of 26 containers** were up, the public route returned **404** for every app, and **no alarm
fired** — the suppression of the backup's own stack stops held. The downtime is real, total, and
correctly silent. What is NOT measured is the backup's duration or its completion, and the round says
so rather than implying a finished backup.
## CORRECTION — I stopped a LEG of the backup, not the backup, and the box carried on by itself
The task list settles what actually happened:
UPID …6AAB104C vzdump 9201 started 23:55:24 (21:55:24Z) **ended 23:59:45 — status `job errors`**
(that is the local-storage leg, and 23:59:45 is when my `pkill` landed)
and then, seconds later, a **different** vzdump was already running:
86528 task UPID:…:vzdump:9201:felhom-agent@pve!agent:
86604 /usr/bin/proxmox-backup-client backup … root.pxar:/mnt/vzsnap0
--include-dev /mnt/vzsnap0/./var/lib/felhom
**--repository felhom@pbs!tester-1@10.77.0.1:felhom-offsite --ns tester-1** --crypt-mode=encrypt
**So the whole-guest backup has two legs, and only one of them needed local space.** The local leg
was writing a ~29 GB source into a filesystem with 3.6 GB left and could never have finished — the
arithmetic that made me stop it stands. But the off-site leg streams to ep0 over WireGuard, encrypted,
and needs no room on `/` at all. **The box's design copes with exactly the problem I thought I was
rescuing it from**, and my note above said „I stopped the backup" when what I stopped was its local
leg. Corrected here rather than left standing.
**What my intervention did and did not cost:** it ended a leg that could not have succeeded, and it
did not stop the off-site leg, which began by itself and is the one that carries this box's data out
of the house. It is still counted as an intervention — I reached in and killed a running job.
## The box came back on its own
22:00:52Z containers = 22
22:01:13Z containers = **26** — back to the full household, unaided
free on `/` recovered 2539 MB -> **7573 MB** once the partial archive was removed
the older completed backup (623 MB, 22:26) was left untouched
## The waiting rule for the off-site leg — decided NOW, before the clock forces it
The off-site leg streams roughly 28 GB to ep0 over WireGuard. It may finish in minutes or run for
an hour; I cannot tell yet, and I do not want to be choosing between „wait" and „press on" at 00:30
with rounds queuing behind me.
**The rule, fixed in advance:**
1. Round 7 does NOT start while the off-site leg is running. Its drawn accident is *internet gone for
ten minutes*, which would cut the very transfer being measured — two experiments spoiling each
other, and neither answerable afterwards.
2. **If the leg is still running at 00:45 CEST**, round 7 starts anyway, and the round records that
its accident interrupted an in-flight off-site backup. That is then a REAL finding about what a
network cut does to a running whole-guest backup — worth having, provided it is labelled as such
and not mistaken for a clean internet-cut round.
3. Either way the remaining rounds keep their ~25-minute spacing from whenever round 7 starts, and
the 05:00 stop is unchanged. If that means fewer than twelve rounds run, the morning verdict says
how many actually ran rather than implying all twelve did.
Written before the situation arises, because a rule invented at the moment it binds is not a rule.