CHAOS NIGHT round 6: I stopped a LEG, and the box carried on by itself
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
Correction to my own account, recorded where the wrong version stood. I wrote "I stopped the backup". What I stopped was its LOCAL leg: the task list shows that job ending 23:59:45 with status "job errors" - my pkill - while a second vzdump was already running, streaming encrypted to ep0 (--repository felhom@pbs!tester-1@...:felhom-offsite --ns tester-1). The arithmetic that forced the intervention still stands: the local leg was writing a ~29GB source into a filesystem with 3.6GB free, falling at ~16MB/s, which gave under four minutes before / filled and the nested PVE wedged - the Phase 0 failure one level up. But the off-site leg needs no local space at all, so the box's design copes with exactly the problem I thought I was rescuing it from. It is still counted as an intervention: I reached in and killed a job. The box then recovered unaided: 22 containers at 22:00:52Z, 26 at 22:01:13Z, and / went from 2539MB free back to 7573MB once the partial archive was removed. The older completed backup was left untouched. Round 6's valuable half stands: during the backup 4 of 26 containers were up, every app returned 404 through the public route, and NO alarm fired - the suppression of the backup's own stack stops held. And the waiting rule for the off-site leg is fixed in advance: round 7 does not start while it runs, unless it is still running at 00:45, in which case round 7 proceeds and records that it cut an in-flight backup - labelled as that, not as a clean internet-cut round. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,3 @@
|
||||
## the OFF-SITE leg of the whole-system backup, watched to its end
|
||||
2026-09-16T22:02:09Z procs=2 apps=26 free=7573M
|
||||
2026-09-16T22:02:40Z procs=2 apps=26 free=7573M
|
||||
@@ -27,3 +27,83 @@ alarm about downtime it caused on purpose. The distinction this project keeps re
|
||||
way to stop an app needs a third suppression set — held here.
|
||||
|
||||
The duration is taken from the round's own runner when it reports, not estimated from this sample.
|
||||
|
||||
## I STOPPED THE BACKUP — a declared intervention, with the arithmetic that forced it
|
||||
Two samples thirty seconds apart, taken because the dump was being written to the VM's own root
|
||||
filesystem:
|
||||
|
||||
21:58:33Z free **4138 MB** dump 3 602 939 904 B (3.36 GiB)
|
||||
21:59:03Z free **3663 MB** dump 4 100 403 200 B (3.82 GiB)
|
||||
-> the archive grew ~497 MB and free space fell ~475 MB in 30 s = **~16 MB/s**
|
||||
-> at that rate `/` (14 G total) would be **FULL in under four minutes**
|
||||
|
||||
And the guest's config says what it was trying to fit:
|
||||
rootfs local-lvm:vm-9201-disk-0, 32G (944 M used)
|
||||
**mp0 local-lvm:vm-9201-disk-1, mp=/var/lib/felhom, backup=1, size=70G** <- 40.58 % used ≈ **28 GB**
|
||||
mp8 /mnt/felhom-drives — no `backup=1`, correctly EXCLUDED (the household's data drive)
|
||||
mp9 bootstrap, read-only
|
||||
|
||||
So a whole-guest dump of ~29 GB of real data was being written into a filesystem with 3.6 GB left.
|
||||
**It could not fit, and finishing was never possible** — the only question was whether it would fill
|
||||
`/` on the nested PVE and wedge the box, which is the Phase 0 failure one level up and would have
|
||||
cost every remaining round of the night.
|
||||
|
||||
**Decision: stopped deliberately, and counted as an intervention.** Not a product failure and not
|
||||
dressed up as one — the box was doing exactly what it was asked to do. What was wrong is the shape of
|
||||
my fixture: a 32 GB system disk gives `pve-root` ~14 GB, while the guest's backed-up volume is 70 GB
|
||||
provisioned with 28 GB used.
|
||||
|
||||
**What round 6 still legitimately measured, and it is the valuable half:** with the backup running,
|
||||
**4 of 26 containers** were up, the public route returned **404** for every app, and **no alarm
|
||||
fired** — the suppression of the backup's own stack stops held. The downtime is real, total, and
|
||||
correctly silent. What is NOT measured is the backup's duration or its completion, and the round says
|
||||
so rather than implying a finished backup.
|
||||
|
||||
## CORRECTION — I stopped a LEG of the backup, not the backup, and the box carried on by itself
|
||||
The task list settles what actually happened:
|
||||
|
||||
UPID …6AAB104C vzdump 9201 started 23:55:24 (21:55:24Z) **ended 23:59:45 — status `job errors`**
|
||||
(that is the local-storage leg, and 23:59:45 is when my `pkill` landed)
|
||||
|
||||
and then, seconds later, a **different** vzdump was already running:
|
||||
|
||||
86528 task UPID:…:vzdump:9201:felhom-agent@pve!agent:
|
||||
86604 /usr/bin/proxmox-backup-client backup … root.pxar:/mnt/vzsnap0
|
||||
--include-dev /mnt/vzsnap0/./var/lib/felhom
|
||||
**--repository felhom@pbs!tester-1@10.77.0.1:felhom-offsite --ns tester-1** --crypt-mode=encrypt
|
||||
|
||||
**So the whole-guest backup has two legs, and only one of them needed local space.** The local leg
|
||||
was writing a ~29 GB source into a filesystem with 3.6 GB left and could never have finished — the
|
||||
arithmetic that made me stop it stands. But the off-site leg streams to ep0 over WireGuard, encrypted,
|
||||
and needs no room on `/` at all. **The box's design copes with exactly the problem I thought I was
|
||||
rescuing it from**, and my note above said „I stopped the backup" when what I stopped was its local
|
||||
leg. Corrected here rather than left standing.
|
||||
|
||||
**What my intervention did and did not cost:** it ended a leg that could not have succeeded, and it
|
||||
did not stop the off-site leg, which began by itself and is the one that carries this box's data out
|
||||
of the house. It is still counted as an intervention — I reached in and killed a running job.
|
||||
|
||||
## The box came back on its own
|
||||
22:00:52Z containers = 22
|
||||
22:01:13Z containers = **26** — back to the full household, unaided
|
||||
free on `/` recovered 2539 MB -> **7573 MB** once the partial archive was removed
|
||||
the older completed backup (623 MB, 22:26) was left untouched
|
||||
|
||||
## The waiting rule for the off-site leg — decided NOW, before the clock forces it
|
||||
The off-site leg streams roughly 28 GB to ep0 over WireGuard. It may finish in minutes or run for
|
||||
an hour; I cannot tell yet, and I do not want to be choosing between „wait" and „press on" at 00:30
|
||||
with rounds queuing behind me.
|
||||
|
||||
**The rule, fixed in advance:**
|
||||
1. Round 7 does NOT start while the off-site leg is running. Its drawn accident is *internet gone for
|
||||
ten minutes*, which would cut the very transfer being measured — two experiments spoiling each
|
||||
other, and neither answerable afterwards.
|
||||
2. **If the leg is still running at 00:45 CEST**, round 7 starts anyway, and the round records that
|
||||
its accident interrupted an in-flight off-site backup. That is then a REAL finding about what a
|
||||
network cut does to a running whole-guest backup — worth having, provided it is labelled as such
|
||||
and not mistaken for a clean internet-cut round.
|
||||
3. Either way the remaining rounds keep their ~25-minute spacing from whenever round 7 starts, and
|
||||
the 05:00 stop is unchanged. If that means fewer than twelve rounds run, the morning verdict says
|
||||
how many actually ran rather than implying all twelve did.
|
||||
|
||||
Written before the situation arises, because a rule invented at the moment it binds is not a rule.
|
||||
|
||||
Reference in New Issue
Block a user