Files
felhom.eu/documentation/audits
admin 36ae3b3191
gates / gates (push) Successful in 21s
CHAOS NIGHT round 6: I stopped a LEG, and the box carried on by itself
Correction to my own account, recorded where the wrong version stood. I wrote
"I stopped the backup". What I stopped was its LOCAL leg: the task list shows
that job ending 23:59:45 with status "job errors" - my pkill - while a second
vzdump was already running, streaming encrypted to ep0
(--repository felhom@pbs!tester-1@...:felhom-offsite --ns tester-1).

The arithmetic that forced the intervention still stands: the local leg was
writing a ~29GB source into a filesystem with 3.6GB free, falling at ~16MB/s,
which gave under four minutes before / filled and the nested PVE wedged - the
Phase 0 failure one level up. But the off-site leg needs no local space at all,
so the box's design copes with exactly the problem I thought I was rescuing it
from. It is still counted as an intervention: I reached in and killed a job.

The box then recovered unaided: 22 containers at 22:00:52Z, 26 at 22:01:13Z,
and / went from 2539MB free back to 7573MB once the partial archive was removed.
The older completed backup was left untouched.

Round 6's valuable half stands: during the backup 4 of 26 containers were up,
every app returned 404 through the public route, and NO alarm fired - the
suppression of the backup's own stack stops held.

And the waiting rule for the off-site leg is fixed in advance: round 7 does not
start while it runs, unless it is still running at 00:45, in which case round 7
proceeds and records that it cut an in-flight backup - labelled as that, not as
a clean internet-cut round.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 00:02:46 +02:00
..
@
2026-07-14 00:45:20 +02:00
@
2026-07-14 10:12:21 +02:00
@
2026-07-14 12:17:26 +02:00
@
2026-07-14 12:17:26 +02:00