Files
felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/diskguard.sh
T
admin 9f40dc3289
gates / gates (push) Successful in 21s
chaos night round 10: a restore leaves no record, and four of my instruments failed
The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.

R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.

Four instrument faults, all mine, all in the evidence:
  * a 'nothing was logged' claim that was unfalsifiable when written - the log
    stream holds zero lines before a reset;
  * an on-disk check against /opt/felhom/data, a directory that does not exist;
  * a household count reporting 0 lines and 0 failures when the truth was one
    line and it WAS a failure - the runner now prints both operands;
  * the disk guard was a TRANSIENT unit reporting 'active' all night, and was
    absent from the reset onward. It is now file-backed and enabled, and its
    script is copied off the box for the first time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 01:35:56 +02:00

21 lines
1.0 KiB
Bash

#!/bin/bash
# SAFETY, not a measurement: the local backup tier retries every ~15 minutes and cannot fit on this
# box (a ~29 GB source into a 14 GB root filesystem). Left alone during a round it would fill `/` and
# wedge the nested PVE, which already nearly cost the night once. This kills ONLY a local-tier dump,
# and only when free space falls under the floor. Every kill is logged so it appears in the evidence.
FLOOR_MB=2500
LOG=/root/diskguard.log
while true; do
FREE=$(df -BM --output=avail / | tail -1 | tr -dc '0-9')
if [ -n "$FREE" ] && [ "$FREE" -lt "$FLOOR_MB" ]; then
if ls /var/lib/vz/dump/*.tar.dat >/dev/null 2>&1; then
echo "$(date -u +%FT%TZ) GUARD: free=${FREE}M below ${FLOOR_MB}M and a local dump is writing — killing it" >> $LOG
pkill -f "[v]zdump"
sleep 3
rm -rf /var/lib/vz/dump/*.tar.dat /var/lib/vz/dump/*.tmp 2>/dev/null
echo "$(date -u +%FT%TZ) GUARD: partial archive removed, free now $(df -BM --output=avail / | tail -1)" >> $LOG
fi
fi
sleep 20
done