chaos night teardown: machine destroyed, host clean, 9202 untouched, ep0 unchanged
gates / gates (push) Successful in 21s

Machine layer: VM 336 destroyed with all three disks purged, gated on its NAME
rather than its number because 9201 and 9202 share the host. qm list now shows
no VMs; /mnt/hdd_1/images/336 is gone; both standing guests still run.

Host layer, before and after: nvme-scratch 6.78% -> 1.61% (~48.5 GB released),
/mnt/hdd_1/images 59G -> 9.4G with only the scratch guest's own 9202 directory
left, free space 827G -> 875G. local-lvm UNCHANGED at 44.75% - the fence that
said 'local-lvm never' held. Firewall back at baseline with 0 physdev rules, so
none of the three network accidents left a rule on a host carrying two standing
guests. Both harness units stopped and disabled before the box died; the disk
guard's log was 0 bytes - it never fired once.

9202: nothing to remove, shown rather than asserted - three infrastructure
containers, 55 catalog TEMPLATES none of which was touched since 21:00, no
app.yaml marked deployed, no offbox config. A false label in my own transcript
is corrected there: I printed '(nothing listed above = no app stacks)' directly
beneath 55 names.

ep0: read again immediately before the delete and identical to the baseline.
The single-host delete handler shows no ep0 cascade, but a grep returning
nothing is the weakest evidence there is, so the store gets a before and an
after rather than an inference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 02:37:49 +02:00
parent 9b44c44f23
commit 0f65c8121d
4 changed files with 109 additions and 0 deletions
@@ -15,3 +15,19 @@ disk /dev/sdb 98G total, 16G used, 83G available, 16% used
At teardown this listing is repeated. The expected result is that every group and snapshot above is
still present; the tester-1 group may have MORE snapshots if a backup ran, and must never have fewer.
## SECOND READING, taken 2026-09-17T00:36:52Z - immediately BEFORE the hub host delete
namespaces demo-felhom, demo-hp, tester-1
groups ns/demo-felhom/ct/9201 -> 2 snapshot(s)
ns/demo-hp/ct/9201 -> 2 snapshot(s)
ns/tester-1/ct/9201 -> 2 snapshot(s)
disk 98G total, 16G used, 83G available, 16%
IDENTICAL to the 00:17:15Z baseline. Nothing changed on ep0 during the whole teardown of the machine
and the host - which is what "read only" is supposed to mean, and is now shown rather than asserted.
WHY A SECOND READING BEFORE THE DELETE: this project already knows that a delete's single removal
operation can destroy a customer's backup groups, not merely release a token. The single-host delete
handler shows no cascade to ep0 - but a grep that returns nothing is the weakest evidence there is,
and tonight has produced four false conclusions built on exactly that. So the store is read again a
few minutes before the delete, and will be read again after it. If the delete touches ep0, the
damage will be measurable against a reading minutes old instead of an hour old.
@@ -0,0 +1,24 @@
# TEARDOWN, scratch guest 9202: there is nothing to remove - measured, 2026-09-17T00:34:39Z
The brief's teardown asks for "scratch 9202 (restored app removed)". No restore was ever possible
(the off-site repository is orphaned and holds zero readable snapshots - see
phase2-offsite-restore-blocked.txt), so nothing was ever put on 9202. That is shown, not asserted:
running containers felhom-controller, filebrowser, traefik <- infrastructure only
/opt/docker/stacks 55 entries
touched since 21:00 NONE (find -newermt "2026-09-16 21:00" returns nothing)
marked deployed NONE (no app.yaml contains desired_state: deployed)
offbox directory absent - an off-site target was never configured on this guest
## A FALSE LABEL IN MY OWN TRANSCRIPT, corrected here
My first listing printed the 55 directory names and then the line
"(nothing listed above = no app stacks)"
directly underneath them. The label contradicted the output immediately above it. The 55 entries are
CATALOG TEMPLATES synced by the controller - the library of apps the box could install - not apps
that are installed. The distinction matters: a reader skimming that transcript would have seen 55
app names under a heading saying there were none.
`bentopdf` is among those templates and is untouched, as the fence requires.
## WHAT THIS MEANS FOR THE TEARDOWN LINE
9202 is returned exactly as it was found: three infrastructure containers, its template library, no
apps, no off-site configuration. Nothing was created on it tonight and nothing was removed from it.
@@ -0,0 +1,36 @@
# TEARDOWN LAYER 2 - the host (demo-hp). Before at 00:30:34Z, after at 00:36:07Z.
## pvesm status, side by side
storage BEFORE used AFTER used change
felhom-pbs 0 0 -
local 29,325,896 KiB 72.49% 29,325,928 KiB 72.49% unchanged (not the drill's storage)
local-lvm 25,278,351 KiB 44.75% 25,278,351 KiB 44.75% UNCHANGED - the fence held, the
drill never used local-lvm
nvme-scratch 66,624,884 KiB 6.78% 15,850,280 KiB 1.61% ~48.5 GB RELEASED
## /mnt/hdd_1, the storage the brief named
images/ 59 G -> 9.4 G and images/336 is GONE ("No such file or directory")
the only directory left there is 9202 - the SCRATCH GUEST's
own disk, which must stay
dump 4.8 G -> 4.8 G untouched
felhom-data 994 M -> 996 M untouched (grew by its own accord, not by the drill)
free space 827 G -> 875 G
## guests
BEFORE VM 336 tester1-chaos-night running + CT 9201 demo-hp + CT 9202 demo-hp-scratch
AFTER no VMs at all + CT 9201 running + CT 9202 running
Both standing guests survived, which is the fence that mattered most on this host.
## the loop, the hog, and the rules
household unit stopped AND disabled on the box before the box was destroyed
diskguard unit stopped AND disabled; its log was 0 bytes - it never fired once all night
leftover files none: no *chaos* or *hog* file under /root or /tmp on demo-hp
firewall -P FORWARD ACCEPT, physdev rules: 0, bridge-nf sysctl: 0
i.e. exactly the baseline, so none of the three network accidents left a rule
behind on a host that also carries two standing guests
## WHAT IS NOT PROVEN HERE
"No leftover drill files" rests on looking at /root and /tmp at depth 1 on demo-hp. It is not an
exhaustive sweep of the host, and it is not claimed to be. The drill's work happened inside VM 336,
which no longer exists, and the accidents that touched the HOST were iptables rules and a `qm set`
disk detach - both verified reverted above.
@@ -0,0 +1,33 @@
# TEARDOWN LAYER 1 - the machine. 2026-09-17T00:35:25Z
## A GUARD BEFORE AN IRREVERSIBLE ACT
The destroy was gated on the VM's NAME, not just its number:
NAME=$(qm config 336 | awk '/^name:/{print $2}')
[ "$NAME" != "tester1-chaos-night" ] && REFUSE
A VMID is a number that can be mistyped, and 9201 and 9202 - a standing demo box and the scratch
guest - live on the same host. The guard passed: name was `tester1-chaos-night`.
## WHAT WAS DONE
qm stop 336 -> status: stopped
qm destroy 336 --purge 1 --destroy-unreferenced-disks 1
-> "purging VM 336 from related configurations.."
## WHAT IS TRUE AFTERWARDS (each one checked, not assumed)
qm list -> NO VMs at all on demo-hp
pct list -> 9201 demo-hp (running), 9202 demo-hp-scratch (running) <- both SURVIVE
/mnt/hdd_1/images/336 -> "No such file or directory"
So all three disks went with it - the 32 G system disk, the 100 G data disk pulled in round 11, and
the 64 G disk I added in Phase 0 to extend the thin pool I had filled. The fixture deviation
recorded in teardown-baseline.txt is therefore gone from the host, but preserved in the record.
## THE FENCES, RE-STATED AGAINST WHAT ACTUALLY HAPPENED
DooPlex never a target; only builds and this repo ran here
Peti's box untouched
ep0 read only, all night; the final listing is taken separately
drill-r50 its hub record still exists (seen in /hosts) - not touched
demo boxes 9201 and 9202 still running, their standing apps and bentopdf untouched
local-lvm never used for the drill box: its disks were on nvme-scratch (/mnt/hdd_1)
prune never run anywhere
tester-1 the CUSTOMER record is never RESET; only tonight's HOST record is deleted,
and that is layer 3, done through the acknowledged flow