From 0f65c8121deb05549cbef94b6b889811db9c0399 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 02:37:49 +0200 Subject: [PATCH] chaos night teardown: machine destroyed, host clean, 9202 untouched, ep0 unchanged Machine layer: VM 336 destroyed with all three disks purged, gated on its NAME rather than its number because 9201 and 9202 share the host. qm list now shows no VMs; /mnt/hdd_1/images/336 is gone; both standing guests still run. Host layer, before and after: nvme-scratch 6.78% -> 1.61% (~48.5 GB released), /mnt/hdd_1/images 59G -> 9.4G with only the scratch guest's own 9202 directory left, free space 827G -> 875G. local-lvm UNCHANGED at 44.75% - the fence that said 'local-lvm never' held. Firewall back at baseline with 0 physdev rules, so none of the three network accidents left a rule on a host carrying two standing guests. Both harness units stopped and disabled before the box died; the disk guard's log was 0 bytes - it never fired once. 9202: nothing to remove, shown rather than asserted - three infrastructure containers, 55 catalog TEMPLATES none of which was touched since 21:00, no app.yaml marked deployed, no offbox config. A false label in my own transcript is corrected there: I printed '(nothing listed above = no app stacks)' directly beneath 55 names. ep0: read again immediately before the delete and identical to the baseline. The single-host delete handler shows no ep0 cascade, but a grep returning nothing is the weakest evidence there is, so the store gets a before and an after rather than an inference. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../ep0-baseline.txt | 16 +++++++++ .../teardown-9202-nothing-to-remove.txt | 24 +++++++++++++ .../teardown-host.txt | 36 +++++++++++++++++++ .../teardown-machine.txt | 33 +++++++++++++++++ 4 files changed, 109 insertions(+) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/teardown-9202-nothing-to-remove.txt create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/teardown-machine.txt diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/ep0-baseline.txt b/documentation/audits/evidence-chaos-night-2026-09-17/ep0-baseline.txt index 16bbde7e..81513d8f 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/ep0-baseline.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/ep0-baseline.txt @@ -15,3 +15,19 @@ disk /dev/sdb 98G total, 16G used, 83G available, 16% used At teardown this listing is repeated. The expected result is that every group and snapshot above is still present; the tester-1 group may have MORE snapshots if a backup ran, and must never have fewer. + +## SECOND READING, taken 2026-09-17T00:36:52Z - immediately BEFORE the hub host delete + namespaces demo-felhom, demo-hp, tester-1 + groups ns/demo-felhom/ct/9201 -> 2 snapshot(s) + ns/demo-hp/ct/9201 -> 2 snapshot(s) + ns/tester-1/ct/9201 -> 2 snapshot(s) + disk 98G total, 16G used, 83G available, 16% +IDENTICAL to the 00:17:15Z baseline. Nothing changed on ep0 during the whole teardown of the machine +and the host - which is what "read only" is supposed to mean, and is now shown rather than asserted. + +WHY A SECOND READING BEFORE THE DELETE: this project already knows that a delete's single removal +operation can destroy a customer's backup groups, not merely release a token. The single-host delete +handler shows no cascade to ep0 - but a grep that returns nothing is the weakest evidence there is, +and tonight has produced four false conclusions built on exactly that. So the store is read again a +few minutes before the delete, and will be read again after it. If the delete touches ep0, the +damage will be measurable against a reading minutes old instead of an hour old. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/teardown-9202-nothing-to-remove.txt b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-9202-nothing-to-remove.txt new file mode 100644 index 00000000..7614ff49 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-9202-nothing-to-remove.txt @@ -0,0 +1,24 @@ +# TEARDOWN, scratch guest 9202: there is nothing to remove - measured, 2026-09-17T00:34:39Z + +The brief's teardown asks for "scratch 9202 (restored app removed)". No restore was ever possible +(the off-site repository is orphaned and holds zero readable snapshots - see +phase2-offsite-restore-blocked.txt), so nothing was ever put on 9202. That is shown, not asserted: + + running containers felhom-controller, filebrowser, traefik <- infrastructure only + /opt/docker/stacks 55 entries + touched since 21:00 NONE (find -newermt "2026-09-16 21:00" returns nothing) + marked deployed NONE (no app.yaml contains desired_state: deployed) + offbox directory absent - an off-site target was never configured on this guest + +## A FALSE LABEL IN MY OWN TRANSCRIPT, corrected here +My first listing printed the 55 directory names and then the line + "(nothing listed above = no app stacks)" +directly underneath them. The label contradicted the output immediately above it. The 55 entries are +CATALOG TEMPLATES synced by the controller - the library of apps the box could install - not apps +that are installed. The distinction matters: a reader skimming that transcript would have seen 55 +app names under a heading saying there were none. +`bentopdf` is among those templates and is untouched, as the fence requires. + +## WHAT THIS MEANS FOR THE TEARDOWN LINE +9202 is returned exactly as it was found: three infrastructure containers, its template library, no +apps, no off-site configuration. Nothing was created on it tonight and nothing was removed from it. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt new file mode 100644 index 00000000..831833a5 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt @@ -0,0 +1,36 @@ +# TEARDOWN LAYER 2 - the host (demo-hp). Before at 00:30:34Z, after at 00:36:07Z. + +## pvesm status, side by side + storage BEFORE used AFTER used change + felhom-pbs 0 0 - + local 29,325,896 KiB 72.49% 29,325,928 KiB 72.49% unchanged (not the drill's storage) + local-lvm 25,278,351 KiB 44.75% 25,278,351 KiB 44.75% UNCHANGED - the fence held, the + drill never used local-lvm + nvme-scratch 66,624,884 KiB 6.78% 15,850,280 KiB 1.61% ~48.5 GB RELEASED + +## /mnt/hdd_1, the storage the brief named + images/ 59 G -> 9.4 G and images/336 is GONE ("No such file or directory") + the only directory left there is 9202 - the SCRATCH GUEST's + own disk, which must stay + dump 4.8 G -> 4.8 G untouched + felhom-data 994 M -> 996 M untouched (grew by its own accord, not by the drill) + free space 827 G -> 875 G + +## guests + BEFORE VM 336 tester1-chaos-night running + CT 9201 demo-hp + CT 9202 demo-hp-scratch + AFTER no VMs at all + CT 9201 running + CT 9202 running +Both standing guests survived, which is the fence that mattered most on this host. + +## the loop, the hog, and the rules + household unit stopped AND disabled on the box before the box was destroyed + diskguard unit stopped AND disabled; its log was 0 bytes - it never fired once all night + leftover files none: no *chaos* or *hog* file under /root or /tmp on demo-hp + firewall -P FORWARD ACCEPT, physdev rules: 0, bridge-nf sysctl: 0 + i.e. exactly the baseline, so none of the three network accidents left a rule + behind on a host that also carries two standing guests + +## WHAT IS NOT PROVEN HERE +"No leftover drill files" rests on looking at /root and /tmp at depth 1 on demo-hp. It is not an +exhaustive sweep of the host, and it is not claimed to be. The drill's work happened inside VM 336, +which no longer exists, and the accidents that touched the HOST were iptables rules and a `qm set` +disk detach - both verified reverted above. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/teardown-machine.txt b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-machine.txt new file mode 100644 index 00000000..268fc6f4 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-machine.txt @@ -0,0 +1,33 @@ +# TEARDOWN LAYER 1 - the machine. 2026-09-17T00:35:25Z + +## A GUARD BEFORE AN IRREVERSIBLE ACT +The destroy was gated on the VM's NAME, not just its number: + NAME=$(qm config 336 | awk '/^name:/{print $2}') + [ "$NAME" != "tester1-chaos-night" ] && REFUSE +A VMID is a number that can be mistyped, and 9201 and 9202 - a standing demo box and the scratch +guest - live on the same host. The guard passed: name was `tester1-chaos-night`. + +## WHAT WAS DONE + qm stop 336 -> status: stopped + qm destroy 336 --purge 1 --destroy-unreferenced-disks 1 + -> "purging VM 336 from related configurations.." + +## WHAT IS TRUE AFTERWARDS (each one checked, not assumed) + qm list -> NO VMs at all on demo-hp + pct list -> 9201 demo-hp (running), 9202 demo-hp-scratch (running) <- both SURVIVE + /mnt/hdd_1/images/336 -> "No such file or directory" + +So all three disks went with it - the 32 G system disk, the 100 G data disk pulled in round 11, and +the 64 G disk I added in Phase 0 to extend the thin pool I had filled. The fixture deviation +recorded in teardown-baseline.txt is therefore gone from the host, but preserved in the record. + +## THE FENCES, RE-STATED AGAINST WHAT ACTUALLY HAPPENED + DooPlex never a target; only builds and this repo ran here + Peti's box untouched + ep0 read only, all night; the final listing is taken separately + drill-r50 its hub record still exists (seen in /hosts) - not touched + demo boxes 9201 and 9202 still running, their standing apps and bentopdf untouched + local-lvm never used for the drill box: its disks were on nvme-scratch (/mnt/hdd_1) + prune never run anywhere + tester-1 the CUSTOMER record is never RESET; only tonight's HOST record is deleted, + and that is layer 3, done through the acknowledged flow