From 69c08b183b46886cb393cc9784a88d4bdaf75d33 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 09:27:10 +0200 Subject: [PATCH] chaos night teardown complete: hub host record deleted, connect mail quoted, ep0 unchanged The acknowledged delete went through at 07:25:13Z (confirm_host_id + delete_escrow=1 -> 303). Every line of the after-state written down before the act matched: the host 404s; drill-r50, both demo hosts and the tester-1 customer still 200; the customer lists zero hosts; ep0 identical across three readings - six snapshots, 16G, nothing removed. The automatic connect mail arrived two seconds later and is quoted with its token redacted. It is provably tonight's: the mailbox held no such mail newer than 18:17:46Z when checked at 00:38Z. Why the hub layer finished six hours late is mine: the retry guard refused to post while the host page contained the word ONLINE, and that word sits in a JavaScript string that is always on the page. The hub's structured status said 'down' from about 00:54Z. It gave up at 01:20Z and nothing ran again until 07:24Z. Eleventh instrument fault of the night. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/DRILL-chaos-night-2026-09-17.md | 34 ++++++++- .../teardown-hub.txt | 69 +++++++++++++++++++ 2 files changed, 102 insertions(+), 1 deletion(-) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/teardown-hub.txt diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index eb3bc719..2e447a34 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -628,7 +628,39 @@ returned by itself in 97 s. ## Teardown — three layers, stated -PENDING +**Machine — gone.** VM 336 stopped and destroyed with all three disks purged (00:35:25Z), gated on +its **name** rather than its number because two standing guests share the host. `qm list` shows no +VMs; `/mnt/hdd_1/images/336` no longer exists. The storage was verified to *be* `/mnt/hdd_1` from its +own definition (`nvme-scratch`, `path /mnt/hdd_1`, `is_mountpoint yes`) rather than assumed. The +machine had **three** disks, not the two the brief asked for — the third was mine, added in Phase 0 +after I filled the thin pool — and that is recorded rather than quietly removed. + +**Host — clean, measured before and after.** `nvme-scratch` 6.78 % → **1.61 %** (~48.5 GB returned); +`local-lvm` **unchanged at 44.75 %**, so the fence that said *never local-lvm* held; free space on +`/mnt/hdd_1` 827 G → 875 G. Guests 9201 and 9202 still running. The household loop and the disk guard +were stopped and disabled **after** their logs were copied off (the guard's log was 0 bytes — it +never fired). Firewall back to `-P FORWARD ACCEPT` with **0** physdev rules, so none of the three +network accidents left a rule behind. Scratch 9202: **nothing to remove**, shown rather than said — +three infrastructure containers, no app deployed, no stack touched since 21:00, no off-site config. + +**Hub — host record deleted through the acknowledged flow.** The first attempt at 00:38:37Z was +**correctly refused** (409, „Host is ONLINE") — the box had died inside the hub's liveness window. The +acknowledged delete went through at **07:25:13Z** (`confirm_host_id` + `delete_escrow=1` → 303). Every +line of the after-state written down *before* the act matched: the host answers 404; `drill-r50`, +both demo hosts and the `tester-1` **customer** still answer 200; the customer now lists zero hosts. +**The automatic connect mail arrived two seconds later** (07:25:15Z, „Kösd össze a Felhom dobozodat"), +quoted in full with its token redacted in `teardown-hub.txt` — and it is provably tonight's, because +the mailbox held no such mail newer than 18:17:46Z when checked at 00:38Z. + +**Why the hub layer finished six hours late — my fault, not the product's.** The retry was guarded +by „don't post while the host page contains ONLINE". That word lives in a JavaScript string that is +**always** on the page, so the guard could never pass. It refused six times and gave up at 01:20Z, +while the hub's structured answer would have said `"status":"down"` from about 00:54Z. Nothing ran +again until 07:24Z. + +**ep0 — backups stayed, nothing removed.** Read three times: 00:17:15Z, 00:36:52Z (just before the +delete) and 07:25:35Z (after). Identical every time — three namespaces, two snapshots each, six in +total, 16 G used. No prune, no verify, no write. ## Claims in the prompt that turned out wrong — named first, as asked diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/teardown-hub.txt b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-hub.txt new file mode 100644 index 00000000..cdd9926f --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-hub.txt @@ -0,0 +1,69 @@ +# TEARDOWN LAYER 3 - the hub. Completed 2026-09-17T07:25:13Z. + +## THE ACT + GET /hosts/tester-1-022354/delete-impact (07:24:58Z) + {"deletable":true,"escrow_present":true,"guests":1,"log_bundles":0,"pbs_secret_present":true, + "recovery_present":true,"reports":20,"status":"down","wg_peer_bound":true} + POST /hosts/tester-1-022354/delete confirm_host_id=tester-1-022354 delete_escrow=1 + -> HTTP 303, Location: /hosts +The escrow acknowledgement moves the escrow to RETAINED CUSTODY; it does not destroy it +(the handler's own refusal text, and R-544). + +## THE AFTER-PICTURE, against the expectation written down at 00:47:30Z BEFORE the act + expected measured + tester-1-022354 gone GET /hosts/tester-1-022354 -> 404 + drill-r50 survives GET /hosts/drill-r50-0a4f9a -> 200 + demo boxes survive GET /hosts/demo-hp-bb76ea -> 200 + GET /hosts/demo-felhom-8363b5 -> 200 + the CUSTOMER is never reset GET /customers/tester-1 -> 200 + three host records left demo-felhom-8363b5, demo-hp-bb76ea, drill-r50-0a4f9a + customer page lists zero hosts 0 links to the deleted host; its 4 remaining + mentions are EVENT HISTORY rows (host_stale, + host_down), which correctly stay + ep0 untouched 3 namespaces, 6 snapshots, 16G used - identical to + the readings of 00:17:15Z and 00:36:52Z +Every line matched. Nothing had to be re-interpreted after the fact. + +## THE AUTOMATIC CONNECT MAIL, quoted (arrived 07:25:15Z - TWO SECONDS after the delete) + From: monitoring@felhom.eu To: tester1@felhom.eu + Subject: [Felhom] Kösd össze a Felhom dobozodat + + Kedves Ügyfél! + + Elkészült a Felhom dobozod, és készen áll az összekötésre. Az alábbi hivatkozáson + tudod te magad összekötni a fiókoddal — nincs szükség bejelentkezésre: + + https://hub.felhom.eu/bind/ + + A hivatkozás megnyitása után két adatot kell megadnod: + + 1. A párosító kódot, amely a doboz képernyőjén (a monitoron) látható. + 2. A tulajdonosi jelmondatodat (az 5 szóból álló kifejezést), amelyet a + Felhom üzemeltetőjétől kaptál — személyesen vagy telefonon, e-mailben + soha. Ez igazolja, hogy a fiók a tiéd. + + A hivatkozás 7 napig érvényes. Biztonsági okból 5 sikertelen próbálkozás után + zárolódik — ilyenkor vedd fel a kapcsolatot az ügyfélszolgálattal. + + Ha nem te kérted ezt, hagyd figyelmen kívül ezt az e-mailt. + + Üdvözlettel, + Felhom.eu + +WHY THIS QUOTE IS PROVABLY TONIGHT'S: the mailbox baseline at 00:38Z held NO self-bind mail newer +than 18:17:46Z. This one is 07:25:15Z. The brief's warning - that an earlier claim about this mail had +been read from yesterday's delete - cannot apply to it. +Also true and delivered while the box was dead: host_down for tester-1-022354 at 01:27:31Z +("no report for 1h"), and host_stale before it (shown "sent" on the customer page). Both TRUE. + +## THE ELEVENTH INSTRUMENT FAULT - and the one that cost the most time +The first delete (00:38:37Z) was correctly refused: the host was still inside its liveness window. +I armed a retry for 00:55Z behind a guard: "do not post while the host page contains ONLINE". +That guard COULD NEVER PASS. The word is inside a JavaScript string that is always on the page: + ... (d.deletable ? '' : ' Host is ONLINE — deletion will be refused.') +So the guard refused six times, 00:55Z to 01:20Z, while the hub's own structured answer would have +said "status":"down" from about 00:54Z. Then the retry loop ran out and exited 0. +The session did not resume until 07:24Z (after a login), so the hub layer finished ~6 hours after the +rest of the teardown. The fix was the obvious one and should have been the first one: ask the +endpoint that returns a STATUS FIELD, not grep a page for a word. Same class as every other fault +tonight - a probe that cannot fail in the way I needed it to is not a measurement.