diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/mail-delivery-proof.txt b/documentation/audits/evidence-chaos-night-2026-09-17/mail-delivery-proof.txt new file mode 100644 index 00000000..eb6b8da5 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/mail-delivery-proof.txt @@ -0,0 +1,54 @@ +# THE ALARMS WERE DELIVERED, NOT JUST RAISED - mailbox read 2026-09-17T00:38Z +# The truth table proves 17 alarms fired and all 17 were true. That is the FEED. This is the +# MAILBOX: proof they actually reached a person. In a project whose recurring trap is "stored is +# not delivered" (91 events once sat in a database having e-mailed nobody), the two are not the +# same claim and only one of them was evidenced before now. + +All mails below are from monitoring@felhom.eu to admin@felhom.eu, subject "[Felhom] ... tester-1". +Times are the mail's own; the box and hub run UTC, the mail body prints CEST. + +## ROUND 11 - the drive pulled out for twenty minutes. The whole set arrived. + 23:51:31Z storage_disconnected "Meghajto varatlanul levalasztva: Adatlemez" + 23:51:49Z app_start_failed "Telepitett alkalmazas nem fut: Paperless-ngx" + 23:51:50Z app_start_failed "... Nextcloud" + 23:51:50Z app_start_failed "... Jellyfin" + 23:51:50Z app_start_failed "... Immich" + 23:53:20Z health_degraded "Rendszer allapot romlott (volt: ok)" +Four apps named individually, and they are EXACTLY the four whose data lived on the pulled drive. + +## ROUND 6 - the whole-guest backup, local tier + 21:59:55Z whole_guest_backup_failed + "Whole-guest backup FAILED on the local tier - retrying with backoff (next attempt in 15m0s)" +It names the TIER in the mail subject line's body, not just in the feed. + +## ROUND 4 and the OOM + 21:43:07Z health_critical "Rendszer allapot kritikus (volt: ok)" + 21:20:26Z app_oom "Alkalmazas memoriaja elfogyott: immich (immich-postgres)" + +## ROUND 1 - the off-site control round + 21:08:58Z backup_run_failures "1 of 12 apps failed to back up in this nightly run: nextcloud" + 21:09:00Z offbox_repo_orphaned "A tavoli mentesi tarolo elarvult: a benne levo mentesek egy + korabbi, mar nem elerheto kulccsal..." +The orphan the household would need to know about was MAILED at 21:09, within seconds of the run. + +## PHASE 0 - my own seeding damage, reported honestly by the product + 20:32:50Z app_deploy_failed "Alkalmazas telepitese nem sikerult: Gokapi - exit code 1" + 20:34:32Z storage_fill_critical "Host tester-1-022354: storage \"local-lvm\" CRITICALLY full at + 100% (threshold 95%)" + 20:39:58Z storage_disconnected an earlier drive event during setup, before round 1 +That storage_fill_critical is the thin pool I filled by firing twelve deploys at once. Note which +host it names: tester-1-022354, the NESTED box - not demo-hp. The fence held and the mail proves it. + +## WHAT IS NOT HERE, AND WHY THAT MATTERS +NO node_stale mail for tonight's box after round 9's hub outage. The node_stale mails in the +mailbox are from 12:27Z and 18:17Z and name DIFFERENT host ids (tester-1-652049, tester-1-33b6a9) - +earlier boxes, not this one. That is exactly what R-549 predicts: the gap was 29m59s against a +30-minute threshold, so the alarm never fired and no mail was sent. One second the other way and +this section would list a page-out for a box that was healthy and had already repaired itself. + +## BASELINE FOR THE TEARDOWN'S CONNECT MAIL +The newest mail in the mailbox at 00:38Z is the 23:53:20Z health_degraded. There is NO self-bind or +connect mail from tonight. The last one was 18:17:46Z, BEFORE the drill began. +So any connect mail appearing after the host delete is provably NEW. The brief specifically flagged +that a previous claim about "the automatic mail waiting in the mailbox" had been read from +yesterday's delete rather than checked - this baseline is what makes tonight's quote checkable. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt index 831833a5..57045b03 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-host.txt @@ -34,3 +34,33 @@ Both standing guests survived, which is the fence that mattered most on this hos exhaustive sweep of the host, and it is not claimed to be. The drill's work happened inside VM 336, which no longer exists, and the accidents that touched the HOST were iptables rules and a `qm set` disk detach - both verified reverted above. + +## LAYER 3 BEGINS - and the first delete attempt was REFUSED, deliberately and correctly +The brief asks for the host record to be deleted "through the acknowledged flow". The flow was read +from the hub's own source rather than guessed (hosts.go:866 onward), which gives the route and the +gates in order: + POST /hosts/{id}/delete + gate 1 host ONLINE -> refused + gate 2 confirm_host_id mismatch -> 400 (type-to-confirm) + gate 3 escrow present, delete_escrow != 1 -> 409, moves NOTHING + +Attempt 1, made on purpose WITHOUT the escrow acknowledgement, 00:38:37Z: + POST /hosts/tester-1-022354/delete confirm_host_id=tester-1-022354 + -> HTTP 409 + -> "Host is ONLINE - deletion is refused (a live agent would receive 401s permanently)." +So gate ONE fired, not gate three. The hub still believed the box was alive: its last report was +~00:23:43Z and the box was destroyed at 00:35:25Z, inside the 30-minute liveness window. + +The impact probe the UI calls before offering a delete agrees, and explains itself: + {"deletable":false,"escrow_present":true,"guests":1,"log_bundles":0, + "pbs_secret_present":true,"recovery_present":true,"reports":20,"status":"ok","wg_peer_bound":true} + +THE RECORD SURVIVED THE REFUSAL, checked properly: GET /hosts/tester-1-022354 -> HTTP 200. +(My first check counted the id on the /hosts page and read "3" then "2", which looked like something +had been removed. It had not: `grep -c` counts LINES CONTAINING a match, not occurrences, and the +page wraps differently between renders. The record's own URL answering 200 is the real test.) + +THE WAIT IS THE PRODUCT'S, NOT MINE TO SHORTCUT. The hub should mark the host stale around +00:53:43Z. The acknowledged delete is armed for 00:55Z with a guard that refuses to post while the +host still reads ONLINE. The fence says the hub is never blocked, stopped or changed - so the +correct response to a liveness gate is to wait for it, not to reach into the hub and move it.