chaos night: the alarms were DELIVERED, and the first delete was correctly refused
gates / gates (push) Successful in 22s

The mailbox closes a gap the truth table could not: 17 alarms fired and all 17
were true, but that was the FEED. The mailbox shows they reached a person.
Round 11's full set arrived - storage_disconnected naming the drive, four
app_start_failed naming exactly the four apps whose data was on it, then
health_degraded. Round 6's whole_guest_backup_failed names the TIER. Round 1's
offbox_repo_orphaned was mailed within seconds of the run.

What is absent matters too: NO node_stale mail for tonight's box after round
9's outage, exactly as R-549 predicts - the gap was 29m59s against a 30-minute
threshold. One second the other way and this would be a page-out for a healthy,
self-repaired box.

Teardown layer 3 began with a DELIBERATE un-acknowledged delete, to see the
gate refuse: HTTP 409, 'Host is ONLINE - deletion is refused'. Gate one fired,
not the escrow gate - the box died inside the hub's 30-minute liveness window.
The record survived, verified by its own URL returning 200 rather than by
counting substrings on a list page (grep -c counts lines, not occurrences - my
'3 then 2' was my error, not a deletion).

The wait is the product's and not mine to shortcut: the acknowledged delete is
armed for 00:55Z behind a guard that will not post while the host reads ONLINE.
The fence says the hub is never changed, so a liveness gate is waited for.

Also recorded: the mailbox baseline proving no connect mail exists from tonight,
so the one quoted after the delete is provably new - the exact check the brief
said had been skipped before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 02:40:38 +02:00
parent 0f65c8121d
commit 6c450bca50
2 changed files with 84 additions and 0 deletions
@@ -0,0 +1,54 @@
# THE ALARMS WERE DELIVERED, NOT JUST RAISED - mailbox read 2026-09-17T00:38Z
# The truth table proves 17 alarms fired and all 17 were true. That is the FEED. This is the
# MAILBOX: proof they actually reached a person. In a project whose recurring trap is "stored is
# not delivered" (91 events once sat in a database having e-mailed nobody), the two are not the
# same claim and only one of them was evidenced before now.
All mails below are from monitoring@felhom.eu to admin@felhom.eu, subject "[Felhom] ... tester-1".
Times are the mail's own; the box and hub run UTC, the mail body prints CEST.
## ROUND 11 - the drive pulled out for twenty minutes. The whole set arrived.
23:51:31Z storage_disconnected "Meghajto varatlanul levalasztva: Adatlemez"
23:51:49Z app_start_failed "Telepitett alkalmazas nem fut: Paperless-ngx"
23:51:50Z app_start_failed "... Nextcloud"
23:51:50Z app_start_failed "... Jellyfin"
23:51:50Z app_start_failed "... Immich"
23:53:20Z health_degraded "Rendszer allapot romlott (volt: ok)"
Four apps named individually, and they are EXACTLY the four whose data lived on the pulled drive.
## ROUND 6 - the whole-guest backup, local tier
21:59:55Z whole_guest_backup_failed
"Whole-guest backup FAILED on the local tier - retrying with backoff (next attempt in 15m0s)"
It names the TIER in the mail subject line's body, not just in the feed.
## ROUND 4 and the OOM
21:43:07Z health_critical "Rendszer allapot kritikus (volt: ok)"
21:20:26Z app_oom "Alkalmazas memoriaja elfogyott: immich (immich-postgres)"
## ROUND 1 - the off-site control round
21:08:58Z backup_run_failures "1 of 12 apps failed to back up in this nightly run: nextcloud"
21:09:00Z offbox_repo_orphaned "A tavoli mentesi tarolo elarvult: a benne levo mentesek egy
korabbi, mar nem elerheto kulccsal..."
The orphan the household would need to know about was MAILED at 21:09, within seconds of the run.
## PHASE 0 - my own seeding damage, reported honestly by the product
20:32:50Z app_deploy_failed "Alkalmazas telepitese nem sikerult: Gokapi - exit code 1"
20:34:32Z storage_fill_critical "Host tester-1-022354: storage \"local-lvm\" CRITICALLY full at
100% (threshold 95%)"
20:39:58Z storage_disconnected an earlier drive event during setup, before round 1
That storage_fill_critical is the thin pool I filled by firing twelve deploys at once. Note which
host it names: tester-1-022354, the NESTED box - not demo-hp. The fence held and the mail proves it.
## WHAT IS NOT HERE, AND WHY THAT MATTERS
NO node_stale mail for tonight's box after round 9's hub outage. The node_stale mails in the
mailbox are from 12:27Z and 18:17Z and name DIFFERENT host ids (tester-1-652049, tester-1-33b6a9) -
earlier boxes, not this one. That is exactly what R-549 predicts: the gap was 29m59s against a
30-minute threshold, so the alarm never fired and no mail was sent. One second the other way and
this section would list a page-out for a box that was healthy and had already repaired itself.
## BASELINE FOR THE TEARDOWN'S CONNECT MAIL
The newest mail in the mailbox at 00:38Z is the 23:53:20Z health_degraded. There is NO self-bind or
connect mail from tonight. The last one was 18:17:46Z, BEFORE the drill began.
So any connect mail appearing after the host delete is provably NEW. The brief specifically flagged
that a previous claim about "the automatic mail waiting in the mailbox" had been read from
yesterday's delete rather than checked - this baseline is what makes tonight's quote checkable.
@@ -34,3 +34,33 @@ Both standing guests survived, which is the fence that mattered most on this hos
exhaustive sweep of the host, and it is not claimed to be. The drill's work happened inside VM 336,
which no longer exists, and the accidents that touched the HOST were iptables rules and a `qm set`
disk detach - both verified reverted above.
## LAYER 3 BEGINS - and the first delete attempt was REFUSED, deliberately and correctly
The brief asks for the host record to be deleted "through the acknowledged flow". The flow was read
from the hub's own source rather than guessed (hosts.go:866 onward), which gives the route and the
gates in order:
POST /hosts/{id}/delete
gate 1 host ONLINE -> refused
gate 2 confirm_host_id mismatch -> 400 (type-to-confirm)
gate 3 escrow present, delete_escrow != 1 -> 409, moves NOTHING
Attempt 1, made on purpose WITHOUT the escrow acknowledgement, 00:38:37Z:
POST /hosts/tester-1-022354/delete confirm_host_id=tester-1-022354
-> HTTP 409
-> "Host is ONLINE - deletion is refused (a live agent would receive 401s permanently)."
So gate ONE fired, not gate three. The hub still believed the box was alive: its last report was
~00:23:43Z and the box was destroyed at 00:35:25Z, inside the 30-minute liveness window.
The impact probe the UI calls before offering a delete agrees, and explains itself:
{"deletable":false,"escrow_present":true,"guests":1,"log_bundles":0,
"pbs_secret_present":true,"recovery_present":true,"reports":20,"status":"ok","wg_peer_bound":true}
THE RECORD SURVIVED THE REFUSAL, checked properly: GET /hosts/tester-1-022354 -> HTTP 200.
(My first check counted the id on the /hosts page and read "3" then "2", which looked like something
had been removed. It had not: `grep -c` counts LINES CONTAINING a match, not occurrences, and the
page wraps differently between renders. The record's own URL answering 200 is the real test.)
THE WAIT IS THE PRODUCT'S, NOT MINE TO SHORTCUT. The hub should mark the host stale around
00:53:43Z. The acknowledged delete is armed for 00:55Z with a guard that refuses to post while the
host still reads ONLINE. The fence says the hub is never blocked, stopped or changed - so the
correct response to a liveness gate is to wait for it, not to reach into the hub and move it.