chaos night: the alarms were DELIVERED, and the first delete was correctly refused
gates / gates (push) Successful in 22s
gates / gates (push) Successful in 22s
The mailbox closes a gap the truth table could not: 17 alarms fired and all 17 were true, but that was the FEED. The mailbox shows they reached a person. Round 11's full set arrived - storage_disconnected naming the drive, four app_start_failed naming exactly the four apps whose data was on it, then health_degraded. Round 6's whole_guest_backup_failed names the TIER. Round 1's offbox_repo_orphaned was mailed within seconds of the run. What is absent matters too: NO node_stale mail for tonight's box after round 9's outage, exactly as R-549 predicts - the gap was 29m59s against a 30-minute threshold. One second the other way and this would be a page-out for a healthy, self-repaired box. Teardown layer 3 began with a DELIBERATE un-acknowledged delete, to see the gate refuse: HTTP 409, 'Host is ONLINE - deletion is refused'. Gate one fired, not the escrow gate - the box died inside the hub's 30-minute liveness window. The record survived, verified by its own URL returning 200 rather than by counting substrings on a list page (grep -c counts lines, not occurrences - my '3 then 2' was my error, not a deletion). The wait is the product's and not mine to shortcut: the acknowledged delete is armed for 00:55Z behind a guard that will not post while the host reads ONLINE. The fence says the hub is never changed, so a liveness gate is waited for. Also recorded: the mailbox baseline proving no connect mail exists from tonight, so the one quoted after the delete is provably new - the exact check the brief said had been skipped before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,54 @@
|
|||||||
|
# THE ALARMS WERE DELIVERED, NOT JUST RAISED - mailbox read 2026-09-17T00:38Z
|
||||||
|
# The truth table proves 17 alarms fired and all 17 were true. That is the FEED. This is the
|
||||||
|
# MAILBOX: proof they actually reached a person. In a project whose recurring trap is "stored is
|
||||||
|
# not delivered" (91 events once sat in a database having e-mailed nobody), the two are not the
|
||||||
|
# same claim and only one of them was evidenced before now.
|
||||||
|
|
||||||
|
All mails below are from monitoring@felhom.eu to admin@felhom.eu, subject "[Felhom] ... tester-1".
|
||||||
|
Times are the mail's own; the box and hub run UTC, the mail body prints CEST.
|
||||||
|
|
||||||
|
## ROUND 11 - the drive pulled out for twenty minutes. The whole set arrived.
|
||||||
|
23:51:31Z storage_disconnected "Meghajto varatlanul levalasztva: Adatlemez"
|
||||||
|
23:51:49Z app_start_failed "Telepitett alkalmazas nem fut: Paperless-ngx"
|
||||||
|
23:51:50Z app_start_failed "... Nextcloud"
|
||||||
|
23:51:50Z app_start_failed "... Jellyfin"
|
||||||
|
23:51:50Z app_start_failed "... Immich"
|
||||||
|
23:53:20Z health_degraded "Rendszer allapot romlott (volt: ok)"
|
||||||
|
Four apps named individually, and they are EXACTLY the four whose data lived on the pulled drive.
|
||||||
|
|
||||||
|
## ROUND 6 - the whole-guest backup, local tier
|
||||||
|
21:59:55Z whole_guest_backup_failed
|
||||||
|
"Whole-guest backup FAILED on the local tier - retrying with backoff (next attempt in 15m0s)"
|
||||||
|
It names the TIER in the mail subject line's body, not just in the feed.
|
||||||
|
|
||||||
|
## ROUND 4 and the OOM
|
||||||
|
21:43:07Z health_critical "Rendszer allapot kritikus (volt: ok)"
|
||||||
|
21:20:26Z app_oom "Alkalmazas memoriaja elfogyott: immich (immich-postgres)"
|
||||||
|
|
||||||
|
## ROUND 1 - the off-site control round
|
||||||
|
21:08:58Z backup_run_failures "1 of 12 apps failed to back up in this nightly run: nextcloud"
|
||||||
|
21:09:00Z offbox_repo_orphaned "A tavoli mentesi tarolo elarvult: a benne levo mentesek egy
|
||||||
|
korabbi, mar nem elerheto kulccsal..."
|
||||||
|
The orphan the household would need to know about was MAILED at 21:09, within seconds of the run.
|
||||||
|
|
||||||
|
## PHASE 0 - my own seeding damage, reported honestly by the product
|
||||||
|
20:32:50Z app_deploy_failed "Alkalmazas telepitese nem sikerult: Gokapi - exit code 1"
|
||||||
|
20:34:32Z storage_fill_critical "Host tester-1-022354: storage \"local-lvm\" CRITICALLY full at
|
||||||
|
100% (threshold 95%)"
|
||||||
|
20:39:58Z storage_disconnected an earlier drive event during setup, before round 1
|
||||||
|
That storage_fill_critical is the thin pool I filled by firing twelve deploys at once. Note which
|
||||||
|
host it names: tester-1-022354, the NESTED box - not demo-hp. The fence held and the mail proves it.
|
||||||
|
|
||||||
|
## WHAT IS NOT HERE, AND WHY THAT MATTERS
|
||||||
|
NO node_stale mail for tonight's box after round 9's hub outage. The node_stale mails in the
|
||||||
|
mailbox are from 12:27Z and 18:17Z and name DIFFERENT host ids (tester-1-652049, tester-1-33b6a9) -
|
||||||
|
earlier boxes, not this one. That is exactly what R-549 predicts: the gap was 29m59s against a
|
||||||
|
30-minute threshold, so the alarm never fired and no mail was sent. One second the other way and
|
||||||
|
this section would list a page-out for a box that was healthy and had already repaired itself.
|
||||||
|
|
||||||
|
## BASELINE FOR THE TEARDOWN'S CONNECT MAIL
|
||||||
|
The newest mail in the mailbox at 00:38Z is the 23:53:20Z health_degraded. There is NO self-bind or
|
||||||
|
connect mail from tonight. The last one was 18:17:46Z, BEFORE the drill began.
|
||||||
|
So any connect mail appearing after the host delete is provably NEW. The brief specifically flagged
|
||||||
|
that a previous claim about "the automatic mail waiting in the mailbox" had been read from
|
||||||
|
yesterday's delete rather than checked - this baseline is what makes tonight's quote checkable.
|
||||||
@@ -34,3 +34,33 @@ Both standing guests survived, which is the fence that mattered most on this hos
|
|||||||
exhaustive sweep of the host, and it is not claimed to be. The drill's work happened inside VM 336,
|
exhaustive sweep of the host, and it is not claimed to be. The drill's work happened inside VM 336,
|
||||||
which no longer exists, and the accidents that touched the HOST were iptables rules and a `qm set`
|
which no longer exists, and the accidents that touched the HOST were iptables rules and a `qm set`
|
||||||
disk detach - both verified reverted above.
|
disk detach - both verified reverted above.
|
||||||
|
|
||||||
|
## LAYER 3 BEGINS - and the first delete attempt was REFUSED, deliberately and correctly
|
||||||
|
The brief asks for the host record to be deleted "through the acknowledged flow". The flow was read
|
||||||
|
from the hub's own source rather than guessed (hosts.go:866 onward), which gives the route and the
|
||||||
|
gates in order:
|
||||||
|
POST /hosts/{id}/delete
|
||||||
|
gate 1 host ONLINE -> refused
|
||||||
|
gate 2 confirm_host_id mismatch -> 400 (type-to-confirm)
|
||||||
|
gate 3 escrow present, delete_escrow != 1 -> 409, moves NOTHING
|
||||||
|
|
||||||
|
Attempt 1, made on purpose WITHOUT the escrow acknowledgement, 00:38:37Z:
|
||||||
|
POST /hosts/tester-1-022354/delete confirm_host_id=tester-1-022354
|
||||||
|
-> HTTP 409
|
||||||
|
-> "Host is ONLINE - deletion is refused (a live agent would receive 401s permanently)."
|
||||||
|
So gate ONE fired, not gate three. The hub still believed the box was alive: its last report was
|
||||||
|
~00:23:43Z and the box was destroyed at 00:35:25Z, inside the 30-minute liveness window.
|
||||||
|
|
||||||
|
The impact probe the UI calls before offering a delete agrees, and explains itself:
|
||||||
|
{"deletable":false,"escrow_present":true,"guests":1,"log_bundles":0,
|
||||||
|
"pbs_secret_present":true,"recovery_present":true,"reports":20,"status":"ok","wg_peer_bound":true}
|
||||||
|
|
||||||
|
THE RECORD SURVIVED THE REFUSAL, checked properly: GET /hosts/tester-1-022354 -> HTTP 200.
|
||||||
|
(My first check counted the id on the /hosts page and read "3" then "2", which looked like something
|
||||||
|
had been removed. It had not: `grep -c` counts LINES CONTAINING a match, not occurrences, and the
|
||||||
|
page wraps differently between renders. The record's own URL answering 200 is the real test.)
|
||||||
|
|
||||||
|
THE WAIT IS THE PRODUCT'S, NOT MINE TO SHORTCUT. The hub should mark the host stale around
|
||||||
|
00:53:43Z. The acknowledged delete is armed for 00:55Z with a guard that refuses to post while the
|
||||||
|
host still reads ONLINE. The fence says the hub is never blocked, stopped or changed - so the
|
||||||
|
correct response to a liveness gate is to wait for it, not to reach into the hub and move it.
|
||||||
|
|||||||
Reference in New Issue
Block a user