Files
felhom.eu/documentation/audits/DRILL-chaos-night-2026-09-17.md
T
admin 69c08b183b
gates / gates (push) Successful in 22s
chaos night teardown complete: hub host record deleted, connect mail quoted, ep0 unchanged
The acknowledged delete went through at 07:25:13Z (confirm_host_id +
delete_escrow=1 -> 303). Every line of the after-state written down before the
act matched: the host 404s; drill-r50, both demo hosts and the tester-1
customer still 200; the customer lists zero hosts; ep0 identical across three
readings - six snapshots, 16G, nothing removed.

The automatic connect mail arrived two seconds later and is quoted with its
token redacted. It is provably tonight's: the mailbox held no such mail newer
than 18:17:46Z when checked at 00:38Z.

Why the hub layer finished six hours late is mine: the retry guard refused to
post while the host page contained the word ONLINE, and that word sits in a
JavaScript string that is always on the page. The hub's structured status said
'down' from about 00:54Z. It gave up at 01:20Z and nothing ran again until
07:24Z. Eleventh instrument fault of the night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 09:27:10 +02:00

54 KiB
Raw Blame History

DRILL — CHAOS NIGHT: random actions on random apps while random things go wrong (2026-09-16/17)

Interventions: 1. One, at 21:59:45Z in round 6: I killed the local leg of a whole-guest backup that could never have fit (a ~29 GB source into a 14 GB root filesystem, falling at ~16 MB/s). The off-site leg then ran by itself from the same snapshot and succeeded, so the data still left the house. Both pre-declared presses went unused: the automatic self-bind mail was already waiting, and the acknowledged-delete path re-issued the PBS credentials on its own. Phase 0's seeding repairs are listed separately in evidence-chaos-night-2026-09-17/interventions.txt — that damage was mine, not the product's, and every repair went through the product's own endpoints.

Ready for a volunteer: still yes. Across twelve rounds — a power cut mid-restore, a hard reset four seconds into another, a full disk, a killed tunnel, a restarted Docker, three severed networks and the data drive pulled out of a running machine for twenty minutes — nothing cost a byte of customer data, and the box healed itself every single time with no human involved. Seventeen alarms fired, all seventeen were true, none were missing, and the mailbox proves every one was delivered rather than merely stored. The honest qualifications: two P2 legibility gaps are filed (an interrupted restore leaves no record; one failed report spends the whole 30-minute staleness budget, measured at 29 m 59 s), and one thing this night could not test — per-app off-site restore, because this box is a rebuild whose restic repository is orphaned by design, which the product surfaced honestly within seconds.

The accident-plus-action pair that hurt most: restore + hard reset (round 10). Not because the box suffered — it was back with 26 of 26 containers in 150 s, boot reconciliation naming the app it recovered, every front door serving. It hurt most because it is the only pair of the night where the household is left not knowing what happened: they pressed restore, were told it had started, the machine went dark four seconds later, and afterwards nothing anywhere tells them whether it finished. The status surface exists and answers with the zero value; the record is in-memory only and does not survive the machine stopping.

Baselines, verified live against Gitea at 21:49 CEST 2026-09-16 (not copied from the brief): felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) · felhom-agent e98b857684f4 v0.131.0 · felhom.eu d124c77e176d hub v0.116.0, ISO 1.28.0 published · app-catalog 94bc5febaca2. All four trees clean and in sync. Highest register row R-545, 212 open. Golden waiver valid to 2026-09-27. Customer tester-1 (enkicsifelhom.hu, tester1@felhom.eu, no host). Venue: demo-hp (Tier 0), a fresh nested VM, disk on the NVMe at its root. Evidence: evidence-chaos-night-2026-09-17/.

The schedule — drawn ONCE, before round 1, and written here first

The point of this section's position in the document is that the night could not be chosen after the fact. chaos_schedule.py is committed beside the evidence; re-running it reproduces this table.

  • seed: 20260917 (the date)
  • script sha256: 4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1
  • generator: evidence-chaos-night-2026-09-17/chaos_schedule.py, stdlib random seeded with the seed
# time X — the action Y — the app Z — the accident
1 23:30 offsite-run adventurelog nothing
2 23:55 restore gokapi power cut
3 00:20 use bookstack disk 95% full
4 00:45 offsite-run mealie tunnel down 10min
5 01:10 use privatebin docker restarted
6 01:35 backup-system adventurelog nothing
7 02:00 update nextcloud internet gone 10min
8 02:25 backup-app nextcloud internet gone 10min
9 02:50 use uptime-kuma internet gone 10min
10 03:15 restore uptime-kuma hard reset
11 03:40 use paperless-ngx drive pulled 20min
12 04:05 use paperless-ngx nothing

Re-draw log — a silent re-draw is a schedule chosen by the person running it, so every one is here:

  • r02 X=reinstall re-drawn (nothing has been removed yet)
  • r06 Z=disk 95% full re-drawn (constraint 4: at most once)
  • r07 Z=nothing re-drawn (constraint 6: never two in a row after r2)
  • r08 X=reinstall re-drawn (nothing has been removed yet)

What the draw happened to give, said plainly before the night judges it: no remove round was ever drawn, so reinstall had nothing to reinstall and was re-drawn twice (rounds 2 and 8). Three internet gone rounds land consecutively (7, 8, 9) — that is the seed's doing, and it makes rounds 7–9 a de-facto endurance test of the same accident against three different actions rather than three independent samples. controller killed, drive pulled 90s, memory pressure, agent restarted and hub unreachable were never drawn at all; this night does not test them, and the morning verdict must not claim it did.

Phase 0 — the golden, the box, the household

0.1 Golden 0.245.0, baked and published. Launched 19:52:47Z as a transient unit in the drill VM, finished 19:58:15Z. Markers: overlay2=1, including mount point=2, upload OK (HTTP 201)=1, FATAL=0, publish-skipped=0. GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626. The teardown was gated on the REGISTRY answering 200, not on an exit code — and that mattered: the wrapper exited 144 while every measured outcome was good. Vouched in the hub and read back from the page (golden currently vouched: 0.245.0). The three bake failures of 2026-09-16 (scp -P, chmod 0700, GITEA_USER=admin) were each guarded and none recurred.

Decision, recorded because silence reads as agreement: the global controller floor was left at 0.244.0. The new box installs golden 0.245.0, which already carries controller 0.245.0, so no floor was needed to deliver anything tonight; raising it would have pushed an update onto demo-felhom, a box not in this drill.

0.2 The box. VM 336 on demo-hp: 8 GiB, 4 cores, 32 G system + 100 G data disk on the NVMe at its root, booted from the published ISO 1.28.0. Boot order set in its own qm set (combining it silently yields order=net0;ide2). Install completion was judged from the disk — blocks used grew 3233 → 6942 MiB then held across three checks — because „Automatically reboot" is ticked and a finished install looks exactly like a stuck one on screen. The summary page was read before pressing Install, and the line that made it safe was „Disk(s): /dev/sda" — the 32 G system disk alone.

The walk, as a volunteer, cost ZERO operator presses. The box registered itself as an unclaimed appliance and polled, visibly, until bound. The bind link came from the waiting mail (minted 18:17:46Z by yesterday's acknowledged host delete), the pairing code off the box's own console (4SY-4TX), the „Tulajdonosi jelmondat" from the hub's customer record: POST /bind/ -> 200, „Sikeres összekötés." Then day-0 ran on its own and the hub recorded, without anyone pressing anything: appliance_bound (customer_selfbind) · appliance_credential_delivered · claim_reissued_reenroll · offsite_reissued · pbsdr_auto_reissue — „Previous key destroyed (acknowledged deletion) — credentials re-issued automatically."

That last event is a first. The brief named „the WG hook provisions by itself after an acknowledged delete" as a claim never measured live. It ran tonight, unprompted. Both pre-declared presses (O1 self-bind, O2 re-issue) were therefore unnecessary.

The dashboard was claimed with the mailed code and proven by logging in with the new password — a claim page that re-renders looks identical to success from the status code alone. The 100 GB data drive was initialised through the wizard's own endpoint (POST /api/storage/init, polled to phase: done), and df shows it mounted at /mnt/felhom-drives/hdd_1 with 93 G free.

The box landed on tonight's golden with no hand upgrade: controller 0.245.0 (healthy), agent 0.131.0, host tester-1-022354 ONLINE. And the R-543 escrow reminder bar shipped hours earlier was live on it, on a box nobody had touched.

0.3 The household — and the first real trouble. Twelve deploys were fired; ten were accepted, two refused for memory with both numbers quoted. Then nine of the ten failed: the guest's disks are thin-provisioned over an ~11.8 GB pool carved from a 32 GB system disk, ten simultaneous image pulls filled it, and the hub recorded storage_fill_critical (100 %) plus nine app_deploy_failed warnings, one per app, each naming the failing pull. Only PrivateBin installed.

The product behaved; the harness did not. R-536's failure event — shipped that same morning so an interrupted install is not silence — fired for all nine within two minutes. The memory guard refused rather than over-committing. The per-stack record stayed honest (deployed: false). The two faults were mine: firing twelve deploys in two seconds is not household behaviour, and a 32 GB system disk was copied from an earlier drill without checking what that drill had installed.

0.4 The escrow ceremony could NOT be completed — and this one is about tonight's own release. See „Finding: the recovery-code step cannot be done when the guide says to do it" below.

Finding: the recovery-code step cannot be done when the guide says to do it (R-546)

The box was at exactly the point of the guide this release added hours earlier — installed, bound with no press, claimed, drive initialised, no apps yet — and the escrow reminder bar was on every page telling the household to create their recovery code. It could not be done.

POST /api/escrow/start   -> 200   {"job_id":"escrow-1789590499667361664","phase":"running"}
GET  /api/escrow/status  -> claimable:false —
     detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"
POST /api/escrow/claim   -> **409** „A folyamat jelenlegi állapotában a kód nem kérhető le."

Both sides agreed on the cause. The hub's own Backup & DR panel read „host enrolled done · WG tunnel peer registered done · descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1) waiting · ceremony possible once the descriptor is applied on the box". The box had no PBS storage (pvesm status: local, local-lvm only) and no escrow section in agent.json at all.

It self-heals, and that was measured rather than assumed. The box was left alone and polled:

20:30:12Z … 20:34:15Z   pbs_storage=none         escrow.pbs_storage_id=none
**20:35:16Z              pbs_storage=felhom-pbs   escrow.pbs_storage_id=felhom-pbs**

~17 minutes after the bind. The retried ceremony passed every preflight item, the claim returned 200 (83-character code, 129.2 bits of entropy, revealed once), and escrow_state became escrowed. The bar then vanished from all four pages checked — the R-543 fix working through its whole lifecycle on a box nobody had set up for the test.

So the defect is timing and wording, not mechanism. For ~17 minutes a volunteer following tonight's guide meets a stderr fragment about a -storage flag, while every page urges them on. Filed R-546 (P2). No product code was changed — this is a validation run.

Phase 1 — the rounds

Rounds run at ~25-minute spacing. The schedule above is fixed; only the wall-clock start moved, because Phase 0 ran long (the storage wizard submits by JavaScript and the endpoint was worked out rather than guessed). Round 1 began 23:07 CEST.

Round 1 — offsite-run / adventurelog / accident: nothing (control round)

the five things
what the customer saw „A távoli mentés elindult — az állapot itt frissül." and, at the end, „A távoli mentési tároló elárvult: a benne lévő mentések egy korábbi, már nem elérhető kulccsal készültek (újratelepítés)."
what the box did by itself walked all twelve apps — stop, dump each volume with real byte counts, restart — captured eleven, could not capture the one that was crash-looping, and finished
time to steady 1m45s (last_duration), last_run 21:09:00Z, progress.active false, last_error empty. A control round: the box never left steady
alarm fired / true? three, all true — app_start_failed named Nextcloud · backup_run_failures „1 of 12 apps failed to back up in this nightly run: nextcloud" · offbox_repo_orphaned, matching the status endpoint's own "orphaned": true
should have fired, did not none

Household loop in the window: 3 lines marked FAILED, all three mine — the loop counted the dashboard's 301 redirect as a failure while accepting the same 301 for app reads. Fixed at 21:11:25Z and marked in the log; only lines after that marker are scored.

What round 1 actually establishes. The off-site tier is armed (escrowed) and the run works end-to-end, but on THIS box — a rebuild for an existing customer — the remote repository was written under a key the box no longer holds, so no snapshot was written. That is the documented rebuild behaviour, surfaced honestly with the route out named in the message rather than reported as success.

Two things that looked like defects in this round and are not, both established with controls rather than inference — five front doors answering 404 (traefik has no route to an unhealthy container; identical byte-for-byte to a no-such-host control) and a crash-looping Nextcloud (image layers corrupted while my thin pool stood at 100 %; „invalid ELF header", repaired by a re-pull). Detail in evidence-chaos-night-2026-09-17/round-1-notes.txt.

Phase 0 postscript — repairing my own damage, and three conclusions I had to retract

Nine of the first ten deploys failed because I fired twelve at once onto a thin pool carved from a 32 GB disk, and the pool hit 100 %. Repairing that took the rest of Phase 0 and produced two distinct faults of mine, with different cures, which only separating them made fixable:

fault symptom cure
image layers written while the pool was full php: … libxml2.so.2: **invalid ELF header**, exit 127 crash loop drop the image, let compose pull it again
my re-seed generated fresh database passwords over volumes already initialised with the first set Postgres auth_failed, MariaDB „Access denied for user … (using password: YES)" remove the app with its data, deploy once with one consistent secret set

A fresh image did not fix gokapi and a fresh database did — that is the evidence the two faults are different things rather than one.

Three conclusions I wrote and then had to retract, each corrected where it stood:

  1. „the five 404s were my mistimed sweep" — wrong for four of them. A negative control (a no-such-host request) returned the identical 404 of 19 bytes, and a positive control returned 200/1200 bytes: traefik simply has no route to an unhealthy container.
  2. „nextcloud is repaired" — wrong. The re-pull fixed the crash, and the app still could not reach its database. The container reported healthy throughout, because the image's own healthcheck asks whether Apache answers, not whether the application works.
  3. „all the broken apps are corrupt layers" — wrong. bookstack logged a clean startup, gokapi logged nothing at all, immich showed a Postgres auth failure.

I also nearly filed a defect against the drive gate, which was working and logging at DEBUG while I read a settings snapshot inside its 30-second tick. An absent log line is not evidence.

None of this is a product defect and none of it is filed as one. What the product did throughout was correct and legible: it refused to route to unhealthy containers, app_start_failed named the app that was down, backup_run_failures said „1 of 12 … nextcloud", the memory guard refused the eleventh and twelfth installs with both numbers quoted, and storage_fill_critical fired at 100 %.

Round 2 — restore gokapi / accident: power cut, 20 s into the restore

the five things
what the customer saw „Visszaállítás elindult — az állapot itt frissül." then the box went dark mid-restore; on return the dashboard and every app were back
what the box did by itself everything — containers 0 → 25 at t+131s → 26 at t+148s, nothing stuck, no shell used. gokapi, the app being restored when the plug came out, returned Up 30 seconds (healthy)
time to steady 148 s, measured against the container count this round took itself before the accident
alarm fired / true? controller_started (info) — true and correct. No false alarm.
should have fired, did not none — per the ladder a 60-second outage yields no node_stale (30 min threshold) and no app_start_failed (90 s boot grace), and neither appeared

Household loop: NOT COLLECTED. The loop was a transient unit on the VM and died with the power cut — the first accident that could have produced household failures instead produced no lines at all. Zero lines is not zero failures, so it is recorded as not collected, and the loop is now a persistent systemd unit that returns with the box.

A finding this round handed over: app_oom (warning) — „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította". That is immich's whole mystery solved: its Postgres was OOM-killed during the reverse-geocoding import, which is why the server saw CONNECTION_CLOSED and crash-looped twelve times. The controller caught an OOM inside an LXC guest and named the exact container — worth recording against this project's standing finding that those signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in the alarm feed, correctly labelled, the whole time.

Round 3 — use bookstack / accident: system disk filled to 96 % for ten minutes

the five things
what the customer saw nothing. wiki, status and paste all answered before, during and after. No banner, no warning, no mail — the household was never told the disk was full
what the box did by itself kept all twelve apps running on a 96 %-full root filesystem and released the space cleanly when the fill was removed (29 G used → 944 M used). The shared thin pool never moved (39.69 %) and the filesystem stayed writable
time to steady the box never left steady — 26 containers before, 26 after, none restarted
alarm fired / true? none fired, checked twice independently after the fill was released
should have fired, did not disk_critical is defined at ≥95 % used and the disk sat at 96 % for ten minutes. But this is the ladder working as designed, not a miss: the fill-watch is a daily sweep (03:30) plus one check ~90 s after a controller start. Predicted before the round, confirmed after

Household loop: 12 operations, 0 failures. The household kept using its apps normally throughout.

The finding is the silence. The honest answer to „would the household be told their disk is full?" is no — unless the controller happens to restart while it is full. Here the timing was almost comic: the controller restarted at 21:28 after round 2's power cut, so its one opportunistic check ran about twenty seconds before the disk filled, and the next is not due until 03:30.

A correction, recorded where it happened: I twice labelled a mid-window reading „end of window", estimating the clock instead of reading it. The readings were unchanged, but „nothing yet, five minutes in" and „nothing in the whole window" are different findings. From here the end-of-window check is taken when the round's own runner reports completion.

Round 4 — offsite-run (mealie) / accident: tunnel killed for ten minutes

The drawn ACTION never ran. The runner aborted with „no session — mine, not the product's": round 2's power cut had rebooted the guest, /tmp is cleared on boot, and the dashboard password file lived there. Round 3 was a use round and never needed it, so round 4 was the first to find it gone. The accident was measured; the off-site run was not. Recorded as half-measured rather than re-run and presented as whole — re-running a round after seeing it fail is how a drill starts choosing its own results. The file now lives in /root, which survives a reboot.

the five things
what the customer saw from outside, the apps vanished for ~90 s (public route 530) and came back on their own (200); from inside the house, nothing — traefik answered 301 throughout
what the box did by itself repaired its own tunnel. cloudflared killed 21:41:33Z, running again 21:43:07.478Z (~97 s), with RestartCount=0 — so Docker's unless-stopped policy did not do it; the controller's protected-infra recovery redeployed it („[infra] deploying cloudflared →…")
time to steady the apps never stopped; the way IN was restored in ~97 s
alarm fired / true? two, correctly paired — health_critical (error) 21:43, health_recovered (info) 21:48. Exactly what the ladder predicts for a missing protected container, and the alarm was not a dead end
should have fired, did not none for the accident

Household loop: 10 operations, 0 failures — and that number is narrower than it looks. The loop does not follow redirects, so it measures „is the app serving on the box", never „can the household reach it from outside". It saw nothing while the public route was returning 530.

Round 5 — use privatebin / accident: docker restarted inside the guest

the five things
what the customer saw a gap well under a minute: 200 before, 404 two seconds after docker returned (traefik had not re-registered routes), serving again by 21:54:19Z — about 40 s of shut doors
what the box did by itself everything. Restart ran 21:53:27Z→21:53:41Z; all 26 containers back at t+16s; the controller returned with them, waited 51 s for the fleet to settle and found nothing boot-orphaned to repair — correct, since every container had already come back on its own policy
time to steady 16 s to 26 of 26 containers; ~40 s until the doors served. The slower number is the one a household feels
alarm fired / true? controller_started (info) — true and correct, and the only line the ladder expects: no app_start_failed (90 s boot grace), no liveness alarm
should have fired, did not none

Household loop: NOT SAMPLED — 0 lines, because the round lasted ~18 s and the loop samples every 2 minutes. Recorded as not sampled, never as a pass.

A trap avoided. The round's own alarm snapshot was taken six seconds after the controller started, and from it controller_started looked missing. A later reading shows it present at 21:53. An alarm cannot be called missing by a measurement taken before it could have fired.

Round 6 — backup-system (whole-guest backup) / accident: none (control round)

the five things
what the customer saw their apps went away and came back: during the backup 4 of 26 containers were up and every app answered 404 publicly; all 26 were serving again by 22:01:13Z. No banner, no mail — correctly
what the box did by itself quiesced the apps, snapshotted, ran the local tier, failed it, announced the failure with a retry schedule, then ran the off-site tier from the snapshot with the apps already back up — 21:59:54Z → 22:08:27Z (~8½ min), encrypted to ep0, taking no local disk at all
time to steady apps down ~21:55:30Z → 22:01:13Z ≈ 5m43s — contaminated by my own intervention (I killed the local leg at 21:59:45), so it is an upper bound on the quiesce window, not a clean measurement
alarm fired / true? whole_guest_backup_failed (error) — true, and better than true: „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)". It names the tier, not „the backup", and says what it will do next. The status surface agreed (target_id:"local", success:false, size_bytes:0). No alarm for the 22 apps it stopped — correct, those stops are suppressed
should have fired, did not none

The finding to carry forward: a whole-guest backup on this box cannot use its local tier — a ~29 GB source into a 14 GB root filesystem — and the product handles that honestly: it fails the tier, says which tier, schedules a retry, and still gets the data out of the house on the off-site tier. The local tier will keep retrying and keep failing on a box shaped like this one.

My intervention, and the correction it needed. I killed the local leg when two samples showed / falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nested PVE. The arithmetic stands, but I first wrote „I stopped the backup", which was wrong: I stopped one leg, and the off-site leg started by itself seconds later and succeeded. Counted as an intervention either way.

Round 7 — use nextcloud / accident: internet cut for ten minutes

Drawn as update; ran as use, because the catalog bump was never pushed — the catalog gates returned INCONCLUSIVE and undetermined is never a pass. Decided and recorded before the round, not after it.

the five things
what the customer saw depends where they stood. At home: nothing — traefik answered 301 throughout and every app kept serving. Away from home: ten minutes of nothing — the public route gave 502, then 200 about a minute after the block lifted
what the box did by itself kept all 26 containers running, needed no repair, and re-established the way in unaided: 530 three seconds after unblocking, all four apps 200 by 22:21:52Z — ~64 s
time to steady the apps never left steady; only the path in broke and healed, in ~64 s
alarm fired / true? none, and none should have — the ladder puts node_stale at 30 minutes and this was ten. Nothing false was raised either
should have fired, did not none

Household loop: 10 operations, 0 failures — narrower than it looks. It does not follow redirects, so it measured the box (up throughout) and was blind to the public outage.

The question this round was meant to answer, and honestly did not. Are alarms raised while the hub is unreachable retried and then silently dropped? Not exercised. The box reports every 15m0s (measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports — the next was due ~22:23:43Z, after it lifted. Nothing was attempted, so nothing could be lost. Recorded as not exercised, never as passed.

The fence held, and it was checked against a baseline taken beforehand. The accident flips a host-wide sysctl on demo-hp and inserts two physdev rules. Afterwards: -P FORWARD ACCEPT, 0 physdev rules, sysctl 0 — identical to the pre-round reading. demo-hp also carries guests 9201 and 9202, so an abandoned rule would have been a fence breach, not an untidy drill.

Round 8 — backup-app nextcloud / accident: internet cut for ten minutes

22:35:48Z–22:47:59Z. The app-data backup ran first and finished in 1 m 55 s (db_dump count 5, success:true, 22:37:48Z). The cut began six seconds later, so the two barely overlapped — and would not have interacted in any case: backup-app is the local app-data tier and needs no internet. The off-site tier is a different action.

the five things
what the customer saw Nothing at home. Every front door kept serving on the LAN for the whole ten minutes. From outside the house the sites were unreachable — the public path was down. 26 apps up before, 26 after, never fewer.
what the box did by itself Kept every container running, kept backing up, kept reporting to the hub, and rebuilt the public path unaided when the link returned. No restart, no intervention, noaction from me.
time to steady ≤43 s after the link returned (public doors 200 again at 22:48:38Z). Containers never left steady at all.
alarm fired / true? none fired, and none should have — node_stale is a 30-minute threshold and this was ten. Nothing false was raised.
should have fired, did not none

And the finding of the round is against my own instrument, not the box.

The hub report due at 22:38:43Z fell inside the cut — and it succeeded: „Hub report pushed successfully (15526 bytes)". It succeeded because hub.felhom.eu resolves to 192.168.0.192, a LAN address (measured from both the guest and the host), and my injector blocks everything except the LAN. So the accident named „internet gone" only ever removed the public path. The box never lost the hub, in round 7 or in round 8.

Two consequences, both stated plainly:

  1. The dropped-event question is still unmeasured after two rounds that appeared to measure it. Events pushed while the hub is unreachable are retried three times and then dropped permanently, with no queue — that behaviour has still never been seen live.
  2. My own memory file carries this exact warning („hub.felhom.eu resolves to the LAN here; an internet-cut drill must block it too") and I did not apply it. A warning that is written down and not read is worth nothing, which is the same class of failure as an unread alarm.

The injector is corrected for round 9's drawn internet cut so that the hub address is blocked too — from the VM's side, at the host's tap rule. The hub itself is never touched; the fence is kept. This is a repair to a broken instrument, not a re-draw: the drawn action, app and accident for every remaining round are unchanged.

A sixth mistimed reading, and the fix is in the runner now. The round's own AFTER step read the public doors three seconds after the unblock and reported 530 on all four. That reading could never have been fair. The re-measurement 43 s later read 200 on all four. The runner measured the recovery before the recovery could begin — so the runner is the thing that gets fixed, not the note.

Round 9 — use uptime-kuma / accident: internet cut for ten minutes, hub included

23:00:50Z–23:12:00Z. The first cut of the night that really removed the hub. The injector was corrected between rounds 8 and 9; the drawn action, app and accident were not changed.

the five things
what the customer saw Nothing at home. uptime-kuma answered 200 on all three reads during the action, and all four front doors read 200 at both post-accident readings. 26 apps before, 26 after, never fewer.
what the box did by itself Kept every container running while it was cut off from the hub and from its own host agent. Built its report on schedule, tried to push it three times over 1 m 40.8 s, gave up, and carried on serving. Both links repaired themselves the instant the block lifted — no restart, no repair action, nothing from me.
time to steady The apps never left steady. The doors read 200 at the first reading, 2 s after the unblock, and again 63 s later.
alarm fired / true? none fired, and none should have — node_stale is a 30-minute threshold. Nothing false was raised.
should have fired, did not none — but see the near-miss below, which is a finding in its own right.

What a lost report costs: measured, not assumed.

23:08:42 [INFO]  [scheduler] Running job: hub-report
23:08:42 [INFO]  [report] Building system report
23:10:23 [WARN]  [report] Push failed: … context deadline exceeded
23:10:23 [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts (took 1m40.813s)

Three attempts, then it stops. Nothing is queued. It gave up 31 seconds before the link returned. And that is correct: a report is a snapshot, so a lost one costs nothing — the next snapshot carries the same truth, and the controller says so itself („backing off (the 15-min cycle still reconciles)"). This is not the event path. A dropped event is a lost fact, not a stale copy of a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop behaviour is still unmeasured after three rounds of internet cuts.

And it came back by itself, on schedule. The very next scheduled report went through — 23:23:43Z, „Hub report pushed successfully (15354 bytes)” — exactly 15 minutes after the cycle that failed, and nothing was done to the box to achieve it. The reading was taken at 23:24:21Z, deliberately after the report was due, so it could not be premature. So the whole shape of a hub outage is now measured end to end: build → three attempts → give up → keep serving → next cycle succeeds → no alarm, no loss.

The near-miss, which is luck and not design. Last good report 22:53:43Z; next scheduled 23:23:42Z; node_stale trips at 30 minutes. The gap is 29 m 59 s. The staleness threshold is exactly twice the report cadence, so a single failed push spends the entire budget — one second of ordinary jitter and the operator is paged about a box that was healthy throughout and had already repaired itself. Filed as a register row.

A second fidelity fault in my accident, named like the first. The cut also severed the controller from its host agent (GET /backup/tiers to the link-local 169.254.253.1 timed out twice). In a real house an ISP outage does not do that — controller and agent share one machine. So „internet gone" as injected is broader than its name: internet, hub and local agent. No further internet cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9.

Round 10 — restore uptime-kuma / accident: hard reset, four seconds into the restore

23:26:06Z–23:29:46Z. The roughest pair drawn. The restore was accepted at 23:26:08Z (302, „Visszaállítás elindítva"); the reset button was pressed at 23:26:12Z, mid-write.

the five things
what the customer saw They pressed restore, were told it had started, and four seconds later the whole machine went dark. About two minutes of nothing. Then every app was back and every front door answered. Nothing ever told them what became of the restore.
what the box did by itself Booted, and brought 26 of 26 containers back with no help. Boot reconciliation named the one app it had to recover („1 app(s) recovered in 1 attempt(s): [paperless-ngx]"), sent a startup hub report at 23:28:26Z, and settled its health probes. No intervention.
time to steady 150 s — 0 containers at t+12 s, 25 at t+133 s, 26 at t+150 s. Doors 200 at both readings (23:28:42Z and 23:29:44Z).
alarm fired / true? one, true — controller_started (info). Exactly what the ladder expects after a reboot. No false alarm.
should have fired, did not none from the alarm ladder — but the restore silence below is a legibility gap, filed as a row.

The restore left no trace anywhere, and the product has no place to leave one. Four candidate status endpoints all 404 (/api/restore/status, /api/backup/restore/status, /backup/restore/status, /api/restore). /api/backup/status carries no restore field at all. On the pages, the only restore text is a button label and a JavaScript label expression. On disk, in the real data directory, there is no restore, lock or state file anywhere — and no file at all was modified in the reset window. An interrupted restore and a restore that never happened are indistinguishable, to the customer and to me.

CORRECTION, 00:24Z — the paragraph above is wrong and stays visible so the correction is too. The four endpoints I called were four I guessed, and all four were wrong. The real route, read out of the restore page's own JavaScript, is /api/backup/restore-status, and it exists: {"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}. So a restore status surface does exist. What is true — and is the better finding — is that after the reboot it is blank: started_at is the Go zero value, and the payload carries no last field at all, while the page's own script renders „ sikertelen." from st.last.message. The restore record is in-memory only and does not survive the machine stopping — precisely the case a hard reset creates, and precisely when a household would want to be told. The register row is corrected to say that instead. I found the real routes by asking the controller for its own rendered links, which is what I should have done before filing anything.

The limit of that measurement, stated rather than glossed. Only four seconds elapsed, so the restore may have finished or may never have written a byte — and I cannot tell, because the controller's log stream holds zero lines before 23:28:00Z (a reset starts it fresh) and the debug ring died with the machine. What is independently verifiable is the absence of any restore record, and that is what is filed; it holds however far the restore got.

Four of my own instruments failed in this round, and all four are recorded in the evidence: a claim that was unfalsifiable when written; an on-disk check against a directory that does not exist; a household count that reported 0 lines and 0 failures when the truth was one line and it was a failure; and a disk guard that reported „active" all night while being a transient unit that vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box.

Round 11 — use paperless-ngx / accident: the data drive pulled out for twenty minutes

23:50:49Z–00:13:09Z. The drive was detached from the running box at 23:50:54Z and put back at 00:10:58Z. This is the round that produced the most alarms of the night, and every one of them was true.

the five things
what the customer saw The four apps whose files live on that drive stopped — Paperless, Jellyfin, Nextcloud, Immich — and Paperless's front door went 404. The other eleven apps kept serving normally throughout. About twenty minutes later everything was back, roughly a minute after the drive was plugged in again.
what the box did by itself Noticed the drive had gone and named it by the label the household sees („Adatlemez"), named each broken app individually, degraded its own health, waited, noticed the drive return, restarted the apps and recovered its health. No restart, no repair, nothing from me.
time to steady 67 s after the drive returned (26 containers again at 00:12:05Z). The door followed at 00:13:06Z.
alarm fired / true? eight, all true, correctly paired at both ends — storage_disconnected (error) → four app_start_failed (warning) → health_degraded (warning) → storage_reconnected (info) → health_recovered (info). The four apps named are exactly the four with data on the pulled drive. Nothing false was raised.
should have fired, did not none

The alarm that looked missing, and was not. The round's own snapshot at 00:13:08Z showed health_degraded with no recovery — which would have been the first missing alarm of the night. The recovery fired at 00:13, and the snapshot missed it by seconds. A re-read at 00:14:13Z, taken after the apps were back, found it. This is the discipline from the earlier mistimed readings earning its keep: when a measurement could have been early, it is re-taken rather than turned into a verdict.

The drive came back clean — /dev/sdd, 98 G, 2 % used, mounted at /mnt/felhom-drives/hdd_1, with the guest's mountpoint config unchanged, 26 containers up and every real front door serving.

Round 12 — use paperless-ngx / accident: nothing (closing control round)

00:15:48Z–00:16:55Z. The night's last round, drawn as a control.

the five things
what the customer saw Nothing at all. The app answered 200 on all three reads, and every front door answered 200 at both readings.
what the box did by itself Nothing needed doing. 26 containers before and after.
time to steady 1 s — it never left steady.
alarm fired / true? none, and none should have. The newest entry in the feed is still round 11's health_recovered at 00:13.
should have fired, did not none

Checked rather than assumed: inject.sh has no „nothing" case — its default branch exits 2 on an unknown accident. The control rounds never reach it, because the runner handles the no-accident case itself and says so („accident: none — control round, deliberately"). This matters because a broken injector produces exactly the same result as a control round, and the only way to tell them apart is to look at which code path ran.

What the closing control round is worth. It shows the quiet is real: after eleven rounds of power cuts, resets, full disks, severed networks and a drive pulled out of a running machine, a round in which nothing was done produced nothing — no alarm, no restart, no drift. The alarm feed is not simply noisy.

Phase 2 — the morning after

Every app answered through its own front door. All twelve real names, measured at 00:19:16Z — cloud, inventory, media, paperless, paste, photos, recipes, share, status, travel, vault, wiki — 200 on the public path, every one. And healthy is not inferred from a door: all 26 containers report Up … (healthy), except traefik and cloudflared, which carry no healthcheck and show a bare Up. The four showing seven minutes are the drive-backed apps restarted after round 11.

The off-site restore could not be done, and two independent instruments agree why. The brief asked for one DB-backed app restored from off-site onto scratch 9202. It has nothing to restore from:

instrument answer
restic, with the box's own key, password file and known_hosts Fatal: wrong password or no key found — exit status 1, read from restic itself rather than from the end of a pipeline
the product's own status surface, which is what a customer sees {"orphaned":true, … ,"snapshots":0,"status":"error"}

The repository is orphaned because this box is a rebuild for an existing customer: on a rebuild the restic password is minted fresh, so snapshots written under the old one can never be opened again. That is a known, documented shape, and the product surfaced it honestly — the true offbox_repo_orphaned alarm of round 1, mailed to the operator within two seconds of the run.

What did leave the house. The whole-guest off-site copy is a different store and it worked: ep0 holds two intact snapshots for this box, including tonight's at 21:59:54Z — round 6's off-site leg, the one that started by itself after I killed the local leg. Its file index is roughly four times the afternoon copy's, consistent with a guest by then carrying twelve apps. Stated as a limit: that is a listing, not a verification. A PBS verify job would prove restorability and it writes verify state, so it was not run — ep0 is read-only for evidence tonight.

The household loop. 204 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of which only two are real events — wiki during round 2's power cut and cloud during round 10's hard reset, each a single sample, each healed before the next probe. The other three were my own classifier counting a 301 redirect as a dashboard failure; the log carries that correction in its own words at 21:11:25Z and the wrong lines were left in place so the correction stays visible. Ten of the twelve rounds left no mark at all in this log, including the twenty minutes with the drive pulled — the loop samples each name every two minutes, so that silence is the instrument's sampling rate and not evidence the household saw nothing.

The catalog bump: verified reverted, not remembered. Clean tree, local exactly level with origin/main (0 ahead, 0 behind), the nextcloud template still on its original redis:7-alpine pin and catalog_since: "2026-07-18", newest commit 2026-09-15. The bump was prepared, refused by the catalog's own gates (image-resolvable and volume-persistence both INCONCLUSIVE — its own canary failed, so the verdict was UNDETERMINED, never a pass) and reverted before any push.

And one delivery proof the truth table could not give. The mailbox shows every alarm arrived at the operator, not merely that it was stored — round 11's whole set, round 6's tier-naming backup failure, round 1's orphan warning, the OOM, the deploy failure and my own thin-pool storage_fill_critical. In a project where 91 events once sat in a database having e-mailed nobody, raised and delivered are two different claims, and only one of them had evidence before tonight.

Interventions — counted, with the reason for each verdict

One. At 21:59:45Z in round 6 I killed the local leg of the whole-guest backup. The arithmetic that forced it: a ~29 GB source being written into a 14 GB root filesystem at ~16 MB/s, i.e. under four minutes to a full / on the nested host, mid-round. What it cost: a leg that could never have succeeded. What happened next without me: the off-site leg started by itself from the same snapshot and succeeded in about eight and a half minutes. Filed as R-548.

My first note on it said „I stopped the backup". That was wrong and is corrected in the evidence: I stopped a leg of it, and the box completed the other one unaided.

Both pre-declared presses went unused. O1, the „Send self-bind link" button, was not needed — the automatic mail was already waiting (18:17:46Z) and the box bound with zero operator presses. O2, „Re-issue PBS credentials", was not needed either — the acknowledged-delete path re-issued them by itself (pbsdr_auto_reissue, 20:19Z). Both prompt claims they were insurance against turned out to be true, and the F-14 path was measured live for the first time.

Counted separately, because it is not a round result: Phase 0's seeding repairs. I filled the LVM thin pool to 100 % by firing twelve deploys at once, then repaired the damage — a guest restart to clear an emergency_ro remount, dropping corrupt image layers, a remove-with-data and one consistent re-deploy after a re-seed minted fresh database passwords over initialised volumes, and a rewritten APP_KEY. The damage was mine, not the product's, the product's behaviour throughout was correct, and every repair went through the product's own endpoints rather than by hand-running compose. Listed in full in evidence-chaos-night-2026-09-17/interventions.txt so the distinction is visible rather than convenient.

Not counted, and why: acts on my own instruments — moving the dashboard password after a power cut cleared /tmp, rewriting the injector and the runner, re-creating the disk guard as a real unit. Counting those would flatter the night in one direction and pad the stop-rule count in the other. Round 4 is the clearest case of declining to intervene: the tunnel was left dead on purpose — „NOT restarting it by hand — whether it returns by itself IS the measurement". It returned by itself in 97 s.

Standing against the stop rule: 1 of 4. The night ran its full twelve rounds.

Teardown — three layers, stated

Machine — gone. VM 336 stopped and destroyed with all three disks purged (00:35:25Z), gated on its name rather than its number because two standing guests share the host. qm list shows no VMs; /mnt/hdd_1/images/336 no longer exists. The storage was verified to be /mnt/hdd_1 from its own definition (nvme-scratch, path /mnt/hdd_1, is_mountpoint yes) rather than assumed. The machine had three disks, not the two the brief asked for — the third was mine, added in Phase 0 after I filled the thin pool — and that is recorded rather than quietly removed.

Host — clean, measured before and after. nvme-scratch 6.78 % → 1.61 % (~48.5 GB returned); local-lvm unchanged at 44.75 %, so the fence that said never local-lvm held; free space on /mnt/hdd_1 827 G → 875 G. Guests 9201 and 9202 still running. The household loop and the disk guard were stopped and disabled after their logs were copied off (the guard's log was 0 bytes — it never fired). Firewall back to -P FORWARD ACCEPT with 0 physdev rules, so none of the three network accidents left a rule behind. Scratch 9202: nothing to remove, shown rather than said — three infrastructure containers, no app deployed, no stack touched since 21:00, no off-site config.

Hub — host record deleted through the acknowledged flow. The first attempt at 00:38:37Z was correctly refused (409, „Host is ONLINE") — the box had died inside the hub's liveness window. The acknowledged delete went through at 07:25:13Z (confirm_host_id + delete_escrow=1 → 303). Every line of the after-state written down before the act matched: the host answers 404; drill-r50, both demo hosts and the tester-1 customer still answer 200; the customer now lists zero hosts. The automatic connect mail arrived two seconds later (07:25:15Z, „Kösd össze a Felhom dobozodat"), quoted in full with its token redacted in teardown-hub.txt — and it is provably tonight's, because the mailbox held no such mail newer than 18:17:46Z when checked at 00:38Z.

Why the hub layer finished six hours late — my fault, not the product's. The retry was guarded by „don't post while the host page contains ONLINE". That word lives in a JavaScript string that is always on the page, so the guard could never pass. It refused six times and gave up at 01:20Z, while the hub's structured answer would have said "status":"down" from about 00:54Z. Nothing ran again until 07:24Z.

ep0 — backups stayed, nothing removed. Read three times: 00:17:15Z, 00:36:52Z (just before the delete) and 07:25:35Z (after). Identical every time — three namespaces, two snapshots each, six in total, 16 G used. No prune, no verify, no write.

Claims in the prompt that turned out wrong — named first, as asked

The two the brief itself flagged both turned out TRUE, and both were checked tonight rather than assumed.

  1. „The automatic mail is waiting in the mailbox." The brief warned this had been read from yesterday's host delete and not verified. It was true. The self-bind mail of 18:17:46Z was in the mailbox, and the box bound with zero operator presses — so the pre-declared press O1 was never needed.
  2. „The WG hook provisions by itself after an acknowledged delete" (the F-14 path). The brief noted this had never been measured live. It was true, and it was measured live for the first time: pbsdr_auto_reissue at 20:19Z — „Previous key destroyed (acknowledged deletion) — credentials re-issued automatically." The second pre-declared press, O2, was never needed either.

Now the ones that were wrong.

  1. WRONG: „restore one DB-backed app from off-site onto scratch 9202." It could not be done at all on this box, and not because anything broke. This box is a rebuild for an existing customer, so its restic password was minted fresh and the snapshots already in the remote store can never be opened by it again. Two independent instruments agree: restic itself (Fatal: wrong password or no key found, exit 1) and the product's own status (orphaned:true, snapshots:0, status:"error"). The brief assumed an off-site app repository this box could open; on a rebuild fixture there is none.
  2. WRONG in effect: „a system disk + one data disk." The machine ended the night with three disks. The third, 64 G, was added by me in Phase 0 to extend the LVM thin pool after I filled it to 100 % by firing twelve deploys at once. The deviation is mine, not the brief's, but the fixture was not the one the brief described and saying so is the point.
  3. WRONG: round 7's drawn action update was not performed as drawn. The catalog's own gates returned image-resolvable INCONCLUSIVE and volume-persistence INCONCLUSIVE — its own canary failed, so the verdict was UNDETERMINED, which is never a pass. The round ran use instead. A deviation from the drawn schedule, logged rather than quietly substituted.
  4. WRONG, and mine rather than the brief's: „an internet cut tests what happens when the hub is unreachable." The accident's name implies it; on this network it was false. hub.felhom.eu resolves to a LAN address here, and my injector allowed the whole LAN — so rounds 7 and 8 cut the public path only, and the box never lost the hub. My own memory file carries that exact warning and I did not apply it. Fixed between rounds 8 and 9 by blocking the hub address from the VM's side, which is what finally made round 9 the measurement it was supposed to be.
  5. WRONG as a description of the night's clock: the schedule table's times. The table drawn from the seed lists rounds at 23:30 through 04:05. Those were nominal. The real spacing was 25 minutes from each round's actual start, and the night's twelve rounds finished at 00:17Z, roughly four hours earlier than the table's own column suggests. Each round's real timestamps are recorded in its own section; the drawn order, apps and accidents were never changed — only the wall-clock the table guessed at.

And one the brief did not make, which the night could not answer. Events pushed while the hub is unreachable are retried three times and then dropped permanently, with no queue. Three ten-minute hub outages happened and no event was raised during any of them, so that path is still unmeasured. What was measured is the report path: built, three attempts over 1 m 40.8 s, given up, and the next scheduled report succeeded — and a report is a snapshot, so nothing was lost.