69c08b183b
gates / gates (push) Successful in 22s
The acknowledged delete went through at 07:25:13Z (confirm_host_id + delete_escrow=1 -> 303). Every line of the after-state written down before the act matched: the host 404s; drill-r50, both demo hosts and the tester-1 customer still 200; the customer lists zero hosts; ep0 identical across three readings - six snapshots, 16G, nothing removed. The automatic connect mail arrived two seconds later and is quoted with its token redacted. It is provably tonight's: the mailbox held no such mail newer than 18:17:46Z when checked at 00:38Z. Why the hub layer finished six hours late is mine: the retry guard refused to post while the host page contained the word ONLINE, and that word sits in a JavaScript string that is always on the page. The hub's structured status said 'down' from about 00:54Z. It gave up at 01:20Z and nothing ran again until 07:24Z. Eleventh instrument fault of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
714 lines
54 KiB
Markdown
714 lines
54 KiB
Markdown
# DRILL — CHAOS NIGHT: random actions on random apps while random things go wrong (2026-09-16/17)
|
||
|
||
**Interventions: 1.** One, at 21:59:45Z in round 6: I killed the **local leg** of a whole-guest
|
||
backup that could never have fit (a ~29 GB source into a 14 GB root filesystem, falling at ~16 MB/s).
|
||
The **off-site leg then ran by itself from the same snapshot and succeeded**, so the data still left
|
||
the house. Both pre-declared presses went **unused**: the automatic self-bind mail was already
|
||
waiting, and the acknowledged-delete path re-issued the PBS credentials on its own. Phase 0's
|
||
seeding repairs are listed separately in `evidence-chaos-night-2026-09-17/interventions.txt` — that
|
||
damage was mine, not the product's, and every repair went through the product's own endpoints.
|
||
|
||
**Ready for a volunteer: still yes.** Across twelve rounds — a power cut mid-restore, a hard reset
|
||
four seconds into another, a full disk, a killed tunnel, a restarted Docker, three severed networks
|
||
and the data drive pulled out of a running machine for twenty minutes — **nothing cost a byte of
|
||
customer data, and the box healed itself every single time with no human involved.** Seventeen
|
||
alarms fired, **all seventeen were true, none were missing**, and the mailbox proves every one was
|
||
**delivered** rather than merely stored. The honest qualifications: two P2 legibility gaps are filed
|
||
(an interrupted restore leaves no record; one failed report spends the whole 30-minute staleness
|
||
budget, measured at 29 m 59 s), and one thing this night could **not** test — per-app off-site
|
||
restore, because this box is a rebuild whose restic repository is orphaned **by design**, which the
|
||
product surfaced honestly within seconds.
|
||
|
||
**The accident-plus-action pair that hurt most: `restore` + hard reset (round 10).** Not because the
|
||
box suffered — it was back with 26 of 26 containers in **150 s**, boot reconciliation naming the app
|
||
it recovered, every front door serving. It hurt most because it is the **only** pair of the night
|
||
where the household is left not knowing what happened: they pressed restore, were told it had
|
||
started, the machine went dark four seconds later, and afterwards **nothing anywhere tells them
|
||
whether it finished.** The status surface exists and answers with the zero value; the record is
|
||
in-memory only and does not survive the machine stopping.
|
||
|
||
> **Baselines, verified live against Gitea at 21:49 CEST 2026-09-16 (not copied from the brief):**
|
||
> felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 ·
|
||
> felhom.eu `d124c77e176d` hub v0.116.0, ISO **1.28.0 published** · app-catalog `94bc5febaca2`.
|
||
> All four trees clean and in sync. Highest register row **R-545**, 212 open. Golden waiver valid to
|
||
> 2026-09-27. Customer **`tester-1`** (`enkicsifelhom.hu`, `tester1@felhom.eu`, no host).
|
||
> Venue: `demo-hp` (Tier 0), a fresh nested VM, disk on the NVMe at its root. Evidence:
|
||
> `evidence-chaos-night-2026-09-17/`.
|
||
|
||
## The schedule — drawn ONCE, before round 1, and written here first
|
||
|
||
The point of this section's position in the document is that the night could not be chosen after the
|
||
fact. `chaos_schedule.py` is committed beside the evidence; re-running it reproduces this table.
|
||
|
||
- **seed:** `20260917` (the date)
|
||
- **script sha256:** `4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1`
|
||
- **generator:** `evidence-chaos-night-2026-09-17/chaos_schedule.py`, stdlib `random` seeded with the seed
|
||
|
||
| # | time | X — the action | Y — the app | Z — the accident |
|
||
|---|---|---|---|---|
|
||
| 1 | 23:30 | offsite-run | adventurelog | nothing |
|
||
| 2 | 23:55 | restore | gokapi | power cut |
|
||
| 3 | 00:20 | use | bookstack | disk 95% full |
|
||
| 4 | 00:45 | offsite-run | mealie | tunnel down 10min |
|
||
| 5 | 01:10 | use | privatebin | docker restarted |
|
||
| 6 | 01:35 | backup-system | adventurelog | nothing |
|
||
| 7 | 02:00 | update | nextcloud | internet gone 10min |
|
||
| 8 | 02:25 | backup-app | nextcloud | internet gone 10min |
|
||
| 9 | 02:50 | use | uptime-kuma | internet gone 10min |
|
||
| 10 | 03:15 | restore | uptime-kuma | hard reset |
|
||
| 11 | 03:40 | use | paperless-ngx | drive pulled 20min |
|
||
| 12 | 04:05 | use | paperless-ngx | nothing |
|
||
|
||
**Re-draw log** — a silent re-draw is a schedule chosen by the person running it, so every one is here:
|
||
|
||
- r02 X=reinstall re-drawn (nothing has been removed yet)
|
||
- r06 Z=disk 95% full re-drawn (constraint 4: at most once)
|
||
- r07 Z=nothing re-drawn (constraint 6: never two in a row after r2)
|
||
- r08 X=reinstall re-drawn (nothing has been removed yet)
|
||
|
||
**What the draw happened to give, said plainly before the night judges it:** no `remove` round was
|
||
ever drawn, so `reinstall` had nothing to reinstall and was re-drawn twice (rounds 2 and 8). Three
|
||
`internet gone` rounds land consecutively (7, 8, 9) — that is the seed's doing, and it makes rounds
|
||
7–9 a de-facto endurance test of the same accident against three different actions rather than three
|
||
independent samples. `controller killed`, `drive pulled 90s`, `memory pressure`, `agent restarted`
|
||
and `hub unreachable` were never drawn at all; **this night does not test them**, and the morning
|
||
verdict must not claim it did.
|
||
|
||
## Phase 0 — the golden, the box, the household
|
||
|
||
**0.1 Golden 0.245.0, baked and published.** Launched 19:52:47Z as a transient unit in the drill VM,
|
||
finished 19:58:15Z. Markers: `overlay2`=1, `including mount point`=2, `upload OK (HTTP 201)`=1,
|
||
FATAL=0, publish-skipped=0. `GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626`.
|
||
**The teardown was gated on the REGISTRY answering 200**, not on an exit code — and that mattered:
|
||
the wrapper exited **144** while every measured outcome was good. Vouched in the hub and read back
|
||
from the page (`golden currently vouched: 0.245.0`). The three bake failures of 2026-09-16 (`scp -P`,
|
||
`chmod 0700`, `GITEA_USER=admin`) were each guarded and none recurred.
|
||
|
||
*Decision, recorded because silence reads as agreement:* the global controller floor was left at
|
||
0.244.0. The new box installs golden 0.245.0, which already carries controller 0.245.0, so no floor
|
||
was needed to deliver anything tonight; raising it would have pushed an update onto demo-felhom, a
|
||
box not in this drill.
|
||
|
||
**0.2 The box.** VM 336 on demo-hp: 8 GiB, 4 cores, 32 G system + 100 G data disk on the NVMe at its
|
||
root, booted from the **published** ISO 1.28.0. Boot order set in its own `qm set` (combining it
|
||
silently yields `order=net0;ide2`). Install completion was judged **from the disk** — blocks used
|
||
grew 3233 → 6942 MiB then held across three checks — because „Automatically reboot" is ticked and a
|
||
finished install looks exactly like a stuck one on screen. The summary page was read before pressing
|
||
Install, and the line that made it safe was **„Disk(s): /dev/sda"** — the 32 G system disk alone.
|
||
|
||
**The walk, as a volunteer, cost ZERO operator presses.** The box registered itself as an unclaimed
|
||
appliance and polled, visibly, until bound. The bind link came from the **waiting mail** (minted
|
||
18:17:46Z by yesterday's acknowledged host delete), the pairing code off the box's own console
|
||
(`4SY-4TX`), the „Tulajdonosi jelmondat" from the hub's customer record:
|
||
POST /bind/<token> -> 200, „Sikeres összekötés."
|
||
Then day-0 ran on its own and the hub recorded, without anyone pressing anything:
|
||
`appliance_bound` (customer_selfbind) · `appliance_credential_delivered` · `claim_reissued_reenroll`
|
||
· `offsite_reissued` · **`pbsdr_auto_reissue` — „Previous key destroyed (acknowledged deletion) —
|
||
credentials re-issued automatically."**
|
||
|
||
**That last event is a first.** The brief named „the WG hook provisions by itself after an
|
||
acknowledged delete" as a claim never measured live. It ran tonight, unprompted. **Both pre-declared
|
||
presses (O1 self-bind, O2 re-issue) were therefore unnecessary.**
|
||
|
||
The dashboard was claimed with the mailed code and **proven by logging in with the new password** —
|
||
a claim page that re-renders looks identical to success from the status code alone. The 100 GB data
|
||
drive was initialised through the wizard's own endpoint (`POST /api/storage/init`, polled to
|
||
`phase: done`), and `df` shows it mounted at `/mnt/felhom-drives/hdd_1` with 93 G free.
|
||
|
||
**The box landed on tonight's golden with no hand upgrade:** controller **0.245.0** (healthy), agent
|
||
**0.131.0**, host `tester-1-022354` ONLINE. And the **R-543 escrow reminder bar shipped hours earlier
|
||
was live on it**, on a box nobody had touched.
|
||
|
||
**0.3 The household — and the first real trouble.** Twelve deploys were fired; **ten were accepted,
|
||
two refused** for memory with both numbers quoted. Then nine of the ten failed: the guest's disks are
|
||
thin-provisioned over an ~11.8 GB pool carved from a 32 GB system disk, ten simultaneous image pulls
|
||
filled it, and the hub recorded `storage_fill_critical` (100 %) plus **nine `app_deploy_failed`
|
||
warnings**, one per app, each naming the failing pull. Only PrivateBin installed.
|
||
|
||
**The product behaved; the harness did not.** R-536's failure event — shipped that same morning so an
|
||
interrupted install is not silence — fired for all nine within two minutes. The memory guard refused
|
||
rather than over-committing. The per-stack record stayed honest (`deployed: false`). The two faults
|
||
were mine: firing twelve deploys in two seconds is not household behaviour, and a 32 GB system disk
|
||
was copied from an earlier drill without checking what that drill had installed.
|
||
|
||
**0.4 The escrow ceremony could NOT be completed — and this one is about tonight's own release.**
|
||
See „Finding: the recovery-code step cannot be done when the guide says to do it" below.
|
||
|
||
|
||
## Finding: the recovery-code step cannot be done when the guide says to do it (R-546)
|
||
|
||
The box was at exactly the point of the guide this release added hours earlier — installed, bound
|
||
with no press, claimed, drive initialised, **no apps yet** — and the escrow reminder bar was on every
|
||
page telling the household to create their recovery code. It could not be done.
|
||
|
||
POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"}
|
||
GET /api/escrow/status -> claimable:false —
|
||
detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"
|
||
POST /api/escrow/claim -> **409** „A folyamat jelenlegi állapotában a kód nem kérhető le."
|
||
|
||
**Both sides agreed on the cause.** The hub's own Backup & DR panel read „host enrolled **done** · WG
|
||
tunnel peer registered **done** · descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1)
|
||
**waiting** · ceremony possible once the descriptor is applied on the box". The box had no PBS storage
|
||
(`pvesm status`: `local`, `local-lvm` only) and no `escrow` section in `agent.json` at all.
|
||
|
||
**It self-heals, and that was measured rather than assumed.** The box was left alone and polled:
|
||
|
||
20:30:12Z … 20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
|
||
**20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs**
|
||
|
||
~17 minutes after the bind. The retried ceremony passed every preflight item, the claim returned
|
||
**200** (83-character code, 129.2 bits of entropy, revealed once), and `escrow_state` became
|
||
**escrowed**. The bar then vanished from all four pages checked — the R-543 fix working through its
|
||
whole lifecycle on a box nobody had set up for the test.
|
||
|
||
**So the defect is timing and wording, not mechanism.** For ~17 minutes a volunteer following
|
||
tonight's guide meets a stderr fragment about a `-storage` flag, while every page urges them on.
|
||
Filed **R-546** (P2). No product code was changed — this is a validation run.
|
||
|
||
## Phase 1 — the rounds
|
||
|
||
Rounds run at ~25-minute spacing. The schedule above is fixed; only the wall-clock start moved,
|
||
because Phase 0 ran long (the storage wizard submits by JavaScript and the endpoint was worked out
|
||
rather than guessed). Round 1 began 23:07 CEST.
|
||
|
||
### Round 1 — `offsite-run` / adventurelog / accident: **nothing** (control round)
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | „A távoli mentés elindult — az állapot itt frissül." and, at the end, „A távoli mentési tároló elárvult: a benne lévő mentések egy korábbi, már nem elérhető kulccsal készültek (újratelepítés)." |
|
||
| what the box did by itself | walked all twelve apps — stop, dump each volume with real byte counts, restart — captured eleven, could not capture the one that was crash-looping, and finished |
|
||
| time to steady | **1m45s** (`last_duration`), `last_run` 21:09:00Z, `progress.active` false, `last_error` empty. A control round: the box never left steady |
|
||
| alarm fired / true? | **three, all true** — `app_start_failed` named Nextcloud · `backup_run_failures` „1 of 12 apps failed to back up in this nightly run: nextcloud" · `offbox_repo_orphaned`, matching the status endpoint's own `"orphaned": true` |
|
||
| should have fired, did not | **none** |
|
||
|
||
**Household loop in the window:** 3 lines marked FAILED, **all three mine** — the loop counted the
|
||
dashboard's 301 redirect as a failure while accepting the same 301 for app reads. Fixed at 21:11:25Z
|
||
and marked in the log; only lines after that marker are scored.
|
||
|
||
**What round 1 actually establishes.** The off-site tier is armed (escrowed) and the run works
|
||
end-to-end, but on THIS box — a rebuild for an existing customer — the remote repository was written
|
||
under a key the box no longer holds, so **no snapshot was written**. That is the documented rebuild
|
||
behaviour, surfaced honestly with the route out named in the message rather than reported as success.
|
||
|
||
**Two things that looked like defects in this round and are not**, both established with controls
|
||
rather than inference — five front doors answering 404 (traefik has no route to an unhealthy
|
||
container; identical byte-for-byte to a no-such-host control) and a crash-looping Nextcloud (image
|
||
layers corrupted while my thin pool stood at 100 %; „invalid ELF header", repaired by a re-pull).
|
||
Detail in `evidence-chaos-night-2026-09-17/round-1-notes.txt`.
|
||
|
||
|
||
### Phase 0 postscript — repairing my own damage, and three conclusions I had to retract
|
||
|
||
Nine of the first ten deploys failed because I fired twelve at once onto a thin pool carved from a
|
||
32 GB disk, and the pool hit 100 %. Repairing that took the rest of Phase 0 and produced **two
|
||
distinct faults of mine, with different cures**, which only separating them made fixable:
|
||
|
||
| fault | symptom | cure |
|
||
|---|---|---|
|
||
| image layers written while the pool was full | `php: … libxml2.so.2: **invalid ELF header**`, exit 127 crash loop | drop the image, let compose pull it again |
|
||
| my re-seed generated **fresh database passwords** over volumes already initialised with the first set | Postgres `auth_failed`, MariaDB „Access denied for user … (using password: YES)" | remove the app **with its data**, deploy once with one consistent secret set |
|
||
|
||
A fresh image did not fix gokapi and a fresh database did — that is the evidence the two faults are
|
||
different things rather than one.
|
||
|
||
**Three conclusions I wrote and then had to retract, each corrected where it stood:**
|
||
1. „the five 404s were my mistimed sweep" — wrong for four of them. A **negative control** (a
|
||
no-such-host request) returned the identical 404 of 19 bytes, and a positive control returned
|
||
200/1200 bytes: traefik simply has **no route to an unhealthy container**.
|
||
2. „nextcloud is repaired" — wrong. The re-pull fixed the crash, and the app still could not reach its
|
||
database. The container reported **`healthy`** throughout, because the image's own healthcheck asks
|
||
whether Apache answers, not whether the application works.
|
||
3. „all the broken apps are corrupt layers" — wrong. bookstack logged a clean startup, gokapi logged
|
||
nothing at all, immich showed a Postgres auth failure.
|
||
|
||
I also nearly filed a defect against the drive gate, which was working and logging at DEBUG while I
|
||
read a settings snapshot inside its 30-second tick. **An absent log line is not evidence.**
|
||
|
||
None of this is a product defect and none of it is filed as one. What the product did throughout was
|
||
correct and legible: it refused to route to unhealthy containers, `app_start_failed` named the app
|
||
that was down, `backup_run_failures` said „1 of 12 … nextcloud", the memory guard refused the
|
||
eleventh and twelfth installs with both numbers quoted, and `storage_fill_critical` fired at 100 %.
|
||
|
||
### Round 2 — `restore` gokapi / accident: **power cut**, 20 s into the restore
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | „Visszaállítás elindult — az állapot itt frissül." then the box went dark mid-restore; on return the dashboard and every app were back |
|
||
| what the box did by itself | everything — containers 0 → **25 at t+131s** → **26 at t+148s**, nothing stuck, no shell used. **gokapi, the app being restored when the plug came out, returned `Up 30 seconds (healthy)`** |
|
||
| time to steady | **148 s**, measured against the container count this round took itself before the accident |
|
||
| alarm fired / true? | `controller_started` (info) — true and correct. **No false alarm.** |
|
||
| should have fired, did not | **none** — per the ladder a 60-second outage yields no `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace), and neither appeared |
|
||
|
||
**Household loop: NOT COLLECTED.** The loop was a transient unit on the VM and died with the power
|
||
cut — the first accident that could have produced household failures instead produced no lines at
|
||
all. Zero lines is not zero failures, so it is recorded as not collected, and the loop is now a
|
||
persistent systemd unit that returns with the box.
|
||
|
||
**A finding this round handed over:** `app_oom` (warning) — „Alkalmazás memóriája elfogyott: immich
|
||
(immich-postgres) — egy folyamatát a memóriakorlát leállította". That is immich's whole mystery
|
||
solved: its Postgres was OOM-killed during the reverse-geocoding import, which is why the server saw
|
||
`CONNECTION_CLOSED` and crash-looped twelve times. **The controller caught an OOM inside an LXC guest
|
||
and named the exact container** — worth recording against this project's standing finding that those
|
||
signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in
|
||
the alarm feed, correctly labelled, the whole time.
|
||
|
||
### Round 3 — `use` bookstack / accident: **system disk filled to 96 % for ten minutes**
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | **nothing.** wiki, status and paste all answered before, during and after. No banner, no warning, no mail — the household was never told the disk was full |
|
||
| what the box did by itself | kept all twelve apps running on a 96 %-full root filesystem and released the space cleanly when the fill was removed (29 G used → 944 M used). The shared thin pool never moved (**39.69 %**) and the filesystem stayed writable |
|
||
| time to steady | the box never left steady — **26 containers before, 26 after**, none restarted |
|
||
| alarm fired / true? | **none fired**, checked twice independently after the fill was released |
|
||
| should have fired, did not | `disk_critical` is defined at ≥95 % used and the disk sat at **96 % for ten minutes**. **But this is the ladder working as designed, not a miss:** the fill-watch is a daily sweep (03:30) plus one check ~90 s after a controller start. Predicted before the round, confirmed after |
|
||
|
||
**Household loop: 12 operations, 0 failures.** The household kept using its apps normally throughout.
|
||
|
||
**The finding is the silence.** The honest answer to „would the household be told their disk is
|
||
full?" is **no** — unless the controller happens to restart while it is full. Here the timing was
|
||
almost comic: the controller restarted at 21:28 after round 2's power cut, so its one opportunistic
|
||
check ran about twenty seconds *before* the disk filled, and the next is not due until 03:30.
|
||
|
||
**A correction, recorded where it happened:** I twice labelled a mid-window reading „end of window",
|
||
estimating the clock instead of reading it. The readings were unchanged, but „nothing yet, five
|
||
minutes in" and „nothing in the whole window" are different findings. From here the end-of-window
|
||
check is taken when the round's own runner reports completion.
|
||
|
||
### Round 4 — `offsite-run` (mealie) / accident: **tunnel killed for ten minutes**
|
||
|
||
**The drawn ACTION never ran.** The runner aborted with „no session — mine, not the product's":
|
||
round 2's power cut had rebooted the guest, `/tmp` is cleared on boot, and the dashboard password
|
||
file lived there. Round 3 was a `use` round and never needed it, so round 4 was the first to find it
|
||
gone. **The accident was measured; the off-site run was not.** Recorded as half-measured rather than
|
||
re-run and presented as whole — re-running a round after seeing it fail is how a drill starts
|
||
choosing its own results. The file now lives in `/root`, which survives a reboot.
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | from outside, the apps vanished for ~90 s (public route **530**) and came back on their own (**200**); from inside the house, nothing — traefik answered **301** throughout |
|
||
| what the box did by itself | **repaired its own tunnel.** cloudflared killed 21:41:33Z, running again **21:43:07.478Z (~97 s)**, with `RestartCount=0` — so Docker's `unless-stopped` policy did *not* do it; the controller's protected-infra recovery redeployed it („[infra] deploying cloudflared →…") |
|
||
| time to steady | the apps never stopped; the way IN was restored in **~97 s** |
|
||
| alarm fired / true? | **two, correctly paired** — `health_critical` (error) 21:43, `health_recovered` (info) 21:48. Exactly what the ladder predicts for a missing protected container, and **the alarm was not a dead end** |
|
||
| should have fired, did not | none for the accident |
|
||
|
||
**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop
|
||
does not follow redirects, so it measures „is the app serving on the box", never „can the household
|
||
reach it from outside". It saw nothing while the public route was returning 530.
|
||
|
||
### Round 5 — `use` privatebin / accident: **docker restarted inside the guest**
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | a gap well under a minute: 200 before, **404** two seconds after docker returned (traefik had not re-registered routes), serving again by 21:54:19Z — about **40 s** of shut doors |
|
||
| what the box did by itself | everything. Restart ran 21:53:27Z→21:53:41Z; **all 26 containers back at t+16s**; the controller returned with them, waited 51 s for the fleet to settle and found **nothing boot-orphaned to repair** — correct, since every container had already come back on its own policy |
|
||
| time to steady | **16 s** to 26 of 26 containers; **~40 s** until the doors served. The slower number is the one a household feels |
|
||
| alarm fired / true? | **`controller_started` (info) — true and correct**, and the only line the ladder expects: no `app_start_failed` (90 s boot grace), no liveness alarm |
|
||
| should have fired, did not | **none** |
|
||
|
||
**Household loop: NOT SAMPLED** — 0 lines, because the round lasted ~18 s and the loop samples every
|
||
2 minutes. Recorded as not sampled, never as a pass.
|
||
|
||
**A trap avoided.** The round's own alarm snapshot was taken **six seconds** after the controller
|
||
started, and from it `controller_started` looked missing. A later reading shows it present at 21:53.
|
||
An alarm cannot be called missing by a measurement taken before it could have fired.
|
||
|
||
### Round 6 — `backup-system` (whole-guest backup) / accident: **none** (control round)
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | their apps went away and came back: during the backup **4 of 26** containers were up and every app answered **404** publicly; all 26 were serving again by 22:01:13Z. No banner, no mail — correctly |
|
||
| what the box did by itself | quiesced the apps, snapshotted, ran the local tier, **failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z (~8½ min)**, encrypted to ep0, taking **no local disk at all** |
|
||
| time to steady | apps down ~21:55:30Z → 22:01:13Z ≈ **5m43s** — **contaminated by my own intervention** (I killed the local leg at 21:59:45), so it is an upper bound on the quiesce window, not a clean measurement |
|
||
| alarm fired / true? | **`whole_guest_backup_failed` (error) — true, and better than true:** „Whole-guest backup FAILED on the **local tier** — retrying with backoff (next attempt in 15m0s)". It names the tier, not „the backup", and says what it will do next. The status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`). **No alarm for the 22 apps it stopped** — correct, those stops are suppressed |
|
||
| should have fired, did not | **none** |
|
||
|
||
**The finding to carry forward:** a whole-guest backup on this box **cannot use its local tier** — a
|
||
~29 GB source into a 14 GB root filesystem — and the product handles that honestly: it fails the tier,
|
||
says which tier, schedules a retry, and still gets the data out of the house on the off-site tier.
|
||
The local tier will keep retrying and keep failing on a box shaped like this one.
|
||
|
||
**My intervention, and the correction it needed.** I killed the local leg when two samples showed `/`
|
||
falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nested PVE. The arithmetic
|
||
stands, but I first wrote „I stopped the backup", which was wrong: I stopped **one leg**, and the
|
||
off-site leg started by itself seconds later and succeeded. Counted as an intervention either way.
|
||
|
||
### Round 7 — `use` nextcloud / accident: **internet cut for ten minutes**
|
||
|
||
*Drawn as `update`; ran as `use`, because the catalog bump was never pushed — the catalog gates
|
||
returned INCONCLUSIVE and undetermined is never a pass. Decided and recorded before the round, not
|
||
after it.*
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | **depends where they stood.** At home: nothing — traefik answered `301` throughout and every app kept serving. Away from home: ten minutes of nothing — the public route gave **502**, then **200** about a minute after the block lifted |
|
||
| what the box did by itself | kept all **26** containers running, needed no repair, and re-established the way in unaided: 530 three seconds after unblocking, **all four apps 200 by 22:21:52Z — ~64 s** |
|
||
| time to steady | the apps never left steady; only the path in broke and healed, in **~64 s** |
|
||
| alarm fired / true? | **none, and none should have** — the ladder puts `node_stale` at 30 minutes and this was ten. Nothing false was raised either |
|
||
| should have fired, did not | **none** |
|
||
|
||
**Household loop: 10 operations, 0 failures — narrower than it looks.** It does not follow redirects,
|
||
so it measured the box (up throughout) and was blind to the public outage.
|
||
|
||
**The question this round was meant to answer, and honestly did not.** Are alarms raised while the hub
|
||
is unreachable retried and then silently dropped? **Not exercised.** The box reports every **15m0s**
|
||
(measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports — the next was due
|
||
~22:23:43Z, after it lifted. Nothing was attempted, so nothing could be lost. Recorded as *not
|
||
exercised*, never as *passed*.
|
||
|
||
**The fence held, and it was checked against a baseline taken beforehand.** The accident flips a
|
||
host-wide sysctl on demo-hp and inserts two `physdev` rules. Afterwards: `-P FORWARD ACCEPT`, **0**
|
||
physdev rules, sysctl **0** — identical to the pre-round reading. demo-hp also carries guests 9201
|
||
and 9202, so an abandoned rule would have been a fence breach, not an untidy drill.
|
||
|
||
### Round 8 — `backup-app` nextcloud / accident: **internet cut for ten minutes**
|
||
|
||
**22:35:48Z–22:47:59Z.** The app-data backup ran first and finished in 1 m 55 s
|
||
(`db_dump` count 5, `success:true`, 22:37:48Z). The cut began six seconds later, so the two barely
|
||
overlapped — and would not have interacted in any case: `backup-app` is the **local** app-data tier
|
||
and needs no internet. The off-site tier is a different action.
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | **Nothing at home.** Every front door kept serving on the LAN for the whole ten minutes. From outside the house the sites were unreachable — the public path was down. 26 apps up before, 26 after, never fewer. |
|
||
| what the box did by itself | Kept every container running, kept backing up, kept reporting to the hub, and rebuilt the public path unaided when the link returned. No restart, no intervention, noaction from me. |
|
||
| time to steady | **≤43 s** after the link returned (public doors 200 again at 22:48:38Z). Containers never left steady at all. |
|
||
| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold and this was ten. Nothing false was raised. |
|
||
| should have fired, did not | **none** |
|
||
|
||
**And the finding of the round is against my own instrument, not the box.**
|
||
|
||
The hub report due at **22:38:43Z fell inside the cut** — and it **succeeded**:
|
||
„Hub report pushed successfully (15526 bytes)". It succeeded because `hub.felhom.eu` resolves to
|
||
**192.168.0.192**, a LAN address (measured from both the guest and the host), and my injector blocks
|
||
everything **except** the LAN. So the accident named „internet gone" only ever removed the **public**
|
||
path. The box never lost the hub, in round 7 or in round 8.
|
||
|
||
Two consequences, both stated plainly:
|
||
|
||
1. **The dropped-event question is still unmeasured** after two rounds that appeared to measure it.
|
||
Events pushed while the hub is unreachable are retried three times and then dropped permanently,
|
||
with no queue — that behaviour has still never been seen live.
|
||
2. **My own memory file carries this exact warning** („hub.felhom.eu resolves to the LAN here; an
|
||
internet-cut drill must block it too") and I did not apply it. A warning that is written down and
|
||
not read is worth nothing, which is the same class of failure as an unread alarm.
|
||
|
||
The injector is corrected for round 9's drawn internet cut so that the hub address is blocked too —
|
||
**from the VM's side, at the host's tap rule.** The hub itself is never touched; the fence is kept.
|
||
This is a repair to a broken instrument, not a re-draw: the drawn action, app and accident for every
|
||
remaining round are unchanged.
|
||
|
||
**A sixth mistimed reading, and the fix is in the runner now.** The round's own AFTER step read the
|
||
public doors **three seconds** after the unblock and reported 530 on all four. That reading could
|
||
never have been fair. The re-measurement 43 s later read 200 on all four. The runner measured the
|
||
recovery before the recovery could begin — so the runner is the thing that gets fixed, not the note.
|
||
|
||
### Round 9 — `use` uptime-kuma / accident: **internet cut for ten minutes, hub included**
|
||
|
||
**23:00:50Z–23:12:00Z.** The first cut of the night that really removed the hub. The injector was
|
||
corrected between rounds 8 and 9; the drawn action, app and accident were **not** changed.
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | **Nothing at home.** uptime-kuma answered 200 on all three reads during the action, and all four front doors read 200 at both post-accident readings. 26 apps before, 26 after, never fewer. |
|
||
| what the box did by itself | Kept every container running while it was cut off from the hub **and** from its own host agent. Built its report on schedule, tried to push it **three times over 1 m 40.8 s**, gave up, and carried on serving. Both links repaired themselves the instant the block lifted — no restart, no repair action, nothing from me. |
|
||
| time to steady | The apps never left steady. The doors read 200 at the first reading, **2 s** after the unblock, and again 63 s later. |
|
||
| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold. Nothing false was raised. |
|
||
| should have fired, did not | **none** — but see the near-miss below, which is a finding in its own right. |
|
||
|
||
**What a lost report costs: measured, not assumed.**
|
||
|
||
```
|
||
23:08:42 [INFO] [scheduler] Running job: hub-report
|
||
23:08:42 [INFO] [report] Building system report
|
||
23:10:23 [WARN] [report] Push failed: … context deadline exceeded
|
||
23:10:23 [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts (took 1m40.813s)
|
||
```
|
||
|
||
Three attempts, then it stops. Nothing is queued. It gave up **31 seconds before** the link returned.
|
||
And that is correct: a report is a **snapshot**, so a lost one costs nothing — the next snapshot
|
||
carries the same truth, and the controller says so itself („backing off (the 15-min cycle still
|
||
reconciles)"). **This is not the event path.** A dropped event is a lost *fact*, not a stale copy of
|
||
a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop
|
||
behaviour is **still unmeasured** after three rounds of internet cuts.
|
||
|
||
**And it came back by itself, on schedule.** The very next scheduled report went through —
|
||
**23:23:43Z, „Hub report pushed successfully (15354 bytes)”** — exactly 15 minutes after the
|
||
cycle that failed, and nothing was done to the box to achieve it. The reading was taken at
|
||
23:24:21Z, deliberately *after* the report was due, so it could not be premature. So the whole
|
||
shape of a hub outage is now measured end to end: build → three attempts → give up →
|
||
keep serving → next cycle succeeds → no alarm, no loss.
|
||
|
||
**The near-miss, which is luck and not design.** Last good report 22:53:43Z; next scheduled
|
||
23:23:42Z; `node_stale` trips at 30 minutes. The gap is **29 m 59 s**. The staleness threshold is
|
||
exactly twice the report cadence, so **a single failed push spends the entire budget** — one second
|
||
of ordinary jitter and the operator is paged about a box that was healthy throughout and had already
|
||
repaired itself. Filed as a register row.
|
||
|
||
**A second fidelity fault in my accident, named like the first.** The cut also severed the controller
|
||
from its host agent (`GET /backup/tiers` to the link-local `169.254.253.1` timed out twice). In a
|
||
real house an ISP outage does not do that — controller and agent share one machine. So „internet
|
||
gone" as injected is **broader than its name**: internet, hub *and* local agent. No further internet
|
||
cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9.
|
||
|
||
### Round 10 — `restore` uptime-kuma / accident: **hard reset, four seconds into the restore**
|
||
|
||
**23:26:06Z–23:29:46Z.** The roughest pair drawn. The restore was accepted at 23:26:08Z
|
||
(302, „Visszaállítás elindítva"); the reset button was pressed at 23:26:12Z, mid-write.
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | They pressed restore, were told it had started, and **four seconds later the whole machine went dark.** About two minutes of nothing. Then every app was back and every front door answered. **Nothing ever told them what became of the restore.** |
|
||
| what the box did by itself | Booted, and brought **26 of 26 containers** back with no help. Boot reconciliation named the one app it had to recover („1 app(s) recovered in 1 attempt(s): [paperless-ngx]"), sent a startup hub report at 23:28:26Z, and settled its health probes. No intervention. |
|
||
| time to steady | **150 s** — 0 containers at t+12 s, 25 at t+133 s, 26 at t+150 s. Doors 200 at both readings (23:28:42Z and 23:29:44Z). |
|
||
| alarm fired / true? | **one, true** — `controller_started` (info). Exactly what the ladder expects after a reboot. No false alarm. |
|
||
| should have fired, did not | **none from the alarm ladder** — but the restore silence below is a legibility gap, filed as a row. |
|
||
|
||
**The restore left no trace anywhere, and the product has no place to leave one.** Four candidate
|
||
status endpoints all 404 (`/api/restore/status`, `/api/backup/restore/status`,
|
||
`/backup/restore/status`, `/api/restore`). `/api/backup/status` carries no restore field at all.
|
||
On the pages, the only restore text is a **button label** and a JavaScript label expression. On disk,
|
||
in the real data directory, there is no restore, lock or state file anywhere — and **no file at all
|
||
was modified in the reset window**. An interrupted restore and a restore that never happened are
|
||
indistinguishable, to the customer and to me.
|
||
|
||
**CORRECTION, 00:24Z — the paragraph above is wrong and stays visible so the correction is too.**
|
||
The four endpoints I called were four I **guessed**, and all four were wrong. The real route, read
|
||
out of the restore page's own JavaScript, is **`/api/backup/restore-status`**, and it exists:
|
||
`{"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}`. So a restore status
|
||
surface **does** exist. What is true — and is the better finding — is that after the reboot it is
|
||
**blank**: `started_at` is the Go zero value, and the payload carries no `last` field at all, while
|
||
the page's own script renders „<operation> sikertelen." from `st.last.message`. **The restore record
|
||
is in-memory only and does not survive the machine stopping** — precisely the case a hard reset
|
||
creates, and precisely when a household would want to be told. The register row is corrected to say
|
||
that instead. I found the real routes by asking the controller for its own rendered links, which is
|
||
what I should have done before filing anything.
|
||
|
||
**The limit of that measurement, stated rather than glossed.** Only four seconds elapsed, so the
|
||
restore may have finished or may never have written a byte — and I cannot tell, because the
|
||
controller's log stream holds **zero lines before 23:28:00Z** (a reset starts it fresh) and the debug
|
||
ring died with the machine. What is independently verifiable is the **absence of any restore record**,
|
||
and that is what is filed; it holds however far the restore got.
|
||
|
||
**Four of my own instruments failed in this round, and all four are recorded in the evidence:** a
|
||
claim that was unfalsifiable when written; an on-disk check against a directory that does not exist;
|
||
a household count that reported 0 lines and 0 failures when the truth was one line and it *was* a
|
||
failure; and a disk guard that reported „active" all night while being a **transient** unit that
|
||
vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box.
|
||
|
||
### Round 11 — `use` paperless-ngx / accident: **the data drive pulled out for twenty minutes**
|
||
|
||
**23:50:49Z–00:13:09Z.** The drive was detached from the **running** box at 23:50:54Z and put back at
|
||
00:10:58Z. This is the round that produced the most alarms of the night, and every one of them was
|
||
true.
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | **The four apps whose files live on that drive stopped** — Paperless, Jellyfin, Nextcloud, Immich — and Paperless's front door went 404. The other eleven apps kept serving normally throughout. About twenty minutes later everything was back, roughly a minute after the drive was plugged in again. |
|
||
| what the box did by itself | Noticed the drive had gone and **named it by the label the household sees** („Adatlemez"), named **each** broken app individually, degraded its own health, waited, noticed the drive return, restarted the apps and recovered its health. No restart, no repair, nothing from me. |
|
||
| time to steady | **67 s after the drive returned** (26 containers again at 00:12:05Z). The door followed at 00:13:06Z. |
|
||
| alarm fired / true? | **eight, all true, correctly paired at both ends** — `storage_disconnected` (error) → four `app_start_failed` (warning) → `health_degraded` (warning) → `storage_reconnected` (info) → `health_recovered` (info). The four apps named are **exactly** the four with data on the pulled drive. Nothing false was raised. |
|
||
| should have fired, did not | **none** |
|
||
|
||
**The alarm that looked missing, and was not.** The round's own snapshot at 00:13:08Z showed
|
||
`health_degraded` with no recovery — which would have been the first missing alarm of the night. The
|
||
recovery fired at **00:13**, and the snapshot missed it by **seconds**. A re-read at 00:14:13Z, taken
|
||
after the apps were back, found it. This is the discipline from the earlier mistimed readings earning
|
||
its keep: when a measurement could have been early, it is re-taken rather than turned into a verdict.
|
||
|
||
**The drive came back clean** — `/dev/sdd`, 98 G, 2 % used, mounted at `/mnt/felhom-drives/hdd_1`,
|
||
with the guest's mountpoint config unchanged, 26 containers up and every real front door serving.
|
||
|
||
### Round 12 — `use` paperless-ngx / accident: **nothing** (closing control round)
|
||
|
||
**00:15:48Z–00:16:55Z.** The night's last round, drawn as a control.
|
||
|
||
| the five things | |
|
||
|---|---|
|
||
| what the customer saw | Nothing at all. The app answered 200 on all three reads, and every front door answered 200 at both readings. |
|
||
| what the box did by itself | Nothing needed doing. 26 containers before and after. |
|
||
| time to steady | **1 s** — it never left steady. |
|
||
| alarm fired / true? | **none, and none should have.** The newest entry in the feed is still round 11's `health_recovered` at 00:13. |
|
||
| should have fired, did not | **none** |
|
||
|
||
**Checked rather than assumed:** `inject.sh` has no „nothing" case — its default branch exits 2 on an
|
||
unknown accident. The control rounds never reach it, because the runner handles the no-accident case
|
||
itself and says so („accident: none — control round, deliberately"). This matters because *a broken
|
||
injector produces exactly the same result as a control round*, and the only way to tell them apart is
|
||
to look at which code path ran.
|
||
|
||
**What the closing control round is worth.** It shows the quiet is real: after eleven rounds of power
|
||
cuts, resets, full disks, severed networks and a drive pulled out of a running machine, a round in
|
||
which nothing was done produced nothing — no alarm, no restart, no drift. The alarm feed is not
|
||
simply noisy.
|
||
|
||
## Phase 2 — the morning after
|
||
|
||
**Every app answered through its own front door.** All twelve real names, measured at 00:19:16Z —
|
||
`cloud`, `inventory`, `media`, `paperless`, `paste`, `photos`, `recipes`, `share`, `status`,
|
||
`travel`, `vault`, `wiki` — **200 on the public path, every one**. And healthy is not inferred from
|
||
a door: all **26 containers** report `Up … (healthy)`, except `traefik` and `cloudflared`, which
|
||
carry no healthcheck and show a bare `Up`. The four showing seven minutes are the drive-backed apps
|
||
restarted after round 11.
|
||
|
||
**The off-site restore could not be done, and two independent instruments agree why.** The brief asked
|
||
for one DB-backed app restored from off-site onto scratch 9202. It has nothing to restore from:
|
||
|
||
| instrument | answer |
|
||
|---|---|
|
||
| `restic`, with the box's own key, password file and `known_hosts` | `Fatal: wrong password or no key found` — exit status **1**, read from restic itself rather than from the end of a pipeline |
|
||
| the product's own status surface, which is what a customer sees | `{"orphaned":true, … ,"snapshots":0,"status":"error"}` |
|
||
|
||
The repository is orphaned **because this box is a rebuild for an existing customer**: on a rebuild
|
||
the restic password is minted fresh, so snapshots written under the old one can never be opened
|
||
again. That is a known, documented shape, and **the product surfaced it honestly** — the true
|
||
`offbox_repo_orphaned` alarm of round 1, mailed to the operator within two seconds of the run.
|
||
|
||
**What did leave the house.** The *whole-guest* off-site copy is a different store and it worked: ep0
|
||
holds two intact snapshots for this box, including tonight's at **21:59:54Z** — round 6's off-site
|
||
leg, the one that started by itself after I killed the local leg. Its file index is roughly four
|
||
times the afternoon copy's, consistent with a guest by then carrying twelve apps. **Stated as a
|
||
limit:** that is a *listing*, not a verification. A PBS verify job would prove restorability and it
|
||
**writes** verify state, so it was not run — ep0 is read-only for evidence tonight.
|
||
|
||
**The household loop.** 204 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of
|
||
which **only two are real events** — `wiki` during round 2's power cut and `cloud` during round 10's
|
||
hard reset, each a single sample, each healed before the next probe. The other three were **my own
|
||
classifier** counting a 301 redirect as a dashboard failure; the log carries that correction in its
|
||
own words at 21:11:25Z and the wrong lines were left in place so the correction stays visible. Ten of
|
||
the twelve rounds left **no mark at all** in this log, including the twenty minutes with the drive
|
||
pulled — the loop samples each name every two minutes, so that silence is the instrument's sampling
|
||
rate and **not** evidence the household saw nothing.
|
||
|
||
**The catalog bump: verified reverted, not remembered.** Clean tree, local exactly level with
|
||
`origin/main` (0 ahead, 0 behind), the nextcloud template still on its original `redis:7-alpine` pin
|
||
and `catalog_since: "2026-07-18"`, newest commit 2026-09-15. The bump was prepared, **refused by the
|
||
catalog's own gates** (image-resolvable and volume-persistence both INCONCLUSIVE — its own canary
|
||
failed, so the verdict was UNDETERMINED, never a pass) and reverted before any push.
|
||
|
||
**And one delivery proof the truth table could not give.** The mailbox shows every alarm **arrived**
|
||
at the operator, not merely that it was stored — round 11's whole set, round 6's tier-naming backup
|
||
failure, round 1's orphan warning, the OOM, the deploy failure and my own thin-pool `storage_fill_critical`.
|
||
In a project where 91 events once sat in a database having e-mailed nobody, *raised* and *delivered*
|
||
are two different claims, and only one of them had evidence before tonight.
|
||
|
||
## Interventions — counted, with the reason for each verdict
|
||
|
||
**One.** At 21:59:45Z in round 6 I killed the **local leg** of the whole-guest backup. The arithmetic
|
||
that forced it: a ~29 GB source being written into a 14 GB root filesystem at ~16 MB/s, i.e. under
|
||
four minutes to a full `/` on the nested host, mid-round. What it cost: a leg that could never have
|
||
succeeded. What happened next without me: **the off-site leg started by itself from the same snapshot
|
||
and succeeded in about eight and a half minutes.** Filed as **R-548**.
|
||
|
||
My first note on it said „I stopped the backup". That was wrong and is corrected in the evidence: I
|
||
stopped **a leg** of it, and the box completed the other one unaided.
|
||
|
||
**Both pre-declared presses went unused.** O1, the „Send self-bind link" button, was not needed —
|
||
the automatic mail was already waiting (18:17:46Z) and the box bound with **zero** operator presses.
|
||
O2, „Re-issue PBS credentials", was not needed either — the acknowledged-delete path re-issued them
|
||
by itself (`pbsdr_auto_reissue`, 20:19Z). **Both prompt claims they were insurance against turned out
|
||
to be true**, and the F-14 path was measured live for the first time.
|
||
|
||
**Counted separately, because it is not a round result:** Phase 0's seeding repairs. I filled the LVM
|
||
thin pool to 100 % by firing twelve deploys at once, then repaired the damage — a guest restart to
|
||
clear an `emergency_ro` remount, dropping corrupt image layers, a remove-with-data and one consistent
|
||
re-deploy after a re-seed minted fresh database passwords over initialised volumes, and a rewritten
|
||
`APP_KEY`. **The damage was mine, not the product's**, the product's behaviour throughout was correct,
|
||
and every repair went through the product's own endpoints rather than by hand-running compose. Listed
|
||
in full in `evidence-chaos-night-2026-09-17/interventions.txt` so the distinction is visible rather
|
||
than convenient.
|
||
|
||
**Not counted, and why:** acts on my **own instruments** — moving the dashboard password after a
|
||
power cut cleared `/tmp`, rewriting the injector and the runner, re-creating the disk guard as a real
|
||
unit. Counting those would flatter the night in one direction and pad the stop-rule count in the
|
||
other. **Round 4 is the clearest case of declining to intervene:** the tunnel was left dead on
|
||
purpose — „NOT restarting it by hand — whether it returns by itself IS the measurement". It
|
||
returned by itself in 97 s.
|
||
|
||
**Standing against the stop rule: 1 of 4.** The night ran its full twelve rounds.
|
||
|
||
## Teardown — three layers, stated
|
||
|
||
**Machine — gone.** VM 336 stopped and destroyed with all three disks purged (00:35:25Z), gated on
|
||
its **name** rather than its number because two standing guests share the host. `qm list` shows no
|
||
VMs; `/mnt/hdd_1/images/336` no longer exists. The storage was verified to *be* `/mnt/hdd_1` from its
|
||
own definition (`nvme-scratch`, `path /mnt/hdd_1`, `is_mountpoint yes`) rather than assumed. The
|
||
machine had **three** disks, not the two the brief asked for — the third was mine, added in Phase 0
|
||
after I filled the thin pool — and that is recorded rather than quietly removed.
|
||
|
||
**Host — clean, measured before and after.** `nvme-scratch` 6.78 % → **1.61 %** (~48.5 GB returned);
|
||
`local-lvm` **unchanged at 44.75 %**, so the fence that said *never local-lvm* held; free space on
|
||
`/mnt/hdd_1` 827 G → 875 G. Guests 9201 and 9202 still running. The household loop and the disk guard
|
||
were stopped and disabled **after** their logs were copied off (the guard's log was 0 bytes — it
|
||
never fired). Firewall back to `-P FORWARD ACCEPT` with **0** physdev rules, so none of the three
|
||
network accidents left a rule behind. Scratch 9202: **nothing to remove**, shown rather than said —
|
||
three infrastructure containers, no app deployed, no stack touched since 21:00, no off-site config.
|
||
|
||
**Hub — host record deleted through the acknowledged flow.** The first attempt at 00:38:37Z was
|
||
**correctly refused** (409, „Host is ONLINE") — the box had died inside the hub's liveness window. The
|
||
acknowledged delete went through at **07:25:13Z** (`confirm_host_id` + `delete_escrow=1` → 303). Every
|
||
line of the after-state written down *before* the act matched: the host answers 404; `drill-r50`,
|
||
both demo hosts and the `tester-1` **customer** still answer 200; the customer now lists zero hosts.
|
||
**The automatic connect mail arrived two seconds later** (07:25:15Z, „Kösd össze a Felhom dobozodat"),
|
||
quoted in full with its token redacted in `teardown-hub.txt` — and it is provably tonight's, because
|
||
the mailbox held no such mail newer than 18:17:46Z when checked at 00:38Z.
|
||
|
||
**Why the hub layer finished six hours late — my fault, not the product's.** The retry was guarded
|
||
by „don't post while the host page contains ONLINE". That word lives in a JavaScript string that is
|
||
**always** on the page, so the guard could never pass. It refused six times and gave up at 01:20Z,
|
||
while the hub's structured answer would have said `"status":"down"` from about 00:54Z. Nothing ran
|
||
again until 07:24Z.
|
||
|
||
**ep0 — backups stayed, nothing removed.** Read three times: 00:17:15Z, 00:36:52Z (just before the
|
||
delete) and 07:25:35Z (after). Identical every time — three namespaces, two snapshots each, six in
|
||
total, 16 G used. No prune, no verify, no write.
|
||
|
||
## Claims in the prompt that turned out wrong — named first, as asked
|
||
|
||
**The two the brief itself flagged both turned out TRUE, and both were checked tonight rather than
|
||
assumed.**
|
||
|
||
1. **„The automatic mail is waiting in the mailbox."** The brief warned this had been read from
|
||
*yesterday's* host delete and not verified. **It was true.** The self-bind mail of **18:17:46Z**
|
||
was in the mailbox, and the box bound with **zero operator presses** — so the pre-declared press
|
||
O1 was never needed.
|
||
2. **„The WG hook provisions by itself after an acknowledged delete" (the F-14 path).** The brief
|
||
noted this had never been measured live. **It was true, and it was measured live for the first
|
||
time:** `pbsdr_auto_reissue` at **20:19Z** — „Previous key destroyed (acknowledged deletion) —
|
||
credentials re-issued automatically." The second pre-declared press, O2, was never needed either.
|
||
|
||
**Now the ones that were wrong.**
|
||
|
||
3. **WRONG: „restore one DB-backed app from off-site onto scratch 9202."** It could not be done at
|
||
all on this box, and not because anything broke. This box is a **rebuild for an existing
|
||
customer**, so its restic password was minted fresh and the snapshots already in the remote store
|
||
can never be opened by it again. Two independent instruments agree: `restic` itself
|
||
(`Fatal: wrong password or no key found`, exit 1) and the product's own status
|
||
(`orphaned:true, snapshots:0, status:"error"`). The brief assumed an off-site app repository this
|
||
box could open; on a rebuild fixture there is none.
|
||
4. **WRONG in effect: „a system disk + one data disk."** The machine ended the night with **three**
|
||
disks. The third, 64 G, was added by me in Phase 0 to extend the LVM thin pool after I filled it
|
||
to 100 % by firing twelve deploys at once. **The deviation is mine, not the brief's**, but the
|
||
fixture was not the one the brief described and saying so is the point.
|
||
5. **WRONG: round 7's drawn action `update` was not performed as drawn.** The catalog's own gates
|
||
returned `image-resolvable INCONCLUSIVE` and `volume-persistence INCONCLUSIVE` — its **own canary
|
||
failed**, so the verdict was UNDETERMINED, which is never a pass. The round ran `use` instead. A
|
||
deviation from the drawn schedule, logged rather than quietly substituted.
|
||
6. **WRONG, and mine rather than the brief's: „an internet cut tests what happens when the hub is
|
||
unreachable."** The accident's *name* implies it; on this network it was false. `hub.felhom.eu`
|
||
resolves to a **LAN** address here, and my injector allowed the whole LAN — so rounds 7 and 8 cut
|
||
the public path only, and the box never lost the hub. My own memory file carries that exact
|
||
warning and I did not apply it. Fixed between rounds 8 and 9 by blocking the hub address **from
|
||
the VM's side**, which is what finally made round 9 the measurement it was supposed to be.
|
||
7. **WRONG as a description of the night's clock: the schedule table's times.** The table drawn from
|
||
the seed lists rounds at 23:30 through 04:05. Those were **nominal**. The real spacing was 25
|
||
minutes from each round's actual start, and the night's twelve rounds finished at **00:17Z**,
|
||
roughly four hours earlier than the table's own column suggests. Each round's real timestamps are
|
||
recorded in its own section; **the drawn order, apps and accidents were never changed** — only
|
||
the wall-clock the table guessed at.
|
||
|
||
**And one the brief did not make, which the night could not answer.** Events pushed while the hub is
|
||
unreachable are retried three times and then dropped permanently, with no queue. Three ten-minute
|
||
hub outages happened and **no event was raised during any of them**, so that path is still
|
||
unmeasured. What *was* measured is the **report** path: built, three attempts over 1 m 40.8 s, given
|
||
up, and the next scheduled report succeeded — and a report is a snapshot, so nothing was lost.
|