Files
felhom.eu/documentation/audits/DRILL-chaos-night-2026-09-17.md
T
admin 69c08b183b
gates / gates (push) Successful in 22s
chaos night teardown complete: hub host record deleted, connect mail quoted, ep0 unchanged
The acknowledged delete went through at 07:25:13Z (confirm_host_id +
delete_escrow=1 -> 303). Every line of the after-state written down before the
act matched: the host 404s; drill-r50, both demo hosts and the tester-1
customer still 200; the customer lists zero hosts; ep0 identical across three
readings - six snapshots, 16G, nothing removed.

The automatic connect mail arrived two seconds later and is quoted with its
token redacted. It is provably tonight's: the mailbox held no such mail newer
than 18:17:46Z when checked at 00:38Z.

Why the hub layer finished six hours late is mine: the retry guard refused to
post while the host page contained the word ONLINE, and that word sits in a
JavaScript string that is always on the page. The hub's structured status said
'down' from about 00:54Z. It gave up at 01:20Z and nothing ran again until
07:24Z. Eleventh instrument fault of the night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 09:27:10 +02:00

714 lines
54 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DRILL — CHAOS NIGHT: random actions on random apps while random things go wrong (2026-09-16/17)
**Interventions: 1.** One, at 21:59:45Z in round 6: I killed the **local leg** of a whole-guest
backup that could never have fit (a ~29 GB source into a 14 GB root filesystem, falling at ~16 MB/s).
The **off-site leg then ran by itself from the same snapshot and succeeded**, so the data still left
the house. Both pre-declared presses went **unused**: the automatic self-bind mail was already
waiting, and the acknowledged-delete path re-issued the PBS credentials on its own. Phase 0's
seeding repairs are listed separately in `evidence-chaos-night-2026-09-17/interventions.txt` — that
damage was mine, not the product's, and every repair went through the product's own endpoints.
**Ready for a volunteer: still yes.** Across twelve rounds — a power cut mid-restore, a hard reset
four seconds into another, a full disk, a killed tunnel, a restarted Docker, three severed networks
and the data drive pulled out of a running machine for twenty minutes — **nothing cost a byte of
customer data, and the box healed itself every single time with no human involved.** Seventeen
alarms fired, **all seventeen were true, none were missing**, and the mailbox proves every one was
**delivered** rather than merely stored. The honest qualifications: two P2 legibility gaps are filed
(an interrupted restore leaves no record; one failed report spends the whole 30-minute staleness
budget, measured at 29 m 59 s), and one thing this night could **not** test — per-app off-site
restore, because this box is a rebuild whose restic repository is orphaned **by design**, which the
product surfaced honestly within seconds.
**The accident-plus-action pair that hurt most: `restore` + hard reset (round 10).** Not because the
box suffered — it was back with 26 of 26 containers in **150 s**, boot reconciliation naming the app
it recovered, every front door serving. It hurt most because it is the **only** pair of the night
where the household is left not knowing what happened: they pressed restore, were told it had
started, the machine went dark four seconds later, and afterwards **nothing anywhere tells them
whether it finished.** The status surface exists and answers with the zero value; the record is
in-memory only and does not survive the machine stopping.
> **Baselines, verified live against Gitea at 21:49 CEST 2026-09-16 (not copied from the brief):**
> felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 ·
> felhom.eu `d124c77e176d` hub v0.116.0, ISO **1.28.0 published** · app-catalog `94bc5febaca2`.
> All four trees clean and in sync. Highest register row **R-545**, 212 open. Golden waiver valid to
> 2026-09-27. Customer **`tester-1`** (`enkicsifelhom.hu`, `tester1@felhom.eu`, no host).
> Venue: `demo-hp` (Tier 0), a fresh nested VM, disk on the NVMe at its root. Evidence:
> `evidence-chaos-night-2026-09-17/`.
## The schedule — drawn ONCE, before round 1, and written here first
The point of this section's position in the document is that the night could not be chosen after the
fact. `chaos_schedule.py` is committed beside the evidence; re-running it reproduces this table.
- **seed:** `20260917` (the date)
- **script sha256:** `4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1`
- **generator:** `evidence-chaos-night-2026-09-17/chaos_schedule.py`, stdlib `random` seeded with the seed
| # | time | X — the action | Y — the app | Z — the accident |
|---|---|---|---|---|
| 1 | 23:30 | offsite-run | adventurelog | nothing |
| 2 | 23:55 | restore | gokapi | power cut |
| 3 | 00:20 | use | bookstack | disk 95% full |
| 4 | 00:45 | offsite-run | mealie | tunnel down 10min |
| 5 | 01:10 | use | privatebin | docker restarted |
| 6 | 01:35 | backup-system | adventurelog | nothing |
| 7 | 02:00 | update | nextcloud | internet gone 10min |
| 8 | 02:25 | backup-app | nextcloud | internet gone 10min |
| 9 | 02:50 | use | uptime-kuma | internet gone 10min |
| 10 | 03:15 | restore | uptime-kuma | hard reset |
| 11 | 03:40 | use | paperless-ngx | drive pulled 20min |
| 12 | 04:05 | use | paperless-ngx | nothing |
**Re-draw log** — a silent re-draw is a schedule chosen by the person running it, so every one is here:
- r02 X=reinstall re-drawn (nothing has been removed yet)
- r06 Z=disk 95% full re-drawn (constraint 4: at most once)
- r07 Z=nothing re-drawn (constraint 6: never two in a row after r2)
- r08 X=reinstall re-drawn (nothing has been removed yet)
**What the draw happened to give, said plainly before the night judges it:** no `remove` round was
ever drawn, so `reinstall` had nothing to reinstall and was re-drawn twice (rounds 2 and 8). Three
`internet gone` rounds land consecutively (7, 8, 9) — that is the seed's doing, and it makes rounds
7–9 a de-facto endurance test of the same accident against three different actions rather than three
independent samples. `controller killed`, `drive pulled 90s`, `memory pressure`, `agent restarted`
and `hub unreachable` were never drawn at all; **this night does not test them**, and the morning
verdict must not claim it did.
## Phase 0 — the golden, the box, the household
**0.1 Golden 0.245.0, baked and published.** Launched 19:52:47Z as a transient unit in the drill VM,
finished 19:58:15Z. Markers: `overlay2`=1, `including mount point`=2, `upload OK (HTTP 201)`=1,
FATAL=0, publish-skipped=0. `GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626`.
**The teardown was gated on the REGISTRY answering 200**, not on an exit code — and that mattered:
the wrapper exited **144** while every measured outcome was good. Vouched in the hub and read back
from the page (`golden currently vouched: 0.245.0`). The three bake failures of 2026-09-16 (`scp -P`,
`chmod 0700`, `GITEA_USER=admin`) were each guarded and none recurred.
*Decision, recorded because silence reads as agreement:* the global controller floor was left at
0.244.0. The new box installs golden 0.245.0, which already carries controller 0.245.0, so no floor
was needed to deliver anything tonight; raising it would have pushed an update onto demo-felhom, a
box not in this drill.
**0.2 The box.** VM 336 on demo-hp: 8 GiB, 4 cores, 32 G system + 100 G data disk on the NVMe at its
root, booted from the **published** ISO 1.28.0. Boot order set in its own `qm set` (combining it
silently yields `order=net0;ide2`). Install completion was judged **from the disk** — blocks used
grew 3233 → 6942 MiB then held across three checks — because „Automatically reboot" is ticked and a
finished install looks exactly like a stuck one on screen. The summary page was read before pressing
Install, and the line that made it safe was **„Disk(s): /dev/sda"** — the 32 G system disk alone.
**The walk, as a volunteer, cost ZERO operator presses.** The box registered itself as an unclaimed
appliance and polled, visibly, until bound. The bind link came from the **waiting mail** (minted
18:17:46Z by yesterday's acknowledged host delete), the pairing code off the box's own console
(`4SY-4TX`), the „Tulajdonosi jelmondat" from the hub's customer record:
POST /bind/<token> -> 200, „Sikeres összekötés."
Then day-0 ran on its own and the hub recorded, without anyone pressing anything:
`appliance_bound` (customer_selfbind) · `appliance_credential_delivered` · `claim_reissued_reenroll`
· `offsite_reissued` · **`pbsdr_auto_reissue` — „Previous key destroyed (acknowledged deletion) —
credentials re-issued automatically."**
**That last event is a first.** The brief named „the WG hook provisions by itself after an
acknowledged delete" as a claim never measured live. It ran tonight, unprompted. **Both pre-declared
presses (O1 self-bind, O2 re-issue) were therefore unnecessary.**
The dashboard was claimed with the mailed code and **proven by logging in with the new password** —
a claim page that re-renders looks identical to success from the status code alone. The 100 GB data
drive was initialised through the wizard's own endpoint (`POST /api/storage/init`, polled to
`phase: done`), and `df` shows it mounted at `/mnt/felhom-drives/hdd_1` with 93 G free.
**The box landed on tonight's golden with no hand upgrade:** controller **0.245.0** (healthy), agent
**0.131.0**, host `tester-1-022354` ONLINE. And the **R-543 escrow reminder bar shipped hours earlier
was live on it**, on a box nobody had touched.
**0.3 The household — and the first real trouble.** Twelve deploys were fired; **ten were accepted,
two refused** for memory with both numbers quoted. Then nine of the ten failed: the guest's disks are
thin-provisioned over an ~11.8 GB pool carved from a 32 GB system disk, ten simultaneous image pulls
filled it, and the hub recorded `storage_fill_critical` (100 %) plus **nine `app_deploy_failed`
warnings**, one per app, each naming the failing pull. Only PrivateBin installed.
**The product behaved; the harness did not.** R-536's failure event — shipped that same morning so an
interrupted install is not silence — fired for all nine within two minutes. The memory guard refused
rather than over-committing. The per-stack record stayed honest (`deployed: false`). The two faults
were mine: firing twelve deploys in two seconds is not household behaviour, and a 32 GB system disk
was copied from an earlier drill without checking what that drill had installed.
**0.4 The escrow ceremony could NOT be completed — and this one is about tonight's own release.**
See „Finding: the recovery-code step cannot be done when the guide says to do it" below.
## Finding: the recovery-code step cannot be done when the guide says to do it (R-546)
The box was at exactly the point of the guide this release added hours earlier — installed, bound
with no press, claimed, drive initialised, **no apps yet** — and the escrow reminder bar was on every
page telling the household to create their recovery code. It could not be done.
POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"}
GET /api/escrow/status -> claimable:false —
detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"
POST /api/escrow/claim -> **409** „A folyamat jelenlegi állapotában a kód nem kérhető le."
**Both sides agreed on the cause.** The hub's own Backup & DR panel read „host enrolled **done** · WG
tunnel peer registered **done** · descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1)
**waiting** · ceremony possible once the descriptor is applied on the box". The box had no PBS storage
(`pvesm status`: `local`, `local-lvm` only) and no `escrow` section in `agent.json` at all.
**It self-heals, and that was measured rather than assumed.** The box was left alone and polled:
20:30:12Z … 20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
**20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs**
~17 minutes after the bind. The retried ceremony passed every preflight item, the claim returned
**200** (83-character code, 129.2 bits of entropy, revealed once), and `escrow_state` became
**escrowed**. The bar then vanished from all four pages checked — the R-543 fix working through its
whole lifecycle on a box nobody had set up for the test.
**So the defect is timing and wording, not mechanism.** For ~17 minutes a volunteer following
tonight's guide meets a stderr fragment about a `-storage` flag, while every page urges them on.
Filed **R-546** (P2). No product code was changed — this is a validation run.
## Phase 1 — the rounds
Rounds run at ~25-minute spacing. The schedule above is fixed; only the wall-clock start moved,
because Phase 0 ran long (the storage wizard submits by JavaScript and the endpoint was worked out
rather than guessed). Round 1 began 23:07 CEST.
### Round 1 — `offsite-run` / adventurelog / accident: **nothing** (control round)
| the five things | |
|---|---|
| what the customer saw | „A távoli mentés elindult — az állapot itt frissül." and, at the end, „A távoli mentési tároló elárvult: a benne lévő mentések egy korábbi, már nem elérhető kulccsal készültek (újratelepítés)." |
| what the box did by itself | walked all twelve apps — stop, dump each volume with real byte counts, restart — captured eleven, could not capture the one that was crash-looping, and finished |
| time to steady | **1m45s** (`last_duration`), `last_run` 21:09:00Z, `progress.active` false, `last_error` empty. A control round: the box never left steady |
| alarm fired / true? | **three, all true** — `app_start_failed` named Nextcloud · `backup_run_failures` „1 of 12 apps failed to back up in this nightly run: nextcloud" · `offbox_repo_orphaned`, matching the status endpoint's own `"orphaned": true` |
| should have fired, did not | **none** |
**Household loop in the window:** 3 lines marked FAILED, **all three mine** — the loop counted the
dashboard's 301 redirect as a failure while accepting the same 301 for app reads. Fixed at 21:11:25Z
and marked in the log; only lines after that marker are scored.
**What round 1 actually establishes.** The off-site tier is armed (escrowed) and the run works
end-to-end, but on THIS box — a rebuild for an existing customer — the remote repository was written
under a key the box no longer holds, so **no snapshot was written**. That is the documented rebuild
behaviour, surfaced honestly with the route out named in the message rather than reported as success.
**Two things that looked like defects in this round and are not**, both established with controls
rather than inference — five front doors answering 404 (traefik has no route to an unhealthy
container; identical byte-for-byte to a no-such-host control) and a crash-looping Nextcloud (image
layers corrupted while my thin pool stood at 100 %; „invalid ELF header", repaired by a re-pull).
Detail in `evidence-chaos-night-2026-09-17/round-1-notes.txt`.
### Phase 0 postscript — repairing my own damage, and three conclusions I had to retract
Nine of the first ten deploys failed because I fired twelve at once onto a thin pool carved from a
32 GB disk, and the pool hit 100 %. Repairing that took the rest of Phase 0 and produced **two
distinct faults of mine, with different cures**, which only separating them made fixable:
| fault | symptom | cure |
|---|---|---|
| image layers written while the pool was full | `php: … libxml2.so.2: **invalid ELF header**`, exit 127 crash loop | drop the image, let compose pull it again |
| my re-seed generated **fresh database passwords** over volumes already initialised with the first set | Postgres `auth_failed`, MariaDB „Access denied for user … (using password: YES)" | remove the app **with its data**, deploy once with one consistent secret set |
A fresh image did not fix gokapi and a fresh database did — that is the evidence the two faults are
different things rather than one.
**Three conclusions I wrote and then had to retract, each corrected where it stood:**
1. „the five 404s were my mistimed sweep" — wrong for four of them. A **negative control** (a
no-such-host request) returned the identical 404 of 19 bytes, and a positive control returned
200/1200 bytes: traefik simply has **no route to an unhealthy container**.
2. „nextcloud is repaired" — wrong. The re-pull fixed the crash, and the app still could not reach its
database. The container reported **`healthy`** throughout, because the image's own healthcheck asks
whether Apache answers, not whether the application works.
3. „all the broken apps are corrupt layers" — wrong. bookstack logged a clean startup, gokapi logged
nothing at all, immich showed a Postgres auth failure.
I also nearly filed a defect against the drive gate, which was working and logging at DEBUG while I
read a settings snapshot inside its 30-second tick. **An absent log line is not evidence.**
None of this is a product defect and none of it is filed as one. What the product did throughout was
correct and legible: it refused to route to unhealthy containers, `app_start_failed` named the app
that was down, `backup_run_failures` said „1 of 12 … nextcloud", the memory guard refused the
eleventh and twelfth installs with both numbers quoted, and `storage_fill_critical` fired at 100 %.
### Round 2 — `restore` gokapi / accident: **power cut**, 20 s into the restore
| the five things | |
|---|---|
| what the customer saw | „Visszaállítás elindult — az állapot itt frissül." then the box went dark mid-restore; on return the dashboard and every app were back |
| what the box did by itself | everything — containers 0 → **25 at t+131s** → **26 at t+148s**, nothing stuck, no shell used. **gokapi, the app being restored when the plug came out, returned `Up 30 seconds (healthy)`** |
| time to steady | **148 s**, measured against the container count this round took itself before the accident |
| alarm fired / true? | `controller_started` (info) — true and correct. **No false alarm.** |
| should have fired, did not | **none** — per the ladder a 60-second outage yields no `node_stale` (30 min threshold) and no `app_start_failed` (90 s boot grace), and neither appeared |
**Household loop: NOT COLLECTED.** The loop was a transient unit on the VM and died with the power
cut — the first accident that could have produced household failures instead produced no lines at
all. Zero lines is not zero failures, so it is recorded as not collected, and the loop is now a
persistent systemd unit that returns with the box.
**A finding this round handed over:** `app_oom` (warning) — „Alkalmazás memóriája elfogyott: immich
(immich-postgres) — egy folyamatát a memóriakorlát leállította". That is immich's whole mystery
solved: its Postgres was OOM-killed during the reverse-geocoding import, which is why the server saw
`CONNECTION_CLOSED` and crash-looped twelve times. **The controller caught an OOM inside an LXC guest
and named the exact container** — worth recording against this project's standing finding that those
signals are usually invisible there. The diagnosis I spent twenty minutes reaching from logs was in
the alarm feed, correctly labelled, the whole time.
### Round 3 — `use` bookstack / accident: **system disk filled to 96 % for ten minutes**
| the five things | |
|---|---|
| what the customer saw | **nothing.** wiki, status and paste all answered before, during and after. No banner, no warning, no mail — the household was never told the disk was full |
| what the box did by itself | kept all twelve apps running on a 96 %-full root filesystem and released the space cleanly when the fill was removed (29 G used → 944 M used). The shared thin pool never moved (**39.69 %**) and the filesystem stayed writable |
| time to steady | the box never left steady — **26 containers before, 26 after**, none restarted |
| alarm fired / true? | **none fired**, checked twice independently after the fill was released |
| should have fired, did not | `disk_critical` is defined at ≥95 % used and the disk sat at **96 % for ten minutes**. **But this is the ladder working as designed, not a miss:** the fill-watch is a daily sweep (03:30) plus one check ~90 s after a controller start. Predicted before the round, confirmed after |
**Household loop: 12 operations, 0 failures.** The household kept using its apps normally throughout.
**The finding is the silence.** The honest answer to „would the household be told their disk is
full?" is **no** — unless the controller happens to restart while it is full. Here the timing was
almost comic: the controller restarted at 21:28 after round 2's power cut, so its one opportunistic
check ran about twenty seconds *before* the disk filled, and the next is not due until 03:30.
**A correction, recorded where it happened:** I twice labelled a mid-window reading „end of window",
estimating the clock instead of reading it. The readings were unchanged, but „nothing yet, five
minutes in" and „nothing in the whole window" are different findings. From here the end-of-window
check is taken when the round's own runner reports completion.
### Round 4 — `offsite-run` (mealie) / accident: **tunnel killed for ten minutes**
**The drawn ACTION never ran.** The runner aborted with „no session — mine, not the product's":
round 2's power cut had rebooted the guest, `/tmp` is cleared on boot, and the dashboard password
file lived there. Round 3 was a `use` round and never needed it, so round 4 was the first to find it
gone. **The accident was measured; the off-site run was not.** Recorded as half-measured rather than
re-run and presented as whole — re-running a round after seeing it fail is how a drill starts
choosing its own results. The file now lives in `/root`, which survives a reboot.
| the five things | |
|---|---|
| what the customer saw | from outside, the apps vanished for ~90 s (public route **530**) and came back on their own (**200**); from inside the house, nothing — traefik answered **301** throughout |
| what the box did by itself | **repaired its own tunnel.** cloudflared killed 21:41:33Z, running again **21:43:07.478Z (~97 s)**, with `RestartCount=0` — so Docker's `unless-stopped` policy did *not* do it; the controller's protected-infra recovery redeployed it („[infra] deploying cloudflared →…") |
| time to steady | the apps never stopped; the way IN was restored in **~97 s** |
| alarm fired / true? | **two, correctly paired** — `health_critical` (error) 21:43, `health_recovered` (info) 21:48. Exactly what the ladder predicts for a missing protected container, and **the alarm was not a dead end** |
| should have fired, did not | none for the accident |
**Household loop: 10 operations, 0 failures — and that number is narrower than it looks.** The loop
does not follow redirects, so it measures „is the app serving on the box", never „can the household
reach it from outside". It saw nothing while the public route was returning 530.
### Round 5 — `use` privatebin / accident: **docker restarted inside the guest**
| the five things | |
|---|---|
| what the customer saw | a gap well under a minute: 200 before, **404** two seconds after docker returned (traefik had not re-registered routes), serving again by 21:54:19Z — about **40 s** of shut doors |
| what the box did by itself | everything. Restart ran 21:53:27Z→21:53:41Z; **all 26 containers back at t+16s**; the controller returned with them, waited 51 s for the fleet to settle and found **nothing boot-orphaned to repair** — correct, since every container had already come back on its own policy |
| time to steady | **16 s** to 26 of 26 containers; **~40 s** until the doors served. The slower number is the one a household feels |
| alarm fired / true? | **`controller_started` (info) — true and correct**, and the only line the ladder expects: no `app_start_failed` (90 s boot grace), no liveness alarm |
| should have fired, did not | **none** |
**Household loop: NOT SAMPLED** — 0 lines, because the round lasted ~18 s and the loop samples every
2 minutes. Recorded as not sampled, never as a pass.
**A trap avoided.** The round's own alarm snapshot was taken **six seconds** after the controller
started, and from it `controller_started` looked missing. A later reading shows it present at 21:53.
An alarm cannot be called missing by a measurement taken before it could have fired.
### Round 6 — `backup-system` (whole-guest backup) / accident: **none** (control round)
| the five things | |
|---|---|
| what the customer saw | their apps went away and came back: during the backup **4 of 26** containers were up and every app answered **404** publicly; all 26 were serving again by 22:01:13Z. No banner, no mail — correctly |
| what the box did by itself | quiesced the apps, snapshotted, ran the local tier, **failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z (~8½ min)**, encrypted to ep0, taking **no local disk at all** |
| time to steady | apps down ~21:55:30Z → 22:01:13Z ≈ **5m43s** — **contaminated by my own intervention** (I killed the local leg at 21:59:45), so it is an upper bound on the quiesce window, not a clean measurement |
| alarm fired / true? | **`whole_guest_backup_failed` (error) — true, and better than true:** „Whole-guest backup FAILED on the **local tier** — retrying with backoff (next attempt in 15m0s)". It names the tier, not „the backup", and says what it will do next. The status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`). **No alarm for the 22 apps it stopped** — correct, those stops are suppressed |
| should have fired, did not | **none** |
**The finding to carry forward:** a whole-guest backup on this box **cannot use its local tier** — a
~29 GB source into a 14 GB root filesystem — and the product handles that honestly: it fails the tier,
says which tier, schedules a retry, and still gets the data out of the house on the off-site tier.
The local tier will keep retrying and keep failing on a box shaped like this one.
**My intervention, and the correction it needed.** I killed the local leg when two samples showed `/`
falling at ~16 MB/s with 3.6 GB left — under four minutes from wedging the nested PVE. The arithmetic
stands, but I first wrote „I stopped the backup", which was wrong: I stopped **one leg**, and the
off-site leg started by itself seconds later and succeeded. Counted as an intervention either way.
### Round 7 — `use` nextcloud / accident: **internet cut for ten minutes**
*Drawn as `update`; ran as `use`, because the catalog bump was never pushed — the catalog gates
returned INCONCLUSIVE and undetermined is never a pass. Decided and recorded before the round, not
after it.*
| the five things | |
|---|---|
| what the customer saw | **depends where they stood.** At home: nothing — traefik answered `301` throughout and every app kept serving. Away from home: ten minutes of nothing — the public route gave **502**, then **200** about a minute after the block lifted |
| what the box did by itself | kept all **26** containers running, needed no repair, and re-established the way in unaided: 530 three seconds after unblocking, **all four apps 200 by 22:21:52Z — ~64 s** |
| time to steady | the apps never left steady; only the path in broke and healed, in **~64 s** |
| alarm fired / true? | **none, and none should have** — the ladder puts `node_stale` at 30 minutes and this was ten. Nothing false was raised either |
| should have fired, did not | **none** |
**Household loop: 10 operations, 0 failures — narrower than it looks.** It does not follow redirects,
so it measured the box (up throughout) and was blind to the public outage.
**The question this round was meant to answer, and honestly did not.** Are alarms raised while the hub
is unreachable retried and then silently dropped? **Not exercised.** The box reports every **15m0s**
(measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports — the next was due
~22:23:43Z, after it lifted. Nothing was attempted, so nothing could be lost. Recorded as *not
exercised*, never as *passed*.
**The fence held, and it was checked against a baseline taken beforehand.** The accident flips a
host-wide sysctl on demo-hp and inserts two `physdev` rules. Afterwards: `-P FORWARD ACCEPT`, **0**
physdev rules, sysctl **0** — identical to the pre-round reading. demo-hp also carries guests 9201
and 9202, so an abandoned rule would have been a fence breach, not an untidy drill.
### Round 8 — `backup-app` nextcloud / accident: **internet cut for ten minutes**
**22:35:48Z–22:47:59Z.** The app-data backup ran first and finished in 1 m 55 s
(`db_dump` count 5, `success:true`, 22:37:48Z). The cut began six seconds later, so the two barely
overlapped — and would not have interacted in any case: `backup-app` is the **local** app-data tier
and needs no internet. The off-site tier is a different action.
| the five things | |
|---|---|
| what the customer saw | **Nothing at home.** Every front door kept serving on the LAN for the whole ten minutes. From outside the house the sites were unreachable — the public path was down. 26 apps up before, 26 after, never fewer. |
| what the box did by itself | Kept every container running, kept backing up, kept reporting to the hub, and rebuilt the public path unaided when the link returned. No restart, no intervention, noaction from me. |
| time to steady | **≤43 s** after the link returned (public doors 200 again at 22:48:38Z). Containers never left steady at all. |
| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold and this was ten. Nothing false was raised. |
| should have fired, did not | **none** |
**And the finding of the round is against my own instrument, not the box.**
The hub report due at **22:38:43Z fell inside the cut** — and it **succeeded**:
„Hub report pushed successfully (15526 bytes)". It succeeded because `hub.felhom.eu` resolves to
**192.168.0.192**, a LAN address (measured from both the guest and the host), and my injector blocks
everything **except** the LAN. So the accident named „internet gone" only ever removed the **public**
path. The box never lost the hub, in round 7 or in round 8.
Two consequences, both stated plainly:
1. **The dropped-event question is still unmeasured** after two rounds that appeared to measure it.
Events pushed while the hub is unreachable are retried three times and then dropped permanently,
with no queue — that behaviour has still never been seen live.
2. **My own memory file carries this exact warning** („hub.felhom.eu resolves to the LAN here; an
internet-cut drill must block it too") and I did not apply it. A warning that is written down and
not read is worth nothing, which is the same class of failure as an unread alarm.
The injector is corrected for round 9's drawn internet cut so that the hub address is blocked too —
**from the VM's side, at the host's tap rule.** The hub itself is never touched; the fence is kept.
This is a repair to a broken instrument, not a re-draw: the drawn action, app and accident for every
remaining round are unchanged.
**A sixth mistimed reading, and the fix is in the runner now.** The round's own AFTER step read the
public doors **three seconds** after the unblock and reported 530 on all four. That reading could
never have been fair. The re-measurement 43 s later read 200 on all four. The runner measured the
recovery before the recovery could begin — so the runner is the thing that gets fixed, not the note.
### Round 9 — `use` uptime-kuma / accident: **internet cut for ten minutes, hub included**
**23:00:50Z–23:12:00Z.** The first cut of the night that really removed the hub. The injector was
corrected between rounds 8 and 9; the drawn action, app and accident were **not** changed.
| the five things | |
|---|---|
| what the customer saw | **Nothing at home.** uptime-kuma answered 200 on all three reads during the action, and all four front doors read 200 at both post-accident readings. 26 apps before, 26 after, never fewer. |
| what the box did by itself | Kept every container running while it was cut off from the hub **and** from its own host agent. Built its report on schedule, tried to push it **three times over 1 m 40.8 s**, gave up, and carried on serving. Both links repaired themselves the instant the block lifted — no restart, no repair action, nothing from me. |
| time to steady | The apps never left steady. The doors read 200 at the first reading, **2 s** after the unblock, and again 63 s later. |
| alarm fired / true? | **none fired, and none should have** — `node_stale` is a 30-minute threshold. Nothing false was raised. |
| should have fired, did not | **none** — but see the near-miss below, which is a finding in its own right. |
**What a lost report costs: measured, not assumed.**
```
23:08:42 [INFO] [scheduler] Running job: hub-report
23:08:42 [INFO] [report] Building system report
23:10:23 [WARN] [report] Push failed: … context deadline exceeded
23:10:23 [ERROR] [scheduler] Job hub-report failed: hub push failed after 3 attempts (took 1m40.813s)
```
Three attempts, then it stops. Nothing is queued. It gave up **31 seconds before** the link returned.
And that is correct: a report is a **snapshot**, so a lost one costs nothing — the next snapshot
carries the same truth, and the controller says so itself („backing off (the 15-min cycle still
reconciles)"). **This is not the event path.** A dropped event is a lost *fact*, not a stale copy of
a picture that will be redrawn. No event happened to be raised during the cut, so the event-drop
behaviour is **still unmeasured** after three rounds of internet cuts.
**And it came back by itself, on schedule.** The very next scheduled report went through —
**23:23:43Z, „Hub report pushed successfully (15354 bytes)”** — exactly 15 minutes after the
cycle that failed, and nothing was done to the box to achieve it. The reading was taken at
23:24:21Z, deliberately *after* the report was due, so it could not be premature. So the whole
shape of a hub outage is now measured end to end: build → three attempts → give up →
keep serving → next cycle succeeds → no alarm, no loss.
**The near-miss, which is luck and not design.** Last good report 22:53:43Z; next scheduled
23:23:42Z; `node_stale` trips at 30 minutes. The gap is **29 m 59 s**. The staleness threshold is
exactly twice the report cadence, so **a single failed push spends the entire budget** — one second
of ordinary jitter and the operator is paged about a box that was healthy throughout and had already
repaired itself. Filed as a register row.
**A second fidelity fault in my accident, named like the first.** The cut also severed the controller
from its host agent (`GET /backup/tiers` to the link-local `169.254.253.1` timed out twice). In a
real house an ISP outage does not do that — controller and agent share one machine. So „internet
gone" as injected is **broader than its name**: internet, hub *and* local agent. No further internet
cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9.
### Round 10 — `restore` uptime-kuma / accident: **hard reset, four seconds into the restore**
**23:26:06Z–23:29:46Z.** The roughest pair drawn. The restore was accepted at 23:26:08Z
(302, „Visszaállítás elindítva"); the reset button was pressed at 23:26:12Z, mid-write.
| the five things | |
|---|---|
| what the customer saw | They pressed restore, were told it had started, and **four seconds later the whole machine went dark.** About two minutes of nothing. Then every app was back and every front door answered. **Nothing ever told them what became of the restore.** |
| what the box did by itself | Booted, and brought **26 of 26 containers** back with no help. Boot reconciliation named the one app it had to recover („1 app(s) recovered in 1 attempt(s): [paperless-ngx]"), sent a startup hub report at 23:28:26Z, and settled its health probes. No intervention. |
| time to steady | **150 s** — 0 containers at t+12 s, 25 at t+133 s, 26 at t+150 s. Doors 200 at both readings (23:28:42Z and 23:29:44Z). |
| alarm fired / true? | **one, true** — `controller_started` (info). Exactly what the ladder expects after a reboot. No false alarm. |
| should have fired, did not | **none from the alarm ladder** — but the restore silence below is a legibility gap, filed as a row. |
**The restore left no trace anywhere, and the product has no place to leave one.** Four candidate
status endpoints all 404 (`/api/restore/status`, `/api/backup/restore/status`,
`/backup/restore/status`, `/api/restore`). `/api/backup/status` carries no restore field at all.
On the pages, the only restore text is a **button label** and a JavaScript label expression. On disk,
in the real data directory, there is no restore, lock or state file anywhere — and **no file at all
was modified in the reset window**. An interrupted restore and a restore that never happened are
indistinguishable, to the customer and to me.
**CORRECTION, 00:24Z — the paragraph above is wrong and stays visible so the correction is too.**
The four endpoints I called were four I **guessed**, and all four were wrong. The real route, read
out of the restore page's own JavaScript, is **`/api/backup/restore-status`**, and it exists:
`{"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}`. So a restore status
surface **does** exist. What is true — and is the better finding — is that after the reboot it is
**blank**: `started_at` is the Go zero value, and the payload carries no `last` field at all, while
the page's own script renders „<operation> sikertelen." from `st.last.message`. **The restore record
is in-memory only and does not survive the machine stopping** — precisely the case a hard reset
creates, and precisely when a household would want to be told. The register row is corrected to say
that instead. I found the real routes by asking the controller for its own rendered links, which is
what I should have done before filing anything.
**The limit of that measurement, stated rather than glossed.** Only four seconds elapsed, so the
restore may have finished or may never have written a byte — and I cannot tell, because the
controller's log stream holds **zero lines before 23:28:00Z** (a reset starts it fresh) and the debug
ring died with the machine. What is independently verifiable is the **absence of any restore record**,
and that is what is filed; it holds however far the restore got.
**Four of my own instruments failed in this round, and all four are recorded in the evidence:** a
claim that was unfalsifiable when written; an on-disk check against a directory that does not exist;
a household count that reported 0 lines and 0 failures when the truth was one line and it *was* a
failure; and a disk guard that reported „active" all night while being a **transient** unit that
vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box.
### Round 11 — `use` paperless-ngx / accident: **the data drive pulled out for twenty minutes**
**23:50:49Z–00:13:09Z.** The drive was detached from the **running** box at 23:50:54Z and put back at
00:10:58Z. This is the round that produced the most alarms of the night, and every one of them was
true.
| the five things | |
|---|---|
| what the customer saw | **The four apps whose files live on that drive stopped** — Paperless, Jellyfin, Nextcloud, Immich — and Paperless's front door went 404. The other eleven apps kept serving normally throughout. About twenty minutes later everything was back, roughly a minute after the drive was plugged in again. |
| what the box did by itself | Noticed the drive had gone and **named it by the label the household sees** („Adatlemez"), named **each** broken app individually, degraded its own health, waited, noticed the drive return, restarted the apps and recovered its health. No restart, no repair, nothing from me. |
| time to steady | **67 s after the drive returned** (26 containers again at 00:12:05Z). The door followed at 00:13:06Z. |
| alarm fired / true? | **eight, all true, correctly paired at both ends** — `storage_disconnected` (error) → four `app_start_failed` (warning) → `health_degraded` (warning) → `storage_reconnected` (info) → `health_recovered` (info). The four apps named are **exactly** the four with data on the pulled drive. Nothing false was raised. |
| should have fired, did not | **none** |
**The alarm that looked missing, and was not.** The round's own snapshot at 00:13:08Z showed
`health_degraded` with no recovery — which would have been the first missing alarm of the night. The
recovery fired at **00:13**, and the snapshot missed it by **seconds**. A re-read at 00:14:13Z, taken
after the apps were back, found it. This is the discipline from the earlier mistimed readings earning
its keep: when a measurement could have been early, it is re-taken rather than turned into a verdict.
**The drive came back clean** — `/dev/sdd`, 98 G, 2 % used, mounted at `/mnt/felhom-drives/hdd_1`,
with the guest's mountpoint config unchanged, 26 containers up and every real front door serving.
### Round 12 — `use` paperless-ngx / accident: **nothing** (closing control round)
**00:15:48Z–00:16:55Z.** The night's last round, drawn as a control.
| the five things | |
|---|---|
| what the customer saw | Nothing at all. The app answered 200 on all three reads, and every front door answered 200 at both readings. |
| what the box did by itself | Nothing needed doing. 26 containers before and after. |
| time to steady | **1 s** — it never left steady. |
| alarm fired / true? | **none, and none should have.** The newest entry in the feed is still round 11's `health_recovered` at 00:13. |
| should have fired, did not | **none** |
**Checked rather than assumed:** `inject.sh` has no „nothing" case — its default branch exits 2 on an
unknown accident. The control rounds never reach it, because the runner handles the no-accident case
itself and says so („accident: none — control round, deliberately"). This matters because *a broken
injector produces exactly the same result as a control round*, and the only way to tell them apart is
to look at which code path ran.
**What the closing control round is worth.** It shows the quiet is real: after eleven rounds of power
cuts, resets, full disks, severed networks and a drive pulled out of a running machine, a round in
which nothing was done produced nothing — no alarm, no restart, no drift. The alarm feed is not
simply noisy.
## Phase 2 — the morning after
**Every app answered through its own front door.** All twelve real names, measured at 00:19:16Z —
`cloud`, `inventory`, `media`, `paperless`, `paste`, `photos`, `recipes`, `share`, `status`,
`travel`, `vault`, `wiki` — **200 on the public path, every one**. And healthy is not inferred from
a door: all **26 containers** report `Up … (healthy)`, except `traefik` and `cloudflared`, which
carry no healthcheck and show a bare `Up`. The four showing seven minutes are the drive-backed apps
restarted after round 11.
**The off-site restore could not be done, and two independent instruments agree why.** The brief asked
for one DB-backed app restored from off-site onto scratch 9202. It has nothing to restore from:
| instrument | answer |
|---|---|
| `restic`, with the box's own key, password file and `known_hosts` | `Fatal: wrong password or no key found` — exit status **1**, read from restic itself rather than from the end of a pipeline |
| the product's own status surface, which is what a customer sees | `{"orphaned":true, … ,"snapshots":0,"status":"error"}` |
The repository is orphaned **because this box is a rebuild for an existing customer**: on a rebuild
the restic password is minted fresh, so snapshots written under the old one can never be opened
again. That is a known, documented shape, and **the product surfaced it honestly** — the true
`offbox_repo_orphaned` alarm of round 1, mailed to the operator within two seconds of the run.
**What did leave the house.** The *whole-guest* off-site copy is a different store and it worked: ep0
holds two intact snapshots for this box, including tonight's at **21:59:54Z** — round 6's off-site
leg, the one that started by itself after I killed the local leg. Its file index is roughly four
times the afternoon copy's, consistent with a guest by then carrying twelve apps. **Stated as a
limit:** that is a *listing*, not a verification. A PBS verify job would prove restorability and it
**writes** verify state, so it was not run — ep0 is read-only for evidence tonight.
**The household loop.** 204 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of
which **only two are real events** — `wiki` during round 2's power cut and `cloud` during round 10's
hard reset, each a single sample, each healed before the next probe. The other three were **my own
classifier** counting a 301 redirect as a dashboard failure; the log carries that correction in its
own words at 21:11:25Z and the wrong lines were left in place so the correction stays visible. Ten of
the twelve rounds left **no mark at all** in this log, including the twenty minutes with the drive
pulled — the loop samples each name every two minutes, so that silence is the instrument's sampling
rate and **not** evidence the household saw nothing.
**The catalog bump: verified reverted, not remembered.** Clean tree, local exactly level with
`origin/main` (0 ahead, 0 behind), the nextcloud template still on its original `redis:7-alpine` pin
and `catalog_since: "2026-07-18"`, newest commit 2026-09-15. The bump was prepared, **refused by the
catalog's own gates** (image-resolvable and volume-persistence both INCONCLUSIVE — its own canary
failed, so the verdict was UNDETERMINED, never a pass) and reverted before any push.
**And one delivery proof the truth table could not give.** The mailbox shows every alarm **arrived**
at the operator, not merely that it was stored — round 11's whole set, round 6's tier-naming backup
failure, round 1's orphan warning, the OOM, the deploy failure and my own thin-pool `storage_fill_critical`.
In a project where 91 events once sat in a database having e-mailed nobody, *raised* and *delivered*
are two different claims, and only one of them had evidence before tonight.
## Interventions — counted, with the reason for each verdict
**One.** At 21:59:45Z in round 6 I killed the **local leg** of the whole-guest backup. The arithmetic
that forced it: a ~29 GB source being written into a 14 GB root filesystem at ~16 MB/s, i.e. under
four minutes to a full `/` on the nested host, mid-round. What it cost: a leg that could never have
succeeded. What happened next without me: **the off-site leg started by itself from the same snapshot
and succeeded in about eight and a half minutes.** Filed as **R-548**.
My first note on it said „I stopped the backup". That was wrong and is corrected in the evidence: I
stopped **a leg** of it, and the box completed the other one unaided.
**Both pre-declared presses went unused.** O1, the „Send self-bind link" button, was not needed —
the automatic mail was already waiting (18:17:46Z) and the box bound with **zero** operator presses.
O2, „Re-issue PBS credentials", was not needed either — the acknowledged-delete path re-issued them
by itself (`pbsdr_auto_reissue`, 20:19Z). **Both prompt claims they were insurance against turned out
to be true**, and the F-14 path was measured live for the first time.
**Counted separately, because it is not a round result:** Phase 0's seeding repairs. I filled the LVM
thin pool to 100 % by firing twelve deploys at once, then repaired the damage — a guest restart to
clear an `emergency_ro` remount, dropping corrupt image layers, a remove-with-data and one consistent
re-deploy after a re-seed minted fresh database passwords over initialised volumes, and a rewritten
`APP_KEY`. **The damage was mine, not the product's**, the product's behaviour throughout was correct,
and every repair went through the product's own endpoints rather than by hand-running compose. Listed
in full in `evidence-chaos-night-2026-09-17/interventions.txt` so the distinction is visible rather
than convenient.
**Not counted, and why:** acts on my **own instruments** — moving the dashboard password after a
power cut cleared `/tmp`, rewriting the injector and the runner, re-creating the disk guard as a real
unit. Counting those would flatter the night in one direction and pad the stop-rule count in the
other. **Round 4 is the clearest case of declining to intervene:** the tunnel was left dead on
purpose — „NOT restarting it by hand — whether it returns by itself IS the measurement". It
returned by itself in 97 s.
**Standing against the stop rule: 1 of 4.** The night ran its full twelve rounds.
## Teardown — three layers, stated
**Machine — gone.** VM 336 stopped and destroyed with all three disks purged (00:35:25Z), gated on
its **name** rather than its number because two standing guests share the host. `qm list` shows no
VMs; `/mnt/hdd_1/images/336` no longer exists. The storage was verified to *be* `/mnt/hdd_1` from its
own definition (`nvme-scratch`, `path /mnt/hdd_1`, `is_mountpoint yes`) rather than assumed. The
machine had **three** disks, not the two the brief asked for — the third was mine, added in Phase 0
after I filled the thin pool — and that is recorded rather than quietly removed.
**Host — clean, measured before and after.** `nvme-scratch` 6.78 % → **1.61 %** (~48.5 GB returned);
`local-lvm` **unchanged at 44.75 %**, so the fence that said *never local-lvm* held; free space on
`/mnt/hdd_1` 827 G → 875 G. Guests 9201 and 9202 still running. The household loop and the disk guard
were stopped and disabled **after** their logs were copied off (the guard's log was 0 bytes — it
never fired). Firewall back to `-P FORWARD ACCEPT` with **0** physdev rules, so none of the three
network accidents left a rule behind. Scratch 9202: **nothing to remove**, shown rather than said —
three infrastructure containers, no app deployed, no stack touched since 21:00, no off-site config.
**Hub — host record deleted through the acknowledged flow.** The first attempt at 00:38:37Z was
**correctly refused** (409, „Host is ONLINE") — the box had died inside the hub's liveness window. The
acknowledged delete went through at **07:25:13Z** (`confirm_host_id` + `delete_escrow=1` → 303). Every
line of the after-state written down *before* the act matched: the host answers 404; `drill-r50`,
both demo hosts and the `tester-1` **customer** still answer 200; the customer now lists zero hosts.
**The automatic connect mail arrived two seconds later** (07:25:15Z, „Kösd össze a Felhom dobozodat"),
quoted in full with its token redacted in `teardown-hub.txt` — and it is provably tonight's, because
the mailbox held no such mail newer than 18:17:46Z when checked at 00:38Z.
**Why the hub layer finished six hours late — my fault, not the product's.** The retry was guarded
by „don't post while the host page contains ONLINE". That word lives in a JavaScript string that is
**always** on the page, so the guard could never pass. It refused six times and gave up at 01:20Z,
while the hub's structured answer would have said `"status":"down"` from about 00:54Z. Nothing ran
again until 07:24Z.
**ep0 — backups stayed, nothing removed.** Read three times: 00:17:15Z, 00:36:52Z (just before the
delete) and 07:25:35Z (after). Identical every time — three namespaces, two snapshots each, six in
total, 16 G used. No prune, no verify, no write.
## Claims in the prompt that turned out wrong — named first, as asked
**The two the brief itself flagged both turned out TRUE, and both were checked tonight rather than
assumed.**
1. **„The automatic mail is waiting in the mailbox."** The brief warned this had been read from
*yesterday's* host delete and not verified. **It was true.** The self-bind mail of **18:17:46Z**
was in the mailbox, and the box bound with **zero operator presses** — so the pre-declared press
O1 was never needed.
2. **„The WG hook provisions by itself after an acknowledged delete" (the F-14 path).** The brief
noted this had never been measured live. **It was true, and it was measured live for the first
time:** `pbsdr_auto_reissue` at **20:19Z** — „Previous key destroyed (acknowledged deletion) —
credentials re-issued automatically." The second pre-declared press, O2, was never needed either.
**Now the ones that were wrong.**
3. **WRONG: „restore one DB-backed app from off-site onto scratch 9202."** It could not be done at
all on this box, and not because anything broke. This box is a **rebuild for an existing
customer**, so its restic password was minted fresh and the snapshots already in the remote store
can never be opened by it again. Two independent instruments agree: `restic` itself
(`Fatal: wrong password or no key found`, exit 1) and the product's own status
(`orphaned:true, snapshots:0, status:"error"`). The brief assumed an off-site app repository this
box could open; on a rebuild fixture there is none.
4. **WRONG in effect: „a system disk + one data disk."** The machine ended the night with **three**
disks. The third, 64 G, was added by me in Phase 0 to extend the LVM thin pool after I filled it
to 100 % by firing twelve deploys at once. **The deviation is mine, not the brief's**, but the
fixture was not the one the brief described and saying so is the point.
5. **WRONG: round 7's drawn action `update` was not performed as drawn.** The catalog's own gates
returned `image-resolvable INCONCLUSIVE` and `volume-persistence INCONCLUSIVE` — its **own canary
failed**, so the verdict was UNDETERMINED, which is never a pass. The round ran `use` instead. A
deviation from the drawn schedule, logged rather than quietly substituted.
6. **WRONG, and mine rather than the brief's: „an internet cut tests what happens when the hub is
unreachable."** The accident's *name* implies it; on this network it was false. `hub.felhom.eu`
resolves to a **LAN** address here, and my injector allowed the whole LAN — so rounds 7 and 8 cut
the public path only, and the box never lost the hub. My own memory file carries that exact
warning and I did not apply it. Fixed between rounds 8 and 9 by blocking the hub address **from
the VM's side**, which is what finally made round 9 the measurement it was supposed to be.
7. **WRONG as a description of the night's clock: the schedule table's times.** The table drawn from
the seed lists rounds at 23:30 through 04:05. Those were **nominal**. The real spacing was 25
minutes from each round's actual start, and the night's twelve rounds finished at **00:17Z**,
roughly four hours earlier than the table's own column suggests. Each round's real timestamps are
recorded in its own section; **the drawn order, apps and accidents were never changed** — only
the wall-clock the table guessed at.
**And one the brief did not make, which the night could not answer.** Events pushed while the hub is
unreachable are retried three times and then dropped permanently, with no queue. Three ten-minute
hub outages happened and **no event was raised during any of them**, so that path is still
unmeasured. What *was* measured is the **report** path: built, three attempts over 1 m 40.8 s, given
up, and the next scheduled report succeeded — and a report is a snapshot, so nothing was lost.