chaos night round 10: a restore leaves no record, and four of my instruments failed
gates / gates (push) Successful in 21s

The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.

R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.

Four instrument faults, all mine, all in the evidence:
  * a 'nothing was logged' claim that was unfalsifiable when written - the log
    stream holds zero lines before a reset;
  * an on-disk check against /opt/felhom/data, a directory that does not exist;
  * a household count reporting 0 lines and 0 failures when the truth was one
    line and it WAS a failure - the runner now prints both operands;
  * the disk guard was a TRANSIENT unit reporting 'active' all night, and was
    absent from the reset onward. It is now file-backed and enabled, and its
    script is copied off the box for the first time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 01:35:56 +02:00
parent f973fd7151
commit 9f40dc3289
5 changed files with 145 additions and 0 deletions
+1
View File
@@ -732,6 +732,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** MEASURED 2026-09-17 (chaos night) on a fresh box (`tester-1-022354`, controller 0.245.0): the customer guest’s root filesystem was held at **96 % for ten minutes** (29 G used, 1.5 G free) and **no alarm of any kind fired** — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: `fillwatch` runs **daily at 03:30** plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about **twenty seconds before** the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. **This is the ladder working as designed, not a missed alarm** — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is **no, unless the controller happens to restart while it is full**, and that answer is not written down anywhere. **Fix shape (one of):** sample the fill more often than daily (a cheap `statfs` on the 5-minute health pass would do it); or say plainly in `08-alarm-ladder.md` that a transient full disk is out of scope. Evidence: `audits/evidence-chaos-night-2026-09-17/round-3.txt`. | **READY — rank P3-LOW; owner: CC** |
| **R-548** | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. | **READY — rank P3-LOW; owner: CC** |
| **R-549** | **[P2-MEDIUM] The staleness alarm's budget is exactly two report cycles, so ONE failed push spends all of it.** MEASURED 2026-09-17 (chaos night, round 9) on `tester-1-022354` (controller 0.245.0): hub reachability was removed for ten minutes from the VM's side. The controller built its 23:08:42Z report, retried the push **three times over 1 m 40.8 s**, and gave up at 23:10:23Z (`hub push failed after 3 attempts`) - **31 seconds before the link returned**. Nothing is queued, and that is correct: a report is a snapshot, so the next one carries the same truth, and the controller says so itself ("backing off (the 15-min cycle still reconciles)"). **The arithmetic is the finding.** Last good report 22:53:43Z; next scheduled 23:23:42Z; measured cadence **15m0s**; `node_stale` trips at **30 minutes**. The gap is **29 m 59 s** - one second inside the threshold. So a single missed push spends the whole staleness budget, and any ordinary jitter pages the operator about a box that is healthy, serving every app, and has already repaired itself unaided. **The product behaved correctly throughout:** no false alarm fired, no app stopped, and both the hub link and the host-agent link recovered by themselves the moment the block lifted. What is filed is the margin, not a misbehaviour. **Fix shape:** either set the staleness threshold to a clear multiple of the cadence (three cycles, not two), or let a push that has failed all three attempts retry once off-cycle instead of waiting for the next scheduled report. **Honest caveat:** the cut was injected by the drill and also severed the controller from its host agent, which a real ISP outage would not do - but the report arithmetic above depends only on the hub being unreachable. Evidence: `audits/evidence-chaos-night-2026-09-17/round-9.txt`. | **READY - rank P2-MEDIUM; owner: CC** |
| **R-550** | **[P2-MEDIUM] A restore leaves no record anywhere, so an interrupted one and one that never happened look identical to the customer.** MEASURED 2026-09-17 (chaos night, round 10) on `tester-1-022354` (controller 0.245.0): an app restore was accepted at 23:26:08Z (`302`, flash "Visszaallitas elindult") and the guest's host was hard-reset **four seconds later**, mid-write. Afterwards, asked through the doors the UI itself uses: `/api/restore/status`, `/api/backup/restore/status`, `/backup/restore/status` and `/api/restore` **all 404**; `/api/backup/status` returns `{enabled,running}` with **no restore field at all**; on `/backups/apps` and `/apps/<name>` the only restore text is a **button label** and a JS label expression (ASCII fragment search, with a negative control returning 0 on every page). On disk, in the real data directory (`/var/lib/docker/volumes/felhom-controller-data/_data`), there is no restore, lock or state file anywhere beneath it, and **no file at all was modified in the reset window**. **The box recovered perfectly** - 26/26 containers back in 150 s, boot reconciliation naming the app it recovered, every front door serving, one true `controller_started` alarm and no false one. What is filed is that the customer pressed a button, was told it had started, and can never learn whether it finished. **Honest limit:** only four seconds elapsed, so the restore may have completed or may never have written a byte - the pre-reset log is unrecoverable (the stream holds zero lines before the reboot) and the debug ring is in memory. The absence of any restore record is verifiable independently of how far it got, and that is the row. **Fix shape:** persist a restore record (started, finished or abandoned-at-boot) the way the backup tiers already persist theirs, and show it on the app's page; boot reconciliation is the natural place to mark an in-flight restore as abandoned. Evidence: `audits/evidence-chaos-night-2026-09-17/round-10.txt`. | **READY - rank P2-MEDIUM; owner: CC** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |