chaos night: R-550 corrected - I guessed four endpoints and all four were wrong
gates / gates (push) Successful in 22s
gates / gates (push) Successful in 22s
The row first claimed a restore leaves no record anywhere, citing four status endpoints that 404'd. All four were paths I guessed. The real route, read out of the restore page's own JavaScript, is /api/backup/restore-status and it exists. The corrected finding is narrower and better: the endpoint answers with the Go zero value (started_at 0001-01-01T00:00:00Z) and carries no 'last' field at all, while the page's own script renders '<operation> sikertelen.' from st.last.message. The restore record is in-memory only and does not survive the machine stopping - exactly the case a hard reset creates. The original wording is left visible in the audit with the correction beside it; the register row is corrected in place because a register must be accurate. The reusable lesson: I found the real routes by asking the controller for its own rendered links. Guessing produced four confident 404s that I then reported as a property of the product. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -448,6 +448,18 @@ in the real data directory, there is no restore, lock or state file anywhere —
|
||||
was modified in the reset window**. An interrupted restore and a restore that never happened are
|
||||
indistinguishable, to the customer and to me.
|
||||
|
||||
**CORRECTION, 00:24Z — the paragraph above is wrong and stays visible so the correction is too.**
|
||||
The four endpoints I called were four I **guessed**, and all four were wrong. The real route, read
|
||||
out of the restore page's own JavaScript, is **`/api/backup/restore-status`**, and it exists:
|
||||
`{"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}`. So a restore status
|
||||
surface **does** exist. What is true — and is the better finding — is that after the reboot it is
|
||||
**blank**: `started_at` is the Go zero value, and the payload carries no `last` field at all, while
|
||||
the page's own script renders „<operation> sikertelen." from `st.last.message`. **The restore record
|
||||
is in-memory only and does not survive the machine stopping** — precisely the case a hard reset
|
||||
creates, and precisely when a household would want to be told. The register row is corrected to say
|
||||
that instead. I found the real routes by asking the controller for its own rendered links, which is
|
||||
what I should have done before filing anything.
|
||||
|
||||
**The limit of that measurement, stated rather than glossed.** Only four seconds elapsed, so the
|
||||
restore may have finished or may never have written a byte — and I cannot tell, because the
|
||||
controller's log stream holds **zero lines before 23:28:00Z** (a reset starts it fresh) and the debug
|
||||
|
||||
@@ -85,3 +85,33 @@ that holds regardless of how far the restore got.
|
||||
vanishes on reboot. It was therefore ABSENT from 23:26:12Z. It is now a file-backed, ENABLED unit
|
||||
(verified: active + enabled, MainPID cmdline "/bin/bash /root/diskguard.sh", 0 kills, 7556 MB free)
|
||||
and its script is copied off the box into this folder, which had also never been done (R-320).
|
||||
|
||||
## CORRECTION TO THIS ROUND'S CENTRAL CLAIM, 2026-09-17T00:24Z — I WAS WRONG
|
||||
Above I wrote that four candidate status endpoints all 404 and that "there is no restore history
|
||||
surface of any kind". That was built on FOUR PATHS I GUESSED, and all four were wrong. The real
|
||||
route, taken from the restore page's own JavaScript rather than from my imagination, is:
|
||||
/api/backup/restore-status
|
||||
It exists, it answers, and this is what it returns now:
|
||||
{"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}
|
||||
|
||||
So the corrected finding is NARROWER and better than the one I filed:
|
||||
* a restore status surface DOES exist;
|
||||
* after the reboot it is EMPTY - `started_at` is the Go zero value 0001-01-01T00:00:00Z;
|
||||
* the payload carries no `last` field at all, yet the page's own script reads `st.last.op` and
|
||||
`st.last.message` to render "<operation> sikertelen." So there is a "last operation" branch in
|
||||
the UI with nothing to populate it after a restart.
|
||||
The record is therefore IN-MEMORY ONLY and does not survive the machine stopping - which is exactly
|
||||
the case a hard reset creates, and exactly when a customer would most want to know.
|
||||
|
||||
### The honest limit is unchanged, and now cuts both ways
|
||||
Only four seconds elapsed. `started_at` may be zero because the restore never really began, not
|
||||
because the reboot erased it. I cannot separate those, because the pre-reset log is unrecoverable.
|
||||
What IS certain: the surface exists, it is blank now, and nothing anywhere tells the customer that a
|
||||
restore they started did not finish.
|
||||
|
||||
### How the error happened, because that is the reusable part
|
||||
I searched for the page at `/apps/uptime-kuma` and guessed API paths by pattern. The controller's
|
||||
actual routes are `/backups`, `/backups/remote`, `/backups/restore` and `/stacks/<name>/backup`.
|
||||
I found them in the end by asking the controller for its OWN rendered links instead of guessing -
|
||||
which is what I should have done first. Guessing produced four confident 404s that I then reported
|
||||
as a property of the product.
|
||||
|
||||
@@ -732,7 +732,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** MEASURED 2026-09-17 (chaos night) on a fresh box (`tester-1-022354`, controller 0.245.0): the customer guest’s root filesystem was held at **96 % for ten minutes** (29 G used, 1.5 G free) and **no alarm of any kind fired** — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: `fillwatch` runs **daily at 03:30** plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about **twenty seconds before** the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. **This is the ladder working as designed, not a missed alarm** — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is **no, unless the controller happens to restart while it is full**, and that answer is not written down anywhere. **Fix shape (one of):** sample the fill more often than daily (a cheap `statfs` on the 5-minute health pass would do it); or say plainly in `08-alarm-ladder.md` that a transient full disk is out of scope. Evidence: `audits/evidence-chaos-night-2026-09-17/round-3.txt`. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-548** | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-549** | **[P2-MEDIUM] The staleness alarm's budget is exactly two report cycles, so ONE failed push spends all of it.** MEASURED 2026-09-17 (chaos night, round 9) on `tester-1-022354` (controller 0.245.0): hub reachability was removed for ten minutes from the VM's side. The controller built its 23:08:42Z report, retried the push **three times over 1 m 40.8 s**, and gave up at 23:10:23Z (`hub push failed after 3 attempts`) - **31 seconds before the link returned**. Nothing is queued, and that is correct: a report is a snapshot, so the next one carries the same truth, and the controller says so itself ("backing off (the 15-min cycle still reconciles)"). **The arithmetic is the finding.** Last good report 22:53:43Z; next scheduled 23:23:42Z; measured cadence **15m0s**; `node_stale` trips at **30 minutes**. The gap is **29 m 59 s** - one second inside the threshold. So a single missed push spends the whole staleness budget, and any ordinary jitter pages the operator about a box that is healthy, serving every app, and has already repaired itself unaided. **The product behaved correctly throughout:** no false alarm fired, no app stopped, and both the hub link and the host-agent link recovered by themselves the moment the block lifted. What is filed is the margin, not a misbehaviour. **Fix shape:** either set the staleness threshold to a clear multiple of the cadence (three cycles, not two), or let a push that has failed all three attempts retry once off-cycle instead of waiting for the next scheduled report. **Honest caveat:** the cut was injected by the drill and also severed the controller from its host agent, which a real ISP outage would not do - but the report arithmetic above depends only on the hub being unreachable. Evidence: `audits/evidence-chaos-night-2026-09-17/round-9.txt`. | **READY - rank P2-MEDIUM; owner: CC** |
|
||||
| **R-550** | **[P2-MEDIUM] A restore leaves no record anywhere, so an interrupted one and one that never happened look identical to the customer.** MEASURED 2026-09-17 (chaos night, round 10) on `tester-1-022354` (controller 0.245.0): an app restore was accepted at 23:26:08Z (`302`, flash "Visszaallitas elindult") and the guest's host was hard-reset **four seconds later**, mid-write. Afterwards, asked through the doors the UI itself uses: `/api/restore/status`, `/api/backup/restore/status`, `/backup/restore/status` and `/api/restore` **all 404**; `/api/backup/status` returns `{enabled,running}` with **no restore field at all**; on `/backups/apps` and `/apps/<name>` the only restore text is a **button label** and a JS label expression (ASCII fragment search, with a negative control returning 0 on every page). On disk, in the real data directory (`/var/lib/docker/volumes/felhom-controller-data/_data`), there is no restore, lock or state file anywhere beneath it, and **no file at all was modified in the reset window**. **The box recovered perfectly** - 26/26 containers back in 150 s, boot reconciliation naming the app it recovered, every front door serving, one true `controller_started` alarm and no false one. What is filed is that the customer pressed a button, was told it had started, and can never learn whether it finished. **Honest limit:** only four seconds elapsed, so the restore may have completed or may never have written a byte - the pre-reset log is unrecoverable (the stream holds zero lines before the reboot) and the debug ring is in memory. The absence of any restore record is verifiable independently of how far it got, and that is the row. **Fix shape:** persist a restore record (started, finished or abandoned-at-boot) the way the backup tiers already persist theirs, and show it on the app's page; boot reconciliation is the natural place to mark an in-flight restore as abandoned. Evidence: `audits/evidence-chaos-night-2026-09-17/round-10.txt`. | **READY - rank P2-MEDIUM; owner: CC** |
|
||||
| **R-550** | **[P2-MEDIUM] The restore record is in-memory only: after the machine stops, the status surface is blank and nothing tells the household their restore did not finish.** MEASURED 2026-09-17 (chaos night, round 10) on `tester-1-022354` (controller 0.245.0): an app restore was accepted at 23:26:08Z (`302`, flash "Visszaallitas elindult") and the guest's host was hard-reset **four seconds later**, mid-write. Afterwards `/api/backup/restore-status` - the endpoint the restore page itself polls - answers `{"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}`: the Go **zero value**, with **no `last` field at all**, while the page's own script renders "<operation> sikertelen." from `st.last.message` and "<operation> folyamatban" from `st.op`. So the UI has a last-operation branch with nothing to populate it after a restart. On disk, in the real data directory, no restore, lock or state file exists and **no file at all was modified in the reset window**. **The box recovered perfectly** - 26/26 containers back in 150 s, boot reconciliation naming the app it recovered, every front door serving, one true `controller_started` alarm and no false one. **CORRECTED 2026-09-17T00:24Z:** this row first claimed no restore surface existed at all, citing four endpoints that 404'd - all four were paths I GUESSED, and all four were wrong. The real routes (`/backups/restore`, `/stacks/<name>/backup`, `/api/backup/restore-status`) came from the controller's own rendered links. The corrected claim is narrower and stands on the endpoint's own answer. **Honest limit:** only four seconds elapsed, so `started_at` may be zero because the restore never truly began rather than because the reboot erased it; the pre-reset log is unrecoverable (the stream holds zero lines before the reboot). Either way nothing tells the customer. **Fix shape:** persist the last restore outcome the way the backup tiers already persist theirs, and have boot reconciliation mark an in-flight restore as abandoned so the page can say so. Evidence: `audits/evidence-chaos-night-2026-09-17/round-10.txt`. | **READY - rank P2-MEDIUM; owner: CC** |
|
||||
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
|
||||
|
||||
Reference in New Issue
Block a user