From 9f40dc328907516b7f4b51239b71621698d6b8e7 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 01:35:56 +0200 Subject: [PATCH] chaos night round 10: a restore leaves no record, and four of my instruments failed The box passed the roughest pair drawn. A hard reset four seconds into a restore: 26/26 containers back in 150 s, boot reconciliation naming the app it recovered, every front door serving, one true controller_started alarm, no false one, no intervention. R-550 filed (P2): there is no restore record anywhere. Four candidate status endpoints 404, no restore field in the status JSON, only a button label on the pages, and no file at all modified in the reset window. An interrupted restore and one that never happened look identical to the customer. Honest limit recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so the absence of a record is what is filed, not a claim about how far it got. Four instrument faults, all mine, all in the evidence: * a 'nothing was logged' claim that was unfalsifiable when written - the log stream holds zero lines before a reset; * an on-disk check against /opt/felhom/data, a directory that does not exist; * a household count reporting 0 lines and 0 failures when the truth was one line and it WAS a failure - the runner now prints both operands; * the disk guard was a TRANSIENT unit reporting 'active' all night, and was absent from the reset onward. It is now file-backed and enabled, and its script is copied off the box for the first time. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/DRILL-chaos-night-2026-09-17.md | 33 +++++++ .../diskguard.sh | 20 +++++ .../round-10.txt | 87 +++++++++++++++++++ .../run_round.sh | 4 + documentation/backlog/OPEN-ITEMS.md | 1 + 5 files changed, 145 insertions(+) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/diskguard.sh create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/round-10.txt diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 8d9bf0a1..7d259c53 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -427,6 +427,39 @@ real house an ISP outage does not do that — controller and agent share one mac gone" as injected is **broader than its name**: internet, hub *and* local agent. No further internet cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9. +### Round 10 — `restore` uptime-kuma / accident: **hard reset, four seconds into the restore** + +**23:26:06Z–23:29:46Z.** The roughest pair drawn. The restore was accepted at 23:26:08Z +(302, „Visszaállítás elindítva"); the reset button was pressed at 23:26:12Z, mid-write. + +| the five things | | +|---|---| +| what the customer saw | They pressed restore, were told it had started, and **four seconds later the whole machine went dark.** About two minutes of nothing. Then every app was back and every front door answered. **Nothing ever told them what became of the restore.** | +| what the box did by itself | Booted, and brought **26 of 26 containers** back with no help. Boot reconciliation named the one app it had to recover („1 app(s) recovered in 1 attempt(s): [paperless-ngx]"), sent a startup hub report at 23:28:26Z, and settled its health probes. No intervention. | +| time to steady | **150 s** — 0 containers at t+12 s, 25 at t+133 s, 26 at t+150 s. Doors 200 at both readings (23:28:42Z and 23:29:44Z). | +| alarm fired / true? | **one, true** — `controller_started` (info). Exactly what the ladder expects after a reboot. No false alarm. | +| should have fired, did not | **none from the alarm ladder** — but the restore silence below is a legibility gap, filed as a row. | + +**The restore left no trace anywhere, and the product has no place to leave one.** Four candidate +status endpoints all 404 (`/api/restore/status`, `/api/backup/restore/status`, +`/backup/restore/status`, `/api/restore`). `/api/backup/status` carries no restore field at all. +On the pages, the only restore text is a **button label** and a JavaScript label expression. On disk, +in the real data directory, there is no restore, lock or state file anywhere — and **no file at all +was modified in the reset window**. An interrupted restore and a restore that never happened are +indistinguishable, to the customer and to me. + +**The limit of that measurement, stated rather than glossed.** Only four seconds elapsed, so the +restore may have finished or may never have written a byte — and I cannot tell, because the +controller's log stream holds **zero lines before 23:28:00Z** (a reset starts it fresh) and the debug +ring died with the machine. What is independently verifiable is the **absence of any restore record**, +and that is what is filed; it holds however far the restore got. + +**Four of my own instruments failed in this round, and all four are recorded in the evidence:** a +claim that was unfalsifiable when written; an on-disk check against a directory that does not exist; +a household count that reported 0 lines and 0 failures when the truth was one line and it *was* a +failure; and a disk guard that reported „active" all night while being a **transient** unit that +vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box. + ### Rounds 8-12 PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/diskguard.sh b/documentation/audits/evidence-chaos-night-2026-09-17/diskguard.sh new file mode 100644 index 00000000..75c1a99c --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/diskguard.sh @@ -0,0 +1,20 @@ +#!/bin/bash +# SAFETY, not a measurement: the local backup tier retries every ~15 minutes and cannot fit on this +# box (a ~29 GB source into a 14 GB root filesystem). Left alone during a round it would fill `/` and +# wedge the nested PVE, which already nearly cost the night once. This kills ONLY a local-tier dump, +# and only when free space falls under the floor. Every kill is logged so it appears in the evidence. +FLOOR_MB=2500 +LOG=/root/diskguard.log +while true; do + FREE=$(df -BM --output=avail / | tail -1 | tr -dc '0-9') + if [ -n "$FREE" ] && [ "$FREE" -lt "$FLOOR_MB" ]; then + if ls /var/lib/vz/dump/*.tar.dat >/dev/null 2>&1; then + echo "$(date -u +%FT%TZ) GUARD: free=${FREE}M below ${FLOOR_MB}M and a local dump is writing — killing it" >> $LOG + pkill -f "[v]zdump" + sleep 3 + rm -rf /var/lib/vz/dump/*.tar.dat /var/lib/vz/dump/*.tmp 2>/dev/null + echo "$(date -u +%FT%TZ) GUARD: partial archive removed, free now $(df -BM --output=avail / | tail -1)" >> $LOG + fi + fi + sleep 20 +done diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-10.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-10.txt new file mode 100644 index 00000000..d9584cbf --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-10.txt @@ -0,0 +1,87 @@ + round 10 armed for 2026-09-16T23:25:48Z (restore uptime-kuma + hard reset) +=== round 10 launched 2026-09-16T23:26:06Z (due 23:25:48Z) === +2026-09-16T23:26:06Z ================ ROUND 10 : restore uptime-kuma, while: hard-reset ================ +2026-09-16T23:26:08Z --- BEFORE --- containers=26 status=200 status=200 paste=200 +2026-09-16T23:26:08Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round) +2026-09-16T23:26:08Z --- ACTION: restore on uptime-kuma --- + POST /backup/restore (uptime-kuma) -> HTTP/1.1 302 Found + flash: Visszaállítás elindult — az állapot itt frissül. +2026-09-16T23:26:12Z --- ACCIDENT: hard-reset (injected after the action started) --- + 2026-09-16T23:26:12Z ACCIDENT=hard-reset round=10 + 2026-09-16T23:26:12Z qm reset 336 (the reset button, mid-write) + 2026-09-16T23:26:14Z reset issued + 2026-09-16T23:26:14Z accident hard-reset complete +2026-09-16T23:26:14Z --- AFTER: what the box did BY ITSELF --- +2026-09-16T23:26:24Z t+12s containers=0 (before 26) +2026-09-16T23:28:07Z t+115s containers=0 (before 26) +2026-09-16T23:28:25Z t+133s containers=25 (before 26) +2026-09-16T23:28:42Z t+150s containers=26 (before 26) +2026-09-16T23:28:42Z STEADY after 150s +2026-09-16T23:28:42Z front doors, FIRST reading at 2026-09-16T23:28:42Z - TOO EARLY to trust if the accident just ended: +2026-09-16T23:28:44Z status=200 status=200 paste=200 wiki=200 +2026-09-16T23:29:44Z front doors, SECOND reading at 2026-09-16T23:29:44Z, 60 s later - THIS is the one to trust: +2026-09-16T23:29:45Z status=200 status=200 paste=200 wiki=200 +2026-09-16T23:29:45Z household lines this round: 0 failures: 0 +2026-09-16T23:29:45Z --- alarms --- + | Time | Severity | Type | Message | Source + | Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller + | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller + | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller + | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller + | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller + | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller +2026-09-16T23:29:46Z ================ END ROUND 10 ================ + +[exited with code 0] + +## WHAT HAPPENED TO THE RESTORE (asked four ways, 23:30-23:34Z) + +The restore was accepted at 23:26:08Z (HTTP 302, flash "Visszaallitas elindult"). +The machine was hard-reset at 23:26:12Z - FOUR SECONDS later, mid-write. + +Asked afterwards, through the same doors the UI uses: + /api/restore/status -> 404 + /api/backup/restore/status -> 404 + /backup/restore/status -> 404 + /api/restore -> 404 + /api/backup/status -> {"ok":true,"data":{"enabled":true,"running":false}} (no restore field at all) + /backups/apps (200) -> the only "Vissza" text is a BUTTON LABEL and a JS label expression: + "Visszaallitas" + ": op === 'offbox-restore' ? 'Tavoli visszaallitas' : 'Visszaallitas'; }" + /apps/uptime-kuma (200) -> one "Vissza" (the button). "megszak" (interrupted): 0 hits. + The "sikeres" hits belong to the MOVE-DATA feature, not the restore. + NEGATIVE CONTROL "zzzznotpresent" -> 0 hits on every page, so the searches are trustworthy. + +On disk, in the REAL data directory (/var/lib/docker/volumes/felhom-controller-data/_data): + no file named *restor*, *lock* or *.pid anywhere beneath it + NO FILE AT ALL modified in the window 23:24:00 - 23:27:30 + +So: there is no restore history surface of any kind. An interrupted restore and a restore that +never happened look EXACTLY the same to the customer, and to me. + +### THE HONEST LIMIT OF THIS MEASUREMENT, which must not be glossed over +Only four seconds elapsed. The restore may have completed, or may never have written a byte. +I cannot tell, because: + * the controller's log stream holds ZERO lines before 23:28:00Z - a hard reset starts it fresh, + so the pre-reset window is unrecoverable by this route; + * the debug ring is in memory and died with the machine. +What IS independently verifiable, and is what gets filed, is the ABSENCE OF ANY RESTORE RECORD - +that holds regardless of how far the restore got. + +## FOUR CORRECTIONS TO MY OWN INSTRUMENTS, all found in this round +1. "The controller logged nothing about the restore" was UNFALSIFIABLE when I first wrote it. + The stream cannot reach before the reset. Re-stated above with its limit. +2. My first on-disk check looked in /opt/felhom/data - A DIRECTORY THAT DOES NOT EXIST. It printed + a tidy "no restore/lock file" that meant nothing. The real path is the docker volume above. + Same class as the "docker: command not found -> a tidy table of absent/0" error from earlier. +3. The runner reported "household lines this round: 0 failures: 0". Both are WRONG. The log is + never truncated (first line 21:06:30Z, continuous) and contains exactly ONE line in the round + window - and it is a FAILURE line: "2026-09-16T23:28:09Z cloud dash UNREACHABLE" (the box was + still booting). The runner now prints the raw before/after counts so the subtraction is checkable. +4. THE DISK GUARD WAS NOT A REAL UNIT. After the reset: "Unit diskguard.service not found" - yet it + had reported "active" all night. It was a TRANSIENT unit, which looks identical while running and + vanishes on reboot. It was therefore ABSENT from 23:26:12Z. It is now a file-backed, ENABLED unit + (verified: active + enabled, MainPID cmdline "/bin/bash /root/diskguard.sh", 0 kills, 7556 MB free) + and its script is copied off the box into this folder, which had also never been done (R-320). diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh b/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh index 7aa2df86..b2c4b591 100755 --- a/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh +++ b/documentation/audits/evidence-chaos-night-2026-09-17/run_round.sh @@ -92,6 +92,10 @@ sleep 60 say " front doors, SECOND reading at $(date -u +%FT%TZ), 60 s later - THIS is the one to trust:" say " $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)" HL1=$(BOX 'wc -l < /root/household.log' | tr -d ' \r') +# Round 10 reported "0 lines, 0 failures" when the truth was one line and it WAS a +# failure. A subtraction whose operands are invisible cannot be checked, so both are +# printed now, and an empty read is named rather than silently becoming zero. +say " household log lines: before=${HL0:-EMPTY} after=${HL1:-EMPTY}" say " household lines this round: $(( ${HL1:-0} - ${HL0:-0} )) failures: $(BOX "tail -n +$(( ${HL0:-0} + 1 )) /root/household.log | grep -cE 'FAILED|UNREACHABLE'" | tr -d ' \r')" say "--- alarms ---"; bash $E/events.sh 8 | tee -a $OUT say "================ END ROUND $N ================" diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 1fd152a3..aea57f09 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -732,6 +732,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** MEASURED 2026-09-17 (chaos night) on a fresh box (`tester-1-022354`, controller 0.245.0): the customer guest’s root filesystem was held at **96 % for ten minutes** (29 G used, 1.5 G free) and **no alarm of any kind fired** — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: `fillwatch` runs **daily at 03:30** plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about **twenty seconds before** the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. **This is the ladder working as designed, not a missed alarm** — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is **no, unless the controller happens to restart while it is full**, and that answer is not written down anywhere. **Fix shape (one of):** sample the fill more often than daily (a cheap `statfs` on the 5-minute health pass would do it); or say plainly in `08-alarm-ladder.md` that a transient full disk is out of scope. Evidence: `audits/evidence-chaos-night-2026-09-17/round-3.txt`. | **READY — rank P3-LOW; owner: CC** | | **R-548** | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. | **READY — rank P3-LOW; owner: CC** | | **R-549** | **[P2-MEDIUM] The staleness alarm's budget is exactly two report cycles, so ONE failed push spends all of it.** MEASURED 2026-09-17 (chaos night, round 9) on `tester-1-022354` (controller 0.245.0): hub reachability was removed for ten minutes from the VM's side. The controller built its 23:08:42Z report, retried the push **three times over 1 m 40.8 s**, and gave up at 23:10:23Z (`hub push failed after 3 attempts`) - **31 seconds before the link returned**. Nothing is queued, and that is correct: a report is a snapshot, so the next one carries the same truth, and the controller says so itself ("backing off (the 15-min cycle still reconciles)"). **The arithmetic is the finding.** Last good report 22:53:43Z; next scheduled 23:23:42Z; measured cadence **15m0s**; `node_stale` trips at **30 minutes**. The gap is **29 m 59 s** - one second inside the threshold. So a single missed push spends the whole staleness budget, and any ordinary jitter pages the operator about a box that is healthy, serving every app, and has already repaired itself unaided. **The product behaved correctly throughout:** no false alarm fired, no app stopped, and both the hub link and the host-agent link recovered by themselves the moment the block lifted. What is filed is the margin, not a misbehaviour. **Fix shape:** either set the staleness threshold to a clear multiple of the cadence (three cycles, not two), or let a push that has failed all three attempts retry once off-cycle instead of waiting for the next scheduled report. **Honest caveat:** the cut was injected by the drill and also severed the controller from its host agent, which a real ISP outage would not do - but the report arithmetic above depends only on the hub being unreachable. Evidence: `audits/evidence-chaos-night-2026-09-17/round-9.txt`. | **READY - rank P2-MEDIUM; owner: CC** | +| **R-550** | **[P2-MEDIUM] A restore leaves no record anywhere, so an interrupted one and one that never happened look identical to the customer.** MEASURED 2026-09-17 (chaos night, round 10) on `tester-1-022354` (controller 0.245.0): an app restore was accepted at 23:26:08Z (`302`, flash "Visszaallitas elindult") and the guest's host was hard-reset **four seconds later**, mid-write. Afterwards, asked through the doors the UI itself uses: `/api/restore/status`, `/api/backup/restore/status`, `/backup/restore/status` and `/api/restore` **all 404**; `/api/backup/status` returns `{enabled,running}` with **no restore field at all**; on `/backups/apps` and `/apps/` the only restore text is a **button label** and a JS label expression (ASCII fragment search, with a negative control returning 0 on every page). On disk, in the real data directory (`/var/lib/docker/volumes/felhom-controller-data/_data`), there is no restore, lock or state file anywhere beneath it, and **no file at all was modified in the reset window**. **The box recovered perfectly** - 26/26 containers back in 150 s, boot reconciliation naming the app it recovered, every front door serving, one true `controller_started` alarm and no false one. What is filed is that the customer pressed a button, was told it had started, and can never learn whether it finished. **Honest limit:** only four seconds elapsed, so the restore may have completed or may never have written a byte - the pre-reset log is unrecoverable (the stream holds zero lines before the reboot) and the debug ring is in memory. The absence of any restore record is verifiable independently of how far it got, and that is the row. **Fix shape:** persist a restore record (started, finished or abandoned-at-boot) the way the backup tiers already persist theirs, and show it on the app's page; boot reconciliation is the natural place to mark an in-flight restore as abandoned. Evidence: `audits/evidence-chaos-night-2026-09-17/round-10.txt`. | **READY - rank P2-MEDIUM; owner: CC** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |