diff --git a/REPORT.md b/REPORT.md index 68107b67..6cee03f9 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,599 +1,348 @@ -# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night) +# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22) -**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no -image built, nothing deployed.** Evidence: +**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of +2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product +where one customer action causes permanent total loss.** + +The drill that found them is preserved at +`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in `documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`. --- -## 1. THE VERDICT +## 0. Baselines, re-established — not trusted from the sheet -**It is a mixture, and the drill's three options are all present — but they belong to different -faults, and only one of them explains the thing you actually saw.** - -### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED** - -Reproduced independently tonight, and it agrees with what R-353 already recorded: - -1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data - drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** — - about an app that was installed, deployed, running and healthy. -2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt - box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it - held `compose/` and nothing else. -3. The restore returned that configuration and reported a bare completion. - -So for that specific event the data was not in the thing that was restored. **R-353 called this -correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong -to think so; its text is more careful than the CHANGELOG's summary of it. - -### What the drill found that nobody had seen: **THE RESTORE LOSES IT** - -This is new, it is worse, and it is not the same fault: - -**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.** -The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.** - -Proven live on `calibre-web` at 22:23:36 with planted files: - -| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore | -|---|---|---|---|---| -| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** | -| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** | - -and the customer was told: - -> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult. -> Ennek az alkalmazásnak nincs adatbázisa." - -Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence -does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue -apps that declare no data drive, that leg is the entire dataset.** - -**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit` -(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the -unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local** -restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring -PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of -them replays it.** - -### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database - -`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always -has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to -`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack -that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the -safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and -then says: - -> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" - -The controller had dumped that database five minutes earlier. - -**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive. +| item | expected | confirmed | +|---|---|---| +| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` | +| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box | +| golden vouched | 0.217.0 | ✔ `