From 877fcd2a385b95bda5752760ccc0d3540f925635 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 22 Aug 2026 10:11:46 +0200 Subject: [PATCH] R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367. --- REPORT.md | 841 ++++++------------ STATUS.md | 42 +- .../REPORT-DRILL-backup-truth-2026-08-21.md | 599 +++++++++++++ documentation/backlog/OPEN-ITEMS.md | 5 +- .../tests/golden-0.218.0-2026-08-22/README.md | 44 + .../tests/golden-0.218.0-2026-08-22/bake.log | 323 +++++++ 6 files changed, 1286 insertions(+), 568 deletions(-) create mode 100644 documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md create mode 100644 documentation/tests/golden-0.218.0-2026-08-22/README.md create mode 100644 documentation/tests/golden-0.218.0-2026-08-22/bake.log diff --git a/REPORT.md b/REPORT.md index 68107b67..6cee03f9 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,599 +1,348 @@ -# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night) +# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22) -**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no -image built, nothing deployed.** Evidence: +**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of +2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product +where one customer action causes permanent total loss.** + +The drill that found them is preserved at +`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in `documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`. --- -## 1. THE VERDICT +## 0. Baselines, re-established — not trusted from the sheet -**It is a mixture, and the drill's three options are all present — but they belong to different -faults, and only one of them explains the thing you actually saw.** - -### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED** - -Reproduced independently tonight, and it agrees with what R-353 already recorded: - -1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data - drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** — - about an app that was installed, deployed, running and healthy. -2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt - box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it - held `compose/` and nothing else. -3. The restore returned that configuration and reported a bare completion. - -So for that specific event the data was not in the thing that was restored. **R-353 called this -correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong -to think so; its text is more careful than the CHANGELOG's summary of it. - -### What the drill found that nobody had seen: **THE RESTORE LOSES IT** - -This is new, it is worse, and it is not the same fault: - -**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.** -The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.** - -Proven live on `calibre-web` at 22:23:36 with planted files: - -| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore | -|---|---|---|---|---| -| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** | -| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** | - -and the customer was told: - -> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult. -> Ennek az alkalmazásnak nincs adatbázisa." - -Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence -does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue -apps that declare no data drive, that leg is the entire dataset.** - -**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit` -(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the -unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local** -restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring -PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of -them replays it.** - -### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database - -`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always -has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to -`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack -that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the -safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and -then says: - -> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" - -The controller had dumped that database five minutes earlier. - -**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive. +| item | expected | confirmed | +|---|---|---| +| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` | +| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box | +| golden vouched | 0.217.0 | ✔ `