Files
felhom.eu/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/FINDING.txt
T
admin f5a4fceeeb
gates / gates (push) Successful in 16s
DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
Diagnostic only — no code changed, no version bumped, nothing deployed.

The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.

Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
  R-354 off-site restore never replays volume dumps
  R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
        has no DB dump, no safety dump is taken, and the customer is told it has none
  R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
        "is not installed", with a remedy those apps make impossible

R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.

Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
2026-08-21 23:30:27 +02:00

35 lines
2.2 KiB
Plaintext

paperless-ngx: the database is dumped into a directory for a stack that does not exist,
so the recovery unit never contains it, and the restore then tells the customer the app
has no database. Proven live 2026-08-21 22:39-22:45 CEST on demo-hp.
MECHANISM (source):
internal/appbackup/dbdump.go:770 deriveStackName("paperless-postgres", known)
1. candidate = suffixStripStackName("paperless-postgres") = "paperless" (line 801)
2. known is non-empty, known["paperless"] is FALSE (the stack is "paperless-ngx")
3. known["paperless-postgres"] is FALSE
4. no known stack name is a prefix of "paperless-postgres" ("paperless-ngx" is not)
5. FALLS THROUGH to `return candidate` (line 797) -> "paperless"
An unresolved mapping is returned as if resolved. There is no warning and no refusal.
OBSERVED CONSEQUENCES (all live):
a) the dump is written to
/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql
284,617 bytes, 72 tables, valid=true -- an orphan directory for a non-existent stack,
on the SYSTEM drive, while the app's real unit is on /mnt/felhom-drives/hdd_1.
b) the real unit /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx/manifest.json
records "db_dumps": null
c) the off-site snapshot therefore carries no .sql at all
(checking folder: `find ... -name "*.sql" | wc -l` = 0)
d) writeSafetyDump (offbox_reconstitute.go:115) filters discovered DBs on
db.StackName == stackName, so `mine` is empty -> returns ("", nil) -> hasDB = false.
NO pre-restore safety dump is taken. Verified: `find /mnt -name "pre-restore-*"`
returned nothing before AND after the destructive restore.
e) the destructive restore ran to completion and reported SUCCESS:
"A(z) paperless-ngx: 0 fájl visszaállítva (mentés: 2026-08-21 22:41)
— az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."
The controller had dumped that same database 5 minutes earlier.
The orphan directory is also invisible to the app's off-site push, because the push
resolves paths from the app's own unit path -- so the only copy of that database dump
is on the system drive of the machine it protects.