Files
felhom.eu/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-part4/results.txt
T
admin f5a4fceeeb
gates / gates (push) Successful in 16s
DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
Diagnostic only — no code changed, no version bumped, nothing deployed.

The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.

Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
  R-354 off-site restore never replays volume dumps
  R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
        has no DB dump, no safety dump is taken, and the customer is told it has none
  R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
        "is not installed", with a remedy those apps make impossible

R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.

Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
2026-08-21 23:30:27 +02:00

65 lines
4.2 KiB
Plaintext

PART 4 — things we claim and have never watched. demo-hp, 2026-08-21, times CEST.
4.1 THE DESTRUCTIVE RESTORE
(a) "nothing is ever deleted" -- PASS, both directions, calibre-web 22:26:59.
POST-SNAPSHOT.txt, created after the snapshot, SURVIVED the restore.
plain.txt, mutated after the snapshot, was OVERWRITTEN back to the snapshot's
content (sha 07e91a98…). Copier is rsync -a, no --delete, no --ignore-existing
(offbox_reconstitute.go:94).
NOTE ON THE COUNT: the message says "2 fájl visszaállítva" because rsync counts
TRANSFERS, not files restored. An identical restore reports "0 fájl visszaállítva",
which is indistinguishable from a restore that did nothing.
(b) "a safety dump is taken and VERIFIED before anything is stopped, and the whole
operation refuses if it cannot be"
HAPPY PATH -- PASS, romm 23:02:44.
21:02:47Z "[offbox] romm: pre-restore safety dump written →
pre-restore-20260821T210246Z-romm-mariadb.sql (60.8 KB)"
21:02:47Z "[stacks] StopStack romm: current state=running"
The dump precedes the stop. Message correctly said
"0 fájl és az adatbázis visszaállítva".
THE REFUSAL -- PASS, romm 23:04:15.
Safety dump made impossible by putting a regular FILE at the db-dumps path.
Result: ok=FALSE,
"A teljes visszaállítás sikertelen: a biztonsági mentés könyvtára nem hozható
létre: mkdir /mnt/felhom-drives/hdd_1/backups/primary/romm/db-dumps:
not a directory"
Nothing changed: plain.txt kept my post-snapshot mutation (sha 3754dfc6…), and
romm's container StartedAt was unchanged (21:03:08) -- the app was never stopped.
JUDGEMENT: honest and it names the path, but it leaks a raw Go mkdir error into
a customer surface.
THE HOLE -- the guard only protects apps whose database the discovery resolves to
the right stack. For paperless-ngx it concludes "no database", so hasDB is false
and the refusal CANNOT fire: the undo is absent rather than refused. See
../phase4-paperless/FINDING.txt.
SIDE EFFECT, NOT PREVIOUSLY FILED -- the safety dump DESTROYS the unit's own DB dump.
DumpOne writes the canonical `<stack>-<dbtype>.sql` (appbackup/dbdump.go:200-202),
i.e. the app's real dump, and only THEN is it renamed to pre-restore-*.
The comment at offbox_reconstitute.go:147-148 states it "can never overwrite the
app's real dump". It does.
PROVEN: romm's db-dumps held romm-mariadb.sql (62,270 B) at 22:59; after one
reconstitute it held ONLY pre-restore-20260821T210246Z-romm-mariadb.sql.
Consequence: until the next backup run the LOCAL restore-from-unit finds no .sql
and tells the customer the app has no database.
4.3 A DAMAGED STORE -- MIXED
Method: one byte flipped inside pack 967853d2… (offset 5,000,000) via the repo's own
SFTP transport; the pack's name is its content hash, so this is genuine corruption.
* `restic check` DOES detect it ("ciphertext verification failed",
"Fatal: repository contains errors"). BUT the controller NEVER RUNS `restic check`:
the only restic verbs in the whole controller are restore, snapshots, backup, unlock,
stats, init, forget, prune, cat. The agent's restore-test is PBS-tier only.
So the off-site store is never verified by any layer, at any time.
* A restore that TOUCHES the damage fails honestly:
ok=FALSE, "A visszaállítás sikertelen: offbox restore paperless-ngx: exit status 1:
… ignoring error for …/documents/originals/0000011.pdf: ciphertext verification failed"
* BUT the failure is not remembered. It left a PARTIAL scratch (78 MB, 54 files,
15 of 16 originals). OffboxFullScratchReady (offbox_restore.go:305) only asks
"does the directory exist and is it non-empty", so the wizard then offered all three
actions including "Teljes visszaállítás indítása".
* Pressing it ran the DESTRUCTIVE restore from that known-incomplete copy and reported
SUCCESS: ok=TRUE, "A(z) paperless-ngx: 0 fájl visszaállítva … Ennek az alkalmazásnak
nincs adatbázisa."
REPO REPAIRED afterwards from the byte-identical originals; `restic check` now says
"no errors were found".