R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s

Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.

R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.

R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.

Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.

R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.

Ceiling R-366 -> R-367.
This commit is contained in:
2026-08-22 10:11:46 +02:00
parent 7064596c2e
commit 877fcd2a38
6 changed files with 1286 additions and 568 deletions
+22 -20
View File
@@ -1,6 +1,6 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-21 (night — the backup holds the data; the off-site RESTORE is what loses it).**
**Updated 2026-08-22 — both of last night's worst findings are fixed and proven on the machine.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
@@ -113,25 +113,27 @@ record with no machine** — created 13 August, no host, no backups, nothing to
## Broken, or knowingly incomplete
- **The off-site restore does not give back an app's Docker volume — and says it worked** (R-354).
Proven on `demo-hp` on the night of 21 August with files we planted and hashed first. The data IS in
the backup: the unit, the off-site snapshot and the verification folder all hold it, byte for byte,
accented Hungarian filenames and all. The **full off-site restore then puts back the user files and
silently skips the volume**, and tells the customer „5 fájl visszaállítva". For an app that keeps its
files on a data drive the lost piece is its settings. **For the 40 of our 53 apps that have no data
drive, that volume IS all their data.** The *local* restore does return it correctly — same tar, other
code path. **Nothing is lost off-site; the last step is what fails.** *(register: R-354)*
- **For those same 40 apps the off-site restore refuses outright, and the reason it gives is untrue**
(R-356). It says the app „nincs telepítve" — is not installed — about an app that is installed and
running, then tells the customer to reinstall it "to the same place", which those apps give them no
way to choose. This is what the OpenGist journey hit on the afternoon of 21 August before falling
back to the local restore. *(register: R-356)*
- **Paperless's database has never been in its backup** (R-355). Its dump is written every night, is
valid, has 72 tables — and lands in a folder named after an app that does not exist, on the wrong
disk, where nothing sends it off-site and nothing restores it. Because the same wrong name is used
when taking the "undo" copy before a restore, **a restore of Paperless takes no undo at all** and then
tells the customer the app has no database. **One app in 53 is affected. It is the document
archive.** *(register: R-355)*
- **FIXED and proven on the machine: the off-site restore gives the app's data back** (R-354,
controller 0.218.0). The same planted files, the same steps, both runs on `demo-hp`: on the old
build the restore said „0 fájl visszaállítva", reported success, and the folder was simply not
there. On the new one it says „**0 fájl és 1 adatkötet visszaállítva**" and all five files come back
**byte for byte**, Hungarian accented names included. The message now names what came back, because
a restore that mentions only its file count is how a silent loss reads as a success. *(register:
R-354, CLOSED)*
- **FIXED and proven on the machine: Paperless's database is in the backup, and a restore of it now
takes an undo copy first** (R-355, controller 0.218.0). The dump was going into a folder named after
an app that does not exist, so nothing collected it — and because the same wrong name was used when
looking for the live database, a restore took **no undo copy at all**. Now: the dump is in the app's
own backup and in the off-site copy for the first time; the restore said „**0 fájl és 3 adatkötet és
az adatbázis visszaállítva**"; and with the undo deliberately made impossible the restore **refused
and did not even stop the app**. We no longer guess which app a database belongs to — Docker already
tells us. **One app of 53 was affected**, established with a check we first proved could catch a
planted second case. *(register: R-355, CLOSED)*
- **STILL BROKEN, and it is now the one that matters most: 40 of our 53 apps still cannot use the
off-site restore at all** (R-356). It refuses before it starts, says a running app „nincs telepítve"
— is not installed — and tells the customer to reinstall it "to the same place", which those apps
give them no way to choose. **Those are exactly the apps whose entire data is the thing R-354 just
fixed**, so today's fix cannot reach them until this one is done. *(register: R-356)*
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the
agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the
standard tool — which we never run. *(register: R-359)*