writeSafetyDump called DumpOne into the app's OWN unit dir and renamed the result
to pre-restore-* afterwards. DumpOne writes <stack>-<dbtype>.sql - the app's
canonical dump - so every safety dump overwrote the app's real backup and then
moved it away, leaving the app with no database backup until the next nightly
run. A local restore-from-unit in that window tells the customer the app never
had a database.
The comment beside it asserted the rename meant it 'can never overwrite the app's
real dump'. False as written, and believed for four months. Measured live before
the fix: docmost and bookstack each held only pre-restore-* files and no
canonical dump.
DumpOneTo takes the final path and derives its own .tmp from it. DumpOne keeps
its signature and calls it with the canonical name. writeSafetyDump asks for its
own name directly; the rename is gone; the comment now states the invariant and
how it is enforced.
db_dumps no longer lists the undo copies. All three consumers of Manifest.DBDumps
were grepped and named - all inside recovery_unit.go, none reads it for recovery.
The files are neither deleted nor hidden.
Tests 1485 -> 1493. FIVE red-proofs, TWO PASSED first time and both are reported:
the behavioural tests inject the dump seam so a mutation inside DumpOneTo was
invisible, and 1.3 had no test at all. Guards added at the layer each defect
lives in; both mutations then convicted.
R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.
The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.
The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.
When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.
--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.
R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.
Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's
PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that
does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The
same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo
copy and the fail-closed refusal was never reached. Fixed by reading the compose project label,
which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no
-p). The old derivation stays as the fallback and an unresolvable attribution is now loud.
Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller,
not the catalogue.
R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside
the unit, so the off-site restore had no volume leg at all — proven live with planted files:
calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking
folder, and the restore reported success without it. For the 40 of 53 apps that declare no data
drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay
with an explicit directory: ONE implementation, two callers. Volumes replay before the database
and inside the stopped window. VolumesReplayed reaches the message.
The comment beside the skip was half false and is corrected; the half that still holds — the
live unit is the local path's source — is named, and scenario D fingerprints the whole live unit
across the operation.
Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the
code: scenario D passed with the unit guard removed because the fingerprint had been narrowed
and was blind to the unit root.