Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367.
20 KiB
REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)
Controller v0.217.0 → v0.218.0. Two fixes, both found by watching a machine on the night of 2026-08-21, both confirmed the same way. Part 1 first, because it is the only place in the product where one customer action causes permanent total loss.
The drill that found them is preserved at
documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md; its evidence is in
documentation/audits/DRILL-backup-truth-2026-08-21/evidence/.
0. Baselines, re-established — not trusted from the sheet
| item | expected | confirmed |
|---|---|---|
| controller | 0.217.0 → 0.218.0 | ✔ repo head v0.217.0; live on demo-hp …:0.217.0 |
| agent | 0.130.0 | ✔ repo head and felhom-agent --version on the box |
| golden vouched | 0.217.0 | ✔ <option value="0.217.0" … selected> |
| floor | 0.217.0 | ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)" |
| min agent | 0.129.0 | ✔ |
| register ceiling | R-366 | ✔ (now R-367) |
ssh hp → 192.168.0.104 |
works | ✔ |
| clean trees, all four repos | HEAD == origin/main, 0 dirty |
✔ |
1. PART 1 — where the wrong name comes from
It comes from a guess, made in a place where an answer was available.
deriveStackName (controller/internal/appbackup/dbdump.go:770-798) is handed a container name and
asked which app owns it. For paperless-postgres it strips the role suffix to paperless; finds
paperless is not a deployed stack; finds the container name is not one either; finds no known
stack is a prefix of it (paperless-ngx is not a prefix of paperless-postgres) — and then returns
the unresolved candidate anyway, on its last line, silently.
Every place that value is used, and what each did with it:
| consumer | file:line | effect |
|---|---|---|
| dump directory | backup/backup.go:482,506 |
wrote to …/primary/**paperless**/db-dumps/ on the SYSTEM drive |
| dump filename | appbackup/dbdump.go:200-212 |
paperless-postgres.sql, not paperless-ngx-postgres.sql |
| unit assembler | backup/recovery_unit.go:131 |
reads …/primary/**paperless-ngx**/db-dumps → finds nothing → db_dumps: null |
| off-site collector | backup/offbox_capture.go:32 |
resolves from the unit path → the orphan is invisible to it |
| restore's DB leg | backup/restore_db.go:69 |
db.StackName != stackName → never replays |
| safety-dump filter | backup/offbox_reconstitute.go:134 |
mine empty → hasDB=false → no undo copy, and the refusal is never reached |
.fab export |
appexport/export.go:600 |
same filter — the bundle carries no database either |
| the "has a DB" flag | web/handlers.go:1245 |
the app displays as having no database |
Corroboration from the code itself: ListDumpFiles (dbdump.go:562) parses a dump filename under
the comment "Parse stack name and DB type from filename: paperless-ngx-postgres.sql" — the reader
was written for a name the writer never produced.
Controller or catalogue? — THE CONTROLLER. No halt.
Every container the controller starts already carries com.docker.compose.project, read live:
paperless-postgres paperless-ngx kimai-db kimai romm-db romm
And that label is the stack name BY CONSTRUCTION, not by luck: composeExecCustomEnv
(internal/stacks/manager.go:1218-1230) runs compose with cmd.Dir set to
/opt/docker/stacks/<stack> and never passes -p, so compose derives the project from that
directory. Renaming the container in the catalogue would have fixed this one app and left the guessing
for the next one. No catalogue change was made.
2. PART 1.2 — the sweep, proven before its answer was trusted
| step | result |
|---|---|
| scratch copy, unplanted | MISMATCHES: 1 (rc 1) — the known one |
plant a second mismatch (kimai-db → timetrack-db) |
MISMATCHES: 2, convicted by name: stack=kimai container=timetrack-db -> derived=timetrack |
| plant removed | back to 1 |
| every mismatch removed | MISMATCHES: 0, rc 0 — "1" is not a stuck value |
| scratch discarded | kimai/docker-compose.yml md5-identical to the real catalogue |
The real count: 1 affected app of 53, from 15 DB containers checked — paperless-ngx.
3. PART 1 — the dumps already written under the wrong name
They stay, untouched, and nothing will ever collect them.
/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql, 312 381 B,
last written 07:38 UTC by the final pre-fix cycle — verified still present after the fix.
Nothing deletes them by design: the F5 stale-primary prune (backup/backup.go:1248) skips any
directory that is not a deployed app, under the guard "an undeployed app's last backup is still its
restore point".
They can be adopted, by hand: move the file to
…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql and it becomes a readable restore point.
Deliberately not automatic — it would sit beside volume tars from a different time (an incoherent
pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a
customer's data on upgrade is a migration, not a fix. Filed as R-367.
4. PART 2 — the mechanism, and the comment corrected
ReconstituteFromOffsite skipped every placement flagged as the unit (offbox_reconstitute.go:344),
and the volume archives live inside the unit. A search of the whole off-site path found no
volume-restore call; the only two callers of restoreDockerVolumes were restore.go:67 and
restore_unit.go:255, both local.
The comment as it stood:
"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the scratch unit instead, so nothing is lost by skipping it."
Which half still holds: the FIRST. The live unit is the local restore path's own source and must never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit across the operation to keep it true.
The second half was false and is corrected, not left. True of the database dump, false of the volume archives, which live in the same unit and were replayed by nothing at all. Skipping the placement is correct; treating the skip as harmless was not — and that sentence is exactly why a reader would not look.
The fix: restoreDockerVolumesFrom(stackName, dumpDir) (int, error) — the local path's own replay
with an explicit directory, the same shape reimportDBDumpsFrom already had beside reimportDBDumps.
One implementation, two callers. Volumes replay before the database (a logical dump must still
win over a volume-tar copy of the same database) and inside the stopped window (Docker will not
replace a volume a container holds).
5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM
R-354 — calibre-web, identical fixture, identical steps, both on demo-hp:
| before (0.217.0, 09:41) | ok=true — „A(z) calibre-web: 0 fájl visszaállítva (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and ls: /vol/R354-2026-08-22: No such file or directory |
| after (0.218.0, 09:56) | ok=true — „A(z) calibre-web: 0 fájl és 1 adatkötet visszaállítva (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and 5/5 files byte-identical |
R-355 — paperless-ngx:
| before (2026-08-21) | „A(z) paperless-ngx: 0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." over a live 72-table PostgreSQL, with no undo copy taken at all |
| after (0.218.0, 09:57) | „A(z) paperless-ngx: 0 fájl és 3 adatkötet és az adatbázis visszaállítva (mentés: 2026-08-22 09:52) — az alkalmazás újraindult." |
A third sentence now exists for the case that had no honest wording — an app that HAS a database
whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak VAN adatbázisa, de a mentés nem
tartalmazott adatbázis-mentést, ezért az adatbázis NEM állt vissza. A visszaállítás előtti állapot
mentése megvan: <undo>". That case previously printed the same confident „nincs adatbázisa".
6. THE HARDWARE WALK
Method: endpoint-level — the exact endpoints the dashboard's JS calls (/api/backup/run,
/backup/offbox/run, /backup/offbox/restore, /backup/offbox/reconstitute), with no shell inside
the machine for any step of the walk. Planting the fixture and reading the result back used a shell
and is setup/verification, not the walk. No browser exists on DooPlex.
The fixture, hashed before anything ran:
7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e SENTINEL.txt
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5 binary-512k.bin
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4 nested/őszibarack.md
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5 plain.txt
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58 árvíztűrő-tükörfúrógép.txt
accented names as RAW BYTES (UTF-8 NFC):
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
The comparator was proved able to convict first: one byte flipped at offset 300 000 of
binary-512k.bin (3b → 00) → binary-512k.bin: FAILED, rc 1, the other four OK; the unmodified
set rc 0. Mutant discarded.
Results
| scenario | result |
|---|---|
| 1-A the dump is inside the unit, and inside the off-site snapshot | PASS. 0.217.0 09:37: db_dumps = None, dump refreshed into the phantom dir. 0.218.0 09:45: db_dumps = ['paperless-ngx-postgres.sql'] (312 669 B) in the app's own unit, and restic ls shows it inside the snapshot for the first time |
| 1-B an undo copy taken and verified before anything stops; the message names the database | PASS. Before: find /mnt -name "pre-restore-*paperless*" → 0. After: pre-restore-20260822T075658Z-paperless-ngx-postgres.sql, 312 957 B, in the app's own unit |
| 1-C the undo cannot be taken → the whole restore refuses, nothing changed | PASS, on the app that could never reach this guard before. ok=false; the marker kept its mutation (51d37b0f…) and StartedAt was unchanged — the app was never stopped |
| 1-D every other app byte-identical | PASS. romm-mariadb.sql, kimai-mariadb.sql unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created |
| 2-A/B the volume comes back and the message says so | PASS. calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database |
| 2-C a snapshot with no volume archives is unchanged | PASS — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected |
| 2-D the live recovery unit is never written | PASS — fingerprinted across the whole operation (see red-proof 6) |
| 2-E a failed replay is reported as a failure naming the volume | PASS (unit test; the live path returns the same error) |
All 15 containers healthy after the walk.
7. RED-PROOFS — seven, each mutation asserted applied and reverted
| # | mutation | outcome |
|---|---|---|
| 1 | resolveStackName ignores the compose label |
FAIL as required. = "paperless", want "paperless-ngx", and the consequence printed both paths — written to …/primary/paperless/db-dumps vs read …/primary/paperless-ngx/db-dumps. The empty database record, returned |
| 2 | outcome message back to the counter-only predicate | FAIL as required, reproducing the 21 August sentence verbatim: "A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." |
| 3 | the fail-closed refusal removed from writeSafetyDump |
FAIL as required — a restore was seen proceeding with no undo copy |
| 4 | the volume leg removed | FAIL as required. VolumesReplayed = 0 and the replay never called — last night's silent loss, returned |
| 5 | the volume count dropped from the message | FAIL as required, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim |
| 6 | the live-unit guard removed | PASSED FIRST — a defect in MY TEST, not the code. The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit root. Widened to the whole unit (excluding only the documented undo copies) it convicts: PLACED:17603ba1… appears. The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory |
| 7 | the volume-replay error swallowed | FAIL as required — a partial replay reporting success |
Green gate: go build ./... && go vet ./... && go test ./... — 28 packages ok, 0 FAIL lines.
Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red
until the bake (§8) — correctly, and never bypassed.
(Instrument note: go test ./... | grep -vE '^ok'; echo rc=$? reports the grep's exit code, which
is 1 when every test passed and nothing was left to print. The verdict above is from an explicit
FAIL-line count, not from that.)
8. WHAT THIS DOES NOT FIX — the blocker for the apps that need it most
The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive (R-356), telling the customer that a running app „nincs telepítve" and to reinstall it "to the same place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset is a named volume, so R-354's fix cannot reach them until R-356 is closed.
That is why the live confirmation used calibre-web and paperless-ngx: they declare a drive and can
actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing
worth doing.
9. REGISTER
- R-354 — CLOSED, shipped + proven-live (v0.218.0).
- R-355 — CLOSED, shipped + proven-live (v0.218.0).
- R-367 — OPENED (LOW): dumps already written under the wrong name are stranded; adoptable by hand, deliberately not automatic, never on an upgrade path.
- Ceiling moved: R-366 → R-367.
10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten
R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives, awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359, R-361–R-363, R-365, R-366.
11. COMMITS, VERSIONS, CI
felhom-controller |
5ce3a44 the fix + tests; 2da259a README + CONTEXT |
| image | gitea.dooplex.hu/admin/felhom-controller:0.218.0 (150M) |
| deployed | demo-hp guest 9201 — …:0.218.0 Up (healthy) |
demo-felhom |
untouched this session, still 0.217.0 |
12. BAKE
RUNBOOK-manual-build.md §4.0/§4.1, in the drill VM. The golden-currency gate went red the moment
v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.
GOLDEN_VERSION |
0.218.0 |
GOLDEN_SHA256 |
8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b |
| package | …/generic/felhom-golden/0.218.0/golden.tar.zst, 657 026 013 B |
| template | debian-13-standard_13.6-1_amd64.tar.zst — listed fresh, not assumed |
Pass markers, each checked, with the negative controls:
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
upload OK (HTTP 201) : 1 -> pre-delete HTTP 404 (404/204 expected)
excluding : 0 <- negative control
FATAL : 0 <- negative control
Token hygiene. Copied file→file and read by a runner script inside the VM, so it never crossed a
shell or a unit property: systemctl show golden-bake … | grep -c -F "$(cat …)" → 0. The leak
grep on the committed log was itself proven before its zero was believed — token appended to a
throwaway copy, grepped (1), copy shredded, then the committed log's 0 accepted.
Teardown: pct destroy 9100 --purge; token, runner, build script and in-VM log shredded after
the log was copied out; VM powered off; qemu exited; qemu-img snapshot -a virgin restored.
Evidence: documentation/tests/golden-0.218.0-2026-08-22/.
13. STOP — THE APPROVAL
Nothing below has been changed. Vouching and the floor are yours.
The three Day-0 values, each with BOTH checks
| field | set to | downloadable | selectable |
|---|---|---|---|
golden_version |
0.218.0 — sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b |
✔ HTTP 200, 657 026 013 B, and the downloaded bytes hash to exactly the bake's reported sha | ✔ offered in the hub dropdown with data-sha=8e427869d13eafb7… |
agent_version |
0.130.0 — sha a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3 |
✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's data-sha exactly |
✔ already the SELECTED option — no change needed |
min_agent |
0.129.0 | — | ✔ already set to 0.129.0 — no change needed |
So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut. The
three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller
being baked declares MinAgent: 0.129.0 (unchanged), the vouched agent is already 0.130.0,
and 0.130.0 ≥ 0.129.0 — so agent_version and min_agent are already correct and only
golden_version moves. Setting min_agent above the vouched agent is the R-216 shape; it is not
happening here.
If the hub refuses the save, the guard is working. The R-120 gate on that POST
(hub/internal/web/configs.go:1165) refuses a golden older than the newest controller the fleet
reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.
Vouching is reversible: re-select 0.217.0 and Save. The old package is never deleted by a bake.
The one line on the floor
Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to.
demo-hp already runs 0.218.0 (deployed directly for the live proof). demo-felhom is still on
0.217.0 and will not move until the floor is raised, so today it still has both defects. I did not
raise it; that is your call, as is vouching.
14. WHAT WAS DROPPED, AND OBSERVATIONS
Dropped: nothing. Both parts completed in the required order, with the live walk and all seven red-proofs. The session did not run short.
Observations, noticed and not acted on:
demo-felhomwas not touched and still runs 0.217.0 — so the fleet is deliberately non-uniform until the floor moves. Its two preserved fixtures were not involved in any step of this session.- The
.fabexport has the same R-355 blind spot and is fixed by the same change —export.go:600filters on the samedb.StackName. Not separately verified live; the unit tests cover the resolver that feeds it. paperless-ngx's orphan directory is now a permanent fixture ondemo-hpuntil R-367 is decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for leaving it alone for now.- The safety dump still overwrites the unit's own DB dump before renaming it (R-361, out of scope)
— visible in this session's own evidence, where
romm'sdb-dumps/holds both apre-restore-*and the regular dump.