Files
felhom.eu/documentation/audits/REPORT-v0.218.0-r354-r355-2026-08-22.md
T
admin 091a4b7444
gates / gates (push) Successful in 16s
Correct the placement mis-framing, and file what we wrote down and never filed (R-368..R-375)
Documentation and survey only. No code, no machine contacted.

THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a
choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast
storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has
been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and
R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands.
R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so
reading it as "not installed" misreads a correct configuration.

The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two
volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy
together is REAL and is what the other tiers exist for. A full data volume stopping the OS is
NOT real and was the overstated one.

THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed.
Its positive control convicted the sweep itself twice before it convicted the corpus - markdown
bold broke the strongest pattern, and the reporter re-searched a truncated line - both false
zeros of the exact class being hunted, and together worth 2 of the 14.

THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122,
M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of
truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day
after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH).

Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373
(20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy
time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go
and never the templates. R-370 records the process failure and is closed by the template change.

PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and
say what it says (with a file->area map and the test "is this something we chose?"), and an
enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count.

Ceiling R-367 -> R-375.
2026-08-22 11:18:08 +02:00

22 KiB
Raw Blame History

REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)

Controller v0.217.0 → v0.218.0. Two fixes, both found by watching a machine on the night of 2026-08-21, both confirmed the same way. Part 1 first, because it is the only place in the product where one customer action causes permanent total loss.

The drill that found them is preserved at documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md; its evidence is in documentation/audits/DRILL-backup-truth-2026-08-21/evidence/.


0. Baselines, re-established — not trusted from the sheet

item expected confirmed
controller 0.217.0 → 0.218.0 ✔ repo head v0.217.0; live on demo-hp …:0.217.0
agent 0.130.0 ✔ repo head and felhom-agent --version on the box
golden vouched 0.217.0 ✔ <option value="0.217.0" … selected>
floor 0.217.0 ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)"
min agent 0.129.0 ✔
register ceiling R-366 ✔ (now R-367)
ssh hp → 192.168.0.104 works ✔
clean trees, all four repos HEAD == origin/main, 0 dirty ✔

1. PART 1 — where the wrong name comes from

It comes from a guess, made in a place where an answer was available.

deriveStackName (controller/internal/appbackup/dbdump.go:770-798) is handed a container name and asked which app owns it. For paperless-postgres it strips the role suffix to paperless; finds paperless is not a deployed stack; finds the container name is not one either; finds no known stack is a prefix of it (paperless-ngx is not a prefix of paperless-postgres) — and then returns the unresolved candidate anyway, on its last line, silently.

Every place that value is used, and what each did with it:

consumer file:line effect
dump directory backup/backup.go:482,506 wrote to …/primary/**paperless**/db-dumps/ on the SYSTEM drive
dump filename appbackup/dbdump.go:200-212 paperless-postgres.sql, not paperless-ngx-postgres.sql
unit assembler backup/recovery_unit.go:131 reads …/primary/**paperless-ngx**/db-dumps → finds nothing → db_dumps: null
off-site collector backup/offbox_capture.go:32 resolves from the unit path → the orphan is invisible to it
restore's DB leg backup/restore_db.go:69 db.StackName != stackName → never replays
safety-dump filter backup/offbox_reconstitute.go:134 mine empty → hasDB=false → no undo copy, and the refusal is never reached
.fab export appexport/export.go:600 same filter — the bundle carries no database either
the "has a DB" flag web/handlers.go:1245 the app displays as having no database

Corroboration from the code itself: ListDumpFiles (dbdump.go:562) parses a dump filename under the comment "Parse stack name and DB type from filename: paperless-ngx-postgres.sql" — the reader was written for a name the writer never produced.

Controller or catalogue? — THE CONTROLLER. No halt.

Every container the controller starts already carries com.docker.compose.project, read live:

paperless-postgres   paperless-ngx    kimai-db   kimai      romm-db   romm

And that label is the stack name BY CONSTRUCTION, not by luck: composeExecCustomEnv (internal/stacks/manager.go:1218-1230) runs compose with cmd.Dir set to /opt/docker/stacks/<stack> and never passes -p, so compose derives the project from that directory. Renaming the container in the catalogue would have fixed this one app and left the guessing for the next one. No catalogue change was made.


2. PART 1.2 — the sweep, proven before its answer was trusted

step result
scratch copy, unplanted MISMATCHES: 1 (rc 1) — the known one
plant a second mismatch (kimai-db → timetrack-db) MISMATCHES: 2, convicted by name: stack=kimai container=timetrack-db -> derived=timetrack
plant removed back to 1
every mismatch removed MISMATCHES: 0, rc 0 — "1" is not a stuck value
scratch discarded kimai/docker-compose.yml md5-identical to the real catalogue

The real count: 1 affected app of 53, from 15 DB containers checked — paperless-ngx.


3. PART 1 — the dumps already written under the wrong name

They stay, untouched, and nothing will ever collect them. /mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql, 312 381 B, last written 07:38 UTC by the final pre-fix cycle — verified still present after the fix.

Nothing deletes them by design: the F5 stale-primary prune (backup/backup.go:1248) skips any directory that is not a deployed app, under the guard "an undeployed app's last backup is still its restore point".

They can be adopted, by hand: move the file to …/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql and it becomes a readable restore point. Deliberately not automatic — it would sit beside volume tars from a different time (an incoherent pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. Filed as R-367.


4. PART 2 — the mechanism, and the comment corrected

ReconstituteFromOffsite skipped every placement flagged as the unit (offbox_reconstitute.go:344), and the volume archives live inside the unit. A search of the whole off-site path found no volume-restore call; the only two callers of restoreDockerVolumes were restore.go:67 and restore_unit.go:255, both local.

The comment as it stood:

"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the scratch unit instead, so nothing is lost by skipping it."

Which half still holds: the FIRST. The live unit is the local restore path's own source and must never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit across the operation to keep it true.

The second half was false and is corrected, not left. True of the database dump, false of the volume archives, which live in the same unit and were replayed by nothing at all. Skipping the placement is correct; treating the skip as harmless was not — and that sentence is exactly why a reader would not look.

The fix: restoreDockerVolumesFrom(stackName, dumpDir) (int, error) — the local path's own replay with an explicit directory, the same shape reimportDBDumpsFrom already had beside reimportDBDumps. One implementation, two callers. Volumes replay before the database (a logical dump must still win over a volume-tar copy of the same database) and inside the stopped window (Docker will not replace a volume a container holds).


5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM

R-354 — calibre-web, identical fixture, identical steps, both on demo-hp:

before (0.217.0, 09:41) ok=true — „A(z) calibre-web: 0 fájl visszaállítva (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and ls: /vol/R354-2026-08-22: No such file or directory
after (0.218.0, 09:56) ok=true — „A(z) calibre-web: 0 fájl és 1 adatkötet visszaállítva (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and 5/5 files byte-identical

R-355 — paperless-ngx:

before (2026-08-21) „A(z) paperless-ngx: 0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." over a live 72-table PostgreSQL, with no undo copy taken at all
after (0.218.0, 09:57) „A(z) paperless-ngx: 0 fájl és 3 adatkötet és az adatbázis visszaállítva (mentés: 2026-08-22 09:52) — az alkalmazás újraindult."

A third sentence now exists for the case that had no honest wording — an app that HAS a database whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak VAN adatbázisa, de a mentés nem tartalmazott adatbázis-mentést, ezért az adatbázis NEM állt vissza. A visszaállítás előtti állapot mentése megvan: <undo>". That case previously printed the same confident „nincs adatbázisa".


6. THE HARDWARE WALK

Method: endpoint-level — the exact endpoints the dashboard's JS calls (/api/backup/run, /backup/offbox/run, /backup/offbox/restore, /backup/offbox/reconstitute), with no shell inside the machine for any step of the walk. Planting the fixture and reading the result back used a shell and is setup/verification, not the walk. No browser exists on DooPlex.

The fixture, hashed before anything ran:

7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e  SENTINEL.txt
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5  binary-512k.bin
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4  nested/őszibarack.md
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5  plain.txt
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58  árvíztűrő-tükörfúrógép.txt

accented names as RAW BYTES (UTF-8 NFC):
  árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
  őszibarack.md              = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64

The comparator was proved able to convict first: one byte flipped at offset 300 000 of binary-512k.bin (3b → 00) → binary-512k.bin: FAILED, rc 1, the other four OK; the unmodified set rc 0. Mutant discarded.

Results

scenario result
1-A the dump is inside the unit, and inside the off-site snapshot PASS. 0.217.0 09:37: db_dumps = None, dump refreshed into the phantom dir. 0.218.0 09:45: db_dumps = ['paperless-ngx-postgres.sql'] (312 669 B) in the app's own unit, and restic ls shows it inside the snapshot for the first time
1-B an undo copy taken and verified before anything stops; the message names the database PASS. Before: find /mnt -name "pre-restore-*paperless*" → 0. After: pre-restore-20260822T075658Z-paperless-ngx-postgres.sql, 312 957 B, in the app's own unit
1-C the undo cannot be taken → the whole restore refuses, nothing changed PASS, on the app that could never reach this guard before. ok=false; the marker kept its mutation (51d37b0f…) and StartedAt was unchanged — the app was never stopped
1-D every other app byte-identical PASS. romm-mariadb.sql, kimai-mariadb.sql unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created
2-A/B the volume comes back and the message says so PASS. calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database
2-C a snapshot with no volume archives is unchanged PASS — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected
2-D the live recovery unit is never written PASS — fingerprinted across the whole operation (see red-proof 6)
2-E a failed replay is reported as a failure naming the volume PASS (unit test; the live path returns the same error)

All 15 containers healthy after the walk.


7. RED-PROOFS — seven, each mutation asserted applied and reverted

# mutation outcome
1 resolveStackName ignores the compose label FAIL as required. = "paperless", want "paperless-ngx", and the consequence printed both paths — written to …/primary/paperless/db-dumps vs read …/primary/paperless-ngx/db-dumps. The empty database record, returned
2 outcome message back to the counter-only predicate FAIL as required, reproducing the 21 August sentence verbatim: "A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."
3 the fail-closed refusal removed from writeSafetyDump FAIL as required — a restore was seen proceeding with no undo copy
4 the volume leg removed FAIL as required. VolumesReplayed = 0 and the replay never called — last night's silent loss, returned
5 the volume count dropped from the message FAIL as required, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim
6 the live-unit guard removed PASSED FIRST — a defect in MY TEST, not the code. The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit root. Widened to the whole unit (excluding only the documented undo copies) it convicts: PLACED:17603ba1… appears. The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory
7 the volume-replay error swallowed FAIL as required — a partial replay reporting success

Green gate: go build ./... && go vet ./... && go test ./... — 28 packages ok, 0 FAIL lines. Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red until the bake (§8) — correctly, and never bypassed.

(Instrument note: go test ./... | grep -vE '^ok'; echo rc=$? reports the grep's exit code, which is 1 when every test passed and nothing was left to print. The verdict above is from an explicit FAIL-line count, not from that.)


8. WHAT THIS DOES NOT FIX — the blocker for the apps that need it most

The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive (R-356), telling the customer that a running app „nincs telepítve" and to reinstall it "to the same place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset is a named volume, so R-354's fix cannot reach them until R-356 is closed.

That is why the live confirmation used calibre-web and paperless-ngx: they declare a drive and can actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing worth doing.


9. REGISTER

  • R-354 — CLOSED, shipped + proven-live (v0.218.0).
  • R-355 — CLOSED, shipped + proven-live (v0.218.0).
  • R-367 — OPENED (LOW): dumps already written under the wrong name are stranded; adoptable by hand, deliberately not automatic, never on an upgrade path.
  • Ceiling moved: R-366 → R-367.

10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten

R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives, awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359, R-361–R-363, R-365, R-366.


11. COMMITS, VERSIONS, CI

felhom-controller 5ce3a44 the fix + tests; 2da259a README + CONTEXT
image gitea.dooplex.hu/admin/felhom-controller:0.218.0 (150M)
deployed demo-hp guest 9201 — …:0.218.0 Up (healthy)
demo-felhom untouched this session, still 0.217.0

12. BAKE

RUNBOOK-manual-build.md §4.0/§4.1, in the drill VM. The golden-currency gate went red the moment v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.

GOLDEN_VERSION 0.218.0
GOLDEN_SHA256 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b
package …/generic/felhom-golden/0.218.0/golden.tar.zst, 657 026 013 B
template debian-13-standard_13.6-1_amd64.tar.zst — listed fresh, not assumed

Pass markers, each checked, with the negative controls:

docker OK (overlay2      : 1     ->  "  docker OK (overlay2; data-root /var/lib/docker)"
including mount point    : 2     ->  rootfs ('/') and mp0 ('/var/lib/felhom')   [there is no mp1]
upload OK (HTTP 201)     : 1     ->  pre-delete HTTP 404 (404/204 expected)
excluding                : 0     <- negative control
FATAL                    : 0     <- negative control

Token hygiene. Copied file→file and read by a runner script inside the VM, so it never crossed a shell or a unit property: systemctl show golden-bake … | grep -c -F "$(cat …)" → 0. The leak grep on the committed log was itself proven before its zero was believed — token appended to a throwaway copy, grepped (1), copy shredded, then the committed log's 0 accepted.

Teardown: pct destroy 9100 --purge; token, runner, build script and in-VM log shredded after the log was copied out; VM powered off; qemu exited; qemu-img snapshot -a virgin restored.

Evidence: documentation/tests/golden-0.218.0-2026-08-22/.


13. STOP — THE APPROVAL

Nothing below has been changed. Vouching and the floor are yours.

The three Day-0 values, each with BOTH checks

field set to downloadable selectable
golden_version 0.218.0 — sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b ✔ HTTP 200, 657 026 013 B, and the downloaded bytes hash to exactly the bake's reported sha ✔ offered in the hub dropdown with data-sha=8e427869d13eafb7…
agent_version 0.130.0 — sha a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3 ✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's data-sha exactly ✔ already the SELECTED option — no change needed
min_agent 0.129.0 — ✔ already set to 0.129.0 — no change needed

So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut. The three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller being baked declares MinAgent: 0.129.0 (unchanged), the vouched agent is already 0.130.0, and 0.130.0 ≥ 0.129.0 — so agent_version and min_agent are already correct and only golden_version moves. Setting min_agent above the vouched agent is the R-216 shape; it is not happening here.

If the hub refuses the save, the guard is working. The R-120 gate on that POST (hub/internal/web/configs.go:1165) refuses a golden older than the newest controller the fleet reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.

Vouching is reversible: re-select 0.217.0 and Save. The old package is never deleted by a bake.

The one line on the floor

Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to. demo-hp already runs 0.218.0 (deployed directly for the live proof). demo-felhom is still on 0.217.0 and will not move until the floor is raised, so today it still has both defects. I did not raise it; that is your call, as is vouching.


14. WHAT WAS DROPPED, AND OBSERVATIONS

Dropped: nothing. Both parts completed in the required order, with the live walk and all seven red-proofs. The session did not run short.

Observations, noticed and not acted on:

  • demo-felhom was not touched and still runs 0.217.0 — so the fleet is deliberately non-uniform until the floor moves. Its two preserved fixtures were not involved in any step of this session.
  • The .fab export has the same R-355 blind spot and is fixed by the same change — export.go:600 filters on the same db.StackName. Not separately verified live; the unit tests cover the resolver that feeds it.
  • paperless-ngx's orphan directory is now a permanent fixture on demo-hp until R-367 is decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for leaving it alone for now.
  • The safety dump still overwrites the unit's own DB dump before renaming it (R-361, out of scope) — visible in this session's own evidence, where romm's db-dumps/ holds both a pre-restore-* and the regular dump.

15. THE PICTURE — where-felhom-stands brought up to date (same day, after the fixes)

Regenerated from where-felhom-stands.yaml (the HTML is a build product; the YAML is the source). Eight claims re-checked, three moved — all downward.

claim was now why
backup.offsite walked partial "18 snapshots, daily, unbroken" was true on 2026-08-09 and false by 2026-08-21 — the next snapshot after it was put there by hand, twelve days later, and nothing reported the gap
fail.wiped-reinstalled.data walked partial a real reinstall orphans BOTH off-premises tiers: restic silently for 12 days (R-193), and the pre-reinstall PBS archives cannot be opened by the rebuilt box at all (R-366)
backup.fill-warning walked partial the warning fires correctly, but once a day — a filesystem filling at 03:31 is unannounced for ~24 h, and it was watched staying silent while a volume sat at 99%

Held, with their notes rewritten: backup.tier1 and recover.byte-identical now carry the R-355/R-354 story and their fixes; backup.restore-proof stays grey for a sharper reason (its archives are orphaned, not merely untested); backup.sikeres gained two fresh instances; fail.customer-self-restore records that R-356 blocks 40 of 53 apps regardless of who is driving.

unproven.py --summary moved, and the checklist asks that it be said: NOT WALKED went from 32 of 55 to 35 of 55. That is the three downgrades above. Nothing was raised — the dataset's own rule forbids raising a status here, and nothing needed it.

A defect found in the page's own toolchain

The rendered header cited the wrong commits and would have gone on doing so. render_stands.py parsed verified_on from the YAML but had the three commit shas hardcoded, so the page printed the 2026-08-09 commits while the dataset said 2026-08-22. That is precisely the stale-build-product failure the renderer's own docstring says it exists to prevent. Parsed now, literals gone from both the script and the output, checked. The count beside it also read "15 status(es) moved in that pass" when 15 was every recorded move ever accumulated; it now separates the two numbers.

check_stands.py was proven able to convict before its OK was accepted: a claim marked missing flipped to walked in a scratch copy fired rule 5 by name (use.dlna: status 'walked' but NO evidence document cited), and the real file still passed. Gates all OK; CI id=380 / run_number=248, green.