Files
felhom.eu/REPORT.md
T
admin 877fcd2a38
gates / gates (push) Successful in 16s
R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.

R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.

R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.

Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.

R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.

Ceiling R-366 -> R-367.
2026-08-22 10:11:46 +02:00

349 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)
**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of
2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product
where one customer action causes permanent total loss.**
The drill that found them is preserved at
`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
---
## 0. Baselines, re-established — not trusted from the sheet
| item | expected | confirmed |
|---|---|---|
| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` |
| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box |
| golden vouched | 0.217.0 | ✔ `<option value="0.217.0" … selected>` |
| floor | 0.217.0 | ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)" |
| min agent | 0.129.0 | ✔ |
| register ceiling | R-366 | ✔ (now **R-367**) |
| `ssh hp` → 192.168.0.104 | works | ✔ |
| clean trees, all four repos | `HEAD == origin/main`, 0 dirty | ✔ |
---
## 1. PART 1 — where the wrong name comes from
**It comes from a guess, made in a place where an answer was available.**
`deriveStackName` (`controller/internal/appbackup/dbdump.go:770-798`) is handed a container name and
asked which app owns it. For `paperless-postgres` it strips the role suffix to `paperless`; finds
`paperless` is **not** a deployed stack; finds the container name is not one either; finds no known
stack is a prefix of it (`paperless-ngx` is not a prefix of `paperless-postgres`) — **and then returns
the unresolved candidate anyway**, on its last line, silently.
**Every place that value is used, and what each did with it:**
| consumer | file:line | effect |
|---|---|---|
| dump directory | `backup/backup.go:482,506` | wrote to `…/primary/**paperless**/db-dumps/` on the SYSTEM drive |
| dump filename | `appbackup/dbdump.go:200-212` | `paperless-postgres.sql`, not `paperless-ngx-postgres.sql` |
| unit assembler | `backup/recovery_unit.go:131` | reads `…/primary/**paperless-ngx**/db-dumps` → finds nothing → `db_dumps: null` |
| off-site collector | `backup/offbox_capture.go:32` | resolves from the unit path → the orphan is invisible to it |
| restore's DB leg | `backup/restore_db.go:69` | `db.StackName != stackName` → never replays |
| **safety-dump filter** | `backup/offbox_reconstitute.go:134` | `mine` empty → `hasDB=false` → **no undo copy, and the refusal is never reached** |
| `.fab` export | `appexport/export.go:600` | same filter — the bundle carries no database either |
| the "has a DB" flag | `web/handlers.go:1245` | the app displays as having no database |
**Corroboration from the code itself:** `ListDumpFiles` (`dbdump.go:562`) parses a dump filename under
the comment *"Parse stack name and DB type from filename: `paperless-ngx-postgres.sql`"* — the reader
was written for a name the writer never produced.
### Controller or catalogue? — **THE CONTROLLER. No halt.**
Every container the controller starts already carries `com.docker.compose.project`, read live:
```
paperless-postgres paperless-ngx kimai-db kimai romm-db romm
```
**And that label is the stack name BY CONSTRUCTION, not by luck:** `composeExecCustomEnv`
(`internal/stacks/manager.go:1218-1230`) runs compose with `cmd.Dir` set to
`/opt/docker/stacks/<stack>` and **never passes `-p`**, so compose derives the project from that
directory. Renaming the container in the catalogue would have fixed this one app and left the guessing
for the next one. **No catalogue change was made.**
---
## 2. PART 1.2 — the sweep, proven before its answer was trusted
| step | result |
|---|---|
| scratch copy, unplanted | `MISMATCHES: 1` (rc 1) — the known one |
| **plant a second mismatch** (`kimai-db` → `timetrack-db`) | **`MISMATCHES: 2`**, convicted by name: `stack=kimai container=timetrack-db -> derived=timetrack` |
| plant removed | back to `1` |
| **every mismatch removed** | **`MISMATCHES: 0`, rc 0** — "1" is not a stuck value |
| scratch discarded | `kimai/docker-compose.yml` md5-identical to the real catalogue |
**The real count: 1 affected app of 53**, from 15 DB containers checked — `paperless-ngx`.
---
## 3. PART 1 — the dumps already written under the wrong name
**They stay, untouched, and nothing will ever collect them.**
`/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql`, 312 381 B,
last written 07:38 UTC by the final pre-fix cycle — **verified still present after the fix.**
Nothing deletes them by design: the F5 stale-primary prune (`backup/backup.go:1248`) skips any
directory that is not a deployed app, under the guard *"an undeployed app's last backup is still its
restore point"*.
**They can be adopted, by hand:** move the file to
`…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point.
**Deliberately not automatic** — it would sit beside volume tars from a different time (an incoherent
pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a
customer's data on upgrade is a migration, not a fix. **Filed as R-367.**
---
## 4. PART 2 — the mechanism, and the comment corrected
`ReconstituteFromOffsite` skipped every placement flagged as the unit (`offbox_reconstitute.go:344`),
and the volume archives live **inside** the unit. A search of the whole off-site path found no
volume-restore call; the only two callers of `restoreDockerVolumes` were `restore.go:67` and
`restore_unit.go:255`, both local.
**The comment as it stood:**
> *"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and
> clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the
> scratch unit instead, so nothing is lost by skipping it."*
**Which half still holds: the FIRST.** The live unit is the local restore path's own source and must
never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit
across the operation to keep it true.
**The second half was false and is corrected, not left.** True of the *database* dump, false of the
*volume* archives, which live in the same unit and were replayed by nothing at all. Skipping the
placement is correct; treating the skip as harmless was not — and that sentence is exactly why a
reader would not look.
**The fix:** `restoreDockerVolumesFrom(stackName, dumpDir) (int, error)` — the local path's own replay
with an explicit directory, the same shape `reimportDBDumpsFrom` already had beside `reimportDBDumps`.
**One implementation, two callers.** Volumes replay **before** the database (a logical dump must still
win over a volume-tar copy of the same database) and **inside the stopped window** (Docker will not
replace a volume a container holds).
---
## 5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM
**R-354 — calibre-web, identical fixture, identical steps, both on `demo-hp`:**
| | |
|---|---|
| **before** (0.217.0, 09:41) | `ok=true` — „A(z) calibre-web: **0 fájl visszaállítva** (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and `ls: /vol/R354-2026-08-22: No such file or directory` |
| **after** (0.218.0, 09:56) | `ok=true` — „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and **5/5 files byte-identical** |
**R-355 — paperless-ngx:**
| | |
|---|---|
| **before** (2026-08-21) | „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" over a live 72-table PostgreSQL, with **no undo copy taken at all** |
| **after** (0.218.0, 09:57) | „A(z) paperless-ngx: **0 fájl és 3 adatkötet és az adatbázis visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult." |
**A third sentence now exists for the case that had no honest wording** — an app that HAS a database
whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak **VAN adatbázisa**, de a mentés nem
tartalmazott adatbázis-mentést, ezért az adatbázis **NEM állt vissza**. A visszaállítás előtti állapot
mentése megvan: `<undo>`". That case previously printed the same confident „nincs adatbázisa".
---
## 6. THE HARDWARE WALK
**Method: endpoint-level — the exact endpoints the dashboard's JS calls (`/api/backup/run`,
`/backup/offbox/run`, `/backup/offbox/restore`, `/backup/offbox/reconstitute`), with no shell inside
the machine for any step of the walk.** Planting the fixture and reading the result back used a shell
and is setup/verification, not the walk. No browser exists on DooPlex.
**The fixture, hashed before anything ran:**
```
7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e SENTINEL.txt
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5 binary-512k.bin
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4 nested/őszibarack.md
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5 plain.txt
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58 árvíztűrő-tükörfúrógép.txt
accented names as RAW BYTES (UTF-8 NFC):
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
```
**The comparator was proved able to convict first:** one byte flipped at offset 300 000 of
`binary-512k.bin` (`3b` → `00`) → `binary-512k.bin: FAILED`, rc 1, the other four `OK`; the unmodified
set rc 0. Mutant discarded.
### Results
| scenario | result |
|---|---|
| **1-A** the dump is inside the unit, and inside the off-site snapshot | **PASS.** 0.217.0 09:37: `db_dumps = None`, dump refreshed into the phantom dir. 0.218.0 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` (312 669 B) in the app's own unit, and `restic ls` shows it inside the snapshot **for the first time** |
| **1-B** an undo copy taken and verified before anything stops; the message names the database | **PASS.** Before: `find /mnt -name "pre-restore-*paperless*"` → **0**. After: `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql`, 312 957 B, in the app's own unit |
| **1-C** the undo cannot be taken → the whole restore refuses, nothing changed | **PASS**, on the app that could never reach this guard before. `ok=false`; the marker kept its mutation (`51d37b0f…`) and `StartedAt` was unchanged — **the app was never stopped** |
| **1-D** every other app byte-identical | **PASS.** `romm-mariadb.sql`, `kimai-mariadb.sql` unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created |
| **2-A/B** the volume comes back and the message says so | **PASS.** calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database |
| **2-C** a snapshot with no volume archives is unchanged | **PASS** — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected |
| **2-D** the live recovery unit is never written | **PASS** — fingerprinted across the whole operation (see red-proof 6) |
| **2-E** a failed replay is reported as a failure naming the volume | **PASS** (unit test; the live path returns the same error) |
All 15 containers healthy after the walk.
---
## 7. RED-PROOFS — seven, each mutation asserted applied and reverted
| # | mutation | outcome |
|---|---|---|
| 1 | `resolveStackName` ignores the compose label | **FAIL as required.** `= "paperless", want "paperless-ngx"`, and the consequence printed both paths — written to `…/primary/paperless/db-dumps` vs read `…/primary/paperless-ngx/db-dumps`. **The empty database record, returned** |
| 2 | outcome message back to the counter-only predicate | **FAIL as required**, reproducing the 21 August sentence verbatim: `"A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."` |
| 3 | the fail-closed refusal removed from `writeSafetyDump` | **FAIL as required** — *a restore was seen proceeding with no undo copy* |
| 4 | the volume leg removed | **FAIL as required.** `VolumesReplayed = 0` and the replay never called — last night's silent loss, returned |
| 5 | the volume count dropped from the message | **FAIL as required**, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim |
| 6 | the live-unit guard removed | **PASSED FIRST — a defect in MY TEST, not the code.** The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit **root**. Widened to the whole unit (excluding only the documented undo copies) it convicts: `PLACED:17603ba1…` appears. **The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory** |
| 7 | the volume-replay error swallowed | **FAIL as required** — a partial replay reporting success |
**Green gate:** `go build ./... && go vet ./... && go test ./...` — **28 packages ok, 0 FAIL lines.**
Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red
until the bake (§8) — correctly, and never bypassed.
*(Instrument note: `go test ./... | grep -vE '^ok'; echo rc=$?` reports the **grep's** exit code, which
is 1 when every test passed and nothing was left to print. The verdict above is from an explicit
`FAIL`-line count, not from that.)*
---
## 8. WHAT THIS DOES **NOT** FIX — the blocker for the apps that need it most
**The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive
(R-356)**, telling the customer that a running app „nincs telepítve" and to reinstall it "to the same
place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset
is a named volume, **so R-354's fix cannot reach them until R-356 is closed.**
That is why the live confirmation used `calibre-web` and `paperless-ngx`: they declare a drive and can
actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing
worth doing.
---
## 9. REGISTER
- **R-354 — CLOSED**, shipped + proven-live (v0.218.0).
- **R-355 — CLOSED**, shipped + proven-live (v0.218.0).
- **R-367 — OPENED (LOW)**: dumps already written under the wrong name are stranded; adoptable by hand,
deliberately not automatic, never on an upgrade path.
- **Ceiling moved: R-366 → R-367.**
---
## 10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten
R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives,
awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification
copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359,
R-361–R-363, R-365, R-366.
---
## 11. COMMITS, VERSIONS, CI
| | |
|---|---|
| `felhom-controller` | `5ce3a44` the fix + tests; `2da259a` README + CONTEXT |
| image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` (150M) |
| deployed | `demo-hp` guest 9201 — `…:0.218.0 Up (healthy)` |
| `demo-felhom` | **untouched this session**, still 0.217.0 |
---
## 12. BAKE
`RUNBOOK-manual-build.md` §4.0/§4.1, in the drill VM. **The golden-currency gate went red the moment
v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.**
| | |
|---|---|
| `GOLDEN_VERSION` | **0.218.0** |
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
| package | `…/generic/felhom-golden/0.218.0/golden.tar.zst`, 657 026 013 B |
| template | `debian-13-standard_13.6-1_amd64.tar.zst` — listed fresh, not assumed |
**Pass markers, each checked, with the negative controls:**
```
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
upload OK (HTTP 201) : 1 -> pre-delete HTTP 404 (404/204 expected)
excluding : 0 <- negative control
FATAL : 0 <- negative control
```
**Token hygiene.** Copied file→file and read by a runner script inside the VM, so it never crossed a
shell or a unit property: `systemctl show golden-bake … | grep -c -F "$(cat …)"` → **0**. **The leak
grep on the committed log was itself proven before its zero was believed** — token appended to a
throwaway copy, grepped (**1**), copy shredded, then the committed log's **0** accepted.
**Teardown:** `pct destroy 9100 --purge`; token, runner, build script and in-VM log shredded **after**
the log was copied out; VM powered off; qemu exited; **`qemu-img snapshot -a virgin` restored.**
Evidence: `documentation/tests/golden-0.218.0-2026-08-22/`.
---
## 13. STOP — THE APPROVAL
**Nothing below has been changed. Vouching and the floor are yours.**
### The three Day-0 values, each with BOTH checks
| field | set to | downloadable | selectable |
|---|---|---|---|
| **`golden_version`** | **0.218.0** — sha `8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b` | ✔ HTTP 200, 657 026 013 B, and the **downloaded bytes** hash to exactly the bake's reported sha | ✔ offered in the hub dropdown with `data-sha=8e427869d13eafb7…` |
| **`agent_version`** | **0.130.0** — sha `a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3` | ✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's `data-sha` exactly | ✔ already the SELECTED option — **no change needed** |
| **`min_agent`** | **0.129.0** | — | ✔ already set to 0.129.0 — **no change needed** |
**So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut.** The
three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller
being baked declares **`MinAgent: 0.129.0` (unchanged)**, the vouched agent is already **0.130.0**,
and `0.130.0 ≥ 0.129.0` — so `agent_version` and `min_agent` are already correct and only
`golden_version` moves. Setting `min_agent` above the vouched agent is the R-216 shape; it is not
happening here.
**If the hub refuses the save, the guard is working.** The R-120 gate on that POST
(`hub/internal/web/configs.go:1165`) refuses a golden older than the newest controller the fleet
reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.
**Vouching is reversible:** re-select 0.217.0 and Save. The old package is never deleted by a bake.
### The one line on the floor
**Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to.**
`demo-hp` already runs 0.218.0 (deployed directly for the live proof). **`demo-felhom` is still on
0.217.0 and will not move until the floor is raised**, so today it still has both defects. I did not
raise it; that is your call, as is vouching.
---
## 14. WHAT WAS DROPPED, AND OBSERVATIONS
**Dropped: nothing.** Both parts completed in the required order, with the live walk and all seven
red-proofs. The session did not run short.
**Observations, noticed and not acted on:**
- **`demo-felhom` was not touched** and still runs 0.217.0 — so the fleet is deliberately non-uniform
until the floor moves. Its two preserved fixtures were not involved in any step of this session.
- **The `.fab` export has the same R-355 blind spot** and is fixed by the same change — `export.go:600`
filters on the same `db.StackName`. Not separately verified live; the unit tests cover the resolver
that feeds it.
- **`paperless-ngx`'s orphan directory is now a permanent fixture on `demo-hp`** until R-367 is
decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for
leaving it alone for now.
- **The safety dump still overwrites the unit's own DB dump before renaming it** (R-361, out of scope)
— visible in this session's own evidence, where `romm`'s `db-dumps/` holds both a `pre-restore-*` and
the regular dump.