R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s

Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.

R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.

R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.

Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.

R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.

Ceiling R-366 -> R-367.
This commit is contained in:
2026-08-22 10:11:46 +02:00
parent 7064596c2e
commit 877fcd2a38
6 changed files with 1286 additions and 568 deletions
+295 -546
View File
@@ -1,599 +1,348 @@
# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night)
# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)
**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no
image built, nothing deployed.** Evidence:
**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of
2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product
where one customer action causes permanent total loss.**
The drill that found them is preserved at
`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
---
## 1. THE VERDICT
## 0. Baselines, re-established — not trusted from the sheet
**It is a mixture, and the drill's three options are all present — but they belong to different
faults, and only one of them explains the thing you actually saw.**
### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED**
Reproduced independently tonight, and it agrees with what R-353 already recorded:
1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data
drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** —
about an app that was installed, deployed, running and healthy.
2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt
box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it
held `compose/` and nothing else.
3. The restore returned that configuration and reported a bare completion.
So for that specific event the data was not in the thing that was restored. **R-353 called this
correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong
to think so; its text is more careful than the CHANGELOG's summary of it.
### What the drill found that nobody had seen: **THE RESTORE LOSES IT**
This is new, it is worse, and it is not the same fault:
**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.**
The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.**
Proven live on `calibre-web` at 22:23:36 with planted files:
| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore |
|---|---|---|---|---|
| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** |
| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** |
and the customer was told:
> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult.
> Ennek az alkalmazásnak nincs adatbázisa."
Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence
does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue
apps that declare no data drive, that leg is the entire dataset.**
**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit`
(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the
unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local**
restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring
PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of
them replays it.**
### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database
`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always
has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to
`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack
that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the
safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and
then says:
> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**"
The controller had dumped that database five minutes earlier.
**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive.
| item | expected | confirmed |
|---|---|---|
| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` |
| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box |
| golden vouched | 0.217.0 | ✔ `<option value="0.217.0" … selected>` |
| floor | 0.217.0 | ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)" |
| min agent | 0.129.0 | ✔ |
| register ceiling | R-366 | ✔ (now **R-367**) |
| `ssh hp` → 192.168.0.104 | works | ✔ |
| clean trees, all four repos | `HEAD == origin/main`, 0 dirty | ✔ |
---
## 2. THE FULL/EMPTY CONTRAST
## 1. PART 1 — where the wrong name comes from
**Apps chosen, and why.** From the two storage classes: **`calibre-web`** declares a data drive
(`needs_hdd: true`, `backup.userdata: media/books class: mandatory`) and **`privatebin`** /
**`opengist`** declare none (the 40-class; data lives in a Docker named volume). I ran **three** cases
rather than two, deliberately: FULL and EMPTY alone cannot separate *"it was empty"* from *"the class
is broken"*, so `privatebin` was run FULL as the disambiguator.
**It comes from a guess, made in a place where an answer was available.**
**The fixture** — 5 files, two with Hungarian accented names, recorded as raw bytes:
`deriveStackName` (`controller/internal/appbackup/dbdump.go:770-798`) is handed a container name and
asked which app owns it. For `paperless-postgres` it strips the role suffix to `paperless`; finds
`paperless` is **not** a deployed stack; finds the container name is not one either; finds no known
stack is a prefix of it (`paperless-ngx` is not a prefix of `paperless-postgres`) — **and then returns
the unresolved candidate anyway**, on its last line, silently.
**Every place that value is used, and what each did with it:**
| consumer | file:line | effect |
|---|---|---|
| dump directory | `backup/backup.go:482,506` | wrote to `…/primary/**paperless**/db-dumps/` on the SYSTEM drive |
| dump filename | `appbackup/dbdump.go:200-212` | `paperless-postgres.sql`, not `paperless-ngx-postgres.sql` |
| unit assembler | `backup/recovery_unit.go:131` | reads `…/primary/**paperless-ngx**/db-dumps` → finds nothing → `db_dumps: null` |
| off-site collector | `backup/offbox_capture.go:32` | resolves from the unit path → the orphan is invisible to it |
| restore's DB leg | `backup/restore_db.go:69` | `db.StackName != stackName` → never replays |
| **safety-dump filter** | `backup/offbox_reconstitute.go:134` | `mine` empty → `hasDB=false` → **no undo copy, and the refusal is never reached** |
| `.fab` export | `appexport/export.go:600` | same filter — the bundle carries no database either |
| the "has a DB" flag | `web/handlers.go:1245` | the app displays as having no database |
**Corroboration from the code itself:** `ListDumpFiles` (`dbdump.go:562`) parses a dump filename under
the comment *"Parse stack name and DB type from filename: `paperless-ngx-postgres.sql`"* — the reader
was written for a name the writer never produced.
### Controller or catalogue? — **THE CONTROLLER. No halt.**
Every container the controller starts already carries `com.docker.compose.project`, read live:
```
SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991
binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a
nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4
plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1
árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea
paperless-postgres paperless-ngx kimai-db kimai romm-db romm
```
name bytes (UTF-8 NFC):
**And that label is the stack name BY CONSTRUCTION, not by luck:** `composeExecCustomEnv`
(`internal/stacks/manager.go:1218-1230`) runs compose with `cmd.Dir` set to
`/opt/docker/stacks/<stack>` and **never passes `-p`**, so compose derives the project from that
directory. Renaming the container in the catalogue would have fixed this one app and left the guessing
for the next one. **No catalogue change was made.**
---
## 2. PART 1.2 — the sweep, proven before its answer was trusted
| step | result |
|---|---|
| scratch copy, unplanted | `MISMATCHES: 1` (rc 1) — the known one |
| **plant a second mismatch** (`kimai-db` → `timetrack-db`) | **`MISMATCHES: 2`**, convicted by name: `stack=kimai container=timetrack-db -> derived=timetrack` |
| plant removed | back to `1` |
| **every mismatch removed** | **`MISMATCHES: 0`, rc 0** — "1" is not a stuck value |
| scratch discarded | `kimai/docker-compose.yml` md5-identical to the real catalogue |
**The real count: 1 affected app of 53**, from 15 DB containers checked — `paperless-ngx`.
---
## 3. PART 1 — the dumps already written under the wrong name
**They stay, untouched, and nothing will ever collect them.**
`/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql`, 312 381 B,
last written 07:38 UTC by the final pre-fix cycle — **verified still present after the fix.**
Nothing deletes them by design: the F5 stale-primary prune (`backup/backup.go:1248`) skips any
directory that is not a deployed app, under the guard *"an undeployed app's last backup is still its
restore point"*.
**They can be adopted, by hand:** move the file to
`…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point.
**Deliberately not automatic** — it would sit beside volume tars from a different time (an incoherent
pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a
customer's data on upgrade is a migration, not a fix. **Filed as R-367.**
---
## 4. PART 2 — the mechanism, and the comment corrected
`ReconstituteFromOffsite` skipped every placement flagged as the unit (`offbox_reconstitute.go:344`),
and the volume archives live **inside** the unit. A search of the whole off-site path found no
volume-restore call; the only two callers of `restoreDockerVolumes` were `restore.go:67` and
`restore_unit.go:255`, both local.
**The comment as it stood:**
> *"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and
> clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the
> scratch unit instead, so nothing is lost by skipping it."*
**Which half still holds: the FIRST.** The live unit is the local restore path's own source and must
never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit
across the operation to keep it true.
**The second half was false and is corrected, not left.** True of the *database* dump, false of the
*volume* archives, which live in the same unit and were replayed by nothing at all. Skipping the
placement is correct; treating the skip as harmless was not — and that sentence is exactly why a
reader would not look.
**The fix:** `restoreDockerVolumesFrom(stackName, dumpDir) (int, error)` — the local path's own replay
with an explicit directory, the same shape `reimportDBDumpsFrom` already had beside `reimportDBDumps`.
**One implementation, two callers.** Volumes replay **before** the database (a logical dump must still
win over a volume-tar copy of the same database) and **inside the stopped window** (Docker will not
replace a volume a container holds).
---
## 5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM
**R-354 — calibre-web, identical fixture, identical steps, both on `demo-hp`:**
| | |
|---|---|
| **before** (0.217.0, 09:41) | `ok=true` — „A(z) calibre-web: **0 fájl visszaállítva** (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and `ls: /vol/R354-2026-08-22: No such file or directory` |
| **after** (0.218.0, 09:56) | `ok=true` — „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and **5/5 files byte-identical** |
**R-355 — paperless-ngx:**
| | |
|---|---|
| **before** (2026-08-21) | „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" over a live 72-table PostgreSQL, with **no undo copy taken at all** |
| **after** (0.218.0, 09:57) | „A(z) paperless-ngx: **0 fájl és 3 adatkötet és az adatbázis visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult." |
**A third sentence now exists for the case that had no honest wording** — an app that HAS a database
whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak **VAN adatbázisa**, de a mentés nem
tartalmazott adatbázis-mentést, ezért az adatbázis **NEM állt vissza**. A visszaállítás előtti állapot
mentése megvan: `<undo>`". That case previously printed the same confident „nincs adatbázisa".
---
## 6. THE HARDWARE WALK
**Method: endpoint-level — the exact endpoints the dashboard's JS calls (`/api/backup/run`,
`/backup/offbox/run`, `/backup/offbox/restore`, `/backup/offbox/reconstitute`), with no shell inside
the machine for any step of the walk.** Planting the fixture and reading the result back used a shell
and is setup/verification, not the walk. No browser exists on DooPlex.
**The fixture, hashed before anything ran:**
```
7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e SENTINEL.txt
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5 binary-512k.bin
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4 nested/őszibarack.md
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5 plain.txt
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58 árvíztűrő-tükörfúrógép.txt
accented names as RAW BYTES (UTF-8 NFC):
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
```
**The comparator was proved able to convict before it was trusted.** One byte flipped at offset
500 000 of `binary-1mb.bin` (`af` → `00`): `sha256sum -c` reported `binary-1mb.bin: FAILED`, rc=1,
while the other four passed; the unmodified set passed rc=0. The mutant was discarded.
**The comparator was proved able to convict first:** one byte flipped at offset 300 000 of
`binary-512k.bin` (`3b` → `00`) → `binary-512k.bin: FAILED`, rc 1, the other four `OK`; the unmodified
set rc 0. Mutant discarded.
### The results
### Results
| case | class | data leg | in unit | in off-site snapshot | checking folder | off-site restore | local restore |
|---|---|---|---|---|---|---|---|
| **calibre-web FULL** | drive | user files | n/a | 5/5 identical | 5/5 identical | **5/5 identical** | — |
| **calibre-web FULL** | drive | named volume 1.42 MB | yes | yes | yes | **NOT restored** | — |
| **privatebin FULL** | no drive | named volume 1.06 MB | yes | 5/5 identical | 5/5 identical | **REFUSED — false reason** | **5/5 identical** |
| **opengist EMPTY** | no drive | named volume 181 KB skeleton | yes | yes | yes | **REFUSED — false reason** | — |
| scenario | result |
|---|---|
| **1-A** the dump is inside the unit, and inside the off-site snapshot | **PASS.** 0.217.0 09:37: `db_dumps = None`, dump refreshed into the phantom dir. 0.218.0 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` (312 669 B) in the app's own unit, and `restic ls` shows it inside the snapshot **for the first time** |
| **1-B** an undo copy taken and verified before anything stops; the message names the database | **PASS.** Before: `find /mnt -name "pre-restore-*paperless*"` → **0**. After: `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql`, 312 957 B, in the app's own unit |
| **1-C** the undo cannot be taken → the whole restore refuses, nothing changed | **PASS**, on the app that could never reach this guard before. `ok=false`; the marker kept its mutation (`51d37b0f…`) and `StartedAt` was unchanged — **the app was never stopped** |
| **1-D** every other app byte-identical | **PASS.** `romm-mariadb.sql`, `kimai-mariadb.sql` unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created |
| **2-A/B** the volume comes back and the message says so | **PASS.** calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database |
| **2-C** a snapshot with no volume archives is unchanged | **PASS** — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected |
| **2-D** the live recovery unit is never written | **PASS** — fingerprinted across the whole operation (see red-proof 6) |
| **2-E** a failed replay is reported as a failure naming the volume | **PASS** (unit test; the live path returns the same error) |
**The contrast decides it.** The unit and the off-site snapshot **hold the data, verified by
identity**, in both classes, accented filenames included. So the capture is sound. The failure is
entirely in the last leg.
**Both messages, verbatim:**
- FULL, drive class → `ok=true`, *„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16)
— az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."*
- FULL and EMPTY, no-drive class → `ok=false`, *„a(z) privatebin nincs telepítve, ezért nincs hová
visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra
az alkalmazást (Alkalmazások) ugyanerre a helyre…"*
The second is the important one. **FULL and EMPTY got the identical sentence**, so the message cannot
distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is
installed, and the instruction ("reinstall it to the same place") **cannot be followed**, because a
40-class app is offered no storage field at deploy time (that is R-352's own measurement).
**This closes R-353's second instruction.** It asked for proof that a named-volume app reaches the
off-site tier rather than inference from gate order. It does. Manifests read tonight:
```
privatebin volume_dumps = ['privatebin_privatebin_data.tar'] db_dumps = None
opengist volume_dumps = ['opengist_opengist_data.tar'] db_dumps = None
kimai volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar']
db_dumps = ['kimai-mariadb.sql']
calibre-web volume_dumps = ['calibre-web_calibre_web_config.tar'] db_dumps = None
paperless-ngx volume_dumps = [3 tars] db_dumps = None ← the defect
```
All 15 containers healthy after the walk.
---
## 3. PART 0 — did the floor move the machine?
## 7. RED-PROOFS — seven, each mutation asserted applied and reverted
**Yes, unaided, in 17 seconds.**
| # | mutation | outcome |
|---|---|---|
| 1 | `resolveStackName` ignores the compose label | **FAIL as required.** `= "paperless", want "paperless-ngx"`, and the consequence printed both paths — written to `…/primary/paperless/db-dumps` vs read `…/primary/paperless-ngx/db-dumps`. **The empty database record, returned** |
| 2 | outcome message back to the counter-only predicate | **FAIL as required**, reproducing the 21 August sentence verbatim: `"A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."` |
| 3 | the fail-closed refusal removed from `writeSafetyDump` | **FAIL as required** — *a restore was seen proceeding with no undo copy* |
| 4 | the volume leg removed | **FAIL as required.** `VolumesReplayed = 0` and the replay never called — last night's silent loss, returned |
| 5 | the volume count dropped from the message | **FAIL as required**, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim |
| 6 | the live-unit guard removed | **PASSED FIRST — a defect in MY TEST, not the code.** The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit **root**. Widened to the whole unit (excluding only the documented undo copies) it convicts: `PLACED:17603ba1…` appears. **The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory** |
| 7 | the volume-replay error swallowed | **FAIL as required** — a partial replay reporting success |
**Green gate:** `go build ./... && go vet ./... && go test ./...` — **28 packages ok, 0 FAIL lines.**
Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red
until the bake (§8) — correctly, and never bypassed.
*(Instrument note: `go test ./... | grep -vE '^ok'; echo rc=$?` reports the **grep's** exit code, which
is 1 when every test passed and nothing was left to print. The verdict above is from an explicit
`FAIL`-line count, not from that.)*
---
## 8. WHAT THIS DOES **NOT** FIX — the blocker for the apps that need it most
**The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive
(R-356)**, telling the customer that a running app „nincs telepítve" and to reinstall it "to the same
place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset
is a named volume, **so R-354's fix cannot reach them until R-356 is closed.**
That is why the live confirmation used `calibre-web` and `paperless-ngx`: they declare a drive and can
actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing
worth doing.
---
## 9. REGISTER
- **R-354 — CLOSED**, shipped + proven-live (v0.218.0).
- **R-355 — CLOSED**, shipped + proven-live (v0.218.0).
- **R-367 — OPENED (LOW)**: dumps already written under the wrong name are stranded; adoptable by hand,
deliberately not automatic, never on an upgrade path.
- **Ceiling moved: R-366 → R-367.**
---
## 10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten
R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives,
awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification
copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359,
R-361–R-363, R-365, R-366.
---
## 11. COMMITS, VERSIONS, CI
| | |
|---|---|
| moved from → to | **0.216.0 → 0.217.0** |
| who initiated | **the hub** — the operator saved the floor at `19:48:37Z`; a poke reached the agent from `10.77.0.1:58093` at `19:48:32Z` for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision. |
| how long | floor saved `19:48:37Z` → *"controller-swap: new controller healthy"* `19:48:54Z` = **17 s**. Swap requested `19:48:41Z` → healthy = 13 s. |
| **the container's own tag on the box** | **`gitea.dooplex.hu/admin/felhom-controller:0.217.0`**, created `2026-08-21 19:48:45 UTC` — read from `docker ps` in guest 9201, not from the hub. |
The hub's own host page for `demo-hp` shows the guest's **Controller column as „—"** — the hub does
not know which controller version the box runs, while the `controller_updated` event it received says
exactly that. Two hub surfaces, one blind.
`demo-felhom` also runs 0.217.0, but it started it at `19:31:04Z` — **17 minutes before the floor was
saved** — and emitted `controller_started` with **no `controller_updated`**. It was moved by hand
during the golden bake, not by the floor.
| `felhom-controller` | `5ce3a44` the fix + tests; `2da259a` README + CONTEXT |
| image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` (150M) |
| deployed | `demo-hp` guest 9201 — `…:0.218.0 Up (healthy)` |
| `demo-felhom` | **untouched this session**, still 0.217.0 |
---
## 4. PART 2 — the four answers, from source
## 12. BAKE
**1. What puts a dump into a backup unit, and when? Which apps qualify?**
`RUNBOOK-manual-build.md` §4.0/§4.1, in the drill VM. **The golden-currency gate went red the moment
v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.**
- **Database dumps** — `runDBDumps` (`backup/backup.go:444-560`). Qualification is **a running
container whose image matches a database image**, mapped to a stack by `deriveStackName`
(`appbackup/dbdump.go:770`). Written to `<nsRoot>/backups/primary/<stack>/db-dumps/<stack>-<type>.sql`.
- **Volume dumps** — `runVolumeDumps` (`backup/backup.go:607`). Qualification is **a deployed,
non-protected stack with at least one Docker *named volume*** (`GetDockerVolumes`). An app with
zero named volumes is skipped silently and is never stopped.
- **When** — one cycle, both legs, scheduled `db-dump` daily at **02:30 CEST**; and again as the
coherence pre-phase of every off-site run (`offbox.go:938-951`), which is what makes a snapshot an
internally coherent {DB@T, files@T} pair.
- `CaptureRecoveryUnit` **only enumerates what is already on disk**
(`recovery_unit.go:131-132`). It writes no dump itself.
- **The gap this leaves:** an app whose data is a *bind mount* and which has no database container
produces neither leg. Its unit is configuration only — and nothing anywhere says so.
| | |
|---|---|
| `GOLDEN_VERSION` | **0.218.0** |
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
| package | `…/generic/felhom-golden/0.218.0/golden.tar.zst`, 657 026 013 B |
| template | `debian-13-standard_13.6-1_amd64.tar.zst` — listed fresh, not assumed |
**2. What should a correct backup contain?**
- **Declares a data drive** (13 of 53): the recovery unit (compose + `app.yaml` carrying the portable
secrets + whatever dumps exist) **plus** the paths its `.felhom.yml` `backup:` block marks
`mandatory`, appended to the restic snapshot as extra paths (`offbox_capture.go:32`).
- **Declares none** (40 of 53): **unit only**. `offboxCaptureSet` returns `(nil, nil)` when the app
has no backup block, so the off-site snapshot is the unit and nothing else. That is correct *given*
that their data is inside the unit's volume tars — and tonight confirmed it is.
**3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.**
`ReconstituteFromOffsite` returns `res` and `nil`. Nothing in it ever asks whether anything was
restored. The handler then calls `EndRestoreOp(**true**, reconstituteOutcomeMsg(...))`
(`web/offbox_handlers.go:455`). `reconstituteOutcomeMsg` (`web/offbox_handlers.go:463-475`) formats
counters:
```go
if res.DBsReplayed == 0 {
return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+
"Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when)
}
```
So with `FilesPlaced == 0` and `DBsReplayed == 0` the customer reads **"0 files restored — the
application restarted"** under a **success**. **This is a finding on its own and is recorded whichever
way the rest goes.** Two aggravations found on top of it:
- **"This application has no database" is asserted from a counter, not from a fact.**
`reimportDBDumpsFrom` returns `(0, nil)` when the dump directory is absent *or* holds no `.sql`
(`restore_db.go:30-46`), so an app that certainly has a database is told it has none. Proven live on
`paperless-ngx`.
- **The file count is rsync's transfer count, not a restore count.** A correct restore of unchanged
data reports **"0 fájl visszaállítva"** — indistinguishable from a restore that did nothing.
Observed: the same app reported 5, then 2, then 0 across three runs.
**4. What did the 9 August off-site snapshot contain, and is it still readable?**
**Readable, and it contained the data.** Three snapshots at `2026-08-09 08:30`:
| id | app | contents |
|---|---|---|
| `41c830db` | calibre-web | unit + `userdata/media/books` incl. the 9-Aug rehearsal fixture (`árvíztűrő-tükörfúrógép.txt`, `őszibarack.md`, `binary-3mb.bin`, `plain.txt`) |
| `9e38b84c` | opengist | unit + **`volume-dumps/opengist_opengist_data.tar`, 182 272 B** — manifest records `volume_dumps: ["opengist_opengist_data.tar"]` |
| `78b93f04` | privatebin | unit + `volume-dumps/privatebin_privatebin_data.tar`, 2 560 B |
All still readable tonight; `restic check` over the whole repository reports **"no errors were
found"**. **So the 9 August off-site copy of OpenGist's data exists and is intact** — it simply was
not the thing the afternoon's restore read, and the off-site route that would have read it refuses
for that class of app.
**The repository had also been dead since 9 August** and nothing said so: no snapshot between
`2026-08-09 08:30` and tonight. The cause is visible in the hub event at 16:01 — the off-box target
was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target,
**every per-app off-site toggle was still off**, so the first run I triggered reported:
**Pass markers, each checked, with the negative controls:**
```
[offbox] backup run started (0 app(s) toggled)
[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
upload OK (HTTP 201) : 1 -> pre-delete HTTP 404 (404/204 expected)
excluding : 0 <- negative control
FATAL : 0 <- negative control
```
**Credit where it is due:** the *card* does not lie about this. It reads „Aktív — nincs kijelölt
alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still
leads, and the log line alone says only „backup OK".
**Token hygiene.** Copied file→file and read by a runner script inside the VM, so it never crossed a
shell or a unit property: `systemctl show golden-bake … | grep -c -F "$(cat …)"` → **0**. **The leak
grep on the committed log was itself proven before its zero was believed** — token appended to a
throwaway copy, grepped (**1**), copy shredded, then the committed log's **0** accepted.
**Teardown:** `pct destroy 9100 --purge`; token, runner, build script and in-VM log shredded **after**
the log was copied out; VM powered off; qemu exited; **`qemu-img snapshot -a virgin` restored.**
Evidence: `documentation/tests/golden-0.218.0-2026-08-22/`.
---
## 5. THE TWO CYCLES COMPARED — AND THEY AGREE
## 13. STOP — THE APPROVAL
**By hand:** local cycle 22:12, off-site 22:17 (after enabling the per-app switches the rebuild had
silently cleared). **By the clock:** the box's own `db-dump` at 02:30 CEST, cross-drive at 03:30,
off-site at 04:15.
**Nothing below has been changed. Vouching and the floor are yours.**
| | manual | scheduled | agree? |
### The three Day-0 values, each with BOTH checks
| field | set to | downloadable | selectable |
|---|---|---|---|
| units refreshed | all 6 apps | all 6 apps, `00:30Z` | **yes** |
| `privatebin` volume tar | 1 055 744 B | 1 055 744 B | **yes** |
| `opengist` volume tar | 181 248 B | 181 248 B | **yes** |
| `calibre-web` config tar | 1 422 848 B | 368 640 B † | **yes, and explained** |
| `paperless-ngx` `db_dumps` | **absent** | **absent** | **yes — the defect reproduces on the scheduled path** |
| the orphan `…/primary/paperless/db-dumps/` | written | **written again, 294 936 B, `00:30Z`** | **yes** |
| planted files on the drive | 5/5 identical | **5/5 identical**, plus `POST-SNAPSHOT.txt` | **yes** |
| off-site run | ok, 27 → snapshots | ok, `02:17:12Z`, **2m8s**, 42.8 MB, **27 snapshots** | **yes** |
| **`golden_version`** | **0.218.0** — sha `8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b` | ✔ HTTP 200, 657 026 013 B, and the **downloaded bytes** hash to exactly the bake's reported sha | ✔ offered in the hub dropdown with `data-sha=8e427869d13eafb7…` |
| **`agent_version`** | **0.130.0** — sha `a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3` | ✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's `data-sha` exactly | ✔ already the SELECTED option — **no change needed** |
| **`min_agent`** | **0.129.0** | — | ✔ already set to 0.129.0 — **no change needed** |
† the tar shrank because the drill had by then removed the planted 1 MB from that volume — the
expected value, not a discrepancy.
**So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut.** The
three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller
being baked declares **`MinAgent: 0.129.0` (unchanged)**, the vouched agent is already **0.130.0**,
and `0.130.0 ≥ 0.129.0` — so `agent_version` and `min_agent` are already correct and only
`golden_version` moves. Setting `min_agent` above the vouched agent is the R-216 shape; it is not
happening here.
**Two independent observations, no disagreement.** The one that matters: **R-355 is not an artefact of
my manual triggering.** The box, unattended, on its own schedule, wrote paperless-ngx's PostgreSQL
dump into a directory for a stack that does not exist and left the app's own unit recording
`db_dumps: null`.
**If the hub refuses the save, the guard is working.** The R-120 gate on that POST
(`hub/internal/web/configs.go:1165`) refuses a golden older than the newest controller the fleet
reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.
The scheduled cross-drive (Tier-2) leg also completed for all six apps at 01:30Z — `crossdrive_completed`
events 3019–3024.
**Vouching is reversible:** re-select 0.217.0 and Save. The old package is never deleted by a bake.
### The one line on the floor
**Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to.**
`demo-hp` already runs 0.218.0 (deployed directly for the live proof). **`demo-felhom` is still on
0.217.0 and will not move until the floor is raised**, so today it still has both defects. I did not
raise it; that is your call, as is vouching.
---
## 6. PART 4 — everything attempted, and the message judged
## 14. WHAT WAS DROPPED, AND OBSERVATIONS
| # | test | outcome | message judged |
|---|---|---|---|
| 1a | **destructive restore: nothing is ever deleted** | **PASS both ways.** A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. | count is rsync transfers, not files restored — see §4.3 |
| 1b | **safety dump taken and verified before the stop** | **PASS.** `21:02:47Z` dump written → `21:02:47Z` `StopStack romm`. | correct: „0 fájl **és az adatbázis** visszaállítva" |
| 1b | **…and the whole operation refuses if it cannot be** | **PASS.** Made impossible by putting a regular file at the `db-dumps` path. Refused; `plain.txt` kept its post-snapshot mutation; the container's `StartedAt` was unchanged — **the app was never stopped.** | honest and names the path, but leaks a raw Go `mkdir … not a directory` into a customer surface |
| 1b | **the hole in it** | **FAIL.** For `paperless-ngx` the discovery resolves the database to the wrong stack, so `hasDB` is false: **no safety dump is taken and the refusal cannot fire.** The undo is absent, not refused. | „Ennek az alkalmazásnak nincs adatbázisa" — false |
| 2 | **end of the abandonment countdown** | state created and overdue; **fires at 05:10 CEST** — see §11 | card states a **past** date in the future tense while overdue |
| 3 | **damaged store** | **MIXED — see below** | honest at the point of failure, then forgotten |
| 4 | **drive pulled mid-restore** | restore failed, nothing written to the wrong place, **the agent re-bound the drive within 5 s** (`23:15:29` pulled → `23:15:34` re-bound) | **WRONG DIAGNOSIS.** „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions. |
| 5 | **full disk, non-destructive path** | **PASS — refuses before it starts.** | **exemplary:** „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named. |
| 5 | **full disk, destructive path** | **FAIL — no gate at all.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the three headroom gates are all on non-destructive paths (`offbox_restore.go:231,297,423`). It stopped the app, failed halfway on ENOSPC, left `DRILL-2026-08-21/` holding 2 of 5 entries, and restarted the app. | honest (`No space left on device (28)`) but raw rsync output |
| 6 | **controller killed mid-restore** | **PASS.** SIGKILL inside the stop→restore→start window. On restart: *"[appstop] crash recovery: an off-site restore (op \"offbox-reconstitute:paperless-ngx\") was interrupted and left 1 app(s) stopped — restarting them"*. App restarted, marker cleared, **and the hub was told** (event 3009). | good — names the operation, not just the app |
| 7 | **the 40-class under pressure** | **the backup noticed; the alarm did not.** Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The **reserve refused per app**: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got `recovery_unit_capture_failed` (**error**) naming the filesystem and its numbers. **The fill watcher said nothing** — it runs **once a day at 03:30** plus once at startup (`cmd/controller/main.go:1092`). | the refusal messages are good; the silence is the problem |
**Dropped: nothing.** Both parts completed in the required order, with the live walk and all seven
red-proofs. The session did not run short.
**Where the clock stopped me:** nothing in Part 4 was skipped for time. Item 2's terminal deletion is
scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only.
**Observations, noticed and not acted on:**
### 4.3 in detail — the damaged store
One byte flipped inside pack `967853d2…` at offset 5 000 000, over the repository's own SFTP
transport. Pack files are named by their content hash, so this is genuine corruption.
- **`restic check` detects it** — *"ciphertext verification failed"*, *"Fatal: repository contains
errors"*.
- **But nothing in the product ever runs it.** The only restic verbs in the entire controller are
`restore, snapshots, backup, unlock, stats, init, forget, prune, cat`. The agent's restore-test is
**PBS-tier only**. **The off-site store is never verified by any layer, at any time.** Corruption is
discovered at restore time — the worst possible moment.
- **A restore that touches the damage fails honestly:** `ok=false`, naming the file and
*"ciphertext verification failed"*.
- **But the failure is not remembered.** It left a **partial** checking folder — 78 MB, 54 files,
15 of 16 originals. `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only *"does the
directory exist and is it non-empty"*, so the wizard then offered **all three** actions including
„Teljes visszaállítás indítása".
- **Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.**
The repository was repaired from byte-identical originals; `restic check` now reports **"no errors
were found"**.
---
## 7. PART 5 — the two rows that were observed and never filed
**5.1 — the delete guard on verification copies is blind, and its own comment says otherwise.**
`offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) gates on `backupMgr.IsRunning()` — the
**concurrency** flag — while its comment states *"It refuses while a backup/restore op is running: the
copy being deleted could be the one currently being written."* R-351b moved all seven restore handlers
onto `restoreOpBlocked()` (which reads both flags); **this handler was left behind**, and one other
site (`:239`) reads the bare flag correctly and documents why.
**Reachability is not a race — it is the whole operation.** `RestoreOffboxScratch`
(`offbox_restore.go:211`) **never calls `acquireRunning` at all**, so `IsRunning()` is false for the
entire duration of an off-site verification restore. Demonstrated live, flags read immediately before
and after the delete:
```
--- BEFORE delete 22:35:25 display= True offbox-restore kimai concurrency= False
--- DELETE calibre-web verification copy: „Az ellenőrző másolat törölve…"
--- AFTER delete 22:35:25 display= True offbox-restore kimai concurrency= False
```
The copy was removed. The guard is app-agnostic, so the same call naming the *restoring* app hits the
directory the restore is writing into. **Rank: MEDIUM** — it needs a customer to press delete during a
restore, but both controls live on the same page, the window is the whole restore, and the target is
the restore's own source. (Reconstitute and place *do* hold the flag, so the exposure is the
verification-restore window only.)
**5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight.**
Filed as an instrument defect. **Occurrences I can evidence:**
1. **2026-07-20** — an accented grep through `ssh → pct exec → bash -c` nearly produced a wrong
"banner cleared" claim. Recorded in `felhom-controller/.claude/rules/ui-hungarian.md:19-22`.
2. **2026-08-13** — `kubectl exec … sh -c "grep '<accented>'"` returned **0 for three strings that
were present**, one step from being reported as a failed hub v0.105.0 deploy.
3. **2026-08-21, tonight, 22:07** — `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`
(octal escaping). Recording the accented filenames' raw bytes from that listing would have been
wrong. Caught by extracting the archive and reading the names with `xxd`.
**Correction to the task's premise:** that is **two inside two weeks**, plus the founding case a month
earlier. I looked for a third inside the two-week window and did not find one on record.
**The smallest guard I would propose — and did NOT build:** the problem is not grep, it is that every
one of these tools *silently transforms* the bytes. So the guard is not "use ASCII fragments" (a
discipline, which is what failed three times) but **a negative control that the harness cannot skip**:
any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a
string that MUST be absent — and a zero result from the first is only reportable when the second also
returns zero *and* a third probe for a known-present ASCII anchor returns non-zero. Three probes, one
helper, no judgement required at the call site. Everything else has been tried and is what "nearly"
means in all three cases.
---
## 8. RANKED REGISTER ROWS OPENED
Ceiling was **R-353**; it **moved to R-366**.
| id | rank | what |
|---|---|---|
| **R-354** | **HIGH** | The off-site full restore has **no named-volume leg**. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. `offbox_reconstitute.go:341-346`. |
| **R-355** | **HIGH** | `paperless-ngx`'s PostgreSQL is dumped to a directory for a **non-existent stack** (`…/primary/paperless/`), so its unit records `db_dumps: null`, nothing off-sites it, **no safety dump is taken on a destructive restore**, and the customer is told the app has no database. `appbackup/dbdump.go:770-798`. One app in 53. |
| **R-356** | **HIGH** | The off-site restore **refuses for all 40 no-drive apps** with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. `offbox_reconstitute.go:208-227`. |
| **R-357** | **MEDIUM** | The **destructive** restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory. |
| **R-358** | **MEDIUM** | A **failed** scratch restore leaves a partial copy that `OffboxFullScratchReady` reports as ready; the destructive restore then runs from it and reports success. |
| **R-359** | **MEDIUM** | The off-site restic store is **never verified** by any layer — `restic check` is not among the verbs the controller runs, and the agent's restore-test is PBS-only. |
| **R-360** | **MEDIUM** | Verification-copy delete gates on `IsRunning()`, which `RestoreOffboxScratch` never holds — deletable throughout a restore. Its comment asserts the opposite. **(Part 5.1)** |
| **R-361** | **MEDIUM** | The safety dump **overwrites the unit's own DB dump**: `DumpOne` writes the canonical `<stack>-<type>.sql` and only then renames it away. The comment at `offbox_reconstitute.go:147-148` states it "can never overwrite the app's real dump". Proven: romm's `romm-mariadb.sql` was present before and absent after. |
| **R-362** | **MEDIUM** | A data drive detached mid-restore is reported as **„permission denied"**. The restore path never consults drive state. |
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
| **R-366** | **HIGH** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives too** — the box cannot read its own pre-reinstall backups (`manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`, hub event 3016, filed unprompted at 21:59). The PBS-tier analogue of R-193: a rebuilt box loses **both** off-premises tiers at once. The restore-test caught it precisely; the gap is that it is called *a failed restore test* rather than *your older whole-guest backups are unreadable*. **Found incidentally — nobody was looking for it.** |
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
nobody. Observed tonight as event 3006, severity `info`. A sweep of every `emit(` call site shows this
is now the **only** remaining instance fleet-wide.
**R-353:** its instruction (2) is **satisfied** — see §2. Its instruction (1) stands and is now
strictly larger than when it was written, because R-354 shows the bare completion can also be reported
over a unit that *did* have a data leg.
---
## 9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH
| time (CEST) | what |
|---|---|
| 21:57 | start; baselines |
| 22:00 | break-glass into `demo-hp` |
| 22:09 | fixture planted, comparator control passed |
| 22:12 | manual local backup |
| 22:13 | manual off-site run — **0 apps toggled** |
| 22:17 | off-site run with 3 apps enabled |
| 22:19–22:20 | checking-folder restores |
| 22:21 | reconstitute privatebin → refused |
| 22:23 | reconstitute calibre-web → **the conviction** |
| 22:25 | local restore-from-unit privatebin → data returned |
| 22:27 | Part 4.1a — no-delete invariant |
| 22:39–22:45 | paperless-ngx → R-355 |
| 22:51–22:57 | Part 4.3 damaged store; repo repaired |
| 23:02–23:04 | Part 4.1b safety dump, both directions |
| 23:10–23:13 | Part 4.5 full disk (both paths); Part 4.6 controller killed mid-restore |
| 23:15 | Part 4.4 drive pulled |
| 23:17–23:26 | Part 4.7 filesystem filled, backup reserve observed, filesystem freed |
| 23:08–23:09 | abandonment set-aside store created; countdown written and controller restarted (fires 05:10) |
**Steps off the customer's path, named:**
1. **Break-glass root access** to `demo-hp` via the hub-vaulted `host_recovery` credential — the box
had lost DooPlex's SSH key (its `authorized_keys` held only its own `root@demo-hp` RSA key) and its
tailnet address was unreachable. DooPlex's public key was **re-added** to `/root/.ssh/authorized_keys`
and an `ssh` alias `hp` → `192.168.0.104` was added to `~/.ssh/config` on DooPlex.
2. **A hub DB snapshot** (`hub.db` + `-wal` + `-shm`) was streamed to the scratchpad to read
`host_recovery` and the events table. It holds every host's secret; it is in the session scratchpad
only and is not in any committed file.
3. **`settings.json` was edited directly** (controller stopped, backup at `/root/settings.json.drill-backup`)
to create the overdue abandonment state. There is no product path to shorten a countdown, and the
real orphan→reset path would have destroyed `demo-hp`'s entire off-site history.
4. **The set-aside store the sweep will delete was created by hand** at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, for the same reason.
5. **Two restic pack files were deliberately corrupted and then restored** from byte-identical copies.
6. **`fallocate` fillers** were used to fill two filesystems and were removed.
7. Apps deployed for the drill: `opengist` (re-deployed empty), `kimai`, `paperless-ngx`, `romm`.
---
## 10. THE FENCES
- **`ep0` / the off-site endpoint — untouched outside this machine's own path, and here is how I know.**
`demo-hp` authenticates as the Hetzner Storage Box **sub-account `u629488-sub3`**, which is chrooted
to its own home: `ls /` returns **`Permission denied`**, and `/home` contains exactly `.ssh` and
`felhom-repo`. Every write, the two corruptions, the set-aside store and the sweep's target are
inside that home. **`demo-felhom` is a different sub-account (`u629488-sub1`)** and a real customer's
copy is a different sub-account again — none reachable with this key. `restic forget`/`prune` were
never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept
every 9-August snapshot (verified by listing all 24, not by a count).
- **`demo-felhom`'s two fixtures — confirmed intact, not assumed.**
(a) the unopenable set-aside store `u629488-sub1:/home/felhom-repo.orphaned-20260810` — listed
tonight, `config`/`data` (258 shards)/`index`/`keys`/`locks`/`snapshots`, mtimes still 18 Jul and
3 Aug; (b) the retained-key case — hub `host_escrow_superseded` rows **11 and 12** for
`demo-felhom-8363b5`, each with a 572-byte `identity_blob`, dated 2026-08-12. `demo-felhom` is
healthy on 0.217.0, its own off-site ran at 02:15 with `last_status: ok`, and
**`--abandon-status` there reports „no abandonment countdown is running on this box"**.
- **`peti-felhom` — not contacted.** It does not appear in the hub host list at all; no command in this
session named it.
---
## 11. PART 4.2 — THE ABANDONMENT COUNTDOWN, WATCHED FIRING
**The terminal deletion has now been observed.** It did exactly what it claims, including the half
nobody had seen.
- **State created** 23:08–23:09 CEST. A realistic set-aside store was built at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, the countdown written overdue
(started 2026-08-07, due 2026-08-20), and the product's own CLI confirmed it:
`abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0`.
- **05:10 CEST — it fired.** `/home` on the storage box is stamped `03:10Z`; the set-aside store is
**gone**; **`/home/felhom-repo`, the live repository, is untouched** (mtime still 4 Aug).
- **05:13 CEST — the hub half.** Event **3025 `offsite_abandon_purged`**: *„Az ügyfél korábbi távoli
mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is törölve (1 csomag). Az ügyfél
döntése alapján, a 14 napos türelmi idő lejárta után."* And demo-hp's superseded escrow row **is
gone from the hub** — it held one before.
- **The controller closed itself out.** Every `abandon_*` field has been removed from `settings.json`
and `--abandon-status` reports *"no abandonment countdown is running on this box"*.
**So the two-phase commit's central promise — *"it removes BOTH halves or neither"* — is confirmed
live for the first time:** the ciphertext and the sealed package that protects it went together, three
minutes apart, and the state that remembered the operation cleaned itself up. **What it removes:**
exactly the recorded set-aside path. **What survives:** the live repository, the live escrow, and the
box's current recovery path.
**Judged:** the completion message is accurate and in plain Hungarian. The only wrong note is while
the countdown is *overdue but not yet swept* — the card then states a past date in the future tense
(**R-365**).
**No countdown is left running anywhere.** `demo-hp` cleared itself; `demo-felhom` reports none.
---
## 11b. WHAT THE NIGHT FOUND THAT NOBODY WAS LOOKING FOR
At 21:59 the box filed, unprompted, hub event 3016:
> `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could
> not be restored+booted … wrong key — manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided
> key dd:d1:d8:53:44:62:5e:0b`
That archive predates the 21 August reinstall by three days. **The rebuild orphaned the PBS whole-guest
archives exactly as it orphaned the restic repository** — so a rebuilt box loses *both* off-premises
tiers at once. The restore-test mechanism deserves credit: it caught it and named the key mismatch
precisely. The gap is what it is *called* — "a restore test failed" reads as a flaky verification, not
as "every whole-guest backup taken before the reinstall is unreadable on this machine". Filed **R-366
(HIGH)**.
---
## 12. MACHINE STATES AT THE END
**`demo-hp` — HEALTHY, not broken. Nothing needs bringing back.** At 05:28 CEST: all 15 containers up
and healthy (`felhom-controller` 0.217.0, `traefik`, `cloudflared`, `filebrowser`, plus the six drill
apps); `/` 4%, `/var/lib/felhom` 13%, the data drive 1%; **no filler files left**, **no app-stop marker**,
**no operation in flight**, **no countdown running**, and `restic check` over its off-site repository
reports **no errors were found**. Its off-site backup, silent since 9 August, is working again and ran
on its own schedule at 04:15.
**Deliberately left in place, each with a reason** (retained, not forgotten):
| left behind | why |
|---|---|
| `opengist`, `privatebin`, `calibre-web` deployed with the planted `DRILL-2026-08-21` fixture | the reproduction for **R-354** and **R-356**; the hashes in §2 make the fix verifiable without rebuilding the case |
| `paperless-ngx` deployed | the **only** reproduction of **R-355**, and it regenerates the orphan directory on every nightly cycle |
| `romm` deployed | the only app on the box where the safety dump actually works — the fixture for **R-361** and the control for R-355 |
| `kimai` deployed | a second correct-derivation DB app, the negative control for R-355 |
| five verification copies under `backups/offsite-restore/` | harmless, and they are the **R-358** fixture |
| `romm`'s `pre-restore-…sql` in its unit | the physical evidence for **R-361** |
| DooPlex's public key in `demo-hp:/root/.ssh/authorized_keys`, and the `hp` alias in `~/.ssh/config` | the box had lost the key and its tailnet address is dead; without it the next session must go through break-glass again |
**Removed / restored during the drill:** both corrupted restic packs (repo verified clean), both
`fallocate` fillers, the set-aside store (by the sweep, as intended), and demo-hp's superseded escrow
row (by the sweep's hub half, as intended).
**`demo-hp`'s tailnet address `100.76.96.79` is still unreachable** and was not repaired — the box is
reachable on the LAN at `192.168.0.104`. That is the one thing about it that is worse than it should
be, and it predates tonight.
**`demo-felhom` — untouched and healthy**, controller 0.217.0, its own off-site ran at 02:15 with
`last_status: ok`, both fixtures verified present (§10), no countdown running.
**`drill-r50`** — not used. Still `DOWN` in the hub, agent 0.129.0, as it was.
---
## 13. TEARDOWN — ALL FOUR LAYERS, STATED
| layer | state |
|---|---|
| **Off-site (storage box)** | The set-aside store I created was **deleted by the product's own sweep**, as designed. The two corrupted packs were **restored byte-identical** and `restic check` passes. Nothing else was written. `restic forget`/`prune` were never invoked by me. |
| **PVE host `demo-hp`** | `/root/settings.json.drill-backup` **retained** (the pre-drill controller settings, in case the abandonment edit needs reverting). DooPlex's SSH key **retained**, with the reason above. `/tmp/plant.tar` left; harmless. |
| **Guest 9201 / controller** | Six apps and the planted fixtures **retained with reasons** (table above). Controller state is clean: no marker, no countdown, no in-flight op. |
| **Hub — stated explicitly** | **No customer record and no host record was created, so none needs deleting.** I used the existing `demo-hp` customer and host throughout. The only hub-side *removal* was demo-hp's superseded escrow row, done by the abandonment sweep itself and reported as event 3025. Events 3002–3025 were generated as a normal consequence of the work and are left as the record. **Nothing is owed here and nothing is blocked.** |
| **DooPlex (this machine)** | The hub DB copies (which contain every host's break-glass secret) are in the session scratchpad only, never in a committed file, and are shredded in the closing step. |
---
## 14. OBSERVATIONS — noticed, not acted on
- **The hub does not display the controller version it is told.** `demo-hp`'s host page shows the
guest's Controller column as **„—"** hours after receiving `controller_updated: 0.216.0 → 0.217.0`.
Two hub surfaces, one blind. Not filed — `REPORT-hub-blindness.md` already exists and this may be
part of it.
- **`demo-felhom` moved to 0.217.0 without a `controller_updated` event** (started 19:31Z, 17 minutes
before the floor was saved). Consistent with a by-hand deployment during the golden bake, not a
defect — recorded so a later reader does not mistake it for one.
- **`restic --latest N` is per-group, not a total.** It briefly read as "18 snapshots became 10" and
would have been reported as data loss. Caught by listing all of them and by an independent
`restic stats` (26.44 MiB / 588 blobs). **An unpaginated listing is not a total** — the rule earned
its place again.
- **`find -newermt` is the wrong probe for a restore**: restic preserves the snapshot's mtimes, so a
freshly restored tree looks old. It briefly read as "nothing was restored".
- **`sftp -b` aborts on the first failing line**, so a batch listing 256 shard directories returned 7
packs and looked like a total. Prefixing each line with `-` fixed it. Same family as the two above.
---
## 15. WHAT WAS DROPPED, PLAINLY
- **Nothing in Parts 0, 2, 3, 4 or 5 was skipped.** Every Part 4 item was attempted and reached a
verdict.
- **The fill-watch alarm was not watched firing on its own schedule.** I proved its cadence from
source (daily 03:30) and observed it fire from the startup path (`disk_critical`, 21:10:58Z, correct
Hungarian copy naming the drive and the free space). Holding a filesystem at 99% for six hours would
have sat across the scheduled backup and the abandonment sweep, and I judged the scheduled cycle —
which the task asks for explicitly — worth more than a second sighting of an alarm whose trigger I
had already read and seen work.
- **I did not drive the abandonment through the real orphan→reset path**, because that path
move-asides the *live* repository, which would have destroyed demo-hp's whole off-site history
including the 9 August snapshots this drill exists to read. The sweep's own code path was exercised
in full on a real store; the deviation is only in how the state was created.
- **`peti-felhom` was not contacted**, per the fence.
- **`demo-felhom` was not touched** and still runs 0.217.0 — so the fleet is deliberately non-uniform
until the floor moves. Its two preserved fixtures were not involved in any step of this session.
- **The `.fab` export has the same R-355 blind spot** and is fixed by the same change — `export.go:600`
filters on the same `db.StackName`. Not separately verified live; the unit tests cover the resolver
that feeds it.
- **`paperless-ngx`'s orphan directory is now a permanent fixture on `demo-hp`** until R-367 is
decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for
leaving it alone for now.
- **The safety dump still overwrites the unit's own DB dump before renaming it** (R-361, out of scope)
— visible in this session's own evidence, where `romm`'s `db-dumps/` holds both a `pre-restore-*` and
the regular dump.