Files
felhom.eu/documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md
T
admin 877fcd2a38
gates / gates (push) Successful in 16s
R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.

R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.

R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.

Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.

R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.

Ceiling R-366 -> R-367.
2026-08-22 10:11:46 +02:00

600 lines
41 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night)
**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no
image built, nothing deployed.** Evidence:
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
---
## 1. THE VERDICT
**It is a mixture, and the drill's three options are all present — but they belong to different
faults, and only one of them explains the thing you actually saw.**
### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED**
Reproduced independently tonight, and it agrees with what R-353 already recorded:
1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data
drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** —
about an app that was installed, deployed, running and healthy.
2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt
box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it
held `compose/` and nothing else.
3. The restore returned that configuration and reported a bare completion.
So for that specific event the data was not in the thing that was restored. **R-353 called this
correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong
to think so; its text is more careful than the CHANGELOG's summary of it.
### What the drill found that nobody had seen: **THE RESTORE LOSES IT**
This is new, it is worse, and it is not the same fault:
**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.**
The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.**
Proven live on `calibre-web` at 22:23:36 with planted files:
| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore |
|---|---|---|---|---|
| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** |
| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** |
and the customer was told:
> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult.
> Ennek az alkalmazásnak nincs adatbázisa."
Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence
does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue
apps that declare no data drive, that leg is the entire dataset.**
**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit`
(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the
unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local**
restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring
PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of
them replays it.**
### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database
`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always
has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to
`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack
that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the
safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and
then says:
> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**"
The controller had dumped that database five minutes earlier.
**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive.
---
## 2. THE FULL/EMPTY CONTRAST
**Apps chosen, and why.** From the two storage classes: **`calibre-web`** declares a data drive
(`needs_hdd: true`, `backup.userdata: media/books class: mandatory`) and **`privatebin`** /
**`opengist`** declare none (the 40-class; data lives in a Docker named volume). I ran **three** cases
rather than two, deliberately: FULL and EMPTY alone cannot separate *"it was empty"* from *"the class
is broken"*, so `privatebin` was run FULL as the disambiguator.
**The fixture** — 5 files, two with Hungarian accented names, recorded as raw bytes:
```
SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991
binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a
nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4
plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1
árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea
name bytes (UTF-8 NFC):
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
```
**The comparator was proved able to convict before it was trusted.** One byte flipped at offset
500 000 of `binary-1mb.bin` (`af` → `00`): `sha256sum -c` reported `binary-1mb.bin: FAILED`, rc=1,
while the other four passed; the unmodified set passed rc=0. The mutant was discarded.
### The results
| case | class | data leg | in unit | in off-site snapshot | checking folder | off-site restore | local restore |
|---|---|---|---|---|---|---|---|
| **calibre-web FULL** | drive | user files | n/a | 5/5 identical | 5/5 identical | **5/5 identical** | — |
| **calibre-web FULL** | drive | named volume 1.42 MB | yes | yes | yes | **NOT restored** | — |
| **privatebin FULL** | no drive | named volume 1.06 MB | yes | 5/5 identical | 5/5 identical | **REFUSED — false reason** | **5/5 identical** |
| **opengist EMPTY** | no drive | named volume 181 KB skeleton | yes | yes | yes | **REFUSED — false reason** | — |
**The contrast decides it.** The unit and the off-site snapshot **hold the data, verified by
identity**, in both classes, accented filenames included. So the capture is sound. The failure is
entirely in the last leg.
**Both messages, verbatim:**
- FULL, drive class → `ok=true`, *„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16)
— az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."*
- FULL and EMPTY, no-drive class → `ok=false`, *„a(z) privatebin nincs telepítve, ezért nincs hová
visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra
az alkalmazást (Alkalmazások) ugyanerre a helyre…"*
The second is the important one. **FULL and EMPTY got the identical sentence**, so the message cannot
distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is
installed, and the instruction ("reinstall it to the same place") **cannot be followed**, because a
40-class app is offered no storage field at deploy time (that is R-352's own measurement).
**This closes R-353's second instruction.** It asked for proof that a named-volume app reaches the
off-site tier rather than inference from gate order. It does. Manifests read tonight:
```
privatebin volume_dumps = ['privatebin_privatebin_data.tar'] db_dumps = None
opengist volume_dumps = ['opengist_opengist_data.tar'] db_dumps = None
kimai volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar']
db_dumps = ['kimai-mariadb.sql']
calibre-web volume_dumps = ['calibre-web_calibre_web_config.tar'] db_dumps = None
paperless-ngx volume_dumps = [3 tars] db_dumps = None ← the defect
```
---
## 3. PART 0 — did the floor move the machine?
**Yes, unaided, in 17 seconds.**
| | |
|---|---|
| moved from → to | **0.216.0 → 0.217.0** |
| who initiated | **the hub** — the operator saved the floor at `19:48:37Z`; a poke reached the agent from `10.77.0.1:58093` at `19:48:32Z` for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision. |
| how long | floor saved `19:48:37Z` → *"controller-swap: new controller healthy"* `19:48:54Z` = **17 s**. Swap requested `19:48:41Z` → healthy = 13 s. |
| **the container's own tag on the box** | **`gitea.dooplex.hu/admin/felhom-controller:0.217.0`**, created `2026-08-21 19:48:45 UTC` — read from `docker ps` in guest 9201, not from the hub. |
The hub's own host page for `demo-hp` shows the guest's **Controller column as „—"** — the hub does
not know which controller version the box runs, while the `controller_updated` event it received says
exactly that. Two hub surfaces, one blind.
`demo-felhom` also runs 0.217.0, but it started it at `19:31:04Z` — **17 minutes before the floor was
saved** — and emitted `controller_started` with **no `controller_updated`**. It was moved by hand
during the golden bake, not by the floor.
---
## 4. PART 2 — the four answers, from source
**1. What puts a dump into a backup unit, and when? Which apps qualify?**
- **Database dumps** — `runDBDumps` (`backup/backup.go:444-560`). Qualification is **a running
container whose image matches a database image**, mapped to a stack by `deriveStackName`
(`appbackup/dbdump.go:770`). Written to `<nsRoot>/backups/primary/<stack>/db-dumps/<stack>-<type>.sql`.
- **Volume dumps** — `runVolumeDumps` (`backup/backup.go:607`). Qualification is **a deployed,
non-protected stack with at least one Docker *named volume*** (`GetDockerVolumes`). An app with
zero named volumes is skipped silently and is never stopped.
- **When** — one cycle, both legs, scheduled `db-dump` daily at **02:30 CEST**; and again as the
coherence pre-phase of every off-site run (`offbox.go:938-951`), which is what makes a snapshot an
internally coherent {DB@T, files@T} pair.
- `CaptureRecoveryUnit` **only enumerates what is already on disk**
(`recovery_unit.go:131-132`). It writes no dump itself.
- **The gap this leaves:** an app whose data is a *bind mount* and which has no database container
produces neither leg. Its unit is configuration only — and nothing anywhere says so.
**2. What should a correct backup contain?**
- **Declares a data drive** (13 of 53): the recovery unit (compose + `app.yaml` carrying the portable
secrets + whatever dumps exist) **plus** the paths its `.felhom.yml` `backup:` block marks
`mandatory`, appended to the restic snapshot as extra paths (`offbox_capture.go:32`).
- **Declares none** (40 of 53): **unit only**. `offboxCaptureSet` returns `(nil, nil)` when the app
has no backup block, so the off-site snapshot is the unit and nothing else. That is correct *given*
that their data is inside the unit's volume tars — and tonight confirmed it is.
**3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.**
`ReconstituteFromOffsite` returns `res` and `nil`. Nothing in it ever asks whether anything was
restored. The handler then calls `EndRestoreOp(**true**, reconstituteOutcomeMsg(...))`
(`web/offbox_handlers.go:455`). `reconstituteOutcomeMsg` (`web/offbox_handlers.go:463-475`) formats
counters:
```go
if res.DBsReplayed == 0 {
return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+
"Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when)
}
```
So with `FilesPlaced == 0` and `DBsReplayed == 0` the customer reads **"0 files restored — the
application restarted"** under a **success**. **This is a finding on its own and is recorded whichever
way the rest goes.** Two aggravations found on top of it:
- **"This application has no database" is asserted from a counter, not from a fact.**
`reimportDBDumpsFrom` returns `(0, nil)` when the dump directory is absent *or* holds no `.sql`
(`restore_db.go:30-46`), so an app that certainly has a database is told it has none. Proven live on
`paperless-ngx`.
- **The file count is rsync's transfer count, not a restore count.** A correct restore of unchanged
data reports **"0 fájl visszaállítva"** — indistinguishable from a restore that did nothing.
Observed: the same app reported 5, then 2, then 0 across three runs.
**4. What did the 9 August off-site snapshot contain, and is it still readable?**
**Readable, and it contained the data.** Three snapshots at `2026-08-09 08:30`:
| id | app | contents |
|---|---|---|
| `41c830db` | calibre-web | unit + `userdata/media/books` incl. the 9-Aug rehearsal fixture (`árvíztűrő-tükörfúrógép.txt`, `őszibarack.md`, `binary-3mb.bin`, `plain.txt`) |
| `9e38b84c` | opengist | unit + **`volume-dumps/opengist_opengist_data.tar`, 182 272 B** — manifest records `volume_dumps: ["opengist_opengist_data.tar"]` |
| `78b93f04` | privatebin | unit + `volume-dumps/privatebin_privatebin_data.tar`, 2 560 B |
All still readable tonight; `restic check` over the whole repository reports **"no errors were
found"**. **So the 9 August off-site copy of OpenGist's data exists and is intact** — it simply was
not the thing the afternoon's restore read, and the off-site route that would have read it refuses
for that class of app.
**The repository had also been dead since 9 August** and nothing said so: no snapshot between
`2026-08-09 08:30` and tonight. The cause is visible in the hub event at 16:01 — the off-box target
was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target,
**every per-app off-site toggle was still off**, so the first run I triggered reported:
```
[offbox] backup run started (0 app(s) toggled)
[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s
```
**Credit where it is due:** the *card* does not lie about this. It reads „Aktív — nincs kijelölt
alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still
leads, and the log line alone says only „backup OK".
---
## 5. THE TWO CYCLES COMPARED — AND THEY AGREE
**By hand:** local cycle 22:12, off-site 22:17 (after enabling the per-app switches the rebuild had
silently cleared). **By the clock:** the box's own `db-dump` at 02:30 CEST, cross-drive at 03:30,
off-site at 04:15.
| | manual | scheduled | agree? |
|---|---|---|---|
| units refreshed | all 6 apps | all 6 apps, `00:30Z` | **yes** |
| `privatebin` volume tar | 1 055 744 B | 1 055 744 B | **yes** |
| `opengist` volume tar | 181 248 B | 181 248 B | **yes** |
| `calibre-web` config tar | 1 422 848 B | 368 640 B † | **yes, and explained** |
| `paperless-ngx` `db_dumps` | **absent** | **absent** | **yes — the defect reproduces on the scheduled path** |
| the orphan `…/primary/paperless/db-dumps/` | written | **written again, 294 936 B, `00:30Z`** | **yes** |
| planted files on the drive | 5/5 identical | **5/5 identical**, plus `POST-SNAPSHOT.txt` | **yes** |
| off-site run | ok, 27 → snapshots | ok, `02:17:12Z`, **2m8s**, 42.8 MB, **27 snapshots** | **yes** |
† the tar shrank because the drill had by then removed the planted 1 MB from that volume — the
expected value, not a discrepancy.
**Two independent observations, no disagreement.** The one that matters: **R-355 is not an artefact of
my manual triggering.** The box, unattended, on its own schedule, wrote paperless-ngx's PostgreSQL
dump into a directory for a stack that does not exist and left the app's own unit recording
`db_dumps: null`.
The scheduled cross-drive (Tier-2) leg also completed for all six apps at 01:30Z — `crossdrive_completed`
events 3019–3024.
---
## 6. PART 4 — everything attempted, and the message judged
| # | test | outcome | message judged |
|---|---|---|---|
| 1a | **destructive restore: nothing is ever deleted** | **PASS both ways.** A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. | count is rsync transfers, not files restored — see §4.3 |
| 1b | **safety dump taken and verified before the stop** | **PASS.** `21:02:47Z` dump written → `21:02:47Z` `StopStack romm`. | correct: „0 fájl **és az adatbázis** visszaállítva" |
| 1b | **…and the whole operation refuses if it cannot be** | **PASS.** Made impossible by putting a regular file at the `db-dumps` path. Refused; `plain.txt` kept its post-snapshot mutation; the container's `StartedAt` was unchanged — **the app was never stopped.** | honest and names the path, but leaks a raw Go `mkdir … not a directory` into a customer surface |
| 1b | **the hole in it** | **FAIL.** For `paperless-ngx` the discovery resolves the database to the wrong stack, so `hasDB` is false: **no safety dump is taken and the refusal cannot fire.** The undo is absent, not refused. | „Ennek az alkalmazásnak nincs adatbázisa" — false |
| 2 | **end of the abandonment countdown** | state created and overdue; **fires at 05:10 CEST** — see §11 | card states a **past** date in the future tense while overdue |
| 3 | **damaged store** | **MIXED — see below** | honest at the point of failure, then forgotten |
| 4 | **drive pulled mid-restore** | restore failed, nothing written to the wrong place, **the agent re-bound the drive within 5 s** (`23:15:29` pulled → `23:15:34` re-bound) | **WRONG DIAGNOSIS.** „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions. |
| 5 | **full disk, non-destructive path** | **PASS — refuses before it starts.** | **exemplary:** „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named. |
| 5 | **full disk, destructive path** | **FAIL — no gate at all.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the three headroom gates are all on non-destructive paths (`offbox_restore.go:231,297,423`). It stopped the app, failed halfway on ENOSPC, left `DRILL-2026-08-21/` holding 2 of 5 entries, and restarted the app. | honest (`No space left on device (28)`) but raw rsync output |
| 6 | **controller killed mid-restore** | **PASS.** SIGKILL inside the stop→restore→start window. On restart: *"[appstop] crash recovery: an off-site restore (op \"offbox-reconstitute:paperless-ngx\") was interrupted and left 1 app(s) stopped — restarting them"*. App restarted, marker cleared, **and the hub was told** (event 3009). | good — names the operation, not just the app |
| 7 | **the 40-class under pressure** | **the backup noticed; the alarm did not.** Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The **reserve refused per app**: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got `recovery_unit_capture_failed` (**error**) naming the filesystem and its numbers. **The fill watcher said nothing** — it runs **once a day at 03:30** plus once at startup (`cmd/controller/main.go:1092`). | the refusal messages are good; the silence is the problem |
**Where the clock stopped me:** nothing in Part 4 was skipped for time. Item 2's terminal deletion is
scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only.
### 4.3 in detail — the damaged store
One byte flipped inside pack `967853d2…` at offset 5 000 000, over the repository's own SFTP
transport. Pack files are named by their content hash, so this is genuine corruption.
- **`restic check` detects it** — *"ciphertext verification failed"*, *"Fatal: repository contains
errors"*.
- **But nothing in the product ever runs it.** The only restic verbs in the entire controller are
`restore, snapshots, backup, unlock, stats, init, forget, prune, cat`. The agent's restore-test is
**PBS-tier only**. **The off-site store is never verified by any layer, at any time.** Corruption is
discovered at restore time — the worst possible moment.
- **A restore that touches the damage fails honestly:** `ok=false`, naming the file and
*"ciphertext verification failed"*.
- **But the failure is not remembered.** It left a **partial** checking folder — 78 MB, 54 files,
15 of 16 originals. `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only *"does the
directory exist and is it non-empty"*, so the wizard then offered **all three** actions including
„Teljes visszaállítás indítása".
- **Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.**
The repository was repaired from byte-identical originals; `restic check` now reports **"no errors
were found"**.
---
## 7. PART 5 — the two rows that were observed and never filed
**5.1 — the delete guard on verification copies is blind, and its own comment says otherwise.**
`offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) gates on `backupMgr.IsRunning()` — the
**concurrency** flag — while its comment states *"It refuses while a backup/restore op is running: the
copy being deleted could be the one currently being written."* R-351b moved all seven restore handlers
onto `restoreOpBlocked()` (which reads both flags); **this handler was left behind**, and one other
site (`:239`) reads the bare flag correctly and documents why.
**Reachability is not a race — it is the whole operation.** `RestoreOffboxScratch`
(`offbox_restore.go:211`) **never calls `acquireRunning` at all**, so `IsRunning()` is false for the
entire duration of an off-site verification restore. Demonstrated live, flags read immediately before
and after the delete:
```
--- BEFORE delete 22:35:25 display= True offbox-restore kimai concurrency= False
--- DELETE calibre-web verification copy: „Az ellenőrző másolat törölve…"
--- AFTER delete 22:35:25 display= True offbox-restore kimai concurrency= False
```
The copy was removed. The guard is app-agnostic, so the same call naming the *restoring* app hits the
directory the restore is writing into. **Rank: MEDIUM** — it needs a customer to press delete during a
restore, but both controls live on the same page, the window is the whole restore, and the target is
the restore's own source. (Reconstitute and place *do* hold the flag, so the exposure is the
verification-restore window only.)
**5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight.**
Filed as an instrument defect. **Occurrences I can evidence:**
1. **2026-07-20** — an accented grep through `ssh → pct exec → bash -c` nearly produced a wrong
"banner cleared" claim. Recorded in `felhom-controller/.claude/rules/ui-hungarian.md:19-22`.
2. **2026-08-13** — `kubectl exec … sh -c "grep '<accented>'"` returned **0 for three strings that
were present**, one step from being reported as a failed hub v0.105.0 deploy.
3. **2026-08-21, tonight, 22:07** — `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`
(octal escaping). Recording the accented filenames' raw bytes from that listing would have been
wrong. Caught by extracting the archive and reading the names with `xxd`.
**Correction to the task's premise:** that is **two inside two weeks**, plus the founding case a month
earlier. I looked for a third inside the two-week window and did not find one on record.
**The smallest guard I would propose — and did NOT build:** the problem is not grep, it is that every
one of these tools *silently transforms* the bytes. So the guard is not "use ASCII fragments" (a
discipline, which is what failed three times) but **a negative control that the harness cannot skip**:
any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a
string that MUST be absent — and a zero result from the first is only reportable when the second also
returns zero *and* a third probe for a known-present ASCII anchor returns non-zero. Three probes, one
helper, no judgement required at the call site. Everything else has been tried and is what "nearly"
means in all three cases.
---
## 8. RANKED REGISTER ROWS OPENED
Ceiling was **R-353**; it **moved to R-366**.
| id | rank | what |
|---|---|---|
| **R-354** | **HIGH** | The off-site full restore has **no named-volume leg**. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. `offbox_reconstitute.go:341-346`. |
| **R-355** | **HIGH** | `paperless-ngx`'s PostgreSQL is dumped to a directory for a **non-existent stack** (`…/primary/paperless/`), so its unit records `db_dumps: null`, nothing off-sites it, **no safety dump is taken on a destructive restore**, and the customer is told the app has no database. `appbackup/dbdump.go:770-798`. One app in 53. |
| **R-356** | **HIGH** | The off-site restore **refuses for all 40 no-drive apps** with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. `offbox_reconstitute.go:208-227`. |
| **R-357** | **MEDIUM** | The **destructive** restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory. |
| **R-358** | **MEDIUM** | A **failed** scratch restore leaves a partial copy that `OffboxFullScratchReady` reports as ready; the destructive restore then runs from it and reports success. |
| **R-359** | **MEDIUM** | The off-site restic store is **never verified** by any layer — `restic check` is not among the verbs the controller runs, and the agent's restore-test is PBS-only. |
| **R-360** | **MEDIUM** | Verification-copy delete gates on `IsRunning()`, which `RestoreOffboxScratch` never holds — deletable throughout a restore. Its comment asserts the opposite. **(Part 5.1)** |
| **R-361** | **MEDIUM** | The safety dump **overwrites the unit's own DB dump**: `DumpOne` writes the canonical `<stack>-<type>.sql` and only then renames it away. The comment at `offbox_reconstitute.go:147-148` states it "can never overwrite the app's real dump". Proven: romm's `romm-mariadb.sql` was present before and absent after. |
| **R-362** | **MEDIUM** | A data drive detached mid-restore is reported as **„permission denied"**. The restore path never consults drive state. |
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
| **R-366** | **HIGH** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives too** — the box cannot read its own pre-reinstall backups (`manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`, hub event 3016, filed unprompted at 21:59). The PBS-tier analogue of R-193: a rebuilt box loses **both** off-premises tiers at once. The restore-test caught it precisely; the gap is that it is called *a failed restore test* rather than *your older whole-guest backups are unreadable*. **Found incidentally — nobody was looking for it.** |
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
nobody. Observed tonight as event 3006, severity `info`. A sweep of every `emit(` call site shows this
is now the **only** remaining instance fleet-wide.
**R-353:** its instruction (2) is **satisfied** — see §2. Its instruction (1) stands and is now
strictly larger than when it was written, because R-354 shows the bare completion can also be reported
over a unit that *did* have a data leg.
---
## 9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH
| time (CEST) | what |
|---|---|
| 21:57 | start; baselines |
| 22:00 | break-glass into `demo-hp` |
| 22:09 | fixture planted, comparator control passed |
| 22:12 | manual local backup |
| 22:13 | manual off-site run — **0 apps toggled** |
| 22:17 | off-site run with 3 apps enabled |
| 22:19–22:20 | checking-folder restores |
| 22:21 | reconstitute privatebin → refused |
| 22:23 | reconstitute calibre-web → **the conviction** |
| 22:25 | local restore-from-unit privatebin → data returned |
| 22:27 | Part 4.1a — no-delete invariant |
| 22:39–22:45 | paperless-ngx → R-355 |
| 22:51–22:57 | Part 4.3 damaged store; repo repaired |
| 23:02–23:04 | Part 4.1b safety dump, both directions |
| 23:10–23:13 | Part 4.5 full disk (both paths); Part 4.6 controller killed mid-restore |
| 23:15 | Part 4.4 drive pulled |
| 23:17–23:26 | Part 4.7 filesystem filled, backup reserve observed, filesystem freed |
| 23:08–23:09 | abandonment set-aside store created; countdown written and controller restarted (fires 05:10) |
**Steps off the customer's path, named:**
1. **Break-glass root access** to `demo-hp` via the hub-vaulted `host_recovery` credential — the box
had lost DooPlex's SSH key (its `authorized_keys` held only its own `root@demo-hp` RSA key) and its
tailnet address was unreachable. DooPlex's public key was **re-added** to `/root/.ssh/authorized_keys`
and an `ssh` alias `hp` → `192.168.0.104` was added to `~/.ssh/config` on DooPlex.
2. **A hub DB snapshot** (`hub.db` + `-wal` + `-shm`) was streamed to the scratchpad to read
`host_recovery` and the events table. It holds every host's secret; it is in the session scratchpad
only and is not in any committed file.
3. **`settings.json` was edited directly** (controller stopped, backup at `/root/settings.json.drill-backup`)
to create the overdue abandonment state. There is no product path to shorten a countdown, and the
real orphan→reset path would have destroyed `demo-hp`'s entire off-site history.
4. **The set-aside store the sweep will delete was created by hand** at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, for the same reason.
5. **Two restic pack files were deliberately corrupted and then restored** from byte-identical copies.
6. **`fallocate` fillers** were used to fill two filesystems and were removed.
7. Apps deployed for the drill: `opengist` (re-deployed empty), `kimai`, `paperless-ngx`, `romm`.
---
## 10. THE FENCES
- **`ep0` / the off-site endpoint — untouched outside this machine's own path, and here is how I know.**
`demo-hp` authenticates as the Hetzner Storage Box **sub-account `u629488-sub3`**, which is chrooted
to its own home: `ls /` returns **`Permission denied`**, and `/home` contains exactly `.ssh` and
`felhom-repo`. Every write, the two corruptions, the set-aside store and the sweep's target are
inside that home. **`demo-felhom` is a different sub-account (`u629488-sub1`)** and a real customer's
copy is a different sub-account again — none reachable with this key. `restic forget`/`prune` were
never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept
every 9-August snapshot (verified by listing all 24, not by a count).
- **`demo-felhom`'s two fixtures — confirmed intact, not assumed.**
(a) the unopenable set-aside store `u629488-sub1:/home/felhom-repo.orphaned-20260810` — listed
tonight, `config`/`data` (258 shards)/`index`/`keys`/`locks`/`snapshots`, mtimes still 18 Jul and
3 Aug; (b) the retained-key case — hub `host_escrow_superseded` rows **11 and 12** for
`demo-felhom-8363b5`, each with a 572-byte `identity_blob`, dated 2026-08-12. `demo-felhom` is
healthy on 0.217.0, its own off-site ran at 02:15 with `last_status: ok`, and
**`--abandon-status` there reports „no abandonment countdown is running on this box"**.
- **`peti-felhom` — not contacted.** It does not appear in the hub host list at all; no command in this
session named it.
---
## 11. PART 4.2 — THE ABANDONMENT COUNTDOWN, WATCHED FIRING
**The terminal deletion has now been observed.** It did exactly what it claims, including the half
nobody had seen.
- **State created** 23:08–23:09 CEST. A realistic set-aside store was built at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, the countdown written overdue
(started 2026-08-07, due 2026-08-20), and the product's own CLI confirmed it:
`abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0`.
- **05:10 CEST — it fired.** `/home` on the storage box is stamped `03:10Z`; the set-aside store is
**gone**; **`/home/felhom-repo`, the live repository, is untouched** (mtime still 4 Aug).
- **05:13 CEST — the hub half.** Event **3025 `offsite_abandon_purged`**: *„Az ügyfél korábbi távoli
mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is törölve (1 csomag). Az ügyfél
döntése alapján, a 14 napos türelmi idő lejárta után."* And demo-hp's superseded escrow row **is
gone from the hub** — it held one before.
- **The controller closed itself out.** Every `abandon_*` field has been removed from `settings.json`
and `--abandon-status` reports *"no abandonment countdown is running on this box"*.
**So the two-phase commit's central promise — *"it removes BOTH halves or neither"* — is confirmed
live for the first time:** the ciphertext and the sealed package that protects it went together, three
minutes apart, and the state that remembered the operation cleaned itself up. **What it removes:**
exactly the recorded set-aside path. **What survives:** the live repository, the live escrow, and the
box's current recovery path.
**Judged:** the completion message is accurate and in plain Hungarian. The only wrong note is while
the countdown is *overdue but not yet swept* — the card then states a past date in the future tense
(**R-365**).
**No countdown is left running anywhere.** `demo-hp` cleared itself; `demo-felhom` reports none.
---
## 11b. WHAT THE NIGHT FOUND THAT NOBODY WAS LOOKING FOR
At 21:59 the box filed, unprompted, hub event 3016:
> `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could
> not be restored+booted … wrong key — manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided
> key dd:d1:d8:53:44:62:5e:0b`
That archive predates the 21 August reinstall by three days. **The rebuild orphaned the PBS whole-guest
archives exactly as it orphaned the restic repository** — so a rebuilt box loses *both* off-premises
tiers at once. The restore-test mechanism deserves credit: it caught it and named the key mismatch
precisely. The gap is what it is *called* — "a restore test failed" reads as a flaky verification, not
as "every whole-guest backup taken before the reinstall is unreadable on this machine". Filed **R-366
(HIGH)**.
---
## 12. MACHINE STATES AT THE END
**`demo-hp` — HEALTHY, not broken. Nothing needs bringing back.** At 05:28 CEST: all 15 containers up
and healthy (`felhom-controller` 0.217.0, `traefik`, `cloudflared`, `filebrowser`, plus the six drill
apps); `/` 4%, `/var/lib/felhom` 13%, the data drive 1%; **no filler files left**, **no app-stop marker**,
**no operation in flight**, **no countdown running**, and `restic check` over its off-site repository
reports **no errors were found**. Its off-site backup, silent since 9 August, is working again and ran
on its own schedule at 04:15.
**Deliberately left in place, each with a reason** (retained, not forgotten):
| left behind | why |
|---|---|
| `opengist`, `privatebin`, `calibre-web` deployed with the planted `DRILL-2026-08-21` fixture | the reproduction for **R-354** and **R-356**; the hashes in §2 make the fix verifiable without rebuilding the case |
| `paperless-ngx` deployed | the **only** reproduction of **R-355**, and it regenerates the orphan directory on every nightly cycle |
| `romm` deployed | the only app on the box where the safety dump actually works — the fixture for **R-361** and the control for R-355 |
| `kimai` deployed | a second correct-derivation DB app, the negative control for R-355 |
| five verification copies under `backups/offsite-restore/` | harmless, and they are the **R-358** fixture |
| `romm`'s `pre-restore-…sql` in its unit | the physical evidence for **R-361** |
| DooPlex's public key in `demo-hp:/root/.ssh/authorized_keys`, and the `hp` alias in `~/.ssh/config` | the box had lost the key and its tailnet address is dead; without it the next session must go through break-glass again |
**Removed / restored during the drill:** both corrupted restic packs (repo verified clean), both
`fallocate` fillers, the set-aside store (by the sweep, as intended), and demo-hp's superseded escrow
row (by the sweep's hub half, as intended).
**`demo-hp`'s tailnet address `100.76.96.79` is still unreachable** and was not repaired — the box is
reachable on the LAN at `192.168.0.104`. That is the one thing about it that is worse than it should
be, and it predates tonight.
**`demo-felhom` — untouched and healthy**, controller 0.217.0, its own off-site ran at 02:15 with
`last_status: ok`, both fixtures verified present (§10), no countdown running.
**`drill-r50`** — not used. Still `DOWN` in the hub, agent 0.129.0, as it was.
---
## 13. TEARDOWN — ALL FOUR LAYERS, STATED
| layer | state |
|---|---|
| **Off-site (storage box)** | The set-aside store I created was **deleted by the product's own sweep**, as designed. The two corrupted packs were **restored byte-identical** and `restic check` passes. Nothing else was written. `restic forget`/`prune` were never invoked by me. |
| **PVE host `demo-hp`** | `/root/settings.json.drill-backup` **retained** (the pre-drill controller settings, in case the abandonment edit needs reverting). DooPlex's SSH key **retained**, with the reason above. `/tmp/plant.tar` left; harmless. |
| **Guest 9201 / controller** | Six apps and the planted fixtures **retained with reasons** (table above). Controller state is clean: no marker, no countdown, no in-flight op. |
| **Hub — stated explicitly** | **No customer record and no host record was created, so none needs deleting.** I used the existing `demo-hp` customer and host throughout. The only hub-side *removal* was demo-hp's superseded escrow row, done by the abandonment sweep itself and reported as event 3025. Events 3002–3025 were generated as a normal consequence of the work and are left as the record. **Nothing is owed here and nothing is blocked.** |
| **DooPlex (this machine)** | The hub DB copies (which contain every host's break-glass secret) are in the session scratchpad only, never in a committed file, and are shredded in the closing step. |
---
## 14. OBSERVATIONS — noticed, not acted on
- **The hub does not display the controller version it is told.** `demo-hp`'s host page shows the
guest's Controller column as **„—"** hours after receiving `controller_updated: 0.216.0 → 0.217.0`.
Two hub surfaces, one blind. Not filed — `REPORT-hub-blindness.md` already exists and this may be
part of it.
- **`demo-felhom` moved to 0.217.0 without a `controller_updated` event** (started 19:31Z, 17 minutes
before the floor was saved). Consistent with a by-hand deployment during the golden bake, not a
defect — recorded so a later reader does not mistake it for one.
- **`restic --latest N` is per-group, not a total.** It briefly read as "18 snapshots became 10" and
would have been reported as data loss. Caught by listing all of them and by an independent
`restic stats` (26.44 MiB / 588 blobs). **An unpaginated listing is not a total** — the rule earned
its place again.
- **`find -newermt` is the wrong probe for a restore**: restic preserves the snapshot's mtimes, so a
freshly restored tree looks old. It briefly read as "nothing was restored".
- **`sftp -b` aborts on the first failing line**, so a batch listing 256 shard directories returned 7
packs and looked like a total. Prefixing each line with `-` fixed it. Same family as the two above.
---
## 15. WHAT WAS DROPPED, PLAINLY
- **Nothing in Parts 0, 2, 3, 4 or 5 was skipped.** Every Part 4 item was attempted and reached a
verdict.
- **The fill-watch alarm was not watched firing on its own schedule.** I proved its cadence from
source (daily 03:30) and observed it fire from the startup path (`disk_critical`, 21:10:58Z, correct
Hungarian copy naming the drive and the free space). Holding a filesystem at 99% for six hours would
have sat across the scheduled backup and the abandonment sweep, and I judged the scheduled cycle —
which the task asks for explicitly — worth more than a second sighting of an alarm whose trigger I
had already read and seen work.
- **I did not drive the abandonment through the real orphan→reset path**, because that path
move-asides the *live* repository, which would have destroyed demo-hp's whole off-site history
including the 9 August snapshots this drill exists to read. The sweep's own code path was exercised
in full on a real store; the deviation is only in how the state was created.
- **`peti-felhom` was not contacted**, per the fence.