d895d9f7dd
gates / gates (push) Successful in 16s
The 2026-08-04 row claimed "a customer's file survives a machine rebuild and comes back — the whole off-site story, end to end". Tonight's drill shows that holds for the declared-userdata leg of a drive-declaring app and for nothing else: the off-site restore has no named-volume leg (R-354), and refuses outright for the 40 apps that declare no data drive (R-356). Since that class keeps ALL its data in named volumes, the end-to-end story is unproven there and disproven for the volume leg generally. The escrow/key half of the row is untouched and still stands. STATUS.md also corrects the fleet pair it still named (0.214.0/0.129.0 -> 0.217.0/0.130.0) and records that demo-hp's off-site had been silent since 9 August.
452 lines
30 KiB
Markdown
452 lines
30 KiB
Markdown
# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night)
|
||
|
||
**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no
|
||
image built, nothing deployed.** Evidence:
|
||
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
|
||
|
||
---
|
||
|
||
## 1. THE VERDICT
|
||
|
||
**It is a mixture, and the drill's three options are all present — but they belong to different
|
||
faults, and only one of them explains the thing you actually saw.**
|
||
|
||
### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED**
|
||
|
||
Reproduced independently tonight, and it agrees with what R-353 already recorded:
|
||
|
||
1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data
|
||
drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** —
|
||
about an app that was installed, deployed, running and healthy.
|
||
2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt
|
||
box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it
|
||
held `compose/` and nothing else.
|
||
3. The restore returned that configuration and reported a bare completion.
|
||
|
||
So for that specific event the data was not in the thing that was restored. **R-353 called this
|
||
correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong
|
||
to think so; its text is more careful than the CHANGELOG's summary of it.
|
||
|
||
### What the drill found that nobody had seen: **THE RESTORE LOSES IT**
|
||
|
||
This is new, it is worse, and it is not the same fault:
|
||
|
||
**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.**
|
||
The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.**
|
||
|
||
Proven live on `calibre-web` at 22:23:36 with planted files:
|
||
|
||
| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore |
|
||
|---|---|---|---|---|
|
||
| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** |
|
||
| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** |
|
||
|
||
and the customer was told:
|
||
|
||
> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult.
|
||
> Ennek az alkalmazásnak nincs adatbázisa."
|
||
|
||
Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence
|
||
does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue
|
||
apps that declare no data drive, that leg is the entire dataset.**
|
||
|
||
**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit`
|
||
(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the
|
||
unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local**
|
||
restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring
|
||
PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of
|
||
them replays it.**
|
||
|
||
### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database
|
||
|
||
`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always
|
||
has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to
|
||
`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack
|
||
that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the
|
||
safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and
|
||
then says:
|
||
|
||
> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**"
|
||
|
||
The controller had dumped that database five minutes earlier.
|
||
|
||
**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive.
|
||
|
||
---
|
||
|
||
## 2. THE FULL/EMPTY CONTRAST
|
||
|
||
**Apps chosen, and why.** From the two storage classes: **`calibre-web`** declares a data drive
|
||
(`needs_hdd: true`, `backup.userdata: media/books class: mandatory`) and **`privatebin`** /
|
||
**`opengist`** declare none (the 40-class; data lives in a Docker named volume). I ran **three** cases
|
||
rather than two, deliberately: FULL and EMPTY alone cannot separate *"it was empty"* from *"the class
|
||
is broken"*, so `privatebin` was run FULL as the disambiguator.
|
||
|
||
**The fixture** — 5 files, two with Hungarian accented names, recorded as raw bytes:
|
||
|
||
```
|
||
SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991
|
||
binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a
|
||
nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4
|
||
plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1
|
||
árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea
|
||
|
||
name bytes (UTF-8 NFC):
|
||
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
|
||
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
|
||
```
|
||
|
||
**The comparator was proved able to convict before it was trusted.** One byte flipped at offset
|
||
500 000 of `binary-1mb.bin` (`af` → `00`): `sha256sum -c` reported `binary-1mb.bin: FAILED`, rc=1,
|
||
while the other four passed; the unmodified set passed rc=0. The mutant was discarded.
|
||
|
||
### The results
|
||
|
||
| case | class | data leg | in unit | in off-site snapshot | checking folder | off-site restore | local restore |
|
||
|---|---|---|---|---|---|---|---|
|
||
| **calibre-web FULL** | drive | user files | n/a | 5/5 identical | 5/5 identical | **5/5 identical** | — |
|
||
| **calibre-web FULL** | drive | named volume 1.42 MB | yes | yes | yes | **NOT restored** | — |
|
||
| **privatebin FULL** | no drive | named volume 1.06 MB | yes | 5/5 identical | 5/5 identical | **REFUSED — false reason** | **5/5 identical** |
|
||
| **opengist EMPTY** | no drive | named volume 181 KB skeleton | yes | yes | yes | **REFUSED — false reason** | — |
|
||
|
||
**The contrast decides it.** The unit and the off-site snapshot **hold the data, verified by
|
||
identity**, in both classes, accented filenames included. So the capture is sound. The failure is
|
||
entirely in the last leg.
|
||
|
||
**Both messages, verbatim:**
|
||
|
||
- FULL, drive class → `ok=true`, *„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16)
|
||
— az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."*
|
||
- FULL and EMPTY, no-drive class → `ok=false`, *„a(z) privatebin nincs telepítve, ezért nincs hová
|
||
visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra
|
||
az alkalmazást (Alkalmazások) ugyanerre a helyre…"*
|
||
|
||
The second is the important one. **FULL and EMPTY got the identical sentence**, so the message cannot
|
||
distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is
|
||
installed, and the instruction ("reinstall it to the same place") **cannot be followed**, because a
|
||
40-class app is offered no storage field at deploy time (that is R-352's own measurement).
|
||
|
||
**This closes R-353's second instruction.** It asked for proof that a named-volume app reaches the
|
||
off-site tier rather than inference from gate order. It does. Manifests read tonight:
|
||
|
||
```
|
||
privatebin volume_dumps = ['privatebin_privatebin_data.tar'] db_dumps = None
|
||
opengist volume_dumps = ['opengist_opengist_data.tar'] db_dumps = None
|
||
kimai volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar']
|
||
db_dumps = ['kimai-mariadb.sql']
|
||
calibre-web volume_dumps = ['calibre-web_calibre_web_config.tar'] db_dumps = None
|
||
paperless-ngx volume_dumps = [3 tars] db_dumps = None ← the defect
|
||
```
|
||
|
||
---
|
||
|
||
## 3. PART 0 — did the floor move the machine?
|
||
|
||
**Yes, unaided, in 17 seconds.**
|
||
|
||
| | |
|
||
|---|---|
|
||
| moved from → to | **0.216.0 → 0.217.0** |
|
||
| who initiated | **the hub** — the operator saved the floor at `19:48:37Z`; a poke reached the agent from `10.77.0.1:58093` at `19:48:32Z` for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision. |
|
||
| how long | floor saved `19:48:37Z` → *"controller-swap: new controller healthy"* `19:48:54Z` = **17 s**. Swap requested `19:48:41Z` → healthy = 13 s. |
|
||
| **the container's own tag on the box** | **`gitea.dooplex.hu/admin/felhom-controller:0.217.0`**, created `2026-08-21 19:48:45 UTC` — read from `docker ps` in guest 9201, not from the hub. |
|
||
|
||
The hub's own host page for `demo-hp` shows the guest's **Controller column as „—"** — the hub does
|
||
not know which controller version the box runs, while the `controller_updated` event it received says
|
||
exactly that. Two hub surfaces, one blind.
|
||
|
||
`demo-felhom` also runs 0.217.0, but it started it at `19:31:04Z` — **17 minutes before the floor was
|
||
saved** — and emitted `controller_started` with **no `controller_updated`**. It was moved by hand
|
||
during the golden bake, not by the floor.
|
||
|
||
---
|
||
|
||
## 4. PART 2 — the four answers, from source
|
||
|
||
**1. What puts a dump into a backup unit, and when? Which apps qualify?**
|
||
|
||
- **Database dumps** — `runDBDumps` (`backup/backup.go:444-560`). Qualification is **a running
|
||
container whose image matches a database image**, mapped to a stack by `deriveStackName`
|
||
(`appbackup/dbdump.go:770`). Written to `<nsRoot>/backups/primary/<stack>/db-dumps/<stack>-<type>.sql`.
|
||
- **Volume dumps** — `runVolumeDumps` (`backup/backup.go:607`). Qualification is **a deployed,
|
||
non-protected stack with at least one Docker *named volume*** (`GetDockerVolumes`). An app with
|
||
zero named volumes is skipped silently and is never stopped.
|
||
- **When** — one cycle, both legs, scheduled `db-dump` daily at **02:30 CEST**; and again as the
|
||
coherence pre-phase of every off-site run (`offbox.go:938-951`), which is what makes a snapshot an
|
||
internally coherent {DB@T, files@T} pair.
|
||
- `CaptureRecoveryUnit` **only enumerates what is already on disk**
|
||
(`recovery_unit.go:131-132`). It writes no dump itself.
|
||
- **The gap this leaves:** an app whose data is a *bind mount* and which has no database container
|
||
produces neither leg. Its unit is configuration only — and nothing anywhere says so.
|
||
|
||
**2. What should a correct backup contain?**
|
||
|
||
- **Declares a data drive** (13 of 53): the recovery unit (compose + `app.yaml` carrying the portable
|
||
secrets + whatever dumps exist) **plus** the paths its `.felhom.yml` `backup:` block marks
|
||
`mandatory`, appended to the restic snapshot as extra paths (`offbox_capture.go:32`).
|
||
- **Declares none** (40 of 53): **unit only**. `offboxCaptureSet` returns `(nil, nil)` when the app
|
||
has no backup block, so the off-site snapshot is the unit and nothing else. That is correct *given*
|
||
that their data is inside the unit's volume tars — and tonight confirmed it is.
|
||
|
||
**3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.**
|
||
|
||
`ReconstituteFromOffsite` returns `res` and `nil`. Nothing in it ever asks whether anything was
|
||
restored. The handler then calls `EndRestoreOp(**true**, reconstituteOutcomeMsg(...))`
|
||
(`web/offbox_handlers.go:455`). `reconstituteOutcomeMsg` (`web/offbox_handlers.go:463-475`) formats
|
||
counters:
|
||
|
||
```go
|
||
if res.DBsReplayed == 0 {
|
||
return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+
|
||
"Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when)
|
||
}
|
||
```
|
||
|
||
So with `FilesPlaced == 0` and `DBsReplayed == 0` the customer reads **"0 files restored — the
|
||
application restarted"** under a **success**. **This is a finding on its own and is recorded whichever
|
||
way the rest goes.** Two aggravations found on top of it:
|
||
|
||
- **"This application has no database" is asserted from a counter, not from a fact.**
|
||
`reimportDBDumpsFrom` returns `(0, nil)` when the dump directory is absent *or* holds no `.sql`
|
||
(`restore_db.go:30-46`), so an app that certainly has a database is told it has none. Proven live on
|
||
`paperless-ngx`.
|
||
- **The file count is rsync's transfer count, not a restore count.** A correct restore of unchanged
|
||
data reports **"0 fájl visszaállítva"** — indistinguishable from a restore that did nothing.
|
||
Observed: the same app reported 5, then 2, then 0 across three runs.
|
||
|
||
**4. What did the 9 August off-site snapshot contain, and is it still readable?**
|
||
|
||
**Readable, and it contained the data.** Three snapshots at `2026-08-09 08:30`:
|
||
|
||
| id | app | contents |
|
||
|---|---|---|
|
||
| `41c830db` | calibre-web | unit + `userdata/media/books` incl. the 9-Aug rehearsal fixture (`árvíztűrő-tükörfúrógép.txt`, `őszibarack.md`, `binary-3mb.bin`, `plain.txt`) |
|
||
| `9e38b84c` | opengist | unit + **`volume-dumps/opengist_opengist_data.tar`, 182 272 B** — manifest records `volume_dumps: ["opengist_opengist_data.tar"]` |
|
||
| `78b93f04` | privatebin | unit + `volume-dumps/privatebin_privatebin_data.tar`, 2 560 B |
|
||
|
||
All still readable tonight; `restic check` over the whole repository reports **"no errors were
|
||
found"**. **So the 9 August off-site copy of OpenGist's data exists and is intact** — it simply was
|
||
not the thing the afternoon's restore read, and the off-site route that would have read it refuses
|
||
for that class of app.
|
||
|
||
**The repository had also been dead since 9 August** and nothing said so: no snapshot between
|
||
`2026-08-09 08:30` and tonight. The cause is visible in the hub event at 16:01 — the off-box target
|
||
was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target,
|
||
**every per-app off-site toggle was still off**, so the first run I triggered reported:
|
||
|
||
```
|
||
[offbox] backup run started (0 app(s) toggled)
|
||
[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s
|
||
```
|
||
|
||
**Credit where it is due:** the *card* does not lie about this. It reads „Aktív — nincs kijelölt
|
||
alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still
|
||
leads, and the log line alone says only „backup OK".
|
||
|
||
---
|
||
|
||
## 5. THE TWO CYCLES COMPARED
|
||
|
||
*(filled in after the scheduled run — see §11)*
|
||
|
||
---
|
||
|
||
## 6. PART 4 — everything attempted, and the message judged
|
||
|
||
| # | test | outcome | message judged |
|
||
|---|---|---|---|
|
||
| 1a | **destructive restore: nothing is ever deleted** | **PASS both ways.** A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. | count is rsync transfers, not files restored — see §4.3 |
|
||
| 1b | **safety dump taken and verified before the stop** | **PASS.** `21:02:47Z` dump written → `21:02:47Z` `StopStack romm`. | correct: „0 fájl **és az adatbázis** visszaállítva" |
|
||
| 1b | **…and the whole operation refuses if it cannot be** | **PASS.** Made impossible by putting a regular file at the `db-dumps` path. Refused; `plain.txt` kept its post-snapshot mutation; the container's `StartedAt` was unchanged — **the app was never stopped.** | honest and names the path, but leaks a raw Go `mkdir … not a directory` into a customer surface |
|
||
| 1b | **the hole in it** | **FAIL.** For `paperless-ngx` the discovery resolves the database to the wrong stack, so `hasDB` is false: **no safety dump is taken and the refusal cannot fire.** The undo is absent, not refused. | „Ennek az alkalmazásnak nincs adatbázisa" — false |
|
||
| 2 | **end of the abandonment countdown** | state created and overdue; **fires at 05:10 CEST** — see §11 | card states a **past** date in the future tense while overdue |
|
||
| 3 | **damaged store** | **MIXED — see below** | honest at the point of failure, then forgotten |
|
||
| 4 | **drive pulled mid-restore** | restore failed, nothing written to the wrong place, **the agent re-bound the drive within 5 s** (`23:15:29` pulled → `23:15:34` re-bound) | **WRONG DIAGNOSIS.** „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions. |
|
||
| 5 | **full disk, non-destructive path** | **PASS — refuses before it starts.** | **exemplary:** „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named. |
|
||
| 5 | **full disk, destructive path** | **FAIL — no gate at all.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the three headroom gates are all on non-destructive paths (`offbox_restore.go:231,297,423`). It stopped the app, failed halfway on ENOSPC, left `DRILL-2026-08-21/` holding 2 of 5 entries, and restarted the app. | honest (`No space left on device (28)`) but raw rsync output |
|
||
| 6 | **controller killed mid-restore** | **PASS.** SIGKILL inside the stop→restore→start window. On restart: *"[appstop] crash recovery: an off-site restore (op \"offbox-reconstitute:paperless-ngx\") was interrupted and left 1 app(s) stopped — restarting them"*. App restarted, marker cleared, **and the hub was told** (event 3009). | good — names the operation, not just the app |
|
||
| 7 | **the 40-class under pressure** | **the backup noticed; the alarm did not.** Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The **reserve refused per app**: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got `recovery_unit_capture_failed` (**error**) naming the filesystem and its numbers. **The fill watcher said nothing** — it runs **once a day at 03:30** plus once at startup (`cmd/controller/main.go:1092`). | the refusal messages are good; the silence is the problem |
|
||
|
||
**Where the clock stopped me:** nothing in Part 4 was skipped for time. Item 2's terminal deletion is
|
||
scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only.
|
||
|
||
### 4.3 in detail — the damaged store
|
||
|
||
One byte flipped inside pack `967853d2…` at offset 5 000 000, over the repository's own SFTP
|
||
transport. Pack files are named by their content hash, so this is genuine corruption.
|
||
|
||
- **`restic check` detects it** — *"ciphertext verification failed"*, *"Fatal: repository contains
|
||
errors"*.
|
||
- **But nothing in the product ever runs it.** The only restic verbs in the entire controller are
|
||
`restore, snapshots, backup, unlock, stats, init, forget, prune, cat`. The agent's restore-test is
|
||
**PBS-tier only**. **The off-site store is never verified by any layer, at any time.** Corruption is
|
||
discovered at restore time — the worst possible moment.
|
||
- **A restore that touches the damage fails honestly:** `ok=false`, naming the file and
|
||
*"ciphertext verification failed"*.
|
||
- **But the failure is not remembered.** It left a **partial** checking folder — 78 MB, 54 files,
|
||
15 of 16 originals. `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only *"does the
|
||
directory exist and is it non-empty"*, so the wizard then offered **all three** actions including
|
||
„Teljes visszaállítás indítása".
|
||
- **Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.**
|
||
|
||
The repository was repaired from byte-identical originals; `restic check` now reports **"no errors
|
||
were found"**.
|
||
|
||
---
|
||
|
||
## 7. PART 5 — the two rows that were observed and never filed
|
||
|
||
**5.1 — the delete guard on verification copies is blind, and its own comment says otherwise.**
|
||
`offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) gates on `backupMgr.IsRunning()` — the
|
||
**concurrency** flag — while its comment states *"It refuses while a backup/restore op is running: the
|
||
copy being deleted could be the one currently being written."* R-351b moved all seven restore handlers
|
||
onto `restoreOpBlocked()` (which reads both flags); **this handler was left behind**, and one other
|
||
site (`:239`) reads the bare flag correctly and documents why.
|
||
|
||
**Reachability is not a race — it is the whole operation.** `RestoreOffboxScratch`
|
||
(`offbox_restore.go:211`) **never calls `acquireRunning` at all**, so `IsRunning()` is false for the
|
||
entire duration of an off-site verification restore. Demonstrated live, flags read immediately before
|
||
and after the delete:
|
||
|
||
```
|
||
--- BEFORE delete 22:35:25 display= True offbox-restore kimai concurrency= False
|
||
--- DELETE calibre-web verification copy: „Az ellenőrző másolat törölve…"
|
||
--- AFTER delete 22:35:25 display= True offbox-restore kimai concurrency= False
|
||
```
|
||
|
||
The copy was removed. The guard is app-agnostic, so the same call naming the *restoring* app hits the
|
||
directory the restore is writing into. **Rank: MEDIUM** — it needs a customer to press delete during a
|
||
restore, but both controls live on the same page, the window is the whole restore, and the target is
|
||
the restore's own source. (Reconstitute and place *do* hold the flag, so the exposure is the
|
||
verification-restore window only.)
|
||
|
||
**5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight.**
|
||
Filed as an instrument defect. **Occurrences I can evidence:**
|
||
|
||
1. **2026-07-20** — an accented grep through `ssh → pct exec → bash -c` nearly produced a wrong
|
||
"banner cleared" claim. Recorded in `felhom-controller/.claude/rules/ui-hungarian.md:19-22`.
|
||
2. **2026-08-13** — `kubectl exec … sh -c "grep '<accented>'"` returned **0 for three strings that
|
||
were present**, one step from being reported as a failed hub v0.105.0 deploy.
|
||
3. **2026-08-21, tonight, 22:07** — `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`
|
||
(octal escaping). Recording the accented filenames' raw bytes from that listing would have been
|
||
wrong. Caught by extracting the archive and reading the names with `xxd`.
|
||
|
||
**Correction to the task's premise:** that is **two inside two weeks**, plus the founding case a month
|
||
earlier. I looked for a third inside the two-week window and did not find one on record.
|
||
|
||
**The smallest guard I would propose — and did NOT build:** the problem is not grep, it is that every
|
||
one of these tools *silently transforms* the bytes. So the guard is not "use ASCII fragments" (a
|
||
discipline, which is what failed three times) but **a negative control that the harness cannot skip**:
|
||
any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a
|
||
string that MUST be absent — and a zero result from the first is only reportable when the second also
|
||
returns zero *and* a third probe for a known-present ASCII anchor returns non-zero. Three probes, one
|
||
helper, no judgement required at the call site. Everything else has been tried and is what "nearly"
|
||
means in all three cases.
|
||
|
||
---
|
||
|
||
## 8. RANKED REGISTER ROWS OPENED
|
||
|
||
Ceiling was **R-353**; it **moved to R-365**.
|
||
|
||
| id | rank | what |
|
||
|---|---|---|
|
||
| **R-354** | **HIGH** | The off-site full restore has **no named-volume leg**. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. `offbox_reconstitute.go:341-346`. |
|
||
| **R-355** | **HIGH** | `paperless-ngx`'s PostgreSQL is dumped to a directory for a **non-existent stack** (`…/primary/paperless/`), so its unit records `db_dumps: null`, nothing off-sites it, **no safety dump is taken on a destructive restore**, and the customer is told the app has no database. `appbackup/dbdump.go:770-798`. One app in 53. |
|
||
| **R-356** | **HIGH** | The off-site restore **refuses for all 40 no-drive apps** with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. `offbox_reconstitute.go:208-227`. |
|
||
| **R-357** | **MEDIUM** | The **destructive** restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory. |
|
||
| **R-358** | **MEDIUM** | A **failed** scratch restore leaves a partial copy that `OffboxFullScratchReady` reports as ready; the destructive restore then runs from it and reports success. |
|
||
| **R-359** | **MEDIUM** | The off-site restic store is **never verified** by any layer — `restic check` is not among the verbs the controller runs, and the agent's restore-test is PBS-only. |
|
||
| **R-360** | **MEDIUM** | Verification-copy delete gates on `IsRunning()`, which `RestoreOffboxScratch` never holds — deletable throughout a restore. Its comment asserts the opposite. **(Part 5.1)** |
|
||
| **R-361** | **MEDIUM** | The safety dump **overwrites the unit's own DB dump**: `DumpOne` writes the canonical `<stack>-<type>.sql` and only then renames it away. The comment at `offbox_reconstitute.go:147-148` states it "can never overwrite the app's real dump". Proven: romm's `romm-mariadb.sql` was present before and absent after. |
|
||
| **R-362** | **MEDIUM** | A data drive detached mid-restore is reported as **„permission denied"**. The restore path never consults drive state. |
|
||
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
|
||
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
|
||
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
|
||
|
||
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
|
||
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
|
||
nobody. Observed tonight as event 3006, severity `info`. A sweep of every `emit(` call site shows this
|
||
is now the **only** remaining instance fleet-wide.
|
||
|
||
**R-353:** its instruction (2) is **satisfied** — see §2. Its instruction (1) stands and is now
|
||
strictly larger than when it was written, because R-354 shows the bare completion can also be reported
|
||
over a unit that *did* have a data leg.
|
||
|
||
---
|
||
|
||
## 9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH
|
||
|
||
| time (CEST) | what |
|
||
|---|---|
|
||
| 21:57 | start; baselines |
|
||
| 22:00 | break-glass into `demo-hp` |
|
||
| 22:09 | fixture planted, comparator control passed |
|
||
| 22:12 | manual local backup |
|
||
| 22:13 | manual off-site run — **0 apps toggled** |
|
||
| 22:17 | off-site run with 3 apps enabled |
|
||
| 22:19–22:20 | checking-folder restores |
|
||
| 22:21 | reconstitute privatebin → refused |
|
||
| 22:23 | reconstitute calibre-web → **the conviction** |
|
||
| 22:25 | local restore-from-unit privatebin → data returned |
|
||
| 22:27 | Part 4.1a — no-delete invariant |
|
||
| 22:39–22:45 | paperless-ngx → R-355 |
|
||
| 22:51–22:57 | Part 4.3 damaged store; repo repaired |
|
||
| 23:02–23:04 | Part 4.1b safety dump, both directions |
|
||
| 23:10–23:13 | Part 4.5 full disk (both paths); Part 4.6 controller killed mid-restore |
|
||
| 23:15 | Part 4.4 drive pulled |
|
||
| 23:17–23:26 | Part 4.7 filesystem filled, backup reserve observed, filesystem freed |
|
||
| 23:08–23:09 | abandonment set-aside store created; countdown written and controller restarted (fires 05:10) |
|
||
|
||
**Steps off the customer's path, named:**
|
||
|
||
1. **Break-glass root access** to `demo-hp` via the hub-vaulted `host_recovery` credential — the box
|
||
had lost DooPlex's SSH key (its `authorized_keys` held only its own `root@demo-hp` RSA key) and its
|
||
tailnet address was unreachable. DooPlex's public key was **re-added** to `/root/.ssh/authorized_keys`
|
||
and an `ssh` alias `hp` → `192.168.0.104` was added to `~/.ssh/config` on DooPlex.
|
||
2. **A hub DB snapshot** (`hub.db` + `-wal` + `-shm`) was streamed to the scratchpad to read
|
||
`host_recovery` and the events table. It holds every host's secret; it is in the session scratchpad
|
||
only and is not in any committed file.
|
||
3. **`settings.json` was edited directly** (controller stopped, backup at `/root/settings.json.drill-backup`)
|
||
to create the overdue abandonment state. There is no product path to shorten a countdown, and the
|
||
real orphan→reset path would have destroyed `demo-hp`'s entire off-site history.
|
||
4. **The set-aside store the sweep will delete was created by hand** at
|
||
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, for the same reason.
|
||
5. **Two restic pack files were deliberately corrupted and then restored** from byte-identical copies.
|
||
6. **`fallocate` fillers** were used to fill two filesystems and were removed.
|
||
7. Apps deployed for the drill: `opengist` (re-deployed empty), `kimai`, `paperless-ngx`, `romm`.
|
||
|
||
---
|
||
|
||
## 10. THE FENCES
|
||
|
||
- **`ep0` / the off-site endpoint — untouched outside this machine's own path, and here is how I know.**
|
||
`demo-hp` authenticates as the Hetzner Storage Box **sub-account `u629488-sub3`**, which is chrooted
|
||
to its own home: `ls /` returns **`Permission denied`**, and `/home` contains exactly `.ssh` and
|
||
`felhom-repo`. Every write, the two corruptions, the set-aside store and the sweep's target are
|
||
inside that home. **`demo-felhom` is a different sub-account (`u629488-sub1`)** and a real customer's
|
||
copy is a different sub-account again — none reachable with this key. `restic forget`/`prune` were
|
||
never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept
|
||
every 9-August snapshot (verified by listing all 24, not by a count).
|
||
- **`demo-felhom`'s two fixtures — confirmed intact, not assumed.**
|
||
(a) the unopenable set-aside store `u629488-sub1:/home/felhom-repo.orphaned-20260810` — listed
|
||
tonight, `config`/`data` (258 shards)/`index`/`keys`/`locks`/`snapshots`, mtimes still 18 Jul and
|
||
3 Aug; (b) the retained-key case — hub `host_escrow_superseded` rows **11 and 12** for
|
||
`demo-felhom-8363b5`, each with a 572-byte `identity_blob`, dated 2026-08-12. `demo-felhom` is
|
||
healthy on 0.217.0, its own off-site ran at 02:15 with `last_status: ok`, and
|
||
**`--abandon-status` there reports „no abandonment countdown is running on this box"**.
|
||
- **`peti-felhom` — not contacted.** It does not appear in the hub host list at all; no command in this
|
||
session named it.
|
||
|
||
---
|
||
|
||
## 11. SCHEDULED CYCLE AND THE COUNTDOWN
|
||
|
||
*(filled in as they fire — 02:30 local backup, 04:15 off-site, 05:10 abandonment sweep, all CEST)*
|
||
|
||
---
|
||
|
||
## 12. MACHINE STATES AT THE END
|
||
|
||
*(final state recorded in §13 after the scheduled runs)*
|