R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s

Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.

R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.

R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.

Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.

R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.

Ceiling R-366 -> R-367.
This commit is contained in:
2026-08-22 10:11:46 +02:00
parent 7064596c2e
commit 877fcd2a38
6 changed files with 1286 additions and 568 deletions
+295 -546
View File
@@ -1,599 +1,348 @@
# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night)
# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)
**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no
image built, nothing deployed.** Evidence:
**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of
2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product
where one customer action causes permanent total loss.**
The drill that found them is preserved at
`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
---
## 1. THE VERDICT
## 0. Baselines, re-established — not trusted from the sheet
**It is a mixture, and the drill's three options are all present — but they belong to different
faults, and only one of them explains the thing you actually saw.**
### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED**
Reproduced independently tonight, and it agrees with what R-353 already recorded:
1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data
drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** —
about an app that was installed, deployed, running and healthy.
2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt
box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it
held `compose/` and nothing else.
3. The restore returned that configuration and reported a bare completion.
So for that specific event the data was not in the thing that was restored. **R-353 called this
correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong
to think so; its text is more careful than the CHANGELOG's summary of it.
### What the drill found that nobody had seen: **THE RESTORE LOSES IT**
This is new, it is worse, and it is not the same fault:
**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.**
The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.**
Proven live on `calibre-web` at 22:23:36 with planted files:
| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore |
|---|---|---|---|---|
| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** |
| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** |
and the customer was told:
> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult.
> Ennek az alkalmazásnak nincs adatbázisa."
Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence
does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue
apps that declare no data drive, that leg is the entire dataset.**
**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit`
(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the
unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local**
restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring
PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of
them replays it.**
### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database
`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always
has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to
`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack
that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the
safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and
then says:
> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**"
The controller had dumped that database five minutes earlier.
**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive.
| item | expected | confirmed |
|---|---|---|
| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` |
| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box |
| golden vouched | 0.217.0 | ✔ `<option value="0.217.0" … selected>` |
| floor | 0.217.0 | ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)" |
| min agent | 0.129.0 | ✔ |
| register ceiling | R-366 | ✔ (now **R-367**) |
| `ssh hp` → 192.168.0.104 | works | ✔ |
| clean trees, all four repos | `HEAD == origin/main`, 0 dirty | ✔ |
---
## 2. THE FULL/EMPTY CONTRAST
## 1. PART 1 — where the wrong name comes from
**Apps chosen, and why.** From the two storage classes: **`calibre-web`** declares a data drive
(`needs_hdd: true`, `backup.userdata: media/books class: mandatory`) and **`privatebin`** /
**`opengist`** declare none (the 40-class; data lives in a Docker named volume). I ran **three** cases
rather than two, deliberately: FULL and EMPTY alone cannot separate *"it was empty"* from *"the class
is broken"*, so `privatebin` was run FULL as the disambiguator.
**It comes from a guess, made in a place where an answer was available.**
**The fixture** — 5 files, two with Hungarian accented names, recorded as raw bytes:
`deriveStackName` (`controller/internal/appbackup/dbdump.go:770-798`) is handed a container name and
asked which app owns it. For `paperless-postgres` it strips the role suffix to `paperless`; finds
`paperless` is **not** a deployed stack; finds the container name is not one either; finds no known
stack is a prefix of it (`paperless-ngx` is not a prefix of `paperless-postgres`) — **and then returns
the unresolved candidate anyway**, on its last line, silently.
**Every place that value is used, and what each did with it:**
| consumer | file:line | effect |
|---|---|---|
| dump directory | `backup/backup.go:482,506` | wrote to `…/primary/**paperless**/db-dumps/` on the SYSTEM drive |
| dump filename | `appbackup/dbdump.go:200-212` | `paperless-postgres.sql`, not `paperless-ngx-postgres.sql` |
| unit assembler | `backup/recovery_unit.go:131` | reads `…/primary/**paperless-ngx**/db-dumps` → finds nothing → `db_dumps: null` |
| off-site collector | `backup/offbox_capture.go:32` | resolves from the unit path → the orphan is invisible to it |
| restore's DB leg | `backup/restore_db.go:69` | `db.StackName != stackName` → never replays |
| **safety-dump filter** | `backup/offbox_reconstitute.go:134` | `mine` empty → `hasDB=false` → **no undo copy, and the refusal is never reached** |
| `.fab` export | `appexport/export.go:600` | same filter — the bundle carries no database either |
| the "has a DB" flag | `web/handlers.go:1245` | the app displays as having no database |
**Corroboration from the code itself:** `ListDumpFiles` (`dbdump.go:562`) parses a dump filename under
the comment *"Parse stack name and DB type from filename: `paperless-ngx-postgres.sql`"* — the reader
was written for a name the writer never produced.
### Controller or catalogue? — **THE CONTROLLER. No halt.**
Every container the controller starts already carries `com.docker.compose.project`, read live:
```
SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991
binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a
nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4
plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1
árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea
paperless-postgres paperless-ngx kimai-db kimai romm-db romm
```
name bytes (UTF-8 NFC):
**And that label is the stack name BY CONSTRUCTION, not by luck:** `composeExecCustomEnv`
(`internal/stacks/manager.go:1218-1230`) runs compose with `cmd.Dir` set to
`/opt/docker/stacks/<stack>` and **never passes `-p`**, so compose derives the project from that
directory. Renaming the container in the catalogue would have fixed this one app and left the guessing
for the next one. **No catalogue change was made.**
---
## 2. PART 1.2 — the sweep, proven before its answer was trusted
| step | result |
|---|---|
| scratch copy, unplanted | `MISMATCHES: 1` (rc 1) — the known one |
| **plant a second mismatch** (`kimai-db` → `timetrack-db`) | **`MISMATCHES: 2`**, convicted by name: `stack=kimai container=timetrack-db -> derived=timetrack` |
| plant removed | back to `1` |
| **every mismatch removed** | **`MISMATCHES: 0`, rc 0** — "1" is not a stuck value |
| scratch discarded | `kimai/docker-compose.yml` md5-identical to the real catalogue |
**The real count: 1 affected app of 53**, from 15 DB containers checked — `paperless-ngx`.
---
## 3. PART 1 — the dumps already written under the wrong name
**They stay, untouched, and nothing will ever collect them.**
`/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql`, 312 381 B,
last written 07:38 UTC by the final pre-fix cycle — **verified still present after the fix.**
Nothing deletes them by design: the F5 stale-primary prune (`backup/backup.go:1248`) skips any
directory that is not a deployed app, under the guard *"an undeployed app's last backup is still its
restore point"*.
**They can be adopted, by hand:** move the file to
`…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point.
**Deliberately not automatic** — it would sit beside volume tars from a different time (an incoherent
pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a
customer's data on upgrade is a migration, not a fix. **Filed as R-367.**
---
## 4. PART 2 — the mechanism, and the comment corrected
`ReconstituteFromOffsite` skipped every placement flagged as the unit (`offbox_reconstitute.go:344`),
and the volume archives live **inside** the unit. A search of the whole off-site path found no
volume-restore call; the only two callers of `restoreDockerVolumes` were `restore.go:67` and
`restore_unit.go:255`, both local.
**The comment as it stood:**
> *"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and
> clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the
> scratch unit instead, so nothing is lost by skipping it."*
**Which half still holds: the FIRST.** The live unit is the local restore path's own source and must
never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit
across the operation to keep it true.
**The second half was false and is corrected, not left.** True of the *database* dump, false of the
*volume* archives, which live in the same unit and were replayed by nothing at all. Skipping the
placement is correct; treating the skip as harmless was not — and that sentence is exactly why a
reader would not look.
**The fix:** `restoreDockerVolumesFrom(stackName, dumpDir) (int, error)` — the local path's own replay
with an explicit directory, the same shape `reimportDBDumpsFrom` already had beside `reimportDBDumps`.
**One implementation, two callers.** Volumes replay **before** the database (a logical dump must still
win over a volume-tar copy of the same database) and **inside the stopped window** (Docker will not
replace a volume a container holds).
---
## 5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM
**R-354 — calibre-web, identical fixture, identical steps, both on `demo-hp`:**
| | |
|---|---|
| **before** (0.217.0, 09:41) | `ok=true` — „A(z) calibre-web: **0 fájl visszaállítva** (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and `ls: /vol/R354-2026-08-22: No such file or directory` |
| **after** (0.218.0, 09:56) | `ok=true` — „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and **5/5 files byte-identical** |
**R-355 — paperless-ngx:**
| | |
|---|---|
| **before** (2026-08-21) | „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" over a live 72-table PostgreSQL, with **no undo copy taken at all** |
| **after** (0.218.0, 09:57) | „A(z) paperless-ngx: **0 fájl és 3 adatkötet és az adatbázis visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult." |
**A third sentence now exists for the case that had no honest wording** — an app that HAS a database
whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak **VAN adatbázisa**, de a mentés nem
tartalmazott adatbázis-mentést, ezért az adatbázis **NEM állt vissza**. A visszaállítás előtti állapot
mentése megvan: `<undo>`". That case previously printed the same confident „nincs adatbázisa".
---
## 6. THE HARDWARE WALK
**Method: endpoint-level — the exact endpoints the dashboard's JS calls (`/api/backup/run`,
`/backup/offbox/run`, `/backup/offbox/restore`, `/backup/offbox/reconstitute`), with no shell inside
the machine for any step of the walk.** Planting the fixture and reading the result back used a shell
and is setup/verification, not the walk. No browser exists on DooPlex.
**The fixture, hashed before anything ran:**
```
7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e SENTINEL.txt
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5 binary-512k.bin
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4 nested/őszibarack.md
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5 plain.txt
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58 árvíztűrő-tükörfúrógép.txt
accented names as RAW BYTES (UTF-8 NFC):
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
```
**The comparator was proved able to convict before it was trusted.** One byte flipped at offset
500 000 of `binary-1mb.bin` (`af` → `00`): `sha256sum -c` reported `binary-1mb.bin: FAILED`, rc=1,
while the other four passed; the unmodified set passed rc=0. The mutant was discarded.
**The comparator was proved able to convict first:** one byte flipped at offset 300 000 of
`binary-512k.bin` (`3b` → `00`) → `binary-512k.bin: FAILED`, rc 1, the other four `OK`; the unmodified
set rc 0. Mutant discarded.
### The results
### Results
| case | class | data leg | in unit | in off-site snapshot | checking folder | off-site restore | local restore |
|---|---|---|---|---|---|---|---|
| **calibre-web FULL** | drive | user files | n/a | 5/5 identical | 5/5 identical | **5/5 identical** | — |
| **calibre-web FULL** | drive | named volume 1.42 MB | yes | yes | yes | **NOT restored** | — |
| **privatebin FULL** | no drive | named volume 1.06 MB | yes | 5/5 identical | 5/5 identical | **REFUSED — false reason** | **5/5 identical** |
| **opengist EMPTY** | no drive | named volume 181 KB skeleton | yes | yes | yes | **REFUSED — false reason** | — |
| scenario | result |
|---|---|
| **1-A** the dump is inside the unit, and inside the off-site snapshot | **PASS.** 0.217.0 09:37: `db_dumps = None`, dump refreshed into the phantom dir. 0.218.0 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` (312 669 B) in the app's own unit, and `restic ls` shows it inside the snapshot **for the first time** |
| **1-B** an undo copy taken and verified before anything stops; the message names the database | **PASS.** Before: `find /mnt -name "pre-restore-*paperless*"` → **0**. After: `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql`, 312 957 B, in the app's own unit |
| **1-C** the undo cannot be taken → the whole restore refuses, nothing changed | **PASS**, on the app that could never reach this guard before. `ok=false`; the marker kept its mutation (`51d37b0f…`) and `StartedAt` was unchanged — **the app was never stopped** |
| **1-D** every other app byte-identical | **PASS.** `romm-mariadb.sql`, `kimai-mariadb.sql` unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created |
| **2-A/B** the volume comes back and the message says so | **PASS.** calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database |
| **2-C** a snapshot with no volume archives is unchanged | **PASS** — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected |
| **2-D** the live recovery unit is never written | **PASS** — fingerprinted across the whole operation (see red-proof 6) |
| **2-E** a failed replay is reported as a failure naming the volume | **PASS** (unit test; the live path returns the same error) |
**The contrast decides it.** The unit and the off-site snapshot **hold the data, verified by
identity**, in both classes, accented filenames included. So the capture is sound. The failure is
entirely in the last leg.
**Both messages, verbatim:**
- FULL, drive class → `ok=true`, *„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16)
— az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."*
- FULL and EMPTY, no-drive class → `ok=false`, *„a(z) privatebin nincs telepítve, ezért nincs hová
visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra
az alkalmazást (Alkalmazások) ugyanerre a helyre…"*
The second is the important one. **FULL and EMPTY got the identical sentence**, so the message cannot
distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is
installed, and the instruction ("reinstall it to the same place") **cannot be followed**, because a
40-class app is offered no storage field at deploy time (that is R-352's own measurement).
**This closes R-353's second instruction.** It asked for proof that a named-volume app reaches the
off-site tier rather than inference from gate order. It does. Manifests read tonight:
```
privatebin volume_dumps = ['privatebin_privatebin_data.tar'] db_dumps = None
opengist volume_dumps = ['opengist_opengist_data.tar'] db_dumps = None
kimai volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar']
db_dumps = ['kimai-mariadb.sql']
calibre-web volume_dumps = ['calibre-web_calibre_web_config.tar'] db_dumps = None
paperless-ngx volume_dumps = [3 tars] db_dumps = None ← the defect
```
All 15 containers healthy after the walk.
---
## 3. PART 0 — did the floor move the machine?
## 7. RED-PROOFS — seven, each mutation asserted applied and reverted
**Yes, unaided, in 17 seconds.**
| # | mutation | outcome |
|---|---|---|
| 1 | `resolveStackName` ignores the compose label | **FAIL as required.** `= "paperless", want "paperless-ngx"`, and the consequence printed both paths — written to `…/primary/paperless/db-dumps` vs read `…/primary/paperless-ngx/db-dumps`. **The empty database record, returned** |
| 2 | outcome message back to the counter-only predicate | **FAIL as required**, reproducing the 21 August sentence verbatim: `"A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."` |
| 3 | the fail-closed refusal removed from `writeSafetyDump` | **FAIL as required** — *a restore was seen proceeding with no undo copy* |
| 4 | the volume leg removed | **FAIL as required.** `VolumesReplayed = 0` and the replay never called — last night's silent loss, returned |
| 5 | the volume count dropped from the message | **FAIL as required**, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim |
| 6 | the live-unit guard removed | **PASSED FIRST — a defect in MY TEST, not the code.** The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit **root**. Widened to the whole unit (excluding only the documented undo copies) it convicts: `PLACED:17603ba1…` appears. **The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory** |
| 7 | the volume-replay error swallowed | **FAIL as required** — a partial replay reporting success |
**Green gate:** `go build ./... && go vet ./... && go test ./...` — **28 packages ok, 0 FAIL lines.**
Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red
until the bake (§8) — correctly, and never bypassed.
*(Instrument note: `go test ./... | grep -vE '^ok'; echo rc=$?` reports the **grep's** exit code, which
is 1 when every test passed and nothing was left to print. The verdict above is from an explicit
`FAIL`-line count, not from that.)*
---
## 8. WHAT THIS DOES **NOT** FIX — the blocker for the apps that need it most
**The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive
(R-356)**, telling the customer that a running app „nincs telepítve" and to reinstall it "to the same
place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset
is a named volume, **so R-354's fix cannot reach them until R-356 is closed.**
That is why the live confirmation used `calibre-web` and `paperless-ngx`: they declare a drive and can
actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing
worth doing.
---
## 9. REGISTER
- **R-354 — CLOSED**, shipped + proven-live (v0.218.0).
- **R-355 — CLOSED**, shipped + proven-live (v0.218.0).
- **R-367 — OPENED (LOW)**: dumps already written under the wrong name are stranded; adoptable by hand,
deliberately not automatic, never on an upgrade path.
- **Ceiling moved: R-366 → R-367.**
---
## 10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten
R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives,
awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification
copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359,
R-361–R-363, R-365, R-366.
---
## 11. COMMITS, VERSIONS, CI
| | |
|---|---|
| moved from → to | **0.216.0 → 0.217.0** |
| who initiated | **the hub** — the operator saved the floor at `19:48:37Z`; a poke reached the agent from `10.77.0.1:58093` at `19:48:32Z` for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision. |
| how long | floor saved `19:48:37Z` → *"controller-swap: new controller healthy"* `19:48:54Z` = **17 s**. Swap requested `19:48:41Z` → healthy = 13 s. |
| **the container's own tag on the box** | **`gitea.dooplex.hu/admin/felhom-controller:0.217.0`**, created `2026-08-21 19:48:45 UTC` — read from `docker ps` in guest 9201, not from the hub. |
The hub's own host page for `demo-hp` shows the guest's **Controller column as „—"** — the hub does
not know which controller version the box runs, while the `controller_updated` event it received says
exactly that. Two hub surfaces, one blind.
`demo-felhom` also runs 0.217.0, but it started it at `19:31:04Z` — **17 minutes before the floor was
saved** — and emitted `controller_started` with **no `controller_updated`**. It was moved by hand
during the golden bake, not by the floor.
| `felhom-controller` | `5ce3a44` the fix + tests; `2da259a` README + CONTEXT |
| image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` (150M) |
| deployed | `demo-hp` guest 9201 — `…:0.218.0 Up (healthy)` |
| `demo-felhom` | **untouched this session**, still 0.217.0 |
---
## 4. PART 2 — the four answers, from source
## 12. BAKE
**1. What puts a dump into a backup unit, and when? Which apps qualify?**
`RUNBOOK-manual-build.md` §4.0/§4.1, in the drill VM. **The golden-currency gate went red the moment
v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.**
- **Database dumps** — `runDBDumps` (`backup/backup.go:444-560`). Qualification is **a running
container whose image matches a database image**, mapped to a stack by `deriveStackName`
(`appbackup/dbdump.go:770`). Written to `<nsRoot>/backups/primary/<stack>/db-dumps/<stack>-<type>.sql`.
- **Volume dumps** — `runVolumeDumps` (`backup/backup.go:607`). Qualification is **a deployed,
non-protected stack with at least one Docker *named volume*** (`GetDockerVolumes`). An app with
zero named volumes is skipped silently and is never stopped.
- **When** — one cycle, both legs, scheduled `db-dump` daily at **02:30 CEST**; and again as the
coherence pre-phase of every off-site run (`offbox.go:938-951`), which is what makes a snapshot an
internally coherent {DB@T, files@T} pair.
- `CaptureRecoveryUnit` **only enumerates what is already on disk**
(`recovery_unit.go:131-132`). It writes no dump itself.
- **The gap this leaves:** an app whose data is a *bind mount* and which has no database container
produces neither leg. Its unit is configuration only — and nothing anywhere says so.
| | |
|---|---|
| `GOLDEN_VERSION` | **0.218.0** |
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
| package | `…/generic/felhom-golden/0.218.0/golden.tar.zst`, 657 026 013 B |
| template | `debian-13-standard_13.6-1_amd64.tar.zst` — listed fresh, not assumed |
**2. What should a correct backup contain?**
- **Declares a data drive** (13 of 53): the recovery unit (compose + `app.yaml` carrying the portable
secrets + whatever dumps exist) **plus** the paths its `.felhom.yml` `backup:` block marks
`mandatory`, appended to the restic snapshot as extra paths (`offbox_capture.go:32`).
- **Declares none** (40 of 53): **unit only**. `offboxCaptureSet` returns `(nil, nil)` when the app
has no backup block, so the off-site snapshot is the unit and nothing else. That is correct *given*
that their data is inside the unit's volume tars — and tonight confirmed it is.
**3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.**
`ReconstituteFromOffsite` returns `res` and `nil`. Nothing in it ever asks whether anything was
restored. The handler then calls `EndRestoreOp(**true**, reconstituteOutcomeMsg(...))`
(`web/offbox_handlers.go:455`). `reconstituteOutcomeMsg` (`web/offbox_handlers.go:463-475`) formats
counters:
```go
if res.DBsReplayed == 0 {
return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+
"Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when)
}
```
So with `FilesPlaced == 0` and `DBsReplayed == 0` the customer reads **"0 files restored — the
application restarted"** under a **success**. **This is a finding on its own and is recorded whichever
way the rest goes.** Two aggravations found on top of it:
- **"This application has no database" is asserted from a counter, not from a fact.**
`reimportDBDumpsFrom` returns `(0, nil)` when the dump directory is absent *or* holds no `.sql`
(`restore_db.go:30-46`), so an app that certainly has a database is told it has none. Proven live on
`paperless-ngx`.
- **The file count is rsync's transfer count, not a restore count.** A correct restore of unchanged
data reports **"0 fájl visszaállítva"** — indistinguishable from a restore that did nothing.
Observed: the same app reported 5, then 2, then 0 across three runs.
**4. What did the 9 August off-site snapshot contain, and is it still readable?**
**Readable, and it contained the data.** Three snapshots at `2026-08-09 08:30`:
| id | app | contents |
|---|---|---|
| `41c830db` | calibre-web | unit + `userdata/media/books` incl. the 9-Aug rehearsal fixture (`árvíztűrő-tükörfúrógép.txt`, `őszibarack.md`, `binary-3mb.bin`, `plain.txt`) |
| `9e38b84c` | opengist | unit + **`volume-dumps/opengist_opengist_data.tar`, 182 272 B** — manifest records `volume_dumps: ["opengist_opengist_data.tar"]` |
| `78b93f04` | privatebin | unit + `volume-dumps/privatebin_privatebin_data.tar`, 2 560 B |
All still readable tonight; `restic check` over the whole repository reports **"no errors were
found"**. **So the 9 August off-site copy of OpenGist's data exists and is intact** — it simply was
not the thing the afternoon's restore read, and the off-site route that would have read it refuses
for that class of app.
**The repository had also been dead since 9 August** and nothing said so: no snapshot between
`2026-08-09 08:30` and tonight. The cause is visible in the hub event at 16:01 — the off-box target
was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target,
**every per-app off-site toggle was still off**, so the first run I triggered reported:
**Pass markers, each checked, with the negative controls:**
```
[offbox] backup run started (0 app(s) toggled)
[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
upload OK (HTTP 201) : 1 -> pre-delete HTTP 404 (404/204 expected)
excluding : 0 <- negative control
FATAL : 0 <- negative control
```
**Credit where it is due:** the *card* does not lie about this. It reads „Aktív — nincs kijelölt
alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still
leads, and the log line alone says only „backup OK".
**Token hygiene.** Copied file→file and read by a runner script inside the VM, so it never crossed a
shell or a unit property: `systemctl show golden-bake … | grep -c -F "$(cat …)"` → **0**. **The leak
grep on the committed log was itself proven before its zero was believed** — token appended to a
throwaway copy, grepped (**1**), copy shredded, then the committed log's **0** accepted.
**Teardown:** `pct destroy 9100 --purge`; token, runner, build script and in-VM log shredded **after**
the log was copied out; VM powered off; qemu exited; **`qemu-img snapshot -a virgin` restored.**
Evidence: `documentation/tests/golden-0.218.0-2026-08-22/`.
---
## 5. THE TWO CYCLES COMPARED — AND THEY AGREE
## 13. STOP — THE APPROVAL
**By hand:** local cycle 22:12, off-site 22:17 (after enabling the per-app switches the rebuild had
silently cleared). **By the clock:** the box's own `db-dump` at 02:30 CEST, cross-drive at 03:30,
off-site at 04:15.
**Nothing below has been changed. Vouching and the floor are yours.**
| | manual | scheduled | agree? |
### The three Day-0 values, each with BOTH checks
| field | set to | downloadable | selectable |
|---|---|---|---|
| units refreshed | all 6 apps | all 6 apps, `00:30Z` | **yes** |
| `privatebin` volume tar | 1 055 744 B | 1 055 744 B | **yes** |
| `opengist` volume tar | 181 248 B | 181 248 B | **yes** |
| `calibre-web` config tar | 1 422 848 B | 368 640 B † | **yes, and explained** |
| `paperless-ngx` `db_dumps` | **absent** | **absent** | **yes — the defect reproduces on the scheduled path** |
| the orphan `…/primary/paperless/db-dumps/` | written | **written again, 294 936 B, `00:30Z`** | **yes** |
| planted files on the drive | 5/5 identical | **5/5 identical**, plus `POST-SNAPSHOT.txt` | **yes** |
| off-site run | ok, 27 → snapshots | ok, `02:17:12Z`, **2m8s**, 42.8 MB, **27 snapshots** | **yes** |
| **`golden_version`** | **0.218.0** — sha `8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b` | ✔ HTTP 200, 657 026 013 B, and the **downloaded bytes** hash to exactly the bake's reported sha | ✔ offered in the hub dropdown with `data-sha=8e427869d13eafb7…` |
| **`agent_version`** | **0.130.0** — sha `a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3` | ✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's `data-sha` exactly | ✔ already the SELECTED option — **no change needed** |
| **`min_agent`** | **0.129.0** | — | ✔ already set to 0.129.0 — **no change needed** |
† the tar shrank because the drill had by then removed the planted 1 MB from that volume — the
expected value, not a discrepancy.
**So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut.** The
three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller
being baked declares **`MinAgent: 0.129.0` (unchanged)**, the vouched agent is already **0.130.0**,
and `0.130.0 ≥ 0.129.0` — so `agent_version` and `min_agent` are already correct and only
`golden_version` moves. Setting `min_agent` above the vouched agent is the R-216 shape; it is not
happening here.
**Two independent observations, no disagreement.** The one that matters: **R-355 is not an artefact of
my manual triggering.** The box, unattended, on its own schedule, wrote paperless-ngx's PostgreSQL
dump into a directory for a stack that does not exist and left the app's own unit recording
`db_dumps: null`.
**If the hub refuses the save, the guard is working.** The R-120 gate on that POST
(`hub/internal/web/configs.go:1165`) refuses a golden older than the newest controller the fleet
reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.
The scheduled cross-drive (Tier-2) leg also completed for all six apps at 01:30Z — `crossdrive_completed`
events 3019–3024.
**Vouching is reversible:** re-select 0.217.0 and Save. The old package is never deleted by a bake.
### The one line on the floor
**Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to.**
`demo-hp` already runs 0.218.0 (deployed directly for the live proof). **`demo-felhom` is still on
0.217.0 and will not move until the floor is raised**, so today it still has both defects. I did not
raise it; that is your call, as is vouching.
---
## 6. PART 4 — everything attempted, and the message judged
## 14. WHAT WAS DROPPED, AND OBSERVATIONS
| # | test | outcome | message judged |
|---|---|---|---|
| 1a | **destructive restore: nothing is ever deleted** | **PASS both ways.** A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. | count is rsync transfers, not files restored — see §4.3 |
| 1b | **safety dump taken and verified before the stop** | **PASS.** `21:02:47Z` dump written → `21:02:47Z` `StopStack romm`. | correct: „0 fájl **és az adatbázis** visszaállítva" |
| 1b | **…and the whole operation refuses if it cannot be** | **PASS.** Made impossible by putting a regular file at the `db-dumps` path. Refused; `plain.txt` kept its post-snapshot mutation; the container's `StartedAt` was unchanged — **the app was never stopped.** | honest and names the path, but leaks a raw Go `mkdir … not a directory` into a customer surface |
| 1b | **the hole in it** | **FAIL.** For `paperless-ngx` the discovery resolves the database to the wrong stack, so `hasDB` is false: **no safety dump is taken and the refusal cannot fire.** The undo is absent, not refused. | „Ennek az alkalmazásnak nincs adatbázisa" — false |
| 2 | **end of the abandonment countdown** | state created and overdue; **fires at 05:10 CEST** — see §11 | card states a **past** date in the future tense while overdue |
| 3 | **damaged store** | **MIXED — see below** | honest at the point of failure, then forgotten |
| 4 | **drive pulled mid-restore** | restore failed, nothing written to the wrong place, **the agent re-bound the drive within 5 s** (`23:15:29` pulled → `23:15:34` re-bound) | **WRONG DIAGNOSIS.** „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions. |
| 5 | **full disk, non-destructive path** | **PASS — refuses before it starts.** | **exemplary:** „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named. |
| 5 | **full disk, destructive path** | **FAIL — no gate at all.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the three headroom gates are all on non-destructive paths (`offbox_restore.go:231,297,423`). It stopped the app, failed halfway on ENOSPC, left `DRILL-2026-08-21/` holding 2 of 5 entries, and restarted the app. | honest (`No space left on device (28)`) but raw rsync output |
| 6 | **controller killed mid-restore** | **PASS.** SIGKILL inside the stop→restore→start window. On restart: *"[appstop] crash recovery: an off-site restore (op \"offbox-reconstitute:paperless-ngx\") was interrupted and left 1 app(s) stopped — restarting them"*. App restarted, marker cleared, **and the hub was told** (event 3009). | good — names the operation, not just the app |
| 7 | **the 40-class under pressure** | **the backup noticed; the alarm did not.** Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The **reserve refused per app**: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got `recovery_unit_capture_failed` (**error**) naming the filesystem and its numbers. **The fill watcher said nothing** — it runs **once a day at 03:30** plus once at startup (`cmd/controller/main.go:1092`). | the refusal messages are good; the silence is the problem |
**Dropped: nothing.** Both parts completed in the required order, with the live walk and all seven
red-proofs. The session did not run short.
**Where the clock stopped me:** nothing in Part 4 was skipped for time. Item 2's terminal deletion is
scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only.
**Observations, noticed and not acted on:**
### 4.3 in detail — the damaged store
One byte flipped inside pack `967853d2…` at offset 5 000 000, over the repository's own SFTP
transport. Pack files are named by their content hash, so this is genuine corruption.
- **`restic check` detects it** — *"ciphertext verification failed"*, *"Fatal: repository contains
errors"*.
- **But nothing in the product ever runs it.** The only restic verbs in the entire controller are
`restore, snapshots, backup, unlock, stats, init, forget, prune, cat`. The agent's restore-test is
**PBS-tier only**. **The off-site store is never verified by any layer, at any time.** Corruption is
discovered at restore time — the worst possible moment.
- **A restore that touches the damage fails honestly:** `ok=false`, naming the file and
*"ciphertext verification failed"*.
- **But the failure is not remembered.** It left a **partial** checking folder — 78 MB, 54 files,
15 of 16 originals. `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only *"does the
directory exist and is it non-empty"*, so the wizard then offered **all three** actions including
„Teljes visszaállítás indítása".
- **Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.**
The repository was repaired from byte-identical originals; `restic check` now reports **"no errors
were found"**.
---
## 7. PART 5 — the two rows that were observed and never filed
**5.1 — the delete guard on verification copies is blind, and its own comment says otherwise.**
`offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) gates on `backupMgr.IsRunning()` — the
**concurrency** flag — while its comment states *"It refuses while a backup/restore op is running: the
copy being deleted could be the one currently being written."* R-351b moved all seven restore handlers
onto `restoreOpBlocked()` (which reads both flags); **this handler was left behind**, and one other
site (`:239`) reads the bare flag correctly and documents why.
**Reachability is not a race — it is the whole operation.** `RestoreOffboxScratch`
(`offbox_restore.go:211`) **never calls `acquireRunning` at all**, so `IsRunning()` is false for the
entire duration of an off-site verification restore. Demonstrated live, flags read immediately before
and after the delete:
```
--- BEFORE delete 22:35:25 display= True offbox-restore kimai concurrency= False
--- DELETE calibre-web verification copy: „Az ellenőrző másolat törölve…"
--- AFTER delete 22:35:25 display= True offbox-restore kimai concurrency= False
```
The copy was removed. The guard is app-agnostic, so the same call naming the *restoring* app hits the
directory the restore is writing into. **Rank: MEDIUM** — it needs a customer to press delete during a
restore, but both controls live on the same page, the window is the whole restore, and the target is
the restore's own source. (Reconstitute and place *do* hold the flag, so the exposure is the
verification-restore window only.)
**5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight.**
Filed as an instrument defect. **Occurrences I can evidence:**
1. **2026-07-20** — an accented grep through `ssh → pct exec → bash -c` nearly produced a wrong
"banner cleared" claim. Recorded in `felhom-controller/.claude/rules/ui-hungarian.md:19-22`.
2. **2026-08-13** — `kubectl exec … sh -c "grep '<accented>'"` returned **0 for three strings that
were present**, one step from being reported as a failed hub v0.105.0 deploy.
3. **2026-08-21, tonight, 22:07** — `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`
(octal escaping). Recording the accented filenames' raw bytes from that listing would have been
wrong. Caught by extracting the archive and reading the names with `xxd`.
**Correction to the task's premise:** that is **two inside two weeks**, plus the founding case a month
earlier. I looked for a third inside the two-week window and did not find one on record.
**The smallest guard I would propose — and did NOT build:** the problem is not grep, it is that every
one of these tools *silently transforms* the bytes. So the guard is not "use ASCII fragments" (a
discipline, which is what failed three times) but **a negative control that the harness cannot skip**:
any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a
string that MUST be absent — and a zero result from the first is only reportable when the second also
returns zero *and* a third probe for a known-present ASCII anchor returns non-zero. Three probes, one
helper, no judgement required at the call site. Everything else has been tried and is what "nearly"
means in all three cases.
---
## 8. RANKED REGISTER ROWS OPENED
Ceiling was **R-353**; it **moved to R-366**.
| id | rank | what |
|---|---|---|
| **R-354** | **HIGH** | The off-site full restore has **no named-volume leg**. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. `offbox_reconstitute.go:341-346`. |
| **R-355** | **HIGH** | `paperless-ngx`'s PostgreSQL is dumped to a directory for a **non-existent stack** (`…/primary/paperless/`), so its unit records `db_dumps: null`, nothing off-sites it, **no safety dump is taken on a destructive restore**, and the customer is told the app has no database. `appbackup/dbdump.go:770-798`. One app in 53. |
| **R-356** | **HIGH** | The off-site restore **refuses for all 40 no-drive apps** with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. `offbox_reconstitute.go:208-227`. |
| **R-357** | **MEDIUM** | The **destructive** restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory. |
| **R-358** | **MEDIUM** | A **failed** scratch restore leaves a partial copy that `OffboxFullScratchReady` reports as ready; the destructive restore then runs from it and reports success. |
| **R-359** | **MEDIUM** | The off-site restic store is **never verified** by any layer — `restic check` is not among the verbs the controller runs, and the agent's restore-test is PBS-only. |
| **R-360** | **MEDIUM** | Verification-copy delete gates on `IsRunning()`, which `RestoreOffboxScratch` never holds — deletable throughout a restore. Its comment asserts the opposite. **(Part 5.1)** |
| **R-361** | **MEDIUM** | The safety dump **overwrites the unit's own DB dump**: `DumpOne` writes the canonical `<stack>-<type>.sql` and only then renames it away. The comment at `offbox_reconstitute.go:147-148` states it "can never overwrite the app's real dump". Proven: romm's `romm-mariadb.sql` was present before and absent after. |
| **R-362** | **MEDIUM** | A data drive detached mid-restore is reported as **„permission denied"**. The restore path never consults drive state. |
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
| **R-366** | **HIGH** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives too** — the box cannot read its own pre-reinstall backups (`manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`, hub event 3016, filed unprompted at 21:59). The PBS-tier analogue of R-193: a rebuilt box loses **both** off-premises tiers at once. The restore-test caught it precisely; the gap is that it is called *a failed restore test* rather than *your older whole-guest backups are unreadable*. **Found incidentally — nobody was looking for it.** |
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
nobody. Observed tonight as event 3006, severity `info`. A sweep of every `emit(` call site shows this
is now the **only** remaining instance fleet-wide.
**R-353:** its instruction (2) is **satisfied** — see §2. Its instruction (1) stands and is now
strictly larger than when it was written, because R-354 shows the bare completion can also be reported
over a unit that *did* have a data leg.
---
## 9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH
| time (CEST) | what |
|---|---|
| 21:57 | start; baselines |
| 22:00 | break-glass into `demo-hp` |
| 22:09 | fixture planted, comparator control passed |
| 22:12 | manual local backup |
| 22:13 | manual off-site run — **0 apps toggled** |
| 22:17 | off-site run with 3 apps enabled |
| 22:19–22:20 | checking-folder restores |
| 22:21 | reconstitute privatebin → refused |
| 22:23 | reconstitute calibre-web → **the conviction** |
| 22:25 | local restore-from-unit privatebin → data returned |
| 22:27 | Part 4.1a — no-delete invariant |
| 22:39–22:45 | paperless-ngx → R-355 |
| 22:51–22:57 | Part 4.3 damaged store; repo repaired |
| 23:02–23:04 | Part 4.1b safety dump, both directions |
| 23:10–23:13 | Part 4.5 full disk (both paths); Part 4.6 controller killed mid-restore |
| 23:15 | Part 4.4 drive pulled |
| 23:17–23:26 | Part 4.7 filesystem filled, backup reserve observed, filesystem freed |
| 23:08–23:09 | abandonment set-aside store created; countdown written and controller restarted (fires 05:10) |
**Steps off the customer's path, named:**
1. **Break-glass root access** to `demo-hp` via the hub-vaulted `host_recovery` credential — the box
had lost DooPlex's SSH key (its `authorized_keys` held only its own `root@demo-hp` RSA key) and its
tailnet address was unreachable. DooPlex's public key was **re-added** to `/root/.ssh/authorized_keys`
and an `ssh` alias `hp` → `192.168.0.104` was added to `~/.ssh/config` on DooPlex.
2. **A hub DB snapshot** (`hub.db` + `-wal` + `-shm`) was streamed to the scratchpad to read
`host_recovery` and the events table. It holds every host's secret; it is in the session scratchpad
only and is not in any committed file.
3. **`settings.json` was edited directly** (controller stopped, backup at `/root/settings.json.drill-backup`)
to create the overdue abandonment state. There is no product path to shorten a countdown, and the
real orphan→reset path would have destroyed `demo-hp`'s entire off-site history.
4. **The set-aside store the sweep will delete was created by hand** at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, for the same reason.
5. **Two restic pack files were deliberately corrupted and then restored** from byte-identical copies.
6. **`fallocate` fillers** were used to fill two filesystems and were removed.
7. Apps deployed for the drill: `opengist` (re-deployed empty), `kimai`, `paperless-ngx`, `romm`.
---
## 10. THE FENCES
- **`ep0` / the off-site endpoint — untouched outside this machine's own path, and here is how I know.**
`demo-hp` authenticates as the Hetzner Storage Box **sub-account `u629488-sub3`**, which is chrooted
to its own home: `ls /` returns **`Permission denied`**, and `/home` contains exactly `.ssh` and
`felhom-repo`. Every write, the two corruptions, the set-aside store and the sweep's target are
inside that home. **`demo-felhom` is a different sub-account (`u629488-sub1`)** and a real customer's
copy is a different sub-account again — none reachable with this key. `restic forget`/`prune` were
never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept
every 9-August snapshot (verified by listing all 24, not by a count).
- **`demo-felhom`'s two fixtures — confirmed intact, not assumed.**
(a) the unopenable set-aside store `u629488-sub1:/home/felhom-repo.orphaned-20260810` — listed
tonight, `config`/`data` (258 shards)/`index`/`keys`/`locks`/`snapshots`, mtimes still 18 Jul and
3 Aug; (b) the retained-key case — hub `host_escrow_superseded` rows **11 and 12** for
`demo-felhom-8363b5`, each with a 572-byte `identity_blob`, dated 2026-08-12. `demo-felhom` is
healthy on 0.217.0, its own off-site ran at 02:15 with `last_status: ok`, and
**`--abandon-status` there reports „no abandonment countdown is running on this box"**.
- **`peti-felhom` — not contacted.** It does not appear in the hub host list at all; no command in this
session named it.
---
## 11. PART 4.2 — THE ABANDONMENT COUNTDOWN, WATCHED FIRING
**The terminal deletion has now been observed.** It did exactly what it claims, including the half
nobody had seen.
- **State created** 23:08–23:09 CEST. A realistic set-aside store was built at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, the countdown written overdue
(started 2026-08-07, due 2026-08-20), and the product's own CLI confirmed it:
`abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0`.
- **05:10 CEST — it fired.** `/home` on the storage box is stamped `03:10Z`; the set-aside store is
**gone**; **`/home/felhom-repo`, the live repository, is untouched** (mtime still 4 Aug).
- **05:13 CEST — the hub half.** Event **3025 `offsite_abandon_purged`**: *„Az ügyfél korábbi távoli
mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is törölve (1 csomag). Az ügyfél
döntése alapján, a 14 napos türelmi idő lejárta után."* And demo-hp's superseded escrow row **is
gone from the hub** — it held one before.
- **The controller closed itself out.** Every `abandon_*` field has been removed from `settings.json`
and `--abandon-status` reports *"no abandonment countdown is running on this box"*.
**So the two-phase commit's central promise — *"it removes BOTH halves or neither"* — is confirmed
live for the first time:** the ciphertext and the sealed package that protects it went together, three
minutes apart, and the state that remembered the operation cleaned itself up. **What it removes:**
exactly the recorded set-aside path. **What survives:** the live repository, the live escrow, and the
box's current recovery path.
**Judged:** the completion message is accurate and in plain Hungarian. The only wrong note is while
the countdown is *overdue but not yet swept* — the card then states a past date in the future tense
(**R-365**).
**No countdown is left running anywhere.** `demo-hp` cleared itself; `demo-felhom` reports none.
---
## 11b. WHAT THE NIGHT FOUND THAT NOBODY WAS LOOKING FOR
At 21:59 the box filed, unprompted, hub event 3016:
> `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could
> not be restored+booted … wrong key — manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided
> key dd:d1:d8:53:44:62:5e:0b`
That archive predates the 21 August reinstall by three days. **The rebuild orphaned the PBS whole-guest
archives exactly as it orphaned the restic repository** — so a rebuilt box loses *both* off-premises
tiers at once. The restore-test mechanism deserves credit: it caught it and named the key mismatch
precisely. The gap is what it is *called* — "a restore test failed" reads as a flaky verification, not
as "every whole-guest backup taken before the reinstall is unreadable on this machine". Filed **R-366
(HIGH)**.
---
## 12. MACHINE STATES AT THE END
**`demo-hp` — HEALTHY, not broken. Nothing needs bringing back.** At 05:28 CEST: all 15 containers up
and healthy (`felhom-controller` 0.217.0, `traefik`, `cloudflared`, `filebrowser`, plus the six drill
apps); `/` 4%, `/var/lib/felhom` 13%, the data drive 1%; **no filler files left**, **no app-stop marker**,
**no operation in flight**, **no countdown running**, and `restic check` over its off-site repository
reports **no errors were found**. Its off-site backup, silent since 9 August, is working again and ran
on its own schedule at 04:15.
**Deliberately left in place, each with a reason** (retained, not forgotten):
| left behind | why |
|---|---|
| `opengist`, `privatebin`, `calibre-web` deployed with the planted `DRILL-2026-08-21` fixture | the reproduction for **R-354** and **R-356**; the hashes in §2 make the fix verifiable without rebuilding the case |
| `paperless-ngx` deployed | the **only** reproduction of **R-355**, and it regenerates the orphan directory on every nightly cycle |
| `romm` deployed | the only app on the box where the safety dump actually works — the fixture for **R-361** and the control for R-355 |
| `kimai` deployed | a second correct-derivation DB app, the negative control for R-355 |
| five verification copies under `backups/offsite-restore/` | harmless, and they are the **R-358** fixture |
| `romm`'s `pre-restore-…sql` in its unit | the physical evidence for **R-361** |
| DooPlex's public key in `demo-hp:/root/.ssh/authorized_keys`, and the `hp` alias in `~/.ssh/config` | the box had lost the key and its tailnet address is dead; without it the next session must go through break-glass again |
**Removed / restored during the drill:** both corrupted restic packs (repo verified clean), both
`fallocate` fillers, the set-aside store (by the sweep, as intended), and demo-hp's superseded escrow
row (by the sweep's hub half, as intended).
**`demo-hp`'s tailnet address `100.76.96.79` is still unreachable** and was not repaired — the box is
reachable on the LAN at `192.168.0.104`. That is the one thing about it that is worse than it should
be, and it predates tonight.
**`demo-felhom` — untouched and healthy**, controller 0.217.0, its own off-site ran at 02:15 with
`last_status: ok`, both fixtures verified present (§10), no countdown running.
**`drill-r50`** — not used. Still `DOWN` in the hub, agent 0.129.0, as it was.
---
## 13. TEARDOWN — ALL FOUR LAYERS, STATED
| layer | state |
|---|---|
| **Off-site (storage box)** | The set-aside store I created was **deleted by the product's own sweep**, as designed. The two corrupted packs were **restored byte-identical** and `restic check` passes. Nothing else was written. `restic forget`/`prune` were never invoked by me. |
| **PVE host `demo-hp`** | `/root/settings.json.drill-backup` **retained** (the pre-drill controller settings, in case the abandonment edit needs reverting). DooPlex's SSH key **retained**, with the reason above. `/tmp/plant.tar` left; harmless. |
| **Guest 9201 / controller** | Six apps and the planted fixtures **retained with reasons** (table above). Controller state is clean: no marker, no countdown, no in-flight op. |
| **Hub — stated explicitly** | **No customer record and no host record was created, so none needs deleting.** I used the existing `demo-hp` customer and host throughout. The only hub-side *removal* was demo-hp's superseded escrow row, done by the abandonment sweep itself and reported as event 3025. Events 3002–3025 were generated as a normal consequence of the work and are left as the record. **Nothing is owed here and nothing is blocked.** |
| **DooPlex (this machine)** | The hub DB copies (which contain every host's break-glass secret) are in the session scratchpad only, never in a committed file, and are shredded in the closing step. |
---
## 14. OBSERVATIONS — noticed, not acted on
- **The hub does not display the controller version it is told.** `demo-hp`'s host page shows the
guest's Controller column as **„—"** hours after receiving `controller_updated: 0.216.0 → 0.217.0`.
Two hub surfaces, one blind. Not filed — `REPORT-hub-blindness.md` already exists and this may be
part of it.
- **`demo-felhom` moved to 0.217.0 without a `controller_updated` event** (started 19:31Z, 17 minutes
before the floor was saved). Consistent with a by-hand deployment during the golden bake, not a
defect — recorded so a later reader does not mistake it for one.
- **`restic --latest N` is per-group, not a total.** It briefly read as "18 snapshots became 10" and
would have been reported as data loss. Caught by listing all of them and by an independent
`restic stats` (26.44 MiB / 588 blobs). **An unpaginated listing is not a total** — the rule earned
its place again.
- **`find -newermt` is the wrong probe for a restore**: restic preserves the snapshot's mtimes, so a
freshly restored tree looks old. It briefly read as "nothing was restored".
- **`sftp -b` aborts on the first failing line**, so a batch listing 256 shard directories returned 7
packs and looked like a total. Prefixing each line with `-` fixed it. Same family as the two above.
---
## 15. WHAT WAS DROPPED, PLAINLY
- **Nothing in Parts 0, 2, 3, 4 or 5 was skipped.** Every Part 4 item was attempted and reached a
verdict.
- **The fill-watch alarm was not watched firing on its own schedule.** I proved its cadence from
source (daily 03:30) and observed it fire from the startup path (`disk_critical`, 21:10:58Z, correct
Hungarian copy naming the drive and the free space). Holding a filesystem at 99% for six hours would
have sat across the scheduled backup and the abandonment sweep, and I judged the scheduled cycle —
which the task asks for explicitly — worth more than a second sighting of an alarm whose trigger I
had already read and seen work.
- **I did not drive the abandonment through the real orphan→reset path**, because that path
move-asides the *live* repository, which would have destroyed demo-hp's whole off-site history
including the 9 August snapshots this drill exists to read. The sweep's own code path was exercised
in full on a real store; the deviation is only in how the state was created.
- **`peti-felhom` was not contacted**, per the fence.
- **`demo-felhom` was not touched** and still runs 0.217.0 — so the fleet is deliberately non-uniform
until the floor moves. Its two preserved fixtures were not involved in any step of this session.
- **The `.fab` export has the same R-355 blind spot** and is fixed by the same change — `export.go:600`
filters on the same `db.StackName`. Not separately verified live; the unit tests cover the resolver
that feeds it.
- **`paperless-ngx`'s orphan directory is now a permanent fixture on `demo-hp`** until R-367 is
decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for
leaving it alone for now.
- **The safety dump still overwrites the unit's own DB dump before renaming it** (R-361, out of scope)
— visible in this session's own evidence, where `romm`'s `db-dumps/` holds both a `pre-restore-*` and
the regular dump.
+22 -20
View File
@@ -1,6 +1,6 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-21 (night — the backup holds the data; the off-site RESTORE is what loses it).**
**Updated 2026-08-22 — both of last night's worst findings are fixed and proven on the machine.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
@@ -113,25 +113,27 @@ record with no machine** — created 13 August, no host, no backups, nothing to
## Broken, or knowingly incomplete
- **The off-site restore does not give back an app's Docker volume — and says it worked** (R-354).
Proven on `demo-hp` on the night of 21 August with files we planted and hashed first. The data IS in
the backup: the unit, the off-site snapshot and the verification folder all hold it, byte for byte,
accented Hungarian filenames and all. The **full off-site restore then puts back the user files and
silently skips the volume**, and tells the customer „5 fájl visszaállítva". For an app that keeps its
files on a data drive the lost piece is its settings. **For the 40 of our 53 apps that have no data
drive, that volume IS all their data.** The *local* restore does return it correctly — same tar, other
code path. **Nothing is lost off-site; the last step is what fails.** *(register: R-354)*
- **For those same 40 apps the off-site restore refuses outright, and the reason it gives is untrue**
(R-356). It says the app „nincs telepítve" — is not installed — about an app that is installed and
running, then tells the customer to reinstall it "to the same place", which those apps give them no
way to choose. This is what the OpenGist journey hit on the afternoon of 21 August before falling
back to the local restore. *(register: R-356)*
- **Paperless's database has never been in its backup** (R-355). Its dump is written every night, is
valid, has 72 tables — and lands in a folder named after an app that does not exist, on the wrong
disk, where nothing sends it off-site and nothing restores it. Because the same wrong name is used
when taking the "undo" copy before a restore, **a restore of Paperless takes no undo at all** and then
tells the customer the app has no database. **One app in 53 is affected. It is the document
archive.** *(register: R-355)*
- **FIXED and proven on the machine: the off-site restore gives the app's data back** (R-354,
controller 0.218.0). The same planted files, the same steps, both runs on `demo-hp`: on the old
build the restore said „0 fájl visszaállítva", reported success, and the folder was simply not
there. On the new one it says „**0 fájl és 1 adatkötet visszaállítva**" and all five files come back
**byte for byte**, Hungarian accented names included. The message now names what came back, because
a restore that mentions only its file count is how a silent loss reads as a success. *(register:
R-354, CLOSED)*
- **FIXED and proven on the machine: Paperless's database is in the backup, and a restore of it now
takes an undo copy first** (R-355, controller 0.218.0). The dump was going into a folder named after
an app that does not exist, so nothing collected it — and because the same wrong name was used when
looking for the live database, a restore took **no undo copy at all**. Now: the dump is in the app's
own backup and in the off-site copy for the first time; the restore said „**0 fájl és 3 adatkötet és
az adatbázis visszaállítva**"; and with the undo deliberately made impossible the restore **refused
and did not even stop the app**. We no longer guess which app a database belongs to — Docker already
tells us. **One app of 53 was affected**, established with a check we first proved could catch a
planted second case. *(register: R-355, CLOSED)*
- **STILL BROKEN, and it is now the one that matters most: 40 of our 53 apps still cannot use the
off-site restore at all** (R-356). It refuses before it starts, says a running app „nincs telepítve"
— is not installed — and tells the customer to reinstall it "to the same place", which those apps
give them no way to choose. **Those are exactly the apps whose entire data is the thing R-354 just
fixed**, so today's fix cannot reach them until this one is done. *(register: R-356)*
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the
agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the
standard tool — which we never run. *(register: R-359)*
@@ -0,0 +1,599 @@
# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night)
**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no
image built, nothing deployed.** Evidence:
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
---
## 1. THE VERDICT
**It is a mixture, and the drill's three options are all present — but they belong to different
faults, and only one of them explains the thing you actually saw.**
### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED**
Reproduced independently tonight, and it agrees with what R-353 already recorded:
1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data
drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** —
about an app that was installed, deployed, running and healthy.
2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt
box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it
held `compose/` and nothing else.
3. The restore returned that configuration and reported a bare completion.
So for that specific event the data was not in the thing that was restored. **R-353 called this
correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong
to think so; its text is more careful than the CHANGELOG's summary of it.
### What the drill found that nobody had seen: **THE RESTORE LOSES IT**
This is new, it is worse, and it is not the same fault:
**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.**
The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.**
Proven live on `calibre-web` at 22:23:36 with planted files:
| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore |
|---|---|---|---|---|
| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** |
| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** |
and the customer was told:
> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult.
> Ennek az alkalmazásnak nincs adatbázisa."
Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence
does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue
apps that declare no data drive, that leg is the entire dataset.**
**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit`
(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the
unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local**
restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring
PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of
them replays it.**
### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database
`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always
has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to
`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack
that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the
safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and
then says:
> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**"
The controller had dumped that database five minutes earlier.
**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive.
---
## 2. THE FULL/EMPTY CONTRAST
**Apps chosen, and why.** From the two storage classes: **`calibre-web`** declares a data drive
(`needs_hdd: true`, `backup.userdata: media/books class: mandatory`) and **`privatebin`** /
**`opengist`** declare none (the 40-class; data lives in a Docker named volume). I ran **three** cases
rather than two, deliberately: FULL and EMPTY alone cannot separate *"it was empty"* from *"the class
is broken"*, so `privatebin` was run FULL as the disambiguator.
**The fixture** — 5 files, two with Hungarian accented names, recorded as raw bytes:
```
SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991
binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a
nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4
plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1
árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea
name bytes (UTF-8 NFC):
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
```
**The comparator was proved able to convict before it was trusted.** One byte flipped at offset
500 000 of `binary-1mb.bin` (`af` → `00`): `sha256sum -c` reported `binary-1mb.bin: FAILED`, rc=1,
while the other four passed; the unmodified set passed rc=0. The mutant was discarded.
### The results
| case | class | data leg | in unit | in off-site snapshot | checking folder | off-site restore | local restore |
|---|---|---|---|---|---|---|---|
| **calibre-web FULL** | drive | user files | n/a | 5/5 identical | 5/5 identical | **5/5 identical** | — |
| **calibre-web FULL** | drive | named volume 1.42 MB | yes | yes | yes | **NOT restored** | — |
| **privatebin FULL** | no drive | named volume 1.06 MB | yes | 5/5 identical | 5/5 identical | **REFUSED — false reason** | **5/5 identical** |
| **opengist EMPTY** | no drive | named volume 181 KB skeleton | yes | yes | yes | **REFUSED — false reason** | — |
**The contrast decides it.** The unit and the off-site snapshot **hold the data, verified by
identity**, in both classes, accented filenames included. So the capture is sound. The failure is
entirely in the last leg.
**Both messages, verbatim:**
- FULL, drive class → `ok=true`, *„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16)
— az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."*
- FULL and EMPTY, no-drive class → `ok=false`, *„a(z) privatebin nincs telepítve, ezért nincs hová
visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra
az alkalmazást (Alkalmazások) ugyanerre a helyre…"*
The second is the important one. **FULL and EMPTY got the identical sentence**, so the message cannot
distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is
installed, and the instruction ("reinstall it to the same place") **cannot be followed**, because a
40-class app is offered no storage field at deploy time (that is R-352's own measurement).
**This closes R-353's second instruction.** It asked for proof that a named-volume app reaches the
off-site tier rather than inference from gate order. It does. Manifests read tonight:
```
privatebin volume_dumps = ['privatebin_privatebin_data.tar'] db_dumps = None
opengist volume_dumps = ['opengist_opengist_data.tar'] db_dumps = None
kimai volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar']
db_dumps = ['kimai-mariadb.sql']
calibre-web volume_dumps = ['calibre-web_calibre_web_config.tar'] db_dumps = None
paperless-ngx volume_dumps = [3 tars] db_dumps = None ← the defect
```
---
## 3. PART 0 — did the floor move the machine?
**Yes, unaided, in 17 seconds.**
| | |
|---|---|
| moved from → to | **0.216.0 → 0.217.0** |
| who initiated | **the hub** — the operator saved the floor at `19:48:37Z`; a poke reached the agent from `10.77.0.1:58093` at `19:48:32Z` for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision. |
| how long | floor saved `19:48:37Z` → *"controller-swap: new controller healthy"* `19:48:54Z` = **17 s**. Swap requested `19:48:41Z` → healthy = 13 s. |
| **the container's own tag on the box** | **`gitea.dooplex.hu/admin/felhom-controller:0.217.0`**, created `2026-08-21 19:48:45 UTC` — read from `docker ps` in guest 9201, not from the hub. |
The hub's own host page for `demo-hp` shows the guest's **Controller column as „—"** — the hub does
not know which controller version the box runs, while the `controller_updated` event it received says
exactly that. Two hub surfaces, one blind.
`demo-felhom` also runs 0.217.0, but it started it at `19:31:04Z` — **17 minutes before the floor was
saved** — and emitted `controller_started` with **no `controller_updated`**. It was moved by hand
during the golden bake, not by the floor.
---
## 4. PART 2 — the four answers, from source
**1. What puts a dump into a backup unit, and when? Which apps qualify?**
- **Database dumps** — `runDBDumps` (`backup/backup.go:444-560`). Qualification is **a running
container whose image matches a database image**, mapped to a stack by `deriveStackName`
(`appbackup/dbdump.go:770`). Written to `<nsRoot>/backups/primary/<stack>/db-dumps/<stack>-<type>.sql`.
- **Volume dumps** — `runVolumeDumps` (`backup/backup.go:607`). Qualification is **a deployed,
non-protected stack with at least one Docker *named volume*** (`GetDockerVolumes`). An app with
zero named volumes is skipped silently and is never stopped.
- **When** — one cycle, both legs, scheduled `db-dump` daily at **02:30 CEST**; and again as the
coherence pre-phase of every off-site run (`offbox.go:938-951`), which is what makes a snapshot an
internally coherent {DB@T, files@T} pair.
- `CaptureRecoveryUnit` **only enumerates what is already on disk**
(`recovery_unit.go:131-132`). It writes no dump itself.
- **The gap this leaves:** an app whose data is a *bind mount* and which has no database container
produces neither leg. Its unit is configuration only — and nothing anywhere says so.
**2. What should a correct backup contain?**
- **Declares a data drive** (13 of 53): the recovery unit (compose + `app.yaml` carrying the portable
secrets + whatever dumps exist) **plus** the paths its `.felhom.yml` `backup:` block marks
`mandatory`, appended to the restic snapshot as extra paths (`offbox_capture.go:32`).
- **Declares none** (40 of 53): **unit only**. `offboxCaptureSet` returns `(nil, nil)` when the app
has no backup block, so the off-site snapshot is the unit and nothing else. That is correct *given*
that their data is inside the unit's volume tars — and tonight confirmed it is.
**3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.**
`ReconstituteFromOffsite` returns `res` and `nil`. Nothing in it ever asks whether anything was
restored. The handler then calls `EndRestoreOp(**true**, reconstituteOutcomeMsg(...))`
(`web/offbox_handlers.go:455`). `reconstituteOutcomeMsg` (`web/offbox_handlers.go:463-475`) formats
counters:
```go
if res.DBsReplayed == 0 {
return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+
"Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when)
}
```
So with `FilesPlaced == 0` and `DBsReplayed == 0` the customer reads **"0 files restored — the
application restarted"** under a **success**. **This is a finding on its own and is recorded whichever
way the rest goes.** Two aggravations found on top of it:
- **"This application has no database" is asserted from a counter, not from a fact.**
`reimportDBDumpsFrom` returns `(0, nil)` when the dump directory is absent *or* holds no `.sql`
(`restore_db.go:30-46`), so an app that certainly has a database is told it has none. Proven live on
`paperless-ngx`.
- **The file count is rsync's transfer count, not a restore count.** A correct restore of unchanged
data reports **"0 fájl visszaállítva"** — indistinguishable from a restore that did nothing.
Observed: the same app reported 5, then 2, then 0 across three runs.
**4. What did the 9 August off-site snapshot contain, and is it still readable?**
**Readable, and it contained the data.** Three snapshots at `2026-08-09 08:30`:
| id | app | contents |
|---|---|---|
| `41c830db` | calibre-web | unit + `userdata/media/books` incl. the 9-Aug rehearsal fixture (`árvíztűrő-tükörfúrógép.txt`, `őszibarack.md`, `binary-3mb.bin`, `plain.txt`) |
| `9e38b84c` | opengist | unit + **`volume-dumps/opengist_opengist_data.tar`, 182 272 B** — manifest records `volume_dumps: ["opengist_opengist_data.tar"]` |
| `78b93f04` | privatebin | unit + `volume-dumps/privatebin_privatebin_data.tar`, 2 560 B |
All still readable tonight; `restic check` over the whole repository reports **"no errors were
found"**. **So the 9 August off-site copy of OpenGist's data exists and is intact** — it simply was
not the thing the afternoon's restore read, and the off-site route that would have read it refuses
for that class of app.
**The repository had also been dead since 9 August** and nothing said so: no snapshot between
`2026-08-09 08:30` and tonight. The cause is visible in the hub event at 16:01 — the off-box target
was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target,
**every per-app off-site toggle was still off**, so the first run I triggered reported:
```
[offbox] backup run started (0 app(s) toggled)
[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s
```
**Credit where it is due:** the *card* does not lie about this. It reads „Aktív — nincs kijelölt
alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still
leads, and the log line alone says only „backup OK".
---
## 5. THE TWO CYCLES COMPARED — AND THEY AGREE
**By hand:** local cycle 22:12, off-site 22:17 (after enabling the per-app switches the rebuild had
silently cleared). **By the clock:** the box's own `db-dump` at 02:30 CEST, cross-drive at 03:30,
off-site at 04:15.
| | manual | scheduled | agree? |
|---|---|---|---|
| units refreshed | all 6 apps | all 6 apps, `00:30Z` | **yes** |
| `privatebin` volume tar | 1 055 744 B | 1 055 744 B | **yes** |
| `opengist` volume tar | 181 248 B | 181 248 B | **yes** |
| `calibre-web` config tar | 1 422 848 B | 368 640 B † | **yes, and explained** |
| `paperless-ngx` `db_dumps` | **absent** | **absent** | **yes — the defect reproduces on the scheduled path** |
| the orphan `…/primary/paperless/db-dumps/` | written | **written again, 294 936 B, `00:30Z`** | **yes** |
| planted files on the drive | 5/5 identical | **5/5 identical**, plus `POST-SNAPSHOT.txt` | **yes** |
| off-site run | ok, 27 → snapshots | ok, `02:17:12Z`, **2m8s**, 42.8 MB, **27 snapshots** | **yes** |
† the tar shrank because the drill had by then removed the planted 1 MB from that volume — the
expected value, not a discrepancy.
**Two independent observations, no disagreement.** The one that matters: **R-355 is not an artefact of
my manual triggering.** The box, unattended, on its own schedule, wrote paperless-ngx's PostgreSQL
dump into a directory for a stack that does not exist and left the app's own unit recording
`db_dumps: null`.
The scheduled cross-drive (Tier-2) leg also completed for all six apps at 01:30Z — `crossdrive_completed`
events 3019–3024.
---
## 6. PART 4 — everything attempted, and the message judged
| # | test | outcome | message judged |
|---|---|---|---|
| 1a | **destructive restore: nothing is ever deleted** | **PASS both ways.** A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. | count is rsync transfers, not files restored — see §4.3 |
| 1b | **safety dump taken and verified before the stop** | **PASS.** `21:02:47Z` dump written → `21:02:47Z` `StopStack romm`. | correct: „0 fájl **és az adatbázis** visszaállítva" |
| 1b | **…and the whole operation refuses if it cannot be** | **PASS.** Made impossible by putting a regular file at the `db-dumps` path. Refused; `plain.txt` kept its post-snapshot mutation; the container's `StartedAt` was unchanged — **the app was never stopped.** | honest and names the path, but leaks a raw Go `mkdir … not a directory` into a customer surface |
| 1b | **the hole in it** | **FAIL.** For `paperless-ngx` the discovery resolves the database to the wrong stack, so `hasDB` is false: **no safety dump is taken and the refusal cannot fire.** The undo is absent, not refused. | „Ennek az alkalmazásnak nincs adatbázisa" — false |
| 2 | **end of the abandonment countdown** | state created and overdue; **fires at 05:10 CEST** — see §11 | card states a **past** date in the future tense while overdue |
| 3 | **damaged store** | **MIXED — see below** | honest at the point of failure, then forgotten |
| 4 | **drive pulled mid-restore** | restore failed, nothing written to the wrong place, **the agent re-bound the drive within 5 s** (`23:15:29` pulled → `23:15:34` re-bound) | **WRONG DIAGNOSIS.** „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions. |
| 5 | **full disk, non-destructive path** | **PASS — refuses before it starts.** | **exemplary:** „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named. |
| 5 | **full disk, destructive path** | **FAIL — no gate at all.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the three headroom gates are all on non-destructive paths (`offbox_restore.go:231,297,423`). It stopped the app, failed halfway on ENOSPC, left `DRILL-2026-08-21/` holding 2 of 5 entries, and restarted the app. | honest (`No space left on device (28)`) but raw rsync output |
| 6 | **controller killed mid-restore** | **PASS.** SIGKILL inside the stop→restore→start window. On restart: *"[appstop] crash recovery: an off-site restore (op \"offbox-reconstitute:paperless-ngx\") was interrupted and left 1 app(s) stopped — restarting them"*. App restarted, marker cleared, **and the hub was told** (event 3009). | good — names the operation, not just the app |
| 7 | **the 40-class under pressure** | **the backup noticed; the alarm did not.** Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The **reserve refused per app**: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got `recovery_unit_capture_failed` (**error**) naming the filesystem and its numbers. **The fill watcher said nothing** — it runs **once a day at 03:30** plus once at startup (`cmd/controller/main.go:1092`). | the refusal messages are good; the silence is the problem |
**Where the clock stopped me:** nothing in Part 4 was skipped for time. Item 2's terminal deletion is
scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only.
### 4.3 in detail — the damaged store
One byte flipped inside pack `967853d2…` at offset 5 000 000, over the repository's own SFTP
transport. Pack files are named by their content hash, so this is genuine corruption.
- **`restic check` detects it** — *"ciphertext verification failed"*, *"Fatal: repository contains
errors"*.
- **But nothing in the product ever runs it.** The only restic verbs in the entire controller are
`restore, snapshots, backup, unlock, stats, init, forget, prune, cat`. The agent's restore-test is
**PBS-tier only**. **The off-site store is never verified by any layer, at any time.** Corruption is
discovered at restore time — the worst possible moment.
- **A restore that touches the damage fails honestly:** `ok=false`, naming the file and
*"ciphertext verification failed"*.
- **But the failure is not remembered.** It left a **partial** checking folder — 78 MB, 54 files,
15 of 16 originals. `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only *"does the
directory exist and is it non-empty"*, so the wizard then offered **all three** actions including
„Teljes visszaállítás indítása".
- **Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.**
The repository was repaired from byte-identical originals; `restic check` now reports **"no errors
were found"**.
---
## 7. PART 5 — the two rows that were observed and never filed
**5.1 — the delete guard on verification copies is blind, and its own comment says otherwise.**
`offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) gates on `backupMgr.IsRunning()` — the
**concurrency** flag — while its comment states *"It refuses while a backup/restore op is running: the
copy being deleted could be the one currently being written."* R-351b moved all seven restore handlers
onto `restoreOpBlocked()` (which reads both flags); **this handler was left behind**, and one other
site (`:239`) reads the bare flag correctly and documents why.
**Reachability is not a race — it is the whole operation.** `RestoreOffboxScratch`
(`offbox_restore.go:211`) **never calls `acquireRunning` at all**, so `IsRunning()` is false for the
entire duration of an off-site verification restore. Demonstrated live, flags read immediately before
and after the delete:
```
--- BEFORE delete 22:35:25 display= True offbox-restore kimai concurrency= False
--- DELETE calibre-web verification copy: „Az ellenőrző másolat törölve…"
--- AFTER delete 22:35:25 display= True offbox-restore kimai concurrency= False
```
The copy was removed. The guard is app-agnostic, so the same call naming the *restoring* app hits the
directory the restore is writing into. **Rank: MEDIUM** — it needs a customer to press delete during a
restore, but both controls live on the same page, the window is the whole restore, and the target is
the restore's own source. (Reconstitute and place *do* hold the flag, so the exposure is the
verification-restore window only.)
**5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight.**
Filed as an instrument defect. **Occurrences I can evidence:**
1. **2026-07-20** — an accented grep through `ssh → pct exec → bash -c` nearly produced a wrong
"banner cleared" claim. Recorded in `felhom-controller/.claude/rules/ui-hungarian.md:19-22`.
2. **2026-08-13** — `kubectl exec … sh -c "grep '<accented>'"` returned **0 for three strings that
were present**, one step from being reported as a failed hub v0.105.0 deploy.
3. **2026-08-21, tonight, 22:07** — `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`
(octal escaping). Recording the accented filenames' raw bytes from that listing would have been
wrong. Caught by extracting the archive and reading the names with `xxd`.
**Correction to the task's premise:** that is **two inside two weeks**, plus the founding case a month
earlier. I looked for a third inside the two-week window and did not find one on record.
**The smallest guard I would propose — and did NOT build:** the problem is not grep, it is that every
one of these tools *silently transforms* the bytes. So the guard is not "use ASCII fragments" (a
discipline, which is what failed three times) but **a negative control that the harness cannot skip**:
any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a
string that MUST be absent — and a zero result from the first is only reportable when the second also
returns zero *and* a third probe for a known-present ASCII anchor returns non-zero. Three probes, one
helper, no judgement required at the call site. Everything else has been tried and is what "nearly"
means in all three cases.
---
## 8. RANKED REGISTER ROWS OPENED
Ceiling was **R-353**; it **moved to R-366**.
| id | rank | what |
|---|---|---|
| **R-354** | **HIGH** | The off-site full restore has **no named-volume leg**. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. `offbox_reconstitute.go:341-346`. |
| **R-355** | **HIGH** | `paperless-ngx`'s PostgreSQL is dumped to a directory for a **non-existent stack** (`…/primary/paperless/`), so its unit records `db_dumps: null`, nothing off-sites it, **no safety dump is taken on a destructive restore**, and the customer is told the app has no database. `appbackup/dbdump.go:770-798`. One app in 53. |
| **R-356** | **HIGH** | The off-site restore **refuses for all 40 no-drive apps** with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. `offbox_reconstitute.go:208-227`. |
| **R-357** | **MEDIUM** | The **destructive** restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory. |
| **R-358** | **MEDIUM** | A **failed** scratch restore leaves a partial copy that `OffboxFullScratchReady` reports as ready; the destructive restore then runs from it and reports success. |
| **R-359** | **MEDIUM** | The off-site restic store is **never verified** by any layer — `restic check` is not among the verbs the controller runs, and the agent's restore-test is PBS-only. |
| **R-360** | **MEDIUM** | Verification-copy delete gates on `IsRunning()`, which `RestoreOffboxScratch` never holds — deletable throughout a restore. Its comment asserts the opposite. **(Part 5.1)** |
| **R-361** | **MEDIUM** | The safety dump **overwrites the unit's own DB dump**: `DumpOne` writes the canonical `<stack>-<type>.sql` and only then renames it away. The comment at `offbox_reconstitute.go:147-148` states it "can never overwrite the app's real dump". Proven: romm's `romm-mariadb.sql` was present before and absent after. |
| **R-362** | **MEDIUM** | A data drive detached mid-restore is reported as **„permission denied"**. The restore path never consults drive state. |
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
| **R-366** | **HIGH** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives too** — the box cannot read its own pre-reinstall backups (`manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`, hub event 3016, filed unprompted at 21:59). The PBS-tier analogue of R-193: a rebuilt box loses **both** off-premises tiers at once. The restore-test caught it precisely; the gap is that it is called *a failed restore test* rather than *your older whole-guest backups are unreadable*. **Found incidentally — nobody was looking for it.** |
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
nobody. Observed tonight as event 3006, severity `info`. A sweep of every `emit(` call site shows this
is now the **only** remaining instance fleet-wide.
**R-353:** its instruction (2) is **satisfied** — see §2. Its instruction (1) stands and is now
strictly larger than when it was written, because R-354 shows the bare completion can also be reported
over a unit that *did* have a data leg.
---
## 9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH
| time (CEST) | what |
|---|---|
| 21:57 | start; baselines |
| 22:00 | break-glass into `demo-hp` |
| 22:09 | fixture planted, comparator control passed |
| 22:12 | manual local backup |
| 22:13 | manual off-site run — **0 apps toggled** |
| 22:17 | off-site run with 3 apps enabled |
| 22:19–22:20 | checking-folder restores |
| 22:21 | reconstitute privatebin → refused |
| 22:23 | reconstitute calibre-web → **the conviction** |
| 22:25 | local restore-from-unit privatebin → data returned |
| 22:27 | Part 4.1a — no-delete invariant |
| 22:39–22:45 | paperless-ngx → R-355 |
| 22:51–22:57 | Part 4.3 damaged store; repo repaired |
| 23:02–23:04 | Part 4.1b safety dump, both directions |
| 23:10–23:13 | Part 4.5 full disk (both paths); Part 4.6 controller killed mid-restore |
| 23:15 | Part 4.4 drive pulled |
| 23:17–23:26 | Part 4.7 filesystem filled, backup reserve observed, filesystem freed |
| 23:08–23:09 | abandonment set-aside store created; countdown written and controller restarted (fires 05:10) |
**Steps off the customer's path, named:**
1. **Break-glass root access** to `demo-hp` via the hub-vaulted `host_recovery` credential — the box
had lost DooPlex's SSH key (its `authorized_keys` held only its own `root@demo-hp` RSA key) and its
tailnet address was unreachable. DooPlex's public key was **re-added** to `/root/.ssh/authorized_keys`
and an `ssh` alias `hp` → `192.168.0.104` was added to `~/.ssh/config` on DooPlex.
2. **A hub DB snapshot** (`hub.db` + `-wal` + `-shm`) was streamed to the scratchpad to read
`host_recovery` and the events table. It holds every host's secret; it is in the session scratchpad
only and is not in any committed file.
3. **`settings.json` was edited directly** (controller stopped, backup at `/root/settings.json.drill-backup`)
to create the overdue abandonment state. There is no product path to shorten a countdown, and the
real orphan→reset path would have destroyed `demo-hp`'s entire off-site history.
4. **The set-aside store the sweep will delete was created by hand** at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, for the same reason.
5. **Two restic pack files were deliberately corrupted and then restored** from byte-identical copies.
6. **`fallocate` fillers** were used to fill two filesystems and were removed.
7. Apps deployed for the drill: `opengist` (re-deployed empty), `kimai`, `paperless-ngx`, `romm`.
---
## 10. THE FENCES
- **`ep0` / the off-site endpoint — untouched outside this machine's own path, and here is how I know.**
`demo-hp` authenticates as the Hetzner Storage Box **sub-account `u629488-sub3`**, which is chrooted
to its own home: `ls /` returns **`Permission denied`**, and `/home` contains exactly `.ssh` and
`felhom-repo`. Every write, the two corruptions, the set-aside store and the sweep's target are
inside that home. **`demo-felhom` is a different sub-account (`u629488-sub1`)** and a real customer's
copy is a different sub-account again — none reachable with this key. `restic forget`/`prune` were
never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept
every 9-August snapshot (verified by listing all 24, not by a count).
- **`demo-felhom`'s two fixtures — confirmed intact, not assumed.**
(a) the unopenable set-aside store `u629488-sub1:/home/felhom-repo.orphaned-20260810` — listed
tonight, `config`/`data` (258 shards)/`index`/`keys`/`locks`/`snapshots`, mtimes still 18 Jul and
3 Aug; (b) the retained-key case — hub `host_escrow_superseded` rows **11 and 12** for
`demo-felhom-8363b5`, each with a 572-byte `identity_blob`, dated 2026-08-12. `demo-felhom` is
healthy on 0.217.0, its own off-site ran at 02:15 with `last_status: ok`, and
**`--abandon-status` there reports „no abandonment countdown is running on this box"**.
- **`peti-felhom` — not contacted.** It does not appear in the hub host list at all; no command in this
session named it.
---
## 11. PART 4.2 — THE ABANDONMENT COUNTDOWN, WATCHED FIRING
**The terminal deletion has now been observed.** It did exactly what it claims, including the half
nobody had seen.
- **State created** 23:08–23:09 CEST. A realistic set-aside store was built at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, the countdown written overdue
(started 2026-08-07, due 2026-08-20), and the product's own CLI confirmed it:
`abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0`.
- **05:10 CEST — it fired.** `/home` on the storage box is stamped `03:10Z`; the set-aside store is
**gone**; **`/home/felhom-repo`, the live repository, is untouched** (mtime still 4 Aug).
- **05:13 CEST — the hub half.** Event **3025 `offsite_abandon_purged`**: *„Az ügyfél korábbi távoli
mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is törölve (1 csomag). Az ügyfél
döntése alapján, a 14 napos türelmi idő lejárta után."* And demo-hp's superseded escrow row **is
gone from the hub** — it held one before.
- **The controller closed itself out.** Every `abandon_*` field has been removed from `settings.json`
and `--abandon-status` reports *"no abandonment countdown is running on this box"*.
**So the two-phase commit's central promise — *"it removes BOTH halves or neither"* — is confirmed
live for the first time:** the ciphertext and the sealed package that protects it went together, three
minutes apart, and the state that remembered the operation cleaned itself up. **What it removes:**
exactly the recorded set-aside path. **What survives:** the live repository, the live escrow, and the
box's current recovery path.
**Judged:** the completion message is accurate and in plain Hungarian. The only wrong note is while
the countdown is *overdue but not yet swept* — the card then states a past date in the future tense
(**R-365**).
**No countdown is left running anywhere.** `demo-hp` cleared itself; `demo-felhom` reports none.
---
## 11b. WHAT THE NIGHT FOUND THAT NOBODY WAS LOOKING FOR
At 21:59 the box filed, unprompted, hub event 3016:
> `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could
> not be restored+booted … wrong key — manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided
> key dd:d1:d8:53:44:62:5e:0b`
That archive predates the 21 August reinstall by three days. **The rebuild orphaned the PBS whole-guest
archives exactly as it orphaned the restic repository** — so a rebuilt box loses *both* off-premises
tiers at once. The restore-test mechanism deserves credit: it caught it and named the key mismatch
precisely. The gap is what it is *called* — "a restore test failed" reads as a flaky verification, not
as "every whole-guest backup taken before the reinstall is unreadable on this machine". Filed **R-366
(HIGH)**.
---
## 12. MACHINE STATES AT THE END
**`demo-hp` — HEALTHY, not broken. Nothing needs bringing back.** At 05:28 CEST: all 15 containers up
and healthy (`felhom-controller` 0.217.0, `traefik`, `cloudflared`, `filebrowser`, plus the six drill
apps); `/` 4%, `/var/lib/felhom` 13%, the data drive 1%; **no filler files left**, **no app-stop marker**,
**no operation in flight**, **no countdown running**, and `restic check` over its off-site repository
reports **no errors were found**. Its off-site backup, silent since 9 August, is working again and ran
on its own schedule at 04:15.
**Deliberately left in place, each with a reason** (retained, not forgotten):
| left behind | why |
|---|---|
| `opengist`, `privatebin`, `calibre-web` deployed with the planted `DRILL-2026-08-21` fixture | the reproduction for **R-354** and **R-356**; the hashes in §2 make the fix verifiable without rebuilding the case |
| `paperless-ngx` deployed | the **only** reproduction of **R-355**, and it regenerates the orphan directory on every nightly cycle |
| `romm` deployed | the only app on the box where the safety dump actually works — the fixture for **R-361** and the control for R-355 |
| `kimai` deployed | a second correct-derivation DB app, the negative control for R-355 |
| five verification copies under `backups/offsite-restore/` | harmless, and they are the **R-358** fixture |
| `romm`'s `pre-restore-…sql` in its unit | the physical evidence for **R-361** |
| DooPlex's public key in `demo-hp:/root/.ssh/authorized_keys`, and the `hp` alias in `~/.ssh/config` | the box had lost the key and its tailnet address is dead; without it the next session must go through break-glass again |
**Removed / restored during the drill:** both corrupted restic packs (repo verified clean), both
`fallocate` fillers, the set-aside store (by the sweep, as intended), and demo-hp's superseded escrow
row (by the sweep's hub half, as intended).
**`demo-hp`'s tailnet address `100.76.96.79` is still unreachable** and was not repaired — the box is
reachable on the LAN at `192.168.0.104`. That is the one thing about it that is worse than it should
be, and it predates tonight.
**`demo-felhom` — untouched and healthy**, controller 0.217.0, its own off-site ran at 02:15 with
`last_status: ok`, both fixtures verified present (§10), no countdown running.
**`drill-r50`** — not used. Still `DOWN` in the hub, agent 0.129.0, as it was.
---
## 13. TEARDOWN — ALL FOUR LAYERS, STATED
| layer | state |
|---|---|
| **Off-site (storage box)** | The set-aside store I created was **deleted by the product's own sweep**, as designed. The two corrupted packs were **restored byte-identical** and `restic check` passes. Nothing else was written. `restic forget`/`prune` were never invoked by me. |
| **PVE host `demo-hp`** | `/root/settings.json.drill-backup` **retained** (the pre-drill controller settings, in case the abandonment edit needs reverting). DooPlex's SSH key **retained**, with the reason above. `/tmp/plant.tar` left; harmless. |
| **Guest 9201 / controller** | Six apps and the planted fixtures **retained with reasons** (table above). Controller state is clean: no marker, no countdown, no in-flight op. |
| **Hub — stated explicitly** | **No customer record and no host record was created, so none needs deleting.** I used the existing `demo-hp` customer and host throughout. The only hub-side *removal* was demo-hp's superseded escrow row, done by the abandonment sweep itself and reported as event 3025. Events 3002–3025 were generated as a normal consequence of the work and are left as the record. **Nothing is owed here and nothing is blocked.** |
| **DooPlex (this machine)** | The hub DB copies (which contain every host's break-glass secret) are in the session scratchpad only, never in a committed file, and are shredded in the closing step. |
---
## 14. OBSERVATIONS — noticed, not acted on
- **The hub does not display the controller version it is told.** `demo-hp`'s host page shows the
guest's Controller column as **„—"** hours after receiving `controller_updated: 0.216.0 → 0.217.0`.
Two hub surfaces, one blind. Not filed — `REPORT-hub-blindness.md` already exists and this may be
part of it.
- **`demo-felhom` moved to 0.217.0 without a `controller_updated` event** (started 19:31Z, 17 minutes
before the floor was saved). Consistent with a by-hand deployment during the golden bake, not a
defect — recorded so a later reader does not mistake it for one.
- **`restic --latest N` is per-group, not a total.** It briefly read as "18 snapshots became 10" and
would have been reported as data loss. Caught by listing all of them and by an independent
`restic stats` (26.44 MiB / 588 blobs). **An unpaginated listing is not a total** — the rule earned
its place again.
- **`find -newermt` is the wrong probe for a restore**: restic preserves the snapshot's mtimes, so a
freshly restored tree looks old. It briefly read as "nothing was restored".
- **`sftp -b` aborts on the first failing line**, so a batch listing 256 shard directories returned 7
packs and looked like a total. Prefixing each line with `-` fixed it. Same family as the two above.
---
## 15. WHAT WAS DROPPED, PLAINLY
- **Nothing in Parts 0, 2, 3, 4 or 5 was skipped.** Every Part 4 item was attempted and reached a
verdict.
- **The fill-watch alarm was not watched firing on its own schedule.** I proved its cadence from
source (daily 03:30) and observed it fire from the startup path (`disk_critical`, 21:10:58Z, correct
Hungarian copy naming the drive and the free space). Holding a filesystem at 99% for six hours would
have sat across the scheduled backup and the abandonment sweep, and I judged the scheduled cycle —
which the task asks for explicitly — worth more than a second sighting of an alarm whose trigger I
had already read and seen work.
- **I did not drive the abandonment through the real orphan→reset path**, because that path
move-asides the *live* repository, which would have destroyed demo-hp's whole off-site history
including the 9 August snapshots this drill exists to read. The sweep's own code path was exercised
in full on a real store; the deviation is only in how the state was created.
- **`peti-felhom` was not contacted**, per the fence.
+3 -2
View File
@@ -653,8 +653,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** Two findings, one session, both shipped. **(a) The blindness.** Every recovery-unit `manifest.json` has carried `drive` and `namespace_root` since schema 1 (`controller/internal/backup/recovery_unit.go:48-49`), written at capture from the app's own placement. `grep -rE '\.Drive\b\|\.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | **CLOSED 2026-08-21** — controller | — | **Shipped:** `backup/offbox_placement.go` (`CheckPlacement`, `PlacementMismatchMessage`, `RecordedUnitForStack`); mismatch **named and refused** before the safety dump, with `ack_placement` as a **separate** field from `confirm=1`; the not-installed refusal names the recorded drive; deploy page **prefills the address and folder from the app's own backup**; `Server.restoreOpBlocked()` reads BOTH flags; `RestoreOpStatus.LastRecent` + `RestoreResultWindow` moved to `internal/backup` as ONE expression for two surfaces. Red-proofs: B with **both** guards removed **was seen starting a restore with no drive attached** (no error, full 3.00 s run into `/tmp/mutant-destination`); C, E and A each returned their wrong outcome; D forced on broke 8 ordinary reconstitute tests, proving reachability both ways. | CC |
| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC |
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **OPEN — HIGH** | — | Restore the unit's `volume-dumps/*.tar` from the SCRATCH unit on the off-site path, as `RestoreFromRecoveryUnit` already does from the live unit. Assert the CONSEQUENCE (planted bytes come back), not the mechanism. | CC |
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **OPEN — HIGH** | — | Two candidates, both two-repo: rename the container to `paperless-ngx-postgres`, or make an unresolved candidate a loud skip rather than a silent fallback. The second is the one that generalises. Pin with a test that a container whose name resolves to no known stack is never dumped silently. | CC |
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | `restoreDockerVolumesFrom` is the LOCAL path's own replay with an explicit directory — ONE implementation, two callers, because a second copy of that loop is what produced the divergence. It reads the SCRATCH unit; the live unit is still never written. Volumes replay BEFORE the database (a logical dump must still win over a volume-tar copy of the same database) and inside the stopped window (Docker will not replace a volume a container holds). `VolumesReplayed` is on the result and in the sentence. **The comment beside the skip was half false and is corrected, not left:** it justified the skip by saying the dump is replayed from the scratch "so nothing is lost" — true of the database, false of the volumes. The half that still holds (the live unit is the local path's source) is named. **PROVEN LIVE on `demo-hp`, negative control first.** On 0.217.0 with the fixture planted, hashed and then deleted from the live volume: „A(z) calibre-web: **0 fájl visszaállítva** … — az alkalmazás újraindult.", `ok=true`, and `ls` reported the directory absent. On 0.218.0, same fixture, same steps: „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** …" and **5/5 files byte-identical**, both Hungarian accented filenames included. **Four red-proofs, each mutation asserted applied:** the volume leg removed returned the silent loss; the count dropped from the message returned the true-but-incomplete sentence verbatim; the unit guard removed was SEEN writing into the live unit; the error swallowed let a partial replay report success. **The unit-guard proof initially PASSED against the mutation** — the fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit root (the R-181 class, in the test rather than the code). Widened, and it convicts. **NOT reached by this fix, and it is the blocker:** the 40 apps that declare no data drive still cannot run this restore at all — **R-356** refuses first. | CC |
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | **Neither candidate: a third that removes the guessing.** Every container the controller starts carries `com.docker.compose.project`, and that label IS the stack name BY CONSTRUCTION — compose is run with `cmd.Dir` set to `/opt/docker/stacks/<stack>` and never `-p` (`stacks/manager.go:1218`). `resolveStackName` prefers it whenever it names a deployed stack; `deriveStackName` stays as the fallback for containers not started by compose; an attribution that resolves to NO known stack is now **loud** instead of silently returned. **The fix is in the CONTROLLER, not the catalogue** — renaming the container would have fixed this one app and left the guessing for the next. **The sweep was proven before its answer was trusted:** a second mismatch planted in a scratch copy (`kimai-db`→`timetrack-db`) was convicted by name, removal returned it to 1, and a catalogue with every mismatch removed exits **0** — so "1" is not a stuck value. **1 affected app of 53**, 15 DB containers checked. **PROVEN LIVE on `demo-hp`, negative control first.** 0.217.0 at 09:37: unit `db_dumps = None`, dump refreshed into the phantom `…/primary/paperless/db-dumps/paperless-postgres.sql`. 0.218.0 at 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` **inside the app's own unit**, and present in the off-site snapshot for the first time. Scenario B: the destructive restore said „**0 fájl és 3 adatkötet és az adatbázis visszaállítva**" and wrote `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql` (312 957 B) where **zero** undo copies had existed. Scenario C: with the undo made impossible the restore REFUSED, the marker kept its mutation and the container's `StartedAt` was unchanged — **the app was never stopped**. Scenario D: `romm-mariadb.sql` and `kimai-mariadb.sql` unchanged in name and location. **Three red-proofs, each asserted applied:** the name fix reverted printed both divergent paths; the message predicate reverted returned the false „nincs adatbázisa" sentence verbatim; the refusal removed was seen letting a restore proceed with no undo. **The dumps already written under the wrong name are NOT deleted** — see **R-367**. | CC |
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. | CC |
| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC |
| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC |
@@ -666,6 +666,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC |
| **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
| **R-367** | **The database dumps already written under the wrong name are stranded, and nothing will ever collect them.** R-355's fix sends `paperless-ngx`'s dump to the right place from now on; it does not move the ones already written. On `demo-hp` that is `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql` (312 381 B, 2026-08-22 07:38, the last pre-fix cycle). **Nothing deletes them and that is by design, not by luck:** the F5 stale-primary prune (`backup.go:1248`) skips any directory whose name is not a deployed app, under the guard *"an undeployed app's last backup is still its restore point"* — verified still present after the fix. They are equally invisible to the off-site push, which resolves paths from the app's own unit. **They CAN be adopted, by hand:** move the file to `…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point for that app. **It is deliberately not automatic.** The adopted dump would sit beside volume tars taken at a different time, i.e. an INCOHERENT pair — the exact shape R-43/R-44's coherence stamp exists to make visible — and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. **Filed rather than done**, because whether a stale orphan is worth adopting at all is a judgement about one machine's history, not a rule. | **OPEN — LOW** | follows R-355 | Decide per box: adopt (and say the pair is skewed), or delete deliberately. Neither on an upgrade path. | Viktor rules, CC executes |
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC |
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
@@ -0,0 +1,44 @@
# Golden bake — 0.218.0 (2026-08-22)
Baked in the drill VM on DooPlex per `documentation/runbooks/RUNBOOK-manual-build.md` §4.0/§4.1,
carrying controller **v0.218.0** (R-354 + R-355).
| | |
|---|---|
| `GOLDEN_VERSION` | **0.218.0** |
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
| package | `https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.218.0/golden.tar.zst` |
| size | 657 026 013 B (archive 626 MB) |
| controller image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` |
| template | `debian-13-standard_13.6-1_amd64.tar.zst` (listed fresh, not assumed) |
| `MinAgent` | **0.129.0** (from the controller CHANGELOG header — unchanged) |
## Pass markers — each checked, with the negative controls
```
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
upload OK (HTTP 201) : 1 -> pre-delete returned HTTP 404 (404/204 expected)
excluding : 0 <- negative control
FATAL : 0 <- negative control
```
## Token hygiene
The Gitea token was copied **file → file** (`scp`) and read by a runner script *inside* the VM, so it
never crossed a shell or a unit property. Verified: `systemctl show golden-bake -p Environment
-p ExecStart | grep -c -F "$(cat /root/.gitea-token)"` → **0**.
**The leak grep on this committed log was itself proven before its zero was believed:** the token was
appended to a throwaway copy, grepped (**1**), the copy shredded, and only then was the committed
log's **0** accepted.
## Teardown
`pct destroy 9100 --purge`; token, runner, build script and in-VM log `shred -u`'d **after** this log
was copied out; VM powered off; qemu exited; `qemu-img snapshot -a virgin` restored.
## NOT vouched
The hub's Day-0 artifact manifest was **not** changed and the floor was **not** raised — both are the
operator's decision. See `REPORT.md` §12.
@@ -0,0 +1,323 @@
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.218.0
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
Logical volume "vm-9100-disk-0" created.
Logical volume pve/vm-9100-disk-0 changed.
Creating filesystem with 8388608 4k blocks and 2097152 inodes
Filesystem UUID: 3334dd11-f031-40d9-902c-cd36f5df48cd
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624
Logical volume "vm-9100-disk-1" created.
Logical volume pve/vm-9100-disk-1 changed.
Creating filesystem with 6291456 4k blocks and 1572864 inodes
Filesystem UUID: 4c45eb14-e05e-42d1-bc5b-a18927aec39a
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
Total bytes read: 553512960 (528MiB, 102MiB/s)
Detected container architecture: amd64
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
done: SHA256:NsUGGcy+RV+rkrR8mWKY2Oihey1ytt92IvzC1uZ8SYE root@felhom-golden
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
done: SHA256:UR39CmAMA7Bt4ljKletHmhKTs32svvU+xnpwCoxwG6U root@felhom-golden
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:kYjjxRfjLMHwSrMHH7h7QEKGXzMwDy1KtD/s/TtaDww root@felhom-golden
[golden] starting + installing Docker (official repo, trixie channel) …
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
Unable to find image 'hello-world:latest' locally
latest: Pulling from library/hello-world
4f55086f7dd0: Pulling fs layer
4f55086f7dd0: Verifying Checksum
4f55086f7dd0: Download complete
4f55086f7dd0: Pull complete
Digest: sha256:5dd0d3e6e255913fc30f90b9f2b1d359cc2cbdb48090cc4b65f1676e203243cc
Status: Downloaded newer image for hello-world:latest
docker OK (overlay2; data-root /var/lib/docker)
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.218.0 (no registry cred at deploy) …
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
Configure a credential helper to remove this warning. See
https://docs.docker.com/go/credential-store/
0.218.0: Pulling from admin/felhom-controller
039e6f9f9752: Pulling fs layer
0094c3ac0914: Pulling fs layer
deca1dac7403: Pulling fs layer
11c19a33d1b8: Pulling fs layer
3ba1be8b4b39: Pulling fs layer
2c80200460b0: Pulling fs layer
11c19a33d1b8: Waiting
3ba1be8b4b39: Waiting
2c80200460b0: Waiting
deca1dac7403: Verifying Checksum
deca1dac7403: Download complete
11c19a33d1b8: Verifying Checksum
11c19a33d1b8: Download complete
3ba1be8b4b39: Verifying Checksum
3ba1be8b4b39: Download complete
2c80200460b0: Verifying Checksum
2c80200460b0: Download complete
0094c3ac0914: Verifying Checksum
0094c3ac0914: Download complete
039e6f9f9752: Download complete
039e6f9f9752: Pull complete
0094c3ac0914: Pull complete
deca1dac7403: Pull complete
11c19a33d1b8: Pull complete
3ba1be8b4b39: Pull complete
2c80200460b0: Pull complete
Digest: sha256:56c35264346e4f4e4eca0f3405d2ca727f29a7b4d1e73fbffbaf1f0835bd035e
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.218.0
gitea.dooplex.hu/admin/felhom-controller:0.218.0
[golden] asking the controller which infra images it manages …
[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
v3.6.7: Pulling from library/traefik
589002ba0eae: Pulling fs layer
ef63511ea6cc: Pulling fs layer
0738e5cb835e: Pulling fs layer
3e6813f70c64: Pulling fs layer
3e6813f70c64: Waiting
589002ba0eae: Verifying Checksum
589002ba0eae: Download complete
ef63511ea6cc: Verifying Checksum
ef63511ea6cc: Download complete
3e6813f70c64: Verifying Checksum
3e6813f70c64: Download complete
0738e5cb835e: Verifying Checksum
0738e5cb835e: Download complete
589002ba0eae: Pull complete
ef63511ea6cc: Pull complete
0738e5cb835e: Pull complete
3e6813f70c64: Pull complete
Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
Status: Downloaded newer image for traefik:v3.6.7
docker.io/library/traefik:v3.6.7
2026.6.0: Pulling from cloudflare/cloudflared
47de5dd0b812: Pulling fs layer
c172f21841df: Pulling fs layer
99515e7b4d35: Pulling fs layer
99ba982a9142: Pulling fs layer
d6b1b89eccac: Pulling fs layer
2780920e5dbf: Pulling fs layer
7c12895b777b: Pulling fs layer
3214acf345c0: Pulling fs layer
52630fc75a18: Pulling fs layer
dd64bf2dd177: Pulling fs layer
b839dfae01f6: Pulling fs layer
ebddc55facdc: Pulling fs layer
bdfd7f7e5bf6: Pulling fs layer
2d4d7adf6272: Pulling fs layer
40008157d8d2: Pulling fs layer
bd8962e29291: Pulling fs layer
cac2ae0193cb: Pulling fs layer
74d1dac84ecc: Pulling fs layer
99ba982a9142: Waiting
b839dfae01f6: Waiting
ebddc55facdc: Waiting
bdfd7f7e5bf6: Waiting
d6b1b89eccac: Waiting
2780920e5dbf: Waiting
7c12895b777b: Waiting
3214acf345c0: Waiting
2d4d7adf6272: Waiting
52630fc75a18: Waiting
40008157d8d2: Waiting
bd8962e29291: Waiting
dd64bf2dd177: Waiting
cac2ae0193cb: Waiting
74d1dac84ecc: Waiting
47de5dd0b812: Download complete
c172f21841df: Verifying Checksum
c172f21841df: Download complete
d6b1b89eccac: Verifying Checksum
d6b1b89eccac: Download complete
2780920e5dbf: Verifying Checksum
2780920e5dbf: Download complete
99ba982a9142: Verifying Checksum
99ba982a9142: Download complete
47de5dd0b812: Pull complete
7c12895b777b: Verifying Checksum
7c12895b777b: Download complete
3214acf345c0: Verifying Checksum
3214acf345c0: Download complete
52630fc75a18: Verifying Checksum
52630fc75a18: Download complete
c172f21841df: Pull complete
dd64bf2dd177: Verifying Checksum
dd64bf2dd177: Download complete
99515e7b4d35: Verifying Checksum
99515e7b4d35: Download complete
b839dfae01f6: Download complete
ebddc55facdc: Verifying Checksum
ebddc55facdc: Download complete
bdfd7f7e5bf6: Verifying Checksum
bdfd7f7e5bf6: Download complete
bd8962e29291: Verifying Checksum
bd8962e29291: Download complete
cac2ae0193cb: Verifying Checksum
cac2ae0193cb: Download complete
99515e7b4d35: Pull complete
2d4d7adf6272: Download complete
99ba982a9142: Pull complete
74d1dac84ecc: Verifying Checksum
74d1dac84ecc: Download complete
40008157d8d2: Verifying Checksum
40008157d8d2: Download complete
d6b1b89eccac: Pull complete
2780920e5dbf: Pull complete
7c12895b777b: Pull complete
3214acf345c0: Pull complete
52630fc75a18: Pull complete
dd64bf2dd177: Pull complete
b839dfae01f6: Pull complete
ebddc55facdc: Pull complete
bdfd7f7e5bf6: Pull complete
2d4d7adf6272: Pull complete
40008157d8d2: Pull complete
bd8962e29291: Pull complete
cac2ae0193cb: Pull complete
74d1dac84ecc: Pull complete
Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0
docker.io/cloudflare/cloudflared:2026.6.0
1.3.3-stable: Pulling from gtstef/filebrowser
6a0ac1617861: Pulling fs layer
ef8806083e82: Pulling fs layer
b74107c861c7: Pulling fs layer
adc935def003: Pulling fs layer
4f4fb700ef54: Pulling fs layer
18695ccc900a: Pulling fs layer
45d119d5c397: Pulling fs layer
dac52db4fc51: Pulling fs layer
6d598f86b2f2: Pulling fs layer
8aa349c8396c: Pulling fs layer
18695ccc900a: Waiting
45d119d5c397: Waiting
dac52db4fc51: Waiting
6d598f86b2f2: Waiting
8aa349c8396c: Waiting
adc935def003: Waiting
4f4fb700ef54: Waiting
6a0ac1617861: Download complete
b74107c861c7: Verifying Checksum
b74107c861c7: Download complete
adc935def003: Download complete
4f4fb700ef54: Verifying Checksum
4f4fb700ef54: Download complete
45d119d5c397: Verifying Checksum
45d119d5c397: Download complete
dac52db4fc51: Verifying Checksum
dac52db4fc51: Download complete
ef8806083e82: Verifying Checksum
ef8806083e82: Download complete
6d598f86b2f2: Verifying Checksum
6d598f86b2f2: Download complete
18695ccc900a: Verifying Checksum
18695ccc900a: Download complete
8aa349c8396c: Verifying Checksum
8aa349c8396c: Download complete
6a0ac1617861: Pull complete
ef8806083e82: Pull complete
b74107c861c7: Pull complete
adc935def003: Pull complete
4f4fb700ef54: Pull complete
18695ccc900a: Pull complete
45d119d5c397: Pull complete
dac52db4fc51: Pull complete
6d598f86b2f2: Pull complete
8aa349c8396c: Pull complete
Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable
docker.io/gtstef/filebrowser:1.3.3-stable
1.1.0: Pulling from admin/felhom-samba
897d797d2723: Pulling fs layer
3051591aa250: Pulling fs layer
ce57a3f93416: Pulling fs layer
fb94eeec2fe1: Pulling fs layer
fb94eeec2fe1: Waiting
ce57a3f93416: Verifying Checksum
ce57a3f93416: Download complete
fb94eeec2fe1: Verifying Checksum
fb94eeec2fe1: Download complete
897d797d2723: Verifying Checksum
897d797d2723: Download complete
3051591aa250: Verifying Checksum
3051591aa250: Download complete
897d797d2723: Pull complete
3051591aa250: Pull complete
ce57a3f93416: Pull complete
fb94eeec2fe1: Pull complete
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
gitea.dooplex.hu/admin/felhom-samba:1.1.0
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
[golden] identity-clean + minimize …
[golden] stop + archive …
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
INFO: archive file size: 626MB
INFO: Finished Backup of VM 9100 (00:00:31)
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_08_22-10_07_08.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
[golden] publishing golden (657026013 bytes, sha256 8e427869d13eafb7…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.218.0/golden.tar.zst
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.218.0
GOLDEN_SHA256=8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.218.0 / 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)