R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367.
This commit is contained in:
@@ -0,0 +1,599 @@
|
||||
# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night)
|
||||
|
||||
**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no
|
||||
image built, nothing deployed.** Evidence:
|
||||
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
|
||||
|
||||
---
|
||||
|
||||
## 1. THE VERDICT
|
||||
|
||||
**It is a mixture, and the drill's three options are all present — but they belong to different
|
||||
faults, and only one of them explains the thing you actually saw.**
|
||||
|
||||
### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED**
|
||||
|
||||
Reproduced independently tonight, and it agrees with what R-353 already recorded:
|
||||
|
||||
1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data
|
||||
drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** —
|
||||
about an app that was installed, deployed, running and healthy.
|
||||
2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt
|
||||
box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it
|
||||
held `compose/` and nothing else.
|
||||
3. The restore returned that configuration and reported a bare completion.
|
||||
|
||||
So for that specific event the data was not in the thing that was restored. **R-353 called this
|
||||
correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong
|
||||
to think so; its text is more careful than the CHANGELOG's summary of it.
|
||||
|
||||
### What the drill found that nobody had seen: **THE RESTORE LOSES IT**
|
||||
|
||||
This is new, it is worse, and it is not the same fault:
|
||||
|
||||
**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.**
|
||||
The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.**
|
||||
|
||||
Proven live on `calibre-web` at 22:23:36 with planted files:
|
||||
|
||||
| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore |
|
||||
|---|---|---|---|---|
|
||||
| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** |
|
||||
| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** |
|
||||
|
||||
and the customer was told:
|
||||
|
||||
> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult.
|
||||
> Ennek az alkalmazásnak nincs adatbázisa."
|
||||
|
||||
Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence
|
||||
does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue
|
||||
apps that declare no data drive, that leg is the entire dataset.**
|
||||
|
||||
**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit`
|
||||
(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the
|
||||
unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local**
|
||||
restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring
|
||||
PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of
|
||||
them replays it.**
|
||||
|
||||
### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database
|
||||
|
||||
`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always
|
||||
has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to
|
||||
`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack
|
||||
that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the
|
||||
safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and
|
||||
then says:
|
||||
|
||||
> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**"
|
||||
|
||||
The controller had dumped that database five minutes earlier.
|
||||
|
||||
**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive.
|
||||
|
||||
---
|
||||
|
||||
## 2. THE FULL/EMPTY CONTRAST
|
||||
|
||||
**Apps chosen, and why.** From the two storage classes: **`calibre-web`** declares a data drive
|
||||
(`needs_hdd: true`, `backup.userdata: media/books class: mandatory`) and **`privatebin`** /
|
||||
**`opengist`** declare none (the 40-class; data lives in a Docker named volume). I ran **three** cases
|
||||
rather than two, deliberately: FULL and EMPTY alone cannot separate *"it was empty"* from *"the class
|
||||
is broken"*, so `privatebin` was run FULL as the disambiguator.
|
||||
|
||||
**The fixture** — 5 files, two with Hungarian accented names, recorded as raw bytes:
|
||||
|
||||
```
|
||||
SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991
|
||||
binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a
|
||||
nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4
|
||||
plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1
|
||||
árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea
|
||||
|
||||
name bytes (UTF-8 NFC):
|
||||
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
|
||||
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
|
||||
```
|
||||
|
||||
**The comparator was proved able to convict before it was trusted.** One byte flipped at offset
|
||||
500 000 of `binary-1mb.bin` (`af` → `00`): `sha256sum -c` reported `binary-1mb.bin: FAILED`, rc=1,
|
||||
while the other four passed; the unmodified set passed rc=0. The mutant was discarded.
|
||||
|
||||
### The results
|
||||
|
||||
| case | class | data leg | in unit | in off-site snapshot | checking folder | off-site restore | local restore |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **calibre-web FULL** | drive | user files | n/a | 5/5 identical | 5/5 identical | **5/5 identical** | — |
|
||||
| **calibre-web FULL** | drive | named volume 1.42 MB | yes | yes | yes | **NOT restored** | — |
|
||||
| **privatebin FULL** | no drive | named volume 1.06 MB | yes | 5/5 identical | 5/5 identical | **REFUSED — false reason** | **5/5 identical** |
|
||||
| **opengist EMPTY** | no drive | named volume 181 KB skeleton | yes | yes | yes | **REFUSED — false reason** | — |
|
||||
|
||||
**The contrast decides it.** The unit and the off-site snapshot **hold the data, verified by
|
||||
identity**, in both classes, accented filenames included. So the capture is sound. The failure is
|
||||
entirely in the last leg.
|
||||
|
||||
**Both messages, verbatim:**
|
||||
|
||||
- FULL, drive class → `ok=true`, *„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16)
|
||||
— az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."*
|
||||
- FULL and EMPTY, no-drive class → `ok=false`, *„a(z) privatebin nincs telepítve, ezért nincs hová
|
||||
visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra
|
||||
az alkalmazást (Alkalmazások) ugyanerre a helyre…"*
|
||||
|
||||
The second is the important one. **FULL and EMPTY got the identical sentence**, so the message cannot
|
||||
distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is
|
||||
installed, and the instruction ("reinstall it to the same place") **cannot be followed**, because a
|
||||
40-class app is offered no storage field at deploy time (that is R-352's own measurement).
|
||||
|
||||
**This closes R-353's second instruction.** It asked for proof that a named-volume app reaches the
|
||||
off-site tier rather than inference from gate order. It does. Manifests read tonight:
|
||||
|
||||
```
|
||||
privatebin volume_dumps = ['privatebin_privatebin_data.tar'] db_dumps = None
|
||||
opengist volume_dumps = ['opengist_opengist_data.tar'] db_dumps = None
|
||||
kimai volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar']
|
||||
db_dumps = ['kimai-mariadb.sql']
|
||||
calibre-web volume_dumps = ['calibre-web_calibre_web_config.tar'] db_dumps = None
|
||||
paperless-ngx volume_dumps = [3 tars] db_dumps = None ← the defect
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. PART 0 — did the floor move the machine?
|
||||
|
||||
**Yes, unaided, in 17 seconds.**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| moved from → to | **0.216.0 → 0.217.0** |
|
||||
| who initiated | **the hub** — the operator saved the floor at `19:48:37Z`; a poke reached the agent from `10.77.0.1:58093` at `19:48:32Z` for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision. |
|
||||
| how long | floor saved `19:48:37Z` → *"controller-swap: new controller healthy"* `19:48:54Z` = **17 s**. Swap requested `19:48:41Z` → healthy = 13 s. |
|
||||
| **the container's own tag on the box** | **`gitea.dooplex.hu/admin/felhom-controller:0.217.0`**, created `2026-08-21 19:48:45 UTC` — read from `docker ps` in guest 9201, not from the hub. |
|
||||
|
||||
The hub's own host page for `demo-hp` shows the guest's **Controller column as „—"** — the hub does
|
||||
not know which controller version the box runs, while the `controller_updated` event it received says
|
||||
exactly that. Two hub surfaces, one blind.
|
||||
|
||||
`demo-felhom` also runs 0.217.0, but it started it at `19:31:04Z` — **17 minutes before the floor was
|
||||
saved** — and emitted `controller_started` with **no `controller_updated`**. It was moved by hand
|
||||
during the golden bake, not by the floor.
|
||||
|
||||
---
|
||||
|
||||
## 4. PART 2 — the four answers, from source
|
||||
|
||||
**1. What puts a dump into a backup unit, and when? Which apps qualify?**
|
||||
|
||||
- **Database dumps** — `runDBDumps` (`backup/backup.go:444-560`). Qualification is **a running
|
||||
container whose image matches a database image**, mapped to a stack by `deriveStackName`
|
||||
(`appbackup/dbdump.go:770`). Written to `<nsRoot>/backups/primary/<stack>/db-dumps/<stack>-<type>.sql`.
|
||||
- **Volume dumps** — `runVolumeDumps` (`backup/backup.go:607`). Qualification is **a deployed,
|
||||
non-protected stack with at least one Docker *named volume*** (`GetDockerVolumes`). An app with
|
||||
zero named volumes is skipped silently and is never stopped.
|
||||
- **When** — one cycle, both legs, scheduled `db-dump` daily at **02:30 CEST**; and again as the
|
||||
coherence pre-phase of every off-site run (`offbox.go:938-951`), which is what makes a snapshot an
|
||||
internally coherent {DB@T, files@T} pair.
|
||||
- `CaptureRecoveryUnit` **only enumerates what is already on disk**
|
||||
(`recovery_unit.go:131-132`). It writes no dump itself.
|
||||
- **The gap this leaves:** an app whose data is a *bind mount* and which has no database container
|
||||
produces neither leg. Its unit is configuration only — and nothing anywhere says so.
|
||||
|
||||
**2. What should a correct backup contain?**
|
||||
|
||||
- **Declares a data drive** (13 of 53): the recovery unit (compose + `app.yaml` carrying the portable
|
||||
secrets + whatever dumps exist) **plus** the paths its `.felhom.yml` `backup:` block marks
|
||||
`mandatory`, appended to the restic snapshot as extra paths (`offbox_capture.go:32`).
|
||||
- **Declares none** (40 of 53): **unit only**. `offboxCaptureSet` returns `(nil, nil)` when the app
|
||||
has no backup block, so the off-site snapshot is the unit and nothing else. That is correct *given*
|
||||
that their data is inside the unit's volume tars — and tonight confirmed it is.
|
||||
|
||||
**3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.**
|
||||
|
||||
`ReconstituteFromOffsite` returns `res` and `nil`. Nothing in it ever asks whether anything was
|
||||
restored. The handler then calls `EndRestoreOp(**true**, reconstituteOutcomeMsg(...))`
|
||||
(`web/offbox_handlers.go:455`). `reconstituteOutcomeMsg` (`web/offbox_handlers.go:463-475`) formats
|
||||
counters:
|
||||
|
||||
```go
|
||||
if res.DBsReplayed == 0 {
|
||||
return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+
|
||||
"Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when)
|
||||
}
|
||||
```
|
||||
|
||||
So with `FilesPlaced == 0` and `DBsReplayed == 0` the customer reads **"0 files restored — the
|
||||
application restarted"** under a **success**. **This is a finding on its own and is recorded whichever
|
||||
way the rest goes.** Two aggravations found on top of it:
|
||||
|
||||
- **"This application has no database" is asserted from a counter, not from a fact.**
|
||||
`reimportDBDumpsFrom` returns `(0, nil)` when the dump directory is absent *or* holds no `.sql`
|
||||
(`restore_db.go:30-46`), so an app that certainly has a database is told it has none. Proven live on
|
||||
`paperless-ngx`.
|
||||
- **The file count is rsync's transfer count, not a restore count.** A correct restore of unchanged
|
||||
data reports **"0 fájl visszaállítva"** — indistinguishable from a restore that did nothing.
|
||||
Observed: the same app reported 5, then 2, then 0 across three runs.
|
||||
|
||||
**4. What did the 9 August off-site snapshot contain, and is it still readable?**
|
||||
|
||||
**Readable, and it contained the data.** Three snapshots at `2026-08-09 08:30`:
|
||||
|
||||
| id | app | contents |
|
||||
|---|---|---|
|
||||
| `41c830db` | calibre-web | unit + `userdata/media/books` incl. the 9-Aug rehearsal fixture (`árvíztűrő-tükörfúrógép.txt`, `őszibarack.md`, `binary-3mb.bin`, `plain.txt`) |
|
||||
| `9e38b84c` | opengist | unit + **`volume-dumps/opengist_opengist_data.tar`, 182 272 B** — manifest records `volume_dumps: ["opengist_opengist_data.tar"]` |
|
||||
| `78b93f04` | privatebin | unit + `volume-dumps/privatebin_privatebin_data.tar`, 2 560 B |
|
||||
|
||||
All still readable tonight; `restic check` over the whole repository reports **"no errors were
|
||||
found"**. **So the 9 August off-site copy of OpenGist's data exists and is intact** — it simply was
|
||||
not the thing the afternoon's restore read, and the off-site route that would have read it refuses
|
||||
for that class of app.
|
||||
|
||||
**The repository had also been dead since 9 August** and nothing said so: no snapshot between
|
||||
`2026-08-09 08:30` and tonight. The cause is visible in the hub event at 16:01 — the off-box target
|
||||
was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target,
|
||||
**every per-app off-site toggle was still off**, so the first run I triggered reported:
|
||||
|
||||
```
|
||||
[offbox] backup run started (0 app(s) toggled)
|
||||
[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s
|
||||
```
|
||||
|
||||
**Credit where it is due:** the *card* does not lie about this. It reads „Aktív — nincs kijelölt
|
||||
alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still
|
||||
leads, and the log line alone says only „backup OK".
|
||||
|
||||
---
|
||||
|
||||
## 5. THE TWO CYCLES COMPARED — AND THEY AGREE
|
||||
|
||||
**By hand:** local cycle 22:12, off-site 22:17 (after enabling the per-app switches the rebuild had
|
||||
silently cleared). **By the clock:** the box's own `db-dump` at 02:30 CEST, cross-drive at 03:30,
|
||||
off-site at 04:15.
|
||||
|
||||
| | manual | scheduled | agree? |
|
||||
|---|---|---|---|
|
||||
| units refreshed | all 6 apps | all 6 apps, `00:30Z` | **yes** |
|
||||
| `privatebin` volume tar | 1 055 744 B | 1 055 744 B | **yes** |
|
||||
| `opengist` volume tar | 181 248 B | 181 248 B | **yes** |
|
||||
| `calibre-web` config tar | 1 422 848 B | 368 640 B † | **yes, and explained** |
|
||||
| `paperless-ngx` `db_dumps` | **absent** | **absent** | **yes — the defect reproduces on the scheduled path** |
|
||||
| the orphan `…/primary/paperless/db-dumps/` | written | **written again, 294 936 B, `00:30Z`** | **yes** |
|
||||
| planted files on the drive | 5/5 identical | **5/5 identical**, plus `POST-SNAPSHOT.txt` | **yes** |
|
||||
| off-site run | ok, 27 → snapshots | ok, `02:17:12Z`, **2m8s**, 42.8 MB, **27 snapshots** | **yes** |
|
||||
|
||||
† the tar shrank because the drill had by then removed the planted 1 MB from that volume — the
|
||||
expected value, not a discrepancy.
|
||||
|
||||
**Two independent observations, no disagreement.** The one that matters: **R-355 is not an artefact of
|
||||
my manual triggering.** The box, unattended, on its own schedule, wrote paperless-ngx's PostgreSQL
|
||||
dump into a directory for a stack that does not exist and left the app's own unit recording
|
||||
`db_dumps: null`.
|
||||
|
||||
The scheduled cross-drive (Tier-2) leg also completed for all six apps at 01:30Z — `crossdrive_completed`
|
||||
events 3019–3024.
|
||||
|
||||
---
|
||||
|
||||
## 6. PART 4 — everything attempted, and the message judged
|
||||
|
||||
| # | test | outcome | message judged |
|
||||
|---|---|---|---|
|
||||
| 1a | **destructive restore: nothing is ever deleted** | **PASS both ways.** A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. | count is rsync transfers, not files restored — see §4.3 |
|
||||
| 1b | **safety dump taken and verified before the stop** | **PASS.** `21:02:47Z` dump written → `21:02:47Z` `StopStack romm`. | correct: „0 fájl **és az adatbázis** visszaállítva" |
|
||||
| 1b | **…and the whole operation refuses if it cannot be** | **PASS.** Made impossible by putting a regular file at the `db-dumps` path. Refused; `plain.txt` kept its post-snapshot mutation; the container's `StartedAt` was unchanged — **the app was never stopped.** | honest and names the path, but leaks a raw Go `mkdir … not a directory` into a customer surface |
|
||||
| 1b | **the hole in it** | **FAIL.** For `paperless-ngx` the discovery resolves the database to the wrong stack, so `hasDB` is false: **no safety dump is taken and the refusal cannot fire.** The undo is absent, not refused. | „Ennek az alkalmazásnak nincs adatbázisa" — false |
|
||||
| 2 | **end of the abandonment countdown** | state created and overdue; **fires at 05:10 CEST** — see §11 | card states a **past** date in the future tense while overdue |
|
||||
| 3 | **damaged store** | **MIXED — see below** | honest at the point of failure, then forgotten |
|
||||
| 4 | **drive pulled mid-restore** | restore failed, nothing written to the wrong place, **the agent re-bound the drive within 5 s** (`23:15:29` pulled → `23:15:34` re-bound) | **WRONG DIAGNOSIS.** „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions. |
|
||||
| 5 | **full disk, non-destructive path** | **PASS — refuses before it starts.** | **exemplary:** „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named. |
|
||||
| 5 | **full disk, destructive path** | **FAIL — no gate at all.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the three headroom gates are all on non-destructive paths (`offbox_restore.go:231,297,423`). It stopped the app, failed halfway on ENOSPC, left `DRILL-2026-08-21/` holding 2 of 5 entries, and restarted the app. | honest (`No space left on device (28)`) but raw rsync output |
|
||||
| 6 | **controller killed mid-restore** | **PASS.** SIGKILL inside the stop→restore→start window. On restart: *"[appstop] crash recovery: an off-site restore (op \"offbox-reconstitute:paperless-ngx\") was interrupted and left 1 app(s) stopped — restarting them"*. App restarted, marker cleared, **and the hub was told** (event 3009). | good — names the operation, not just the app |
|
||||
| 7 | **the 40-class under pressure** | **the backup noticed; the alarm did not.** Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The **reserve refused per app**: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got `recovery_unit_capture_failed` (**error**) naming the filesystem and its numbers. **The fill watcher said nothing** — it runs **once a day at 03:30** plus once at startup (`cmd/controller/main.go:1092`). | the refusal messages are good; the silence is the problem |
|
||||
|
||||
**Where the clock stopped me:** nothing in Part 4 was skipped for time. Item 2's terminal deletion is
|
||||
scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only.
|
||||
|
||||
### 4.3 in detail — the damaged store
|
||||
|
||||
One byte flipped inside pack `967853d2…` at offset 5 000 000, over the repository's own SFTP
|
||||
transport. Pack files are named by their content hash, so this is genuine corruption.
|
||||
|
||||
- **`restic check` detects it** — *"ciphertext verification failed"*, *"Fatal: repository contains
|
||||
errors"*.
|
||||
- **But nothing in the product ever runs it.** The only restic verbs in the entire controller are
|
||||
`restore, snapshots, backup, unlock, stats, init, forget, prune, cat`. The agent's restore-test is
|
||||
**PBS-tier only**. **The off-site store is never verified by any layer, at any time.** Corruption is
|
||||
discovered at restore time — the worst possible moment.
|
||||
- **A restore that touches the damage fails honestly:** `ok=false`, naming the file and
|
||||
*"ciphertext verification failed"*.
|
||||
- **But the failure is not remembered.** It left a **partial** checking folder — 78 MB, 54 files,
|
||||
15 of 16 originals. `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only *"does the
|
||||
directory exist and is it non-empty"*, so the wizard then offered **all three** actions including
|
||||
„Teljes visszaállítás indítása".
|
||||
- **Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.**
|
||||
|
||||
The repository was repaired from byte-identical originals; `restic check` now reports **"no errors
|
||||
were found"**.
|
||||
|
||||
---
|
||||
|
||||
## 7. PART 5 — the two rows that were observed and never filed
|
||||
|
||||
**5.1 — the delete guard on verification copies is blind, and its own comment says otherwise.**
|
||||
`offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) gates on `backupMgr.IsRunning()` — the
|
||||
**concurrency** flag — while its comment states *"It refuses while a backup/restore op is running: the
|
||||
copy being deleted could be the one currently being written."* R-351b moved all seven restore handlers
|
||||
onto `restoreOpBlocked()` (which reads both flags); **this handler was left behind**, and one other
|
||||
site (`:239`) reads the bare flag correctly and documents why.
|
||||
|
||||
**Reachability is not a race — it is the whole operation.** `RestoreOffboxScratch`
|
||||
(`offbox_restore.go:211`) **never calls `acquireRunning` at all**, so `IsRunning()` is false for the
|
||||
entire duration of an off-site verification restore. Demonstrated live, flags read immediately before
|
||||
and after the delete:
|
||||
|
||||
```
|
||||
--- BEFORE delete 22:35:25 display= True offbox-restore kimai concurrency= False
|
||||
--- DELETE calibre-web verification copy: „Az ellenőrző másolat törölve…"
|
||||
--- AFTER delete 22:35:25 display= True offbox-restore kimai concurrency= False
|
||||
```
|
||||
|
||||
The copy was removed. The guard is app-agnostic, so the same call naming the *restoring* app hits the
|
||||
directory the restore is writing into. **Rank: MEDIUM** — it needs a customer to press delete during a
|
||||
restore, but both controls live on the same page, the window is the whole restore, and the target is
|
||||
the restore's own source. (Reconstitute and place *do* hold the flag, so the exposure is the
|
||||
verification-restore window only.)
|
||||
|
||||
**5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight.**
|
||||
Filed as an instrument defect. **Occurrences I can evidence:**
|
||||
|
||||
1. **2026-07-20** — an accented grep through `ssh → pct exec → bash -c` nearly produced a wrong
|
||||
"banner cleared" claim. Recorded in `felhom-controller/.claude/rules/ui-hungarian.md:19-22`.
|
||||
2. **2026-08-13** — `kubectl exec … sh -c "grep '<accented>'"` returned **0 for three strings that
|
||||
were present**, one step from being reported as a failed hub v0.105.0 deploy.
|
||||
3. **2026-08-21, tonight, 22:07** — `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`
|
||||
(octal escaping). Recording the accented filenames' raw bytes from that listing would have been
|
||||
wrong. Caught by extracting the archive and reading the names with `xxd`.
|
||||
|
||||
**Correction to the task's premise:** that is **two inside two weeks**, plus the founding case a month
|
||||
earlier. I looked for a third inside the two-week window and did not find one on record.
|
||||
|
||||
**The smallest guard I would propose — and did NOT build:** the problem is not grep, it is that every
|
||||
one of these tools *silently transforms* the bytes. So the guard is not "use ASCII fragments" (a
|
||||
discipline, which is what failed three times) but **a negative control that the harness cannot skip**:
|
||||
any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a
|
||||
string that MUST be absent — and a zero result from the first is only reportable when the second also
|
||||
returns zero *and* a third probe for a known-present ASCII anchor returns non-zero. Three probes, one
|
||||
helper, no judgement required at the call site. Everything else has been tried and is what "nearly"
|
||||
means in all three cases.
|
||||
|
||||
---
|
||||
|
||||
## 8. RANKED REGISTER ROWS OPENED
|
||||
|
||||
Ceiling was **R-353**; it **moved to R-366**.
|
||||
|
||||
| id | rank | what |
|
||||
|---|---|---|
|
||||
| **R-354** | **HIGH** | The off-site full restore has **no named-volume leg**. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. `offbox_reconstitute.go:341-346`. |
|
||||
| **R-355** | **HIGH** | `paperless-ngx`'s PostgreSQL is dumped to a directory for a **non-existent stack** (`…/primary/paperless/`), so its unit records `db_dumps: null`, nothing off-sites it, **no safety dump is taken on a destructive restore**, and the customer is told the app has no database. `appbackup/dbdump.go:770-798`. One app in 53. |
|
||||
| **R-356** | **HIGH** | The off-site restore **refuses for all 40 no-drive apps** with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. `offbox_reconstitute.go:208-227`. |
|
||||
| **R-357** | **MEDIUM** | The **destructive** restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory. |
|
||||
| **R-358** | **MEDIUM** | A **failed** scratch restore leaves a partial copy that `OffboxFullScratchReady` reports as ready; the destructive restore then runs from it and reports success. |
|
||||
| **R-359** | **MEDIUM** | The off-site restic store is **never verified** by any layer — `restic check` is not among the verbs the controller runs, and the agent's restore-test is PBS-only. |
|
||||
| **R-360** | **MEDIUM** | Verification-copy delete gates on `IsRunning()`, which `RestoreOffboxScratch` never holds — deletable throughout a restore. Its comment asserts the opposite. **(Part 5.1)** |
|
||||
| **R-361** | **MEDIUM** | The safety dump **overwrites the unit's own DB dump**: `DumpOne` writes the canonical `<stack>-<type>.sql` and only then renames it away. The comment at `offbox_reconstitute.go:147-148` states it "can never overwrite the app's real dump". Proven: romm's `romm-mariadb.sql` was present before and absent after. |
|
||||
| **R-362** | **MEDIUM** | A data drive detached mid-restore is reported as **„permission denied"**. The restore path never consults drive state. |
|
||||
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
|
||||
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
|
||||
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
|
||||
| **R-366** | **HIGH** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives too** — the box cannot read its own pre-reinstall backups (`manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`, hub event 3016, filed unprompted at 21:59). The PBS-tier analogue of R-193: a rebuilt box loses **both** off-premises tiers at once. The restore-test caught it precisely; the gap is that it is called *a failed restore test* rather than *your older whole-guest backups are unreadable*. **Found incidentally — nobody was looking for it.** |
|
||||
|
||||
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
|
||||
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
|
||||
nobody. Observed tonight as event 3006, severity `info`. A sweep of every `emit(` call site shows this
|
||||
is now the **only** remaining instance fleet-wide.
|
||||
|
||||
**R-353:** its instruction (2) is **satisfied** — see §2. Its instruction (1) stands and is now
|
||||
strictly larger than when it was written, because R-354 shows the bare completion can also be reported
|
||||
over a unit that *did* have a data leg.
|
||||
|
||||
---
|
||||
|
||||
## 9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH
|
||||
|
||||
| time (CEST) | what |
|
||||
|---|---|
|
||||
| 21:57 | start; baselines |
|
||||
| 22:00 | break-glass into `demo-hp` |
|
||||
| 22:09 | fixture planted, comparator control passed |
|
||||
| 22:12 | manual local backup |
|
||||
| 22:13 | manual off-site run — **0 apps toggled** |
|
||||
| 22:17 | off-site run with 3 apps enabled |
|
||||
| 22:19–22:20 | checking-folder restores |
|
||||
| 22:21 | reconstitute privatebin → refused |
|
||||
| 22:23 | reconstitute calibre-web → **the conviction** |
|
||||
| 22:25 | local restore-from-unit privatebin → data returned |
|
||||
| 22:27 | Part 4.1a — no-delete invariant |
|
||||
| 22:39–22:45 | paperless-ngx → R-355 |
|
||||
| 22:51–22:57 | Part 4.3 damaged store; repo repaired |
|
||||
| 23:02–23:04 | Part 4.1b safety dump, both directions |
|
||||
| 23:10–23:13 | Part 4.5 full disk (both paths); Part 4.6 controller killed mid-restore |
|
||||
| 23:15 | Part 4.4 drive pulled |
|
||||
| 23:17–23:26 | Part 4.7 filesystem filled, backup reserve observed, filesystem freed |
|
||||
| 23:08–23:09 | abandonment set-aside store created; countdown written and controller restarted (fires 05:10) |
|
||||
|
||||
**Steps off the customer's path, named:**
|
||||
|
||||
1. **Break-glass root access** to `demo-hp` via the hub-vaulted `host_recovery` credential — the box
|
||||
had lost DooPlex's SSH key (its `authorized_keys` held only its own `root@demo-hp` RSA key) and its
|
||||
tailnet address was unreachable. DooPlex's public key was **re-added** to `/root/.ssh/authorized_keys`
|
||||
and an `ssh` alias `hp` → `192.168.0.104` was added to `~/.ssh/config` on DooPlex.
|
||||
2. **A hub DB snapshot** (`hub.db` + `-wal` + `-shm`) was streamed to the scratchpad to read
|
||||
`host_recovery` and the events table. It holds every host's secret; it is in the session scratchpad
|
||||
only and is not in any committed file.
|
||||
3. **`settings.json` was edited directly** (controller stopped, backup at `/root/settings.json.drill-backup`)
|
||||
to create the overdue abandonment state. There is no product path to shorten a countdown, and the
|
||||
real orphan→reset path would have destroyed `demo-hp`'s entire off-site history.
|
||||
4. **The set-aside store the sweep will delete was created by hand** at
|
||||
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, for the same reason.
|
||||
5. **Two restic pack files were deliberately corrupted and then restored** from byte-identical copies.
|
||||
6. **`fallocate` fillers** were used to fill two filesystems and were removed.
|
||||
7. Apps deployed for the drill: `opengist` (re-deployed empty), `kimai`, `paperless-ngx`, `romm`.
|
||||
|
||||
---
|
||||
|
||||
## 10. THE FENCES
|
||||
|
||||
- **`ep0` / the off-site endpoint — untouched outside this machine's own path, and here is how I know.**
|
||||
`demo-hp` authenticates as the Hetzner Storage Box **sub-account `u629488-sub3`**, which is chrooted
|
||||
to its own home: `ls /` returns **`Permission denied`**, and `/home` contains exactly `.ssh` and
|
||||
`felhom-repo`. Every write, the two corruptions, the set-aside store and the sweep's target are
|
||||
inside that home. **`demo-felhom` is a different sub-account (`u629488-sub1`)** and a real customer's
|
||||
copy is a different sub-account again — none reachable with this key. `restic forget`/`prune` were
|
||||
never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept
|
||||
every 9-August snapshot (verified by listing all 24, not by a count).
|
||||
- **`demo-felhom`'s two fixtures — confirmed intact, not assumed.**
|
||||
(a) the unopenable set-aside store `u629488-sub1:/home/felhom-repo.orphaned-20260810` — listed
|
||||
tonight, `config`/`data` (258 shards)/`index`/`keys`/`locks`/`snapshots`, mtimes still 18 Jul and
|
||||
3 Aug; (b) the retained-key case — hub `host_escrow_superseded` rows **11 and 12** for
|
||||
`demo-felhom-8363b5`, each with a 572-byte `identity_blob`, dated 2026-08-12. `demo-felhom` is
|
||||
healthy on 0.217.0, its own off-site ran at 02:15 with `last_status: ok`, and
|
||||
**`--abandon-status` there reports „no abandonment countdown is running on this box"**.
|
||||
- **`peti-felhom` — not contacted.** It does not appear in the hub host list at all; no command in this
|
||||
session named it.
|
||||
|
||||
---
|
||||
|
||||
## 11. PART 4.2 — THE ABANDONMENT COUNTDOWN, WATCHED FIRING
|
||||
|
||||
**The terminal deletion has now been observed.** It did exactly what it claims, including the half
|
||||
nobody had seen.
|
||||
|
||||
- **State created** 23:08–23:09 CEST. A realistic set-aside store was built at
|
||||
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, the countdown written overdue
|
||||
(started 2026-08-07, due 2026-08-20), and the product's own CLI confirmed it:
|
||||
`abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0`.
|
||||
- **05:10 CEST — it fired.** `/home` on the storage box is stamped `03:10Z`; the set-aside store is
|
||||
**gone**; **`/home/felhom-repo`, the live repository, is untouched** (mtime still 4 Aug).
|
||||
- **05:13 CEST — the hub half.** Event **3025 `offsite_abandon_purged`**: *„Az ügyfél korábbi távoli
|
||||
mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is törölve (1 csomag). Az ügyfél
|
||||
döntése alapján, a 14 napos türelmi idő lejárta után."* And demo-hp's superseded escrow row **is
|
||||
gone from the hub** — it held one before.
|
||||
- **The controller closed itself out.** Every `abandon_*` field has been removed from `settings.json`
|
||||
and `--abandon-status` reports *"no abandonment countdown is running on this box"*.
|
||||
|
||||
**So the two-phase commit's central promise — *"it removes BOTH halves or neither"* — is confirmed
|
||||
live for the first time:** the ciphertext and the sealed package that protects it went together, three
|
||||
minutes apart, and the state that remembered the operation cleaned itself up. **What it removes:**
|
||||
exactly the recorded set-aside path. **What survives:** the live repository, the live escrow, and the
|
||||
box's current recovery path.
|
||||
|
||||
**Judged:** the completion message is accurate and in plain Hungarian. The only wrong note is while
|
||||
the countdown is *overdue but not yet swept* — the card then states a past date in the future tense
|
||||
(**R-365**).
|
||||
|
||||
**No countdown is left running anywhere.** `demo-hp` cleared itself; `demo-felhom` reports none.
|
||||
|
||||
---
|
||||
|
||||
## 11b. WHAT THE NIGHT FOUND THAT NOBODY WAS LOOKING FOR
|
||||
|
||||
At 21:59 the box filed, unprompted, hub event 3016:
|
||||
|
||||
> `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could
|
||||
> not be restored+booted … wrong key — manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided
|
||||
> key dd:d1:d8:53:44:62:5e:0b`
|
||||
|
||||
That archive predates the 21 August reinstall by three days. **The rebuild orphaned the PBS whole-guest
|
||||
archives exactly as it orphaned the restic repository** — so a rebuilt box loses *both* off-premises
|
||||
tiers at once. The restore-test mechanism deserves credit: it caught it and named the key mismatch
|
||||
precisely. The gap is what it is *called* — "a restore test failed" reads as a flaky verification, not
|
||||
as "every whole-guest backup taken before the reinstall is unreadable on this machine". Filed **R-366
|
||||
(HIGH)**.
|
||||
|
||||
---
|
||||
|
||||
## 12. MACHINE STATES AT THE END
|
||||
|
||||
**`demo-hp` — HEALTHY, not broken. Nothing needs bringing back.** At 05:28 CEST: all 15 containers up
|
||||
and healthy (`felhom-controller` 0.217.0, `traefik`, `cloudflared`, `filebrowser`, plus the six drill
|
||||
apps); `/` 4%, `/var/lib/felhom` 13%, the data drive 1%; **no filler files left**, **no app-stop marker**,
|
||||
**no operation in flight**, **no countdown running**, and `restic check` over its off-site repository
|
||||
reports **no errors were found**. Its off-site backup, silent since 9 August, is working again and ran
|
||||
on its own schedule at 04:15.
|
||||
|
||||
**Deliberately left in place, each with a reason** (retained, not forgotten):
|
||||
|
||||
| left behind | why |
|
||||
|---|---|
|
||||
| `opengist`, `privatebin`, `calibre-web` deployed with the planted `DRILL-2026-08-21` fixture | the reproduction for **R-354** and **R-356**; the hashes in §2 make the fix verifiable without rebuilding the case |
|
||||
| `paperless-ngx` deployed | the **only** reproduction of **R-355**, and it regenerates the orphan directory on every nightly cycle |
|
||||
| `romm` deployed | the only app on the box where the safety dump actually works — the fixture for **R-361** and the control for R-355 |
|
||||
| `kimai` deployed | a second correct-derivation DB app, the negative control for R-355 |
|
||||
| five verification copies under `backups/offsite-restore/` | harmless, and they are the **R-358** fixture |
|
||||
| `romm`'s `pre-restore-…sql` in its unit | the physical evidence for **R-361** |
|
||||
| DooPlex's public key in `demo-hp:/root/.ssh/authorized_keys`, and the `hp` alias in `~/.ssh/config` | the box had lost the key and its tailnet address is dead; without it the next session must go through break-glass again |
|
||||
|
||||
**Removed / restored during the drill:** both corrupted restic packs (repo verified clean), both
|
||||
`fallocate` fillers, the set-aside store (by the sweep, as intended), and demo-hp's superseded escrow
|
||||
row (by the sweep's hub half, as intended).
|
||||
|
||||
**`demo-hp`'s tailnet address `100.76.96.79` is still unreachable** and was not repaired — the box is
|
||||
reachable on the LAN at `192.168.0.104`. That is the one thing about it that is worse than it should
|
||||
be, and it predates tonight.
|
||||
|
||||
**`demo-felhom` — untouched and healthy**, controller 0.217.0, its own off-site ran at 02:15 with
|
||||
`last_status: ok`, both fixtures verified present (§10), no countdown running.
|
||||
|
||||
**`drill-r50`** — not used. Still `DOWN` in the hub, agent 0.129.0, as it was.
|
||||
|
||||
---
|
||||
|
||||
## 13. TEARDOWN — ALL FOUR LAYERS, STATED
|
||||
|
||||
| layer | state |
|
||||
|---|---|
|
||||
| **Off-site (storage box)** | The set-aside store I created was **deleted by the product's own sweep**, as designed. The two corrupted packs were **restored byte-identical** and `restic check` passes. Nothing else was written. `restic forget`/`prune` were never invoked by me. |
|
||||
| **PVE host `demo-hp`** | `/root/settings.json.drill-backup` **retained** (the pre-drill controller settings, in case the abandonment edit needs reverting). DooPlex's SSH key **retained**, with the reason above. `/tmp/plant.tar` left; harmless. |
|
||||
| **Guest 9201 / controller** | Six apps and the planted fixtures **retained with reasons** (table above). Controller state is clean: no marker, no countdown, no in-flight op. |
|
||||
| **Hub — stated explicitly** | **No customer record and no host record was created, so none needs deleting.** I used the existing `demo-hp` customer and host throughout. The only hub-side *removal* was demo-hp's superseded escrow row, done by the abandonment sweep itself and reported as event 3025. Events 3002–3025 were generated as a normal consequence of the work and are left as the record. **Nothing is owed here and nothing is blocked.** |
|
||||
| **DooPlex (this machine)** | The hub DB copies (which contain every host's break-glass secret) are in the session scratchpad only, never in a committed file, and are shredded in the closing step. |
|
||||
|
||||
---
|
||||
|
||||
## 14. OBSERVATIONS — noticed, not acted on
|
||||
|
||||
- **The hub does not display the controller version it is told.** `demo-hp`'s host page shows the
|
||||
guest's Controller column as **„—"** hours after receiving `controller_updated: 0.216.0 → 0.217.0`.
|
||||
Two hub surfaces, one blind. Not filed — `REPORT-hub-blindness.md` already exists and this may be
|
||||
part of it.
|
||||
- **`demo-felhom` moved to 0.217.0 without a `controller_updated` event** (started 19:31Z, 17 minutes
|
||||
before the floor was saved). Consistent with a by-hand deployment during the golden bake, not a
|
||||
defect — recorded so a later reader does not mistake it for one.
|
||||
- **`restic --latest N` is per-group, not a total.** It briefly read as "18 snapshots became 10" and
|
||||
would have been reported as data loss. Caught by listing all of them and by an independent
|
||||
`restic stats` (26.44 MiB / 588 blobs). **An unpaginated listing is not a total** — the rule earned
|
||||
its place again.
|
||||
- **`find -newermt` is the wrong probe for a restore**: restic preserves the snapshot's mtimes, so a
|
||||
freshly restored tree looks old. It briefly read as "nothing was restored".
|
||||
- **`sftp -b` aborts on the first failing line**, so a batch listing 256 shard directories returned 7
|
||||
packs and looked like a total. Prefixing each line with `-` fixed it. Same family as the two above.
|
||||
|
||||
---
|
||||
|
||||
## 15. WHAT WAS DROPPED, PLAINLY
|
||||
|
||||
- **Nothing in Parts 0, 2, 3, 4 or 5 was skipped.** Every Part 4 item was attempted and reached a
|
||||
verdict.
|
||||
- **The fill-watch alarm was not watched firing on its own schedule.** I proved its cadence from
|
||||
source (daily 03:30) and observed it fire from the startup path (`disk_critical`, 21:10:58Z, correct
|
||||
Hungarian copy naming the drive and the free space). Holding a filesystem at 99% for six hours would
|
||||
have sat across the scheduled backup and the abandonment sweep, and I judged the scheduled cycle —
|
||||
which the task asks for explicitly — worth more than a second sighting of an alarm whose trigger I
|
||||
had already read and seen work.
|
||||
- **I did not drive the abandonment through the real orphan→reset path**, because that path
|
||||
move-asides the *live* repository, which would have destroyed demo-hp's whole off-site history
|
||||
including the 9 August snapshots this drill exists to read. The sweep's own code path was exercised
|
||||
in full on a real store; the deviation is only in how the state was created.
|
||||
- **`peti-felhom` was not contacted**, per the fence.
|
||||
@@ -653,8 +653,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** Two findings, one session, both shipped. **(a) The blindness.** Every recovery-unit `manifest.json` has carried `drive` and `namespace_root` since schema 1 (`controller/internal/backup/recovery_unit.go:48-49`), written at capture from the app's own placement. `grep -rE '\.Drive\b\|\.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | **CLOSED 2026-08-21** — controller | — | **Shipped:** `backup/offbox_placement.go` (`CheckPlacement`, `PlacementMismatchMessage`, `RecordedUnitForStack`); mismatch **named and refused** before the safety dump, with `ack_placement` as a **separate** field from `confirm=1`; the not-installed refusal names the recorded drive; deploy page **prefills the address and folder from the app's own backup**; `Server.restoreOpBlocked()` reads BOTH flags; `RestoreOpStatus.LastRecent` + `RestoreResultWindow` moved to `internal/backup` as ONE expression for two surfaces. Red-proofs: B with **both** guards removed **was seen starting a restore with no drive attached** (no error, full 3.00 s run into `/tmp/mutant-destination`); C, E and A each returned their wrong outcome; D forced on broke 8 ordinary reconstitute tests, proving reachability both ways. | CC |
|
||||
| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
|
||||
| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC |
|
||||
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **OPEN — HIGH** | — | Restore the unit's `volume-dumps/*.tar` from the SCRATCH unit on the off-site path, as `RestoreFromRecoveryUnit` already does from the live unit. Assert the CONSEQUENCE (planted bytes come back), not the mechanism. | CC |
|
||||
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **OPEN — HIGH** | — | Two candidates, both two-repo: rename the container to `paperless-ngx-postgres`, or make an unresolved candidate a loud skip rather than a silent fallback. The second is the one that generalises. Pin with a test that a container whose name resolves to no known stack is never dumped silently. | CC |
|
||||
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | `restoreDockerVolumesFrom` is the LOCAL path's own replay with an explicit directory — ONE implementation, two callers, because a second copy of that loop is what produced the divergence. It reads the SCRATCH unit; the live unit is still never written. Volumes replay BEFORE the database (a logical dump must still win over a volume-tar copy of the same database) and inside the stopped window (Docker will not replace a volume a container holds). `VolumesReplayed` is on the result and in the sentence. **The comment beside the skip was half false and is corrected, not left:** it justified the skip by saying the dump is replayed from the scratch "so nothing is lost" — true of the database, false of the volumes. The half that still holds (the live unit is the local path's source) is named. **PROVEN LIVE on `demo-hp`, negative control first.** On 0.217.0 with the fixture planted, hashed and then deleted from the live volume: „A(z) calibre-web: **0 fájl visszaállítva** … — az alkalmazás újraindult.", `ok=true`, and `ls` reported the directory absent. On 0.218.0, same fixture, same steps: „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** …" and **5/5 files byte-identical**, both Hungarian accented filenames included. **Four red-proofs, each mutation asserted applied:** the volume leg removed returned the silent loss; the count dropped from the message returned the true-but-incomplete sentence verbatim; the unit guard removed was SEEN writing into the live unit; the error swallowed let a partial replay report success. **The unit-guard proof initially PASSED against the mutation** — the fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit root (the R-181 class, in the test rather than the code). Widened, and it convicts. **NOT reached by this fix, and it is the blocker:** the 40 apps that declare no data drive still cannot run this restore at all — **R-356** refuses first. | CC |
|
||||
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | **Neither candidate: a third that removes the guessing.** Every container the controller starts carries `com.docker.compose.project`, and that label IS the stack name BY CONSTRUCTION — compose is run with `cmd.Dir` set to `/opt/docker/stacks/<stack>` and never `-p` (`stacks/manager.go:1218`). `resolveStackName` prefers it whenever it names a deployed stack; `deriveStackName` stays as the fallback for containers not started by compose; an attribution that resolves to NO known stack is now **loud** instead of silently returned. **The fix is in the CONTROLLER, not the catalogue** — renaming the container would have fixed this one app and left the guessing for the next. **The sweep was proven before its answer was trusted:** a second mismatch planted in a scratch copy (`kimai-db`→`timetrack-db`) was convicted by name, removal returned it to 1, and a catalogue with every mismatch removed exits **0** — so "1" is not a stuck value. **1 affected app of 53**, 15 DB containers checked. **PROVEN LIVE on `demo-hp`, negative control first.** 0.217.0 at 09:37: unit `db_dumps = None`, dump refreshed into the phantom `…/primary/paperless/db-dumps/paperless-postgres.sql`. 0.218.0 at 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` **inside the app's own unit**, and present in the off-site snapshot for the first time. Scenario B: the destructive restore said „**0 fájl és 3 adatkötet és az adatbázis visszaállítva**" and wrote `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql` (312 957 B) where **zero** undo copies had existed. Scenario C: with the undo made impossible the restore REFUSED, the marker kept its mutation and the container's `StartedAt` was unchanged — **the app was never stopped**. Scenario D: `romm-mariadb.sql` and `kimai-mariadb.sql` unchanged in name and location. **Three red-proofs, each asserted applied:** the name fix reverted printed both divergent paths; the message predicate reverted returned the false „nincs adatbázisa" sentence verbatim; the refusal removed was seen letting a restore proceed with no undo. **The dumps already written under the wrong name are NOT deleted** — see **R-367**. | CC |
|
||||
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. | CC |
|
||||
| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC |
|
||||
| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC |
|
||||
@@ -666,6 +666,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
|
||||
| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC |
|
||||
| **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
|
||||
| **R-367** | **The database dumps already written under the wrong name are stranded, and nothing will ever collect them.** R-355's fix sends `paperless-ngx`'s dump to the right place from now on; it does not move the ones already written. On `demo-hp` that is `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql` (312 381 B, 2026-08-22 07:38, the last pre-fix cycle). **Nothing deletes them and that is by design, not by luck:** the F5 stale-primary prune (`backup.go:1248`) skips any directory whose name is not a deployed app, under the guard *"an undeployed app's last backup is still its restore point"* — verified still present after the fix. They are equally invisible to the off-site push, which resolves paths from the app's own unit. **They CAN be adopted, by hand:** move the file to `…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point for that app. **It is deliberately not automatic.** The adopted dump would sit beside volume tars taken at a different time, i.e. an INCOHERENT pair — the exact shape R-43/R-44's coherence stamp exists to make visible — and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. **Filed rather than done**, because whether a stale orphan is worth adopting at all is a judgement about one machine's history, not a rule. | **OPEN — LOW** | follows R-355 | Decide per box: adopt (and say the pair is skewed), or delete deliberately. Neither on an upgrade path. | Viktor rules, CC executes |
|
||||
|
||||
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC |
|
||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
# Golden bake — 0.218.0 (2026-08-22)
|
||||
|
||||
Baked in the drill VM on DooPlex per `documentation/runbooks/RUNBOOK-manual-build.md` §4.0/§4.1,
|
||||
carrying controller **v0.218.0** (R-354 + R-355).
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `GOLDEN_VERSION` | **0.218.0** |
|
||||
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
|
||||
| package | `https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.218.0/golden.tar.zst` |
|
||||
| size | 657 026 013 B (archive 626 MB) |
|
||||
| controller image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` |
|
||||
| template | `debian-13-standard_13.6-1_amd64.tar.zst` (listed fresh, not assumed) |
|
||||
| `MinAgent` | **0.129.0** (from the controller CHANGELOG header — unchanged) |
|
||||
|
||||
## Pass markers — each checked, with the negative controls
|
||||
|
||||
```
|
||||
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
|
||||
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
|
||||
upload OK (HTTP 201) : 1 -> pre-delete returned HTTP 404 (404/204 expected)
|
||||
excluding : 0 <- negative control
|
||||
FATAL : 0 <- negative control
|
||||
```
|
||||
|
||||
## Token hygiene
|
||||
|
||||
The Gitea token was copied **file → file** (`scp`) and read by a runner script *inside* the VM, so it
|
||||
never crossed a shell or a unit property. Verified: `systemctl show golden-bake -p Environment
|
||||
-p ExecStart | grep -c -F "$(cat /root/.gitea-token)"` → **0**.
|
||||
|
||||
**The leak grep on this committed log was itself proven before its zero was believed:** the token was
|
||||
appended to a throwaway copy, grepped (**1**), the copy shredded, and only then was the committed
|
||||
log's **0** accepted.
|
||||
|
||||
## Teardown
|
||||
|
||||
`pct destroy 9100 --purge`; token, runner, build script and in-VM log `shred -u`'d **after** this log
|
||||
was copied out; VM powered off; qemu exited; `qemu-img snapshot -a virgin` restored.
|
||||
|
||||
## NOT vouched
|
||||
|
||||
The hub's Day-0 artifact manifest was **not** changed and the floor was **not** raised — both are the
|
||||
operator's decision. See `REPORT.md` §12.
|
||||
@@ -0,0 +1,323 @@
|
||||
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.218.0
|
||||
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
|
||||
Logical volume "vm-9100-disk-0" created.
|
||||
Logical volume pve/vm-9100-disk-0 changed.
|
||||
Creating filesystem with 8388608 4k blocks and 2097152 inodes
|
||||
Filesystem UUID: 3334dd11-f031-40d9-902c-cd36f5df48cd
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
4096000, 7962624
|
||||
Logical volume "vm-9100-disk-1" created.
|
||||
Logical volume pve/vm-9100-disk-1 changed.
|
||||
Creating filesystem with 6291456 4k blocks and 1572864 inodes
|
||||
Filesystem UUID: 4c45eb14-e05e-42d1-bc5b-a18927aec39a
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
|
||||
Total bytes read: 553512960 (528MiB, 102MiB/s)
|
||||
Detected container architecture: amd64
|
||||
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
|
||||
done: SHA256:NsUGGcy+RV+rkrR8mWKY2Oihey1ytt92IvzC1uZ8SYE root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
|
||||
done: SHA256:UR39CmAMA7Bt4ljKletHmhKTs32svvU+xnpwCoxwG6U root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
|
||||
done: SHA256:kYjjxRfjLMHwSrMHH7h7QEKGXzMwDy1KtD/s/TtaDww root@felhom-golden
|
||||
[golden] starting + installing Docker (official repo, trixie channel) …
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
|
||||
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
|
||||
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
|
||||
Unable to find image 'hello-world:latest' locally
|
||||
latest: Pulling from library/hello-world
|
||||
4f55086f7dd0: Pulling fs layer
|
||||
4f55086f7dd0: Verifying Checksum
|
||||
4f55086f7dd0: Download complete
|
||||
4f55086f7dd0: Pull complete
|
||||
Digest: sha256:5dd0d3e6e255913fc30f90b9f2b1d359cc2cbdb48090cc4b65f1676e203243cc
|
||||
Status: Downloaded newer image for hello-world:latest
|
||||
docker OK (overlay2; data-root /var/lib/docker)
|
||||
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
|
||||
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
|
||||
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
|
||||
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.218.0 (no registry cred at deploy) …
|
||||
|
||||
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
|
||||
Configure a credential helper to remove this warning. See
|
||||
https://docs.docker.com/go/credential-store/
|
||||
|
||||
0.218.0: Pulling from admin/felhom-controller
|
||||
039e6f9f9752: Pulling fs layer
|
||||
0094c3ac0914: Pulling fs layer
|
||||
deca1dac7403: Pulling fs layer
|
||||
11c19a33d1b8: Pulling fs layer
|
||||
3ba1be8b4b39: Pulling fs layer
|
||||
2c80200460b0: Pulling fs layer
|
||||
11c19a33d1b8: Waiting
|
||||
3ba1be8b4b39: Waiting
|
||||
2c80200460b0: Waiting
|
||||
deca1dac7403: Verifying Checksum
|
||||
deca1dac7403: Download complete
|
||||
11c19a33d1b8: Verifying Checksum
|
||||
11c19a33d1b8: Download complete
|
||||
3ba1be8b4b39: Verifying Checksum
|
||||
3ba1be8b4b39: Download complete
|
||||
2c80200460b0: Verifying Checksum
|
||||
2c80200460b0: Download complete
|
||||
0094c3ac0914: Verifying Checksum
|
||||
0094c3ac0914: Download complete
|
||||
039e6f9f9752: Download complete
|
||||
039e6f9f9752: Pull complete
|
||||
0094c3ac0914: Pull complete
|
||||
deca1dac7403: Pull complete
|
||||
11c19a33d1b8: Pull complete
|
||||
3ba1be8b4b39: Pull complete
|
||||
2c80200460b0: Pull complete
|
||||
Digest: sha256:56c35264346e4f4e4eca0f3405d2ca727f29a7b4d1e73fbffbaf1f0835bd035e
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.218.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.218.0
|
||||
[golden] asking the controller which infra images it manages …
|
||||
[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
|
||||
v3.6.7: Pulling from library/traefik
|
||||
589002ba0eae: Pulling fs layer
|
||||
ef63511ea6cc: Pulling fs layer
|
||||
0738e5cb835e: Pulling fs layer
|
||||
3e6813f70c64: Pulling fs layer
|
||||
3e6813f70c64: Waiting
|
||||
589002ba0eae: Verifying Checksum
|
||||
589002ba0eae: Download complete
|
||||
ef63511ea6cc: Verifying Checksum
|
||||
ef63511ea6cc: Download complete
|
||||
3e6813f70c64: Verifying Checksum
|
||||
3e6813f70c64: Download complete
|
||||
0738e5cb835e: Verifying Checksum
|
||||
0738e5cb835e: Download complete
|
||||
589002ba0eae: Pull complete
|
||||
ef63511ea6cc: Pull complete
|
||||
0738e5cb835e: Pull complete
|
||||
3e6813f70c64: Pull complete
|
||||
Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
|
||||
Status: Downloaded newer image for traefik:v3.6.7
|
||||
docker.io/library/traefik:v3.6.7
|
||||
2026.6.0: Pulling from cloudflare/cloudflared
|
||||
47de5dd0b812: Pulling fs layer
|
||||
c172f21841df: Pulling fs layer
|
||||
99515e7b4d35: Pulling fs layer
|
||||
99ba982a9142: Pulling fs layer
|
||||
d6b1b89eccac: Pulling fs layer
|
||||
2780920e5dbf: Pulling fs layer
|
||||
7c12895b777b: Pulling fs layer
|
||||
3214acf345c0: Pulling fs layer
|
||||
52630fc75a18: Pulling fs layer
|
||||
dd64bf2dd177: Pulling fs layer
|
||||
b839dfae01f6: Pulling fs layer
|
||||
ebddc55facdc: Pulling fs layer
|
||||
bdfd7f7e5bf6: Pulling fs layer
|
||||
2d4d7adf6272: Pulling fs layer
|
||||
40008157d8d2: Pulling fs layer
|
||||
bd8962e29291: Pulling fs layer
|
||||
cac2ae0193cb: Pulling fs layer
|
||||
74d1dac84ecc: Pulling fs layer
|
||||
99ba982a9142: Waiting
|
||||
b839dfae01f6: Waiting
|
||||
ebddc55facdc: Waiting
|
||||
bdfd7f7e5bf6: Waiting
|
||||
d6b1b89eccac: Waiting
|
||||
2780920e5dbf: Waiting
|
||||
7c12895b777b: Waiting
|
||||
3214acf345c0: Waiting
|
||||
2d4d7adf6272: Waiting
|
||||
52630fc75a18: Waiting
|
||||
40008157d8d2: Waiting
|
||||
bd8962e29291: Waiting
|
||||
dd64bf2dd177: Waiting
|
||||
cac2ae0193cb: Waiting
|
||||
74d1dac84ecc: Waiting
|
||||
47de5dd0b812: Download complete
|
||||
c172f21841df: Verifying Checksum
|
||||
c172f21841df: Download complete
|
||||
d6b1b89eccac: Verifying Checksum
|
||||
d6b1b89eccac: Download complete
|
||||
2780920e5dbf: Verifying Checksum
|
||||
2780920e5dbf: Download complete
|
||||
99ba982a9142: Verifying Checksum
|
||||
99ba982a9142: Download complete
|
||||
47de5dd0b812: Pull complete
|
||||
7c12895b777b: Verifying Checksum
|
||||
7c12895b777b: Download complete
|
||||
3214acf345c0: Verifying Checksum
|
||||
3214acf345c0: Download complete
|
||||
52630fc75a18: Verifying Checksum
|
||||
52630fc75a18: Download complete
|
||||
c172f21841df: Pull complete
|
||||
dd64bf2dd177: Verifying Checksum
|
||||
dd64bf2dd177: Download complete
|
||||
99515e7b4d35: Verifying Checksum
|
||||
99515e7b4d35: Download complete
|
||||
b839dfae01f6: Download complete
|
||||
ebddc55facdc: Verifying Checksum
|
||||
ebddc55facdc: Download complete
|
||||
bdfd7f7e5bf6: Verifying Checksum
|
||||
bdfd7f7e5bf6: Download complete
|
||||
bd8962e29291: Verifying Checksum
|
||||
bd8962e29291: Download complete
|
||||
cac2ae0193cb: Verifying Checksum
|
||||
cac2ae0193cb: Download complete
|
||||
99515e7b4d35: Pull complete
|
||||
2d4d7adf6272: Download complete
|
||||
99ba982a9142: Pull complete
|
||||
74d1dac84ecc: Verifying Checksum
|
||||
74d1dac84ecc: Download complete
|
||||
40008157d8d2: Verifying Checksum
|
||||
40008157d8d2: Download complete
|
||||
d6b1b89eccac: Pull complete
|
||||
2780920e5dbf: Pull complete
|
||||
7c12895b777b: Pull complete
|
||||
3214acf345c0: Pull complete
|
||||
52630fc75a18: Pull complete
|
||||
dd64bf2dd177: Pull complete
|
||||
b839dfae01f6: Pull complete
|
||||
ebddc55facdc: Pull complete
|
||||
bdfd7f7e5bf6: Pull complete
|
||||
2d4d7adf6272: Pull complete
|
||||
40008157d8d2: Pull complete
|
||||
bd8962e29291: Pull complete
|
||||
cac2ae0193cb: Pull complete
|
||||
74d1dac84ecc: Pull complete
|
||||
Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
|
||||
Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0
|
||||
docker.io/cloudflare/cloudflared:2026.6.0
|
||||
1.3.3-stable: Pulling from gtstef/filebrowser
|
||||
6a0ac1617861: Pulling fs layer
|
||||
ef8806083e82: Pulling fs layer
|
||||
b74107c861c7: Pulling fs layer
|
||||
adc935def003: Pulling fs layer
|
||||
4f4fb700ef54: Pulling fs layer
|
||||
18695ccc900a: Pulling fs layer
|
||||
45d119d5c397: Pulling fs layer
|
||||
dac52db4fc51: Pulling fs layer
|
||||
6d598f86b2f2: Pulling fs layer
|
||||
8aa349c8396c: Pulling fs layer
|
||||
18695ccc900a: Waiting
|
||||
45d119d5c397: Waiting
|
||||
dac52db4fc51: Waiting
|
||||
6d598f86b2f2: Waiting
|
||||
8aa349c8396c: Waiting
|
||||
adc935def003: Waiting
|
||||
4f4fb700ef54: Waiting
|
||||
6a0ac1617861: Download complete
|
||||
b74107c861c7: Verifying Checksum
|
||||
b74107c861c7: Download complete
|
||||
adc935def003: Download complete
|
||||
4f4fb700ef54: Verifying Checksum
|
||||
4f4fb700ef54: Download complete
|
||||
45d119d5c397: Verifying Checksum
|
||||
45d119d5c397: Download complete
|
||||
dac52db4fc51: Verifying Checksum
|
||||
dac52db4fc51: Download complete
|
||||
ef8806083e82: Verifying Checksum
|
||||
ef8806083e82: Download complete
|
||||
6d598f86b2f2: Verifying Checksum
|
||||
6d598f86b2f2: Download complete
|
||||
18695ccc900a: Verifying Checksum
|
||||
18695ccc900a: Download complete
|
||||
8aa349c8396c: Verifying Checksum
|
||||
8aa349c8396c: Download complete
|
||||
6a0ac1617861: Pull complete
|
||||
ef8806083e82: Pull complete
|
||||
b74107c861c7: Pull complete
|
||||
adc935def003: Pull complete
|
||||
4f4fb700ef54: Pull complete
|
||||
18695ccc900a: Pull complete
|
||||
45d119d5c397: Pull complete
|
||||
dac52db4fc51: Pull complete
|
||||
6d598f86b2f2: Pull complete
|
||||
8aa349c8396c: Pull complete
|
||||
Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
|
||||
Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable
|
||||
docker.io/gtstef/filebrowser:1.3.3-stable
|
||||
1.1.0: Pulling from admin/felhom-samba
|
||||
897d797d2723: Pulling fs layer
|
||||
3051591aa250: Pulling fs layer
|
||||
ce57a3f93416: Pulling fs layer
|
||||
fb94eeec2fe1: Pulling fs layer
|
||||
fb94eeec2fe1: Waiting
|
||||
ce57a3f93416: Verifying Checksum
|
||||
ce57a3f93416: Download complete
|
||||
fb94eeec2fe1: Verifying Checksum
|
||||
fb94eeec2fe1: Download complete
|
||||
897d797d2723: Verifying Checksum
|
||||
897d797d2723: Download complete
|
||||
3051591aa250: Verifying Checksum
|
||||
3051591aa250: Download complete
|
||||
897d797d2723: Pull complete
|
||||
3051591aa250: Pull complete
|
||||
ce57a3f93416: Pull complete
|
||||
fb94eeec2fe1: Pull complete
|
||||
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
|
||||
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
|
||||
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
|
||||
[golden] identity-clean + minimize …
|
||||
[golden] stop + archive …
|
||||
INFO: including mount point rootfs ('/') in backup
|
||||
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||
INFO: archive file size: 626MB
|
||||
INFO: Finished Backup of VM 9100 (00:00:31)
|
||||
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_08_22-10_07_08.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
|
||||
[golden] publishing golden (657026013 bytes, sha256 8e427869d13eafb7…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.218.0/golden.tar.zst
|
||||
[golden] pre-delete existing: HTTP 404 (404/204 expected)
|
||||
[golden] upload OK (HTTP 201)
|
||||
GOLDEN_VERSION=0.218.0
|
||||
GOLDEN_SHA256=8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b
|
||||
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.218.0 / 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b
|
||||
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
|
||||
Reference in New Issue
Block a user