diff --git a/REPORT.md b/REPORT.md index cccfdaba..69436100 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,190 +1,451 @@ -# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday) +# REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night) -**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade -anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change, -which is the predicted result.** Two dated checks filed as **R-341**. - -**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only. -Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`. +**Unattended diagnostic drill on `demo-hp`. No code changed in any repository. No version bumped, no +image built, nothing deployed.** Evidence: +`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`. --- -## 1. Baselines re-confirmed on the machine +## 1. THE VERDICT -Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook. +**It is a mixture, and the drill's three options are all present — but they belong to different +faults, and only one of them explains the thing you actually saw.** + +### What explains YOUR observation (OpenGist, 2026-08-21 afternoon): **THE BACKUP IS EMPTY**, and then **THE MESSAGE LIED** + +Reproduced independently tonight, and it agrees with what R-353 already recorded: + +1. The off-site restore for OpenGist **never ran**. It refused, because OpenGist declares no data + drive, and the refusal says *„a(z) opengist nincs telepítve"* — **"OpenGist is not installed"** — + about an app that was installed, deployed, running and healthy. +2. The person therefore used the **local** restore-from-unit. The local unit on the freshly rebuilt + box was **genuinely empty of data** — no dump cycle had run yet on a one-hour-old machine — so it + held `compose/` and nothing else. +3. The restore returned that configuration and reported a bare completion. + +So for that specific event the data was not in the thing that was restored. **R-353 called this +correctly and I did not find an error in it.** I nearly filed a correction against it and was wrong +to think so; its text is more careful than the CHANGELOG's summary of it. + +### What the drill found that nobody had seen: **THE RESTORE LOSES IT** + +This is new, it is worse, and it is not the same fault: + +**When the off-site snapshot DOES hold the data, the off-site full restore still does not return it.** +The off-site restore has a files leg and a database leg. **It has no named-volume leg at all.** + +Proven live on `calibre-web` at 22:23:36 with planted files: + +| leg | in the unit | in the off-site snapshot | in the checking folder | returned by the off-site restore | +|---|---|---|---|---| +| declared user files (`media/books`) | n/a | yes, 5/5 byte-identical | yes, 5/5 byte-identical | **yes, 5/5 byte-identical** | +| named volume `calibre_web_config` | yes, 1 422 848 B | yes | yes | **NO — silently skipped** | + +and the customer was told: + +> „A(z) calibre-web: **5 fájl visszaállítva** (mentés: 2026-08-21 22:16) — az alkalmazás újraindult. +> Ennek az alkalmazásnak nincs adatbázisa." + +Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence +does not mention it. **For `calibre-web` the lost leg is the app's settings. For the 40 catalogue +apps that declare no data drive, that leg is the entire dataset.** + +**Why:** `ReconstituteFromOffsite` skips every placement flagged `isUnit` +(`controller/internal/backup/offbox_reconstitute.go:341-346`), and the volume tars live *inside* the +unit. `grep` for a volume-restore call in the whole off-site path returns nothing. The **local** +restore does have one (`restore.go:99 restoreDockerVolumes`) — proven tonight by restoring +PrivateBin's planted 1 MB from its volume tar, byte-identical. **Two code paths, the same tar, one of +them replays it.** + +### And a third, separate: **THE BACKUP IS EMPTY** — really empty — for `paperless-ngx`'s database + +`paperless-ngx` runs a 72-table PostgreSQL. Its recovery unit records **`db_dumps: null`**. It always +has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to +`/mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/`, a directory named after a stack +that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the +safety-dump code filters on the same wrong name, **the destructive restore takes no undo at all** and +then says: + +> „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" + +The controller had dumped that database five minutes earlier. + +**One app in 53 is affected** (catalogue-wide sweep in §6). It is the document archive. + +--- + +## 2. THE FULL/EMPTY CONTRAST + +**Apps chosen, and why.** From the two storage classes: **`calibre-web`** declares a data drive +(`needs_hdd: true`, `backup.userdata: media/books class: mandatory`) and **`privatebin`** / +**`opengist`** declare none (the 40-class; data lives in a Docker named volume). I ran **three** cases +rather than two, deliberately: FULL and EMPTY alone cannot separate *"it was empty"* from *"the class +is broken"*, so `privatebin` was run FULL as the disambiguator. + +**The fixture** — 5 files, two with Hungarian accented names, recorded as raw bytes: + +``` +SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991 +binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a +nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4 +plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1 +árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea + +name bytes (UTF-8 NFC): + árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74 + őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64 +``` + +**The comparator was proved able to convict before it was trusted.** One byte flipped at offset +500 000 of `binary-1mb.bin` (`af` → `00`): `sha256sum -c` reported `binary-1mb.bin: FAILED`, rc=1, +while the other four passed; the unmodified set passed rc=0. The mutant was discarded. + +### The results + +| case | class | data leg | in unit | in off-site snapshot | checking folder | off-site restore | local restore | +|---|---|---|---|---|---|---|---| +| **calibre-web FULL** | drive | user files | n/a | 5/5 identical | 5/5 identical | **5/5 identical** | — | +| **calibre-web FULL** | drive | named volume 1.42 MB | yes | yes | yes | **NOT restored** | — | +| **privatebin FULL** | no drive | named volume 1.06 MB | yes | 5/5 identical | 5/5 identical | **REFUSED — false reason** | **5/5 identical** | +| **opengist EMPTY** | no drive | named volume 181 KB skeleton | yes | yes | yes | **REFUSED — false reason** | — | + +**The contrast decides it.** The unit and the off-site snapshot **hold the data, verified by +identity**, in both classes, accented filenames included. So the capture is sound. The failure is +entirely in the last leg. + +**Both messages, verbatim:** + +- FULL, drive class → `ok=true`, *„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16) + — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."* +- FULL and EMPTY, no-drive class → `ok=false`, *„a(z) privatebin nincs telepítve, ezért nincs hová + visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra + az alkalmazást (Alkalmazások) ugyanerre a helyre…"* + +The second is the important one. **FULL and EMPTY got the identical sentence**, so the message cannot +distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is +installed, and the instruction ("reinstall it to the same place") **cannot be followed**, because a +40-class app is offered no storage field at deploy time (that is R-352's own measurement). + +**This closes R-353's second instruction.** It asked for proof that a named-volume app reaches the +off-site tier rather than inference from gate order. It does. Manifests read tonight: + +``` +privatebin volume_dumps = ['privatebin_privatebin_data.tar'] db_dumps = None +opengist volume_dumps = ['opengist_opengist_data.tar'] db_dumps = None +kimai volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar'] + db_dumps = ['kimai-mariadb.sql'] +calibre-web volume_dumps = ['calibre-web_calibre_web_config.tar'] db_dumps = None +paperless-ngx volume_dumps = [3 tars] db_dumps = None ← the defect +``` + +--- + +## 3. PART 0 — did the floor move the machine? + +**Yes, unaided, in 17 seconds.** | | | |---|---| -| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** | -| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) | -| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident | -| Effective `open files` | **65536 / 65536** — the morning's drop-in in force | -| `Recv-Q` / loopback | **0** / **`200` in 11 ms** | -| Datastore | 3.7 G of 98 G, 4% | -| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline | +| moved from → to | **0.216.0 → 0.217.0** | +| who initiated | **the hub** — the operator saved the floor at `19:48:37Z`; a poke reached the agent from `10.77.0.1:58093` at `19:48:32Z` for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision. | +| how long | floor saved `19:48:37Z` → *"controller-swap: new controller healthy"* `19:48:54Z` = **17 s**. Swap requested `19:48:41Z` → healthy = 13 s. | +| **the container's own tag on the box** | **`gitea.dooplex.hu/admin/felhom-controller:0.217.0`**, created `2026-08-21 19:48:45 UTC` — read from `docker ps` in guest 9201, not from the hub. | -## 2. The before-slope — and a correction I owe the morning's report +The hub's own host page for `demo-hp` shows the guest's **Controller column as „—"** — the hub does +not know which controller version the box runs, while the `controller_updated` event it received says +exactly that. Two hub surfaces, one blind. -Two independent windows on the same proxy generation: +`demo-felhom` also runs 0.217.0, but it started it at `19:31:04Z` — **17 minutes before the floor was +saved** — and emitted `controller_started` with **no `controller_updated`**. It was moved by hand +during the golden bake, not by the floor. -| window | from → to | delta | rate | +--- + +## 4. PART 2 — the four answers, from source + +**1. What puts a dump into a backup unit, and when? Which apps qualify?** + +- **Database dumps** — `runDBDumps` (`backup/backup.go:444-560`). Qualification is **a running + container whose image matches a database image**, mapped to a stack by `deriveStackName` + (`appbackup/dbdump.go:770`). Written to `/backups/primary//db-dumps/-.sql`. +- **Volume dumps** — `runVolumeDumps` (`backup/backup.go:607`). Qualification is **a deployed, + non-protected stack with at least one Docker *named volume*** (`GetDockerVolumes`). An app with + zero named volumes is skipped silently and is never stopped. +- **When** — one cycle, both legs, scheduled `db-dump` daily at **02:30 CEST**; and again as the + coherence pre-phase of every off-site run (`offbox.go:938-951`), which is what makes a snapshot an + internally coherent {DB@T, files@T} pair. +- `CaptureRecoveryUnit` **only enumerates what is already on disk** + (`recovery_unit.go:131-132`). It writes no dump itself. +- **The gap this leaves:** an app whose data is a *bind mount* and which has no database container + produces neither leg. Its unit is configuration only — and nothing anywhere says so. + +**2. What should a correct backup contain?** + +- **Declares a data drive** (13 of 53): the recovery unit (compose + `app.yaml` carrying the portable + secrets + whatever dumps exist) **plus** the paths its `.felhom.yml` `backup:` block marks + `mandatory`, appended to the restic snapshot as extra paths (`offbox_capture.go:32`). +- **Declares none** (40 of 53): **unit only**. `offboxCaptureSet` returns `(nil, nil)` when the app + has no backup block, so the off-site snapshot is the unit and nothing else. That is correct *given* + that their data is inside the unit's volume tars — and tonight confirmed it is. + +**3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.** + +`ReconstituteFromOffsite` returns `res` and `nil`. Nothing in it ever asks whether anything was +restored. The handler then calls `EndRestoreOp(**true**, reconstituteOutcomeMsg(...))` +(`web/offbox_handlers.go:455`). `reconstituteOutcomeMsg` (`web/offbox_handlers.go:463-475`) formats +counters: + +```go +if res.DBsReplayed == 0 { + return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+ + "Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when) +} +``` + +So with `FilesPlaced == 0` and `DBsReplayed == 0` the customer reads **"0 files restored — the +application restarted"** under a **success**. **This is a finding on its own and is recorded whichever +way the rest goes.** Two aggravations found on top of it: + +- **"This application has no database" is asserted from a counter, not from a fact.** + `reimportDBDumpsFrom` returns `(0, nil)` when the dump directory is absent *or* holds no `.sql` + (`restore_db.go:30-46`), so an app that certainly has a database is told it has none. Proven live on + `paperless-ngx`. +- **The file count is rsync's transfer count, not a restore count.** A correct restore of unchanged + data reports **"0 fájl visszaállítva"** — indistinguishable from a restore that did nothing. + Observed: the same app reported 5, then 2, then 0 across three runs. + +**4. What did the 9 August off-site snapshot contain, and is it still readable?** + +**Readable, and it contained the data.** Three snapshots at `2026-08-09 08:30`: + +| id | app | contents | +|---|---|---| +| `41c830db` | calibre-web | unit + `userdata/media/books` incl. the 9-Aug rehearsal fixture (`árvíztűrő-tükörfúrógép.txt`, `őszibarack.md`, `binary-3mb.bin`, `plain.txt`) | +| `9e38b84c` | opengist | unit + **`volume-dumps/opengist_opengist_data.tar`, 182 272 B** — manifest records `volume_dumps: ["opengist_opengist_data.tar"]` | +| `78b93f04` | privatebin | unit + `volume-dumps/privatebin_privatebin_data.tar`, 2 560 B | + +All still readable tonight; `restic check` over the whole repository reports **"no errors were +found"**. **So the 9 August off-site copy of OpenGist's data exists and is intact** — it simply was +not the thing the afternoon's restore read, and the off-site route that would have read it refuses +for that class of app. + +**The repository had also been dead since 9 August** and nothing said so: no snapshot between +`2026-08-09 08:30` and tonight. The cause is visible in the hub event at 16:01 — the off-box target +was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target, +**every per-app off-site toggle was still off**, so the first run I triggered reported: + +``` +[offbox] backup run started (0 app(s) toggled) +[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s +``` + +**Credit where it is due:** the *card* does not lie about this. It reads „Aktív — nincs kijelölt +alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still +leads, and the log line alone says only „backup OK". + +--- + +## 5. THE TWO CYCLES COMPARED + +*(filled in after the scheduled run — see §11)* + +--- + +## 6. PART 4 — everything attempted, and the message judged + +| # | test | outcome | message judged | |---|---|---|---| -| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** | -| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** | +| 1a | **destructive restore: nothing is ever deleted** | **PASS both ways.** A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. | count is rsync transfers, not files restored — see §4.3 | +| 1b | **safety dump taken and verified before the stop** | **PASS.** `21:02:47Z` dump written → `21:02:47Z` `StopStack romm`. | correct: „0 fájl **és az adatbázis** visszaállítva" | +| 1b | **…and the whole operation refuses if it cannot be** | **PASS.** Made impossible by putting a regular file at the `db-dumps` path. Refused; `plain.txt` kept its post-snapshot mutation; the container's `StartedAt` was unchanged — **the app was never stopped.** | honest and names the path, but leaks a raw Go `mkdir … not a directory` into a customer surface | +| 1b | **the hole in it** | **FAIL.** For `paperless-ngx` the discovery resolves the database to the wrong stack, so `hasDB` is false: **no safety dump is taken and the refusal cannot fire.** The undo is absent, not refused. | „Ennek az alkalmazásnak nincs adatbázisa" — false | +| 2 | **end of the abandonment countdown** | state created and overdue; **fires at 05:10 CEST** — see §11 | card states a **past** date in the future tense while overdue | +| 3 | **damaged store** | **MIXED — see below** | honest at the point of failure, then forgotten | +| 4 | **drive pulled mid-restore** | restore failed, nothing written to the wrong place, **the agent re-bound the drive within 5 s** (`23:15:29` pulled → `23:15:34` re-bound) | **WRONG DIAGNOSIS.** „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions. | +| 5 | **full disk, non-destructive path** | **PASS — refuses before it starts.** | **exemplary:** „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named. | +| 5 | **full disk, destructive path** | **FAIL — no gate at all.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the three headroom gates are all on non-destructive paths (`offbox_restore.go:231,297,423`). It stopped the app, failed halfway on ENOSPC, left `DRILL-2026-08-21/` holding 2 of 5 entries, and restarted the app. | honest (`No space left on device (28)`) but raw rsync output | +| 6 | **controller killed mid-restore** | **PASS.** SIGKILL inside the stop→restore→start window. On restart: *"[appstop] crash recovery: an off-site restore (op \"offbox-reconstitute:paperless-ngx\") was interrupted and left 1 app(s) stopped — restarting them"*. App restarted, marker cleared, **and the hub was told** (event 3009). | good — names the operation, not just the app | +| 7 | **the 40-class under pressure** | **the backup noticed; the alarm did not.** Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The **reserve refused per app**: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got `recovery_unit_capture_failed` (**error**) naming the filesystem and its numbers. **The fill watcher said nothing** — it runs **once a day at 03:30** plus once at startup (`cmd/controller/main.go:1092`). | the refusal messages are good; the silence is the problem | -**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were -extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one -cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was -coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days, -not two years.** Corrected in the incident document and in R-336 rather than left standing. +**Where the clock stopped me:** nothing in Part 4 was skipped for time. Item 2's terminal deletion is +scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only. -**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the -window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge -the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target -connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** +### 4.3 in detail — the damaged store -## 3. The changelog, verbatim — the run's primary deliverable +One byte flipped inside pack `967853d2…` at offset 5 000 000, over the repository's own SFTP +transport. Pack files are named by their content hash, so this is genuine corruption. -All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for -`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen| -backlog|hyper|tokio`. +- **`restic check` detects it** — *"ciphertext verification failed"*, *"Fatal: repository contains + errors"*. +- **But nothing in the product ever runs it.** The only restic verbs in the entire controller are + `restore, snapshots, backup, unlock, stats, init, forget, prune, cat`. The agent's restore-test is + **PBS-tier only**. **The off-site store is never verified by any layer, at any time.** Corruption is + discovered at restore time — the worst possible moment. +- **A restore that touches the damage fails honestly:** `ok=false`, naming the file and + *"ciphertext verification failed"*. +- **But the failure is not remembered.** It left a **partial** checking folder — 78 MB, 54 files, + 15 of 16 originals. `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only *"does the + directory exist and is it non-empty"*, so the wizard then offered **all three** actions including + „Teljes visszaállítás indítása". +- **Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.** -**Exactly one hit, and it is a false positive:** +The repository was repaired from byte-identical originals; `restic check` now reports **"no errors +were found"**. -> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints` -> ` honor the node's proxy settings or not` +--- -HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon. +## 7. PART 5 — the two rows that were observed and never filed -What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests: +**5.1 — the delete guard on verification copies is blind, and its own comment says otherwise.** +`offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) gates on `backupMgr.IsRunning()` — the +**concurrency** flag — while its comment states *"It refuses while a backup/restore op is running: the +copy being deleted could be the one currently being written."* R-351b moved all seven restore handlers +onto `restoreOpBlocked()` (which reads both flags); **this handler was left behind**, and one other +site (`:239`) reads the bare flag correctly and documents why. -> `* backup: harden the handling of client supplied backup manifests:` -> ` - only accept archive names that are plain file names carrying a server side type extension. A` -> ` crafted name in a manifest could previously make a sync job read or write outside of the` -> ` snapshot directory, running as the unprivileged 'backup' user.` -> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every` -> ` archive it lists was really uploaded during that session and that the checksums match` +**Reachability is not a race — it is the whole operation.** `RestoreOffboxScratch` +(`offbox_restore.go:211`) **never calls `acquireRunning` at all**, so `IsRunning()` is false for the +entire duration of an off-site verification restore. Demonstrated live, flags read immediately before +and after the delete: -plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate -limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency, -docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes. +``` +--- BEFORE delete 22:35:25 display= True offbox-restore kimai concurrency= False +--- DELETE calibre-web verification copy: „Az ellenőrző másolat törölve…" +--- AFTER delete 22:35:25 display= True offbox-restore kimai concurrency= False +``` -**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade -for this reason.** +The copy was removed. The guard is app-agnostic, so the same call naming the *restoring* app hits the +directory the restore is writing into. **Rank: MEDIUM** — it needs a customer to press delete during a +restore, but both controls live on the same page, the window is the whole restore, and the target is +the restore's own source. (Reconstitute and place *do* hold the flag, so the exposure is the +verification-restore window only.) -## 4. STOP 1 — the ruling +**5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight.** +Filed as an instrument defect. **Occurrences I can evidence:** -**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we -need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the -upgrade procedure on a Tier-2 protected machine, not a fix for the leak**. +1. **2026-07-20** — an accented grep through `ssh → pct exec → bash -c` nearly produced a wrong + "banner cleared" claim. Recorded in `felhom-controller/.claude/rules/ui-hungarian.md:19-22`. +2. **2026-08-13** — `kubectl exec … sh -c "grep ''"` returned **0 for three strings that + were present**, one step from being reported as a failed hub v0.105.0 deploy. +3. **2026-08-21, tonight, 22:07** — `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md` + (octal escaping). Recording the accented filenames' raw bytes from that listing would have been + wrong. Caught by extracting the archive and reading the names with `xxd`. -**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`): -unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing -explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither -outcome could be rationalised into a success. +**Correction to the task's premise:** that is **two inside two weeks**, plus the founding case a month +earlier. I looked for a third inside the two-week window and did not find one on record. -## 5. STOP 2 — the snapshot, and what it does not cover +**The smallest guard I would propose — and did NOT build:** the problem is not grep, it is that every +one of these tools *silently transforms* the bytes. So the guard is not "use ASCII fragments" (a +discipline, which is what failed three times) but **a negative control that the harness cannot skip**: +any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a +string that MUST be absent — and a zero result from the first is only reportable when the second also +returns zero *and* a third probe for a known-present ASCII anchor returns non-zero. Three probes, one +helper, no judgement required at the call site. Everything else has been tried and is what "nearly" +means in all three cases. -Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not -merely started), server #147604682, project 15217960. +--- -**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and -Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state -(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup -data**. Fine for a package install that writes no datastore content; **it must not be remembered as -datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace -package install would take the only off-premises copy offline. +## 8. RANKED REGISTER ROWS OPENED -## 6. The upgrade and its verification +Ceiling was **R-353**; it **moved to R-365**. -Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then -`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded -server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` — -**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and -what apt did. +| id | rank | what | +|---|---|---| +| **R-354** | **HIGH** | The off-site full restore has **no named-volume leg**. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. `offbox_reconstitute.go:341-346`. | +| **R-355** | **HIGH** | `paperless-ngx`'s PostgreSQL is dumped to a directory for a **non-existent stack** (`…/primary/paperless/`), so its unit records `db_dumps: null`, nothing off-sites it, **no safety dump is taken on a destructive restore**, and the customer is told the app has no database. `appbackup/dbdump.go:770-798`. One app in 53. | +| **R-356** | **HIGH** | The off-site restore **refuses for all 40 no-drive apps** with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. `offbox_reconstitute.go:208-227`. | +| **R-357** | **MEDIUM** | The **destructive** restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory. | +| **R-358** | **MEDIUM** | A **failed** scratch restore leaves a partial copy that `OffboxFullScratchReady` reports as ready; the destructive restore then runs from it and reports success. | +| **R-359** | **MEDIUM** | The off-site restic store is **never verified** by any layer — `restic check` is not among the verbs the controller runs, and the agent's restore-test is PBS-only. | +| **R-360** | **MEDIUM** | Verification-copy delete gates on `IsRunning()`, which `RestoreOffboxScratch` never holds — deletable throughout a restore. Its comment asserts the opposite. **(Part 5.1)** | +| **R-361** | **MEDIUM** | The safety dump **overwrites the unit's own DB dump**: `DumpOne` writes the canonical `-.sql` and only then renames it away. The comment at `offbox_reconstitute.go:147-148` states it "can never overwrite the app's real dump". Proven: romm's `romm-mariadb.sql` was present before and absent after. | +| **R-362** | **MEDIUM** | A data drive detached mid-restore is reported as **„permission denied"**. The restore path never consults drive state. | +| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. | +| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** | +| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). | -| check | result | +**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"` +(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails +nobody. Observed tonight as event 3006, severity `info`. A sweep of every `emit(` call site shows this +is now the **only** remaining instance fleet-wide. + +**R-353:** its instruction (2) is **satisfied** — see §2. Its instruction (1) stands and is now +strictly larger than when it was written, because R-354 shows the bare completion can also be reported +over a unit that *did* have a data leg. + +--- + +## 9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH + +| time (CEST) | what | |---|---| -| installed | server / client / docs **4.2.5-1** | -| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** | -| proxy restarted | 542065 → **551655** @ 09:51:04 | -| **effective `open files`** | **65536 / 65536** — survived the new package | -| drop-ins on disk | both present, unmodified | -| `Recv-Q` | **0** | -| loopback | **`200` in 12 ms** | -| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` | -| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` | -| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` | +| 21:57 | start; baselines | +| 22:00 | break-glass into `demo-hp` | +| 22:09 | fixture planted, comparator control passed | +| 22:12 | manual local backup | +| 22:13 | manual off-site run — **0 apps toggled** | +| 22:17 | off-site run with 3 apps enabled | +| 22:19–22:20 | checking-folder restores | +| 22:21 | reconstitute privatebin → refused | +| 22:23 | reconstitute calibre-web → **the conviction** | +| 22:25 | local restore-from-unit privatebin → data returned | +| 22:27 | Part 4.1a — no-delete invariant | +| 22:39–22:45 | paperless-ngx → R-355 | +| 22:51–22:57 | Part 4.3 damaged store; repo repaired | +| 23:02–23:04 | Part 4.1b safety dump, both directions | +| 23:10–23:12 | Part 4.5 full disk; Part 4.6 controller killed | +| 23:15 | Part 4.4 drive pulled | +| 23:17–00:16 | Part 4.7 filesystem filled and freed | +| 23:35 | abandonment countdown created (fires 05:10) | -`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently -settles the morning's confusion**: that string was never reporting a stale daemon, and now that -installed and running genuinely match, both halves agree. +**Steps off the customer's path, named:** -**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`. -**That unit does not exist** — `systemctl cat` returns *"No files found for -proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server") -and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because -for as long as it took to check, it looked exactly like one. +1. **Break-glass root access** to `demo-hp` via the hub-vaulted `host_recovery` credential — the box + had lost DooPlex's SSH key (its `authorized_keys` held only its own `root@demo-hp` RSA key) and its + tailnet address was unreachable. DooPlex's public key was **re-added** to `/root/.ssh/authorized_keys` + and an `ssh` alias `hp` → `192.168.0.104` was added to `~/.ssh/config` on DooPlex. +2. **A hub DB snapshot** (`hub.db` + `-wal` + `-shm`) was streamed to the scratchpad to read + `host_recovery` and the events table. It holds every host's secret; it is in the session scratchpad + only and is not in any committed file. +3. **`settings.json` was edited directly** (controller stopped, backup at `/root/settings.json.drill-backup`) + to create the overdue abandonment state. There is no product path to shorten a countdown, and the + real orphan→reset path would have destroyed `demo-hp`'s entire off-site history. +4. **The set-aside store the sweep will delete was created by hand** at + `u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, for the same reason. +5. **Two restic pack files were deliberately corrupted and then restored** from byte-identical copies. +6. **`fallocate` fillers** were used to fill two filesystems and were removed. +7. Apps deployed for the drill: `opengist` (re-deployed empty), `kimai`, `paperless-ngx`, `romm`. -**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads -above answer it without mutating a protected datastore. +--- -## 7. The after-slope — unchanged, as predicted +## 10. THE FENCES -| | window | delta | rate | -|---|---|---|---| -| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** | -| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** | +- **`ep0` / the off-site endpoint — untouched outside this machine's own path, and here is how I know.** + `demo-hp` authenticates as the Hetzner Storage Box **sub-account `u629488-sub3`**, which is chrooted + to its own home: `ls /` returns **`Permission denied`**, and `/home` contains exactly `.ssh` and + `felhom-repo`. Every write, the two corruptions, the set-aside store and the sweep's target are + inside that home. **`demo-felhom` is a different sub-account (`u629488-sub1`)** and a real customer's + copy is a different sub-account again — none reachable with this key. `restic forget`/`prune` were + never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept + every 9-August snapshot (verified by listing all 24, not by a count). +- **`demo-felhom`'s two fixtures — confirmed intact, not assumed.** + (a) the unopenable set-aside store `u629488-sub1:/home/felhom-repo.orphaned-20260810` — listed + tonight, `config`/`data` (258 shards)/`index`/`keys`/`locks`/`snapshots`, mtimes still 18 Jul and + 3 Aug; (b) the retained-key case — hub `host_escrow_superseded` rows **11 and 12** for + `demo-felhom-8363b5`, each with a 572-byte `identity_blob`, dated 2026-08-12. `demo-felhom` is + healthy on 0.217.0, its own off-site ran at 02:15 with `last_status: ok`, and + **`--abandon-status` there reports „no abandonment countdown is running on this box"**. +- **`peti-felhom` — not contacted.** It does not appear in the hub host list at all; no command in this + session named it. -**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on -n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure -being numerically higher is noise, not a regression — and certainly not an improvement.** Composition -repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.** +--- -**Thirty minutes cannot settle this in either direction, and nothing here claims it does.** +## 11. SCHEDULED CYCLE AND THE COUNTDOWN -## 8. Register +*(filled in as they fire — 02:30 local backup, 04:15 off-site, 05:10 abandonment sweep, all CEST)* -- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and - **+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact - command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the - split is what identifies which leak it is. If the PID has changed, the window is void. -- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the - mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its - own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a - weekly-write DR endpoint correct. -- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS - version literal, only package names, so there was nothing to update (and per `docs.md`, version - literals do not belong in a current-state doc anyway). +--- -## 9. Teardown +## 12. MACHINE STATES AT THE END -**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text -directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as -the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month. - -## 10. Observations, not acted on - -- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while - status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the - real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity - signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here. -- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on - this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the - thing under test. -- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the - inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run. - -## 11. CI, checked by run ID - -**Run `351`, `head_sha 3e50902a9`, conclusion `failure`, elapsed 15 s** (10:27:00→10:27:15Z). -Inherited, not caused: CI's only step is `python3 scripts/repo_gates.py --fast`, which convicts -`golden-currency` on **R-334** (controller 0.216.0 vs newest golden 0.214.0) — files this run did not -touch. 15 s is the workflow's own honest-failure band, not the R-265 reap band. The previous run -`350` on `435e044` failed identically, before this run began. - -**Expect one `[felhom CI] gates FAILED` mail for run 351.** Same cause as run 348 this morning; not a -new fault, and not related to the upgrade. - -`python3 scripts/unproven.py --summary` — **unchanged: 23 walked, 32 not walked of 55.** No number -moved, correctly: this run proved an operational fact, not a product claim. +*(final state recorded in §13 after the scheduled runs)* diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.listing.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.listing.txt new file mode 100644 index 00000000..a4750ee0 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.listing.txt @@ -0,0 +1,42 @@ +drwxr-xr-x root/root 0 2026-08-21 18:33 offsite-restore/ +drwxr-xr-x root/root 0 2026-08-21 18:33 offsite-restore/opengist/ +drwxr-xr-x root/root 0 2026-08-03 08:11 offsite-restore/opengist/mnt/ +drwxr-xr-x root/root 0 2026-08-04 14:51 offsite-restore/opengist/mnt/sys_drive/ +drwxr-xr-x root/root 0 2026-08-03 08:29 offsite-restore/opengist/mnt/sys_drive/felhom-data/ +drwxr-xr-x root/root 0 2026-08-04 23:13 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/ +drwxr-xr-x root/root 0 2026-08-06 22:02 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/ +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/ +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/ +-rw-r--r-- root/root 1750 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/.felhom.yml +-rw-r--r-- root/root 1260 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/docker-compose.yml +-rw------- root/root 287 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/app.yaml +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/volume-dumps/ +-rw-r--r-- root/root 182272 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/volume-dumps/opengist_opengist_data.tar +-rw-r--r-- root/root 1121 2026-08-09 10:30 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json +drwxr-xr-x root/root 0 2026-08-21 18:30 offsite-restore/calibre-web/ +drwxr-xr-x root/root 0 2026-08-03 08:11 offsite-restore/calibre-web/mnt/ +drwxr-xr-x root/root 0 2026-08-04 14:51 offsite-restore/calibre-web/mnt/sys_drive/ +drwxr-xr-x root/root 0 2026-08-03 08:29 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/ +drwxrwsr-x root/1000 0 2026-08-04 18:45 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/ +drwxr-sr-x root/1000 0 2026-08-04 18:50 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/ +drwxrwsr-x 1000/1000 0 2026-08-09 10:15 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/ +-rw-r--r-- 1000/1000 181 2026-08-04 14:53 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt +drwxrwxr-x 1000/1000 0 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/ +-rw-rw-r-- 1000/1000 21 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/plain.txt +-rw-rw-r-- 1000/1000 3145728 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/binary-3mb.bin +drwxrwxr-x 1000/1000 0 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/nested/ +-rw-rw-r-- 1000/1000 25 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/nested/\305\221szibarack.md +-rw-rw-r-- 1000/1000 59 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/\303\241rv\303\255zt\305\261r\305\221-t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt +-rw-r--r-- 1000/1000 32768 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/metadata.db-shm +-rw-r--r-- 1000/1000 0 2026-08-09 04:15 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/metadata.db-wal +-rw-r--r-- 1000/1000 413696 2026-08-04 14:52 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/metadata.db +drwxr-xr-x root/root 0 2026-08-04 23:13 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/ +drwxr-xr-x root/root 0 2026-08-06 22:02 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/ +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/ +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/ +-rw-r--r-- root/root 3035 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/.felhom.yml +-rw-r--r-- root/root 2122 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/docker-compose.yml +-rw------- root/root 317 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/app.yaml +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/volume-dumps/ +-rw-r--r-- root/root 368640 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/volume-dumps/calibre-web_calibre_web_config.tar +-rw-r--r-- root/root 1157 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/manifest.json diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.sha256 b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.sha256 new file mode 100644 index 00000000..a28d4e18 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.sha256 @@ -0,0 +1 @@ +c9498bfba3dab7b8c59196be8a6c792f0044356d969b2d81ce4b00b61950aa96 documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.tar diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.tar b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.tar new file mode 100644 index 00000000..752614d6 Binary files /dev/null and b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase0-preexisting-scratch/offsite-restore-scratch.tar differ diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/RESULT-MATRIX.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/RESULT-MATRIX.txt new file mode 100644 index 00000000..c2e1deb2 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/RESULT-MATRIX.txt @@ -0,0 +1,22 @@ +DRILL 2026-08-21 — FULL/EMPTY contrast on demo-hp. All hashes sha256, byte-for-byte. +Comparator positive control: one byte flipped at offset 500000 of binary-1mb.bin +(af -> 00) => sha256sum -c FAILED rc=1; original PASSED rc=0. Mutant discarded. + +Planted fixture (5 files, incl. two UTF-8 Hungarian accented names): + SENTINEL.txt 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991 + binary-1mb.bin 725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a + nested/őszibarack.md a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4 + plain.txt 07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1 + árvíztűrő-tükörfúrógép.txt 0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea +Name bytes (UTF-8 NFC): + árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e747874 + őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64 + + | class | data leg | in unit | in offsite snap | checking folder | OFF-SITE restore | LOCAL restore + calibre-web FULL | drive | userdata files | n/a | YES 5/5 ident. | YES 5/5 ident. | YES 5/5 ident. | - + calibre-web FULL | drive | named volume | YES 1.42MB | YES | YES | NO (silent) | - + privatebin FULL | no-drive | named volume | YES 1.06MB | YES 5/5 ident.| YES 5/5 ident. | REFUSED (false) | YES 5/5 ident. + opengist EMPTY | no-drive | named volume | YES 181KB skeleton | YES | YES | REFUSED (false) | - + +CONCLUSION: the unit and the off-site snapshot HOLD the data, verified by identity. +The off-site full restore has no named-volume leg at all; the local restore has one. diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.listing.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.listing.txt new file mode 100644 index 00000000..94e9dbdc --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.listing.txt @@ -0,0 +1,88 @@ +drwxr-xr-x root/root 0 2026-08-21 22:19 offsite-restore/ +drwxr-xr-x root/root 0 2026-08-21 18:33 offsite-restore/opengist/ +drwxr-xr-x root/root 0 2026-08-21 18:00 offsite-restore/opengist/mnt/ +drwxr-xr-x root/root 0 2026-08-21 18:01 offsite-restore/opengist/mnt/sys_drive/ +drwxr-xr-x root/root 0 2026-08-21 18:31 offsite-restore/opengist/mnt/sys_drive/felhom-data/ +drwxr-xr-x root/root 0 2026-08-21 18:31 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/ +drwxr-xr-x root/root 0 2026-08-21 18:31 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/ +drwxr-xr-x root/root 0 2026-08-21 22:17 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/ +drwxr-xr-x root/root 0 2026-08-21 22:17 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/ +-rw-r--r-- root/root 1750 2026-08-21 22:17 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/.felhom.yml +-rw-r--r-- root/root 1260 2026-08-21 22:17 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/docker-compose.yml +-rw------- root/root 287 2026-08-21 22:17 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/compose/app.yaml +drwxr-xr-x root/root 0 2026-08-21 22:16 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/volume-dumps/ +-rw-r--r-- root/root 181248 2026-08-21 22:16 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/volume-dumps/opengist_opengist_data.tar +-rw-r--r-- root/root 1121 2026-08-21 22:17 offsite-restore/opengist/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json +drwxr-xr-x root/root 0 2026-08-21 22:19 offsite-restore/privatebin/ +drwxr-xr-x root/root 0 2026-08-21 18:00 offsite-restore/privatebin/mnt/ +drwxr-xr-x root/root 0 2026-08-21 18:01 offsite-restore/privatebin/mnt/sys_drive/ +drwxr-xr-x root/root 0 2026-08-21 18:31 offsite-restore/privatebin/mnt/sys_drive/felhom-data/ +drwxr-xr-x root/root 0 2026-08-21 18:31 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/ +drwxr-xr-x root/root 0 2026-08-21 18:31 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/ +drwxr-xr-x root/root 0 2026-08-21 22:17 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/ +drwxr-xr-x root/root 0 2026-08-21 22:17 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/ +-rw-r--r-- root/root 1723 2026-08-21 22:17 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml +-rw-r--r-- root/root 1223 2026-08-21 22:17 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml +-rw------- root/root 288 2026-08-21 22:17 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml +drwxr-xr-x root/root 0 2026-08-21 22:16 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/ +-rw-r--r-- root/root 1055744 2026-08-21 22:16 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar +-rw-r--r-- root/root 1118 2026-08-21 22:17 offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json +drwxr-xr-x root/root 0 2026-08-21 18:30 offsite-restore/calibre-web/ +drwxr-xr-x root/root 0 2026-08-21 18:00 offsite-restore/calibre-web/mnt/ +drwxr-xr-x nobody/nogroup 0 2026-08-21 18:26 offsite-restore/calibre-web/mnt/felhom-drives/ +drwxr-xr-x root/root 0 2026-07-22 03:30 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/ +drwxrwsr-x root/1000 0 2026-07-21 19:08 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/ +drwxrwsr-x root/1000 0 2026-07-26 08:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/ +drwxrwsr-x 1000/1000 0 2026-08-21 22:10 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/ +-rw-r--r-- 1000/1000 181 2026-08-04 14:53 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-SENTINEL.txt +drwxrwxr-x 1000/1000 0 2026-08-21 22:09 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/ +-rw-rw-r-- 1000/1000 35 2026-08-21 22:09 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/plain.txt +-rw-rw-r-- 1000/1000 1048576 2026-08-21 22:09 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/binary-1mb.bin +drwxrwxr-x 1000/1000 0 2026-08-21 22:09 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/nested/ +-rw-rw-r-- 1000/1000 21 2026-08-21 22:09 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/nested/\305\221szibarack.md +-rw-rw-r-- 1000/1000 46 2026-08-21 22:09 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/SENTINEL.txt +-rw-rw-r-- 1000/1000 52 2026-08-21 22:09 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/\303\241rv\303\255zt\305\261r\305\221-t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt +drwxrwxr-x 1000/1000 0 2026-08-09 09:59 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/ +-rw-rw-r-- 1000/1000 21 2026-08-09 09:59 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/plain.txt +-rw-rw-r-- 1000/1000 3145728 2026-08-09 09:59 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/binary-3mb.bin +drwxrwxr-x 1000/1000 0 2026-08-09 09:59 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/nested/ +-rw-rw-r-- 1000/1000 25 2026-08-09 09:59 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/nested/\305\221szibarack.md +-rw-rw-r-- 1000/1000 59 2026-08-09 09:59 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/\303\241rv\303\255zt\305\261r\305\221-t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt +-rw-r--r-- 1000/1000 32768 2026-08-21 22:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-shm +-rw-r--r-- 1000/1000 0 2026-08-09 04:15 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-wal +-rw-r--r-- 1000/1000 413696 2026-08-04 14:52 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db +drwxr-xr-x root/root 0 2026-07-23 12:10 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/ +drwxr-xr-x root/root 0 2026-08-21 18:31 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/ +drwxr-xr-x root/root 0 2026-08-21 22:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/ +drwxr-xr-x root/root 0 2026-08-21 22:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/ +-rw-r--r-- root/root 3035 2026-08-21 22:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/.felhom.yml +-rw-r--r-- root/root 2122 2026-08-21 22:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/docker-compose.yml +-rw------- root/root 327 2026-08-21 22:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/app.yaml +drwxr-xr-x root/root 0 2026-08-21 22:16 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/volume-dumps/ +-rw-r--r-- root/root 1422848 2026-08-21 22:16 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/volume-dumps/calibre-web_calibre_web_config.tar +-rw-r--r-- root/root 1165 2026-08-21 22:17 offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json +drwxr-xr-x root/root 0 2026-08-04 14:51 offsite-restore/calibre-web/mnt/sys_drive/ +drwxr-xr-x root/root 0 2026-08-03 08:29 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/ +drwxrwsr-x root/1000 0 2026-08-04 18:45 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/ +drwxr-sr-x root/1000 0 2026-08-04 18:50 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/ +drwxrwsr-x 1000/1000 0 2026-08-09 10:15 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/ +-rw-r--r-- 1000/1000 181 2026-08-04 14:53 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt +drwxrwxr-x 1000/1000 0 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/ +-rw-rw-r-- 1000/1000 21 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/plain.txt +-rw-rw-r-- 1000/1000 3145728 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/binary-3mb.bin +drwxrwxr-x 1000/1000 0 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/nested/ +-rw-rw-r-- 1000/1000 25 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/nested/\305\221szibarack.md +-rw-rw-r-- 1000/1000 59 2026-08-09 09:59 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/rehearsal-2026-08-09/\303\241rv\303\255zt\305\261r\305\221-t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt +-rw-r--r-- 1000/1000 32768 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/metadata.db-shm +-rw-r--r-- 1000/1000 0 2026-08-09 04:15 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/metadata.db-wal +-rw-r--r-- 1000/1000 413696 2026-08-04 14:52 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/userdata/media/books/metadata.db +drwxr-xr-x root/root 0 2026-08-04 23:13 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/ +drwxr-xr-x root/root 0 2026-08-06 22:02 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/ +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/ +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/ +-rw-r--r-- root/root 3035 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/.felhom.yml +-rw-r--r-- root/root 2122 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/docker-compose.yml +-rw------- root/root 317 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/compose/app.yaml +drwxr-xr-x root/root 0 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/volume-dumps/ +-rw-r--r-- root/root 368640 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/volume-dumps/calibre-web_calibre_web_config.tar +-rw-r--r-- root/root 1157 2026-08-09 10:30 offsite-restore/calibre-web/mnt/sys_drive/felhom-data/backups/primary/calibre-web/manifest.json diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.sha256 b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.sha256 new file mode 100644 index 00000000..f4262a67 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.sha256 @@ -0,0 +1 @@ +7e59e57d6458d28f530dbaddbee0f2314ea1ef885052701531f57bad2529849d documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.tar diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.tar b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.tar new file mode 100644 index 00000000..f5c63393 Binary files /dev/null and b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/checking-folders-after-full-restore.tar differ diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/messages-verbatim.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/messages-verbatim.txt new file mode 100644 index 00000000..beb3e6e5 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/messages-verbatim.txt @@ -0,0 +1,32 @@ +VERBATIM customer-facing outcome strings, read from /api/backup/restore-status +(the same value the wizard banner renders). Times CEST. + +1) 22:21:51 off-site FULL RESTORE (reconstitute), app = privatebin [40-class, no HDD_PATH] + ok = FALSE + "A teljes visszaállítás sikertelen: a(z) privatebin nincs telepítve, ezért nincs hová + visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. + Telepítsd újra az alkalmazást (Alkalmazások) ugyanerre a helyre, utána ez a + visszaállítás működni fog" + FACT AT THAT MOMENT: privatebin deployed=true, state=running, container healthy. + +2) 22:23:36 off-site FULL RESTORE (reconstitute), app = calibre-web [drive class] + ok = TRUE + "A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16) — az alkalmazás + újraindult. Ennek az alkalmazásnak nincs adatbázisa." + FACT: the 5 declared-userdata files came back byte-identical. + The named volume calibre-web_calibre_web_config was NOT restored, though its + 1,422,848-byte tar was in the snapshot, in the checking folder, and named in + manifest.json volume_dumps. The message does not mention it. + +3) 22:25:31 LOCAL restore from recovery unit, app = privatebin + ok = TRUE + "privatebin visszaállítva (helyi)." + FACT: the named volume WAS restored, all 5 planted files byte-identical. + The message carries no counts at all - it reads the same whatever happened. + +4) 22:13:15 off-site backup run with ZERO apps selected + log: "[offbox] backup run started (0 app(s) toggled)" + "[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s" + card: badge "Aktív — nincs kijelölt alkalmazás" / "✓ Rendben" / + "Sikeres — nincs mentésre jelölt alkalmazás" / + "Nincs távoli mentésre jelölt alkalmazás — jelölj ki legalább egyet." diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/plant-manifest.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/plant-manifest.txt new file mode 100644 index 00000000..50ffb4c5 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase3-experiment/plant-manifest.txt @@ -0,0 +1,5 @@ +0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991 DRILL-2026-08-21/SENTINEL.txt +725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a DRILL-2026-08-21/binary-1mb.bin +a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4 DRILL-2026-08-21/nested/őszibarack.md +07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1 DRILL-2026-08-21/plain.txt +0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-abandon/repo-inventory-before.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-abandon/repo-inventory-before.txt new file mode 100644 index 00000000..69bc4f75 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-abandon/repo-inventory-before.txt @@ -0,0 +1,59 @@ +ID Time Host Tags Paths +------------------------------------------------------------------------------------------------------------------------------ +e6132ae5 2026-08-04 19:36:26 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +9ac78c98 2026-08-04 19:36:31 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +92212ff8 2026-08-04 19:36:36 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +9d6d233f 2026-08-05 09:12:42 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +29e7b245 2026-08-05 09:12:48 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +cd1db049 2026-08-05 09:12:53 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +0ef7a006 2026-08-06 20:00:44 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +8662a8c1 2026-08-06 20:00:50 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +6dfa6602 2026-08-06 20:00:54 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +3635f945 2026-08-07 02:15:13 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +7b1fa8b0 2026-08-07 02:15:17 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +38edf5b8 2026-08-07 02:15:21 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +a4d03ee3 2026-08-08 02:15:12 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +f534bffe 2026-08-08 02:15:17 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +a685d30e 2026-08-08 02:15:22 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +41c830db 2026-08-09 08:30:38 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +9e38b84c 2026-08-09 08:30:49 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +78b93f04 2026-08-09 08:30:53 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +8c44bd4c 2026-08-21 21:00:52 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +07bac5ad 2026-08-21 21:00:55 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +7e4b703b 2026-08-21 21:00:58 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +16cb8ce7 2026-08-21 21:01:08 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +2a891149 2026-08-21 21:01:13 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +49b9f317 2026-08-21 21:01:22 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist +------------------------------------------------------------------------------------------------------------------------------ +24 snapshots diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/FINDING.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/FINDING.txt new file mode 100644 index 00000000..29e7c1c3 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/FINDING.txt @@ -0,0 +1,34 @@ +paperless-ngx: the database is dumped into a directory for a stack that does not exist, +so the recovery unit never contains it, and the restore then tells the customer the app +has no database. Proven live 2026-08-21 22:39-22:45 CEST on demo-hp. + +MECHANISM (source): + internal/appbackup/dbdump.go:770 deriveStackName("paperless-postgres", known) + 1. candidate = suffixStripStackName("paperless-postgres") = "paperless" (line 801) + 2. known is non-empty, known["paperless"] is FALSE (the stack is "paperless-ngx") + 3. known["paperless-postgres"] is FALSE + 4. no known stack name is a prefix of "paperless-postgres" ("paperless-ngx" is not) + 5. FALLS THROUGH to `return candidate` (line 797) -> "paperless" + An unresolved mapping is returned as if resolved. There is no warning and no refusal. + +OBSERVED CONSEQUENCES (all live): + a) the dump is written to + /mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql + 284,617 bytes, 72 tables, valid=true -- an orphan directory for a non-existent stack, + on the SYSTEM drive, while the app's real unit is on /mnt/felhom-drives/hdd_1. + b) the real unit /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx/manifest.json + records "db_dumps": null + c) the off-site snapshot therefore carries no .sql at all + (checking folder: `find ... -name "*.sql" | wc -l` = 0) + d) writeSafetyDump (offbox_reconstitute.go:115) filters discovered DBs on + db.StackName == stackName, so `mine` is empty -> returns ("", nil) -> hasDB = false. + NO pre-restore safety dump is taken. Verified: `find /mnt -name "pre-restore-*"` + returned nothing before AND after the destructive restore. + e) the destructive restore ran to completion and reported SUCCESS: + "A(z) paperless-ngx: 0 fájl visszaállítva (mentés: 2026-08-21 22:41) + — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." + The controller had dumped that same database 5 minutes earlier. + +The orphan directory is also invisible to the app's off-site push, because the push +resolves paths from the app's own unit path -- so the only copy of that database dump +is on the system drive of the machine it protects. diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/discovery-log.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/discovery-log.txt new file mode 100644 index 00000000..30b551c0 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/discovery-log.txt @@ -0,0 +1,9 @@ +{"timestamp":"2026-08-21T20:38:46Z","level":"DEBUG","message":"DiscoverDatabases: paperless-postgres → stack=paperless, dbUser=paperless, dbName=paperless","source":"dbdump.go:150"} +{"timestamp":"2026-08-21T20:39:37Z","level":"DEBUG","message":"DiscoverDatabases: paperless-postgres → stack=paperless, dbUser=paperless, dbName=paperless","source":"dbdump.go:150"} +{"timestamp":"2026-08-21T20:39:37Z","level":"DEBUG","message":"DumpOne: starting dump for container=paperless-postgres, stack=paperless, dbType=postgres, dumpDir=/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps","source":"dbdump.go:189"} +{"timestamp":"2026-08-21T20:39:37Z","level":"DEBUG","message":"DumpOne: completed paperless-postgres → paperless-postgres.sql (size=277.9 KB, valid=true, tables=72, duration=313ms)","source":"dbdump.go:330"} +{"timestamp":"2026-08-21T20:41:23Z","level":"DEBUG","message":"DiscoverDatabases: paperless-postgres → stack=paperless, dbUser=paperless, dbName=paperless","source":"dbdump.go:150"} +{"timestamp":"2026-08-21T20:41:23Z","level":"DEBUG","message":"DumpOne: starting dump for container=paperless-postgres, stack=paperless, dbType=postgres, dumpDir=/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps","source":"dbdump.go:189"} +{"timestamp":"2026-08-21T20:41:23Z","level":"DEBUG","message":"DumpOne: completed paperless-postgres → paperless-postgres.sql (size=278.2 KB, valid=true, tables=72, duration=314ms)","source":"dbdump.go:330"} +{"timestamp":"2026-08-21T20:43:46Z","level":"DEBUG","message":"DiscoverDatabases: paperless-postgres → stack=paperless, dbUser=paperless, dbName=paperless","source":"dbdump.go:150"} +{"timestamp":"2026-08-21T20:44:33Z","level":"DEBUG","message":"DiscoverDatabases: paperless-postgres → stack=paperless, dbUser=paperless, dbName=paperless","source":"dbdump.go:150"} diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/orphan-dump-and-manifest.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/orphan-dump-and-manifest.txt new file mode 100644 index 00000000..a69ecf3e --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-paperless/orphan-dump-and-manifest.txt @@ -0,0 +1,48 @@ +total 288 +drwxr-xr-x 2 root root 4096 Aug 21 20:41 . +drwxr-xr-x 3 root root 4096 Aug 21 20:39 .. +-rw-r--r-- 1 root root 284902 Aug 21 20:41 paperless-postgres.sql +--- real unit: +{ + "schema_version": 2, + "app_name": "paperless-ngx", + "display_name": "Paperless-ngx", + "controller_version": "0.217.0", + "created_at": "2026-08-21T20:42:01Z", + "drive": "/mnt/felhom-drives/hdd_1", + "namespace_root": "/mnt/felhom-drives/hdd_1", + "image_pins": [ + "ghcr.io/paperless-ngx/paperless-ngx:2.20.15", + "postgres:16-alpine", + "redis:7-alpine" + ], + "secret_env_vars": [ + "DB_PASSWORD", + "PAPERLESS_SECRET_KEY", + "PAPERLESS_ADMIN_PASSWORD" + ], + "data_key_env_vars": null, + "secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore", + "config_files": [ + "docker-compose.yml", + ".felhom.yml", + "app.yaml" + ], + "db_dumps": null, + "volume_dumps": [ + "paperless-ngx_paperless_data.tar", + "paperless-ngx_paperless_postgres_data.tar", + "paperless-ngx_paperless_redis_data.tar" + ], + "checksums": { + ".felhom.yml": "a7cce0a557fd151e6385a137f4721366dd2cd0aa3876783f0f9f0acc9a78dd23", + "app.yaml": "8dfec452dfc00c7e3dce26864bf97165acac44f32470d2425c96e2bd86649cf0", + "docker-compose.yml": "b112952565ef3928f192358ea58fdf4a5a26d788593bf2614e36050c77539b3d" + }, + "portable_secret_env_vars": [ + "DB_PASSWORD", + "PAPERLESS_SECRET_KEY" + ], + "offsite_run_id": "20260821T204123Z", + "dumps_at": "2026-08-21T20:41:23Z" +} diff --git a/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-part4/results.txt b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-part4/results.txt new file mode 100644 index 00000000..35154bb9 --- /dev/null +++ b/documentation/audits/DRILL-backup-truth-2026-08-21/evidence/phase4-part4/results.txt @@ -0,0 +1,64 @@ +PART 4 — things we claim and have never watched. demo-hp, 2026-08-21, times CEST. + +4.1 THE DESTRUCTIVE RESTORE + (a) "nothing is ever deleted" -- PASS, both directions, calibre-web 22:26:59. + POST-SNAPSHOT.txt, created after the snapshot, SURVIVED the restore. + plain.txt, mutated after the snapshot, was OVERWRITTEN back to the snapshot's + content (sha 07e91a98…). Copier is rsync -a, no --delete, no --ignore-existing + (offbox_reconstitute.go:94). + NOTE ON THE COUNT: the message says "2 fájl visszaállítva" because rsync counts + TRANSFERS, not files restored. An identical restore reports "0 fájl visszaállítva", + which is indistinguishable from a restore that did nothing. + + (b) "a safety dump is taken and VERIFIED before anything is stopped, and the whole + operation refuses if it cannot be" + HAPPY PATH -- PASS, romm 23:02:44. + 21:02:47Z "[offbox] romm: pre-restore safety dump written → + pre-restore-20260821T210246Z-romm-mariadb.sql (60.8 KB)" + 21:02:47Z "[stacks] StopStack romm: current state=running" + The dump precedes the stop. Message correctly said + "0 fájl és az adatbázis visszaállítva". + THE REFUSAL -- PASS, romm 23:04:15. + Safety dump made impossible by putting a regular FILE at the db-dumps path. + Result: ok=FALSE, + "A teljes visszaállítás sikertelen: a biztonsági mentés könyvtára nem hozható + létre: mkdir /mnt/felhom-drives/hdd_1/backups/primary/romm/db-dumps: + not a directory" + Nothing changed: plain.txt kept my post-snapshot mutation (sha 3754dfc6…), and + romm's container StartedAt was unchanged (21:03:08) -- the app was never stopped. + JUDGEMENT: honest and it names the path, but it leaks a raw Go mkdir error into + a customer surface. + THE HOLE -- the guard only protects apps whose database the discovery resolves to + the right stack. For paperless-ngx it concludes "no database", so hasDB is false + and the refusal CANNOT fire: the undo is absent rather than refused. See + ../phase4-paperless/FINDING.txt. + SIDE EFFECT, NOT PREVIOUSLY FILED -- the safety dump DESTROYS the unit's own DB dump. + DumpOne writes the canonical `-.sql` (appbackup/dbdump.go:200-202), + i.e. the app's real dump, and only THEN is it renamed to pre-restore-*. + The comment at offbox_reconstitute.go:147-148 states it "can never overwrite the + app's real dump". It does. + PROVEN: romm's db-dumps held romm-mariadb.sql (62,270 B) at 22:59; after one + reconstitute it held ONLY pre-restore-20260821T210246Z-romm-mariadb.sql. + Consequence: until the next backup run the LOCAL restore-from-unit finds no .sql + and tells the customer the app has no database. + +4.3 A DAMAGED STORE -- MIXED + Method: one byte flipped inside pack 967853d2… (offset 5,000,000) via the repo's own + SFTP transport; the pack's name is its content hash, so this is genuine corruption. + * `restic check` DOES detect it ("ciphertext verification failed", + "Fatal: repository contains errors"). BUT the controller NEVER RUNS `restic check`: + the only restic verbs in the whole controller are restore, snapshots, backup, unlock, + stats, init, forget, prune, cat. The agent's restore-test is PBS-tier only. + So the off-site store is never verified by any layer, at any time. + * A restore that TOUCHES the damage fails honestly: + ok=FALSE, "A visszaállítás sikertelen: offbox restore paperless-ngx: exit status 1: + … ignoring error for …/documents/originals/0000011.pdf: ciphertext verification failed" + * BUT the failure is not remembered. It left a PARTIAL scratch (78 MB, 54 files, + 15 of 16 originals). OffboxFullScratchReady (offbox_restore.go:305) only asks + "does the directory exist and is it non-empty", so the wizard then offered all three + actions including "Teljes visszaállítás indítása". + * Pressing it ran the DESTRUCTIVE restore from that known-incomplete copy and reported + SUCCESS: ok=TRUE, "A(z) paperless-ngx: 0 fájl visszaállítva … Ennek az alkalmazásnak + nincs adatbázisa." + REPO REPAIRED afterwards from the byte-identical originals; `restic check` now says + "no errors were found". diff --git a/documentation/audits/REPORT-ep0-pbs-upgrade-2026-08-18.md b/documentation/audits/REPORT-ep0-pbs-upgrade-2026-08-18.md new file mode 100644 index 00000000..cccfdaba --- /dev/null +++ b/documentation/audits/REPORT-ep0-pbs-upgrade-2026-08-18.md @@ -0,0 +1,190 @@ +# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday) + +**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade +anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change, +which is the predicted result.** Two dated checks filed as **R-341**. + +**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only. +Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`. + +--- + +## 1. Baselines re-confirmed on the machine + +Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook. + +| | | +|---|---| +| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** | +| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) | +| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident | +| Effective `open files` | **65536 / 65536** — the morning's drop-in in force | +| `Recv-Q` / loopback | **0** / **`200` in 11 ms** | +| Datastore | 3.7 G of 98 G, 4% | +| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline | + +## 2. The before-slope — and a correction I owe the morning's report + +Two independent windows on the same proxy generation: + +| window | from → to | delta | rate | +|---|---|---|---| +| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** | +| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** | + +**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were +extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one +cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was +coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days, +not two years.** Corrected in the incident document and in R-336 rather than left standing. + +**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the +window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge +the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target +connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** + +## 3. The changelog, verbatim — the run's primary deliverable + +All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for +`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen| +backlog|hyper|tokio`. + +**Exactly one hit, and it is a false positive:** + +> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints` +> ` honor the node's proxy settings or not` + +HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon. + +What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests: + +> `* backup: harden the handling of client supplied backup manifests:` +> ` - only accept archive names that are plain file names carrying a server side type extension. A` +> ` crafted name in a manifest could previously make a sync job read or write outside of the` +> ` snapshot directory, running as the unprivileged 'backup' user.` +> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every` +> ` archive it lists was really uploaded during that session and that the checksums match` + +plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate +limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency, +docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes. + +**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade +for this reason.** + +## 4. STOP 1 — the ruling + +**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we +need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the +upgrade procedure on a Tier-2 protected machine, not a fix for the leak**. + +**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`): +unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing +explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither +outcome could be rationalised into a success. + +## 5. STOP 2 — the snapshot, and what it does not cover + +Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not +merely started), server #147604682, project 15217960. + +**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and +Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state +(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup +data**. Fine for a package install that writes no datastore content; **it must not be remembered as +datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace +package install would take the only off-premises copy offline. + +## 6. The upgrade and its verification + +Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then +`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded +server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` — +**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and +what apt did. + +| check | result | +|---|---| +| installed | server / client / docs **4.2.5-1** | +| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** | +| proxy restarted | 542065 → **551655** @ 09:51:04 | +| **effective `open files`** | **65536 / 65536** — survived the new package | +| drop-ins on disk | both present, unmodified | +| `Recv-Q` | **0** | +| loopback | **`200` in 12 ms** | +| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` | +| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` | +| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` | + +`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently +settles the morning's confusion**: that string was never reporting a stale daemon, and now that +installed and running genuinely match, both halves agree. + +**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`. +**That unit does not exist** — `systemctl cat` returns *"No files found for +proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server") +and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because +for as long as it took to check, it looked exactly like one. + +**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads +above answer it without mutating a protected datastore. + +## 7. The after-slope — unchanged, as predicted + +| | window | delta | rate | +|---|---|---|---| +| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** | +| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** | + +**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on +n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure +being numerically higher is noise, not a regression — and certainly not an improvement.** Composition +repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.** + +**Thirty minutes cannot settle this in either direction, and nothing here claims it does.** + +## 8. Register + +- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and + **+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact + command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the + split is what identifies which leak it is. If the PID has changed, the window is void. +- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the + mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its + own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a + weekly-write DR endpoint correct. +- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS + version literal, only package names, so there was nothing to update (and per `docs.md`, version + literals do not belong in a current-state doc anyway). + +## 9. Teardown + +**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text +directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as +the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month. + +## 10. Observations, not acted on + +- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while + status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the + real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity + signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here. +- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on + this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the + thing under test. +- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the + inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run. + +## 11. CI, checked by run ID + +**Run `351`, `head_sha 3e50902a9`, conclusion `failure`, elapsed 15 s** (10:27:00→10:27:15Z). +Inherited, not caused: CI's only step is `python3 scripts/repo_gates.py --fast`, which convicts +`golden-currency` on **R-334** (controller 0.216.0 vs newest golden 0.214.0) — files this run did not +touch. 15 s is the workflow's own honest-failure band, not the R-265 reap band. The previous run +`350` on `435e044` failed identically, before this run began. + +**Expect one `[felhom CI] gates FAILED` mail for run 351.** Same cause as run 348 this morning; not a +new fault, and not related to the upgrade. + +`python3 scripts/unproven.py --summary` — **unchanged: 23 walked, 32 not walked of 55.** No number +moved, correctly: this run proved an operational fact, not a product claim. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 4d6bd24c..579b373c 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -652,7 +652,19 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes | | **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** Two findings, one session, both shipped. **(a) The blindness.** Every recovery-unit `manifest.json` has carried `drive` and `namespace_root` since schema 1 (`controller/internal/backup/recovery_unit.go:48-49`), written at capture from the app's own placement. `grep -rE '\.Drive\b\|\.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | **CLOSED 2026-08-21** — controller | — | **Shipped:** `backup/offbox_placement.go` (`CheckPlacement`, `PlacementMismatchMessage`, `RecordedUnitForStack`); mismatch **named and refused** before the safety dump, with `ack_placement` as a **separate** field from `confirm=1`; the not-installed refusal names the recorded drive; deploy page **prefills the address and folder from the app's own backup**; `Server.restoreOpBlocked()` reads BOTH flags; `RestoreOpStatus.LastRecent` + `RestoreResultWindow` moved to `internal/backup` as ONE expression for two surfaces. Red-proofs: B with **both** guards removed **was seen starting a restore with no drive attached** (no error, full 3.00 s run into `/tmp/mutant-destination`); C, E and A each returned their wrong outcome; D forced on broke 8 ordinary reconstitute tests, proving reachability both ways. | CC | | **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. | CC | +| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC | +| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **OPEN — HIGH** | — | Restore the unit's `volume-dumps/*.tar` from the SCRATCH unit on the off-site path, as `RestoreFromRecoveryUnit` already does from the live unit. Assert the CONSEQUENCE (planted bytes come back), not the mechanism. | CC | +| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **OPEN — HIGH** | — | Two candidates, both two-repo: rename the container to `paperless-ngx-postgres`, or make an unresolved candidate a loud skip rather than a silent fallback. The second is the one that generalises. Pin with a test that a container whose name resolves to no known stack is never dumped silently. | CC | +| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. | CC | +| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC | +| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC | +| **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC | +| **R-360** | **The verification-copy delete gates on the concurrency flag, which the verification restore never holds — so the copy is deletable for the whole restore, and the handler's own comment claims the opposite.** `offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) reads `backupMgr.IsRunning()`; `RestoreOffboxScratch` (`offbox_restore.go:211`) **never calls `acquireRunning`**. R-351b moved all seven restore handlers onto `restoreOpBlocked()` (both flags) and left this one behind. Demonstrated 2026-08-21 22:35 with the flags read immediately before and after: `display=True offbox-restore kimai / concurrency=False` on both sides, and the delete succeeded. The guard does not compare app names, so the same call naming the restoring app removes the directory the restore is writing into. **Observed in two consecutive reports and filed neither time; filed now.** | **OPEN — MEDIUM** | — | `restoreOpBlocked()`, and a test that asserts the CONSEQUENCE — the copy survives a delete attempt mid-restore. | CC | +| **R-361** | **The pre-restore safety dump overwrites the app's own DB dump, and the comment beside it says it cannot.** `writeSafetyDump` calls `DumpOne`, which writes the canonical `-.sql` (`appbackup/dbdump.go:200-202`) — i.e. the unit's real dump — and only THEN renames it to `pre-restore-*`. The comment at `offbox_reconstitute.go:147-148` states the rename means it "can never overwrite the app's real dump". Proven 2026-08-21: `romm`'s `db-dumps/` held `romm-mariadb.sql` (62 270 B) at 22:59 and held ONLY `pre-restore-20260821T210246Z-romm-mariadb.sql` after one reconstitute. Until the next backup run the local restore-from-unit finds no `.sql` and reports the app has no database. The `pre-restore-*` file is also enumerated into the manifest's `db_dumps` and shipped off-site. | **OPEN — MEDIUM** | — | Dump to the safety name directly, or to a temp name. Pin the invariant the comment already asserts. | CC | +| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC | +| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC | +| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC | +| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC | | **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC | | **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc//fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |