# RECON — what DooPlex's own backup actually protects **2026-08-06, read-only.** Nothing was modified, stopped, pruned or deleted. One scratch restore into `/tmp/recon-restore-20260806-121056`, deleted afterwards. No customer machine touched. **State at start:** no backup running, no restic process. `dooplex-backup.timer` **enabled + active**, next fire 2026-08-07 03:18:49. The service showing `disabled` is **expected** for a timer-driven unit and is **not a finding** — recorded here so nobody chases it later. Last run 03:06:07 → 03:12:51, `Result=success`, `ExecMainStatus=0`. --- ## 1. Set inventory Five sets run from `backup-all.sh`. Two more things write into the same tree and are **not** part of that service. Every "last success" below is from the journal or the repo, not from a config line. | Set | Source paths | Source dev | Mechanism | Destination | Dest dev | Schedule | Last success | Snaps | Oldest → newest | Verified? | Off-site | |---|---|---|---|---|---|---|---|---|---|---|---| | **K3s State** | `/var/lib/rancher/k3s/server/{db,token,node-token,cred,tls}` · `/etc/rancher/k3s` | `sdb1` | restic | `backup/k3s/restic-repo` | **`sda1`** | nightly 03:00 | 2026-08-06 03:06:07 | 26 (13+13) | 03-31 → 08-06 | **yes** — `check --read-data-subset=5%` every run | no | | **Kubernetes Secrets** | k8s API (all secrets + configmaps, per-namespace) | n/a | **GPG files — NO restic** | `backup/secrets/exports/` | **`sda1`** | nightly | 2026-08-06 03:06 | **0** | 07-06 → 08-06 (96 files, 134 MB, 30-day `find -mtime +30 -delete`) | no | no | | **Git Manifests** | Gitea repos via API token | n/a | clone/mirror | `backup/homelab-manifests/` | **`sda1`** | nightly | 2026-08-06 03:06 | — | — | no | no | | **User Data** | `/mnt/4_hdd/data` · `/mnt/5_hdd/felhom.eu/git/.claude-memory` | `sdg2` · **`sda1`** | restic | `backup/data/restic-repo` | **`sda1`** | nightly | 2026-08-06 10:51:17 (manual) / 03:0x nightly | 14 | 03-31 → 08-06 | **no** | no | | **PostgreSQL Dumps** | `/mnt/5_hdd/backup/postgresql/dumps` | **`sda1`** | restic | `backup/postgresql/restic-repo` | **`sda1`** | nightly | 2026-08-06 03:09:52 | 11 | 06-25 → 08-06 | **no** | no | | *(separate)* **Longhorn PVCs** | 77 volumes, replicas under `/var/lib/longhorn` | `sdb1` | Longhorn backupstore | `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` | **`sda1`** | daily 04:00 + weekly, **retain=1** | 2026-08-06 04:10 (sync 10:08) | 77 vols / 88 GB | — | no | no | | *(separate)* **jarrs.eu** | Hetzner → DooPlex | remote | rsync **inbound** | `backup/jarrs.eu/` | `sda1` | cron 04:00 | 2026-08-06 04:00 | 33 files / 19 MB | — | n/a | **inbound only** | **Dormant destination directories** (nothing writes them any more): `demo-felhom` 25 GB, newest **2026-07-10** (27 days); `sparkyfitness` 652 KB, newest **2026-06-24** (43 days); `secrets/restic-repo` — initialised 2026-01-05, **zero snapshots ever**. Tree total ≈ 167 GB. `/mnt/5_hdd` is 24% used, 6.6 TB free — **no space pressure**. **Per-set reliability: 14 runs in 14 days, all five sets `✓ completed successfully` every run.** The routine works; the questions below are about what it would survive, not whether it runs. --- ## 2. The same-disk answer, stated directly **The destination for every set is `/dev/sda1`** (WDC WD103KFBX-68CCLN0, serial `WD-B1036H8J`). | Set | Source device | Same disk as destination? | |---|---|---| | K3s State | `sdb1` (Apacer SSD, root) | **No** — cross-disk | | Longhorn PVCs | `sdb1` (replicas under `/var/lib/longhorn`) | **No** — cross-disk | | User Data → `/mnt/4_hdd/data` | `sdg2` (ST3000VX010 `Z6A07P2G`) | **No** — cross-disk | | **User Data → `.claude-memory`** | **`sda1`** | **YES — same physical disk** | | **PostgreSQL Dumps** | **`sda1`** (`/mnt/5_hdd/backup/postgresql/dumps`) | **YES — and the source is *inside the backup tree*** | | Kubernetes Secrets | k8s API (no disk source) | destination `sda1`, single copy | | Git Manifests | Gitea (which is a Longhorn PVC on `sdb1`) | No — cross-disk | **The operator's suspicion was right, and it is narrower than feared.** Two sets are same-disk: the memory store (added 2026-08-06) and the PostgreSQL dumps. The postgres case is the stranger of the two — the dump directory lives *inside* `/mnt/5_hdd/backup/`, so restic is deduplicating a copy of files that are already in the backup tree, onto the same disk. It protects against a bad dump, not against losing `sda1`. Everything else is genuinely cross-disk. **But cross-disk is not off-site — see §4.** --- ## 3. Failure matrix **"Demonstrated" means someone has actually done it and checked the result.** Only one cell in this column is a yes, and it was earned today. | Failure | What is lost | What is recoverable | From where | How long | Ever demonstrated? | |---|---|---|---|---|---| | **A file deleted by mistake** | the file | **all of it**, to the previous night (or to the 6-hourly dump for a DB) | the relevant restic repo / Longhorn backup on `sda1` | minutes | **YES — today.** A file was restored from the `data` repo and matched the live original byte for byte (§5) | | **`sda1` (5_hdd) fails** | **every backup, all of it** — all four restic repos, the Longhorn backupstore, the secrets exports, the manifest mirror | the *live system* keeps running: k3s, Longhorn replicas, Gitea, the registry and the hub are all on `sdb1`. `/mnt/4_hdd/data` survives on `sdg2`. **`.claude-memory` is lost** (source and only copy were both on `sda1`) | nothing — there is no second copy of the backups | n/a for the backups; the live system is unaffected | no | | **`sdg2` (4_hdd) fails** | the live `/mnt/4_hdd/data` | **all of it** | `data` restic repo on `sda1` | hours (405 GiB) | **partially** — the single-file probe worked; a full-set restore has never been run | | **DooPlex's board or PSU dies** | nothing permanently, if the disks survive | **everything**, by moving the disks or rebuilding and reattaching | `sda1` + `sdb1` in hand | days — and see the note below | no | | **The flat burns** | **everything** — every disk is in the same box, and **nothing leaves it** | **nothing from this backup system.** Separately: customer data has a genuine off-site tier (Hetzner + ep0); the felhom *repos* exist on Gitea, which is on DooPlex; some repos have working clones on this same machine | — | — | no | | **A bad script or ransomware writes over `/mnt/5_hdd/backup/`** | every backup at once | the live system only | nothing — the tree is a single writable path, and the restic repos are **not** `--append-only` | n/a | no | **The note on "board or PSU dies":** recovery needs both the disks *and* the restic passphrase. The on-box copy is `/etc/backup/restic-password` on **`sdb1`** — a different disk from the backups, which is good. **The operator has confirmed an offline copy exists out of band**, so this row is recoverable rather than fatal. See §6 for the one residual that remains. --- ## 4. Does anything leave the box? **No. Nothing.** Stated without softening, because it is the row the matrix turns on. - No `rclone`, no remote restic repository, no rsync-to-remote, no S3/B2/SFTP target anywhere in `/opt/backup/scripts/`. - The only outbound-looking job, `sync-hetzner-backups.sh` (cron 04:00), is **inbound**: it pulls *jarrs.eu* database backups **from** Hetzner **to** DooPlex. It is a different project and it moves data the wrong way to help here. - Longhorn's backup target is `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — **DooPlex writing to itself over NFS**, landing on `sda1`. The felhom *product* has a real off-site chain (Hetzner Storage Box + ep0/PBS-DR) for customer boxes. **DooPlex, which runs the hub that manages that chain, has none of it.** --- ## 5. The restore probe | | | |---|---| | Repo | `/mnt/5_hdd/backup/data/restic-repo` | | Snapshot | `b587f775` (2026-08-06 10:51:17) | | File | `.claude-memory/hetzner-storagebox-api.md` (3.812 KiB) | | Restored sha256 | `5b4d66b0341fe38c21b3a77f2f0af59ad4389fc21a8ce941c729fb807a3dd953` | | Live sha256 | `5b4d66b0341fe38c21b3a77f2f0af59ad4389fc21a8ce941c729fb807a3dd953` | | **Result** | **MATCH — a file came back, byte for byte** | Scratch directory `/tmp/recon-restore-20260806-121056` **deleted**. A first attempt used `MEMORY.md` and the hashes **differed** — correctly: the restored copy was 17,977 bytes, the file's exact size *at snapshot time*, and the live file had since been edited to 18,078. That is the snapshot behaving as a point-in-time copy, not a fault, and the probe was redone against a file this session never touched. --- ## 6. The password question **Measured on the box:** `/etc/backup/restic-password` — 45 bytes, `0600 root:root`, on **`/dev/sdb1`** (the root SSD, *not* the disk holding the backups). It is **not a source path in any set**. The same file is the passphrase for the GPG-encrypted secrets exports (`--passphrase-file "${RESTIC_PASSWORD_FILE}"`), so one secret opens both the restic repos and the secrets archive. **Off the box:** the operator confirms **an offline copy exists out of band.** A recon on the machine cannot observe that, and without it this section would have concluded that an `sdb1` loss makes every repo permanently unopenable — which would have been wrong. The failure matrix above is written on the corrected basis. **The residual, stated narrowly and without alarm:** the offline copy is *operator-held*, not *system-held*. The felhom product solves the equivalent customer problem with hub-vaulted escrow and a zero-knowledge recovery code; the operator case has no equivalent mechanism, so recovery time in the "board dies" and "flat burns" rows depends on that copy being **current** and **findable by someone other than the operator**. Worth confirming both; not worth engineering around. --- ## 7. Coverage gaps | Item | Status | Evidence | |---|---|---| | **Gitea repositories** | **covered** (twice) | `gitea-data` is a 50 Gi Longhorn PVC → backupstore; plus `backup-manifests.sh` mirrors repos via the API to `backup/homelab-manifests/` | | **Container registry** | **covered** | part of `gitea-data` (Gitea's built-in registry) | | **Gitea database** | **covered** | CNPG → 6-hourly dumps → `postgresql` restic repo | | **k3s state** | **covered, and the only verified set** | `db`, `token`, `node-token`, `cred`, `tls` + `/etc/rancher/k3s`, `check --read-data-subset=5%` each run | | **Longhorn volumes** | **covered, thinly** | 77 volumes, but **`retain=1`** — one generation per volume. A corruption noticed after the next run has no earlier copy | | **Hub database** | **covered** | `hub-data` (1 Gi) is a Longhorn PVC | | **TLS material / secrets** | **partially** | exported and GPG-encrypted nightly, 30-day retention, on `sda1` only, **never entered into any restic repo** — `backup-secrets.sh` contains no `restic` call at all | | **`/opt/backup` itself** | **NOT covered** | on `sdb1`, in no set. The scripts that perform every backup are backed up by nothing (R-231) | | **`~/.claude/settings.json`** | **NOT covered** | on `sdb1`, in no set. Now reconstructible via `install_workspace.py` | | **ISO build inputs** | **NOT covered** | `/mnt/5_hdd/felhom.eu/build`, on `sda1`, in no set | | **Local git clones** | **NOT covered — and circular** | `/mnt/5_hdd/felhom.eu/git` on `sda1`, in no set. Recoverable from Gitea — **which runs on this machine.** If DooPlex is gone, both copies are gone together | | **`.claude-memory`** | covered, **same-disk only** | §2 | --- ## 8. Observations — noticed, not acted on 1. **`/opt/backup/docs/` is EMPTY.** The systemd unit advertises `Documentation=file:///opt/backup/docs/BACKUP-RESTORE.md` and that file does not exist. The restore procedure you would reach for while the machine is on fire is absent, and the unit points at it anyway. 2. **`notify_failure` is a no-op.** `NOTIFY_ON_FAILURE="true"` but `NOTIFY_WEBHOOK_URL` is commented out, and the function requires it to be non-empty. **A failed backup notifies nobody.** The five sets have succeeded 14/14, so this has never mattered yet — which is exactly when it is cheapest to fix. 3. **`secrets/restic-repo` exists and has never held a snapshot.** A directory that looks like a backup and is not one. `backup-secrets.sh` never calls restic; its only match for the word is a comment about the password file. 4. **A stale lock is sitting in the `data` repo, and it is MINE.** `pid 3524524`, non-exclusive, created 11:20:26, process gone — left by a `timeout`-wrapped `restic ls` during the previous task's verification. **I am fenced from clearing it (§4 of the task) and have not.** Its impact is concrete: `restic backup` needs only a non-exclusive lock and will succeed, but `restic check` already refused (§ the check run below), and `apply_retention` → `forget --prune` needs an **exclusive** lock. With `set -euo pipefail` and the `ERR` trap, tonight's User Data set would therefore fail — and per observation 2, silently. One command clears it: `sudo RESTIC_PASSWORD_FILE=/etc/backup/restic-password restic -r /mnt/5_hdd/backup/data/restic-repo unlock` 5. **Integrity checks:** `k3s` and `postgresql` pass `restic check` cleanly; `secrets` is empty so passes trivially; **`data` could not be checked** because of observation 4. Only `k3s` is verified on a schedule — the others have never been verified except by hand. 6. **Two dormant destination directories** hold 25 GB (`demo-felhom`, last written 2026-07-10) and 652 KB (`sparkyfitness`, 2026-06-24). Nothing writes them; nothing prunes them. 7. **The postgres dumps are produced by a k8s CronJob** (`database-system/postgresql-backup`, every 6 h) that the backup service does not control. `backup-postgresql.sh` refuses only if the directory is *missing* — if the CronJob stopped, the set would keep succeeding over stale dumps. Currently fresh: newest dump `20260806-100000` at 12:00 today. 8. **`restic check --read-data-subset` was not run** beyond what `backup-k3s.sh` does itself, per the task's fence. No repository has ever had its data blocks read back in full. --- ## 9. Recommendations, ranked — recommend only, nothing implemented | # | Recommendation | Cost | Failure it removes | |---|---|---|---| | 1 | **Clear the stale lock in the `data` repo** (command in §8.4) | one command | tonight's User Data set failing, silently | | 2 | **Give `notify_failure` somewhere to go** — the project already has a working transactional mail path (Resend, used by CI and the hub) | ~10 lines | a backup failing and nobody learning it. This is the cheapest item here and the one that makes every other failure visible | | 3 | **Get one copy off the box.** Even one set — k3s + secrets + postgres dumps is ~1.2 GB — to the Storage Box or ep0 | small; the credentials and the tooling already exist | the "flat burns" row, which is currently total loss | | 4 | **Make `/mnt/5_hdd/backup` non-trivially destroyable** — restic `--append-only` via a REST/rclone-served repo, or a second target the writer cannot reach | moderate | the ransomware / bad-script row | | 5 | **Move the postgres dump directory out of the backup tree**, and add `.claude-memory` to a cross-disk destination | small | the two same-disk cases in §2 | | 6 | **Raise Longhorn `retain` above 1** | disk only, and there is 6.6 TB free | a corruption noticed one day late having no earlier copy | | 7 | **Write `/opt/backup/docs/BACKUP-RESTORE.md`**, or drop the `Documentation=` line | an hour | reaching for a restore procedure that is not there | | 8 | **Do one full-set restore drill** and record it | half a day | the entire "Ever demonstrated?" column being empty | | 9 | **Put the secrets exports into their empty restic repo**, or delete the repo so it stops looking like a backup | small | a directory that reads as protection and is not | --- ## 10. The one-sentence answer **The backup runs reliably, restores a file correctly, and would survive any single disk failing — but every copy it makes lives in one box, and if that box is gone, so is all of it.**