diff --git a/documentation/audits/dooplex-survival-2026-10-09/PLAN.md b/documentation/audits/dooplex-survival-2026-10-09/PLAN.md new file mode 100644 index 00000000..efe3262e --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/PLAN.md @@ -0,0 +1,97 @@ +# PLAN — an encrypted off-site copy of Gitea and DooPlex's secrets (R-232 (b), (h)) + +2026-10-09, Part A of the DooPlex-survival brief. Everything below was **measured read-only**; nothing was created. +Raw readings: `partA/readings.txt`. + +## What is measured + +| Thing | Size | Growth | Where it lives today | +|---|---|---|---| +| Gitea repositories (`/data/git/repositories`, 10 repos under `admin/`) | **614 MB** | ~600 MB in 8 months (oldest product repo 2026-02-11); felhom.eu.git 188 MB | Longhorn PVC `gitea-data` on `sdb1`; Longhorn backup on `sda1`, **retain=1**. Nothing else. | +| Gitea database (CNPG `postgresql`, db `gitea`) | 64 MB live, **6.2 MB** as a `pg_dump -Fc` | 5.62 → 6.22 MB in 4 days (~150 KB/day) | 6-hourly dumps in `/mnt/5_hdd/backup/postgresql/dumps/` (same disk as every backup) | +| Gitea config `app.ini` (holds `SECRET_KEY`, `INTERNAL_TOKEN`, JWT secrets) | 2 KB | — | in the PVC only | +| Gitea registry (`/data/gitea/packages`) | 27.7 GB | — | **Left out on purpose:** the images rebuild from the code. | +| DooPlex secrets set: the nightly GPG files (`secrets-`, `configmaps-`, `by-namespace-*.gpg`) | **4.4 MB per night** | flat (2,055,542 → 2,056,090 B in 12 days) | `sda1` only, 30 days | + +**Correction to the 2026-08-06 recon (§7):** it said the Gitea repositories are "covered twice". The API mirror +(`backup/homelab-manifests/git-mirrors/`) holds **one** repository, `homelab-manifests.git` (7.1 MB). The four +product repositories have exactly one backup: the Longhorn copy with retain=1, on the same machine. + +## The two places measured + +| | ep0, namespace `operator` (where the hub DB goes) | A new Hetzner Storage Box sub-account | +|---|---|---| +| Free space | 83 GB of 98 GB (`/mnt/pbs-datastore`, 16 % used) | depends on the box; not measured | +| Write-only key | **Already exists.** Token `dooplex-hub@pbs!push` has `DatastoreBackup` on `/operator` only: it cannot forget a snapshot or list the households (measured 2026-10-05). A second read-only token `!restore` exists for the restore test. | the R-820 append-only pin; the sub-account shell can still `rm` unless the pin is right | +| Changes needed there | **None.** A new backup group (`host/dooplex-gitea`) in the same namespace needs no new token, ACL or prune job. The existing prune job `prune-operator-hubdb` (keep-daily 14, keep-weekly 8) covers every group in `operator`. | a new sub-account + key custody | +| Cost | 0 € extra | a sub-account (same Storage Box fee) | +| Return copy | ep0's nightly pull to DooPlex (`ep0-copy`, 05:00) copies it back too — ciphertext only | none | + +**Pick: ep0 `operator`.** No change on ep0 at all, the same write-only token as the hub DB, the same prune and +the same alarm family. Worst case on ep0: a git repack every night makes each copy new chunks — 22 kept copies × +0.65 GB ≈ 14 GB, inside the 83 GB free. + +## What runs + +A script and a timer on DooPlex, `felhom-dooplex-offsite` (versioned in `felhom.eu/scripts/dooplex-offsite/`, +installed by its `install.sh`, tests by hand like the hub DB's), **daily 00:20** (after the 00:00 database dump, +before the hub DB at 02:30): + +1. **The database first:** the newest **complete** `gitea.dump` + `globals.sql` from the 6-hourly dumps (folder with + `SUCCESS`; refused if older than 7 h). This is PostgreSQL's own consistent snapshot. +2. **Then the files**, read from the running Gitea pod with `tar` over `kubectl exec`: `/data/git/repositories`, + `/data/git/lfs`, `/data/gitea/conf/app.ini`, attachments, avatars, `jwt`. **Not** packages, logs, indexers, queues, tmp. +3. **Then the secrets set:** the newest night's three `.gpg` files (already GPG-encrypted with DooPlex's restic + passphrase, which the operator holds offline). +4. A `MANIFEST.sha256` of everything, then `proxmox-backup-client backup` → `host/dooplex-gitea` in `operator`, + `--crypt-mode encrypt` with a **new key** `/etc/felhom-dooplex-offsite/enc.key`. +5. Stage shredded; a success timestamp (written only after the push returns 0) for Prometheus. + +**Why not Gitea's own `gitea dump`:** Gitea's documentation says the instance "must be shutdown during backup" for +consistency, and `gitea dump` does not stop it either. A nightly stop costs CI and the registry a gap and adds a +restart that must come back. Instead the order makes the copy safe: the database is older than the repositories, so +every commit the database names is in the copy (git writes objects before it moves a ref). A repository newer than +its database is the normal state after a push and Gitea reads it. **The restore test catches the rest:** it runs +`git fsck` on every repository. + +## Encryption and the key + +- New key, generated on DooPlex as root: `proxmox-backup-client key create --kdf none` → `0600` file in + `/etc/felhom-dooplex-offsite/`, never printed by CC, never in chat. +- **The operator copies its paper form into the password manager** at his own terminal (not through `!` here): + `sudo proxmox-backup-client key paperkey /etc/felhom-dooplex-offsite/enc.key --output-format text` — the `data` + field is enough (as for the hub DB key, runbook Step 0). +- ep0 never sees the key. The secrets files inside are encrypted twice (GPG + PBS). + +## Retention and who can delete + +- ep0's existing `prune-operator-hubdb` (03:45): 14 daily + 8 weekly, per group. ep0's GC frees chunks. +- **Can delete:** only ep0's root (and that prune job). The push token cannot forget a snapshot; the restore token + cannot write. DooPlex's root cannot delete on ep0 with what it holds. + +## Alarms and failure mail + +- `DooplexGiteaOffsiteStale` (26 h, `absent()` included) and `DooplexGiteaRestoreTestStale` (8 days), same group as + `HubDBBackupStale`, proven with `promtool test rules` and a red-proof. +- Failure mail: an `OnFailure=` unit that calls the R-232 (a) `notify_failure` (Resend → admin@felhom.eu). The same + `OnFailure=` is added to the two hub-DB units. Proven with one dry failure (a unit that fails on purpose, no backup). + +## The restore test + +- **Weekly, Sunday 05:30**, with the read-only token: restore the newest copy to a temp dir, check `MANIFEST.sha256`, + `git fsck` every repository, `pg_restore --list gitea.dump`, then shred; success timestamp. +- **Once now (Part C):** restore into a throwaway machine, start Gitea there with no outside network, compare every + product repository's `main` with live Gitea, read one file byte for byte, log in with a throwaway admin. Steps + become `runbooks/gitea-restore.md`. The throwaway and every copy are deleted. + +## What changes on DooPlex (needs the operator's yes) + +1. `/usr/local/sbin/felhom-dooplex-offsite` + `felhom-dooplex-offsite-restore-test` (scripts, from git). +2. Units: `felhom-dooplex-offsite.{service,timer}`, `felhom-dooplex-offsite-restore-test.{service,timer}`, + `felhom-backup-failmail@.service`; one `OnFailure=` line added to the two hub-DB services. +3. `/etc/felhom-dooplex-offsite/` (root 0700): `env` (no secret) and the new `enc.key`. The tokens are the existing + `/etc/felhom-hub-backup/token-push` and `token-restore` (read, not copied). +4. Two alarm rules in `homelab-manifests` `prometheus-rules` + the Prometheus reload. +5. One manual run, one manual restore test, one dry failure mail. + +**On ep0: nothing.** Reads only (the snapshot list). diff --git a/documentation/audits/dooplex-survival-2026-10-09/partA/readings.txt b/documentation/audits/dooplex-survival-2026-10-09/partA/readings.txt new file mode 100644 index 00000000..3d37fbed --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partA/readings.txt @@ -0,0 +1,60 @@ +# Part A readings, 2026-10-09T07:54:41Z, read-only +## Gitea +gitea version 1.26.2 built with go1.26.3-X:jsonv2 : bindata, timetzdata, sqlite, sqlite_unlock_notify +613.5M /data/git/repositories +4.0K /data/git/lfs +27.7G /data/gitea/packages +12.0K /data/gitea/conf +4.0K /data/gitea/attachments +12.0K /data/gitea/avatars +8.0K /data/gitea/jwt +app-catalog-drill.git +app-catalog-felhom.eu.git +felhom-agent.git +felhom-controller.git +felhom.eu.git +homelab-manifests.git +jarr.git +misc-scripts.git +recipe-importer.git +revfulop-calendar.git +188.1M /data/git/repositories/admin/felhom.eu.git +73.7M /data/git/repositories/admin/felhom-controller.git +61.4M /data/git/repositories/admin/felhom-agent.git +19.2M /data/git/repositories/admin/app-catalog-felhom.eu.git +DB_TYPE = postgres +HOST = postgresql-rw.database-system.svc.cluster.local:5432 +NAME = gitea +## Gitea DB +64 MB +/mnt/5_hdd/backup/postgresql/dumps/20261005-100001/gitea.dump 5624428 +/mnt/5_hdd/backup/postgresql/dumps/20261009-040001/gitea.dump 6218173 +## secrets set +-rw-r--r-- 1 root root 1926239 Sep 28 03:07 by-namespace-20260928_030725.tar.gz.gpg +-rw-r--r-- 1 root root 1944545 Oct 9 03:10 by-namespace-20261009_031011.tar.gz.gpg +-rw-r--r-- 1 root root 401829 Sep 28 03:07 configmaps-20260928_030725.yaml.gpg +-rw-r--r-- 1 root root 404486 Oct 9 03:10 configmaps-20261009_031011.yaml.gpg +-rw-r--r-- 1 root root 2055542 Sep 28 03:07 secrets-20260928_030725.yaml.gpg +-rw-r--r-- 1 root root 2056090 Oct 9 03:10 secrets-20261009_031011.yaml.gpg +130M /mnt/5_hdd/backup/secrets/exports +## API mirror +homelab-manifests.git +## ep0 (read-only ssh) +Filesystem Size Used Avail Use% Mounted on +/dev/sdb 98G 15G 83G 16% /mnt/pbs-datastore + "comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)", + "id": "prune-demo-hp", + "keep-last": 2, + "ns": "demo-hp", + "id": "prune-operator-hubdb", + "keep-daily": 14, + "keep-weekly": 8, + "ns": "operator", + "comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)", + "id": "prune-demo-felhom", + "keep-last": 2, + "ns": "demo-felhom", +/datastore/felhom-offsite/operator dooplex-hub@pbs DatastoreBackup +/datastore/felhom-offsite/operator dooplex-hub@pbs DatastoreReader +/datastore/felhom-offsite/operator dooplex-hub@pbs!push DatastoreBackup +/datastore/felhom-offsite/operator dooplex-hub@pbs!restore DatastoreReader diff --git a/documentation/audits/dooplex-survival-2026-10-09/partB/red-proof.txt b/documentation/audits/dooplex-survival-2026-10-09/partB/red-proof.txt new file mode 100644 index 00000000..4f1eddb9 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partB/red-proof.txt @@ -0,0 +1,31 @@ +### RED: link-check (removed: LINKS) +FAIL: test_a_link_in_the_pods_archive_refuses (__main__.Push.test_a_link_in_the_pods_archive_refuses) +Ran 1 test in 0.283s +FAILED (failures=1) +### RED: dump-age (removed: DAGE) +FAIL: test_stale_dump_refuses (__main__.Push.test_stale_dump_refuses) +Ran 1 test in 0.289s +FAILED (failures=1) +### RED: repo-count (removed: GOT-ge-WANT) +FAIL: test_fewer_repositories_than_the_pod_lists_refuses (__main__.Push.test_fewer_repositories_than_the_pod_lists_refuses) +Ran 1 test in 0.271s +FAILED (failures=1) +### RED: fsck (removed: fsck) +FAIL: test_a_broken_repository_fails_git_fsck (__main__.RestoreTest.test_a_broken_repository_fails_git_fsck) +Ran 1 test in 0.462s +FAILED (failures=1) +### RED: manifest (removed: sha256-check) +FAIL: test_a_changed_file_fails_the_manifest (__main__.RestoreTest.test_a_changed_file_fails_the_manifest) +Ran 1 test in 0.457s +FAILED (failures=1) +### GREEN after restore +Ran 19 tests in 34.596s +OK +### RED: git reads the repo config (old fsck line) +FAIL: test_git_never_reads_a_repositorys_own_config (__main__.RestoreTest.test_git_never_reads_a_repositorys_own_config) +AssertionError: 1 != 0 : felhom-dooplex-offsite-restore-test: FAILED: git fsck admin/felhom.eu.git: fatal: Expected git repo version <= 1, found 99 +Ran 1 test in 0.441s +FAILED (failures=1) +### GREEN, full suite +Ran 20 tests in 35.065s +OK diff --git a/documentation/audits/dooplex-survival-2026-10-09/partE/R-518.txt b/documentation/audits/dooplex-survival-2026-10-09/partE/R-518.txt new file mode 100644 index 00000000..b37a7e40 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partE/R-518.txt @@ -0,0 +1,151 @@ +# R-518 read-back — the first night with BOTH tiers due on demo-hp. Read-only, collected from DooPlex 2026-10-09T08:02:41Z. +# demo-hp host journal = CEST (UTC+2). Controller metrics, docker, PVE UPID and ep0 = UTC. + +################ CH1 — demo-hp agent journal (felhom-agent), backup lines 2026-10-08 22:15-22:45 local (= 20:15-20:45Z) +Oct 08 22:20:43 demo-hp felhom-agent[666190]: time=2026-10-08T22:20:43.713+02:00 level=INFO msg="backup: space preflight passed" vmid=9201 target=local last_archive_bytes=4885939687 need_bytes=7181166432 avail_bytes=17724211200 +Oct 08 22:20:44 demo-hp felhom-agent[666190]: time=2026-10-08T22:20:44.762+02:00 level=INFO msg="local-api: backup reached snapshotted (app may resume)" vmid=9201 target=local job=backup-9201-1791490843121326096 +Oct 08 22:25:40 demo-hp felhom-agent[666190]: time=2026-10-08T22:25:40.590+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_08-22_20_43.tar.zst size_bytes=4887664196 uncovered_volumes=2 +Oct 08 22:25:40 demo-hp felhom-agent[666190]: time=2026-10-08T22:25:40.590+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1791490843121326096 archive=local:backup/vzdump-lxc-9201-2026_10_08-22_20_43.tar.zst +Oct 08 22:26:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:26:14.221+02:00 level=INFO msg="local-api: backup refused — a heavy operation is already in flight" vmid=9201 requested_target=felhom-pbs busy=backup:local +Oct 08 22:27:10 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:10.591+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=guest vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z +Oct 08 22:27:11 demo-hp felhom-os-apply[1209621]: os-apply: START release=ring0-20261008T202710Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0 +Oct 08 22:27:16 demo-hp felhom-os-apply[1209908]: os-apply: REPAIR configured=0 journal=0 fixed=0 +Oct 08 22:27:26 demo-hp felhom-os-apply[1210615]: os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0 +Oct 08 22:27:37 demo-hp felhom-os-apply[1211512]: os-apply: DONE rc=0 seconds=4.5 upgraded=2 restart-needed=- reboot-needed=no +Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0" +Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0" +Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0" +Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=4.5 upgraded=2 restart-needed=- reboot-needed=no" +Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.958+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=guest vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=2 pending=1 not_covered=1 restart_needed=0 reboot_needed=false wrapper_seconds=35.3 +Oct 08 22:27:46 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:46.035+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=host vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z +Oct 08 22:27:47 demo-hp felhom-os-apply[1211951]: os-apply: START release=ring0-20261008T202710Z layer=host lane=fast mode=apply select=pending-fast packages=0 +Oct 08 22:27:51 demo-hp felhom-os-apply[1212307]: os-apply: REPAIR configured=0 journal=0 fixed=0 +Oct 08 22:27:56 demo-hp felhom-os-apply[1212564]: os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0 +Oct 08 22:28:02 demo-hp felhom-os-apply[1213497]: os-apply: DONE rc=0 seconds=3.5 upgraded=2 restart-needed=kvm,pmxcfs,pve-firewall,pve-ha-crm,pve-ha-lrm,pvedaemon,pvedaemon worke,pveproxy,pveproxy worker,pvescheduler,pvestatd,rrdcached,spiceproxy,spiceproxy work reboot-needed=no +Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=host lane=fast mode=apply select=pending-fast packages=0" +Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0" +Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0" +Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=3.5 upgraded=2 restart-needed=kvm,pmxcfs,pve-firewall,pve-ha-crm,pve-ha-lrm,pvedaemon,pvedaemon worke,pveproxy,pveproxy worker,pvescheduler,pvestatd,rrdcached,spiceproxy,spiceproxy work reboot-needed=no" +Oct 08 22:28:12 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:12.164+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=host vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=2 pending=80 not_covered=80 restart_needed=14 reboot_needed=false wrapper_seconds=25 +Oct 08 22:28:14 demo-hp felhom-os-apply[1214081]: os-apply: LIVE-RESTORE already on +Oct 08 22:28:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:14.220+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: LIVE-RESTORE already on" +Oct 08 22:28:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:14.220+02:00 level=INFO msg="osupdate: live-restore" vmid=9201 result="{\"result\": \"already on\"}" +Oct 08 22:28:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:14.220+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=docker vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z +Oct 08 22:28:16 demo-hp felhom-os-apply[1214207]: os-apply: START release=ring0-20261008T202710Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0 +Oct 08 22:28:21 demo-hp felhom-os-apply[1214648]: os-apply: REPAIR configured=0 journal=0 fixed=0 +Oct 08 22:28:27 demo-hp felhom-os-apply[1214956]: os-apply: PLAN upgrade=1 already=0 not-installed=0 from-snapshot=0 +Oct 08 22:29:01 demo-hp felhom-os-apply[1217263]: os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket) +Oct 08 22:29:03 demo-hp felhom-os-apply[1218011]: os-apply: DONE rc=0 seconds=5.1 upgraded=1 restart-needed=- reboot-needed=no +Oct 08 22:29:21 demo-hp felhom-os-apply[1219347]: os-apply: OOM-CHECK result=pass oom_killed=True oom_event=True exit=137 image=gitea.dooplex.hu/admin/felhom-controller:0.303.0 — the engine reported the memory kill: OOMKilled=true and the oom event +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0" +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0" +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=0 not-installed=0 from-snapshot=0" +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket)" +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=5.1 upgraded=1 restart-needed=- reboot-needed=no" +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: OOM-CHECK result=pass oom_killed=True oom_event=True exit=137 image=gitea.dooplex.hu/admin/felhom-controller:0.303.0 — the engine reported the memory kill: OOMKilled=true and the oom event" +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.330+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=docker vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=1 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=67 +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.342+02:00 level=INFO msg="osupdate: pve step holds the /etc/pve write gate — the agent's own writes wait until it ends" run=20261008T202710Z layer=pve vmid=9201 trigger=night +Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.342+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=pve vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z +Oct 08 22:29:22 demo-hp felhom-os-apply[1219396]: os-apply: START release=ring0-20261008T202710Z layer=pve:9201 lane=slow mode=apply select=pending-pve packages=0 authority=ring0 +Oct 08 22:29:26 demo-hp felhom-os-apply[1219654]: os-apply: REPAIR configured=0 journal=0 fixed=0 +Oct 08 22:29:30 demo-hp felhom-os-apply[1219888]: os-apply: PLAN upgrade=70 already=0 not-installed=0 from-snapshot=0 +Oct 08 22:29:33 demo-hp felhom-os-apply[1220017]: os-apply: REFUSED: R6 the plan would touch shim-signed-common, which is not in the plan +Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=pve:9201 lane=slow mode=apply select=pending-pve packages=0 authority=ring0" +Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0" +Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=70 already=0 not-installed=0 from-snapshot=0" +Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REFUSED: R6 the plan would touch shim-signed-common, which is not in the plan" +Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.403+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=pve vmid=9201 ring=0 trigger=night outcome=refused healthy=false reason="" upgraded=0 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=0 +Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.782+02:00 level=INFO msg="osupdate: pve step released the /etc/pve write gate" run=20261008T202710Z layer=pve vmid=9201 trigger=night +Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.783+02:00 level=WARN msg="osupdate: kernel step skipped — the Proxmox step did not end healthy" run=20261008T202710Z vmid=9201 trigger=night pve_outcome=refused +Oct 08 22:34:28 demo-hp felhom-agent[666190]: time=2026-10-08T22:34:28.003+02:00 level=INFO msg="local-api: backup reached snapshotted (app may resume)" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1791491666431534232 +Oct 08 22:37:38 demo-hp felhom-agent[666190]: time=2026-10-08T22:37:38.865+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-10-08T20:34:27Z size_bytes=15886927313 uncovered_volumes=2 +Oct 08 22:37:38 demo-hp felhom-agent[666190]: time=2026-10-08T22:37:38.865+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1791491666431534232 archive=felhom-pbs:backup/ct/9201/2026-10-08T20:34:27Z +== vzdump task log +2026/10/09 05:05:43 main.go:3684: [INFO] local-api: mount mp8 → /mnt/felhom-drives (storage=/mnt/felhom-drives, class=, backup=false) + +################ CH1 context — tier cadences as armed (agent journal, latest start) +Oct 09 07:16:47 demo-hp felhom-agent[2730121]: time=2026-10-09T07:16:47.622+02:00 level=INFO msg="backup tier armed" target=local cadence=24h0m0s keep_last=1 wait_timeout=30m0s prune_pbs_allowed=false primary=true +Oct 09 07:16:47 demo-hp felhom-agent[2730121]: time=2026-10-09T07:16:47.622+02:00 level=INFO msg="backup tier armed" target=felhom-pbs cadence=168h0m0s keep_last=0 wait_timeout=12h0m0s prune_pbs_allowed=false primary=false + +################ CH1b — PVE task index on demo-hp (pvedaemon task log; a different writer from the agent): vzdump tasks 2026-10-07..09 +UPID:demo-hp:0031920A:00EA3C89:6AC5AFDF:vzdump:9201:felhom-agent@pve!agent: start=2026-10-07T02:35:11Z end=2026-10-07T02:40:45Z OK +UPID:demo-hp:003C6995:01011C3B:6AC5EA6D:vzdump:9201:felhom-agent@pve!agent: start=2026-10-07T06:45:01Z end=2026-10-07T06:49:49Z OK +UPID:demo-hp:00120671:00AABC2D:6AC7FB1B:vzdump:9201:felhom-agent@pve!agent: start=2026-10-08T20:20:43Z end=2026-10-08T20:25:39Z OK +UPID:demo-hp:0012D84F:00ABFDC1:6AC7FE52:vzdump:9201:felhom-agent@pve!agent: start=2026-10-08T20:34:26Z end=2026-10-08T20:37:35Z OK +-- local vzdump log of the local tier +2026-10-08 22:20:43 INFO: Starting Backup of VM 9201 (lxc) +2026-10-08 22:20:43 INFO: status = running +2026-10-08 22:20:43 INFO: CT Name: demo-hp +2026-10-08 22:20:43 INFO: including mount point rootfs ('/') in backup +2026-10-08 22:20:43 INFO: including mount point mp0 ('/var/lib/felhom') in backup +2026-10-08 22:20:43 INFO: excluding bind mount point mp8 ('/mnt/felhom-drives') from backup (not a volume) +2026-10-08 22:20:43 INFO: excluding bind mount point mp9 ('/etc/felhom-bootstrap') from backup (not a volume) +2026-10-08 22:20:43 INFO: backup mode: snapshot +2026-10-08 22:20:43 INFO: ionice priority: 7 +2026-10-08 22:20:43 INFO: suspend vm to make snapshot +2026-10-08 22:20:43 INFO: create storage snapshot 'vzdump' +2026-10-08 22:20:45 INFO: resume vm +2026-10-08 22:20:45 INFO: guest is online again after 2 seconds +2026-10-08 22:20:45 INFO: creating vzdump archive '/var/lib/vz/dump/vzdump-lxc-9201-2026_10_08-22_20_43.tar.zst' +2026-10-08 22:25:33 INFO: Total bytes written: 16625551360 (16GiB, 56MiB/s) +2026-10-08 22:25:33 INFO: archive file size: 4.55GB +2026-10-08 22:25:33 INFO: adding notes to backup +2026-10-08 22:25:33 INFO: prune older backups with retention: keep-last=1 +2026-10-08 22:25:33 INFO: removing backup 'local:backup/vzdump-lxc-9201-2026_10_07-08_45_01.tar.zst' +2026-10-08 22:25:33 INFO: pruned 1 backup(s) not covered by keep-retention policy +2026-10-08 22:25:35 INFO: cleanup temporary 'vzdump' snapshot +2026-10-08 22:25:38 INFO: Finished Backup of VM 9201 (00:04:55) + +################ CH2 — controller metrics.db in guest 9201 (one row per RUNNING container per ~60 s sample; a stopped container has no row) +# query: per sample, running count and the containers present at 20:14-20:20:30Z that are missing. Opened ?mode=ro via the controller container's sqlite3. +2026-10-08T20:14:19Z|21| +2026-10-08T20:15:19Z|21| +2026-10-08T20:16:19Z|21| +2026-10-08T20:17:19Z|21| +2026-10-08T20:18:19Z|21| +2026-10-08T20:19:19Z|21| +2026-10-08T20:20:19Z|21| +2026-10-08T20:21:21Z|15|privatebin,paperless-webserver,paperless-postgres,paperless-redis,opengist,kimai +2026-10-08T20:22:19Z|21| +2026-10-08T20:23:19Z|21| +2026-10-08T20:24:19Z|21| +2026-10-08T20:25:19Z|21| +2026-10-08T20:26:19Z|6|privatebin,paperless-webserver,paperless-postgres,paperless-redis,opengist,kimai,kimai-db,docmost,docmost-postgres,docmost-redis,calibre-web,bookstack,bookstack-db,adventurelog,bentopdf +2026-10-08T20:27:19Z|21| +2026-10-08T20:28:19Z|21| +2026-10-08T20:29:04Z|21| +2026-10-08T20:30:04Z|21| +2026-10-08T20:31:04Z|21| +2026-10-08T20:32:04Z|21| +2026-10-08T20:33:04Z|21| +2026-10-08T20:34:04Z|21| +2026-10-08T20:35:04Z|15|privatebin,paperless-webserver,paperless-postgres,paperless-redis,opengist,kimai +2026-10-08T20:36:04Z|21| +2026-10-08T20:37:04Z|21| +2026-10-08T20:38:04Z|21| +2026-10-08T20:39:04Z|21| +2026-10-08T20:40:04Z|21| +2026-10-08T20:41:04Z|21| +total distinct containers 20:14-20:20:30Z: 21 + +################ CH2b — guest docker: container start/create times (bentopdf recreated 15 s after the off-site snapshot) +/bentopdf started=2026-10-08T20:34:43.277725548Z created=2026-10-08T20:34:43.157907984Z +/traefik started=2026-10-08T20:29:01.412754011Z created=2026-10-04T09:01:48.259555986Z +/bookstack started=2026-10-09T02:15:27.95387884Z created=2026-10-09T02:15:22.165968876Z +/felhom-controller started=2026-10-09T05:05:42.178985741Z created=2026-10-09T05:05:42.122744205Z +# all other app containers were recreated 2026-10-09T02:15-02:16Z (after the window), so their StartedAt cannot show the 20:2x stops. + +################ CH3 — ep0 PBS datastore, namespace demo-hp, ct/9201 (ls, read-only) +felhom-hetzner +2026-10-09T08:02:47Z +total 20 +drwxr-xr-x 4 backup backup 4096 2026-10-09 03:29:59.997681474 +0000 . +drwxr-xr-x 3 backup backup 4096 2026-07-26 15:42:44.833483940 +0000 .. +drwxr-xr-x 2 backup backup 4096 2026-10-09 05:17:18.883235046 +0000 2026-10-01T20:15:29Z +drwxr-xr-x 2 backup backup 4096 2026-10-09 05:17:15.547227803 +0000 2026-10-08T20:34:27Z +-rw-r--r-- 1 backup backup 19 2026-07-26 15:42:44.833483940 +0000 owner + +################ What could NOT be read (tried) +# controller docker log for the night: 'docker logs felhom-controller' starts 2026-10-09 05:05:43 — the 0.304.0 swap replaced the container. +# controller persisted debug ring (data/debug-ring.log): 5000 lines, oldest 2026-10-09T07:30:53Z — does not reach the night. +# guest docker events: 'docker events --since 2026-10-08T20:15Z --until 20:45Z' returned nothing; the event buffer's oldest entry is 2026-10-09T07:58Z (healthcheck execs fill it). diff --git a/documentation/audits/dooplex-survival-2026-10-09/partE/R-861.txt b/documentation/audits/dooplex-survival-2026-10-09/partE/R-861.txt new file mode 100644 index 00000000..1cdf3fb6 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partE/R-861.txt @@ -0,0 +1,116 @@ +# R-861 (a) read-back — read-only, collected from DooPlex 2026-10-09T07:57:32Z. Host journals are CEST (UTC+2); docker StartedAt is UTC. +# Release context: felhom.eu/documentation/audits/release-2026-10-09/delivery/floors-0.304.0.txt (floors set 2026-10-09T05:05:30Z) + +################ demo-hp +demo-hp +felhom-agent 0.154.0 +== CH1a agent journal 06:50-07:40 local (swap lines) +Oct 09 07:05:37 demo-hp felhom-agent[666190]: time=2026-10-09T07:05:37.552+02:00 level=WARN msg="local-api: controller swap requested" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 previous=gitea.dooplex.hu/admin/felhom-controller:0.303.0 +Oct 09 07:05:39 demo-hp sudo[2696732]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:05:40 demo-hp felhom-priv-apply[2696817]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:05:40 demo-hp felhom-agent[666190]: time=2026-10-09T07:05:40.585+02:00 level=INFO msg="controller-swap: image file written, restarting bootstrap" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:05:47 demo-hp felhom-agent[666190]: time=2026-10-09T07:05:47.706+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:16:49 demo-hp sudo[2730461]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config +Oct 09 07:16:49 demo-hp felhom-priv-apply[2730492]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config +Oct 09 07:16:54 demo-hp sudo[2730994]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key +Oct 09 07:16:54 demo-hp felhom-priv-apply[2731003]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op +Oct 09 07:32:01 demo-hp felhom-agent[2730121]: time=2026-10-09T07:32:01.215+02:00 level=WARN msg="osupdate: config bundle INSTALLED" op=agent_config_update agent_version=0.154.0 bundle="{\"agent_version\": \"0.154.0\", \"authority\": \"signed\", \"kept\": [\"/etc/felhom/crash-guard.conf\"], \"prev_dir\": \"/var/lib/felhom-os-apply/bundle-prev/20261009T053200Z-before-0.154.0\", \"same\": 26, \"self_check\": {\"crash_guard\": {\"armed\": true, \"kernel_panic\": 10}, \"os_apply\": \"felhom-os-apply ok bundle-format=1 files=28\", \"priv_apply\": \"felhom-priv-apply ok verbs=unit,dnsmasq,wg,sshd-config,sshd-key,controller-image\", \"selfupdate\": \"usage ok\", \"sudo_l\": \"felhom-os-apply listed\", \"visudo\": \"ok\"}, \"sha256\": \"93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0\", \"signers_created\": false, \"skipped\": [], \"written\": [\"/usr/local/sbin/felhom-os-apply\"]}" pass_seconds=1 +== CH1b sudo log 2026-10-09 06:50-07:40 local +Oct 09 07:05:36 demo-hp sudo[2696598]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image +Oct 09 07:05:37 demo-hp sudo[2696651]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image +Oct 09 07:05:39 demo-hp sudo[2696732]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:16:49 demo-hp sudo[2730461]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config +Oct 09 07:16:54 demo-hp sudo[2730994]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key +== CH1c priv-apply wrapper log 2026-10-09 +Oct 09 07:05:40 demo-hp felhom-priv-apply[2696817]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0 +== tee route since 2026-10-07 00:00 (count, then lines) +2 +Oct 07 09:09:27 demo-hp sudo[4038406]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image +Oct 07 13:36:34 demo-hp sudo[102535]: root : PWD=/root ; USER=felhom-agent ; COMMAND=/usr/bin/sudo -n /usr/sbin/pct exec 9202 -- tee /etc/felhom-controller-image +== tee route on 2026-10-09 (count) +0 +== verb runs since 2026-10-07 +Oct 07 13:36:33 demo-hp sudo[102481]: root : PWD=/root ; USER=felhom-agent ; COMMAND=/usr/bin/sudo -n /usr/local/sbin/felhom-priv-apply controller-image 9202 +Oct 07 13:36:33 demo-hp sudo[102483]: felhom-agent : PWD=/root ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9202 +Oct 07 18:50:13 demo-hp sudo[631922]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:05:39 demo-hp sudo[2696732]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +== CH2 guest 9201 (docker in the guest) +gitea.dooplex.hu/admin/felhom-controller:0.304.0 +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 3 hours (healthy) +/felhom-controller image=gitea.dooplex.hu/admin/felhom-controller:0.304.0 started=2026-10-09T05:05:42.178985741Z restarts=0 + +################ felhom-pve +demo-felhom +felhom-agent 0.154.0 +== CH1a agent journal 06:50-07:40 local (swap lines) +Oct 09 07:05:35 demo-felhom felhom-agent[1202]: time=2026-10-09T07:05:35.089+02:00 level=WARN msg="local-api: controller swap requested" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 previous=gitea.dooplex.hu/admin/felhom-controller:0.303.0 +Oct 09 07:05:36 demo-felhom sudo[130104]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:05:37 demo-felhom felhom-priv-apply[130111]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:05:37 demo-felhom felhom-agent[1202]: time=2026-10-09T07:05:37.382+02:00 level=INFO msg="controller-swap: image file written, restarting bootstrap" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:05:46 demo-felhom felhom-agent[1202]: time=2026-10-09T07:05:46.966+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:16:07 demo-felhom sudo[141297]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config +Oct 09 07:16:07 demo-felhom felhom-priv-apply[141326]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config +Oct 09 07:16:11 demo-felhom sudo[141738]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key +Oct 09 07:16:11 demo-felhom felhom-priv-apply[141741]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op +Oct 09 07:31:16 demo-felhom felhom-agent[141055]: time=2026-10-09T07:31:16.948+02:00 level=WARN msg="osupdate: config bundle INSTALLED" op=agent_config_update agent_version=0.154.0 bundle="{\"agent_version\": \"0.154.0\", \"authority\": \"signed\", \"kept\": [\"/etc/felhom/crash-guard.conf\"], \"prev_dir\": \"/var/lib/felhom-os-apply/bundle-prev/20261009T053116Z-before-0.154.0\", \"same\": 26, \"self_check\": {\"crash_guard\": {\"armed\": true, \"kernel_panic\": 10}, \"os_apply\": \"felhom-os-apply ok bundle-format=1 files=28\", \"priv_apply\": \"felhom-priv-apply ok verbs=unit,dnsmasq,wg,sshd-config,sshd-key,controller-image\", \"selfupdate\": \"usage ok\", \"sudo_l\": \"felhom-os-apply listed\", \"visudo\": \"ok\"}, \"sha256\": \"93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0\", \"signers_created\": false, \"skipped\": [], \"written\": [\"/usr/local/sbin/felhom-os-apply\"]}" pass_seconds=0.6 +== CH1b sudo log 2026-10-09 06:50-07:40 local +Oct 09 07:05:34 demo-felhom sudo[130062]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image +Oct 09 07:05:35 demo-felhom sudo[130074]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image +Oct 09 07:05:36 demo-felhom sudo[130104]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:16:07 demo-felhom sudo[141297]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config +Oct 09 07:16:11 demo-felhom sudo[141738]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key +== CH1c priv-apply wrapper log 2026-10-09 +Oct 09 07:05:37 demo-felhom felhom-priv-apply[130111]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0 +== tee route since 2026-10-07 00:00 (count, then lines) +1 +Oct 07 09:09:26 demo-felhom sudo[2302098]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image +== tee route on 2026-10-09 (count) +0 +== verb runs since 2026-10-07 +Oct 07 18:50:12 demo-felhom sudo[199002]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:05:36 demo-felhom sudo[130104]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +== CH2 guest 9201 (docker in the guest) +gitea.dooplex.hu/admin/felhom-controller:0.304.0 +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 3 hours (healthy) +/felhom-controller image=gitea.dooplex.hu/admin/felhom-controller:0.304.0 started=2026-10-09T05:05:38.555475195Z restarts=0 + +################ root@192.168.0.154 +felhom +felhom-agent 0.154.0 +== CH1a agent journal 06:50-07:40 local (swap lines) +Oct 09 07:05:37 felhom felhom-agent[144633]: time=2026-10-09T07:05:37.458+02:00 level=WARN msg="local-api: controller swap requested" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 previous=gitea.dooplex.hu/admin/felhom-controller:0.303.0 +Oct 09 07:05:39 felhom sudo[2992327]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:05:40 felhom felhom-priv-apply[2992353]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:05:40 felhom felhom-agent[144633]: time=2026-10-09T07:05:40.580+02:00 level=INFO msg="controller-swap: image file written, restarting bootstrap" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:05:47 felhom felhom-agent[144633]: time=2026-10-09T07:05:47.589+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 +Oct 09 07:18:20 felhom sudo[3010887]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config +Oct 09 07:18:20 felhom felhom-priv-apply[3010891]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config +Oct 09 07:18:24 felhom sudo[3011416]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key +Oct 09 07:18:24 felhom felhom-priv-apply[3011425]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op +Oct 09 07:33:30 felhom felhom-agent[3010618]: time=2026-10-09T07:33:30.409+02:00 level=WARN msg="osupdate: config bundle INSTALLED" op=agent_config_update agent_version=0.154.0 bundle="{\"agent_version\": \"0.154.0\", \"authority\": \"signed\", \"kept\": [\"/etc/felhom/crash-guard.conf\"], \"prev_dir\": \"/var/lib/felhom-os-apply/bundle-prev/20261009T053329Z-before-0.154.0\", \"same\": 26, \"self_check\": {\"crash_guard\": {\"armed\": true, \"kernel_panic\": 10}, \"os_apply\": \"felhom-os-apply ok bundle-format=1 files=28\", \"priv_apply\": \"felhom-priv-apply ok verbs=unit,dnsmasq,wg,sshd-config,sshd-key,controller-image\", \"selfupdate\": \"usage ok\", \"sudo_l\": \"felhom-os-apply listed\", \"visudo\": \"ok\"}, \"sha256\": \"93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0\", \"signers_created\": false, \"skipped\": [], \"written\": [\"/usr/local/sbin/felhom-os-apply\"]}" pass_seconds=0.9 +== CH1b sudo log 2026-10-09 06:50-07:40 local +Oct 09 07:05:36 felhom sudo[2992261]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image +Oct 09 07:05:37 felhom sudo[2992286]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image +Oct 09 07:05:39 felhom sudo[2992327]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:18:20 felhom sudo[3010887]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config +Oct 09 07:18:24 felhom sudo[3011416]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key +== CH1c priv-apply wrapper log 2026-10-09 +Oct 09 07:05:40 felhom felhom-priv-apply[2992353]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0 +== tee route since 2026-10-07 00:00 (count, then lines) +1 +Oct 07 09:09:27 felhom sudo[3745733]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image +== tee route on 2026-10-09 (count) +0 +== verb runs since 2026-10-07 +Oct 07 18:50:13 felhom sudo[126979]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +Oct 09 07:05:39 felhom sudo[2992327]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +== CH2 guest 9201 (docker in the guest) +gitea.dooplex.hu/admin/felhom-controller:0.304.0 +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 2 hours (healthy) +/felhom-controller image=gitea.dooplex.hu/admin/felhom-controller:0.304.0 started=2026-10-09T05:58:31.221358279Z restarts=0 + +################ note: Tester 1 controller StartedAt 05:58:31Z (not 05:05:47Z) +# The Tester 1 VM rebooted at 07:58:11 local (kernel line "Linux version 7.0.14-22-pve" at Oct 09 07:58:11; +# agent "daemon starting version=0.154.0" at 07:58:20; os-apply plan "plan-boot-...-kernel-kernel-boot.json"). +# The guest controller log's first lines after that: "2026/10/09 05:58:31 [INFO] felhom-controller 0.304.0 starting (customer: tester-1 ...)". +# No second controller-image verb run and no tee on 2026-10-09 (counts above). The later StartedAt is the reboot, not a second swap. diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index 0df0c6fc..9a0a8f1a 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,19 @@ +## 2026-10-09 — dooplex-offsite: Gitea + DooPlex's secrets leave DooPlex nightly, encrypted (R-232 (b), (h)) + +- New `scripts/dooplex-offsite/`: `felhom-dooplex-offsite` (daily 00:20) copies the newest complete `gitea.dump` + (taken FIRST), then Gitea's repositories, LFS, `app.ini`, attachments, avatars and `jwt` out of the pod (NOT the + 27.7 GB registry), then the newest night's three GPG secrets files, writes `MANIFEST.sha256` + `REPOS`, and pushes one + encrypted `dooplex.pxar` to ep0 (`operator` ns, `host/dooplex-gitea`) with the hub-DB's write-only token and a NEW key. + Refuses (no push, no success signal) on: no complete dump / dump > 7 h; file copy failing twice; any symlink or + hardlink in the pod's archive; fewer repositories than the pod lists; no `app.ini`; secrets > 30 h. +- `felhom-dooplex-offsite-restore-test` (Sun 05:30, read-only token): manifest check, repo count, `git fsck` on every + repository through a scratch repo with a known config (the pod's `config` files are never read by root git), + `pg_restore --list`, secrets present. +- `felhom-backup-failmail@.service`: `OnFailure=` mail through R-232 (a)'s `notify_failure`; also added to both + hub-DB units (`scripts/hub-db-backup/*.service`). +- 20 tests (`test_dooplex_offsite.py`; fakes for kubectl/PBS/pg_restore, real git/tar), green with GNU and BusyBox + tools; red-proofs for 6 checks in `audits/dooplex-survival-2026-10-09/partB/red-proof.txt`. + ## 2026-10-09 — ep0-copy-gc reads ep0's real namespace answer (found on the first live dry run) - `scripts/ep0-copy-gc/felhom-ep0-copy-gc`: `proxmox-backup-client namespace list --output-format json` answers diff --git a/scripts/dooplex-offsite/felhom-backup-failmail b/scripts/dooplex-offsite/felhom-backup-failmail new file mode 100755 index 00000000..f41753f5 --- /dev/null +++ b/scripts/dooplex-offsite/felhom-backup-failmail @@ -0,0 +1,10 @@ +#!/bin/bash +# felhom-backup-failmail — mail admin@ that a Felhom backup unit on DooPlex failed (R-232 (a) route, reused). +# Called by felhom-backup-failmail@.service, which the backup units name in OnFailure=. +# Reuses notify_failure from /opt/backup/scripts/backup-config.sh (Resend; never prints the key; never fails). +set -u +UNIT=${1:?unit name} +CONFIG=${FELHOM_FAILMAIL_CONFIG:-/opt/backup/scripts/backup-config.sh} +# shellcheck source=/dev/null +. "$CONFIG" +notify_failure "systemd unit ${UNIT} failed on $(hostname) — see: journalctl -u ${UNIT}" diff --git a/scripts/dooplex-offsite/felhom-backup-failmail@.service b/scripts/dooplex-offsite/felhom-backup-failmail@.service new file mode 100644 index 00000000..449fcedc --- /dev/null +++ b/scripts/dooplex-offsite/felhom-backup-failmail@.service @@ -0,0 +1,8 @@ +# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh. +# Named by OnFailure= in felhom-dooplex-offsite*.service and felhom-hub-db-*.service; %i is the failed unit. +[Unit] +Description=Felhom: mail admin@ that %i failed (R-232) + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/felhom-backup-failmail %i diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite b/scripts/dooplex-offsite/felhom-dooplex-offsite new file mode 100755 index 00000000..58e6bc35 --- /dev/null +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite @@ -0,0 +1,109 @@ +#!/bin/sh +# felhom-dooplex-offsite — push Gitea (repositories, database dump, config) and DooPlex's nightly secrets export to +# ep0's PBS, encrypted on DooPlex (R-232 (b)). Runs on DooPlex as root from felhom-dooplex-offsite.timer (00:20). +# Plan: documentation/audits/dooplex-survival-2026-10-09/PLAN.md. Restore: documentation/runbooks/gitea-restore.md. +# Pinned by test_dooplex_offsite.py. +# +# Order is the consistency argument (PLAN.md): the database dump is taken FIRST (the newest complete 6-hourly dump, +# PostgreSQL's own snapshot), the repositories AFTER it, so every commit the database names is in the copy. +# +# Refuses to push — and so never writes the success signal — when: no complete dump exists or the newest is older than +# DUMP_MAX_AGE_H; the dump is empty; the file copy from the pod fails twice; the copy holds no repository or fewer +# repositories than the pod lists; app.ini is missing; no secrets export exists or it is older than SECRETS_MAX_AGE_H. +# The success timestamp is written ONLY after the push returns 0 (CLAUDE.md "presence is not success"). +set -eu +CONF=${FELHOM_DXOFF_CONF:-/etc/felhom-dooplex-offsite} +TOKENS=${FELHOM_DXOFF_TOKENS:-/etc/felhom-hub-backup} +STATE=${FELHOM_DXOFF_STATE:-/var/lib/felhom-dooplex-offsite} +TEXTFILE_DIR=${FELHOM_DXOFF_TEXTFILE_DIR:-/var/lib/node_exporter/textfile_collector} +DUMPS=${FELHOM_DXOFF_DUMPS:-/mnt/5_hdd/backup/postgresql/dumps} +SECRETS=${FELHOM_DXOFF_SECRETS:-/mnt/5_hdd/backup/secrets/exports} +DUMP_MAX_AGE_H=${FELHOM_DXOFF_DUMP_MAX_AGE_H:-7} +SECRETS_MAX_AGE_H=${FELHOM_DXOFF_SECRETS_MAX_AGE_H:-30} +NOW=${FELHOM_DXOFF_NOW:-$(date +%s)} +RETRY_SLEEP=${FELHOM_DXOFF_RETRY_SLEEP:-30} +. "$CONF/env" # PBS_REPOSITORY_PUSH, PBS_FINGERPRINT (no secrets in this file) + +log() { echo "felhom-dooplex-offsite: $*"; } +die() { echo "felhom-dooplex-offsite: FAILED: $*" >&2; exit 1; } +K() { kubectl -n gitea-system exec deploy/gitea -c gitea -- "$@"; } + +umask 077 +STAGE="$STATE/stage" +mkdir -p "$STATE"; chmod 700 "$STATE" +rm -rf "$STAGE"; mkdir -p "$STAGE/root/gitea" "$STAGE/root/db" "$STAGE/root/secrets" +cleanup() { + for f in "$STAGE/root/gitea/gitea/conf/app.ini" "$STAGE/root/db/gitea.dump" "$STAGE/root/db/globals.sql"; do + [ -f "$f" ] && [ ! -L "$f" ] && { shred -u "$f" 2>/dev/null || rm -f "$f"; } + done + rm -rf "$STAGE" +} +trap cleanup EXIT + +# 1. the database — the newest COMPLETE dump (its folder carries SUCCESS), taken before the files +DUMPDIR="" +for d in $(ls -1d "$DUMPS"/[0-9]*-[0-9]* 2>/dev/null | sort -r); do + if [ -f "$d/SUCCESS" ] && [ -s "$d/gitea.dump" ]; then DUMPDIR=$d; break; fi +done +[ -n "$DUMPDIR" ] || die "no complete gitea.dump under $DUMPS" +DAGE=$((NOW - $(stat -c %Y "$DUMPDIR/SUCCESS"))) +[ "$DAGE" -le $((DUMP_MAX_AGE_H * 3600)) ] || die "newest complete dump ${DUMPDIR##*/} is $((DAGE / 3600)) h old (limit ${DUMP_MAX_AGE_H} h) — the dump CronJob stopped" +cp "$DUMPDIR/gitea.dump" "$STAGE/root/db/gitea.dump" +[ -f "$DUMPDIR/globals.sql" ] && cp "$DUMPDIR/globals.sql" "$STAGE/root/db/globals.sql" +echo "${DUMPDIR##*/}" > "$STAGE/root/db/DUMP-FOLDER" +log "database: ${DUMPDIR##*/}, $(wc -c < "$STAGE/root/db/gitea.dump" | tr -d ' ') bytes, $((DAGE / 60)) min old" + +# 2. the files — after the dump. Not the registry (packages: rebuilt from the code), logs, indexers, queues, tmp. +PATHS="git/repositories git/lfs gitea/conf/app.ini gitea/attachments gitea/avatars gitea/repo-avatars gitea/jwt" +n=0 +until K tar -cf - -C /data $PATHS > "$STAGE/files.tar"; do + n=$((n + 1)); [ "$n" -lt 2 ] || die "copying Gitea's files out of the pod failed twice" + log "file copy failed once (a file moved under a push?) — retrying in ${RETRY_SLEEP} s"; sleep "$RETRY_SLEEP" +done +# The pod's archive is untrusted input to a root process: refuse any symlink or hardlink in it (a bare repository holds +# none), and extract without the pod's owners. +LINKS=$(tar -tvf "$STAGE/files.tar" | grep -c '^[lh]') || LINKS=0 +[ "$LINKS" -eq 0 ] || die "the pod's archive holds $LINKS link(s) — refused" +tar -xof "$STAGE/files.tar" -C "$STAGE/root/gitea" || die "unpacking the file copy" +rm -f "$STAGE/files.tar" +[ -s "$STAGE/root/gitea/gitea/conf/app.ini" ] && [ ! -L "$STAGE/root/gitea/gitea/conf/app.ini" ] || die "app.ini missing from the copy" +LISTING=$(K find /data/git/repositories -mindepth 2 -maxdepth 2 -type d -name '*.git') || die "listing repositories in the pod" +WANT=$(printf '%s\n' "$LISTING" | grep -c '\.git$') || WANT=0 +GOT=$(find "$STAGE/root/gitea/git/repositories" -mindepth 2 -maxdepth 2 -type d -name '*.git' | wc -l | tr -d ' ') +[ "$GOT" -gt 0 ] || die "the copy holds no repository" +[ "$GOT" -ge "$WANT" ] || die "the copy holds $GOT repositories, the pod lists $WANT" +log "files: $GOT repositories, $(du -sm "$STAGE/root/gitea" | cut -f1) MB" + +# 3. the secrets — the newest night's GPG files (already encrypted with DooPlex's restic passphrase) +NEWEST=$(ls -1 "$SECRETS"/secrets-*.yaml.gpg 2>/dev/null | sort | tail -n 1) +[ -n "$NEWEST" ] || die "no secrets export under $SECRETS" +STAMP=${NEWEST##*/secrets-}; STAMP=${STAMP%.yaml.gpg} +SAGE=$((NOW - $(stat -c %Y "$NEWEST"))) +[ "$SAGE" -le $((SECRETS_MAX_AGE_H * 3600)) ] || die "newest secrets export is $((SAGE / 3600)) h old (limit ${SECRETS_MAX_AGE_H} h)" +cp "$SECRETS"/*-"$STAMP".*gpg "$STAGE/root/secrets/" +log "secrets: $(ls "$STAGE/root/secrets" | wc -l | tr -d ' ') file(s) of $STAMP" + +# 4. the manifest the restore test checks, then the push +echo "$GOT" > "$STAGE/root/REPOS" +(cd "$STAGE/root" && find . -type f ! -name MANIFEST.sha256 -print0 | sort -z | xargs -0 sha256sum > MANIFEST.sha256) \ + || die "writing the manifest" +[ "$(wc -l < "$STAGE/root/MANIFEST.sha256")" -gt "$GOT" ] || die "the manifest is short" +START=$(date +%s) +PBS_PASSWORD_FILE="$TOKENS/token-push" PBS_FINGERPRINT="$PBS_FINGERPRINT" \ + proxmox-backup-client backup dooplex.pxar:"$STAGE/root" --ns operator --backup-type host --backup-id dooplex-gitea \ + --keyfile "$CONF/enc.key" --crypt-mode encrypt --repository "$PBS_REPOSITORY_PUSH" \ + || die "proxmox-backup-client backup" +BYTES=$(du -sb "$STAGE/root" | cut -f1) +log "pushed to ep0 (ns operator, host/dooplex-gitea) in $(( $(date +%s) - START )) s" + +TMP="$TEXTFILE_DIR/felhom_dooplex_offsite.prom.$$" +{ + echo "# HELP felhom_dooplex_offsite_last_success_timestamp_seconds Last successful push of Gitea + DooPlex secrets to ep0 (R-232)." + echo "# TYPE felhom_dooplex_offsite_last_success_timestamp_seconds gauge" + echo "felhom_dooplex_offsite_last_success_timestamp_seconds $(date +%s)" + echo "felhom_dooplex_offsite_last_success_bytes $BYTES" + echo "felhom_dooplex_offsite_last_success_repositories $GOT" +} > "$TMP" +chmod 644 "$TMP" +mv "$TMP" "$TEXTFILE_DIR/felhom_dooplex_offsite.prom" +log "success signal written" diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test new file mode 100755 index 00000000..4c38e874 --- /dev/null +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test @@ -0,0 +1,79 @@ +#!/bin/sh +# felhom-dooplex-offsite-restore-test — restore the newest Gitea + secrets copy from ep0 with the READ-ONLY token and +# check it (R-232 (h)). Runs on DooPlex as root from felhom-dooplex-offsite-restore-test.timer (Sun 05:30). +# Pinned by test_dooplex_offsite.py. +# +# The success timestamp is written ONLY when: the newest copy on ep0 is at most MAX_AGE_H old; it restores and +# decrypts; every file matches MANIFEST.sha256; the repository count matches REPOS; `git fsck` passes on EVERY +# repository; `pg_restore --list` reads gitea.dump; app.ini is there; at least one secrets file is there. +set -eu +CONF=${FELHOM_DXOFF_CONF:-/etc/felhom-dooplex-offsite} +TOKENS=${FELHOM_DXOFF_TOKENS:-/etc/felhom-hub-backup} +STATE=${FELHOM_DXOFF_STATE:-/var/lib/felhom-dooplex-offsite} +TEXTFILE_DIR=${FELHOM_DXOFF_TEXTFILE_DIR:-/var/lib/node_exporter/textfile_collector} +MAX_AGE_H=${FELHOM_DXOFF_RESTORE_MAX_AGE_H:-50} +NOW=${FELHOM_DXOFF_NOW:-$(date +%s)} +. "$CONF/env" # PBS_REPOSITORY_RESTORE, PBS_FINGERPRINT + +log() { echo "felhom-dooplex-offsite-restore-test: $*"; } +die() { echo "felhom-dooplex-offsite-restore-test: FAILED: $*" >&2; exit 1; } + +umask 077 +mkdir -p "$STATE"; chmod 700 "$STATE" +T=$(mktemp -d "$STATE/restore.XXXXXX") +trap 'for f in "$T"/out/gitea/gitea/conf/app.ini "$T"/out/db/gitea.dump "$T"/out/db/globals.sql; do [ -f "$f" ] && shred -u "$f" 2>/dev/null; done; rm -rf "$T"' EXIT +export PBS_PASSWORD_FILE="$TOKENS/token-restore" PBS_FINGERPRINT + +LIST=$(proxmox-backup-client snapshot list host/dooplex-gitea --ns operator --output-format json --repository "$PBS_REPOSITORY_RESTORE") \ + || die "listing snapshots on ep0" +NEWEST=$(printf '%s' "$LIST" | python3 -c ' +import json, sys +s = [x for x in json.load(sys.stdin) if x.get("backup-type") == "host" and x.get("backup-id") == "dooplex-gitea"] +if s: + print(max(x["backup-time"] for x in s)) +') || die "reading the snapshot list" +[ -n "$NEWEST" ] || die "no Gitea copy on ep0" +AGE=$((NOW - NEWEST)) +[ "$AGE" -le $((MAX_AGE_H * 3600)) ] || die "newest copy on ep0 is $((AGE / 3600)) h old (limit ${MAX_AGE_H} h)" +SNAPSHOT="host/dooplex-gitea/$(date -u -d "@$NEWEST" +%Y-%m-%dT%H:%M:%SZ)" +log "restoring $SNAPSHOT" + +proxmox-backup-client restore "$SNAPSHOT" dooplex.pxar "$T/out" --ns operator \ + --keyfile "$CONF/enc.key" --repository "$PBS_REPOSITORY_RESTORE" || die "restore of $SNAPSHOT" +O="$T/out" +[ -s "$O/MANIFEST.sha256" ] || die "the copy holds no MANIFEST.sha256" +(cd "$O" && sha256sum -c MANIFEST.sha256 >/dev/null) || die "a file does not match MANIFEST.sha256" +FILES=$(wc -l < "$O/MANIFEST.sha256" | tr -d ' ') + +WANT=$(cat "$O/REPOS" 2>/dev/null) || die "the copy holds no REPOS count" +REPOS=$(find "$O/gitea/git/repositories" -mindepth 2 -maxdepth 2 -type d -name '*.git' | sort) +GOT=$(printf '%s\n' "$REPOS" | grep -c '\.git$') || GOT=0 +[ "$GOT" -gt 0 ] && [ "$GOT" = "$WANT" ] || die "the copy holds $GOT repositories, REPOS says $WANT" +# The repositories' own `config` files came from the Gitea pod — untrusted input to a root git. So git never reads +# them: each repository is checked through a fresh scratch repository with a known config, holding a copy of its +# HEAD and refs, with the copy's objects as its object store. No system or global config either. +export GIT_CONFIG_NOSYSTEM=1 GIT_CONFIG_GLOBAL=/dev/null +for r in $REPOS; do + S="$T/fsck.git"; rm -rf "$S" + git init -q --bare "$S" || die "git init for the fsck scratch repository" + cp "$r/HEAD" "$S/HEAD"; [ -f "$r/packed-refs" ] && cp "$r/packed-refs" "$S/packed-refs" + [ -d "$r/refs" ] && cp -R "$r/refs/." "$S/refs/" + GIT_OBJECT_DIRECTORY="$r/objects" git --git-dir="$S" -c core.hooksPath=/dev/null -c core.fsmonitor=false \ + fsck --no-progress --no-dangling >/dev/null 2>"$T/fsck.err" \ + || die "git fsck ${r#"$O"/gitea/git/repositories/}: $(head -n 3 "$T/fsck.err")" +done +rm -rf "$T/fsck.git" +[ -s "$O/gitea/gitea/conf/app.ini" ] || die "app.ini missing" +pg_restore --list "$O/db/gitea.dump" >/dev/null || die "pg_restore cannot read gitea.dump" +ls "$O"/secrets/*.gpg >/dev/null 2>&1 || die "no secrets file in the copy" +log "checked: $FILES files match the manifest, $GOT repositories pass git fsck, gitea.dump readable, $(ls "$O"/secrets | wc -l | tr -d ' ') secrets file(s)" + +TMP="$TEXTFILE_DIR/felhom_dooplex_offsite_restore.prom.$$" +{ + echo "# HELP felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds Last successful restore test of the Gitea + secrets copy on ep0 (R-232)." + echo "# TYPE felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds gauge" + echo "felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds $(date +%s)" +} > "$TMP" +chmod 644 "$TMP" +mv "$TMP" "$TEXTFILE_DIR/felhom_dooplex_offsite_restore.prom" +log "success signal written" diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test.service b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test.service new file mode 100644 index 00000000..47ab9c83 --- /dev/null +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test.service @@ -0,0 +1,20 @@ +# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh. Plan: audits/dooplex-survival-2026-10-09/PLAN.md. +[Unit] +Description=Felhom: restore-test the Gitea + secrets copy on ep0 (R-232) +Wants=network-online.target felhom-ep0-pbs-tunnel.service +After=network-online.target felhom-ep0-pbs-tunnel.service +OnFailure=felhom-backup-failmail@%n.service + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/felhom-dooplex-offsite-restore-test +Environment=HOME=/var/lib/felhom-dooplex-offsite KUBECONFIG=/etc/rancher/k3s/k3s.yaml +UMask=0077 +TimeoutStartSec=90min +Nice=10 +IOSchedulingClass=idle +PrivateTmp=yes +ProtectSystem=strict +ProtectHome=yes +ReadWritePaths=/var/lib/felhom-dooplex-offsite /var/lib/node_exporter/textfile_collector +NoNewPrivileges=yes diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test.timer b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test.timer new file mode 100644 index 00000000..35239033 --- /dev/null +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test.timer @@ -0,0 +1,11 @@ +# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh. +[Unit] +Description=Felhom: restore-test the Gitea + secrets copy on ep0 (R-232) — schedule + +[Timer] +OnCalendar=Sun *-*-* 05:30:00 +Persistent=true +RandomizedDelaySec=2min + +[Install] +WantedBy=timers.target diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite.service b/scripts/dooplex-offsite/felhom-dooplex-offsite.service new file mode 100644 index 00000000..5da6918b --- /dev/null +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite.service @@ -0,0 +1,20 @@ +# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh. Plan: audits/dooplex-survival-2026-10-09/PLAN.md. +[Unit] +Description=Felhom: push Gitea + DooPlex secrets to ep0, encrypted (R-232) +Wants=network-online.target felhom-ep0-pbs-tunnel.service +After=network-online.target felhom-ep0-pbs-tunnel.service +OnFailure=felhom-backup-failmail@%n.service + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/felhom-dooplex-offsite +Environment=HOME=/var/lib/felhom-dooplex-offsite KUBECONFIG=/etc/rancher/k3s/k3s.yaml +UMask=0077 +TimeoutStartSec=90min +Nice=10 +IOSchedulingClass=idle +PrivateTmp=yes +ProtectSystem=strict +ProtectHome=yes +ReadWritePaths=/var/lib/felhom-dooplex-offsite /var/lib/node_exporter/textfile_collector +NoNewPrivileges=yes diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite.timer b/scripts/dooplex-offsite/felhom-dooplex-offsite.timer new file mode 100644 index 00000000..0e3a9edc --- /dev/null +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite.timer @@ -0,0 +1,11 @@ +# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh. +[Unit] +Description=Felhom: push Gitea + DooPlex secrets to ep0 (R-232) — schedule + +[Timer] +OnCalendar=*-*-* 00:20:00 +Persistent=true +RandomizedDelaySec=2min + +[Install] +WantedBy=timers.target diff --git a/scripts/dooplex-offsite/install.sh b/scripts/dooplex-offsite/install.sh new file mode 100755 index 00000000..2c89e19c --- /dev/null +++ b/scripts/dooplex-offsite/install.sh @@ -0,0 +1,24 @@ +#!/bin/sh +# install.sh — install the Gitea + secrets off-site units on DooPlex (R-232). Root. Idempotent. +# Installs the three scripts and five units, writes /etc/felhom-dooplex-offsite/env (no secrets) when absent, and does +# NOT enable the timers — enable them by hand after the first manual run: +# systemctl enable --now felhom-dooplex-offsite.timer felhom-dooplex-offsite-restore-test.timer +# The tokens are the hub-DB ones (/etc/felhom-hub-backup/token-push, token-restore), read in place. enc.key is created +# separately (proxmox-backup-client key create --kdf none), never by this script. +set -eu +HERE=$(cd "$(dirname "$0")" && pwd) +[ "$(id -u)" = 0 ] || { echo "install.sh: run as root" >&2; exit 1; } +for s in felhom-dooplex-offsite felhom-dooplex-offsite-restore-test felhom-backup-failmail; do + install -m 0755 "$HERE/$s" "/usr/local/sbin/$s" +done +for u in felhom-dooplex-offsite.service felhom-dooplex-offsite.timer felhom-dooplex-offsite-restore-test.service \ + felhom-dooplex-offsite-restore-test.timer felhom-backup-failmail@.service; do + install -m 0644 "$HERE/$u" "/etc/systemd/system/$u" +done +install -d -m 0700 /etc/felhom-dooplex-offsite /var/lib/felhom-dooplex-offsite +if [ ! -f /etc/felhom-dooplex-offsite/env ]; then + umask 077 + grep -E '^(PBS_REPOSITORY_PUSH|PBS_REPOSITORY_RESTORE|PBS_FINGERPRINT)=' /etc/felhom-hub-backup/env > /etc/felhom-dooplex-offsite/env +fi +systemctl daemon-reload +echo "install.sh: installed; timers NOT enabled (see the header)" diff --git a/scripts/dooplex-offsite/test_dooplex_offsite.py b/scripts/dooplex-offsite/test_dooplex_offsite.py new file mode 100644 index 00000000..1ef828d6 --- /dev/null +++ b/scripts/dooplex-offsite/test_dooplex_offsite.py @@ -0,0 +1,332 @@ +#!/usr/bin/env python3 +"""Tests for felhom-dooplex-offsite, its restore test and felhom-backup-failmail (R-232). + +No test reaches Gitea, k3s, PBS, ep0 or Resend: `kubectl`, `proxmox-backup-client` and `pg_restore` are fakes on PATH +(the Gitea pod's /data is a temp dir; the PBS "server" is a temp dir; the fake pg_restore accepts a file that starts +with PostgreSQL's custom-dump magic "PGDMP"). `git`, `tar`, `sha256sum`, `shred`, `find` are the real tools, so +`git fsck` runs on real repositories. Each test asserts the CONSEQUENCE: whether a push happened and whether the +success signal (the file the alarm reads) was written. Run: python3 scripts/dooplex-offsite/test_dooplex_offsite.py +""" +import json +import os +import shutil +import subprocess +import tempfile +import time +import unittest + +HERE = os.path.dirname(os.path.abspath(__file__)) +PUSH = os.path.join(HERE, "felhom-dooplex-offsite") +RESTORE = os.path.join(HERE, "felhom-dooplex-offsite-restore-test") +FAILMAIL = os.path.join(HERE, "felhom-backup-failmail") + +FAKE_KUBECTL = r'''#!/usr/bin/env python3 +import os, subprocess, sys +a = sys.argv[1:] +pod = os.environ["FAKE_POD_DATA"] +cmd = a[a.index("--") + 1:] +if cmd[0] == "tar": + cnt = os.path.join(os.environ["FAKE_STATE"], "tar-calls") + n = int(open(cnt).read()) if os.path.exists(cnt) else 0 + open(cnt, "w").write(str(n + 1)) + if n < int(os.environ.get("FAKE_TAR_FAILS", "0")): + sys.stderr.write("tar: file vanished\n"); sys.exit(1) + c = ["tar" if x == "tar" else (pod if x == "/data" else x) for x in cmd] + sys.exit(subprocess.call(c)) +if cmd[0] == "find": + out = subprocess.run(["find", pod + cmd[1][len("/data"):]] + cmd[2:], + capture_output=True, text=True) + for line in out.stdout.splitlines(): + print("/data" + line[len(pod):]) + for extra in filter(None, os.environ.get("FAKE_EXTRA_REPOS", "").split(",")): + print("/data/git/repositories/admin/" + extra) + sys.exit(out.returncode) +sys.exit(97) +''' + +FAKE_PBS = r'''#!/usr/bin/env python3 +import json, os, shutil, sys, time +a = sys.argv[1:] +srv = os.environ["FAKE_PBS_DIR"] +open(os.path.join(srv, "calls.log"), "a").write(json.dumps({"argv": a, "pw": os.environ.get("PBS_PASSWORD_FILE", "")}) + "\n") +base = os.path.join(srv, "snaps") +if a[0] == "backup": + if os.environ.get("FAKE_PBS_FAIL"): sys.exit(1) + src = a[1].split(":", 1)[1] + t = int(os.environ.get("FAKE_PBS_TIME", time.time())) + shutil.copytree(src, os.path.join(base, str(t))) + sys.exit(0) +if a[0] == "snapshot" and a[1] == "list": + out = [{"backup-type": "host", "backup-id": "dooplex-gitea", "backup-time": int(x)} for x in (os.listdir(base) if os.path.isdir(base) else [])] + print(json.dumps(out)); sys.exit(0) +if a[0] == "restore": + import calendar + t = calendar.timegm(time.strptime(a[1].split("/")[-1], "%Y-%m-%dT%H:%M:%SZ")) + shutil.copytree(os.path.join(base, str(t)), a[3]) + sys.exit(0) +sys.exit(98) +''' + +FAKE_PG_RESTORE = r'''#!/bin/sh +[ "$1" = "--list" ] || exit 97 +head -c 5 "$2" | grep -q '^PGDMP' || { echo "pg_restore: error: input file does not appear to be a valid archive" >&2; exit 1; } +''' + + +def git(*a, cwd=None): + env = dict(os.environ, GIT_AUTHOR_NAME="t", GIT_AUTHOR_EMAIL="t@t", GIT_COMMITTER_NAME="t", GIT_COMMITTER_EMAIL="t@t") + subprocess.run(["git", *a], cwd=cwd, check=True, capture_output=True, env=env) + + +class Base(unittest.TestCase): + def setUp(self): + self.t = tempfile.mkdtemp() + j = lambda *p: os.path.join(self.t, *p) + self.pod, self.pbs, self.state, self.text = j("pod"), j("pbs"), j("state"), j("textfile") + self.conf, self.tokens, self.dumps, self.secrets, self.bin = j("conf"), j("tokens"), j("dumps"), j("secrets"), j("bin") + for d in (self.pbs, self.state, self.text, self.conf, self.tokens, self.dumps, self.secrets, self.bin): + os.makedirs(d) + # the Gitea pod: two real repositories with a commit each, app.ini, the other paths + for name in ("felhom.eu", "felhom-agent"): + bare = j("pod", "git", "repositories", "admin", name + ".git") + os.makedirs(os.path.dirname(bare), exist_ok=True) + git("init", "-q", "--bare", bare) + work = j("work-" + name) + git("init", "-q", work) + open(os.path.join(work, "README"), "w").write(name + "\n") + git("add", "README", cwd=work); git("commit", "-q", "-m", "c", cwd=work) + git("push", "-q", bare, "HEAD:refs/heads/main", cwd=work) + for p in ("git/lfs", "gitea/conf", "gitea/attachments", "gitea/avatars", "gitea/repo-avatars", "gitea/jwt", "gitea/packages"): + os.makedirs(j("pod", p), exist_ok=True) + open(j("pod", "gitea/conf/app.ini"), "w").write("[database]\nDB_TYPE = postgres\n") + open(j("pod", "gitea/packages/blob"), "w").write("registry - must not be copied\n") + # a complete dump (00:00 local) and the nightly secrets export + self.now = int(time.time()) + self.add_dump("20261008-220001", self.now - 1200) + self.add_secrets("20261008_031011", self.now - 21 * 3600) + for n in ("token-push", "token-restore"): + open(os.path.join(self.tokens, n), "w").write("x") + open(os.path.join(self.conf, "enc.key"), "w").write("{}") + open(os.path.join(self.conf, "env"), "w").write( + "PBS_REPOSITORY_PUSH='u!push@h:1:s'\nPBS_REPOSITORY_RESTORE='u!restore@h:1:s'\nPBS_FINGERPRINT='aa'\n") + for name, body in (("kubectl", FAKE_KUBECTL), ("proxmox-backup-client", FAKE_PBS), ("pg_restore", FAKE_PG_RESTORE)): + p = os.path.join(self.bin, name) + open(p, "w").write(body); os.chmod(p, 0o755) + + def tearDown(self): + shutil.rmtree(self.t, ignore_errors=True) + + def add_dump(self, name, mtime, complete=True, magic=b"PGDMP"): + d = os.path.join(self.dumps, name); os.makedirs(d) + open(os.path.join(d, "gitea.dump"), "wb").write(magic + b"\x01dump") + open(os.path.join(d, "globals.sql"), "w").write("-- roles\n") + if complete: + open(os.path.join(d, "SUCCESS"), "w").close(); os.utime(os.path.join(d, "SUCCESS"), (mtime, mtime)) + + def add_secrets(self, stamp, mtime): + for k, ext in (("secrets", "yaml.gpg"), ("configmaps", "yaml.gpg"), ("by-namespace", "tar.gz.gpg")): + p = os.path.join(self.secrets, "%s-%s.%s" % (k, stamp, ext)) + open(p, "wb").write(b"gpg-" + k.encode()); os.utime(p, (mtime, mtime)) + + def env(self, **kw): + e = dict(os.environ, PATH=self.bin + ":" + os.environ["PATH"], FAKE_POD_DATA=self.pod, FAKE_PBS_DIR=self.pbs, + FAKE_STATE=self.t, FELHOM_DXOFF_CONF=self.conf, FELHOM_DXOFF_TOKENS=self.tokens, + FELHOM_DXOFF_STATE=self.state, FELHOM_DXOFF_TEXTFILE_DIR=self.text, FELHOM_DXOFF_DUMPS=self.dumps, + FELHOM_DXOFF_SECRETS=self.secrets) + e.update({k: str(v) for k, v in kw.items()}) + return e + + def push(self, **kw): + return subprocess.run([PUSH], env=self.env(**kw), capture_output=True, text=True) + + def restore(self, **kw): + return subprocess.run([RESTORE], env=self.env(**kw), capture_output=True, text=True) + + def pushed(self): + d = os.path.join(self.pbs, "snaps") + return sorted(os.listdir(d)) if os.path.isdir(d) else [] + + def signal(self, name="felhom_dooplex_offsite.prom"): + return os.path.exists(os.path.join(self.text, name)) + + +class Push(Base): + def test_happy_path_pushes_everything_but_the_registry_and_writes_the_signal(self): + r = self.push() + self.assertEqual(r.returncode, 0, r.stderr) + self.assertEqual(len(self.pushed()), 1) + snap = os.path.join(self.pbs, "snaps", self.pushed()[0]) + self.assertTrue(os.path.isfile(os.path.join(snap, "gitea/git/repositories/admin/felhom.eu.git/HEAD"))) + self.assertTrue(os.path.isfile(os.path.join(snap, "gitea/gitea/conf/app.ini"))) + self.assertFalse(os.path.exists(os.path.join(snap, "gitea/gitea/packages")), "the registry must stay out") + self.assertEqual(open(os.path.join(snap, "db/DUMP-FOLDER")).read().strip(), "20261008-220001") + self.assertEqual(len(os.listdir(os.path.join(snap, "secrets"))), 3) + self.assertEqual(open(os.path.join(snap, "REPOS")).read().strip(), "2") + self.assertTrue(self.signal()) + call = json.loads(open(os.path.join(self.pbs, "calls.log")).readline()) + self.assertIn("--crypt-mode", call["argv"]); self.assertIn("encrypt", call["argv"]) + self.assertTrue(call["pw"].endswith("token-push")) + self.assertFalse(os.path.exists(os.path.join(self.state, "stage")), "the stage must be removed") + + def test_newest_incomplete_dump_is_skipped_for_the_complete_one(self): + self.add_dump("20261009-040001", self.now - 60, complete=False) + r = self.push() + self.assertEqual(r.returncode, 0, r.stderr) + snap = os.path.join(self.pbs, "snaps", self.pushed()[0]) + self.assertEqual(open(os.path.join(snap, "db/DUMP-FOLDER")).read().strip(), "20261008-220001") + + def test_stale_dump_refuses(self): + shutil.rmtree(self.dumps); os.makedirs(self.dumps) + self.add_dump("20261008-040001", self.now - 8 * 3600) + r = self.push() + self.assertNotEqual(r.returncode, 0) + self.assertIn("dump CronJob stopped", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_no_complete_dump_refuses(self): + shutil.rmtree(self.dumps); os.makedirs(self.dumps) + self.add_dump("20261009-040001", self.now, complete=False) + r = self.push() + self.assertNotEqual(r.returncode, 0) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_one_failed_file_copy_is_retried(self): + r = self.push(FAKE_TAR_FAILS=1, FELHOM_DXOFF_RETRY_SLEEP=0) + self.assertEqual(r.returncode, 0, r.stderr) + self.assertTrue(self.signal()) + + def test_two_failed_file_copies_refuse(self): + r = self.push(FAKE_TAR_FAILS=2, FELHOM_DXOFF_RETRY_SLEEP=0) + self.assertNotEqual(r.returncode, 0) + self.assertIn("failed twice", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_fewer_repositories_than_the_pod_lists_refuses(self): + r = self.push(FAKE_EXTRA_REPOS="ghost.git") + self.assertNotEqual(r.returncode, 0) + self.assertIn("the pod lists 3", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_missing_app_ini_refuses(self): + os.remove(os.path.join(self.pod, "gitea/conf/app.ini")) + r = self.push() + self.assertNotEqual(r.returncode, 0) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_stale_secrets_refuse(self): + shutil.rmtree(self.secrets); os.makedirs(self.secrets) + self.add_secrets("20261007_031011", self.now - 31 * 3600) + r = self.push() + self.assertNotEqual(r.returncode, 0) + self.assertIn("secrets export", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_a_link_in_the_pods_archive_refuses(self): + os.symlink("/etc/passwd", os.path.join(self.pod, "gitea/avatars/evil")) + r = self.push() + self.assertNotEqual(r.returncode, 0) + self.assertIn("link(s)", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_failed_push_writes_no_signal(self): + r = self.push(FAKE_PBS_FAIL=1) + self.assertNotEqual(r.returncode, 0) + self.assertFalse(self.signal()) + + +class RestoreTest(Base): + def pushed_copy(self, **kw): + r = self.push(FAKE_PBS_TIME=self.now - 3600, **kw) + self.assertEqual(r.returncode, 0, r.stderr) + return os.path.join(self.pbs, "snaps", self.pushed()[0]) + + def test_happy_path_writes_the_signal_with_the_readonly_token(self): + self.pushed_copy() + r = self.restore() + self.assertEqual(r.returncode, 0, r.stderr) + self.assertIn("2 repositories pass git fsck", r.stdout) + self.assertTrue(self.signal("felhom_dooplex_offsite_restore.prom")) + calls = [json.loads(x) for x in open(os.path.join(self.pbs, "calls.log"))] + self.assertTrue(all(c["pw"].endswith("token-restore") for c in calls if c["argv"][0] != "backup")) + + def test_a_changed_file_fails_the_manifest(self): + snap = self.pushed_copy() + open(os.path.join(snap, "secrets", sorted(os.listdir(os.path.join(snap, "secrets")))[0]), "ab").write(b"x") + r = self.restore() + self.assertNotEqual(r.returncode, 0) + self.assertIn("MANIFEST", r.stderr) + self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom")) + + def test_a_broken_repository_fails_git_fsck(self): + snap = self.pushed_copy() + repo = os.path.join(snap, "gitea/git/repositories/admin/felhom-agent.git") + objs = [os.path.join(dp, f) for dp, _, fs in os.walk(os.path.join(repo, "objects")) for f in fs + if len(os.path.basename(dp)) == 2] + os.remove(objs[0]) + # keep the manifest honest about the removal, so ONLY git fsck can catch it + man = os.path.join(snap, "MANIFEST.sha256") + rel = "./" + os.path.relpath(objs[0], snap) + kept = [l for l in open(man).readlines() if not l.rstrip().endswith(rel)] + open(man, "w").writelines(kept) + r = self.restore() + self.assertNotEqual(r.returncode, 0) + self.assertIn("git fsck", r.stderr) + self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom")) + + def test_git_never_reads_a_repositorys_own_config(self): + # The pod's `config` files are untrusted input to a root git. A config git would refuse to open + # (repositoryformatversion 99) proves git never read it: the old `git -C fsck` failed here (red-proof). + snap = self.pushed_copy() + cfg = os.path.join(snap, "gitea/git/repositories/admin/felhom.eu.git/config") + open(cfg, "w").write("[core]\n\trepositoryformatversion = 99\n\tbare = true\n\tfsmonitor = touch /nonexistent/x\n") + man = os.path.join(snap, "MANIFEST.sha256") + kept = [l for l in open(man).readlines() if not l.rstrip().endswith("felhom.eu.git/config")] + open(man, "w").writelines(kept) + r = self.restore() + self.assertEqual(r.returncode, 0, r.stderr) + + def test_an_unreadable_dump_fails(self): + shutil.rmtree(self.dumps); os.makedirs(self.dumps) + self.add_dump("20261008-220001", self.now - 1200, magic=b"XXXXX") + self.pushed_copy() + r = self.restore() + self.assertNotEqual(r.returncode, 0) + self.assertIn("pg_restore", r.stderr) + self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom")) + + def test_an_old_copy_fails(self): + r = self.push(FAKE_PBS_TIME=self.now - 51 * 3600) + self.assertEqual(r.returncode, 0, r.stderr) + r = self.restore() + self.assertNotEqual(r.returncode, 0) + self.assertIn("h old", r.stderr) + self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom")) + + def test_no_copy_fails(self): + r = self.restore() + self.assertNotEqual(r.returncode, 0) + self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom")) + + +class FailMail(unittest.TestCase): + def test_failmail_calls_notify_failure_with_the_unit_name(self): + t = tempfile.mkdtemp() + try: + cfg = os.path.join(t, "cfg.sh"); out = os.path.join(t, "out") + open(cfg, "w").write('notify_failure() { echo "$1" > %s; return 0; }\n' % out) + r = subprocess.run([FAILMAIL, "felhom-dooplex-offsite.service"], env=dict(os.environ, FELHOM_FAILMAIL_CONFIG=cfg), + capture_output=True, text=True) + self.assertEqual(r.returncode, 0, r.stderr) + self.assertIn("felhom-dooplex-offsite.service failed", open(out).read()) + finally: + shutil.rmtree(t) + + def test_every_backup_unit_names_the_failmail(self): + units = [os.path.join(HERE, u) for u in ("felhom-dooplex-offsite.service", "felhom-dooplex-offsite-restore-test.service")] + units += [os.path.join(HERE, "..", "hub-db-backup", u) for u in ("felhom-hub-db-backup.service", "felhom-hub-db-restore-test.service")] + for u in units: + self.assertIn("OnFailure=felhom-backup-failmail@%n.service", open(u).read(), u) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/scripts/hub-db-backup/felhom-hub-db-backup.service b/scripts/hub-db-backup/felhom-hub-db-backup.service index 2f63fc5f..3d4749dd 100644 --- a/scripts/hub-db-backup/felhom-hub-db-backup.service +++ b/scripts/hub-db-backup/felhom-hub-db-backup.service @@ -3,6 +3,7 @@ Description=Felhom: push the hub DB snapshot to ep0 (R-173) Wants=network-online.target felhom-ep0-pbs-tunnel.service After=network-online.target felhom-ep0-pbs-tunnel.service +OnFailure=felhom-backup-failmail@%n.service [Service] Type=oneshot diff --git a/scripts/hub-db-backup/felhom-hub-db-restore-test.service b/scripts/hub-db-backup/felhom-hub-db-restore-test.service index 8217f282..7f95270c 100644 --- a/scripts/hub-db-backup/felhom-hub-db-restore-test.service +++ b/scripts/hub-db-backup/felhom-hub-db-restore-test.service @@ -3,6 +3,7 @@ Description=Felhom: restore-test the hub DB copy on ep0 (R-173) Wants=network-online.target felhom-ep0-pbs-tunnel.service After=network-online.target felhom-ep0-pbs-tunnel.service +OnFailure=felhom-backup-failmail@%n.service [Service] Type=oneshot