dooplex-offsite: nightly encrypted copy of Gitea + DooPlex secrets to ep0, restore test, failure mail (R-232 b/h)
gates / gates (push) Successful in 5m4s

Part A plan + readings, Part E read-backs (R-861 a, R-518) in audits/dooplex-survival-2026-10-09/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-09 10:08:31 +02:00
parent 9a55f0bbc7
commit 1707c928a9
18 changed files with 1097 additions and 0 deletions
@@ -0,0 +1,97 @@
# PLAN — an encrypted off-site copy of Gitea and DooPlex's secrets (R-232 (b), (h))
2026-10-09, Part A of the DooPlex-survival brief. Everything below was **measured read-only**; nothing was created.
Raw readings: `partA/readings.txt`.
## What is measured
| Thing | Size | Growth | Where it lives today |
|---|---|---|---|
| Gitea repositories (`/data/git/repositories`, 10 repos under `admin/`) | **614 MB** | ~600 MB in 8 months (oldest product repo 2026-02-11); felhom.eu.git 188 MB | Longhorn PVC `gitea-data` on `sdb1`; Longhorn backup on `sda1`, **retain=1**. Nothing else. |
| Gitea database (CNPG `postgresql`, db `gitea`) | 64 MB live, **6.2 MB** as a `pg_dump -Fc` | 5.62 → 6.22 MB in 4 days (~150 KB/day) | 6-hourly dumps in `/mnt/5_hdd/backup/postgresql/dumps/` (same disk as every backup) |
| Gitea config `app.ini` (holds `SECRET_KEY`, `INTERNAL_TOKEN`, JWT secrets) | 2 KB | — | in the PVC only |
| Gitea registry (`/data/gitea/packages`) | 27.7 GB | — | **Left out on purpose:** the images rebuild from the code. |
| DooPlex secrets set: the nightly GPG files (`secrets-`, `configmaps-`, `by-namespace-*.gpg`) | **4.4 MB per night** | flat (2,055,542 → 2,056,090 B in 12 days) | `sda1` only, 30 days |
**Correction to the 2026-08-06 recon (§7):** it said the Gitea repositories are "covered twice". The API mirror
(`backup/homelab-manifests/git-mirrors/`) holds **one** repository, `homelab-manifests.git` (7.1 MB). The four
product repositories have exactly one backup: the Longhorn copy with retain=1, on the same machine.
## The two places measured
| | ep0, namespace `operator` (where the hub DB goes) | A new Hetzner Storage Box sub-account |
|---|---|---|
| Free space | 83 GB of 98 GB (`/mnt/pbs-datastore`, 16 % used) | depends on the box; not measured |
| Write-only key | **Already exists.** Token `dooplex-hub@pbs!push` has `DatastoreBackup` on `/operator` only: it cannot forget a snapshot or list the households (measured 2026-10-05). A second read-only token `!restore` exists for the restore test. | the R-820 append-only pin; the sub-account shell can still `rm` unless the pin is right |
| Changes needed there | **None.** A new backup group (`host/dooplex-gitea`) in the same namespace needs no new token, ACL or prune job. The existing prune job `prune-operator-hubdb` (keep-daily 14, keep-weekly 8) covers every group in `operator`. | a new sub-account + key custody |
| Cost | 0 € extra | a sub-account (same Storage Box fee) |
| Return copy | ep0's nightly pull to DooPlex (`ep0-copy`, 05:00) copies it back too — ciphertext only | none |
**Pick: ep0 `operator`.** No change on ep0 at all, the same write-only token as the hub DB, the same prune and
the same alarm family. Worst case on ep0: a git repack every night makes each copy new chunks — 22 kept copies ×
0.65 GB ≈ 14 GB, inside the 83 GB free.
## What runs
A script and a timer on DooPlex, `felhom-dooplex-offsite` (versioned in `felhom.eu/scripts/dooplex-offsite/`,
installed by its `install.sh`, tests by hand like the hub DB's), **daily 00:20** (after the 00:00 database dump,
before the hub DB at 02:30):
1. **The database first:** the newest **complete** `gitea.dump` + `globals.sql` from the 6-hourly dumps (folder with
`SUCCESS`; refused if older than 7 h). This is PostgreSQL's own consistent snapshot.
2. **Then the files**, read from the running Gitea pod with `tar` over `kubectl exec`: `/data/git/repositories`,
`/data/git/lfs`, `/data/gitea/conf/app.ini`, attachments, avatars, `jwt`. **Not** packages, logs, indexers, queues, tmp.
3. **Then the secrets set:** the newest night's three `.gpg` files (already GPG-encrypted with DooPlex's restic
passphrase, which the operator holds offline).
4. A `MANIFEST.sha256` of everything, then `proxmox-backup-client backup` → `host/dooplex-gitea` in `operator`,
`--crypt-mode encrypt` with a **new key** `/etc/felhom-dooplex-offsite/enc.key`.
5. Stage shredded; a success timestamp (written only after the push returns 0) for Prometheus.
**Why not Gitea's own `gitea dump`:** Gitea's documentation says the instance "must be shutdown during backup" for
consistency, and `gitea dump` does not stop it either. A nightly stop costs CI and the registry a gap and adds a
restart that must come back. Instead the order makes the copy safe: the database is older than the repositories, so
every commit the database names is in the copy (git writes objects before it moves a ref). A repository newer than
its database is the normal state after a push and Gitea reads it. **The restore test catches the rest:** it runs
`git fsck` on every repository.
## Encryption and the key
- New key, generated on DooPlex as root: `proxmox-backup-client key create --kdf none` → `0600` file in
`/etc/felhom-dooplex-offsite/`, never printed by CC, never in chat.
- **The operator copies its paper form into the password manager** at his own terminal (not through `!` here):
`sudo proxmox-backup-client key paperkey /etc/felhom-dooplex-offsite/enc.key --output-format text` — the `data`
field is enough (as for the hub DB key, runbook Step 0).
- ep0 never sees the key. The secrets files inside are encrypted twice (GPG + PBS).
## Retention and who can delete
- ep0's existing `prune-operator-hubdb` (03:45): 14 daily + 8 weekly, per group. ep0's GC frees chunks.
- **Can delete:** only ep0's root (and that prune job). The push token cannot forget a snapshot; the restore token
cannot write. DooPlex's root cannot delete on ep0 with what it holds.
## Alarms and failure mail
- `DooplexGiteaOffsiteStale` (26 h, `absent()` included) and `DooplexGiteaRestoreTestStale` (8 days), same group as
`HubDBBackupStale`, proven with `promtool test rules` and a red-proof.
- Failure mail: an `OnFailure=` unit that calls the R-232 (a) `notify_failure` (Resend → admin@felhom.eu). The same
`OnFailure=` is added to the two hub-DB units. Proven with one dry failure (a unit that fails on purpose, no backup).
## The restore test
- **Weekly, Sunday 05:30**, with the read-only token: restore the newest copy to a temp dir, check `MANIFEST.sha256`,
`git fsck` every repository, `pg_restore --list gitea.dump`, then shred; success timestamp.
- **Once now (Part C):** restore into a throwaway machine, start Gitea there with no outside network, compare every
product repository's `main` with live Gitea, read one file byte for byte, log in with a throwaway admin. Steps
become `runbooks/gitea-restore.md`. The throwaway and every copy are deleted.
## What changes on DooPlex (needs the operator's yes)
1. `/usr/local/sbin/felhom-dooplex-offsite` + `felhom-dooplex-offsite-restore-test` (scripts, from git).
2. Units: `felhom-dooplex-offsite.{service,timer}`, `felhom-dooplex-offsite-restore-test.{service,timer}`,
`felhom-backup-failmail@.service`; one `OnFailure=` line added to the two hub-DB services.
3. `/etc/felhom-dooplex-offsite/` (root 0700): `env` (no secret) and the new `enc.key`. The tokens are the existing
`/etc/felhom-hub-backup/token-push` and `token-restore` (read, not copied).
4. Two alarm rules in `homelab-manifests` `prometheus-rules` + the Prometheus reload.
5. One manual run, one manual restore test, one dry failure mail.
**On ep0: nothing.** Reads only (the snapshot list).
@@ -0,0 +1,60 @@
# Part A readings, 2026-10-09T07:54:41Z, read-only
## Gitea
gitea version 1.26.2 built with go1.26.3-X:jsonv2 : bindata, timetzdata, sqlite, sqlite_unlock_notify
613.5M /data/git/repositories
4.0K /data/git/lfs
27.7G /data/gitea/packages
12.0K /data/gitea/conf
4.0K /data/gitea/attachments
12.0K /data/gitea/avatars
8.0K /data/gitea/jwt
app-catalog-drill.git
app-catalog-felhom.eu.git
felhom-agent.git
felhom-controller.git
felhom.eu.git
homelab-manifests.git
jarr.git
misc-scripts.git
recipe-importer.git
revfulop-calendar.git
188.1M /data/git/repositories/admin/felhom.eu.git
73.7M /data/git/repositories/admin/felhom-controller.git
61.4M /data/git/repositories/admin/felhom-agent.git
19.2M /data/git/repositories/admin/app-catalog-felhom.eu.git
DB_TYPE = postgres
HOST = postgresql-rw.database-system.svc.cluster.local:5432
NAME = gitea
## Gitea DB
64 MB
/mnt/5_hdd/backup/postgresql/dumps/20261005-100001/gitea.dump 5624428
/mnt/5_hdd/backup/postgresql/dumps/20261009-040001/gitea.dump 6218173
## secrets set
-rw-r--r-- 1 root root 1926239 Sep 28 03:07 by-namespace-20260928_030725.tar.gz.gpg
-rw-r--r-- 1 root root 1944545 Oct 9 03:10 by-namespace-20261009_031011.tar.gz.gpg
-rw-r--r-- 1 root root 401829 Sep 28 03:07 configmaps-20260928_030725.yaml.gpg
-rw-r--r-- 1 root root 404486 Oct 9 03:10 configmaps-20261009_031011.yaml.gpg
-rw-r--r-- 1 root root 2055542 Sep 28 03:07 secrets-20260928_030725.yaml.gpg
-rw-r--r-- 1 root root 2056090 Oct 9 03:10 secrets-20261009_031011.yaml.gpg
130M /mnt/5_hdd/backup/secrets/exports
## API mirror
homelab-manifests.git
## ep0 (read-only ssh)
Filesystem Size Used Avail Use% Mounted on
/dev/sdb 98G 15G 83G 16% /mnt/pbs-datastore
"comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)",
"id": "prune-demo-hp",
"keep-last": 2,
"ns": "demo-hp",
"id": "prune-operator-hubdb",
"keep-daily": 14,
"keep-weekly": 8,
"ns": "operator",
"comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)",
"id": "prune-demo-felhom",
"keep-last": 2,
"ns": "demo-felhom",
/datastore/felhom-offsite/operator dooplex-hub@pbs DatastoreBackup
/datastore/felhom-offsite/operator dooplex-hub@pbs DatastoreReader
/datastore/felhom-offsite/operator dooplex-hub@pbs!push DatastoreBackup
/datastore/felhom-offsite/operator dooplex-hub@pbs!restore DatastoreReader
@@ -0,0 +1,31 @@
### RED: link-check (removed: LINKS)
FAIL: test_a_link_in_the_pods_archive_refuses (__main__.Push.test_a_link_in_the_pods_archive_refuses)
Ran 1 test in 0.283s
FAILED (failures=1)
### RED: dump-age (removed: DAGE)
FAIL: test_stale_dump_refuses (__main__.Push.test_stale_dump_refuses)
Ran 1 test in 0.289s
FAILED (failures=1)
### RED: repo-count (removed: GOT-ge-WANT)
FAIL: test_fewer_repositories_than_the_pod_lists_refuses (__main__.Push.test_fewer_repositories_than_the_pod_lists_refuses)
Ran 1 test in 0.271s
FAILED (failures=1)
### RED: fsck (removed: fsck)
FAIL: test_a_broken_repository_fails_git_fsck (__main__.RestoreTest.test_a_broken_repository_fails_git_fsck)
Ran 1 test in 0.462s
FAILED (failures=1)
### RED: manifest (removed: sha256-check)
FAIL: test_a_changed_file_fails_the_manifest (__main__.RestoreTest.test_a_changed_file_fails_the_manifest)
Ran 1 test in 0.457s
FAILED (failures=1)
### GREEN after restore
Ran 19 tests in 34.596s
OK
### RED: git reads the repo config (old fsck line)
FAIL: test_git_never_reads_a_repositorys_own_config (__main__.RestoreTest.test_git_never_reads_a_repositorys_own_config)
AssertionError: 1 != 0 : felhom-dooplex-offsite-restore-test: FAILED: git fsck admin/felhom.eu.git: fatal: Expected git repo version <= 1, found 99
Ran 1 test in 0.441s
FAILED (failures=1)
### GREEN, full suite
Ran 20 tests in 35.065s
OK
@@ -0,0 +1,151 @@
# R-518 read-back — the first night with BOTH tiers due on demo-hp. Read-only, collected from DooPlex 2026-10-09T08:02:41Z.
# demo-hp host journal = CEST (UTC+2). Controller metrics, docker, PVE UPID and ep0 = UTC.
################ CH1 — demo-hp agent journal (felhom-agent), backup lines 2026-10-08 22:15-22:45 local (= 20:15-20:45Z)
Oct 08 22:20:43 demo-hp felhom-agent[666190]: time=2026-10-08T22:20:43.713+02:00 level=INFO msg="backup: space preflight passed" vmid=9201 target=local last_archive_bytes=4885939687 need_bytes=7181166432 avail_bytes=17724211200
Oct 08 22:20:44 demo-hp felhom-agent[666190]: time=2026-10-08T22:20:44.762+02:00 level=INFO msg="local-api: backup reached snapshotted (app may resume)" vmid=9201 target=local job=backup-9201-1791490843121326096
Oct 08 22:25:40 demo-hp felhom-agent[666190]: time=2026-10-08T22:25:40.590+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_08-22_20_43.tar.zst size_bytes=4887664196 uncovered_volumes=2
Oct 08 22:25:40 demo-hp felhom-agent[666190]: time=2026-10-08T22:25:40.590+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1791490843121326096 archive=local:backup/vzdump-lxc-9201-2026_10_08-22_20_43.tar.zst
Oct 08 22:26:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:26:14.221+02:00 level=INFO msg="local-api: backup refused — a heavy operation is already in flight" vmid=9201 requested_target=felhom-pbs busy=backup:local
Oct 08 22:27:10 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:10.591+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=guest vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z
Oct 08 22:27:11 demo-hp felhom-os-apply[1209621]: os-apply: START release=ring0-20261008T202710Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0
Oct 08 22:27:16 demo-hp felhom-os-apply[1209908]: os-apply: REPAIR configured=0 journal=0 fixed=0
Oct 08 22:27:26 demo-hp felhom-os-apply[1210615]: os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0
Oct 08 22:27:37 demo-hp felhom-os-apply[1211512]: os-apply: DONE rc=0 seconds=4.5 upgraded=2 restart-needed=- reboot-needed=no
Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0"
Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0"
Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.957+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=4.5 upgraded=2 restart-needed=- reboot-needed=no"
Oct 08 22:27:45 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:45.958+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=guest vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=2 pending=1 not_covered=1 restart_needed=0 reboot_needed=false wrapper_seconds=35.3
Oct 08 22:27:46 demo-hp felhom-agent[666190]: time=2026-10-08T22:27:46.035+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=host vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z
Oct 08 22:27:47 demo-hp felhom-os-apply[1211951]: os-apply: START release=ring0-20261008T202710Z layer=host lane=fast mode=apply select=pending-fast packages=0
Oct 08 22:27:51 demo-hp felhom-os-apply[1212307]: os-apply: REPAIR configured=0 journal=0 fixed=0
Oct 08 22:27:56 demo-hp felhom-os-apply[1212564]: os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0
Oct 08 22:28:02 demo-hp felhom-os-apply[1213497]: os-apply: DONE rc=0 seconds=3.5 upgraded=2 restart-needed=kvm,pmxcfs,pve-firewall,pve-ha-crm,pve-ha-lrm,pvedaemon,pvedaemon worke,pveproxy,pveproxy worker,pvescheduler,pvestatd,rrdcached,spiceproxy,spiceproxy work reboot-needed=no
Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=host lane=fast mode=apply select=pending-fast packages=0"
Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0"
Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0"
Oct 08 22:28:11 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:11.095+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=3.5 upgraded=2 restart-needed=kvm,pmxcfs,pve-firewall,pve-ha-crm,pve-ha-lrm,pvedaemon,pvedaemon worke,pveproxy,pveproxy worker,pvescheduler,pvestatd,rrdcached,spiceproxy,spiceproxy work reboot-needed=no"
Oct 08 22:28:12 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:12.164+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=host vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=2 pending=80 not_covered=80 restart_needed=14 reboot_needed=false wrapper_seconds=25
Oct 08 22:28:14 demo-hp felhom-os-apply[1214081]: os-apply: LIVE-RESTORE already on
Oct 08 22:28:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:14.220+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: LIVE-RESTORE already on"
Oct 08 22:28:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:14.220+02:00 level=INFO msg="osupdate: live-restore" vmid=9201 result="{\"result\": \"already on\"}"
Oct 08 22:28:14 demo-hp felhom-agent[666190]: time=2026-10-08T22:28:14.220+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=docker vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z
Oct 08 22:28:16 demo-hp felhom-os-apply[1214207]: os-apply: START release=ring0-20261008T202710Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0
Oct 08 22:28:21 demo-hp felhom-os-apply[1214648]: os-apply: REPAIR configured=0 journal=0 fixed=0
Oct 08 22:28:27 demo-hp felhom-os-apply[1214956]: os-apply: PLAN upgrade=1 already=0 not-installed=0 from-snapshot=0
Oct 08 22:29:01 demo-hp felhom-os-apply[1217263]: os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket)
Oct 08 22:29:03 demo-hp felhom-os-apply[1218011]: os-apply: DONE rc=0 seconds=5.1 upgraded=1 restart-needed=- reboot-needed=no
Oct 08 22:29:21 demo-hp felhom-os-apply[1219347]: os-apply: OOM-CHECK result=pass oom_killed=True oom_event=True exit=137 image=gitea.dooplex.hu/admin/felhom-controller:0.303.0 — the engine reported the memory kill: OOMKilled=true and the oom event
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0"
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0"
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=0 not-installed=0 from-snapshot=0"
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket)"
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=5.1 upgraded=1 restart-needed=- reboot-needed=no"
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.329+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: OOM-CHECK result=pass oom_killed=True oom_event=True exit=137 image=gitea.dooplex.hu/admin/felhom-controller:0.303.0 — the engine reported the memory kill: OOMKilled=true and the oom event"
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.330+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=docker vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=1 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=67
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.342+02:00 level=INFO msg="osupdate: pve step holds the /etc/pve write gate — the agent's own writes wait until it ends" run=20261008T202710Z layer=pve vmid=9201 trigger=night
Oct 08 22:29:21 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:21.342+02:00 level=INFO msg="osupdate: START" run=20261008T202710Z layer=pve vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T202710Z
Oct 08 22:29:22 demo-hp felhom-os-apply[1219396]: os-apply: START release=ring0-20261008T202710Z layer=pve:9201 lane=slow mode=apply select=pending-pve packages=0 authority=ring0
Oct 08 22:29:26 demo-hp felhom-os-apply[1219654]: os-apply: REPAIR configured=0 journal=0 fixed=0
Oct 08 22:29:30 demo-hp felhom-os-apply[1219888]: os-apply: PLAN upgrade=70 already=0 not-installed=0 from-snapshot=0
Oct 08 22:29:33 demo-hp felhom-os-apply[1220017]: os-apply: REFUSED: R6 the plan would touch shim-signed-common, which is not in the plan
Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T202710Z layer=pve:9201 lane=slow mode=apply select=pending-pve packages=0 authority=ring0"
Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0"
Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=70 already=0 not-installed=0 from-snapshot=0"
Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.402+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REFUSED: R6 the plan would touch shim-signed-common, which is not in the plan"
Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.403+02:00 level=INFO msg="osupdate: DONE" run=20261008T202710Z layer=pve vmid=9201 ring=0 trigger=night outcome=refused healthy=false reason="" upgraded=0 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=0
Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.782+02:00 level=INFO msg="osupdate: pve step released the /etc/pve write gate" run=20261008T202710Z layer=pve vmid=9201 trigger=night
Oct 08 22:29:33 demo-hp felhom-agent[666190]: time=2026-10-08T22:29:33.783+02:00 level=WARN msg="osupdate: kernel step skipped — the Proxmox step did not end healthy" run=20261008T202710Z vmid=9201 trigger=night pve_outcome=refused
Oct 08 22:34:28 demo-hp felhom-agent[666190]: time=2026-10-08T22:34:28.003+02:00 level=INFO msg="local-api: backup reached snapshotted (app may resume)" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1791491666431534232
Oct 08 22:37:38 demo-hp felhom-agent[666190]: time=2026-10-08T22:37:38.865+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-10-08T20:34:27Z size_bytes=15886927313 uncovered_volumes=2
Oct 08 22:37:38 demo-hp felhom-agent[666190]: time=2026-10-08T22:37:38.865+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1791491666431534232 archive=felhom-pbs:backup/ct/9201/2026-10-08T20:34:27Z
== vzdump task log
2026/10/09 05:05:43 main.go:3684: [INFO] local-api: mount mp8 → /mnt/felhom-drives (storage=/mnt/felhom-drives, class=, backup=false)
################ CH1 context — tier cadences as armed (agent journal, latest start)
Oct 09 07:16:47 demo-hp felhom-agent[2730121]: time=2026-10-09T07:16:47.622+02:00 level=INFO msg="backup tier armed" target=local cadence=24h0m0s keep_last=1 wait_timeout=30m0s prune_pbs_allowed=false primary=true
Oct 09 07:16:47 demo-hp felhom-agent[2730121]: time=2026-10-09T07:16:47.622+02:00 level=INFO msg="backup tier armed" target=felhom-pbs cadence=168h0m0s keep_last=0 wait_timeout=12h0m0s prune_pbs_allowed=false primary=false
################ CH1b — PVE task index on demo-hp (pvedaemon task log; a different writer from the agent): vzdump tasks 2026-10-07..09
UPID:demo-hp:0031920A:00EA3C89:6AC5AFDF:vzdump:9201:felhom-agent@pve!agent: start=2026-10-07T02:35:11Z end=2026-10-07T02:40:45Z OK
UPID:demo-hp:003C6995:01011C3B:6AC5EA6D:vzdump:9201:felhom-agent@pve!agent: start=2026-10-07T06:45:01Z end=2026-10-07T06:49:49Z OK
UPID:demo-hp:00120671:00AABC2D:6AC7FB1B:vzdump:9201:felhom-agent@pve!agent: start=2026-10-08T20:20:43Z end=2026-10-08T20:25:39Z OK
UPID:demo-hp:0012D84F:00ABFDC1:6AC7FE52:vzdump:9201:felhom-agent@pve!agent: start=2026-10-08T20:34:26Z end=2026-10-08T20:37:35Z OK
-- local vzdump log of the local tier
2026-10-08 22:20:43 INFO: Starting Backup of VM 9201 (lxc)
2026-10-08 22:20:43 INFO: status = running
2026-10-08 22:20:43 INFO: CT Name: demo-hp
2026-10-08 22:20:43 INFO: including mount point rootfs ('/') in backup
2026-10-08 22:20:43 INFO: including mount point mp0 ('/var/lib/felhom') in backup
2026-10-08 22:20:43 INFO: excluding bind mount point mp8 ('/mnt/felhom-drives') from backup (not a volume)
2026-10-08 22:20:43 INFO: excluding bind mount point mp9 ('/etc/felhom-bootstrap') from backup (not a volume)
2026-10-08 22:20:43 INFO: backup mode: snapshot
2026-10-08 22:20:43 INFO: ionice priority: 7
2026-10-08 22:20:43 INFO: suspend vm to make snapshot
2026-10-08 22:20:43 INFO: create storage snapshot 'vzdump'
2026-10-08 22:20:45 INFO: resume vm
2026-10-08 22:20:45 INFO: guest is online again after 2 seconds
2026-10-08 22:20:45 INFO: creating vzdump archive '/var/lib/vz/dump/vzdump-lxc-9201-2026_10_08-22_20_43.tar.zst'
2026-10-08 22:25:33 INFO: Total bytes written: 16625551360 (16GiB, 56MiB/s)
2026-10-08 22:25:33 INFO: archive file size: 4.55GB
2026-10-08 22:25:33 INFO: adding notes to backup
2026-10-08 22:25:33 INFO: prune older backups with retention: keep-last=1
2026-10-08 22:25:33 INFO: removing backup 'local:backup/vzdump-lxc-9201-2026_10_07-08_45_01.tar.zst'
2026-10-08 22:25:33 INFO: pruned 1 backup(s) not covered by keep-retention policy
2026-10-08 22:25:35 INFO: cleanup temporary 'vzdump' snapshot
2026-10-08 22:25:38 INFO: Finished Backup of VM 9201 (00:04:55)
################ CH2 — controller metrics.db in guest 9201 (one row per RUNNING container per ~60 s sample; a stopped container has no row)
# query: per sample, running count and the containers present at 20:14-20:20:30Z that are missing. Opened ?mode=ro via the controller container's sqlite3.
2026-10-08T20:14:19Z|21|
2026-10-08T20:15:19Z|21|
2026-10-08T20:16:19Z|21|
2026-10-08T20:17:19Z|21|
2026-10-08T20:18:19Z|21|
2026-10-08T20:19:19Z|21|
2026-10-08T20:20:19Z|21|
2026-10-08T20:21:21Z|15|privatebin,paperless-webserver,paperless-postgres,paperless-redis,opengist,kimai
2026-10-08T20:22:19Z|21|
2026-10-08T20:23:19Z|21|
2026-10-08T20:24:19Z|21|
2026-10-08T20:25:19Z|21|
2026-10-08T20:26:19Z|6|privatebin,paperless-webserver,paperless-postgres,paperless-redis,opengist,kimai,kimai-db,docmost,docmost-postgres,docmost-redis,calibre-web,bookstack,bookstack-db,adventurelog,bentopdf
2026-10-08T20:27:19Z|21|
2026-10-08T20:28:19Z|21|
2026-10-08T20:29:04Z|21|
2026-10-08T20:30:04Z|21|
2026-10-08T20:31:04Z|21|
2026-10-08T20:32:04Z|21|
2026-10-08T20:33:04Z|21|
2026-10-08T20:34:04Z|21|
2026-10-08T20:35:04Z|15|privatebin,paperless-webserver,paperless-postgres,paperless-redis,opengist,kimai
2026-10-08T20:36:04Z|21|
2026-10-08T20:37:04Z|21|
2026-10-08T20:38:04Z|21|
2026-10-08T20:39:04Z|21|
2026-10-08T20:40:04Z|21|
2026-10-08T20:41:04Z|21|
total distinct containers 20:14-20:20:30Z: 21
################ CH2b — guest docker: container start/create times (bentopdf recreated 15 s after the off-site snapshot)
/bentopdf started=2026-10-08T20:34:43.277725548Z created=2026-10-08T20:34:43.157907984Z
/traefik started=2026-10-08T20:29:01.412754011Z created=2026-10-04T09:01:48.259555986Z
/bookstack started=2026-10-09T02:15:27.95387884Z created=2026-10-09T02:15:22.165968876Z
/felhom-controller started=2026-10-09T05:05:42.178985741Z created=2026-10-09T05:05:42.122744205Z
# all other app containers were recreated 2026-10-09T02:15-02:16Z (after the window), so their StartedAt cannot show the 20:2x stops.
################ CH3 — ep0 PBS datastore, namespace demo-hp, ct/9201 (ls, read-only)
felhom-hetzner
2026-10-09T08:02:47Z
total 20
drwxr-xr-x 4 backup backup 4096 2026-10-09 03:29:59.997681474 +0000 .
drwxr-xr-x 3 backup backup 4096 2026-07-26 15:42:44.833483940 +0000 ..
drwxr-xr-x 2 backup backup 4096 2026-10-09 05:17:18.883235046 +0000 2026-10-01T20:15:29Z
drwxr-xr-x 2 backup backup 4096 2026-10-09 05:17:15.547227803 +0000 2026-10-08T20:34:27Z
-rw-r--r-- 1 backup backup 19 2026-07-26 15:42:44.833483940 +0000 owner
################ What could NOT be read (tried)
# controller docker log for the night: 'docker logs felhom-controller' starts 2026-10-09 05:05:43 — the 0.304.0 swap replaced the container.
# controller persisted debug ring (data/debug-ring.log): 5000 lines, oldest 2026-10-09T07:30:53Z — does not reach the night.
# guest docker events: 'docker events --since 2026-10-08T20:15Z --until 20:45Z' returned nothing; the event buffer's oldest entry is 2026-10-09T07:58Z (healthcheck execs fill it).
@@ -0,0 +1,116 @@
# R-861 (a) read-back — read-only, collected from DooPlex 2026-10-09T07:57:32Z. Host journals are CEST (UTC+2); docker StartedAt is UTC.
# Release context: felhom.eu/documentation/audits/release-2026-10-09/delivery/floors-0.304.0.txt (floors set 2026-10-09T05:05:30Z)
################ demo-hp
demo-hp
felhom-agent 0.154.0
== CH1a agent journal 06:50-07:40 local (swap lines)
Oct 09 07:05:37 demo-hp felhom-agent[666190]: time=2026-10-09T07:05:37.552+02:00 level=WARN msg="local-api: controller swap requested" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 previous=gitea.dooplex.hu/admin/felhom-controller:0.303.0
Oct 09 07:05:39 demo-hp sudo[2696732]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:05:40 demo-hp felhom-priv-apply[2696817]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:05:40 demo-hp felhom-agent[666190]: time=2026-10-09T07:05:40.585+02:00 level=INFO msg="controller-swap: image file written, restarting bootstrap" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:05:47 demo-hp felhom-agent[666190]: time=2026-10-09T07:05:47.706+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:16:49 demo-hp sudo[2730461]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config
Oct 09 07:16:49 demo-hp felhom-priv-apply[2730492]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config
Oct 09 07:16:54 demo-hp sudo[2730994]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key
Oct 09 07:16:54 demo-hp felhom-priv-apply[2731003]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op
Oct 09 07:32:01 demo-hp felhom-agent[2730121]: time=2026-10-09T07:32:01.215+02:00 level=WARN msg="osupdate: config bundle INSTALLED" op=agent_config_update agent_version=0.154.0 bundle="{\"agent_version\": \"0.154.0\", \"authority\": \"signed\", \"kept\": [\"/etc/felhom/crash-guard.conf\"], \"prev_dir\": \"/var/lib/felhom-os-apply/bundle-prev/20261009T053200Z-before-0.154.0\", \"same\": 26, \"self_check\": {\"crash_guard\": {\"armed\": true, \"kernel_panic\": 10}, \"os_apply\": \"felhom-os-apply ok bundle-format=1 files=28\", \"priv_apply\": \"felhom-priv-apply ok verbs=unit,dnsmasq,wg,sshd-config,sshd-key,controller-image\", \"selfupdate\": \"usage ok\", \"sudo_l\": \"felhom-os-apply listed\", \"visudo\": \"ok\"}, \"sha256\": \"93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0\", \"signers_created\": false, \"skipped\": [], \"written\": [\"/usr/local/sbin/felhom-os-apply\"]}" pass_seconds=1
== CH1b sudo log 2026-10-09 06:50-07:40 local
Oct 09 07:05:36 demo-hp sudo[2696598]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image
Oct 09 07:05:37 demo-hp sudo[2696651]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image
Oct 09 07:05:39 demo-hp sudo[2696732]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:16:49 demo-hp sudo[2730461]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config
Oct 09 07:16:54 demo-hp sudo[2730994]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key
== CH1c priv-apply wrapper log 2026-10-09
Oct 09 07:05:40 demo-hp felhom-priv-apply[2696817]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0
== tee route since 2026-10-07 00:00 (count, then lines)
2
Oct 07 09:09:27 demo-hp sudo[4038406]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image
Oct 07 13:36:34 demo-hp sudo[102535]: root : PWD=/root ; USER=felhom-agent ; COMMAND=/usr/bin/sudo -n /usr/sbin/pct exec 9202 -- tee /etc/felhom-controller-image
== tee route on 2026-10-09 (count)
0
== verb runs since 2026-10-07
Oct 07 13:36:33 demo-hp sudo[102481]: root : PWD=/root ; USER=felhom-agent ; COMMAND=/usr/bin/sudo -n /usr/local/sbin/felhom-priv-apply controller-image 9202
Oct 07 13:36:33 demo-hp sudo[102483]: felhom-agent : PWD=/root ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9202
Oct 07 18:50:13 demo-hp sudo[631922]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:05:39 demo-hp sudo[2696732]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
== CH2 guest 9201 (docker in the guest)
gitea.dooplex.hu/admin/felhom-controller:0.304.0
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 3 hours (healthy)
/felhom-controller image=gitea.dooplex.hu/admin/felhom-controller:0.304.0 started=2026-10-09T05:05:42.178985741Z restarts=0
################ felhom-pve
demo-felhom
felhom-agent 0.154.0
== CH1a agent journal 06:50-07:40 local (swap lines)
Oct 09 07:05:35 demo-felhom felhom-agent[1202]: time=2026-10-09T07:05:35.089+02:00 level=WARN msg="local-api: controller swap requested" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 previous=gitea.dooplex.hu/admin/felhom-controller:0.303.0
Oct 09 07:05:36 demo-felhom sudo[130104]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:05:37 demo-felhom felhom-priv-apply[130111]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:05:37 demo-felhom felhom-agent[1202]: time=2026-10-09T07:05:37.382+02:00 level=INFO msg="controller-swap: image file written, restarting bootstrap" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:05:46 demo-felhom felhom-agent[1202]: time=2026-10-09T07:05:46.966+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:16:07 demo-felhom sudo[141297]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config
Oct 09 07:16:07 demo-felhom felhom-priv-apply[141326]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config
Oct 09 07:16:11 demo-felhom sudo[141738]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key
Oct 09 07:16:11 demo-felhom felhom-priv-apply[141741]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op
Oct 09 07:31:16 demo-felhom felhom-agent[141055]: time=2026-10-09T07:31:16.948+02:00 level=WARN msg="osupdate: config bundle INSTALLED" op=agent_config_update agent_version=0.154.0 bundle="{\"agent_version\": \"0.154.0\", \"authority\": \"signed\", \"kept\": [\"/etc/felhom/crash-guard.conf\"], \"prev_dir\": \"/var/lib/felhom-os-apply/bundle-prev/20261009T053116Z-before-0.154.0\", \"same\": 26, \"self_check\": {\"crash_guard\": {\"armed\": true, \"kernel_panic\": 10}, \"os_apply\": \"felhom-os-apply ok bundle-format=1 files=28\", \"priv_apply\": \"felhom-priv-apply ok verbs=unit,dnsmasq,wg,sshd-config,sshd-key,controller-image\", \"selfupdate\": \"usage ok\", \"sudo_l\": \"felhom-os-apply listed\", \"visudo\": \"ok\"}, \"sha256\": \"93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0\", \"signers_created\": false, \"skipped\": [], \"written\": [\"/usr/local/sbin/felhom-os-apply\"]}" pass_seconds=0.6
== CH1b sudo log 2026-10-09 06:50-07:40 local
Oct 09 07:05:34 demo-felhom sudo[130062]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image
Oct 09 07:05:35 demo-felhom sudo[130074]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image
Oct 09 07:05:36 demo-felhom sudo[130104]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:16:07 demo-felhom sudo[141297]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config
Oct 09 07:16:11 demo-felhom sudo[141738]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key
== CH1c priv-apply wrapper log 2026-10-09
Oct 09 07:05:37 demo-felhom felhom-priv-apply[130111]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0
== tee route since 2026-10-07 00:00 (count, then lines)
1
Oct 07 09:09:26 demo-felhom sudo[2302098]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image
== tee route on 2026-10-09 (count)
0
== verb runs since 2026-10-07
Oct 07 18:50:12 demo-felhom sudo[199002]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:05:36 demo-felhom sudo[130104]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
== CH2 guest 9201 (docker in the guest)
gitea.dooplex.hu/admin/felhom-controller:0.304.0
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 3 hours (healthy)
/felhom-controller image=gitea.dooplex.hu/admin/felhom-controller:0.304.0 started=2026-10-09T05:05:38.555475195Z restarts=0
################ root@192.168.0.154
felhom
felhom-agent 0.154.0
== CH1a agent journal 06:50-07:40 local (swap lines)
Oct 09 07:05:37 felhom felhom-agent[144633]: time=2026-10-09T07:05:37.458+02:00 level=WARN msg="local-api: controller swap requested" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0 previous=gitea.dooplex.hu/admin/felhom-controller:0.303.0
Oct 09 07:05:39 felhom sudo[2992327]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:05:40 felhom felhom-priv-apply[2992353]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:05:40 felhom felhom-agent[144633]: time=2026-10-09T07:05:40.580+02:00 level=INFO msg="controller-swap: image file written, restarting bootstrap" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:05:47 felhom felhom-agent[144633]: time=2026-10-09T07:05:47.589+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.304.0
Oct 09 07:18:20 felhom sudo[3010887]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config
Oct 09 07:18:20 felhom felhom-priv-apply[3010891]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config
Oct 09 07:18:24 felhom sudo[3011416]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key
Oct 09 07:18:24 felhom felhom-priv-apply[3011425]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op
Oct 09 07:33:30 felhom felhom-agent[3010618]: time=2026-10-09T07:33:30.409+02:00 level=WARN msg="osupdate: config bundle INSTALLED" op=agent_config_update agent_version=0.154.0 bundle="{\"agent_version\": \"0.154.0\", \"authority\": \"signed\", \"kept\": [\"/etc/felhom/crash-guard.conf\"], \"prev_dir\": \"/var/lib/felhom-os-apply/bundle-prev/20261009T053329Z-before-0.154.0\", \"same\": 26, \"self_check\": {\"crash_guard\": {\"armed\": true, \"kernel_panic\": 10}, \"os_apply\": \"felhom-os-apply ok bundle-format=1 files=28\", \"priv_apply\": \"felhom-priv-apply ok verbs=unit,dnsmasq,wg,sshd-config,sshd-key,controller-image\", \"selfupdate\": \"usage ok\", \"sudo_l\": \"felhom-os-apply listed\", \"visudo\": \"ok\"}, \"sha256\": \"93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0\", \"signers_created\": false, \"skipped\": [], \"written\": [\"/usr/local/sbin/felhom-os-apply\"]}" pass_seconds=0.9
== CH1b sudo log 2026-10-09 06:50-07:40 local
Oct 09 07:05:36 felhom sudo[2992261]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image
Oct 09 07:05:37 felhom sudo[2992286]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- cat /etc/felhom-controller-image
Oct 09 07:05:39 felhom sudo[2992327]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:18:20 felhom sudo[3010887]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-config
Oct 09 07:18:24 felhom sudo[3011416]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply sshd-key
== CH1c priv-apply wrapper log 2026-10-09
Oct 09 07:05:40 felhom felhom-priv-apply[2992353]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.304.0
== tee route since 2026-10-07 00:00 (count, then lines)
1
Oct 07 09:09:27 felhom sudo[3745733]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image
== tee route on 2026-10-09 (count)
0
== verb runs since 2026-10-07
Oct 07 18:50:13 felhom sudo[126979]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
Oct 09 07:05:39 felhom sudo[2992327]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201
== CH2 guest 9201 (docker in the guest)
gitea.dooplex.hu/admin/felhom-controller:0.304.0
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 2 hours (healthy)
/felhom-controller image=gitea.dooplex.hu/admin/felhom-controller:0.304.0 started=2026-10-09T05:58:31.221358279Z restarts=0
################ note: Tester 1 controller StartedAt 05:58:31Z (not 05:05:47Z)
# The Tester 1 VM rebooted at 07:58:11 local (kernel line "Linux version 7.0.14-22-pve" at Oct 09 07:58:11;
# agent "daemon starting version=0.154.0" at 07:58:20; os-apply plan "plan-boot-...-kernel-kernel-boot.json").
# The guest controller log's first lines after that: "2026/10/09 05:58:31 [INFO] felhom-controller 0.304.0 starting (customer: tester-1 ...)".
# No second controller-image verb run and no tee on 2026-10-09 (counts above). The later StartedAt is the reboot, not a second swap.
+16
View File
@@ -1,3 +1,19 @@
## 2026-10-09 — dooplex-offsite: Gitea + DooPlex's secrets leave DooPlex nightly, encrypted (R-232 (b), (h))
- New `scripts/dooplex-offsite/`: `felhom-dooplex-offsite` (daily 00:20) copies the newest complete `gitea.dump`
(taken FIRST), then Gitea's repositories, LFS, `app.ini`, attachments, avatars and `jwt` out of the pod (NOT the
27.7 GB registry), then the newest night's three GPG secrets files, writes `MANIFEST.sha256` + `REPOS`, and pushes one
encrypted `dooplex.pxar` to ep0 (`operator` ns, `host/dooplex-gitea`) with the hub-DB's write-only token and a NEW key.
Refuses (no push, no success signal) on: no complete dump / dump > 7 h; file copy failing twice; any symlink or
hardlink in the pod's archive; fewer repositories than the pod lists; no `app.ini`; secrets > 30 h.
- `felhom-dooplex-offsite-restore-test` (Sun 05:30, read-only token): manifest check, repo count, `git fsck` on every
repository through a scratch repo with a known config (the pod's `config` files are never read by root git),
`pg_restore --list`, secrets present.
- `felhom-backup-failmail@.service`: `OnFailure=` mail through R-232 (a)'s `notify_failure`; also added to both
hub-DB units (`scripts/hub-db-backup/*.service`).
- 20 tests (`test_dooplex_offsite.py`; fakes for kubectl/PBS/pg_restore, real git/tar), green with GNU and BusyBox
tools; red-proofs for 6 checks in `audits/dooplex-survival-2026-10-09/partB/red-proof.txt`.
## 2026-10-09 — ep0-copy-gc reads ep0's real namespace answer (found on the first live dry run)
- `scripts/ep0-copy-gc/felhom-ep0-copy-gc`: `proxmox-backup-client namespace list --output-format json` answers
+10
View File
@@ -0,0 +1,10 @@
#!/bin/bash
# felhom-backup-failmail — mail admin@ that a Felhom backup unit on DooPlex failed (R-232 (a) route, reused).
# Called by felhom-backup-failmail@<unit>.service, which the backup units name in OnFailure=.
# Reuses notify_failure from /opt/backup/scripts/backup-config.sh (Resend; never prints the key; never fails).
set -u
UNIT=${1:?unit name}
CONFIG=${FELHOM_FAILMAIL_CONFIG:-/opt/backup/scripts/backup-config.sh}
# shellcheck source=/dev/null
. "$CONFIG"
notify_failure "systemd unit ${UNIT} failed on $(hostname) — see: journalctl -u ${UNIT}"
@@ -0,0 +1,8 @@
# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh.
# Named by OnFailure= in felhom-dooplex-offsite*.service and felhom-hub-db-*.service; %i is the failed unit.
[Unit]
Description=Felhom: mail admin@ that %i failed (R-232)
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/felhom-backup-failmail %i
+109
View File
@@ -0,0 +1,109 @@
#!/bin/sh
# felhom-dooplex-offsite — push Gitea (repositories, database dump, config) and DooPlex's nightly secrets export to
# ep0's PBS, encrypted on DooPlex (R-232 (b)). Runs on DooPlex as root from felhom-dooplex-offsite.timer (00:20).
# Plan: documentation/audits/dooplex-survival-2026-10-09/PLAN.md. Restore: documentation/runbooks/gitea-restore.md.
# Pinned by test_dooplex_offsite.py.
#
# Order is the consistency argument (PLAN.md): the database dump is taken FIRST (the newest complete 6-hourly dump,
# PostgreSQL's own snapshot), the repositories AFTER it, so every commit the database names is in the copy.
#
# Refuses to push — and so never writes the success signal — when: no complete dump exists or the newest is older than
# DUMP_MAX_AGE_H; the dump is empty; the file copy from the pod fails twice; the copy holds no repository or fewer
# repositories than the pod lists; app.ini is missing; no secrets export exists or it is older than SECRETS_MAX_AGE_H.
# The success timestamp is written ONLY after the push returns 0 (CLAUDE.md "presence is not success").
set -eu
CONF=${FELHOM_DXOFF_CONF:-/etc/felhom-dooplex-offsite}
TOKENS=${FELHOM_DXOFF_TOKENS:-/etc/felhom-hub-backup}
STATE=${FELHOM_DXOFF_STATE:-/var/lib/felhom-dooplex-offsite}
TEXTFILE_DIR=${FELHOM_DXOFF_TEXTFILE_DIR:-/var/lib/node_exporter/textfile_collector}
DUMPS=${FELHOM_DXOFF_DUMPS:-/mnt/5_hdd/backup/postgresql/dumps}
SECRETS=${FELHOM_DXOFF_SECRETS:-/mnt/5_hdd/backup/secrets/exports}
DUMP_MAX_AGE_H=${FELHOM_DXOFF_DUMP_MAX_AGE_H:-7}
SECRETS_MAX_AGE_H=${FELHOM_DXOFF_SECRETS_MAX_AGE_H:-30}
NOW=${FELHOM_DXOFF_NOW:-$(date +%s)}
RETRY_SLEEP=${FELHOM_DXOFF_RETRY_SLEEP:-30}
. "$CONF/env" # PBS_REPOSITORY_PUSH, PBS_FINGERPRINT (no secrets in this file)
log() { echo "felhom-dooplex-offsite: $*"; }
die() { echo "felhom-dooplex-offsite: FAILED: $*" >&2; exit 1; }
K() { kubectl -n gitea-system exec deploy/gitea -c gitea -- "$@"; }
umask 077
STAGE="$STATE/stage"
mkdir -p "$STATE"; chmod 700 "$STATE"
rm -rf "$STAGE"; mkdir -p "$STAGE/root/gitea" "$STAGE/root/db" "$STAGE/root/secrets"
cleanup() {
for f in "$STAGE/root/gitea/gitea/conf/app.ini" "$STAGE/root/db/gitea.dump" "$STAGE/root/db/globals.sql"; do
[ -f "$f" ] && [ ! -L "$f" ] && { shred -u "$f" 2>/dev/null || rm -f "$f"; }
done
rm -rf "$STAGE"
}
trap cleanup EXIT
# 1. the database — the newest COMPLETE dump (its folder carries SUCCESS), taken before the files
DUMPDIR=""
for d in $(ls -1d "$DUMPS"/[0-9]*-[0-9]* 2>/dev/null | sort -r); do
if [ -f "$d/SUCCESS" ] && [ -s "$d/gitea.dump" ]; then DUMPDIR=$d; break; fi
done
[ -n "$DUMPDIR" ] || die "no complete gitea.dump under $DUMPS"
DAGE=$((NOW - $(stat -c %Y "$DUMPDIR/SUCCESS")))
[ "$DAGE" -le $((DUMP_MAX_AGE_H * 3600)) ] || die "newest complete dump ${DUMPDIR##*/} is $((DAGE / 3600)) h old (limit ${DUMP_MAX_AGE_H} h) — the dump CronJob stopped"
cp "$DUMPDIR/gitea.dump" "$STAGE/root/db/gitea.dump"
[ -f "$DUMPDIR/globals.sql" ] && cp "$DUMPDIR/globals.sql" "$STAGE/root/db/globals.sql"
echo "${DUMPDIR##*/}" > "$STAGE/root/db/DUMP-FOLDER"
log "database: ${DUMPDIR##*/}, $(wc -c < "$STAGE/root/db/gitea.dump" | tr -d ' ') bytes, $((DAGE / 60)) min old"
# 2. the files — after the dump. Not the registry (packages: rebuilt from the code), logs, indexers, queues, tmp.
PATHS="git/repositories git/lfs gitea/conf/app.ini gitea/attachments gitea/avatars gitea/repo-avatars gitea/jwt"
n=0
until K tar -cf - -C /data $PATHS > "$STAGE/files.tar"; do
n=$((n + 1)); [ "$n" -lt 2 ] || die "copying Gitea's files out of the pod failed twice"
log "file copy failed once (a file moved under a push?) — retrying in ${RETRY_SLEEP} s"; sleep "$RETRY_SLEEP"
done
# The pod's archive is untrusted input to a root process: refuse any symlink or hardlink in it (a bare repository holds
# none), and extract without the pod's owners.
LINKS=$(tar -tvf "$STAGE/files.tar" | grep -c '^[lh]') || LINKS=0
[ "$LINKS" -eq 0 ] || die "the pod's archive holds $LINKS link(s) — refused"
tar -xof "$STAGE/files.tar" -C "$STAGE/root/gitea" || die "unpacking the file copy"
rm -f "$STAGE/files.tar"
[ -s "$STAGE/root/gitea/gitea/conf/app.ini" ] && [ ! -L "$STAGE/root/gitea/gitea/conf/app.ini" ] || die "app.ini missing from the copy"
LISTING=$(K find /data/git/repositories -mindepth 2 -maxdepth 2 -type d -name '*.git') || die "listing repositories in the pod"
WANT=$(printf '%s\n' "$LISTING" | grep -c '\.git$') || WANT=0
GOT=$(find "$STAGE/root/gitea/git/repositories" -mindepth 2 -maxdepth 2 -type d -name '*.git' | wc -l | tr -d ' ')
[ "$GOT" -gt 0 ] || die "the copy holds no repository"
[ "$GOT" -ge "$WANT" ] || die "the copy holds $GOT repositories, the pod lists $WANT"
log "files: $GOT repositories, $(du -sm "$STAGE/root/gitea" | cut -f1) MB"
# 3. the secrets — the newest night's GPG files (already encrypted with DooPlex's restic passphrase)
NEWEST=$(ls -1 "$SECRETS"/secrets-*.yaml.gpg 2>/dev/null | sort | tail -n 1)
[ -n "$NEWEST" ] || die "no secrets export under $SECRETS"
STAMP=${NEWEST##*/secrets-}; STAMP=${STAMP%.yaml.gpg}
SAGE=$((NOW - $(stat -c %Y "$NEWEST")))
[ "$SAGE" -le $((SECRETS_MAX_AGE_H * 3600)) ] || die "newest secrets export is $((SAGE / 3600)) h old (limit ${SECRETS_MAX_AGE_H} h)"
cp "$SECRETS"/*-"$STAMP".*gpg "$STAGE/root/secrets/"
log "secrets: $(ls "$STAGE/root/secrets" | wc -l | tr -d ' ') file(s) of $STAMP"
# 4. the manifest the restore test checks, then the push
echo "$GOT" > "$STAGE/root/REPOS"
(cd "$STAGE/root" && find . -type f ! -name MANIFEST.sha256 -print0 | sort -z | xargs -0 sha256sum > MANIFEST.sha256) \
|| die "writing the manifest"
[ "$(wc -l < "$STAGE/root/MANIFEST.sha256")" -gt "$GOT" ] || die "the manifest is short"
START=$(date +%s)
PBS_PASSWORD_FILE="$TOKENS/token-push" PBS_FINGERPRINT="$PBS_FINGERPRINT" \
proxmox-backup-client backup dooplex.pxar:"$STAGE/root" --ns operator --backup-type host --backup-id dooplex-gitea \
--keyfile "$CONF/enc.key" --crypt-mode encrypt --repository "$PBS_REPOSITORY_PUSH" \
|| die "proxmox-backup-client backup"
BYTES=$(du -sb "$STAGE/root" | cut -f1)
log "pushed to ep0 (ns operator, host/dooplex-gitea) in $(( $(date +%s) - START )) s"
TMP="$TEXTFILE_DIR/felhom_dooplex_offsite.prom.$$"
{
echo "# HELP felhom_dooplex_offsite_last_success_timestamp_seconds Last successful push of Gitea + DooPlex secrets to ep0 (R-232)."
echo "# TYPE felhom_dooplex_offsite_last_success_timestamp_seconds gauge"
echo "felhom_dooplex_offsite_last_success_timestamp_seconds $(date +%s)"
echo "felhom_dooplex_offsite_last_success_bytes $BYTES"
echo "felhom_dooplex_offsite_last_success_repositories $GOT"
} > "$TMP"
chmod 644 "$TMP"
mv "$TMP" "$TEXTFILE_DIR/felhom_dooplex_offsite.prom"
log "success signal written"
@@ -0,0 +1,79 @@
#!/bin/sh
# felhom-dooplex-offsite-restore-test — restore the newest Gitea + secrets copy from ep0 with the READ-ONLY token and
# check it (R-232 (h)). Runs on DooPlex as root from felhom-dooplex-offsite-restore-test.timer (Sun 05:30).
# Pinned by test_dooplex_offsite.py.
#
# The success timestamp is written ONLY when: the newest copy on ep0 is at most MAX_AGE_H old; it restores and
# decrypts; every file matches MANIFEST.sha256; the repository count matches REPOS; `git fsck` passes on EVERY
# repository; `pg_restore --list` reads gitea.dump; app.ini is there; at least one secrets file is there.
set -eu
CONF=${FELHOM_DXOFF_CONF:-/etc/felhom-dooplex-offsite}
TOKENS=${FELHOM_DXOFF_TOKENS:-/etc/felhom-hub-backup}
STATE=${FELHOM_DXOFF_STATE:-/var/lib/felhom-dooplex-offsite}
TEXTFILE_DIR=${FELHOM_DXOFF_TEXTFILE_DIR:-/var/lib/node_exporter/textfile_collector}
MAX_AGE_H=${FELHOM_DXOFF_RESTORE_MAX_AGE_H:-50}
NOW=${FELHOM_DXOFF_NOW:-$(date +%s)}
. "$CONF/env" # PBS_REPOSITORY_RESTORE, PBS_FINGERPRINT
log() { echo "felhom-dooplex-offsite-restore-test: $*"; }
die() { echo "felhom-dooplex-offsite-restore-test: FAILED: $*" >&2; exit 1; }
umask 077
mkdir -p "$STATE"; chmod 700 "$STATE"
T=$(mktemp -d "$STATE/restore.XXXXXX")
trap 'for f in "$T"/out/gitea/gitea/conf/app.ini "$T"/out/db/gitea.dump "$T"/out/db/globals.sql; do [ -f "$f" ] && shred -u "$f" 2>/dev/null; done; rm -rf "$T"' EXIT
export PBS_PASSWORD_FILE="$TOKENS/token-restore" PBS_FINGERPRINT
LIST=$(proxmox-backup-client snapshot list host/dooplex-gitea --ns operator --output-format json --repository "$PBS_REPOSITORY_RESTORE") \
|| die "listing snapshots on ep0"
NEWEST=$(printf '%s' "$LIST" | python3 -c '
import json, sys
s = [x for x in json.load(sys.stdin) if x.get("backup-type") == "host" and x.get("backup-id") == "dooplex-gitea"]
if s:
print(max(x["backup-time"] for x in s))
') || die "reading the snapshot list"
[ -n "$NEWEST" ] || die "no Gitea copy on ep0"
AGE=$((NOW - NEWEST))
[ "$AGE" -le $((MAX_AGE_H * 3600)) ] || die "newest copy on ep0 is $((AGE / 3600)) h old (limit ${MAX_AGE_H} h)"
SNAPSHOT="host/dooplex-gitea/$(date -u -d "@$NEWEST" +%Y-%m-%dT%H:%M:%SZ)"
log "restoring $SNAPSHOT"
proxmox-backup-client restore "$SNAPSHOT" dooplex.pxar "$T/out" --ns operator \
--keyfile "$CONF/enc.key" --repository "$PBS_REPOSITORY_RESTORE" || die "restore of $SNAPSHOT"
O="$T/out"
[ -s "$O/MANIFEST.sha256" ] || die "the copy holds no MANIFEST.sha256"
(cd "$O" && sha256sum -c MANIFEST.sha256 >/dev/null) || die "a file does not match MANIFEST.sha256"
FILES=$(wc -l < "$O/MANIFEST.sha256" | tr -d ' ')
WANT=$(cat "$O/REPOS" 2>/dev/null) || die "the copy holds no REPOS count"
REPOS=$(find "$O/gitea/git/repositories" -mindepth 2 -maxdepth 2 -type d -name '*.git' | sort)
GOT=$(printf '%s\n' "$REPOS" | grep -c '\.git$') || GOT=0
[ "$GOT" -gt 0 ] && [ "$GOT" = "$WANT" ] || die "the copy holds $GOT repositories, REPOS says $WANT"
# The repositories' own `config` files came from the Gitea pod — untrusted input to a root git. So git never reads
# them: each repository is checked through a fresh scratch repository with a known config, holding a copy of its
# HEAD and refs, with the copy's objects as its object store. No system or global config either.
export GIT_CONFIG_NOSYSTEM=1 GIT_CONFIG_GLOBAL=/dev/null
for r in $REPOS; do
S="$T/fsck.git"; rm -rf "$S"
git init -q --bare "$S" || die "git init for the fsck scratch repository"
cp "$r/HEAD" "$S/HEAD"; [ -f "$r/packed-refs" ] && cp "$r/packed-refs" "$S/packed-refs"
[ -d "$r/refs" ] && cp -R "$r/refs/." "$S/refs/"
GIT_OBJECT_DIRECTORY="$r/objects" git --git-dir="$S" -c core.hooksPath=/dev/null -c core.fsmonitor=false \
fsck --no-progress --no-dangling >/dev/null 2>"$T/fsck.err" \
|| die "git fsck ${r#"$O"/gitea/git/repositories/}: $(head -n 3 "$T/fsck.err")"
done
rm -rf "$T/fsck.git"
[ -s "$O/gitea/gitea/conf/app.ini" ] || die "app.ini missing"
pg_restore --list "$O/db/gitea.dump" >/dev/null || die "pg_restore cannot read gitea.dump"
ls "$O"/secrets/*.gpg >/dev/null 2>&1 || die "no secrets file in the copy"
log "checked: $FILES files match the manifest, $GOT repositories pass git fsck, gitea.dump readable, $(ls "$O"/secrets | wc -l | tr -d ' ') secrets file(s)"
TMP="$TEXTFILE_DIR/felhom_dooplex_offsite_restore.prom.$$"
{
echo "# HELP felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds Last successful restore test of the Gitea + secrets copy on ep0 (R-232)."
echo "# TYPE felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds gauge"
echo "felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds $(date +%s)"
} > "$TMP"
chmod 644 "$TMP"
mv "$TMP" "$TEXTFILE_DIR/felhom_dooplex_offsite_restore.prom"
log "success signal written"
@@ -0,0 +1,20 @@
# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh. Plan: audits/dooplex-survival-2026-10-09/PLAN.md.
[Unit]
Description=Felhom: restore-test the Gitea + secrets copy on ep0 (R-232)
Wants=network-online.target felhom-ep0-pbs-tunnel.service
After=network-online.target felhom-ep0-pbs-tunnel.service
OnFailure=felhom-backup-failmail@%n.service
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/felhom-dooplex-offsite-restore-test
Environment=HOME=/var/lib/felhom-dooplex-offsite KUBECONFIG=/etc/rancher/k3s/k3s.yaml
UMask=0077
TimeoutStartSec=90min
Nice=10
IOSchedulingClass=idle
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
ReadWritePaths=/var/lib/felhom-dooplex-offsite /var/lib/node_exporter/textfile_collector
NoNewPrivileges=yes
@@ -0,0 +1,11 @@
# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh.
[Unit]
Description=Felhom: restore-test the Gitea + secrets copy on ep0 (R-232) — schedule
[Timer]
OnCalendar=Sun *-*-* 05:30:00
Persistent=true
RandomizedDelaySec=2min
[Install]
WantedBy=timers.target
@@ -0,0 +1,20 @@
# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh. Plan: audits/dooplex-survival-2026-10-09/PLAN.md.
[Unit]
Description=Felhom: push Gitea + DooPlex secrets to ep0, encrypted (R-232)
Wants=network-online.target felhom-ep0-pbs-tunnel.service
After=network-online.target felhom-ep0-pbs-tunnel.service
OnFailure=felhom-backup-failmail@%n.service
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/felhom-dooplex-offsite
Environment=HOME=/var/lib/felhom-dooplex-offsite KUBECONFIG=/etc/rancher/k3s/k3s.yaml
UMask=0077
TimeoutStartSec=90min
Nice=10
IOSchedulingClass=idle
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
ReadWritePaths=/var/lib/felhom-dooplex-offsite /var/lib/node_exporter/textfile_collector
NoNewPrivileges=yes
@@ -0,0 +1,11 @@
# Versioned in felhom.eu/scripts/dooplex-offsite/ (R-232); installed by install.sh.
[Unit]
Description=Felhom: push Gitea + DooPlex secrets to ep0 (R-232) — schedule
[Timer]
OnCalendar=*-*-* 00:20:00
Persistent=true
RandomizedDelaySec=2min
[Install]
WantedBy=timers.target
+24
View File
@@ -0,0 +1,24 @@
#!/bin/sh
# install.sh — install the Gitea + secrets off-site units on DooPlex (R-232). Root. Idempotent.
# Installs the three scripts and five units, writes /etc/felhom-dooplex-offsite/env (no secrets) when absent, and does
# NOT enable the timers — enable them by hand after the first manual run:
# systemctl enable --now felhom-dooplex-offsite.timer felhom-dooplex-offsite-restore-test.timer
# The tokens are the hub-DB ones (/etc/felhom-hub-backup/token-push, token-restore), read in place. enc.key is created
# separately (proxmox-backup-client key create --kdf none), never by this script.
set -eu
HERE=$(cd "$(dirname "$0")" && pwd)
[ "$(id -u)" = 0 ] || { echo "install.sh: run as root" >&2; exit 1; }
for s in felhom-dooplex-offsite felhom-dooplex-offsite-restore-test felhom-backup-failmail; do
install -m 0755 "$HERE/$s" "/usr/local/sbin/$s"
done
for u in felhom-dooplex-offsite.service felhom-dooplex-offsite.timer felhom-dooplex-offsite-restore-test.service \
felhom-dooplex-offsite-restore-test.timer felhom-backup-failmail@.service; do
install -m 0644 "$HERE/$u" "/etc/systemd/system/$u"
done
install -d -m 0700 /etc/felhom-dooplex-offsite /var/lib/felhom-dooplex-offsite
if [ ! -f /etc/felhom-dooplex-offsite/env ]; then
umask 077
grep -E '^(PBS_REPOSITORY_PUSH|PBS_REPOSITORY_RESTORE|PBS_FINGERPRINT)=' /etc/felhom-hub-backup/env > /etc/felhom-dooplex-offsite/env
fi
systemctl daemon-reload
echo "install.sh: installed; timers NOT enabled (see the header)"
@@ -0,0 +1,332 @@
#!/usr/bin/env python3
"""Tests for felhom-dooplex-offsite, its restore test and felhom-backup-failmail (R-232).
No test reaches Gitea, k3s, PBS, ep0 or Resend: `kubectl`, `proxmox-backup-client` and `pg_restore` are fakes on PATH
(the Gitea pod's /data is a temp dir; the PBS "server" is a temp dir; the fake pg_restore accepts a file that starts
with PostgreSQL's custom-dump magic "PGDMP"). `git`, `tar`, `sha256sum`, `shred`, `find` are the real tools, so
`git fsck` runs on real repositories. Each test asserts the CONSEQUENCE: whether a push happened and whether the
success signal (the file the alarm reads) was written. Run: python3 scripts/dooplex-offsite/test_dooplex_offsite.py
"""
import json
import os
import shutil
import subprocess
import tempfile
import time
import unittest
HERE = os.path.dirname(os.path.abspath(__file__))
PUSH = os.path.join(HERE, "felhom-dooplex-offsite")
RESTORE = os.path.join(HERE, "felhom-dooplex-offsite-restore-test")
FAILMAIL = os.path.join(HERE, "felhom-backup-failmail")
FAKE_KUBECTL = r'''#!/usr/bin/env python3
import os, subprocess, sys
a = sys.argv[1:]
pod = os.environ["FAKE_POD_DATA"]
cmd = a[a.index("--") + 1:]
if cmd[0] == "tar":
cnt = os.path.join(os.environ["FAKE_STATE"], "tar-calls")
n = int(open(cnt).read()) if os.path.exists(cnt) else 0
open(cnt, "w").write(str(n + 1))
if n < int(os.environ.get("FAKE_TAR_FAILS", "0")):
sys.stderr.write("tar: file vanished\n"); sys.exit(1)
c = ["tar" if x == "tar" else (pod if x == "/data" else x) for x in cmd]
sys.exit(subprocess.call(c))
if cmd[0] == "find":
out = subprocess.run(["find", pod + cmd[1][len("/data"):]] + cmd[2:],
capture_output=True, text=True)
for line in out.stdout.splitlines():
print("/data" + line[len(pod):])
for extra in filter(None, os.environ.get("FAKE_EXTRA_REPOS", "").split(",")):
print("/data/git/repositories/admin/" + extra)
sys.exit(out.returncode)
sys.exit(97)
'''
FAKE_PBS = r'''#!/usr/bin/env python3
import json, os, shutil, sys, time
a = sys.argv[1:]
srv = os.environ["FAKE_PBS_DIR"]
open(os.path.join(srv, "calls.log"), "a").write(json.dumps({"argv": a, "pw": os.environ.get("PBS_PASSWORD_FILE", "")}) + "\n")
base = os.path.join(srv, "snaps")
if a[0] == "backup":
if os.environ.get("FAKE_PBS_FAIL"): sys.exit(1)
src = a[1].split(":", 1)[1]
t = int(os.environ.get("FAKE_PBS_TIME", time.time()))
shutil.copytree(src, os.path.join(base, str(t)))
sys.exit(0)
if a[0] == "snapshot" and a[1] == "list":
out = [{"backup-type": "host", "backup-id": "dooplex-gitea", "backup-time": int(x)} for x in (os.listdir(base) if os.path.isdir(base) else [])]
print(json.dumps(out)); sys.exit(0)
if a[0] == "restore":
import calendar
t = calendar.timegm(time.strptime(a[1].split("/")[-1], "%Y-%m-%dT%H:%M:%SZ"))
shutil.copytree(os.path.join(base, str(t)), a[3])
sys.exit(0)
sys.exit(98)
'''
FAKE_PG_RESTORE = r'''#!/bin/sh
[ "$1" = "--list" ] || exit 97
head -c 5 "$2" | grep -q '^PGDMP' || { echo "pg_restore: error: input file does not appear to be a valid archive" >&2; exit 1; }
'''
def git(*a, cwd=None):
env = dict(os.environ, GIT_AUTHOR_NAME="t", GIT_AUTHOR_EMAIL="t@t", GIT_COMMITTER_NAME="t", GIT_COMMITTER_EMAIL="t@t")
subprocess.run(["git", *a], cwd=cwd, check=True, capture_output=True, env=env)
class Base(unittest.TestCase):
def setUp(self):
self.t = tempfile.mkdtemp()
j = lambda *p: os.path.join(self.t, *p)
self.pod, self.pbs, self.state, self.text = j("pod"), j("pbs"), j("state"), j("textfile")
self.conf, self.tokens, self.dumps, self.secrets, self.bin = j("conf"), j("tokens"), j("dumps"), j("secrets"), j("bin")
for d in (self.pbs, self.state, self.text, self.conf, self.tokens, self.dumps, self.secrets, self.bin):
os.makedirs(d)
# the Gitea pod: two real repositories with a commit each, app.ini, the other paths
for name in ("felhom.eu", "felhom-agent"):
bare = j("pod", "git", "repositories", "admin", name + ".git")
os.makedirs(os.path.dirname(bare), exist_ok=True)
git("init", "-q", "--bare", bare)
work = j("work-" + name)
git("init", "-q", work)
open(os.path.join(work, "README"), "w").write(name + "\n")
git("add", "README", cwd=work); git("commit", "-q", "-m", "c", cwd=work)
git("push", "-q", bare, "HEAD:refs/heads/main", cwd=work)
for p in ("git/lfs", "gitea/conf", "gitea/attachments", "gitea/avatars", "gitea/repo-avatars", "gitea/jwt", "gitea/packages"):
os.makedirs(j("pod", p), exist_ok=True)
open(j("pod", "gitea/conf/app.ini"), "w").write("[database]\nDB_TYPE = postgres\n")
open(j("pod", "gitea/packages/blob"), "w").write("registry - must not be copied\n")
# a complete dump (00:00 local) and the nightly secrets export
self.now = int(time.time())
self.add_dump("20261008-220001", self.now - 1200)
self.add_secrets("20261008_031011", self.now - 21 * 3600)
for n in ("token-push", "token-restore"):
open(os.path.join(self.tokens, n), "w").write("x")
open(os.path.join(self.conf, "enc.key"), "w").write("{}")
open(os.path.join(self.conf, "env"), "w").write(
"PBS_REPOSITORY_PUSH='u!push@h:1:s'\nPBS_REPOSITORY_RESTORE='u!restore@h:1:s'\nPBS_FINGERPRINT='aa'\n")
for name, body in (("kubectl", FAKE_KUBECTL), ("proxmox-backup-client", FAKE_PBS), ("pg_restore", FAKE_PG_RESTORE)):
p = os.path.join(self.bin, name)
open(p, "w").write(body); os.chmod(p, 0o755)
def tearDown(self):
shutil.rmtree(self.t, ignore_errors=True)
def add_dump(self, name, mtime, complete=True, magic=b"PGDMP"):
d = os.path.join(self.dumps, name); os.makedirs(d)
open(os.path.join(d, "gitea.dump"), "wb").write(magic + b"\x01dump")
open(os.path.join(d, "globals.sql"), "w").write("-- roles\n")
if complete:
open(os.path.join(d, "SUCCESS"), "w").close(); os.utime(os.path.join(d, "SUCCESS"), (mtime, mtime))
def add_secrets(self, stamp, mtime):
for k, ext in (("secrets", "yaml.gpg"), ("configmaps", "yaml.gpg"), ("by-namespace", "tar.gz.gpg")):
p = os.path.join(self.secrets, "%s-%s.%s" % (k, stamp, ext))
open(p, "wb").write(b"gpg-" + k.encode()); os.utime(p, (mtime, mtime))
def env(self, **kw):
e = dict(os.environ, PATH=self.bin + ":" + os.environ["PATH"], FAKE_POD_DATA=self.pod, FAKE_PBS_DIR=self.pbs,
FAKE_STATE=self.t, FELHOM_DXOFF_CONF=self.conf, FELHOM_DXOFF_TOKENS=self.tokens,
FELHOM_DXOFF_STATE=self.state, FELHOM_DXOFF_TEXTFILE_DIR=self.text, FELHOM_DXOFF_DUMPS=self.dumps,
FELHOM_DXOFF_SECRETS=self.secrets)
e.update({k: str(v) for k, v in kw.items()})
return e
def push(self, **kw):
return subprocess.run([PUSH], env=self.env(**kw), capture_output=True, text=True)
def restore(self, **kw):
return subprocess.run([RESTORE], env=self.env(**kw), capture_output=True, text=True)
def pushed(self):
d = os.path.join(self.pbs, "snaps")
return sorted(os.listdir(d)) if os.path.isdir(d) else []
def signal(self, name="felhom_dooplex_offsite.prom"):
return os.path.exists(os.path.join(self.text, name))
class Push(Base):
def test_happy_path_pushes_everything_but_the_registry_and_writes_the_signal(self):
r = self.push()
self.assertEqual(r.returncode, 0, r.stderr)
self.assertEqual(len(self.pushed()), 1)
snap = os.path.join(self.pbs, "snaps", self.pushed()[0])
self.assertTrue(os.path.isfile(os.path.join(snap, "gitea/git/repositories/admin/felhom.eu.git/HEAD")))
self.assertTrue(os.path.isfile(os.path.join(snap, "gitea/gitea/conf/app.ini")))
self.assertFalse(os.path.exists(os.path.join(snap, "gitea/gitea/packages")), "the registry must stay out")
self.assertEqual(open(os.path.join(snap, "db/DUMP-FOLDER")).read().strip(), "20261008-220001")
self.assertEqual(len(os.listdir(os.path.join(snap, "secrets"))), 3)
self.assertEqual(open(os.path.join(snap, "REPOS")).read().strip(), "2")
self.assertTrue(self.signal())
call = json.loads(open(os.path.join(self.pbs, "calls.log")).readline())
self.assertIn("--crypt-mode", call["argv"]); self.assertIn("encrypt", call["argv"])
self.assertTrue(call["pw"].endswith("token-push"))
self.assertFalse(os.path.exists(os.path.join(self.state, "stage")), "the stage must be removed")
def test_newest_incomplete_dump_is_skipped_for_the_complete_one(self):
self.add_dump("20261009-040001", self.now - 60, complete=False)
r = self.push()
self.assertEqual(r.returncode, 0, r.stderr)
snap = os.path.join(self.pbs, "snaps", self.pushed()[0])
self.assertEqual(open(os.path.join(snap, "db/DUMP-FOLDER")).read().strip(), "20261008-220001")
def test_stale_dump_refuses(self):
shutil.rmtree(self.dumps); os.makedirs(self.dumps)
self.add_dump("20261008-040001", self.now - 8 * 3600)
r = self.push()
self.assertNotEqual(r.returncode, 0)
self.assertIn("dump CronJob stopped", r.stderr)
self.assertEqual(self.pushed(), []); self.assertFalse(self.signal())
def test_no_complete_dump_refuses(self):
shutil.rmtree(self.dumps); os.makedirs(self.dumps)
self.add_dump("20261009-040001", self.now, complete=False)
r = self.push()
self.assertNotEqual(r.returncode, 0)
self.assertEqual(self.pushed(), []); self.assertFalse(self.signal())
def test_one_failed_file_copy_is_retried(self):
r = self.push(FAKE_TAR_FAILS=1, FELHOM_DXOFF_RETRY_SLEEP=0)
self.assertEqual(r.returncode, 0, r.stderr)
self.assertTrue(self.signal())
def test_two_failed_file_copies_refuse(self):
r = self.push(FAKE_TAR_FAILS=2, FELHOM_DXOFF_RETRY_SLEEP=0)
self.assertNotEqual(r.returncode, 0)
self.assertIn("failed twice", r.stderr)
self.assertEqual(self.pushed(), []); self.assertFalse(self.signal())
def test_fewer_repositories_than_the_pod_lists_refuses(self):
r = self.push(FAKE_EXTRA_REPOS="ghost.git")
self.assertNotEqual(r.returncode, 0)
self.assertIn("the pod lists 3", r.stderr)
self.assertEqual(self.pushed(), []); self.assertFalse(self.signal())
def test_missing_app_ini_refuses(self):
os.remove(os.path.join(self.pod, "gitea/conf/app.ini"))
r = self.push()
self.assertNotEqual(r.returncode, 0)
self.assertEqual(self.pushed(), []); self.assertFalse(self.signal())
def test_stale_secrets_refuse(self):
shutil.rmtree(self.secrets); os.makedirs(self.secrets)
self.add_secrets("20261007_031011", self.now - 31 * 3600)
r = self.push()
self.assertNotEqual(r.returncode, 0)
self.assertIn("secrets export", r.stderr)
self.assertEqual(self.pushed(), []); self.assertFalse(self.signal())
def test_a_link_in_the_pods_archive_refuses(self):
os.symlink("/etc/passwd", os.path.join(self.pod, "gitea/avatars/evil"))
r = self.push()
self.assertNotEqual(r.returncode, 0)
self.assertIn("link(s)", r.stderr)
self.assertEqual(self.pushed(), []); self.assertFalse(self.signal())
def test_failed_push_writes_no_signal(self):
r = self.push(FAKE_PBS_FAIL=1)
self.assertNotEqual(r.returncode, 0)
self.assertFalse(self.signal())
class RestoreTest(Base):
def pushed_copy(self, **kw):
r = self.push(FAKE_PBS_TIME=self.now - 3600, **kw)
self.assertEqual(r.returncode, 0, r.stderr)
return os.path.join(self.pbs, "snaps", self.pushed()[0])
def test_happy_path_writes_the_signal_with_the_readonly_token(self):
self.pushed_copy()
r = self.restore()
self.assertEqual(r.returncode, 0, r.stderr)
self.assertIn("2 repositories pass git fsck", r.stdout)
self.assertTrue(self.signal("felhom_dooplex_offsite_restore.prom"))
calls = [json.loads(x) for x in open(os.path.join(self.pbs, "calls.log"))]
self.assertTrue(all(c["pw"].endswith("token-restore") for c in calls if c["argv"][0] != "backup"))
def test_a_changed_file_fails_the_manifest(self):
snap = self.pushed_copy()
open(os.path.join(snap, "secrets", sorted(os.listdir(os.path.join(snap, "secrets")))[0]), "ab").write(b"x")
r = self.restore()
self.assertNotEqual(r.returncode, 0)
self.assertIn("MANIFEST", r.stderr)
self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom"))
def test_a_broken_repository_fails_git_fsck(self):
snap = self.pushed_copy()
repo = os.path.join(snap, "gitea/git/repositories/admin/felhom-agent.git")
objs = [os.path.join(dp, f) for dp, _, fs in os.walk(os.path.join(repo, "objects")) for f in fs
if len(os.path.basename(dp)) == 2]
os.remove(objs[0])
# keep the manifest honest about the removal, so ONLY git fsck can catch it
man = os.path.join(snap, "MANIFEST.sha256")
rel = "./" + os.path.relpath(objs[0], snap)
kept = [l for l in open(man).readlines() if not l.rstrip().endswith(rel)]
open(man, "w").writelines(kept)
r = self.restore()
self.assertNotEqual(r.returncode, 0)
self.assertIn("git fsck", r.stderr)
self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom"))
def test_git_never_reads_a_repositorys_own_config(self):
# The pod's `config` files are untrusted input to a root git. A config git would refuse to open
# (repositoryformatversion 99) proves git never read it: the old `git -C <repo> fsck` failed here (red-proof).
snap = self.pushed_copy()
cfg = os.path.join(snap, "gitea/git/repositories/admin/felhom.eu.git/config")
open(cfg, "w").write("[core]\n\trepositoryformatversion = 99\n\tbare = true\n\tfsmonitor = touch /nonexistent/x\n")
man = os.path.join(snap, "MANIFEST.sha256")
kept = [l for l in open(man).readlines() if not l.rstrip().endswith("felhom.eu.git/config")]
open(man, "w").writelines(kept)
r = self.restore()
self.assertEqual(r.returncode, 0, r.stderr)
def test_an_unreadable_dump_fails(self):
shutil.rmtree(self.dumps); os.makedirs(self.dumps)
self.add_dump("20261008-220001", self.now - 1200, magic=b"XXXXX")
self.pushed_copy()
r = self.restore()
self.assertNotEqual(r.returncode, 0)
self.assertIn("pg_restore", r.stderr)
self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom"))
def test_an_old_copy_fails(self):
r = self.push(FAKE_PBS_TIME=self.now - 51 * 3600)
self.assertEqual(r.returncode, 0, r.stderr)
r = self.restore()
self.assertNotEqual(r.returncode, 0)
self.assertIn("h old", r.stderr)
self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom"))
def test_no_copy_fails(self):
r = self.restore()
self.assertNotEqual(r.returncode, 0)
self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom"))
class FailMail(unittest.TestCase):
def test_failmail_calls_notify_failure_with_the_unit_name(self):
t = tempfile.mkdtemp()
try:
cfg = os.path.join(t, "cfg.sh"); out = os.path.join(t, "out")
open(cfg, "w").write('notify_failure() { echo "$1" > %s; return 0; }\n' % out)
r = subprocess.run([FAILMAIL, "felhom-dooplex-offsite.service"], env=dict(os.environ, FELHOM_FAILMAIL_CONFIG=cfg),
capture_output=True, text=True)
self.assertEqual(r.returncode, 0, r.stderr)
self.assertIn("felhom-dooplex-offsite.service failed", open(out).read())
finally:
shutil.rmtree(t)
def test_every_backup_unit_names_the_failmail(self):
units = [os.path.join(HERE, u) for u in ("felhom-dooplex-offsite.service", "felhom-dooplex-offsite-restore-test.service")]
units += [os.path.join(HERE, "..", "hub-db-backup", u) for u in ("felhom-hub-db-backup.service", "felhom-hub-db-restore-test.service")]
for u in units:
self.assertIn("OnFailure=felhom-backup-failmail@%n.service", open(u).read(), u)
if __name__ == "__main__":
unittest.main(verbosity=2)
@@ -3,6 +3,7 @@
Description=Felhom: push the hub DB snapshot to ep0 (R-173)
Wants=network-online.target felhom-ep0-pbs-tunnel.service
After=network-online.target felhom-ep0-pbs-tunnel.service
OnFailure=felhom-backup-failmail@%n.service
[Service]
Type=oneshot
@@ -3,6 +3,7 @@
Description=Felhom: restore-test the hub DB copy on ep0 (R-173)
Wants=network-online.target felhom-ep0-pbs-tunnel.service
After=network-online.target felhom-ep0-pbs-tunnel.service
OnFailure=felhom-backup-failmail@%n.service
[Service]
Type=oneshot