R-232: Gitea restored from the ep0 copy into a throwaway (bench 9401) — runbooks/gitea-restore.md; Part B alarm + install + failmail evidence
gates / gates (push) Successful in 5m33s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-09 10:27:20 +02:00
parent 97d3c29f9c
commit 994826ec28
10 changed files with 296 additions and 0 deletions
@@ -0,0 +1,25 @@
## promtool, in pod/prometheus-55b675779d-t8c74, 2026-10-09T08:18:23Z
### green
SUCCESS
### red: threshold 26h -> 260h
FAILED:
alertname: DooplexGiteaOffsiteStale, time: 1d2h40m,
Labels:{alertname="DooplexGiteaOffsiteStale", component="backup", instance="dooplex", severity="critical"}
### red: absent() removed from both
0
FAILED:
alertname: DooplexGiteaOffsiteStale, time: 40m,
Labels:{alertname="DooplexGiteaOffsiteStale", component="backup", severity="critical"}
alertname: DooplexGiteaRestoreTestStale, time: 2h,
Labels:{alertname="DooplexGiteaRestoreTestStale", component="backup", severity="warning"}
### green again
SUCCESS
## 2026-10-09T08:20:27Z reload + /api/v1/rules (homelab-manifests 691db39)
reload http 200
backup-freshness HubDBBackupStale unknown unknown
backup-freshness HubDBRestoreTestStale unknown unknown
backup-freshness DooplexGiteaOffsiteStale unknown unknown
backup-freshness DooplexGiteaRestoreTestStale unknown unknown
## the metrics Prometheus scrapes
[[1791534030.947, '1791533573']]
[[1791534030.979, '1791533623']]
@@ -0,0 +1,55 @@
rule_files:
- bf.yml
evaluation_interval: 1m
tests:
- interval: 5m
input_series:
- series: felhom_dooplex_offsite_last_success_timestamp_seconds{instance="dooplex"}
values: 0+0x400
- series: felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds{instance="dooplex"}
values: 0+0x400
alert_rule_test:
- eval_time: 26h
alertname: DooplexGiteaOffsiteStale
exp_alerts: []
- eval_time: 26h40m
alertname: DooplexGiteaOffsiteStale
exp_alerts:
- exp_labels:
severity: critical
component: backup
instance: dooplex
exp_annotations: &id001
summary: Gitea and DooPlex's secrets have not reached ep0 for 26 h (R-232)
description: No successful off-site push of Gitea (repositories, database dump, config) and the secrets export for >26h (daily at 00:20). Check `journalctl -u felhom-dooplex-offsite.service` and the tunnel `systemctl status felhom-ep0-pbs-tunnel`.
- interval: 1h
input_series:
- series: felhom_dooplex_offsite_last_success_timestamp_seconds{instance="dooplex"}
values: 0x23 86400x23 172800x3
- series: felhom_dooplex_offsite_restore_test_last_success_timestamp_seconds{instance="dooplex"}
values: 0+0x50
alert_rule_test:
- eval_time: 50h
alertname: DooplexGiteaOffsiteStale
exp_alerts: []
- interval: 5m
input_series:
- series: up{job="node"}
values: 1+0x30
alert_rule_test:
- eval_time: 40m
alertname: DooplexGiteaOffsiteStale
exp_alerts:
- exp_labels:
severity: critical
component: backup
exp_annotations: *id001
- eval_time: 2h
alertname: DooplexGiteaRestoreTestStale
exp_alerts:
- exp_labels:
severity: warning
component: backup
exp_annotations:
summary: The Gitea copy on ep0 has not passed a restore test for 8 days (R-232)
description: The weekly restore test (Sun 05:30) has not succeeded for >8 days. Check `journalctl -u felhom-dooplex-offsite-restore-test.service`.
@@ -10,3 +10,14 @@
2026-10-09T10:14:08+02:00 dooplex systemd[1]: felhom-backup-failmail@felhom-failmail-dryrun-r232.service.service: Main process exited, code=exited, status=1/FAILURE
2026-10-09T10:14:08+02:00 dooplex systemd[1]: felhom-backup-failmail@felhom-failmail-dryrun-r232.service.service: Failed with result 'exit-code'.
2026-10-09T10:14:08+02:00 dooplex systemd[1]: Failed to start felhom-backup-failmail@felhom-failmail-dryrun-r232.service.service - Felhom: mail admin@ that felhom-failmail-dryrun-r232.service failed (R-232).
## 2026-10-09T08:17:11Z dry failure 2 (felhom.eu 02a54e26)
2026-10-09T10:16:59+02:00 dooplex systemd[1]: Started felhom-failmail-dryrun-r232b.service - [systemd-run] /bin/false.
2026-10-09T10:16:59+02:00 dooplex systemd[1]: felhom-failmail-dryrun-r232b.service: Main process exited, code=exited, status=1/FAILURE
2026-10-09T10:16:59+02:00 dooplex systemd[1]: felhom-failmail-dryrun-r232b.service: Failed with result 'exit-code'.
2026-10-09T10:16:59+02:00 dooplex systemd[1]: felhom-failmail-dryrun-r232b.service: Triggering OnFailure= dependencies.
2026-10-09T10:16:59+02:00 dooplex systemd[1]: Starting felhom-backup-failmail@felhom-failmail-dryrun-r232b.service.service - Felhom: mail admin@ that felhom-failmail-dryrun-r232b.service failed (R-232)...
2026-10-09T10:16:59+02:00 dooplex felhom-backup-failmail[3059936]: notify_failure: mail accepted id=01a11fbc-a750-7b78-9901-0f98776939f3
2026-10-09T10:16:59+02:00 dooplex felhom-backup-failmail[3059976]: [2026-10-09 10:16:59] [INFO] notify_failure: failure mail sent to admin@felhom.eu
2026-10-09T10:16:59+02:00 dooplex systemd[1]: felhom-backup-failmail@felhom-failmail-dryrun-r232b.service.service: Deactivated successfully.
2026-10-09T10:16:59+02:00 dooplex systemd[1]: Finished felhom-backup-failmail@felhom-failmail-dryrun-r232b.service.service - Felhom: mail admin@ that felhom-failmail-dryrun-r232b.service failed (R-232).
## inbox (second channel, Gmail connector): 2026-10-09T08:17:00Z from monitoring@felhom.eu to admin@felhom.eu, subject '[DooPlex backup] FAILED: systemd unit felhom-failmail-dryrun-r232b.service failed on dooplex — ...'
@@ -14,3 +14,12 @@ PBS_FINGERPRINT=<pinned, as hub-backup>
OnFailure=felhom-backup-failmail@felhom-hub-db-backup.service.service
OnFailure=felhom-backup-failmail@felhom-dooplex-offsite.service.service
## 2026-10-09T08:20:39Z timers enabled
Sat 2026-10-10 00:21:25 CEST 14h - - felhom-dooplex-offsite.timer felhom-dooplex-offsite.service
Sat 2026-10-10 02:30:23 CEST 16h Fri 2026-10-09 02:31:36 CEST 7h ago felhom-hub-db-backup.timer felhom-hub-db-backup.service
Sun 2026-10-11 04:30:24 CEST 1 day 18h - - felhom-hub-db-restore-test.timer felhom-hub-db-restore-test.service
Sun 2026-10-11 05:31:46 CEST 1 day 19h - - felhom-dooplex-offsite-restore-test.timer felhom-dooplex-offsite-restore-test.service
HubDBBackupStale inactive ok
HubDBRestoreTestStale inactive ok
DooplexGiteaOffsiteStale inactive ok
DooplexGiteaRestoreTestStale inactive ok
@@ -0,0 +1,28 @@
## live Gitea main, 2026-10-09T08:22:32Z (git ls-remote against gitea.dooplex.hu)
felhom.eu 02a54e26f5ab118f8dd19c79452ce88d17307b8e
felhom-controller 13eeb484abdfec2fae97559cb1a374234fe27565
felhom-agent 24ea960229e16c77bb312bcb439c023c267f8f6a
app-catalog-felhom.eu 31e96515910816ab6c1b59cc38c1567e8b347771
## restore on DooPlex with the READ-ONLY token
| snapshot | size | files |
| host/dooplex-gitea/2026-10-09T08:11:56Z | 545.424 MiB | catalog.pcat1 dooplex.pxar index.json |
Using encryption key from '/etc/felhom-dooplex-offsite/enc.key'..
Fingerprint: 93:03:bf:d7:1f:4c:9e:fe
progress 19% (107.599 MiB of 544.173 MiB in 5s, 21.39 MiB/s)
progress 66% (360.433 MiB of 544.173 MiB in 15.1s, 25.056 MiB/s)
restore complete (544.173 MiB processed in 21.9s, average 24.795 MiB/s)
5.55user 4.03system 0:22.12elapsed 43%CPU (0avgtext+0avgdata 206676maxresident)k
440inputs+1268544outputs (0major+583210minor)pagefaults 0swaps
manifest OK: 27805 files
20261009-040001
10
app-catalog-drill.git
app-catalog-felhom.eu.git
felhom-agent.git
felhom-controller.git
felhom.eu.git
homelab-manifests.git
jarr.git
misc-scripts.git
recipe-importer.git
revfulop-calendar.git
@@ -0,0 +1,7 @@
## 2026-10-09T08:23:08Z copy to bench 9401 (no secrets/)
625M /root/gr
MANIFEST.sha256
REPOS
db
gitea
bench-manifest-OK
@@ -0,0 +1,35 @@
## 2026-10-09T08:23:47Z pull (the only network step)
docker.io/library/postgres:17.2
docker.io/gitea/gitea:1.26.2
gr-net internal=true
## postgres
repository rows|10
user rows|1
## gitea config: database -> gr-db, throwaway DB password, mailer off
DB_TYPE = postgres
HOST = gr-db:5432
NAME = gitea
USER = gitea
USER = <operator mail address, redacted>
mailer ENABLED = false
## healthz after ~6 s
{
"status": "pass",
"description": "Gitea: Git with a cup of tea",
"checks": {
"cache:ping": [
{
"status": "pass",
"time": "2026-10-09T08:24:16Z"
}
],
"database:ping": [
{
"status": "pass",
"time": "2026-10-09T08:24:16Z"
}
## no route out (from inside the Gitea container)
outside unreachable (rc=1)
wget: bad address 'gitea.com'
172.19.0.0/16 dev eth0 scope link src 172.19.0.3
@@ -0,0 +1,17 @@
New user 'restore-check' has been successfully created!
## login: /user as the throwaway admin
login ok: restore-check is_admin= True
## every repository listed
10 repositories: admin/app-catalog-drill admin/app-catalog-felhom.eu admin/felhom-agent admin/felhom-controller admin/felhom.eu admin/homelab-manifests admin/jarr admin/misc-scripts admin/recipe-importer admin/revfulop-calendar
## main of the four product repositories (restored)
felhom.eu 1707c928a981992597b5cd4af9e10c3bd747b3b7
felhom-controller 13eeb484abdfec2fae97559cb1a374234fe27565
felhom-agent 24ea960229e16c77bb312bcb439c023c267f8f6a
app-catalog-felhom.eu 31e96515910816ab6c1b59cc38c1567e8b347771
## one file byte for byte: felhom.eu CLAUDE.md at the restored main
7866140c4b83a610a3e418a1269d69e2d8384810f8144d34eb2343a6b6190b5f -
## live side (DooPlex clone, fetched from live Gitea 2026-10-09T08:24:41Z)
copy main 1707c928 is ancestor of live main: yes
live commits after the copy:
02a54e26 2026-10-09T10:15:10+02:00 dooplex-offsite: failure mail survives the shared config's unset variable (found by t
live CLAUDE.md at 1707c928: 7866140c4b83a610a3e418a1269d69e2d8384810f8144d34eb2343a6b6190b5f -
@@ -0,0 +1,23 @@
## 2026-10-09T08:24:52Z teardown — bench 9401
gr-gitea
gr-db
gr-net
images-removed
total 20
drwx------ 3 root root 4096 Oct 9 08:24 .
drwxr-xr-x 18 root root 4096 Oct 9 08:22 ..
-rw-r--r-- 1 root root 607 Jul 4 09:05 .bashrc
-rw-r--r-- 1 root root 132 Jul 4 09:05 .profile
drwx------ 2 root root 4096 Jul 14 19:32 .ssh
1
## teardown — DooPlex
total 16
drwx------ 4 root root 4096 Oct 9 10:24 .
drwxr-xr-x 59 root root 4096 Oct 9 10:10 ..
drwx------ 3 root root 4096 Oct 9 10:11 .cache
drwx------ 3 root root 4096 Oct 9 10:10 .kube
9ebdb233c72e9ef9daa61bd9a28e6c76cc02625eb1787dc3550b50cd975481a5 created=2026-10-09T08:24:05Z labels=map[com.docker.volume.anonymous:]
## the anonymous volume made by gr-db at 08:24:05Z (holds the restored database) — removed by name
9ebdb233c72e9ef9daa61bd9a28e6c76cc02625eb1787dc3550b50cd975481a5
volumes-left=0
containers-left=0
+86
View File
@@ -0,0 +1,86 @@
# Runbook — bring Gitea back from the off-site copy on ep0 (R-232)
> **TESTED 2026-10-09** into a throwaway (the bench, LXC 9401 on demo-hp): 10 of 10 repositories listed, the four
> product repositories' `main` equal to live Gitea (one was one commit behind: that commit was pushed three minutes
> after the copy, and the copy's commit is its parent), one file byte for byte, a throwaway admin logged in. Restore
> from ep0: 544 MB in 22 s. Evidence: `audits/dooplex-survival-2026-10-09/partC/`. The copy itself:
> `audits/dooplex-survival-2026-10-09/PLAN.md` and `scripts/dooplex-offsite/`.
## What the copy holds
One encrypted archive `dooplex.pxar` per night in ep0's PBS, namespace `operator`, group `host/dooplex-gitea`
(14 daily + 8 weekly kept). Inside:
| Path | What |
|---|---|
| `db/gitea.dump` | `pg_dump -Fc` of the `gitea` database (PostgreSQL 17.2), taken BEFORE the files |
| `db/globals.sql`, `db/DUMP-FOLDER` | all roles of the CNPG cluster (password hashes — not needed for this restore); which dump |
| `gitea/git/repositories/<owner>/<repo>.git` | the bare repositories |
| `gitea/git/lfs`, `gitea/gitea/{attachments,avatars,repo-avatars,jwt}` | the rest of Gitea's data |
| `gitea/gitea/conf/app.ini` | the config, **with Gitea's secrets** (`SECRET_KEY`, `INTERNAL_TOKEN`, JWT, the DB password) |
| `secrets/*.gpg` | DooPlex's nightly k8s Secrets/ConfigMaps export, GPG-encrypted with DooPlex's restic passphrase |
| `MANIFEST.sha256`, `REPOS` | a checksum of every file; the repository count |
**Not in it:** the container registry (`/data/gitea/packages`, 27.7 GB). The images rebuild from the code.
## What you need
- **The key**: the `data` field of the paper key from the password manager („DooPlex off-site (Gitea) key"). Write
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
to `enc.key` (root, `umask 077`). On DooPlex it is `/etc/felhom-dooplex-offsite/enc.key`.
- **A read-only token** for `dooplex-hub@pbs!restore` (DooPlex: `/etc/felhom-hub-backup/token-restore`), or ep0 root to
mint one (`RUNBOOK-hub-db-offsite-backup.md` Step 2).
- **A route to ep0's PBS** (`127.0.0.1:18007` through DooPlex's tunnel, or ep0's 8007 over the WireGuard).
- A machine with Docker. For the secrets files: DooPlex's restic passphrase (operator, offline).
## Steps
1. **Restore the newest copy** (any machine with `proxmox-backup-client`):
```bash
export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
R='dooplex-hub@pbs!restore@<ep0 PBS>:felhom-offsite'
proxmox-backup-client snapshot list host/dooplex-gitea --ns operator --repository "$R" # pick the newest
umask 077; proxmox-backup-client restore host/dooplex-gitea/<time> dooplex.pxar ./out --ns operator --keyfile enc.key --repository "$R"
(cd out && sha256sum -c MANIFEST.sha256 >/dev/null && echo manifest OK)
```
2. **The database.** `pg_restore` must be 17 (the dump is from 17.2):
```bash
docker network create --internal gr-net # a test: no route out. A real rebuild: a normal network
docker run -d --name gr-db --network gr-net -e POSTGRES_PASSWORD=<pw> postgres:17.2
docker exec -i gr-db psql -U postgres -c "CREATE ROLE gitea LOGIN PASSWORD '<gitea pw>'" -c "CREATE DATABASE gitea OWNER gitea"
docker cp out/db/gitea.dump gr-db:/tmp/ && docker exec gr-db pg_restore -U postgres -d gitea --no-owner --role=gitea --exit-on-error /tmp/gitea.dump
```
On a rebuilt DooPlex: restore into the CNPG cluster's `gitea` database instead (same `pg_restore` line).
3. **The config.** In `out/gitea/gitea/conf/app.ini`, `[database]`: `HOST` → the new database, `PASSWD` → `<gitea pw>`.
For a test also `[mailer] ENABLED = false`. Keep every other key — `SECRET_KEY` and `INTERNAL_TOKEN` must be the old
ones or Gitea cannot read its own stored secrets (2FA, tokens).
4. **Start Gitea** on the data, owned by uid 1000 (the image's `git` user):
```bash
chown -R 1000:1000 out/gitea
docker run -d --name gr-gitea --network gr-net -v "$PWD/out/gitea:/data" gitea/gitea:<the version live ran>
docker exec gr-gitea wget -q -O - http://127.0.0.1:3000/api/healthz # "status": "pass"
```
5. **Check it** (a throwaway admin for a test; the real admin's password works on a real rebuild):
```bash
docker exec -u git gr-gitea gitea admin user create --admin --username restore-check --password <pw> \
--email restore-check@example.invalid --must-change-password=false
# /api/v1/repos/search?limit=50&private=true → the count equals out/REPOS
# /api/v1/repos/admin/<repo>/branches/main → equals the last known main (git ls-remote of any clone)
# /api/v1/repos/admin/felhom.eu/raw/CLAUDE.md?ref=<main> | sha256sum → equals `git show <main>:CLAUDE.md | sha256sum`
```
6. **A test ends with teardown** — the copy holds Gitea's secrets:
```bash
docker rm -f gr-gitea gr-db; docker network rm gr-net; docker rmi postgres:17.2 gitea/gitea:<ver>
docker volume ls # ⚠ postgres leaves an ANONYMOUS volume holding the restored database — remove it BY NAME
find out -type f \( -name app.ini -o -name gitea.dump -o -name globals.sql -o -name '*.gpg' \) -exec shred -u {} +; rm -rf out enc.key
```
Never `docker volume prune`: on a shared machine it deletes other volumes too.
## Gotchas found on 2026-10-09
- **The anonymous Postgres volume** survives `docker rm -f` (no `-v`). It held the restored database; found by counting
volumes after the teardown, removed by name.
- `pg_restore --no-owner --role=gitea`: the dump's objects belong to `gitea` in the source too, but `--no-owner` avoids
needing every role from `globals.sql`.
- The weekly restore test on DooPlex (`felhom-dooplex-offsite-restore-test`, Sun 05:30) checks the manifest, `git fsck`
on every repository and `pg_restore --list` — it does not start Gitea. This runbook is the full test.