R-232/R-173: hub restored into a throwaway k3s (runbook §3 steps 4–5 proven, corrected); R-173, R-861, R-518 closed; R-921, R-922 filed; STATUS, capability map, report
gates / gates (push) Successful in 5m23s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-09 10:54:59 +02:00
parent 59f1ad2b86
commit 81c835827e
12 changed files with 448 additions and 13 deletions
+92
View File
@@ -0,0 +1,92 @@
# REPORT — can the business survive losing DooPlex? (2026-10-09)
| Part | What | Result |
|---|---|---|
| A | Measure; plan | DONE — `audits/dooplex-survival-2026-10-09/PLAN.md`; operator said **yes** in chat |
| B | Nightly encrypted copy of Gitea + secrets to ep0 | DONE, LIVE — first push 57 s; read back; alarms in force; failure mail proven |
| C | Gitea restored into a throwaway | DONE — 10/10 repos, `main` = live, a file byte for byte, a login; deleted |
| D | Hub restored into a throwaway k3s (R-173 steps 4–5) | DONE — customers 4/4, hosts 4/4, console passwords 4/4; deleted; R-173 closed |
| E | R-861 (a), R-518, the mail-address row | R-861 closed, R-518 closed (+ R-921), R-922 filed |
**Register: before 135 · after 135 · opened 2 (R-921, R-922) · closed 3 (R-173, R-861, R-518).** A parallel session
added one row meanwhile (the file read 136 before this session's edit).
Architecture read before any claim: `07-backup-architecture.md` (trust model, encryption), `06-offsite-connectivity.md`,
`01-topology-and-trust.md`, `runbooks/target-selection.md`, R-232 + its recon, the R-173 runbook.
## Part A — measured (read only)
- Gitea 1.26.2: repositories **614 MB** (10 repos), database `gitea` 64 MB live = **6.2 MB** as a dump (~150 KB/day),
`app.ini` 2 KB (holds Gitea's secrets). Registry 27.7 GB — **left out on purpose** (images rebuild from code).
- Secrets set: three GPG files per night, **4.4 MB**, flat.
- **Correction to the recon (§7):** the API mirror holds ONE repo (`homelab-manifests`). The product repos had one
backup only: Longhorn, retain=1, on DooPlex.
- ep0 `operator`: 83 GB free of 98 GB; the hub-DB's write-only token already reaches it. **Nothing to create on ep0.**
## Part B — what was installed on DooPlex (operator yes)
- `/usr/local/sbin/felhom-dooplex-offsite` (+ restore test, + `felhom-backup-failmail`), from `scripts/dooplex-offsite/`
at felhom.eu `1707c928`/`02a54e26`. Timers: daily 00:20 and Sun 05:30.
- New key `/etc/felhom-dooplex-offsite/enc.key` (root 0600, `--kdf none`, never printed). Tokens: the hub-DB's, read in place.
- `OnFailure=felhom-backup-failmail@%n.service` on the two new units **and the two hub-DB units**.
- homelab-manifests `691db39`: `DooplexGiteaOffsiteStale` (26 h) + `DooplexGiteaRestoreTestStale` (8 days), `absent()`
included; only the rules ConfigMap synced; `/-/reload` 200; `/api/v1/rules` reads both `ok`.
- **Consistency choice:** Gitea's docs say stop it for a consistent backup. Not chosen — a nightly stop costs CI and the
registry. Instead: the newest complete database dump (taken first), then the files; the restore test runs `git fsck`
on every repo.
- Proof: first push 544 MB in 57 s; restore test with the read-only token: 27 805 files match, 10 repos pass `git fsck`;
the push token asked to forget a snapshot → `permission check failed`; ep0's disk shows the new group.
- **Found and fixed live:** the failure-mail script died under `set -u` on the shared config's unset variable (the first
dry run sent no mail). Red test, fix, second dry run → Resend id, mail in the inbox 08:17:00Z.
- Security review findings, both fixed and tested: links in the pod's archive are refused; root `git fsck` never reads a
repo's own `config` (red-proved with a config git refuses).
- Tests: 20, green with GNU and BusyBox tools; red-proofs for 7 checks (`partB/red-proof.txt`); promtool green + 2 reds.
## Part C — Gitea from ep0 into a throwaway (the bench, LXC 9401 on demo-hp)
Restored on DooPlex with the read-only token (22 s), streamed to the bench **without the secrets files**, Postgres 17.2
and Gitea 1.26.2 on an `--internal` Docker network (no route out: `wget gitea.com` → bad address). 10/10 repos; three
product `main` equal live; `felhom.eu` one commit behind — that commit (`02a54e26`) was pushed 3 min after the copy and
the copy's `1707c928` is its parent; `CLAUDE.md` sha256 equal; throwaway admin logged in. Runbook: `runbooks/gitea-restore.md`.
## Part D — the hub into a throwaway k3s (same bench)
k3s v1.33.6 in a Docker container on an internal network (two fixes needed: the cgroup v2 move, `--flannel-iface eth0`).
Hub image of the live digest, by file. Runbook §3 steps 4–5 as written. Start-up: `console passwords sealed at rest
(0 legacy …)`. Customers 4/4 and hosts 4/4 equal live (`GET /configs`, `/hosts`, live read GET-only). 4/4 reveals →
32-char passwords (length only); second channel `hubdb-check`: 4/4 with the saved key, 0/4 random. **Found:** the
restored hub tried to mail two households a pending kernel notice within a minute — only the cut network stopped it;
operator-mail-off does not. Runbook §3 corrected (and: three more Secrets are not optional).
## Teardown (three layers)
- **Machine (bench 9401):** containers 0, volumes 0, networks back to the original four, images removed, `/root` as before,
stopped again (it was stopped). **Found:** Postgres and k3s leave anonymous volumes holding the data; 1 + 8 found by
counting, each checked by creation time, removed by name (no prune).
- **Host (DooPlex):** both scratch dirs shredded and removed; the seal-key and password files shredded.
- **Hub:** provisioned nothing. The live hub was only read (GET).
## Part E
- **R-861 (a) CLOSED:** all three 0.304.0 swaps went through `felhom-priv-apply controller-image` (sudo log + wrapper log
+ agent journal) and the guests run 0.304.0 (Docker). No `tee` that day.
- **R-518 CLOSED:** 2026-10-08 on demo-hp both tiers ran, each with its own stop under ~2 min. A third stop of ~1 min
with no copy happened when the off-site tier answered BUSY → **R-921** (P3, needs a release).
- **R-922 filed** (P2): a household's cleared mail address stays on the hub (the no-clobber guard). Rule: operator.
## Other
- CI: `1707c928`, `02a54e26` first failed with every step red and no runner log (R-887, dropped jobs); re-run → success.
`994826ec` success. I pushed the second commit before reading the first run — the rule says check first.
- `unproven.py --summary`: not walked 35 of 55 (unchanged).
- Instruction files: none edited.
## For the operator
1. **Save the new key's paper copy** (at your own terminal, not through `!` here):
`sudo proxmox-backup-client key paperkey /etc/felhom-dooplex-offsite/enc.key --output-format text` → the `data` field
into the password manager as „DooPlex off-site (Gitea) key". **If you do nothing:** after a DooPlex loss the copy on
ep0 cannot be opened. Pick: do it today.
2. **R-922 — a household clears its mail address:** (A) the hub drops the address when the household clears it (the
controller says so explicitly); (B) keep it, and write the keep time into the privacy notice. **Pick A.** **If you do
nothing:** the address stays on the hub with no end date.
+16 -2
View File
@@ -2,8 +2,22 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
**Updated 2026-10-09 (morning): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller
0.304.0. The open-items list is at 135. Report: `REPORT-day4-2026-10-09.md`.**
**Updated 2026-10-09 (midday): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller
0.304.0. The open-items list is at 135. Reports: `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.**
## Midday (2026-10-09): the code and the hub can now survive losing DooPlex
- **Every night at 00:20, all the code (Gitea) and DooPlex's secrets go to ep0, encrypted.** You said yes in chat.
The key that writes the copy cannot delete or read old copies. Nothing changed on ep0.
- **Gitea was brought back from that copy on a throwaway machine** — the first real restore. All 10 repositories are
there, the newest commits match, a file matched byte for byte, a login worked.
- **The hub was brought back from its copy into a throwaway k3s.** Same 4 customers, same 4 boxes, all 4 console
passwords open. Found: a restored hub mails households at once. The test had no network, so nothing went out.
The runbook now says so.
- **Every backup job on DooPlex now mails you when it fails.** A test mail reached the inbox.
- **You need to do one thing:** save the new key's paper copy in your password manager (the report has the command).
If you do nothing, the copy on ep0 cannot be opened after DooPlex is lost.
- **Not copied on purpose:** the container registry (27.7 GB). The images rebuild from the code.
## Morning (2026-10-09): released, delivered, and four of your answers proven live
@@ -235,7 +235,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) |
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs ; **2026-10-09: the whole recovery (§3 steps 1–5) PROVEN on a throwaway k3s** — the live hub image started on the restored copy, customers 4/4 and hosts 4/4 equal live, 4/4 console passwords revealed | `audits/hub-db-offsite-2026-10-05/`; `audits/dooplex-survival-2026-10-09/partD-*.txt`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | A restored hub mails households pending notices at once — a test restore has no network (runbook §3) |
| **Gitea (all code) and DooPlex's secrets survive the loss of DooPlex: a nightly encrypted copy on ep0; restore-tested weekly; an alarm when either stops; a failure mail** | `scripts/dooplex-offsite/` (R-232), homelab-manifests rules | **PROVEN-LIVE (2026-10-09)** — first push 57 s, read back with the read-only token (27 805 files, 10 repos pass `git fsck`); Gitea restored into a throwaway and started: 10/10 repos, product `main` = live, a file byte for byte, a login; alarm by `promtool` + red-proofs; failure mail reached the inbox | `audits/dooplex-survival-2026-10-09/`; `runbooks/gitea-restore.md` | The container registry is NOT copied (images rebuild from the code); the on-box backup tree is still one writable path (R-232 c) |
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
@@ -0,0 +1,27 @@
## 2026-10-09T08:42:44Z fetch the k3s image + airgap set (bench network, before isolation)
docker.io/rancher/k3s:v1.33.6-k3s1
total 140296
drwxr-xr-x 2 root root 4096 Oct 9 08:36 .
drwxr-xr-x 3 root root 4096 Oct 9 08:36 ..
-rw-r--r-- 1 root root 143649404 Oct 9 08:36 k3s-airgap-images-amd64.tar.zst
pd-net internal=true
## after ~15 s
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME
pd-k3s Ready control-plane,master 4s v1.33.6+k3s1 172.30.0.10 <none> K3s v1.33.6+k3s1 7.0.14-20-pve containerd://2.1.5-k3s1.33
No resources found
time="2026-10-09T08:42:48Z" level=error msg="Sending HTTP/1.1 503 response to 127.0.0.1:48848: runtime core not ready"
## 2026-10-09T08:44:46Z fetch the k3s image + airgap set (bench network, before isolation)
docker.io/rancher/k3s:v1.33.6-k3s1
total 140296
drwxr-xr-x 2 root root 4096 Oct 9 08:36 .
drwxr-xr-x 3 root root 4096 Oct 9 08:44 ..
-rw-r--r-- 1 root root 143649404 Oct 9 08:36 k3s-airgap-images-amd64.tar.zst
pd-net internal=true
pd-k3s Up 55 seconds
## after ~15 s
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME
pd-k3s Ready control-plane,master 49s v1.33.6+k3s1 172.30.0.10 <none> K3s v1.33.6+k3s1 7.0.14-20-pve containerd://2.1.5-k3s1.33
NAMESPACE NAME READY STATUS RESTARTS AGE
kube-system coredns-6d668d687-bvh4d 1/1 Running 0 43s
kube-system local-path-provisioner-869c44bfbd-rf27x 1/1 Running 0 43s
time="2026-10-09T08:44:50Z" level=error msg="Sending HTTP/1.1 503 response to 127.0.0.1:57484: runtime core not ready"
@@ -0,0 +1,19 @@
## 2026-10-09T08:43:12Z hub DB restore on DooPlex, read-only token (as felhom-hub-db-restore-test)
newest: host/dooplex-hub/2026-10-09T00:31:46Z
restore complete (380.047 MiB processed in 4.5s, average 84.169 MiB/s)
-rw------- 1 root root 398508032 Oct 9 02:31 /var/lib/felhom-hub-backup/partd/out/hub.db
integrity: ok
Error: in prepare, no such table: customers
hosts: 4 customers:
Error: in prepare, no such table: customers
Error: in prepare, no such column: id
select id from hosts order by id
^--- error here
sealed console passwords: 4
seal key file bytes: 64 (from Secret/offsite-secret-key, the value the operator saved 2026-10-05; not printed)
## (corrected queries: no customers table; customers live in customer_configs)
customer_configs: 4
customer ids: Tester-2 demo-felhom demo-hp tester-1
hosts: Tester-2-be8404→Tester-2 demo-felhom-8363b5→demo-felhom demo-hp-bb76ea→demo-hp tester-1-d70be4→tester-1
@@ -0,0 +1,5 @@
## 2026-10-09T08:44:01Z transfer to bench 9401
-rw------- 1 root root 26974208 Oct 9 08:44 /root/pd/hub-image.tar
95b67d69afdeddaf
DooPlex side: 95b67d69afdeddaf
64
@@ -0,0 +1,49 @@
## 2026-10-09T08:45:51Z system pods
coredns-6d668d687-bvh4d Running
local-path-provisioner-869c44bfbd-rf27x Running
## import the hub image (by file; k3s cannot pull)
Importing elapsed: 1.7 s total: 0.0 B (0.0 B/s)
helper image: docker.io/rancher/mirrored-library-busybox:1.36.1
## secrets: the REAL seal key (step 4: same value); throwaway dummies for mail, report API, registry
secret/offsite-secret-key created
secret/resend-api created
secret/report-api created
secret/gitea-creds created
persistentvolumeclaim/hub-data created
configmap/hub-config created
deployment.apps/hub created
service/hub created
## first start on an EMPTY volume (after ~20 s)
hub-7f775576fd-fmddj 1/1 Running 0 16s
## step 5: scale to 0, copy the restored hub.db in with a helper pod, delete -wal/-shm, scale to 1
deployment.apps/hub scaled
pod/hub-7f775576fd-fmddj condition met
pod/hubdb-copy created
pod/hubdb-copy condition met
total 412
drwxrwxrwx 4 root root 4096 Oct 9 08:46 .
drwxr-xr-x 1 root root 4096 Oct 9 08:46 ..
drwxr-xr-x 2 root root 20480 Oct 9 08:46 assets
-rw-r--r-- 1 root root 389120 Oct 9 08:46 hub.db
drwx------ 2 root root 4096 Oct 9 08:46 snapshots
total 389200
drwxrwxrwx 4 root root 4096 Oct 9 08:46 .
drwxr-xr-x 1 root root 4096 Oct 9 08:46 ..
drwxr-xr-x 2 root root 20480 Oct 9 08:46 assets
-rw------- 1 root root 398508032 Oct 9 08:46 hub.db
drwx------ 2 root root 4096 Oct 9 08:46 snapshots
95b67d69afdeddaf
bench copy: 95b67d69afdeddaf
pod "hubdb-copy" deleted
deployment.apps/hub scaled
## started on the restored DB (after ~20 s)
hub-7f775576fd-cp9nj 1/1 Running 0 16s
## start-up log (sealed / version / errors)
2026/10/09 10:46:52 [INFO] off-site secrets sealed at rest (0 legacy plaintext row(s) sealed now)
2026/10/09 10:46:52 [INFO] console passwords sealed at rest (0 legacy plaintext row(s) sealed now)
2026/10/09 10:46:52 [INFO] box secrets sealed at rest (0 legacy plaintext value(s) sealed now)
2026/10/09 10:46:52 [INFO] Default controller-version floor: 0.120.0
2026/10/09 10:46:52 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000
2026/10/09 10:46:52 [INFO] Registry version checker started (every 6h)
2026/10/09 10:46:52 [INFO] Listening on :8080
2026/10/09 10:46:56 [WARN] Registry version check failed: HTTP request failed: Get "https://gitea.dooplex.hu/v2/admin/felhom-controller/tags/list": dial tcp: lookup gitea.dooplex.hu on 10.43.0.10:53: server misbehaving
@@ -0,0 +1,26 @@
## LIVE hub (GET only; /configs + /hosts) 2026-10-09T08:48:00Z
customers: 4 Tester-2 demo-felhom demo-hp tester-1
hosts: 4 Tester-2-be8404 demo-felhom-8363b5 demo-hp-bb76ea tester-1-d70be4
## THROWAWAY hub on the bench k3s 2026-10-09T08:48:12Z (port-forward inside pd-k3s; same queries + a reveal per host)
customers: 4 Tester-2 demo-felhom demo-hp tester-1
hosts: 4 Tester-2-be8404 demo-felhom-8363b5 demo-hp-bb76ea tester-1-d70be4
reveal Tester-2-be8404: http 200, password 32 chars, sealed-form=False (value not printed)
reveal demo-felhom-8363b5: http 200, password 32 chars, sealed-form=False (value not printed)
reveal demo-hp-bb76ea: http 200, password 32 chars, sealed-form=False (value not printed)
reveal tester-1-d70be4: http 200, password 32 chars, sealed-form=False (value not printed)
## the throwaway hub log for the reveals
2026/10/09 10:46:56 [WARN] Registry version check failed: HTTP request failed: Get "https://gitea.dooplex.hu/v2/admin/felhom-controller/tags/list": dial tcp: lookup gitea.dooplex.hu on 10.43.0.10:53:
2026/10/09 10:47:54 [ERROR] kernel notice (7.0.14-22-pve) to customer demo-felhom failed: sending request: Post "https://api.resend.com/emails": dial tcp: lookup api.resend.com on 10.43.0.10:53: serve
2026/10/09 10:47:58 [ERROR] kernel notice (7.0.14-22-pve) to customer demo-hp failed: sending request: Post "https://api.resend.com/emails": dial tcp: lookup api.resend.com on 10.43.0.10:53: server mi
2026/10/09 10:48:17 [INFO] operator revealed break-glass console credential for host Tester-2-be8404 (user=root@pam, secret 32 chars)
2026/10/09 10:48:17 [INFO] operator revealed break-glass console credential for host demo-felhom-8363b5 (user=root@pam, secret 32 chars)
2026/10/09 10:48:18 [INFO] operator revealed break-glass console credential for host demo-hp-bb76ea (user=root@pam, secret 32 chars)
2026/10/09 10:48:18 [INFO] operator revealed break-glass console credential for host tester-1-d70be4 (user=root@pam, secret 32 chars)
## hubdb-check on DooPlex (second channel), on a COPY of the restored DB 2026-10-09T08:48:30Z
hosts=4 console_passwords_opened=4 failed=0 absent=0
rc=0
control (a random key):
hosts=4 console_passwords_opened=0 failed=4 absent=0
hubdb-check: FAILED: not every console password opened with this key
rc=1
@@ -0,0 +1,175 @@
## 2026-10-09T08:48:46Z teardown — bench 9401
pd-k3s volumes: c57fb08341e5b6fa44a389ddec08f866774ad7acde602870e2e574d044ab33d4 403bfbca4a80250b0e49a8131844b17e59552746513dac0d01f409f942a6e95d d8c1d37af50fe0473724a093c518ad86ec878c2ed96018c8a2ba908346f97daf 585175f25b48d2f86c1a782ee532cac86f099c94c83fe7e84cfd9da986251698
pd-k3s
c57fb08341e5b6fa44a389ddec08f866774ad7acde602870e2e574d044ab33d4
403bfbca4a80250b0e49a8131844b17e59552746513dac0d01f409f942a6e95d
d8c1d37af50fe0473724a093c518ad86ec878c2ed96018c8a2ba908346f97daf
585175f25b48d2f86c1a782ee532cac86f099c94c83fe7e84cfd9da986251698
pd-net
image-removed
containers=0 volumes=8 networks=bridge host none traefik-public
.bashrc
.profile
.ssh
/dev/loop2 59G 3.8G 53G 7% /
are supported and installed on your system.
are supported and installed on your system.
status: stopped
## teardown — DooPlex
.cache
.kube
stage
-home-kisfenyo-git-homelab-manifests
-home-kisfenyo-git-jarr
-mnt-5-hdd-felhom-eu-build-wt-kept
-mnt-5-hdd-felhom-eu-drill-app-catalog-drill
-mnt-5-hdd-felhom-eu-git
-mnt-5-hdd-felhom-eu-git-app-catalog-felhom-eu
-mnt-5-hdd-felhom-eu-git-felhom-agent
-mnt-5-hdd-felhom-eu-git-felhom-controller
-mnt-5-hdd-felhom-eu-git-felhom-eu
-mnt-5-hdd-felhom-eu-git-wt-catalog
-mnt-5-hdd-felhom-eu-worktrees-ctrl-c-msgref
-mnt-5-hdd-felhom-eu-worktrees-ctrl-d4-sessions
-mnt-5-hdd-felhom-eu-worktrees-ctrl-integ
-mnt-5-hdd-felhom-eu-worktrees-hub-cf-mail-drop
ag.txt
au_tests.py
backups_degraded.old
backups_empty.old
backups_full.old
backups_interrupted_run.old
backups_nobackup_yet.old
backups_tier_due.old
bash-edit-diff
bb
bh.bak
bh.new
bookstackfix.py
cache-break-state-35d4e820-990d-465a-b5ee-c7b72649d29a.json
cache-break-state-55a70710-3ae0-4d71-b946-d9db725be672.json
cache-break-state-9a980473-7466-41a7-9c4c-d5f6c1785531.json
cache-break-state-c6c2aaef-766c-4184-8225-3610d20724ca.json
cache-break-state-e6fbe918-56ae-4fc3-af26-4a8a10574e30.json
cache-break-state-f6c29d80-39ec-4c54-960a-646ec49d5e71.json
cat_gates.txt
cfg.bak
cg.txt
cg2.txt
cg3.txt
cgi.txt
cgi2.txt
ci-1594.log
ci-ctl.log
ci.py
ci1357.log
ci1358.log
ci1361.log
ciwait.sh
claperfix.py
cmd_controller_main.go.bak
configs.go.bak
crg.bak
ctl-build.txt
ctlbuild.txt
ctrlgates.txt
deb.I8Wk
deb2.y8f3
deb3.XEFL
docs45.py
dr.bak
drafts
dump_test.go
dumpgo
en.old
felhom-opsign
g.txt
g1.txt
g2.txt
g3.txt
gates-b.txt
gates.txt
gg.bak
gr.txt
h.bak
hb.new
hc.bak
hr.bak
hu.old
hubbak
hubbuild.txt
hubgate.txt
hubrepogates.txt
internal_backup_backup.go.bak
internal_i18n_locales_en.json.bak
internal_i18n_locales_hu.json.bak
isogate.5cI4
m.go
main.bak
main.go.bak
mk.cur
mk.new
mk.old
mnt-hdd_1.mount
newrules.txt
o.bak
off.go
ok.bak
op.bak
opa.bak
osa.bak
pa.bak
partC-ids.json
partC.md
partd.py
partd2.py
partd3.py
promtest
r.bak
r262.bak
r705a.py
r706.py
rec.bak
reg3.py
rel147.txt
rg.bak
rg.txt
rgi.txt
rgi2.txt
rp
rp.bak
rp2
rp2.bak
rr.txt
server.go.bak
service.go.bak
sg.bak
snap.bak
srv.bak
ss.bak
st.bak
states.py
stg.bak
sudoers.v0145
sudotest
t.bak
t.go
t.txt
test_bb.py
uf.bak
unchecked-table.md
v279docs.py
vik.bak
w.bak
x.go
## 2026-10-09T08:49:18Z leftover volumes from the two failed pd-k3s starts (08:36, 08:42)
9bf479a381ed84a255645772d9424cc60a73d024bfc537485536b4e897831ac1 2026-10-09T08:36:57Z map[com.docker.volume.anonymous:]
88e8070b44ca866db658800c9c013b7d6095880b86033d2200d99c91ea97ce0e 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
7119dcb4fa0e1ada468bd477c3bf7f00bda498d3cf8d984f8ff903d674bf5fb5 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
54334eaba740d0cd68c33610c767f361b5882b2a1711ab6abbd6de22a246f416 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
71831f601ff91a65952659dc60d05c5b7c6712158505d20c9665b3e0561bb5df 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:]
a7d0645cf17a95b33018daeb558756bb01f54ab3a59d8553e09b4b8933649d78 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:]
b922c6e0612a135afcdc7d9eb49a964682e4940b45b4fefe311b2258967e275d 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:]
cfabb1005614015b82c12090ad7520c6ba877a4dcfafba317803dae13c0079ef 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
volumes=0 containers=0
status: stopped
+10
View File
@@ -26,6 +26,16 @@
---
## 2026-10-09 — can the business survive losing DooPlex: Gitea + secrets off-site, Gitea and the hub restored into throwaways
The full text of every row below: `git show 59f1ad2b86:documentation/backlog/OPEN-ITEMS.md`.
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** (P2) | CLOSED 2026-10-09 — the last steps of the hub's recovery runbook (§3 steps 4–5) proven on a throwaway single-node k3s (the bench, an INTERNAL Docker network, no route out): the newest ep0 copy restored with the read-only token, the live hub image (same digest) deployed, scaled to 0, `hub.db` copied into the PVC by a helper pod (-wal/-shm removed), scaled to 1 → `console passwords sealed at rest (0 legacy …)`; customers 4/4 and hosts 4/4 equal live; 4/4 console passwords revealed (length only) and, second channel, `hubdb-check` 4/4 with the saved key, 0/4 with a random one. Found: a restored hub at once mails customers their pending notices (operator mail off does NOT stop them) — runbook §3 corrected. Deleted: the k3s container, its volumes, every copy (shredded). | `audits/dooplex-survival-2026-10-09/partD-*.txt`, `runbooks/RUNBOOK-hub-db-offsite-backup.md` §3 |
| **R-861** | **The agent's sudoers lets the agent user reach root without the operator key.** (P2) | CLOSED 2026-10-09 — (a) proven: all three controller swaps of 0.304.0 (demo-hp, demo-felhom, Tester 1, ~05:05Z) ran `felhom-priv-apply controller-image 9201` (sudo log + the wrapper's `WROTE … 0.304.0` + the agent's "new controller healthy"), and the guests run 0.304.0 (Docker); the `tee` route 0 times that day. (b) B2 delivered + B3 accepted, (c) C2 accepted — `09` §3 decision 165. | `audits/dooplex-survival-2026-10-09/partE/R-861.txt` |
| **R-518** | **„Mentés most" stopped every app for ~8 min.** (P2) | CLOSED 2026-10-09 — the two-tier night under one-stop-per-tier read back on demo-hp (2026-10-08): local tier 20:20–20:25Z, off-site tier 20:34–20:37Z, each with its own stop under ~2 min (agent journal + Proxmox task log; controller metrics.db; ep0 listing `2026-10-08T20:34:27Z`, 7-day cadence held). A third, useless stop when the off-site tier answered BUSY → new row **R-921**. | `audits/dooplex-survival-2026-10-09/partE/R-518.txt` |
## 2026-10-09 — the release day: D1–D4 proven live, the ep0-copy job installed
The full text of every row below: `git show b9073e8fb6:documentation/backlog/OPEN-ITEMS.md`.
File diff suppressed because one or more lines are too long
@@ -153,11 +153,19 @@ Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
start it again.
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05 (steps 1–3) and 2026-10-09 (steps 4–5)
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. **Steps 4–5 were run on
2026-10-09 into a throwaway single-node k3s** (the bench, LXC 9401, k3s v1.33.6 in a Docker container on an INTERNAL
network — `audits/dooplex-survival-2026-10-09/partD-*.txt`): the hub image of the live version started on the restored
copy, the customer list and host list equal live (4/4, 4/4), all 4 console passwords revealed (R-173 closed).
> **⚠ Corrected 2026-10-09 — a restored hub mails households AT ONCE.** Within a minute of starting on the copy it
> tried to send two households a pending „kernel notice" (blocked only because the test had no network). Turning
> operator mail off (`operator_enabled: false`) does NOT stop household mail. So: **a test restore has no network at
> all** (no route to Resend, to ep0, to the boxes). **A real recovery** should expect the notices the copy still holds
> to go out once — check the hub's pending notices before giving it a network if that matters.
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
the read-only token (or ep0 root to mint a new one: Step 2).
@@ -183,9 +191,19 @@ the read-only token (or ep0 root to mint a new one: Step 2).
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
**Corrected 2026-10-09:** the Deployment also needs, NOT optional, `Secret/resend-api` (`RESEND_API_KEY`),
`Secret/report-api` (`REPORT_API_KEY`) and `Secret/gitea-creds` (`username`, `password`); without them the pod does
not start. On a real rebuild their values come from DooPlex's nightly secrets export, which since 2026-10-09 is also
off-site (`runbooks/gitea-restore.md`, `secrets/*.gpg`, opened with DooPlex's restic passphrase). A test uses dummies.
The hub's image is pulled from Gitea's registry — on a rebuild with no registry, `docker save` it from any machine
that has it, or build it from the restored code.
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
As run 2026-10-09: `kubectl scale deploy/hub --replicas=0`; a `busybox` pod mounting `hub-data` at `/data`;
`kubectl cp hub.db hubdb-copy:/data/hub.db.new`, then in the pod `rm -f /data/hub.db-wal /data/hub.db-shm && mv
/data/hub.db.new /data/hub.db`; delete the pod; scale to 1 (20 s to Ready). The customer list is `GET /configs`
(rows link to `/customers/<id>`), the hosts `GET /hosts`.
6. Shred `k`, `enc.key` copies and `out/` when done.
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account