R-232/R-173: hub restored into a throwaway k3s (runbook §3 steps 4–5 proven, corrected); R-173, R-861, R-518 closed; R-921, R-922 filed; STATUS, capability map, report
gates / gates (push) Successful in 5m23s
gates / gates (push) Successful in 5m23s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,92 @@
|
||||
# REPORT — can the business survive losing DooPlex? (2026-10-09)
|
||||
|
||||
| Part | What | Result |
|
||||
|---|---|---|
|
||||
| A | Measure; plan | DONE — `audits/dooplex-survival-2026-10-09/PLAN.md`; operator said **yes** in chat |
|
||||
| B | Nightly encrypted copy of Gitea + secrets to ep0 | DONE, LIVE — first push 57 s; read back; alarms in force; failure mail proven |
|
||||
| C | Gitea restored into a throwaway | DONE — 10/10 repos, `main` = live, a file byte for byte, a login; deleted |
|
||||
| D | Hub restored into a throwaway k3s (R-173 steps 4–5) | DONE — customers 4/4, hosts 4/4, console passwords 4/4; deleted; R-173 closed |
|
||||
| E | R-861 (a), R-518, the mail-address row | R-861 closed, R-518 closed (+ R-921), R-922 filed |
|
||||
|
||||
**Register: before 135 · after 135 · opened 2 (R-921, R-922) · closed 3 (R-173, R-861, R-518).** A parallel session
|
||||
added one row meanwhile (the file read 136 before this session's edit).
|
||||
|
||||
Architecture read before any claim: `07-backup-architecture.md` (trust model, encryption), `06-offsite-connectivity.md`,
|
||||
`01-topology-and-trust.md`, `runbooks/target-selection.md`, R-232 + its recon, the R-173 runbook.
|
||||
|
||||
## Part A — measured (read only)
|
||||
|
||||
- Gitea 1.26.2: repositories **614 MB** (10 repos), database `gitea` 64 MB live = **6.2 MB** as a dump (~150 KB/day),
|
||||
`app.ini` 2 KB (holds Gitea's secrets). Registry 27.7 GB — **left out on purpose** (images rebuild from code).
|
||||
- Secrets set: three GPG files per night, **4.4 MB**, flat.
|
||||
- **Correction to the recon (§7):** the API mirror holds ONE repo (`homelab-manifests`). The product repos had one
|
||||
backup only: Longhorn, retain=1, on DooPlex.
|
||||
- ep0 `operator`: 83 GB free of 98 GB; the hub-DB's write-only token already reaches it. **Nothing to create on ep0.**
|
||||
|
||||
## Part B — what was installed on DooPlex (operator yes)
|
||||
|
||||
- `/usr/local/sbin/felhom-dooplex-offsite` (+ restore test, + `felhom-backup-failmail`), from `scripts/dooplex-offsite/`
|
||||
at felhom.eu `1707c928`/`02a54e26`. Timers: daily 00:20 and Sun 05:30.
|
||||
- New key `/etc/felhom-dooplex-offsite/enc.key` (root 0600, `--kdf none`, never printed). Tokens: the hub-DB's, read in place.
|
||||
- `OnFailure=felhom-backup-failmail@%n.service` on the two new units **and the two hub-DB units**.
|
||||
- homelab-manifests `691db39`: `DooplexGiteaOffsiteStale` (26 h) + `DooplexGiteaRestoreTestStale` (8 days), `absent()`
|
||||
included; only the rules ConfigMap synced; `/-/reload` 200; `/api/v1/rules` reads both `ok`.
|
||||
- **Consistency choice:** Gitea's docs say stop it for a consistent backup. Not chosen — a nightly stop costs CI and the
|
||||
registry. Instead: the newest complete database dump (taken first), then the files; the restore test runs `git fsck`
|
||||
on every repo.
|
||||
- Proof: first push 544 MB in 57 s; restore test with the read-only token: 27 805 files match, 10 repos pass `git fsck`;
|
||||
the push token asked to forget a snapshot → `permission check failed`; ep0's disk shows the new group.
|
||||
- **Found and fixed live:** the failure-mail script died under `set -u` on the shared config's unset variable (the first
|
||||
dry run sent no mail). Red test, fix, second dry run → Resend id, mail in the inbox 08:17:00Z.
|
||||
- Security review findings, both fixed and tested: links in the pod's archive are refused; root `git fsck` never reads a
|
||||
repo's own `config` (red-proved with a config git refuses).
|
||||
- Tests: 20, green with GNU and BusyBox tools; red-proofs for 7 checks (`partB/red-proof.txt`); promtool green + 2 reds.
|
||||
|
||||
## Part C — Gitea from ep0 into a throwaway (the bench, LXC 9401 on demo-hp)
|
||||
|
||||
Restored on DooPlex with the read-only token (22 s), streamed to the bench **without the secrets files**, Postgres 17.2
|
||||
and Gitea 1.26.2 on an `--internal` Docker network (no route out: `wget gitea.com` → bad address). 10/10 repos; three
|
||||
product `main` equal live; `felhom.eu` one commit behind — that commit (`02a54e26`) was pushed 3 min after the copy and
|
||||
the copy's `1707c928` is its parent; `CLAUDE.md` sha256 equal; throwaway admin logged in. Runbook: `runbooks/gitea-restore.md`.
|
||||
|
||||
## Part D — the hub into a throwaway k3s (same bench)
|
||||
|
||||
k3s v1.33.6 in a Docker container on an internal network (two fixes needed: the cgroup v2 move, `--flannel-iface eth0`).
|
||||
Hub image of the live digest, by file. Runbook §3 steps 4–5 as written. Start-up: `console passwords sealed at rest
|
||||
(0 legacy …)`. Customers 4/4 and hosts 4/4 equal live (`GET /configs`, `/hosts`, live read GET-only). 4/4 reveals →
|
||||
32-char passwords (length only); second channel `hubdb-check`: 4/4 with the saved key, 0/4 random. **Found:** the
|
||||
restored hub tried to mail two households a pending kernel notice within a minute — only the cut network stopped it;
|
||||
operator-mail-off does not. Runbook §3 corrected (and: three more Secrets are not optional).
|
||||
|
||||
## Teardown (three layers)
|
||||
|
||||
- **Machine (bench 9401):** containers 0, volumes 0, networks back to the original four, images removed, `/root` as before,
|
||||
stopped again (it was stopped). **Found:** Postgres and k3s leave anonymous volumes holding the data; 1 + 8 found by
|
||||
counting, each checked by creation time, removed by name (no prune).
|
||||
- **Host (DooPlex):** both scratch dirs shredded and removed; the seal-key and password files shredded.
|
||||
- **Hub:** provisioned nothing. The live hub was only read (GET).
|
||||
|
||||
## Part E
|
||||
|
||||
- **R-861 (a) CLOSED:** all three 0.304.0 swaps went through `felhom-priv-apply controller-image` (sudo log + wrapper log
|
||||
+ agent journal) and the guests run 0.304.0 (Docker). No `tee` that day.
|
||||
- **R-518 CLOSED:** 2026-10-08 on demo-hp both tiers ran, each with its own stop under ~2 min. A third stop of ~1 min
|
||||
with no copy happened when the off-site tier answered BUSY → **R-921** (P3, needs a release).
|
||||
- **R-922 filed** (P2): a household's cleared mail address stays on the hub (the no-clobber guard). Rule: operator.
|
||||
|
||||
## Other
|
||||
|
||||
- CI: `1707c928`, `02a54e26` first failed with every step red and no runner log (R-887, dropped jobs); re-run → success.
|
||||
`994826ec` success. I pushed the second commit before reading the first run — the rule says check first.
|
||||
- `unproven.py --summary`: not walked 35 of 55 (unchanged).
|
||||
- Instruction files: none edited.
|
||||
|
||||
## For the operator
|
||||
|
||||
1. **Save the new key's paper copy** (at your own terminal, not through `!` here):
|
||||
`sudo proxmox-backup-client key paperkey /etc/felhom-dooplex-offsite/enc.key --output-format text` → the `data` field
|
||||
into the password manager as „DooPlex off-site (Gitea) key". **If you do nothing:** after a DooPlex loss the copy on
|
||||
ep0 cannot be opened. Pick: do it today.
|
||||
2. **R-922 — a household clears its mail address:** (A) the hub drops the address when the household clears it (the
|
||||
controller says so explicitly); (B) keep it, and write the keep time into the privacy notice. **Pick A.** **If you do
|
||||
nothing:** the address stays on the hub with no end date.
|
||||
@@ -2,8 +2,22 @@
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
|
||||
|
||||
**Updated 2026-10-09 (morning): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller
|
||||
0.304.0. The open-items list is at 135. Report: `REPORT-day4-2026-10-09.md`.**
|
||||
**Updated 2026-10-09 (midday): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller
|
||||
0.304.0. The open-items list is at 135. Reports: `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.**
|
||||
|
||||
## Midday (2026-10-09): the code and the hub can now survive losing DooPlex
|
||||
|
||||
- **Every night at 00:20, all the code (Gitea) and DooPlex's secrets go to ep0, encrypted.** You said yes in chat.
|
||||
The key that writes the copy cannot delete or read old copies. Nothing changed on ep0.
|
||||
- **Gitea was brought back from that copy on a throwaway machine** — the first real restore. All 10 repositories are
|
||||
there, the newest commits match, a file matched byte for byte, a login worked.
|
||||
- **The hub was brought back from its copy into a throwaway k3s.** Same 4 customers, same 4 boxes, all 4 console
|
||||
passwords open. Found: a restored hub mails households at once. The test had no network, so nothing went out.
|
||||
The runbook now says so.
|
||||
- **Every backup job on DooPlex now mails you when it fails.** A test mail reached the inbox.
|
||||
- **You need to do one thing:** save the new key's paper copy in your password manager (the report has the command).
|
||||
If you do nothing, the copy on ep0 cannot be opened after DooPlex is lost.
|
||||
- **Not copied on purpose:** the container registry (27.7 GB). The images rebuild from the code.
|
||||
|
||||
## Morning (2026-10-09): released, delivered, and four of your answers proven live
|
||||
|
||||
|
||||
@@ -235,7 +235,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
|
||||
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) |
|
||||
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs ; **2026-10-09: the whole recovery (§3 steps 1–5) PROVEN on a throwaway k3s** — the live hub image started on the restored copy, customers 4/4 and hosts 4/4 equal live, 4/4 console passwords revealed | `audits/hub-db-offsite-2026-10-05/`; `audits/dooplex-survival-2026-10-09/partD-*.txt`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | A restored hub mails households pending notices at once — a test restore has no network (runbook §3) |
|
||||
| **Gitea (all code) and DooPlex's secrets survive the loss of DooPlex: a nightly encrypted copy on ep0; restore-tested weekly; an alarm when either stops; a failure mail** | `scripts/dooplex-offsite/` (R-232), homelab-manifests rules | **PROVEN-LIVE (2026-10-09)** — first push 57 s, read back with the read-only token (27 805 files, 10 repos pass `git fsck`); Gitea restored into a throwaway and started: 10/10 repos, product `main` = live, a file byte for byte, a login; alarm by `promtool` + red-proofs; failure mail reached the inbox | `audits/dooplex-survival-2026-10-09/`; `runbooks/gitea-restore.md` | The container registry is NOT copied (images rebuild from the code); the on-box backup tree is still one writable path (R-232 c) |
|
||||
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
## 2026-10-09T08:42:44Z fetch the k3s image + airgap set (bench network, before isolation)
|
||||
docker.io/rancher/k3s:v1.33.6-k3s1
|
||||
total 140296
|
||||
drwxr-xr-x 2 root root 4096 Oct 9 08:36 .
|
||||
drwxr-xr-x 3 root root 4096 Oct 9 08:36 ..
|
||||
-rw-r--r-- 1 root root 143649404 Oct 9 08:36 k3s-airgap-images-amd64.tar.zst
|
||||
pd-net internal=true
|
||||
## after ~15 s
|
||||
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME
|
||||
pd-k3s Ready control-plane,master 4s v1.33.6+k3s1 172.30.0.10 <none> K3s v1.33.6+k3s1 7.0.14-20-pve containerd://2.1.5-k3s1.33
|
||||
No resources found
|
||||
time="2026-10-09T08:42:48Z" level=error msg="Sending HTTP/1.1 503 response to 127.0.0.1:48848: runtime core not ready"
|
||||
## 2026-10-09T08:44:46Z fetch the k3s image + airgap set (bench network, before isolation)
|
||||
docker.io/rancher/k3s:v1.33.6-k3s1
|
||||
total 140296
|
||||
drwxr-xr-x 2 root root 4096 Oct 9 08:36 .
|
||||
drwxr-xr-x 3 root root 4096 Oct 9 08:44 ..
|
||||
-rw-r--r-- 1 root root 143649404 Oct 9 08:36 k3s-airgap-images-amd64.tar.zst
|
||||
pd-net internal=true
|
||||
pd-k3s Up 55 seconds
|
||||
## after ~15 s
|
||||
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME
|
||||
pd-k3s Ready control-plane,master 49s v1.33.6+k3s1 172.30.0.10 <none> K3s v1.33.6+k3s1 7.0.14-20-pve containerd://2.1.5-k3s1.33
|
||||
NAMESPACE NAME READY STATUS RESTARTS AGE
|
||||
kube-system coredns-6d668d687-bvh4d 1/1 Running 0 43s
|
||||
kube-system local-path-provisioner-869c44bfbd-rf27x 1/1 Running 0 43s
|
||||
time="2026-10-09T08:44:50Z" level=error msg="Sending HTTP/1.1 503 response to 127.0.0.1:57484: runtime core not ready"
|
||||
@@ -0,0 +1,19 @@
|
||||
## 2026-10-09T08:43:12Z hub DB restore on DooPlex, read-only token (as felhom-hub-db-restore-test)
|
||||
newest: host/dooplex-hub/2026-10-09T00:31:46Z
|
||||
restore complete (380.047 MiB processed in 4.5s, average 84.169 MiB/s)
|
||||
-rw------- 1 root root 398508032 Oct 9 02:31 /var/lib/felhom-hub-backup/partd/out/hub.db
|
||||
integrity: ok
|
||||
Error: in prepare, no such table: customers
|
||||
hosts: 4 customers:
|
||||
Error: in prepare, no such table: customers
|
||||
|
||||
Error: in prepare, no such column: id
|
||||
select id from hosts order by id
|
||||
^--- error here
|
||||
|
||||
sealed console passwords: 4
|
||||
seal key file bytes: 64 (from Secret/offsite-secret-key, the value the operator saved 2026-10-05; not printed)
|
||||
## (corrected queries: no customers table; customers live in customer_configs)
|
||||
customer_configs: 4
|
||||
customer ids: Tester-2 demo-felhom demo-hp tester-1
|
||||
hosts: Tester-2-be8404→Tester-2 demo-felhom-8363b5→demo-felhom demo-hp-bb76ea→demo-hp tester-1-d70be4→tester-1
|
||||
@@ -0,0 +1,5 @@
|
||||
## 2026-10-09T08:44:01Z transfer to bench 9401
|
||||
-rw------- 1 root root 26974208 Oct 9 08:44 /root/pd/hub-image.tar
|
||||
95b67d69afdeddaf
|
||||
DooPlex side: 95b67d69afdeddaf
|
||||
64
|
||||
@@ -0,0 +1,49 @@
|
||||
## 2026-10-09T08:45:51Z system pods
|
||||
coredns-6d668d687-bvh4d Running
|
||||
local-path-provisioner-869c44bfbd-rf27x Running
|
||||
## import the hub image (by file; k3s cannot pull)
|
||||
Importing elapsed: 1.7 s total: 0.0 B (0.0 B/s)
|
||||
helper image: docker.io/rancher/mirrored-library-busybox:1.36.1
|
||||
## secrets: the REAL seal key (step 4: same value); throwaway dummies for mail, report API, registry
|
||||
secret/offsite-secret-key created
|
||||
secret/resend-api created
|
||||
secret/report-api created
|
||||
secret/gitea-creds created
|
||||
persistentvolumeclaim/hub-data created
|
||||
configmap/hub-config created
|
||||
deployment.apps/hub created
|
||||
service/hub created
|
||||
## first start on an EMPTY volume (after ~20 s)
|
||||
hub-7f775576fd-fmddj 1/1 Running 0 16s
|
||||
## step 5: scale to 0, copy the restored hub.db in with a helper pod, delete -wal/-shm, scale to 1
|
||||
deployment.apps/hub scaled
|
||||
pod/hub-7f775576fd-fmddj condition met
|
||||
pod/hubdb-copy created
|
||||
pod/hubdb-copy condition met
|
||||
total 412
|
||||
drwxrwxrwx 4 root root 4096 Oct 9 08:46 .
|
||||
drwxr-xr-x 1 root root 4096 Oct 9 08:46 ..
|
||||
drwxr-xr-x 2 root root 20480 Oct 9 08:46 assets
|
||||
-rw-r--r-- 1 root root 389120 Oct 9 08:46 hub.db
|
||||
drwx------ 2 root root 4096 Oct 9 08:46 snapshots
|
||||
total 389200
|
||||
drwxrwxrwx 4 root root 4096 Oct 9 08:46 .
|
||||
drwxr-xr-x 1 root root 4096 Oct 9 08:46 ..
|
||||
drwxr-xr-x 2 root root 20480 Oct 9 08:46 assets
|
||||
-rw------- 1 root root 398508032 Oct 9 08:46 hub.db
|
||||
drwx------ 2 root root 4096 Oct 9 08:46 snapshots
|
||||
95b67d69afdeddaf
|
||||
bench copy: 95b67d69afdeddaf
|
||||
pod "hubdb-copy" deleted
|
||||
deployment.apps/hub scaled
|
||||
## started on the restored DB (after ~20 s)
|
||||
hub-7f775576fd-cp9nj 1/1 Running 0 16s
|
||||
## start-up log (sealed / version / errors)
|
||||
2026/10/09 10:46:52 [INFO] off-site secrets sealed at rest (0 legacy plaintext row(s) sealed now)
|
||||
2026/10/09 10:46:52 [INFO] console passwords sealed at rest (0 legacy plaintext row(s) sealed now)
|
||||
2026/10/09 10:46:52 [INFO] box secrets sealed at rest (0 legacy plaintext value(s) sealed now)
|
||||
2026/10/09 10:46:52 [INFO] Default controller-version floor: 0.120.0
|
||||
2026/10/09 10:46:52 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000
|
||||
2026/10/09 10:46:52 [INFO] Registry version checker started (every 6h)
|
||||
2026/10/09 10:46:52 [INFO] Listening on :8080
|
||||
2026/10/09 10:46:56 [WARN] Registry version check failed: HTTP request failed: Get "https://gitea.dooplex.hu/v2/admin/felhom-controller/tags/list": dial tcp: lookup gitea.dooplex.hu on 10.43.0.10:53: server misbehaving
|
||||
@@ -0,0 +1,26 @@
|
||||
## LIVE hub (GET only; /configs + /hosts) 2026-10-09T08:48:00Z
|
||||
customers: 4 Tester-2 demo-felhom demo-hp tester-1
|
||||
hosts: 4 Tester-2-be8404 demo-felhom-8363b5 demo-hp-bb76ea tester-1-d70be4
|
||||
## THROWAWAY hub on the bench k3s 2026-10-09T08:48:12Z (port-forward inside pd-k3s; same queries + a reveal per host)
|
||||
customers: 4 Tester-2 demo-felhom demo-hp tester-1
|
||||
hosts: 4 Tester-2-be8404 demo-felhom-8363b5 demo-hp-bb76ea tester-1-d70be4
|
||||
reveal Tester-2-be8404: http 200, password 32 chars, sealed-form=False (value not printed)
|
||||
reveal demo-felhom-8363b5: http 200, password 32 chars, sealed-form=False (value not printed)
|
||||
reveal demo-hp-bb76ea: http 200, password 32 chars, sealed-form=False (value not printed)
|
||||
reveal tester-1-d70be4: http 200, password 32 chars, sealed-form=False (value not printed)
|
||||
|
||||
## the throwaway hub log for the reveals
|
||||
2026/10/09 10:46:56 [WARN] Registry version check failed: HTTP request failed: Get "https://gitea.dooplex.hu/v2/admin/felhom-controller/tags/list": dial tcp: lookup gitea.dooplex.hu on 10.43.0.10:53:
|
||||
2026/10/09 10:47:54 [ERROR] kernel notice (7.0.14-22-pve) to customer demo-felhom failed: sending request: Post "https://api.resend.com/emails": dial tcp: lookup api.resend.com on 10.43.0.10:53: serve
|
||||
2026/10/09 10:47:58 [ERROR] kernel notice (7.0.14-22-pve) to customer demo-hp failed: sending request: Post "https://api.resend.com/emails": dial tcp: lookup api.resend.com on 10.43.0.10:53: server mi
|
||||
2026/10/09 10:48:17 [INFO] operator revealed break-glass console credential for host Tester-2-be8404 (user=root@pam, secret 32 chars)
|
||||
2026/10/09 10:48:17 [INFO] operator revealed break-glass console credential for host demo-felhom-8363b5 (user=root@pam, secret 32 chars)
|
||||
2026/10/09 10:48:18 [INFO] operator revealed break-glass console credential for host demo-hp-bb76ea (user=root@pam, secret 32 chars)
|
||||
2026/10/09 10:48:18 [INFO] operator revealed break-glass console credential for host tester-1-d70be4 (user=root@pam, secret 32 chars)
|
||||
## hubdb-check on DooPlex (second channel), on a COPY of the restored DB 2026-10-09T08:48:30Z
|
||||
hosts=4 console_passwords_opened=4 failed=0 absent=0
|
||||
rc=0
|
||||
control (a random key):
|
||||
hosts=4 console_passwords_opened=0 failed=4 absent=0
|
||||
hubdb-check: FAILED: not every console password opened with this key
|
||||
rc=1
|
||||
@@ -0,0 +1,175 @@
|
||||
## 2026-10-09T08:48:46Z teardown — bench 9401
|
||||
pd-k3s volumes: c57fb08341e5b6fa44a389ddec08f866774ad7acde602870e2e574d044ab33d4 403bfbca4a80250b0e49a8131844b17e59552746513dac0d01f409f942a6e95d d8c1d37af50fe0473724a093c518ad86ec878c2ed96018c8a2ba908346f97daf 585175f25b48d2f86c1a782ee532cac86f099c94c83fe7e84cfd9da986251698
|
||||
pd-k3s
|
||||
c57fb08341e5b6fa44a389ddec08f866774ad7acde602870e2e574d044ab33d4
|
||||
403bfbca4a80250b0e49a8131844b17e59552746513dac0d01f409f942a6e95d
|
||||
d8c1d37af50fe0473724a093c518ad86ec878c2ed96018c8a2ba908346f97daf
|
||||
585175f25b48d2f86c1a782ee532cac86f099c94c83fe7e84cfd9da986251698
|
||||
pd-net
|
||||
image-removed
|
||||
containers=0 volumes=8 networks=bridge host none traefik-public
|
||||
.bashrc
|
||||
.profile
|
||||
.ssh
|
||||
/dev/loop2 59G 3.8G 53G 7% /
|
||||
are supported and installed on your system.
|
||||
are supported and installed on your system.
|
||||
status: stopped
|
||||
## teardown — DooPlex
|
||||
.cache
|
||||
.kube
|
||||
stage
|
||||
-home-kisfenyo-git-homelab-manifests
|
||||
-home-kisfenyo-git-jarr
|
||||
-mnt-5-hdd-felhom-eu-build-wt-kept
|
||||
-mnt-5-hdd-felhom-eu-drill-app-catalog-drill
|
||||
-mnt-5-hdd-felhom-eu-git
|
||||
-mnt-5-hdd-felhom-eu-git-app-catalog-felhom-eu
|
||||
-mnt-5-hdd-felhom-eu-git-felhom-agent
|
||||
-mnt-5-hdd-felhom-eu-git-felhom-controller
|
||||
-mnt-5-hdd-felhom-eu-git-felhom-eu
|
||||
-mnt-5-hdd-felhom-eu-git-wt-catalog
|
||||
-mnt-5-hdd-felhom-eu-worktrees-ctrl-c-msgref
|
||||
-mnt-5-hdd-felhom-eu-worktrees-ctrl-d4-sessions
|
||||
-mnt-5-hdd-felhom-eu-worktrees-ctrl-integ
|
||||
-mnt-5-hdd-felhom-eu-worktrees-hub-cf-mail-drop
|
||||
ag.txt
|
||||
au_tests.py
|
||||
backups_degraded.old
|
||||
backups_empty.old
|
||||
backups_full.old
|
||||
backups_interrupted_run.old
|
||||
backups_nobackup_yet.old
|
||||
backups_tier_due.old
|
||||
bash-edit-diff
|
||||
bb
|
||||
bh.bak
|
||||
bh.new
|
||||
bookstackfix.py
|
||||
cache-break-state-35d4e820-990d-465a-b5ee-c7b72649d29a.json
|
||||
cache-break-state-55a70710-3ae0-4d71-b946-d9db725be672.json
|
||||
cache-break-state-9a980473-7466-41a7-9c4c-d5f6c1785531.json
|
||||
cache-break-state-c6c2aaef-766c-4184-8225-3610d20724ca.json
|
||||
cache-break-state-e6fbe918-56ae-4fc3-af26-4a8a10574e30.json
|
||||
cache-break-state-f6c29d80-39ec-4c54-960a-646ec49d5e71.json
|
||||
cat_gates.txt
|
||||
cfg.bak
|
||||
cg.txt
|
||||
cg2.txt
|
||||
cg3.txt
|
||||
cgi.txt
|
||||
cgi2.txt
|
||||
ci-1594.log
|
||||
ci-ctl.log
|
||||
ci.py
|
||||
ci1357.log
|
||||
ci1358.log
|
||||
ci1361.log
|
||||
ciwait.sh
|
||||
claperfix.py
|
||||
cmd_controller_main.go.bak
|
||||
configs.go.bak
|
||||
crg.bak
|
||||
ctl-build.txt
|
||||
ctlbuild.txt
|
||||
ctrlgates.txt
|
||||
deb.I8Wk
|
||||
deb2.y8f3
|
||||
deb3.XEFL
|
||||
docs45.py
|
||||
dr.bak
|
||||
drafts
|
||||
dump_test.go
|
||||
dumpgo
|
||||
en.old
|
||||
felhom-opsign
|
||||
g.txt
|
||||
g1.txt
|
||||
g2.txt
|
||||
g3.txt
|
||||
gates-b.txt
|
||||
gates.txt
|
||||
gg.bak
|
||||
gr.txt
|
||||
h.bak
|
||||
hb.new
|
||||
hc.bak
|
||||
hr.bak
|
||||
hu.old
|
||||
hubbak
|
||||
hubbuild.txt
|
||||
hubgate.txt
|
||||
hubrepogates.txt
|
||||
internal_backup_backup.go.bak
|
||||
internal_i18n_locales_en.json.bak
|
||||
internal_i18n_locales_hu.json.bak
|
||||
isogate.5cI4
|
||||
m.go
|
||||
main.bak
|
||||
main.go.bak
|
||||
mk.cur
|
||||
mk.new
|
||||
mk.old
|
||||
mnt-hdd_1.mount
|
||||
newrules.txt
|
||||
o.bak
|
||||
off.go
|
||||
ok.bak
|
||||
op.bak
|
||||
opa.bak
|
||||
osa.bak
|
||||
pa.bak
|
||||
partC-ids.json
|
||||
partC.md
|
||||
partd.py
|
||||
partd2.py
|
||||
partd3.py
|
||||
promtest
|
||||
r.bak
|
||||
r262.bak
|
||||
r705a.py
|
||||
r706.py
|
||||
rec.bak
|
||||
reg3.py
|
||||
rel147.txt
|
||||
rg.bak
|
||||
rg.txt
|
||||
rgi.txt
|
||||
rgi2.txt
|
||||
rp
|
||||
rp.bak
|
||||
rp2
|
||||
rp2.bak
|
||||
rr.txt
|
||||
server.go.bak
|
||||
service.go.bak
|
||||
sg.bak
|
||||
snap.bak
|
||||
srv.bak
|
||||
ss.bak
|
||||
st.bak
|
||||
states.py
|
||||
stg.bak
|
||||
sudoers.v0145
|
||||
sudotest
|
||||
t.bak
|
||||
t.go
|
||||
t.txt
|
||||
test_bb.py
|
||||
uf.bak
|
||||
unchecked-table.md
|
||||
v279docs.py
|
||||
vik.bak
|
||||
w.bak
|
||||
x.go
|
||||
## 2026-10-09T08:49:18Z leftover volumes from the two failed pd-k3s starts (08:36, 08:42)
|
||||
9bf479a381ed84a255645772d9424cc60a73d024bfc537485536b4e897831ac1 2026-10-09T08:36:57Z map[com.docker.volume.anonymous:]
|
||||
88e8070b44ca866db658800c9c013b7d6095880b86033d2200d99c91ea97ce0e 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
|
||||
7119dcb4fa0e1ada468bd477c3bf7f00bda498d3cf8d984f8ff903d674bf5fb5 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
|
||||
54334eaba740d0cd68c33610c767f361b5882b2a1711ab6abbd6de22a246f416 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
|
||||
71831f601ff91a65952659dc60d05c5b7c6712158505d20c9665b3e0561bb5df 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:]
|
||||
a7d0645cf17a95b33018daeb558756bb01f54ab3a59d8553e09b4b8933649d78 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:]
|
||||
b922c6e0612a135afcdc7d9eb49a964682e4940b45b4fefe311b2258967e275d 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:]
|
||||
cfabb1005614015b82c12090ad7520c6ba877a4dcfafba317803dae13c0079ef 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:]
|
||||
volumes=0 containers=0
|
||||
status: stopped
|
||||
@@ -26,6 +26,16 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-09 — can the business survive losing DooPlex: Gitea + secrets off-site, Gitea and the hub restored into throwaways
|
||||
|
||||
The full text of every row below: `git show 59f1ad2b86:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** (P2) | CLOSED 2026-10-09 — the last steps of the hub's recovery runbook (§3 steps 4–5) proven on a throwaway single-node k3s (the bench, an INTERNAL Docker network, no route out): the newest ep0 copy restored with the read-only token, the live hub image (same digest) deployed, scaled to 0, `hub.db` copied into the PVC by a helper pod (-wal/-shm removed), scaled to 1 → `console passwords sealed at rest (0 legacy …)`; customers 4/4 and hosts 4/4 equal live; 4/4 console passwords revealed (length only) and, second channel, `hubdb-check` 4/4 with the saved key, 0/4 with a random one. Found: a restored hub at once mails customers their pending notices (operator mail off does NOT stop them) — runbook §3 corrected. Deleted: the k3s container, its volumes, every copy (shredded). | `audits/dooplex-survival-2026-10-09/partD-*.txt`, `runbooks/RUNBOOK-hub-db-offsite-backup.md` §3 |
|
||||
| **R-861** | **The agent's sudoers lets the agent user reach root without the operator key.** (P2) | CLOSED 2026-10-09 — (a) proven: all three controller swaps of 0.304.0 (demo-hp, demo-felhom, Tester 1, ~05:05Z) ran `felhom-priv-apply controller-image 9201` (sudo log + the wrapper's `WROTE … 0.304.0` + the agent's "new controller healthy"), and the guests run 0.304.0 (Docker); the `tee` route 0 times that day. (b) B2 delivered + B3 accepted, (c) C2 accepted — `09` §3 decision 165. | `audits/dooplex-survival-2026-10-09/partE/R-861.txt` |
|
||||
| **R-518** | **„Mentés most" stopped every app for ~8 min.** (P2) | CLOSED 2026-10-09 — the two-tier night under one-stop-per-tier read back on demo-hp (2026-10-08): local tier 20:20–20:25Z, off-site tier 20:34–20:37Z, each with its own stop under ~2 min (agent journal + Proxmox task log; controller metrics.db; ep0 listing `2026-10-08T20:34:27Z`, 7-day cadence held). A third, useless stop when the off-site tier answered BUSY → new row **R-921**. | `audits/dooplex-survival-2026-10-09/partE/R-518.txt` |
|
||||
|
||||
## 2026-10-09 — the release day: D1–D4 proven live, the ep0-copy job installed
|
||||
|
||||
The full text of every row below: `git show b9073e8fb6:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -153,11 +153,19 @@ Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot
|
||||
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
|
||||
start it again.
|
||||
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05 (steps 1–3) and 2026-10-09 (steps 4–5)
|
||||
|
||||
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
|
||||
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
|
||||
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
|
||||
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. **Steps 4–5 were run on
|
||||
2026-10-09 into a throwaway single-node k3s** (the bench, LXC 9401, k3s v1.33.6 in a Docker container on an INTERNAL
|
||||
network — `audits/dooplex-survival-2026-10-09/partD-*.txt`): the hub image of the live version started on the restored
|
||||
copy, the customer list and host list equal live (4/4, 4/4), all 4 console passwords revealed (R-173 closed).
|
||||
|
||||
> **⚠ Corrected 2026-10-09 — a restored hub mails households AT ONCE.** Within a minute of starting on the copy it
|
||||
> tried to send two households a pending „kernel notice" (blocked only because the test had no network). Turning
|
||||
> operator mail off (`operator_enabled: false`) does NOT stop household mail. So: **a test restore has no network at
|
||||
> all** (no route to Resend, to ep0, to the boxes). **A real recovery** should expect the notices the copy still holds
|
||||
> to go out once — check the hub's pending notices before giving it a network if that matters.
|
||||
|
||||
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
|
||||
the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
@@ -183,9 +191,19 @@ the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
|
||||
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
|
||||
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
|
||||
**Corrected 2026-10-09:** the Deployment also needs, NOT optional, `Secret/resend-api` (`RESEND_API_KEY`),
|
||||
`Secret/report-api` (`REPORT_API_KEY`) and `Secret/gitea-creds` (`username`, `password`); without them the pod does
|
||||
not start. On a real rebuild their values come from DooPlex's nightly secrets export, which since 2026-10-09 is also
|
||||
off-site (`runbooks/gitea-restore.md`, `secrets/*.gpg`, opened with DooPlex's restic passphrase). A test uses dummies.
|
||||
The hub's image is pulled from Gitea's registry — on a rebuild with no registry, `docker save` it from any machine
|
||||
that has it, or build it from the restored code.
|
||||
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
|
||||
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
|
||||
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
|
||||
As run 2026-10-09: `kubectl scale deploy/hub --replicas=0`; a `busybox` pod mounting `hub-data` at `/data`;
|
||||
`kubectl cp hub.db hubdb-copy:/data/hub.db.new`, then in the pod `rm -f /data/hub.db-wal /data/hub.db-shm && mv
|
||||
/data/hub.db.new /data/hub.db`; delete the pod; scale to 1 (20 s to Ready). The customer list is `GET /configs`
|
||||
(rows link to `/customers/<id>`), the hosts `GET /hosts`.
|
||||
6. Shred `k`, `enc.key` copies and `out/` when done.
|
||||
|
||||
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
|
||||
|
||||
Reference in New Issue
Block a user