From 4c388a398ba82f24b61fb09bed3b308a019e15ed Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 9 Oct 2026 12:13:41 +0200 Subject: [PATCH] R-923/R-924: total-loss-of-dooplex runbook (every recovery secret, where it lives, walked on paper), break-glass sheet (names only), Vaultwarden off-site proven (push, restore test, throwaway start); 'off DooPlex' claims corrected; R-924 filed Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- STATUS.md | 4 +- .../04-control-plane-authorization.md | 4 +- .../vaultwarden/push-1.txt | 21 ++++ .../vaultwarden/restore-test-1.txt | 4 + .../vaultwarden/throwaway-1.txt | 37 +++++++ documentation/backlog/OPEN-ITEMS.md | 3 +- .../runbooks/RUNBOOK-hub-db-offsite-backup.md | 9 +- documentation/runbooks/break-glass-sheet.md | 62 +++++++++++ documentation/runbooks/gitea-restore.md | 2 +- documentation/runbooks/secrets.md | 5 + .../runbooks/total-loss-of-dooplex.md | 100 ++++++++++++++++++ 11 files changed, 244 insertions(+), 7 deletions(-) create mode 100644 documentation/audits/dooplex-survival-2026-10-09/vaultwarden/push-1.txt create mode 100644 documentation/audits/dooplex-survival-2026-10-09/vaultwarden/restore-test-1.txt create mode 100644 documentation/audits/dooplex-survival-2026-10-09/vaultwarden/throwaway-1.txt create mode 100644 documentation/runbooks/break-glass-sheet.md create mode 100644 documentation/runbooks/total-loss-of-dooplex.md diff --git a/STATUS.md b/STATUS.md index b9ad25f6..fc31f66e 100644 --- a/STATUS.md +++ b/STATUS.md @@ -15,8 +15,8 @@ passwords open. Found: a restored hub mails households at once. The test had no network, so nothing went out. The runbook now says so. - **Every backup job on DooPlex now mails you when it fails.** A test mail reached the inbox. -- **You need to do one thing:** save the new key's paper copy in your password manager (the report has the command). - If you do nothing, the copy on ep0 cannot be opened after DooPlex is lost. +- **You need to do one thing:** print the new key on paper (see the afternoon entry: the password manager runs on + DooPlex, so it is not enough). If you do nothing, the copy on ep0 cannot be opened after DooPlex is lost. - **Not copied on purpose:** the container registry (27.7 GB). The images rebuild from the code. ## Morning (2026-10-09): released, delivered, and four of your answers proven live diff --git a/documentation/architecture/04-control-plane-authorization.md b/documentation/architecture/04-control-plane-authorization.md index 97b696fa..041628bb 100644 --- a/documentation/architecture/04-control-plane-authorization.md +++ b/documentation/architecture/04-control-plane-authorization.md @@ -109,7 +109,9 @@ only**: both keys sit on DooPlex at `/mnt/5_hdd/felhom.eu/felhom-op-operational` 2026-09-15 and were tightened the same day (R-533). **CC may sign `agent_update` ops with the operational key until the first PAYING customer exists; testers do not count.** Every signature is per box (the blob binds `host_id`), so one box moves at a time and a fenced box cannot be swept along. -Proven twice: `demo-hp-bb76ea` 2026-09-15, `demo-felhom-8363b5` 2026-09-16. **Revisit on the first +Proven twice: `demo-hp-bb76ea` 2026-09-15, `demo-felhom-8363b5` 2026-09-16. **Measured 2026-10-09 (R-924):** all three +key files open with an EMPTY passphrase, and `/mnt/5_hdd/felhom.eu` is in no backup — so today neither the "passphrase- +protected" nor the "kept offline" of §1 holds, and a loss of DooPlex loses both keys at once. **Revisit on the first sale** — at that point the key belongs behind the operator (or a hardware key, §7), and a fleet rollout step still has to be designed (R-530). diff --git a/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/push-1.txt b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/push-1.txt new file mode 100644 index 00000000..628451b5 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/push-1.txt @@ -0,0 +1,21 @@ +## 2026-10-09T10:02:43Z install from felhom.eu 4cee21ac +push_rc=0 +2026-10-09T12:02:44+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: database: 20261009-100001, 6294981 bytes, 2 min old +2026-10-09T12:04:15+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: files: 10 repositories, 616 MB +2026-10-09T12:04:15+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: secrets: 3 file(s) of 20261009_031011 +2026-10-09T12:04:16+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: vaultwarden: integrity ok, 1 user(s), files: db.sqlite3 rsa_key.pem +2026-10-09T12:04:31+02:00 dooplex felhom-dooplex-offsite[4152837]: dooplex.pxar: had to backup 138.87 MiB of 546.853 MiB (compressed 133.173 MiB) in 11.79 s (average 11.783 MiB/s) +2026-10-09T12:04:31+02:00 dooplex felhom-dooplex-offsite[4152837]: dooplex.pxar: backup was done incrementally, reused 407.983 MiB (74.6%) +2026-10-09T12:04:32+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: pushed to ep0 (ns operator, host/dooplex-gitea) in 13 s +2026-10-09T12:04:32+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: success signal written +db.sqlite3 +db.sqlite3-shm +db.sqlite3-wal +icon_cache +lost+found +rsa_key.pem +tmp +felhom_dooplex_offsite_last_success_timestamp_seconds 1791540272 +felhom_dooplex_offsite_last_success_bytes 568894843 +felhom_dooplex_offsite_last_success_repositories 10 +felhom_dooplex_offsite_last_success_vaultwarden_users 1 diff --git a/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/restore-test-1.txt b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/restore-test-1.txt new file mode 100644 index 00000000..a2107da9 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/restore-test-1.txt @@ -0,0 +1,4 @@ +rc=0 +2026-10-09T12:04:40+02:00 dooplex felhom-dooplex-offsite-restore-test[4155906]: felhom-dooplex-offsite-restore-test: restoring host/dooplex-gitea/2026-10-09T10:04:19Z +2026-10-09T12:05:27+02:00 dooplex felhom-dooplex-offsite-restore-test[4155906]: felhom-dooplex-offsite-restore-test: checked: 27969 files match the manifest, 10 repositories pass git fsck, gitea.dump readable, 3 secrets file(s), Vaultwarden 1 user(s) / 797 item(s) +2026-10-09T12:05:27+02:00 dooplex felhom-dooplex-offsite-restore-test[4155906]: felhom-dooplex-offsite-restore-test: success signal written diff --git a/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/throwaway-1.txt b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/throwaway-1.txt new file mode 100644 index 00000000..520fc397 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/throwaway-1.txt @@ -0,0 +1,37 @@ +## 2026-10-09T10:05:39Z restore on DooPlex (read-only token), Vaultwarden part only to the bench +restore complete (546.853 MiB processed in 50.5s, average 10.821 MiB/s) +total 1136 +drwx------ 2 root root 4096 Oct 9 12:04 . +drwx------ 6 root root 4096 Oct 9 12:04 .. +-rw------- 1 root root 2 Oct 9 12:04 USERS +-rw------- 1 root root 1146880 Oct 9 12:04 db.sqlite3 +-rw-r--r-- 1 root root 1679 Dec 9 2025 rsa_key.pem +USERS +db.sqlite3 +rsa_key.pem +## row counts (python sqlite3, read-only) +integrity: ok +users 1 rows +ciphers 797 rows +folders 5 rows +organizations 0 rows +attachments 0 rows +sends 0 rows +twofactor 0 rows +sso_users 1 rows +devices 13 rows +docker.io/vaultwarden/server:1.37.3 +## started on the restored data after ~2 s: /alive -> 200 +## /api/config (public, no login): version 2026.6.0 server Vaultwarden +## no route out: unreachable +## 2026-10-09T10:07:13Z teardown — bench +anon-volumes-of-vw-check=[] +image-removed +containers=0 volumes=0 +.bashrc +.profile +.ssh +status: stopped +## teardown — DooPlex +.cache +.kube diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index fdd45928..39f381dc 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -192,10 +192,11 @@ stopping line that lies. | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -## Security & access — 14 rows (P3 10, P4 4) +## Security & access — 15 rows (P2 1, P3 10, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| +| **R-924** | Security & access | P2 | **The operator signing keys exist only on DooPlex, with no passphrase and in no backup — a loss of DooPlex means no box ever accepts a signed agent update, config bundle or OS step again without an on-site re-enrollment.** MEASURED 2026-10-09: `/mnt/5_hdd/felhom.eu/felhom-op-operational`, `felhom-rec-recovery`, `felhom_op_ed25519` (0600 `kisfenyo`); `ssh-keygen -y -P ''` opens all three (no passphrase); `/mnt/5_hdd/felhom.eu` is in no backup set, the off-site job does not carry them. `04` §1 says the operational key is "passphrase-protected" and the recovery key "kept offline" — today neither holds (custody ruling 1 of 2026-09-16 put both on DooPlex for this phase; `04` §3.1). Both keys share one fate, so the cold key cannot do its one job (re-pin after losing the operational key). `runbooks/total-loss-of-dooplex.md` S6. | **WAITING-ON-OPERATOR** — custody: print both (the break-glass sheet has the lines and commands) and/or add them to the nightly encrypted copy (a DooPlex change), or move the recovery key off DooPlex entirely. | operator decision | Operator: pick; CC: the copy change if picked | operator | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | diff --git a/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md b/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md index bf199841..351c4178 100644 --- a/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md +++ b/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md @@ -22,7 +22,12 @@ ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches (127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places, one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees. -### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved) +### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved) — **CORRECTED 2026-10-09: NOT off DooPlex** + +> **Corrected 2026-10-09 (operator finding, R-923):** "the password manager" here is Vaultwarden, which **runs on +> DooPlex**. A key kept only there is lost with DooPlex — so on 2026-10-05 both keys were **not** off DooPlex, and R-173's +> "both keys are off DooPlex" was wrong in that sense. The fix: Vaultwarden itself now rides the nightly off-site copy +> (behind the DooPlex off-site key), and the keys go on paper: `break-glass-sheet.md` (S2 = this backup key, S4 = the seal key). Without these, every later step backs up something nobody can open after a DooPlex loss. @@ -167,7 +172,7 @@ copy, the customer list and host list equal live (4/4, 4/4), all 4 console passw > all** (no route to Resend, to ep0, to the boxes). **A real recovery** should expect the notices the copy still holds > to go out once — check the hub's pending notices before giving it a network if that matters. -What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and +What you need, from the break-glass sheet (`break-glass-sheet.md`; the password manager only if it survived or was restored first — `total-loss-of-dooplex.md` step 5): the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and the read-only token (or ep0 root to mint a new one: Step 2). 1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write diff --git a/documentation/runbooks/break-glass-sheet.md b/documentation/runbooks/break-glass-sheet.md new file mode 100644 index 00000000..bc9f0633 --- /dev/null +++ b/documentation/runbooks/break-glass-sheet.md @@ -0,0 +1,62 @@ +# Break-glass sheet — the keys that must exist outside DooPlex (R-923) + +> **A template. It holds NO key value, and none may ever be written into this file, a commit, a log or a chat.** +> Print this page, then write or stick each value onto the paper copy only. Keep the paper away from home (DooPlex is +> at home: a fire takes both). Why each key is here and what it opens: `total-loss-of-dooplex.md`. +> +> **A key in the password manager is not enough:** Vaultwarden runs on DooPlex (operator ruling, 2026-10-09). + +## How to print a key without it landing anywhere + +Run these from your **own workstation** (not through Claude Code, not through `!`). Each command writes one file on +the workstation; open it, print it, then destroy the file. Nothing is stored on DooPlex or in any log. + +```bash +D=kisfenyo@192.168.0.180 # DooPlex +# S1 and S2 — PBS keys: Proxmox's own paper form (text + QR code) +ssh -t $D 'sudo proxmox-backup-client key paperkey /etc/felhom-dooplex-offsite/enc.key --output-format html --subject "S1 DooPlex off-site key"' > s1.html +ssh -t $D 'sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format html --subject "S2 Hub-DB off-site key"' > s2.html +# S4, S5, S7 — one line each +ssh -t $D 'sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath="{.data.OFFSITE_SECRET_KEY}" | base64 -d' > s4.txt +ssh -t $D 'sudo cat /etc/backup/restic-password' > s5.txt +ssh -t $D 'sudo cat /etc/felhom-hub-backup/token-restore' > s7.txt +# S6 — the signing keys (OpenSSH text, ~7 lines each) +ssh $D 'cat /mnt/5_hdd/felhom.eu/felhom-rec-recovery' > s6-recovery.txt +ssh $D 'cat /mnt/5_hdd/felhom.eu/felhom-op-operational' > s6-operational.txt +# open each, print, check the print is readable, then: +shred -u s1.html s2.html s4.txt s5.txt s7.txt s6-recovery.txt s6-operational.txt +``` + +On Windows without `shred`: delete the files and empty the recycle bin. The `-t` adds a carriage return to the +captured line on some systems; strip it when typing the value back. + +**Check a print later without exposing it:** for S1/S2 the only reliable check is one restore with a key file rebuilt +from the printed `data` field (hub-DB runbook Step 0 note) — `key show` cannot check a rebuilt key. + +--- + +## The sheet (print from here) + +**FELHOM — break-glass keys.** Printed on: ____________ Stored at: ______________________________ + +| # | Key | Public fingerprint (to match the right key) | Value — write or stick here | +|---|---|---|---| +| S1 | DooPlex off-site key — opens Gitea, the password manager, the k8s Secrets export (ep0 `operator`, `host/dooplex-gitea`) | PBS key `93:03:bf:d7:1f:4c:9e:fe…` | `data`: ______________________________________________ | +| S2 | Hub-DB off-site key — opens the hub database (ep0 `operator`, `host/dooplex-hub`) | PBS key `b2:19:bf:36:3b:97:3d:6c…` | `data`: ______________________________________________ | +| S4 | Hub seal key `OFFSITE_SECRET_KEY` (64 hex) | — | ______________________________________________________ | +| S5 | DooPlex restic / Secrets-export passphrase | — | ______________________________________________________ | +| S6a | Signing RECOVERY key `felhom-rec-recovery` | `SHA256:/ixgTesZqykAGJpFUUd4kLAiHFgKOkYFLNC3AQXWP+k` | (staple the printout) | +| S6b | Signing OPERATIONAL key `felhom-op-operational` | `SHA256:7YqN4rXO08yixTeOO+UtQ8jHyIGycICuctQgRYVGnWw` | (staple the printout) | +| S7 | ep0 read-only token `dooplex-hub@pbs!restore` | — | ______________________________________________________ | +| S8 | Hetzner account: login e-mail · password · two-factor recovery codes | — | ______________________________________________________ | +| S11 | Cloudflare account: login · password · two-factor recovery codes | — | ______________________________________________________ | +| S12 | Gmail `felhom.eu@gmail.com`: password · two-factor recovery codes | — | ______________________________________________________ | +| S3 | Vaultwarden master password — **in your head**; write it here only if you decide to | — | ______________________________________________________ | + +**Not secret, needed with the keys:** +- ep0: `167.233.158.164` (Hetzner, `felhom-hetzner`); PBS datastore `felhom-offsite`, namespace `operator`; + its certificate fingerprint `c6:07:28:3f:5b:7b:5a:41:90:28:d7:ca:4f:37:14:70:56:39:2e:2f:0b:71:e8:06:ca:60:4a:d5:56:5f:3c:fd`. +- Restore order: `total-loss-of-dooplex.md` (in the restored Gitea, repo `felhom.eu`, `documentation/runbooks/`). + **Print that page too** — the runbook itself is inside the copy it explains how to open. + +**Re-print when** a key above is rotated, and once a year. diff --git a/documentation/runbooks/gitea-restore.md b/documentation/runbooks/gitea-restore.md index 40f909e5..a3d26ab4 100644 --- a/documentation/runbooks/gitea-restore.md +++ b/documentation/runbooks/gitea-restore.md @@ -25,7 +25,7 @@ One encrypted archive `dooplex.pxar` per night in ep0's PBS, namespace `operator ## What you need -- **The key**: the `data` field of the paper key from the password manager („DooPlex off-site (Gitea) key"). Write +- **The key**: the `data` field of the paper key — **S1 on the break-glass sheet** (`break-glass-sheet.md`). Not the password manager: it runs on DooPlex and is itself inside this copy (R-923). Write `{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": ""}` to `enc.key` (root, `umask 077`). On DooPlex it is `/etc/felhom-dooplex-offsite/enc.key`. - **A read-only token** for `dooplex-hub@pbs!restore` (DooPlex: `/etc/felhom-hub-backup/token-restore`), or ep0 root to diff --git a/documentation/runbooks/secrets.md b/documentation/runbooks/secrets.md index 360d1e1a..5fec9d3f 100644 --- a/documentation/runbooks/secrets.md +++ b/documentation/runbooks/secrets.md @@ -5,6 +5,11 @@ committed to git. The manifests in `manifests/` carry only **placeholders + comm The live values live in the operator's out-of-band store (password manager / escrow) and in the cluster as Secrets created imperatively with the commands below. +> **Corrected 2026-10-09 (R-923):** the password manager is Vaultwarden, which **runs on DooPlex** — it is not out of +> band for a loss of DooPlex. Every Secret below is also in DooPlex's nightly Secrets export, which since 2026-10-09 rides +> the encrypted off-site copy; the keys that open it are on the break-glass sheet (`break-glass-sheet.md`, +> `total-loss-of-dooplex.md`). + > The `felhom` ArgoCD app syncs the **entire** `manifests/` dir, so any Secret committed there is > re-applied on every sync. Secrets that must stay out of git are therefore created as objects that are > **not present in `manifests/`** (ArgoCD does not manage them), e.g. `Secret/resend-api` below. diff --git a/documentation/runbooks/total-loss-of-dooplex.md b/documentation/runbooks/total-loss-of-dooplex.md new file mode 100644 index 00000000..f0b3b2a6 --- /dev/null +++ b/documentation/runbooks/total-loss-of-dooplex.md @@ -0,0 +1,100 @@ +# Runbook — total loss of DooPlex: bring the business back on new hardware (R-232, R-923) + +> **Written 2026-10-09. Not walked end to end.** Its parts were walked separately on throwaways: Gitea +> (`gitea-restore.md`), the hub (`RUNBOOK-hub-db-offsite-backup.md` §3), Vaultwarden (rows + start, +> `audits/dooplex-survival-2026-10-09/vaultwarden/`). **This page names secrets, never values.** The printable page that +> holds the few values that must exist outside DooPlex is `break-glass-sheet.md`. + +## What a total loss means, and what survives + +DooPlex (192.168.0.180, the home server) holds: Gitea (all code + the container registry), the hub (k3s), Vaultwarden +(the operator's password manager), DooPlex's own PBS, every local backup, the signing keys, the CI runner, the +build tree. **Assume all of it is gone, and the house with it.** + +| Survives (off DooPlex) | What it holds | Opened by | +|---|---|---| +| ep0 (Hetzner, `167.233.158.164`), PBS `felhom-offsite`, namespace `operator` | `host/dooplex-hub` — the hub database, nightly; `host/dooplex-gitea` — Gitea (repos, DB dump, `app.ini`), **Vaultwarden** (DB + `rsa_key.pem`), DooPlex's nightly k8s Secrets export (GPG), nightly | S1 (+ S2 for the hub DB), a PBS token or ep0 root | +| ep0, the households' namespaces | each household's whole-box copies | each household's own key (escrow) — not the operator's | +| Hetzner Storage Boxes | each household's restic copy | each household's own password (in the box, and in its whole-box copies) | +| The boxes themselves | they keep running and backing up; they cannot report, update or mail until the hub is back | — | +| The operator's workstation (if it has them) | the WireGuard peer to ep0 (`10.77.0.250`), maybe one of ep0's two root SSH keys, maybe a Bitwarden client's offline vault cache | the operator | + +**Lost with DooPlex and NOT in any copy** (measured 2026-10-09): the container registry (rebuild the images from +code); **the operator signing keys** `felhom-op-operational`, `felhom-rec-recovery`, `felhom_op_ed25519` +(`/mnt/5_hdd/felhom.eu/`, no passphrase, no backup — R-924); DooPlex's own SSH key (one of ep0's two root keys); +the build tree and drills (`/mnt/5_hdd/felhom.eu/{build,drills}`); `.claude-memory` (same-disk restic only); the +local restic repos; Prometheus history; DooPlex's other homelab apps (not Felhom; Longhorn on DooPlex only). + +## The secrets a full recovery needs — where each lives today + +**S = must exist outside DooPlex (on the sheet).** "DooPlex only" = a total loss loses it unless the sheet carries it. +"Vaultwarden" = the password manager **on DooPlex** — a key kept only there is **not** off DooPlex (operator ruling, +2026-10-09); it comes back only after S1 and the master password. + +| # | Secret / login | Needed for | Lives today | Lost in a total loss? | +|---|---|---|---|---| +| **S1** | **DooPlex off-site key** (`/etc/felhom-dooplex-offsite/enc.key`, its paper `data` field) | opening `host/dooplex-gitea`: Gitea, Vaultwarden, the Secrets export | DooPlex; Vaultwarden if the operator saved it there | **YES unless on paper** — and it is the key to everything else | +| **S2** | **Hub-DB off-site key** (`/etc/felhom-hub-backup/enc.key`, `data` field) | opening `host/dooplex-hub` | DooPlex; Vaultwarden (saved 2026-10-05) | yes unless on paper (Vaultwarden is behind S1 + S3) | +| **S3** | **Vaultwarden master password** (user count 1; **no two-factor rows**, SSO-linked to Authentik but password login allowed) | every item in the password manager | the operator's head | no — if remembered | +| **S4** | **Hub seal key** `OFFSITE_SECRET_KEY` | the hub opening its console passwords and off-site passwords | DooPlex (`Secret/offsite-secret-key`); inside the Secrets export (behind S1 + S5); Vaultwarden | yes unless on paper or S1+S5 work | +| **S5** | **DooPlex restic/GPG passphrase** (`/etc/backup/restic-password`) | opening the nightly Secrets export (`secrets/*.gpg`: every k8s Secret — Resend, report-api, Hetzner token, the hub's SSH keys, gitea-creds, TLS …) | DooPlex (`sdb1`); "an offline copy, out of band" (operator, recon 2026-08-06 — form unknown to CC). **Checked 2026-10-09:** the newest export opens with it (symmetric GPG, AES256) and holds 428 Secrets, among them all eight the hub uses (names only read) | depends on that offline copy | +| **S6** | **Signing recovery key** `felhom-rec-recovery` (and the operational key `felhom-op-operational`) | signing agent updates / config bundles / OS steps for every box; the recovery key authorizes rotating the operational key | **DooPlex only**, no passphrase, in no backup | **YES** — then every box needs on-site re-enrollment to accept a new key (`04` §4) | +| **S7** | **ep0 read-only restore token** `dooplex-hub@pbs!restore` | reading both copies without ep0 root | DooPlex only (`/etc/felhom-hub-backup/token-restore`) | yes — but ep0 root can mint a new one (hub-DB runbook Step 2) | +| **S8** | **Hetzner account login + its two-factor** | ep0's console (rescue, root reset), ep0 itself, Storage Boxes, the Hetzner API token | the operator; probably Vaultwarden | **circle**: if only in Vaultwarden, it is behind the copy it is needed to reach — paper or head | +| S9 | ep0 root SSH | the PBS tunnel; minting tokens | ep0 `authorized_keys`: DooPlex's key (lost) + one more key, owner not recorded (operator workstation?) | partly — Hetzner console (S8) is the fallback | +| S10 | WireGuard operator peer (`10.77.0.250`) | reaching ep0's PBS (`10.77.0.1:8007`) without SSH | the operator's workstation | no, if the workstation survives | +| S11 | **Cloudflare account login** (+2FA) | DNS of `felhom.eu` and `dooplex.hu` (both on Cloudflare); pointing `hub.felhom.eu` at the new home; households' tunnels | the operator; Vaultwarden | circle unless in head/paper | +| S12 | **Gmail `felhom.eu@gmail.com`** (+2FA) | the `@felhom.eu` catch-all inbox; account recovery for other services | the operator | — | +| S13 | Domain registrar(s) for `felhom.eu`, `dooplex.hu`; No-IP (`dooplex.hopto.org`, the home address CNAME) | moving DNS / the home address | the operator; Vaultwarden | — | +| S14 | Resend account / API key | all mail | Secrets export (S1+S5); Vaultwarden; the Resend dashboard (re-mint) | no — re-mint at Resend | +| S15 | Gitea admin password; registry credentials (`gitea-creds`) | Gitea UI; image push/pull | the restored Gitea DB (password hash) + Vaultwarden; the Secrets export | no, after S1 (+S3) | +| S16 | The hub operator password (`HUB_PW`) | the hub UI | the restored hub DB (hash); `~/.config/credentials` on DooPlex; Vaultwarden | no, after S2 (+S3) | +| S17 | Authentik admin | SSO for Vaultwarden and DooPlex apps | Authentik's DB (Longhorn on DooPlex, local PostgreSQL dumps only) | **yes** — not in any copy; Vaultwarden does not need it (password login) | +| S18 | Households' recovery codes, PBS keys, restic passwords | households' restores | the households / escrow / the boxes | not the operator's | + +## The order of steps (new hardware) + +1. **A machine and the sheet.** Any Linux machine with Docker; for the hub also k3s. Install `proxmox-backup-client`. +2. **Reach ep0's PBS.** Either the operator's WireGuard peer (S10) → `10.77.0.1:8007`; or ep0 root SSH (S9) and the + tunnel `ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@167.233.158.164`; or the Hetzner console (S8) → root → add a + new SSH key. Fingerprint pin: `proxmox-backup-manager cert info` on ep0. +3. **A read token.** S7 from the sheet, or mint a new one on ep0 (hub-DB runbook Step 2: `token/restore`, ACL + `DatastoreReader` on `/datastore/felhom-offsite/operator`). +4. **Restore `host/dooplex-gitea`** with S1 (key file rebuilt from the `data` field — hub-DB runbook Step 0 note). +5. **Vaultwarden first** (it holds the rest): `docker run -v /vaultwarden:/data vaultwarden/server:` + (`DOMAIN`, `SIGNUPS_ALLOWED=false`); log in with the e-mail and S3. Or, if a Bitwarden client still holds an offline + copy, use that. Every later login comes from here. +6. **The Secrets export:** `gpg --batch --passphrase-file -d secrets-.yaml.gpg > secrets.yaml` — + every k8s Secret of the old cluster (shred it after use). +7. **Gitea:** `gitea-restore.md` (on a normal network this time). Then rebuild the images from the code into the new + registry: hub, controller, agent (`RUNBOOK-manual-build.md`). +8. **The hub:** k3s, then `RUNBOOK-hub-db-offsite-backup.md` §3 with S2 and S4 — **start it with mail held** + (`MAIL-HOLD`, from the next hub release; until then: no network until the pending notices are understood). +9. **DNS:** in Cloudflare (S11) point `hub.felhom.eu` (today a CNAME to `dooplex.hopto.org`) and `gitea.dooplex.hu` at + the new place. The boxes reconnect by themselves: their API keys are in the restored hub DB. +10. **Signing:** with S6 on paper, put the keys back and sign as before. Without it, no box accepts an agent update or a + bundle until it is re-enrolled on site (`04` §4). +11. **Backups again:** reinstall `scripts/hub-db-backup/` and `scripts/dooplex-offsite/` from the restored code, the + tunnel and the tokens, and check the next night's copies. + +## Walked on paper (2026-10-09): what still needs DooPlex if the operator has only the sheet, memory and a new machine + +| Step | Needs | From the sheet / memory? | +|---|---|---| +| reach ep0 | WireGuard peer **or** ep0 root **or** Hetzner login | WireGuard peer is on the workstation, **not** on the sheet; ep0 root key: one of two keys is DooPlex's; **Hetzner login (S8) must be on the sheet** | +| PBS fingerprint | ep0's cert fingerprint | not secret; on the sheet as a line, or read on ep0 | +| a read token | S7 or ep0 root | on the sheet (S7), or via S8 | +| open the Gitea/Vaultwarden copy | S1 | **sheet** | +| open the vault | S3 | memory | +| open the Secrets export | S5 | **sheet** (today: "offline copy", form unknown) | +| the hub's seal key | S4 (or S1+S5) | sheet, or via S5 | +| the hub DB | S2 (or Vaultwarden) | sheet, or via S3 | +| sign a box update | S6 | **sheet — today exists on DooPlex only** | +| hub image | the registry (lost) | rebuild from restored code — needs a Go/Docker build machine, no secret | +| Gitea `main` → `felhom.eu` website, ISO host | Cloudflare (S11), the build tree (lost) | S11 on the sheet; the build tree is rebuilt from code | +| Authentik (S17) | not in any copy | not needed for the business; DooPlex homelab apps that use it are out of scope | + +**Result:** with a filled sheet (S1, S2, S4, S5, S6, S7, S8, S11 + the ep0 fingerprint) and the master password in +memory, nothing in the Felhom recovery still needs DooPlex. **Without the sheet, three things block:** S1 (no copy +opens), S6 (no box can be signed for), S8 (ep0 cannot be reached once DooPlex's key is gone, unless the workstation's +WireGuard peer survives).