R-923/R-924: total-loss-of-dooplex runbook (every recovery secret, where it lives, walked on paper), break-glass sheet (names only), Vaultwarden off-site proven (push, restore test, throwaway start); 'off DooPlex' claims corrected; R-924 filed
gates / gates (push) Successful in 5m41s
gates / gates (push) Successful in 5m41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -15,8 +15,8 @@
|
||||
passwords open. Found: a restored hub mails households at once. The test had no network, so nothing went out.
|
||||
The runbook now says so.
|
||||
- **Every backup job on DooPlex now mails you when it fails.** A test mail reached the inbox.
|
||||
- **You need to do one thing:** save the new key's paper copy in your password manager (the report has the command).
|
||||
If you do nothing, the copy on ep0 cannot be opened after DooPlex is lost.
|
||||
- **You need to do one thing:** print the new key on paper (see the afternoon entry: the password manager runs on
|
||||
DooPlex, so it is not enough). If you do nothing, the copy on ep0 cannot be opened after DooPlex is lost.
|
||||
- **Not copied on purpose:** the container registry (27.7 GB). The images rebuild from the code.
|
||||
|
||||
## Morning (2026-10-09): released, delivered, and four of your answers proven live
|
||||
|
||||
@@ -109,7 +109,9 @@ only**: both keys sit on DooPlex at `/mnt/5_hdd/felhom.eu/felhom-op-operational`
|
||||
2026-09-15 and were tightened the same day (R-533). **CC may sign `agent_update` ops with the
|
||||
operational key until the first PAYING customer exists; testers do not count.** Every signature is
|
||||
per box (the blob binds `host_id`), so one box moves at a time and a fenced box cannot be swept along.
|
||||
Proven twice: `demo-hp-bb76ea` 2026-09-15, `demo-felhom-8363b5` 2026-09-16. **Revisit on the first
|
||||
Proven twice: `demo-hp-bb76ea` 2026-09-15, `demo-felhom-8363b5` 2026-09-16. **Measured 2026-10-09 (R-924):** all three
|
||||
key files open with an EMPTY passphrase, and `/mnt/5_hdd/felhom.eu` is in no backup — so today neither the "passphrase-
|
||||
protected" nor the "kept offline" of §1 holds, and a loss of DooPlex loses both keys at once. **Revisit on the first
|
||||
sale** — at that point the key belongs behind the operator (or a hardware key, §7), and a fleet
|
||||
rollout step still has to be designed (R-530).
|
||||
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
## 2026-10-09T10:02:43Z install from felhom.eu 4cee21ac
|
||||
push_rc=0
|
||||
2026-10-09T12:02:44+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: database: 20261009-100001, 6294981 bytes, 2 min old
|
||||
2026-10-09T12:04:15+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: files: 10 repositories, 616 MB
|
||||
2026-10-09T12:04:15+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: secrets: 3 file(s) of 20261009_031011
|
||||
2026-10-09T12:04:16+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: vaultwarden: integrity ok, 1 user(s), files: db.sqlite3 rsa_key.pem
|
||||
2026-10-09T12:04:31+02:00 dooplex felhom-dooplex-offsite[4152837]: dooplex.pxar: had to backup 138.87 MiB of 546.853 MiB (compressed 133.173 MiB) in 11.79 s (average 11.783 MiB/s)
|
||||
2026-10-09T12:04:31+02:00 dooplex felhom-dooplex-offsite[4152837]: dooplex.pxar: backup was done incrementally, reused 407.983 MiB (74.6%)
|
||||
2026-10-09T12:04:32+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: pushed to ep0 (ns operator, host/dooplex-gitea) in 13 s
|
||||
2026-10-09T12:04:32+02:00 dooplex felhom-dooplex-offsite[4134357]: felhom-dooplex-offsite: success signal written
|
||||
db.sqlite3
|
||||
db.sqlite3-shm
|
||||
db.sqlite3-wal
|
||||
icon_cache
|
||||
lost+found
|
||||
rsa_key.pem
|
||||
tmp
|
||||
felhom_dooplex_offsite_last_success_timestamp_seconds 1791540272
|
||||
felhom_dooplex_offsite_last_success_bytes 568894843
|
||||
felhom_dooplex_offsite_last_success_repositories 10
|
||||
felhom_dooplex_offsite_last_success_vaultwarden_users 1
|
||||
@@ -0,0 +1,4 @@
|
||||
rc=0
|
||||
2026-10-09T12:04:40+02:00 dooplex felhom-dooplex-offsite-restore-test[4155906]: felhom-dooplex-offsite-restore-test: restoring host/dooplex-gitea/2026-10-09T10:04:19Z
|
||||
2026-10-09T12:05:27+02:00 dooplex felhom-dooplex-offsite-restore-test[4155906]: felhom-dooplex-offsite-restore-test: checked: 27969 files match the manifest, 10 repositories pass git fsck, gitea.dump readable, 3 secrets file(s), Vaultwarden 1 user(s) / 797 item(s)
|
||||
2026-10-09T12:05:27+02:00 dooplex felhom-dooplex-offsite-restore-test[4155906]: felhom-dooplex-offsite-restore-test: success signal written
|
||||
@@ -0,0 +1,37 @@
|
||||
## 2026-10-09T10:05:39Z restore on DooPlex (read-only token), Vaultwarden part only to the bench
|
||||
restore complete (546.853 MiB processed in 50.5s, average 10.821 MiB/s)
|
||||
total 1136
|
||||
drwx------ 2 root root 4096 Oct 9 12:04 .
|
||||
drwx------ 6 root root 4096 Oct 9 12:04 ..
|
||||
-rw------- 1 root root 2 Oct 9 12:04 USERS
|
||||
-rw------- 1 root root 1146880 Oct 9 12:04 db.sqlite3
|
||||
-rw-r--r-- 1 root root 1679 Dec 9 2025 rsa_key.pem
|
||||
USERS
|
||||
db.sqlite3
|
||||
rsa_key.pem
|
||||
## row counts (python sqlite3, read-only)
|
||||
integrity: ok
|
||||
users 1 rows
|
||||
ciphers 797 rows
|
||||
folders 5 rows
|
||||
organizations 0 rows
|
||||
attachments 0 rows
|
||||
sends 0 rows
|
||||
twofactor 0 rows
|
||||
sso_users 1 rows
|
||||
devices 13 rows
|
||||
docker.io/vaultwarden/server:1.37.3
|
||||
## started on the restored data after ~2 s: /alive -> 200
|
||||
## /api/config (public, no login): version 2026.6.0 server Vaultwarden
|
||||
## no route out: unreachable
|
||||
## 2026-10-09T10:07:13Z teardown — bench
|
||||
anon-volumes-of-vw-check=[]
|
||||
image-removed
|
||||
containers=0 volumes=0
|
||||
.bashrc
|
||||
.profile
|
||||
.ssh
|
||||
status: stopped
|
||||
## teardown — DooPlex
|
||||
.cache
|
||||
.kube
|
||||
@@ -192,10 +192,11 @@ stopping line that lies.
|
||||
| **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC |
|
||||
| **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
|
||||
|
||||
## Security & access — 14 rows (P3 10, P4 4)
|
||||
## Security & access — 15 rows (P2 1, P3 10, P4 4)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-924** | Security & access | P2 | **The operator signing keys exist only on DooPlex, with no passphrase and in no backup — a loss of DooPlex means no box ever accepts a signed agent update, config bundle or OS step again without an on-site re-enrollment.** MEASURED 2026-10-09: `/mnt/5_hdd/felhom.eu/felhom-op-operational`, `felhom-rec-recovery`, `felhom_op_ed25519` (0600 `kisfenyo`); `ssh-keygen -y -P ''` opens all three (no passphrase); `/mnt/5_hdd/felhom.eu` is in no backup set, the off-site job does not carry them. `04` §1 says the operational key is "passphrase-protected" and the recovery key "kept offline" — today neither holds (custody ruling 1 of 2026-09-16 put both on DooPlex for this phase; `04` §3.1). Both keys share one fate, so the cold key cannot do its one job (re-pin after losing the operational key). `runbooks/total-loss-of-dooplex.md` S6. | **WAITING-ON-OPERATOR** — custody: print both (the break-glass sheet has the lines and commands) and/or add them to the nightly encrypted copy (a DooPlex change), or move the recovery key off DooPlex entirely. | operator decision | Operator: pick; CC: the copy change if picked | operator |
|
||||
| **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor |
|
||||
| **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with "<domain>"` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC |
|
||||
| **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator |
|
||||
|
||||
@@ -22,7 +22,12 @@ ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches
|
||||
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
|
||||
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
|
||||
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved)
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved) — **CORRECTED 2026-10-09: NOT off DooPlex**
|
||||
|
||||
> **Corrected 2026-10-09 (operator finding, R-923):** "the password manager" here is Vaultwarden, which **runs on
|
||||
> DooPlex**. A key kept only there is lost with DooPlex — so on 2026-10-05 both keys were **not** off DooPlex, and R-173's
|
||||
> "both keys are off DooPlex" was wrong in that sense. The fix: Vaultwarden itself now rides the nightly off-site copy
|
||||
> (behind the DooPlex off-site key), and the keys go on paper: `break-glass-sheet.md` (S2 = this backup key, S4 = the seal key).
|
||||
|
||||
Without these, every later step backs up something nobody can open after a DooPlex loss.
|
||||
|
||||
@@ -167,7 +172,7 @@ copy, the customer list and host list equal live (4/4, 4/4), all 4 console passw
|
||||
> all** (no route to Resend, to ep0, to the boxes). **A real recovery** should expect the notices the copy still holds
|
||||
> to go out once — check the hub's pending notices before giving it a network if that matters.
|
||||
|
||||
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
|
||||
What you need, from the break-glass sheet (`break-glass-sheet.md`; the password manager only if it survived or was restored first — `total-loss-of-dooplex.md` step 5): the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
|
||||
the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
|
||||
1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# Break-glass sheet — the keys that must exist outside DooPlex (R-923)
|
||||
|
||||
> **A template. It holds NO key value, and none may ever be written into this file, a commit, a log or a chat.**
|
||||
> Print this page, then write or stick each value onto the paper copy only. Keep the paper away from home (DooPlex is
|
||||
> at home: a fire takes both). Why each key is here and what it opens: `total-loss-of-dooplex.md`.
|
||||
>
|
||||
> **A key in the password manager is not enough:** Vaultwarden runs on DooPlex (operator ruling, 2026-10-09).
|
||||
|
||||
## How to print a key without it landing anywhere
|
||||
|
||||
Run these from your **own workstation** (not through Claude Code, not through `!`). Each command writes one file on
|
||||
the workstation; open it, print it, then destroy the file. Nothing is stored on DooPlex or in any log.
|
||||
|
||||
```bash
|
||||
D=kisfenyo@192.168.0.180 # DooPlex
|
||||
# S1 and S2 — PBS keys: Proxmox's own paper form (text + QR code)
|
||||
ssh -t $D 'sudo proxmox-backup-client key paperkey /etc/felhom-dooplex-offsite/enc.key --output-format html --subject "S1 DooPlex off-site key"' > s1.html
|
||||
ssh -t $D 'sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format html --subject "S2 Hub-DB off-site key"' > s2.html
|
||||
# S4, S5, S7 — one line each
|
||||
ssh -t $D 'sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath="{.data.OFFSITE_SECRET_KEY}" | base64 -d' > s4.txt
|
||||
ssh -t $D 'sudo cat /etc/backup/restic-password' > s5.txt
|
||||
ssh -t $D 'sudo cat /etc/felhom-hub-backup/token-restore' > s7.txt
|
||||
# S6 — the signing keys (OpenSSH text, ~7 lines each)
|
||||
ssh $D 'cat /mnt/5_hdd/felhom.eu/felhom-rec-recovery' > s6-recovery.txt
|
||||
ssh $D 'cat /mnt/5_hdd/felhom.eu/felhom-op-operational' > s6-operational.txt
|
||||
# open each, print, check the print is readable, then:
|
||||
shred -u s1.html s2.html s4.txt s5.txt s7.txt s6-recovery.txt s6-operational.txt
|
||||
```
|
||||
|
||||
On Windows without `shred`: delete the files and empty the recycle bin. The `-t` adds a carriage return to the
|
||||
captured line on some systems; strip it when typing the value back.
|
||||
|
||||
**Check a print later without exposing it:** for S1/S2 the only reliable check is one restore with a key file rebuilt
|
||||
from the printed `data` field (hub-DB runbook Step 0 note) — `key show` cannot check a rebuilt key.
|
||||
|
||||
---
|
||||
|
||||
## The sheet (print from here)
|
||||
|
||||
**FELHOM — break-glass keys.** Printed on: ____________ Stored at: ______________________________
|
||||
|
||||
| # | Key | Public fingerprint (to match the right key) | Value — write or stick here |
|
||||
|---|---|---|---|
|
||||
| S1 | DooPlex off-site key — opens Gitea, the password manager, the k8s Secrets export (ep0 `operator`, `host/dooplex-gitea`) | PBS key `93:03:bf:d7:1f:4c:9e:fe…` | `data`: ______________________________________________ |
|
||||
| S2 | Hub-DB off-site key — opens the hub database (ep0 `operator`, `host/dooplex-hub`) | PBS key `b2:19:bf:36:3b:97:3d:6c…` | `data`: ______________________________________________ |
|
||||
| S4 | Hub seal key `OFFSITE_SECRET_KEY` (64 hex) | — | ______________________________________________________ |
|
||||
| S5 | DooPlex restic / Secrets-export passphrase | — | ______________________________________________________ |
|
||||
| S6a | Signing RECOVERY key `felhom-rec-recovery` | `SHA256:/ixgTesZqykAGJpFUUd4kLAiHFgKOkYFLNC3AQXWP+k` | (staple the printout) |
|
||||
| S6b | Signing OPERATIONAL key `felhom-op-operational` | `SHA256:7YqN4rXO08yixTeOO+UtQ8jHyIGycICuctQgRYVGnWw` | (staple the printout) |
|
||||
| S7 | ep0 read-only token `dooplex-hub@pbs!restore` | — | ______________________________________________________ |
|
||||
| S8 | Hetzner account: login e-mail · password · two-factor recovery codes | — | ______________________________________________________ |
|
||||
| S11 | Cloudflare account: login · password · two-factor recovery codes | — | ______________________________________________________ |
|
||||
| S12 | Gmail `felhom.eu@gmail.com`: password · two-factor recovery codes | — | ______________________________________________________ |
|
||||
| S3 | Vaultwarden master password — **in your head**; write it here only if you decide to | — | ______________________________________________________ |
|
||||
|
||||
**Not secret, needed with the keys:**
|
||||
- ep0: `167.233.158.164` (Hetzner, `felhom-hetzner`); PBS datastore `felhom-offsite`, namespace `operator`;
|
||||
its certificate fingerprint `c6:07:28:3f:5b:7b:5a:41:90:28:d7:ca:4f:37:14:70:56:39:2e:2f:0b:71:e8:06:ca:60:4a:d5:56:5f:3c:fd`.
|
||||
- Restore order: `total-loss-of-dooplex.md` (in the restored Gitea, repo `felhom.eu`, `documentation/runbooks/`).
|
||||
**Print that page too** — the runbook itself is inside the copy it explains how to open.
|
||||
|
||||
**Re-print when** a key above is rotated, and once a year.
|
||||
@@ -25,7 +25,7 @@ One encrypted archive `dooplex.pxar` per night in ep0's PBS, namespace `operator
|
||||
|
||||
## What you need
|
||||
|
||||
- **The key**: the `data` field of the paper key from the password manager („DooPlex off-site (Gitea) key"). Write
|
||||
- **The key**: the `data` field of the paper key — **S1 on the break-glass sheet** (`break-glass-sheet.md`). Not the password manager: it runs on DooPlex and is itself inside this copy (R-923). Write
|
||||
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
|
||||
to `enc.key` (root, `umask 077`). On DooPlex it is `/etc/felhom-dooplex-offsite/enc.key`.
|
||||
- **A read-only token** for `dooplex-hub@pbs!restore` (DooPlex: `/etc/felhom-hub-backup/token-restore`), or ep0 root to
|
||||
|
||||
@@ -5,6 +5,11 @@ committed to git. The manifests in `manifests/` carry only **placeholders + comm
|
||||
The live values live in the operator's out-of-band store (password manager / escrow) and in the cluster
|
||||
as Secrets created imperatively with the commands below.
|
||||
|
||||
> **Corrected 2026-10-09 (R-923):** the password manager is Vaultwarden, which **runs on DooPlex** — it is not out of
|
||||
> band for a loss of DooPlex. Every Secret below is also in DooPlex's nightly Secrets export, which since 2026-10-09 rides
|
||||
> the encrypted off-site copy; the keys that open it are on the break-glass sheet (`break-glass-sheet.md`,
|
||||
> `total-loss-of-dooplex.md`).
|
||||
|
||||
> The `felhom` ArgoCD app syncs the **entire** `manifests/` dir, so any Secret committed there is
|
||||
> re-applied on every sync. Secrets that must stay out of git are therefore created as objects that are
|
||||
> **not present in `manifests/`** (ArgoCD does not manage them), e.g. `Secret/resend-api` below.
|
||||
|
||||
@@ -0,0 +1,100 @@
|
||||
# Runbook — total loss of DooPlex: bring the business back on new hardware (R-232, R-923)
|
||||
|
||||
> **Written 2026-10-09. Not walked end to end.** Its parts were walked separately on throwaways: Gitea
|
||||
> (`gitea-restore.md`), the hub (`RUNBOOK-hub-db-offsite-backup.md` §3), Vaultwarden (rows + start,
|
||||
> `audits/dooplex-survival-2026-10-09/vaultwarden/`). **This page names secrets, never values.** The printable page that
|
||||
> holds the few values that must exist outside DooPlex is `break-glass-sheet.md`.
|
||||
|
||||
## What a total loss means, and what survives
|
||||
|
||||
DooPlex (192.168.0.180, the home server) holds: Gitea (all code + the container registry), the hub (k3s), Vaultwarden
|
||||
(the operator's password manager), DooPlex's own PBS, every local backup, the signing keys, the CI runner, the
|
||||
build tree. **Assume all of it is gone, and the house with it.**
|
||||
|
||||
| Survives (off DooPlex) | What it holds | Opened by |
|
||||
|---|---|---|
|
||||
| ep0 (Hetzner, `167.233.158.164`), PBS `felhom-offsite`, namespace `operator` | `host/dooplex-hub` — the hub database, nightly; `host/dooplex-gitea` — Gitea (repos, DB dump, `app.ini`), **Vaultwarden** (DB + `rsa_key.pem`), DooPlex's nightly k8s Secrets export (GPG), nightly | S1 (+ S2 for the hub DB), a PBS token or ep0 root |
|
||||
| ep0, the households' namespaces | each household's whole-box copies | each household's own key (escrow) — not the operator's |
|
||||
| Hetzner Storage Boxes | each household's restic copy | each household's own password (in the box, and in its whole-box copies) |
|
||||
| The boxes themselves | they keep running and backing up; they cannot report, update or mail until the hub is back | — |
|
||||
| The operator's workstation (if it has them) | the WireGuard peer to ep0 (`10.77.0.250`), maybe one of ep0's two root SSH keys, maybe a Bitwarden client's offline vault cache | the operator |
|
||||
|
||||
**Lost with DooPlex and NOT in any copy** (measured 2026-10-09): the container registry (rebuild the images from
|
||||
code); **the operator signing keys** `felhom-op-operational`, `felhom-rec-recovery`, `felhom_op_ed25519`
|
||||
(`/mnt/5_hdd/felhom.eu/`, no passphrase, no backup — R-924); DooPlex's own SSH key (one of ep0's two root keys);
|
||||
the build tree and drills (`/mnt/5_hdd/felhom.eu/{build,drills}`); `.claude-memory` (same-disk restic only); the
|
||||
local restic repos; Prometheus history; DooPlex's other homelab apps (not Felhom; Longhorn on DooPlex only).
|
||||
|
||||
## The secrets a full recovery needs — where each lives today
|
||||
|
||||
**S = must exist outside DooPlex (on the sheet).** "DooPlex only" = a total loss loses it unless the sheet carries it.
|
||||
"Vaultwarden" = the password manager **on DooPlex** — a key kept only there is **not** off DooPlex (operator ruling,
|
||||
2026-10-09); it comes back only after S1 and the master password.
|
||||
|
||||
| # | Secret / login | Needed for | Lives today | Lost in a total loss? |
|
||||
|---|---|---|---|---|
|
||||
| **S1** | **DooPlex off-site key** (`/etc/felhom-dooplex-offsite/enc.key`, its paper `data` field) | opening `host/dooplex-gitea`: Gitea, Vaultwarden, the Secrets export | DooPlex; Vaultwarden if the operator saved it there | **YES unless on paper** — and it is the key to everything else |
|
||||
| **S2** | **Hub-DB off-site key** (`/etc/felhom-hub-backup/enc.key`, `data` field) | opening `host/dooplex-hub` | DooPlex; Vaultwarden (saved 2026-10-05) | yes unless on paper (Vaultwarden is behind S1 + S3) |
|
||||
| **S3** | **Vaultwarden master password** (user count 1; **no two-factor rows**, SSO-linked to Authentik but password login allowed) | every item in the password manager | the operator's head | no — if remembered |
|
||||
| **S4** | **Hub seal key** `OFFSITE_SECRET_KEY` | the hub opening its console passwords and off-site passwords | DooPlex (`Secret/offsite-secret-key`); inside the Secrets export (behind S1 + S5); Vaultwarden | yes unless on paper or S1+S5 work |
|
||||
| **S5** | **DooPlex restic/GPG passphrase** (`/etc/backup/restic-password`) | opening the nightly Secrets export (`secrets/*.gpg`: every k8s Secret — Resend, report-api, Hetzner token, the hub's SSH keys, gitea-creds, TLS …) | DooPlex (`sdb1`); "an offline copy, out of band" (operator, recon 2026-08-06 — form unknown to CC). **Checked 2026-10-09:** the newest export opens with it (symmetric GPG, AES256) and holds 428 Secrets, among them all eight the hub uses (names only read) | depends on that offline copy |
|
||||
| **S6** | **Signing recovery key** `felhom-rec-recovery` (and the operational key `felhom-op-operational`) | signing agent updates / config bundles / OS steps for every box; the recovery key authorizes rotating the operational key | **DooPlex only**, no passphrase, in no backup | **YES** — then every box needs on-site re-enrollment to accept a new key (`04` §4) |
|
||||
| **S7** | **ep0 read-only restore token** `dooplex-hub@pbs!restore` | reading both copies without ep0 root | DooPlex only (`/etc/felhom-hub-backup/token-restore`) | yes — but ep0 root can mint a new one (hub-DB runbook Step 2) |
|
||||
| **S8** | **Hetzner account login + its two-factor** | ep0's console (rescue, root reset), ep0 itself, Storage Boxes, the Hetzner API token | the operator; probably Vaultwarden | **circle**: if only in Vaultwarden, it is behind the copy it is needed to reach — paper or head |
|
||||
| S9 | ep0 root SSH | the PBS tunnel; minting tokens | ep0 `authorized_keys`: DooPlex's key (lost) + one more key, owner not recorded (operator workstation?) | partly — Hetzner console (S8) is the fallback |
|
||||
| S10 | WireGuard operator peer (`10.77.0.250`) | reaching ep0's PBS (`10.77.0.1:8007`) without SSH | the operator's workstation | no, if the workstation survives |
|
||||
| S11 | **Cloudflare account login** (+2FA) | DNS of `felhom.eu` and `dooplex.hu` (both on Cloudflare); pointing `hub.felhom.eu` at the new home; households' tunnels | the operator; Vaultwarden | circle unless in head/paper |
|
||||
| S12 | **Gmail `felhom.eu@gmail.com`** (+2FA) | the `@felhom.eu` catch-all inbox; account recovery for other services | the operator | — |
|
||||
| S13 | Domain registrar(s) for `felhom.eu`, `dooplex.hu`; No-IP (`dooplex.hopto.org`, the home address CNAME) | moving DNS / the home address | the operator; Vaultwarden | — |
|
||||
| S14 | Resend account / API key | all mail | Secrets export (S1+S5); Vaultwarden; the Resend dashboard (re-mint) | no — re-mint at Resend |
|
||||
| S15 | Gitea admin password; registry credentials (`gitea-creds`) | Gitea UI; image push/pull | the restored Gitea DB (password hash) + Vaultwarden; the Secrets export | no, after S1 (+S3) |
|
||||
| S16 | The hub operator password (`HUB_PW`) | the hub UI | the restored hub DB (hash); `~/.config/credentials` on DooPlex; Vaultwarden | no, after S2 (+S3) |
|
||||
| S17 | Authentik admin | SSO for Vaultwarden and DooPlex apps | Authentik's DB (Longhorn on DooPlex, local PostgreSQL dumps only) | **yes** — not in any copy; Vaultwarden does not need it (password login) |
|
||||
| S18 | Households' recovery codes, PBS keys, restic passwords | households' restores | the households / escrow / the boxes | not the operator's |
|
||||
|
||||
## The order of steps (new hardware)
|
||||
|
||||
1. **A machine and the sheet.** Any Linux machine with Docker; for the hub also k3s. Install `proxmox-backup-client`.
|
||||
2. **Reach ep0's PBS.** Either the operator's WireGuard peer (S10) → `10.77.0.1:8007`; or ep0 root SSH (S9) and the
|
||||
tunnel `ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@167.233.158.164`; or the Hetzner console (S8) → root → add a
|
||||
new SSH key. Fingerprint pin: `proxmox-backup-manager cert info` on ep0.
|
||||
3. **A read token.** S7 from the sheet, or mint a new one on ep0 (hub-DB runbook Step 2: `token/restore`, ACL
|
||||
`DatastoreReader` on `/datastore/felhom-offsite/operator`).
|
||||
4. **Restore `host/dooplex-gitea`** with S1 (key file rebuilt from the `data` field — hub-DB runbook Step 0 note).
|
||||
5. **Vaultwarden first** (it holds the rest): `docker run -v <out>/vaultwarden:/data vaultwarden/server:<version>`
|
||||
(`DOMAIN`, `SIGNUPS_ALLOWED=false`); log in with the e-mail and S3. Or, if a Bitwarden client still holds an offline
|
||||
copy, use that. Every later login comes from here.
|
||||
6. **The Secrets export:** `gpg --batch --passphrase-file <S5 file> -d secrets-<stamp>.yaml.gpg > secrets.yaml` —
|
||||
every k8s Secret of the old cluster (shred it after use).
|
||||
7. **Gitea:** `gitea-restore.md` (on a normal network this time). Then rebuild the images from the code into the new
|
||||
registry: hub, controller, agent (`RUNBOOK-manual-build.md`).
|
||||
8. **The hub:** k3s, then `RUNBOOK-hub-db-offsite-backup.md` §3 with S2 and S4 — **start it with mail held**
|
||||
(`MAIL-HOLD`, from the next hub release; until then: no network until the pending notices are understood).
|
||||
9. **DNS:** in Cloudflare (S11) point `hub.felhom.eu` (today a CNAME to `dooplex.hopto.org`) and `gitea.dooplex.hu` at
|
||||
the new place. The boxes reconnect by themselves: their API keys are in the restored hub DB.
|
||||
10. **Signing:** with S6 on paper, put the keys back and sign as before. Without it, no box accepts an agent update or a
|
||||
bundle until it is re-enrolled on site (`04` §4).
|
||||
11. **Backups again:** reinstall `scripts/hub-db-backup/` and `scripts/dooplex-offsite/` from the restored code, the
|
||||
tunnel and the tokens, and check the next night's copies.
|
||||
|
||||
## Walked on paper (2026-10-09): what still needs DooPlex if the operator has only the sheet, memory and a new machine
|
||||
|
||||
| Step | Needs | From the sheet / memory? |
|
||||
|---|---|---|
|
||||
| reach ep0 | WireGuard peer **or** ep0 root **or** Hetzner login | WireGuard peer is on the workstation, **not** on the sheet; ep0 root key: one of two keys is DooPlex's; **Hetzner login (S8) must be on the sheet** |
|
||||
| PBS fingerprint | ep0's cert fingerprint | not secret; on the sheet as a line, or read on ep0 |
|
||||
| a read token | S7 or ep0 root | on the sheet (S7), or via S8 |
|
||||
| open the Gitea/Vaultwarden copy | S1 | **sheet** |
|
||||
| open the vault | S3 | memory |
|
||||
| open the Secrets export | S5 | **sheet** (today: "offline copy", form unknown) |
|
||||
| the hub's seal key | S4 (or S1+S5) | sheet, or via S5 |
|
||||
| the hub DB | S2 (or Vaultwarden) | sheet, or via S3 |
|
||||
| sign a box update | S6 | **sheet — today exists on DooPlex only** |
|
||||
| hub image | the registry (lost) | rebuild from restored code — needs a Go/Docker build machine, no secret |
|
||||
| Gitea `main` → `felhom.eu` website, ISO host | Cloudflare (S11), the build tree (lost) | S11 on the sheet; the build tree is rebuilt from code |
|
||||
| Authentik (S17) | not in any copy | not needed for the business; DooPlex homelab apps that use it are out of scope |
|
||||
|
||||
**Result:** with a filled sheet (S1, S2, S4, S5, S6, S7, S8, S11 + the ep0 fingerprint) and the master password in
|
||||
memory, nothing in the Felhom recovery still needs DooPlex. **Without the sheet, three things block:** S1 (no copy
|
||||
opens), S6 (no box can be signed for), S8 (ep0 cannot be reached once DooPlex's key is gone, unless the workstation's
|
||||
WireGuard peer survives).
|
||||
Reference in New Issue
Block a user