R-232/R-173: hub restored into a throwaway k3s (runbook §3 steps 4–5 proven, corrected); R-173, R-861, R-518 closed; R-921, R-922 filed; STATUS, capability map, report
gates / gates (push) Successful in 5m23s
gates / gates (push) Successful in 5m23s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -153,11 +153,19 @@ Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot
|
||||
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
|
||||
start it again.
|
||||
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05 (steps 1–3) and 2026-10-09 (steps 4–5)
|
||||
|
||||
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
|
||||
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
|
||||
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
|
||||
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. **Steps 4–5 were run on
|
||||
2026-10-09 into a throwaway single-node k3s** (the bench, LXC 9401, k3s v1.33.6 in a Docker container on an INTERNAL
|
||||
network — `audits/dooplex-survival-2026-10-09/partD-*.txt`): the hub image of the live version started on the restored
|
||||
copy, the customer list and host list equal live (4/4, 4/4), all 4 console passwords revealed (R-173 closed).
|
||||
|
||||
> **⚠ Corrected 2026-10-09 — a restored hub mails households AT ONCE.** Within a minute of starting on the copy it
|
||||
> tried to send two households a pending „kernel notice" (blocked only because the test had no network). Turning
|
||||
> operator mail off (`operator_enabled: false`) does NOT stop household mail. So: **a test restore has no network at
|
||||
> all** (no route to Resend, to ep0, to the boxes). **A real recovery** should expect the notices the copy still holds
|
||||
> to go out once — check the hub's pending notices before giving it a network if that matters.
|
||||
|
||||
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
|
||||
the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
@@ -183,9 +191,19 @@ the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
|
||||
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
|
||||
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
|
||||
**Corrected 2026-10-09:** the Deployment also needs, NOT optional, `Secret/resend-api` (`RESEND_API_KEY`),
|
||||
`Secret/report-api` (`REPORT_API_KEY`) and `Secret/gitea-creds` (`username`, `password`); without them the pod does
|
||||
not start. On a real rebuild their values come from DooPlex's nightly secrets export, which since 2026-10-09 is also
|
||||
off-site (`runbooks/gitea-restore.md`, `secrets/*.gpg`, opened with DooPlex's restic passphrase). A test uses dummies.
|
||||
The hub's image is pulled from Gitea's registry — on a rebuild with no registry, `docker save` it from any machine
|
||||
that has it, or build it from the restored code.
|
||||
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
|
||||
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
|
||||
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
|
||||
As run 2026-10-09: `kubectl scale deploy/hub --replicas=0`; a `busybox` pod mounting `hub-data` at `/data`;
|
||||
`kubectl cp hub.db hubdb-copy:/data/hub.db.new`, then in the pod `rm -f /data/hub.db-wal /data/hub.db-shm && mv
|
||||
/data/hub.db.new /data/hub.db`; delete the pod; scale to 1 (20 s to Ready). The customer list is `GET /configs`
|
||||
(rows link to `/customers/<id>`), the hosts `GET /hosts`.
|
||||
6. Shred `k`, `enc.key` copies and `out/` when done.
|
||||
|
||||
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
|
||||
|
||||
Reference in New Issue
Block a user