hub (unreleased): R-922 option A — a household's clear deletes its notification address (email_cleared); MAIL-HOLD — a restored hub sends no mail until released; two log lines drop the address; runbooks: mail hold is restore step 1; 07 §6.4 R-921 pre-check; R-921/R-922 narrowed
gates / gates (push) Successful in 5m25s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-09 12:37:28 +02:00
parent 2c48feb325
commit d55c590c5a
35 changed files with 861 additions and 23 deletions
@@ -751,6 +751,12 @@ the apps' own stop and start (demo-hp 2026-10-05: stop 21 s, start 46 s). **Not
rule (the scratch guest has no agent connection; the demo boxes take only deliveries and read-backs) — the first
night run on the demo boxes is the proof owed (R-518). The page states about 1–1.5 minutes (an estimate from the
parts above, not a measured press).
**Check first, stop second (R-921, built 2026-10-09, ships with the next controller release).** Measured 2026-10-08 on
demo-hp: the off-site tier stopped every app for ~1 min and the agent then refused it (BUSY: the night OS step held the
agent's host-wide lock). Before stopping anything the controller now asks the agent which tiers have a job in flight
(`GET /backup/status`); another tier's job running or snapshotted → no stop, the tier stays due. A refusal after the
stop still resumes the apps at once (pinned by test). **Not covered:** the host-wide busy lock (OS step, restore test,
fstrim) — the agent serves it on no endpoint, so the measured case needs an agent change too (R-921, LEFT).
**Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the
local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the
longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both
File diff suppressed because one or more lines are too long
@@ -113,6 +113,7 @@ A központi rendszer (`hub.felhom.eu`) a Felhom saját szerverén fut (k3s fürt
| Adat | Mire kell | Meddig marad meg |
|---|---|---|
| Ügyfél azonosítója, neve, domainje, e-mail-címe, nyelve | szerződés teljesítése, értesítések | az ügyfél törléséig <!-- source: hub/internal/store/store.go:161-171 (customer_configs: customer_id, customer_name, domain, email); :223 (language); store.go DELETE FROM customer_configs; documentation/architecture/00-capability-map.md:97 (Customer DELETE cascade) --> |
| Értesítési e-mail-cím (amelyre a háztartás a figyelmeztetéseket kéri) | értesítések | amíg a háztartás meg nem változtatja; ha a háztartás a dashboardon kitörli, a központi rendszer a következő szinkronizáláskor törli; az ügyfél törlésekor törlődik <!-- source: hub/internal/api/handler.go handleSavePreferences (R-922, email_cleared → address deleted; operator ruling 2026-10-09 option A, from the next hub + controller release); hub/internal/store/customer_delete.go:60 (customer_notifications purged at delete) --> |
| A szerver állapotjelentései (gépnév, processzor-, memória-, lemezhasználat, hőmérséklet, a telepített alkalmazások neve és állapota, mentések állapota, a dashboard nyelve) | felügyelet, hibajelzés | **90 nap** <!-- source: felhom-controller/controller/internal/report/types.go:13-56, :66-79 (report fields); hub/internal/store/store.go:1528-1540 (Prune: reports + host_reports); manifests/hub.yaml:78 (retention.max_days: 90); hub/cmd/hub/main.go:930 (default 90) --> |
| Események (pl. „mentés sikertelen", „lemez megtelt") | felügyelet, ügyfélnek látható napló | **90 nap** <!-- source: hub/internal/store/store.go:2695-2702 (PruneEvents); manifests/hub.yaml:78 --> |
| Alkalmazásonkénti erőforrás-statisztika | kapacitástervezés | **90 nap** <!-- source: hub/cmd/hub/main.go:1022 (PruneAppTelemetry 90 days) --> |
@@ -158,7 +158,7 @@ Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
start it again.
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05 (steps 1–3) and 2026-10-09 (steps 4–5)
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05 (steps 1–3) and 2026-10-09 (steps 4–5) — *steps renumbered 2026-10-09: the mail hold became step 1, so the old 1–3 are now 2–4 and the old 4–5 are 5–6*
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. **Steps 4–5 were run on
@@ -175,10 +175,17 @@ copy, the customer list and host list equal live (4/4, 4/4), all 4 console passw
What you need, from the break-glass sheet (`break-glass-sheet.md`; the password manager only if it survived or was restored first — `total-loss-of-dooplex.md` step 5): the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
the read-only token (or ep0 root to mint a new one: Step 2).
1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
1. **Mail held — before the copy is ever started (from the next hub release, R-923 follow-up).** Create the empty
marker `MAIL-HOLD` in the hub's data directory (`/data/MAIL-HOLD`, in the PVC, by the helper pod of step 6) BEFORE
the restored hub first starts. While it exists the hub sends no e-mail at all (households and operator), logs each
dropped mail as `[WARN] MAIL-HOLD: not sending <kind> to <who>`, and every operator page shows a banner. Held mails
are dropped, not queued (they were decided from a snapshot). Release it on the Configuration page (or `POST
/configuration/mail-hold/release`) once the restored hub is the only hub and its pending notices are understood.
Until that release is deployed: start a test restore with no network at all.
2. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
to `enc.key` (Step 0 note). A copy of DooPlex's `/etc/felhom-hub-backup/enc.key` works as is.
2. **Restore the newest copy** (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
3. **Restore the newest copy** (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
```bash
export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
R='dooplex-hub@pbs!restore@<ep0>:8007:felhom-offsite'
@@ -187,14 +194,14 @@ the read-only token (or ep0 root to mint a new one: Step 2).
sqlite3 -readonly out/hub.db 'PRAGMA integrity_check' # must print: ok
```
If DooPlex's `ep0-copy` datastore survived, the same copy is there too (pulled nightly).
3. **Prove the seal key matches BEFORE putting the copy in place** — on a COPY of `out/hub.db` (the check migrates it):
4. **Prove the seal key matches BEFORE putting the copy in place** — on a COPY of `out/hub.db` (the check migrates it):
```bash
cd felhom.eu/hub && go build -o hubdb-check ./cmd/hubdb-check
printf '%s' "<OFFSITE_SECRET_KEY>" > k; chmod 600 k # from the password manager — not on a command line in a shared shell
./hubdb-check copy-of-hub.db k # want: hosts=N console_passwords_opened=N failed=0; exit 0
```
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
5. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
**Corrected 2026-10-09:** the Deployment also needs, NOT optional, `Secret/resend-api` (`RESEND_API_KEY`),
`Secret/report-api` (`REPORT_API_KEY`) and `Secret/gitea-creds` (`username`, `password`); without them the pod does
@@ -202,7 +209,7 @@ the read-only token (or ep0 root to mint a new one: Step 2).
off-site (`runbooks/gitea-restore.md`, `secrets/*.gpg`, opened with DooPlex's restic passphrase). A test uses dummies.
The hub's image is pulled from Gitea's registry — on a rebuild with no registry, `docker save` it from any machine
that has it, or build it from the restored code.
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
6. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
As run 2026-10-09: `kubectl scale deploy/hub --replicas=0`; a `busybox` pod mounting `hub-data` at `/data`;
@@ -69,7 +69,7 @@ local restic repos; Prometheus history; DooPlex's other homelab apps (not Felhom
7. **Gitea:** `gitea-restore.md` (on a normal network this time). Then rebuild the images from the code into the new
registry: hub, controller, agent (`RUNBOOK-manual-build.md`).
8. **The hub:** k3s, then `RUNBOOK-hub-db-offsite-backup.md` §3 with S2 and S4 — **start it with mail held**
(`MAIL-HOLD`, from the next hub release; until then: no network until the pending notices are understood).
(§3 step 1: the `MAIL-HOLD` marker, from the next hub release; until then: no network until the pending notices are understood).
9. **DNS:** in Cloudflare (S11) point `hub.felhom.eu` (today a CNAME to `dooplex.hopto.org`) and `gitea.dooplex.hu` at
the new place. The boxes reconnect by themselves: their API keys are in the restored hub DB.
10. **Signing:** with S6 on paper, put the keys back and sign as before. Without it, no box accepts an agent update or a