From 53d8131b2082ed07cd3fdaded2ac0d2361f7588c Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 5 Oct 2026 16:25:48 +0200 Subject: [PATCH] R-886 opened (336); alarm proven end to end; REPORT-hub-db-offsite-2026-10-05 Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 2 +- REPORT-hub-db-offsite-2026-10-05.md | 130 ++++++++++++++++++ STATUS.md | 10 +- .../partD/alarm-drill.txt | 10 ++ documentation/backlog/OPEN-ITEMS.md | 3 +- 5 files changed, 151 insertions(+), 4 deletions(-) create mode 100644 REPORT-hub-db-offsite-2026-10-05.md diff --git a/CONTEXT.md b/CONTEXT.md index ba76795d..988a9a22 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -25,7 +25,7 @@ > (03:45, keep-daily 14, keep-weekly 8, max-depth 0). Success-only textfile metrics → `HubDBBackupStale` (26 h, critical) > and `HubDBRestoreTestStale` (8 d) in homelab-manifests `backup-freshness`. `hub/cmd/hubdb-check` opens a restored copy > with the seal key (runbook §3). Hub PVC 2 Gi + label `enabled`. Collateral: Longhorn instance-manager restart (R-882), -> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Register 332 → 335. +> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Alarm proven end to end (fired 14:22Z, Alertmanager email sent, 0 failed; cleared by a real push); Alertmanager nflog/silences unwritable since the restart (R-886). Register 332 → 336. > Report: `REPORT-hub-db-offsite-2026-10-05.md`. > **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller diff --git a/REPORT-hub-db-offsite-2026-10-05.md b/REPORT-hub-db-offsite-2026-10-05.md new file mode 100644 index 00000000..5ec96801 --- /dev/null +++ b/REPORT-hub-db-offsite-2026-10-05.md @@ -0,0 +1,130 @@ +# REPORT — the hub database off DooPlex (R-173 option A) and the torn-backup live test (R-519) — 2026-10-05 (evening) + +| Part | Result | Releases / changes | +|---|---|---| +| **A** — PVC label + nightly `VACUUM INTO` snapshot | **done** — label `enabled`, volume 2 Gi, snapshot 353 MiB in 44 s, keep 2, mutex; 6 red-proofs | hub **v0.136.0** (one release) | +| **B** — ep0: namespace, two tokens, prune job | **done** — `operator` ns, `dooplex-hub@pbs` with `!push` (DatastoreBackup) and `!restore` (DatastoreReader), `prune-operator-hubdb` (ns `operator` only); household jobs unchanged | ep0 config only | +| **C** — DooPlex push / restore-test units, key, alarm | **done** — versioned scripts + units, 15 tests, 9 red-proofs; `enc.key` made, operator saved the paper key; two alarm rules, `promtool`-proven, 2 red-proofs | `scripts/hub-db-backup/`, homelab-manifests rules | +| **D** — live proof | **done** — push, ep0 listing, restore test, token limits, paper-key restore, alarm fired + mailed + cleared; R-519 live on 9202; runbook §3 tested (steps 1–3) | `hub/cmd/hubdb-check` (tool, not in the image) | + +**Side events (all with the operator's word):** the Longhorn instance-manager on DooPlex was restarted (77/77 volumes +healthy in 110 s); my failed offline-grow attempt kept the hub down ~9.5 min (12:53–13:02Z); zipline came back on a +newer release and was pinned to 4.7.0. + +## Wrong claims (in the brief and the runbook) + +1. **Runbook §1: "the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any + sync."** The Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did not + copy it down (`partA/step1-labels.txt`, before-state). Which one Longhorn reads was not measured; both say `enabled` now. +2. **Brief + runbook: "keep 2 snapshots" on the 1 Gi volume.** Two 353 MiB snapshots plus the 370 MB live database did + not fit (594 MB free). The operator chose to grow to 2 Gi. +3. **Runbook Step 3 implied the snapshot is quick (0.63 s measured on a scratch copy).** On the live Longhorn volume: 44 s. +4. **Runbook Step 2: `--schedule 'daily 03:45'`** — not a PBS calendar event (`unable to parse … 'daily'`); `03:45` is. +5. **Runbook Step 2: `proxmox-backup-client namespace create … root@pam`** needs a password CC does not have; + `proxmox-backup-debug api create …/namespace` works (the CLI then crashes printing the result — the namespace exists). + **`user generate-token … --output-format json`** is refused; the API path is `…/access/users//token/`. +6. **Brief: the push token "DatastoreBackup only (no prune, no read)".** No prune: TRUE (`missing Datastore.Modify| + Datastore.Prune`). **No read: FALSE** — the push token restored its own copy (`restore complete … 352.676 MiB`): PBS + lets a backup's owner read it. Scope: ns `operator` only, and the content is encrypted with a key ep0 never sees. +7. **Runbook Step 2: a token's ACL alone grants it** — PBS cuts a token's rights down by its user's, so the user needs + both roles on the path too (done; the user has no password). +8. **The 1 Gi → 2 Gi growth was assumed routine** (`allowVolumeExpansion=true`). It failed: the instance-manager called a + vanished host PID (`nsenter: cannot open /host/proc/196610/ns/mnt`), and offline growth was blocked by the expansion's + own attachment ticket (`partA/step1-expansion-failure.txt`, `step1-offline-expansion.txt`). +9. **Runbook §3 as proposed: "a working reveal proves the key matches"** — it needs a running hub on the copy. A check + that needs no hub now exists (`hubdb-check`), and the procedure says how to rebuild the key from the paper `data` field. +10. **My own: "the hub is down ~2 minutes" for the offline grow** — it was ~9.5 minutes, and it did not work. + +## Part A — evidence (`documentation/audits/hub-db-offsite-2026-10-05/partA/`) + +- Architecture read first: `05-hub-architecture.md` §16 (new §16.3), `07-backup-architecture.md`, the runbook. +- Hub baseline `cea8502f`-era main at v0.135.0; commits `d4be9f6` (code, docs, manifest), `5d8922e` (image tag 0.136.0). +- `internal/store/snapshot.go` (`SnapshotInto` = `VACUUM INTO ?`), `internal/dbsnap` (`.tmp` → rename, 0600, keep 2, + `ErrBusy`, `NeedsCatchUp`), `cmd/hub/main.go` (02:00 Budapest + start-up catch-up). +- Tests: `dbsnap_test.go` — integrity_check `ok` and equal row counts in EVERY table vs. the live DB with rows still in + the WAL (precondition asserted: a plain file copy held 0 of 150 hosts); keep 2; never two at once; a failed write + leaves no file; catch-up age. `cmd/hub/r173_wiring_test.go` (AST). Red-proofs R1–R6 convict (`red-proof.txt`, + `red-proof-run1.txt`; R2's first run convicted by HANGING — the test now fails in 2 s). +- Live: `db snapshot written: hub-20261005T123037Z.db (369807360 bytes, 44.276s)`, mode `-rw-------`. +- PVC/Volume labels both `enabled`; PVC capacity 2Gi; `/data` 2,028,392 KB, 37 % used (`step1-im-restart.txt`). +- Full hub suite green (`go build/vet/test ./...`). + +## Part B — evidence (`partB/`) + +- Before (`ep0-before.txt`): prune jobs `prune-demo-hp`, `prune-demo-felhom` (ns each, keep-last 2, 03:30); GC + `sun 04:30`; users `root@pam`, `felhom@pbs`; no `operator` ns. +- After (`ep0-after.txt`): the same two jobs unchanged + `prune-operator-hubdb` (ns `operator`, max-depth 0, 03:45, + keep-daily 14, keep-weekly 8); GC unchanged; user `dooplex-hub@pbs`; tokens `!push`, `!restore`; ACLs only on + `/datastore/felhom-offsite/operator`. +- Token secrets: ep0 → pipe → root 0600 files on DooPlex (36 bytes each), never on a command line or printed. No + household namespace was listed or read (the `operator` check read only whether that one name exists). + +## Part C — evidence (`partC/`) + +- Tools present: `sqlite3` 3.46.1, `proxmox-backup-client` 4.2.3, `kubectl` (k3s) as root, textfile dir + `/var/lib/node_exporter/textfile_collector` (already scraped), tunnel `felhom-ep0-pbs-tunnel` active on 127.0.0.1:18007. +- `scripts/hub-db-backup/` (commit `cea8502f`), installed with `install.sh`; config + `/etc/felhom-hub-backup/{env,token-push,token-restore,enc.key}` all root 0600. Timers enabled after the manual runs: + next push Tue 02:31, next restore test Sun 2026-10-11 04:31. +- `enc.key` created (`kdf none`, fingerprint `b2:19:bf:36:…`); **STOPPED for the paper key; the operator saved the `data` + field** before the first push. +- Tests: 15 (`test_hub_db_backup.py`, fakes for `kubectl` and `proxmox-backup-client`). Red-proofs P1–P9 all convict + (`red-proof.txt`; P4 first did NOT convict — masked by integrity_check — and P5 errored; both tests were strengthened). + **Not in CI** (R-885). +- Alarm: `homelab-manifests` `ebc14b0`; `promtool check rules` (4 rules) and `promtool test rules` green; red: threshold + 260 h → FAILED; `absent()` removed → FAILED (the first attempt at that mutation did not apply — recorded) + (`alarm-rule-test.txt`, `bf_test.yml`). Only `ConfigMap/prometheus-rules` was synced; `POST /-/reload` → 200; rules + listed in `/api/v1/rules`. + +## Part D — evidence (`partD/`) + +- **Push** (`push-1.txt`): unit `Result=success`; `checked: 369807360 bytes, integrity ok, 4 host(s)`; `Encryption key + fingerprint: b2:19:bf:36:…`; 352.676 MiB (17.242 MiB compressed) in 7.19 s; success timestamp written. +- **ep0 listing with the read-only token:** `host/dooplex-hub/2026-10-05T13:50:18Z 352.677 MiB`. +- **Restore test** (`restore-test-1.txt`): `checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable`; + success timestamp written; no restored file left. +- **Token limits + paper key** (`token-limits-and-paperkey.txt`): forget refused for both tokens; restore token cannot + back up; neither can list the datastore root; push token CAN restore its own copy (claim 6); a key file rebuilt from + the `data` field only → restore `rc=0`, integrity `ok`, 4 hosts; without a key → `missing key`. +- **Alarm** (`alarm-drill.txt`): success file hidden 13:51:46Z → pending 13:52:06Z → **firing 14:22Z** → Alertmanager: + 1 active alert, receiver `email-notifications`, `notifications_total{email}` 5 and `failed_total` 0 → a real push + 14:23 (2 s) → inactive 14:23:42Z, email counter 6. **The mail itself was not read**: the Gmail connector reads the + felhom.eu catch-all, and Alertmanager mails the operator's personal address — the operator is asked to confirm. +- **R-519 on 9202** (`r519/`): 9202 raised to controller 0.296.0 by hand (`00-upgrade-9202.txt`); bookstack installed; + run 1 complete (33 s); run 2 cut by `docker restart felhom-controller` at 14:05:59.67Z, 0.8 s after `Volume dump: + bookstack/bookstack_bookstack_config` and before its database volume. After: the run record turned `interrupted`; + the controller restarted bookstack itself; **/backups and /backups/apps both carry `data-interrupted-run`** ("A + legutóbbi mentés (2026-10-05 16:05) megszakadt …"; negative control 0); **the restore point reads 14:04:44Z** — the + run-1 database volume, its oldest part (config 14:05:58, SQL 14:05:54). Run 3 complete → notice gone from both pages, + record clear, the torn `.tar.tmp` replaced. Bookstack removed through the product: 0 containers, 0 volumes, no + folder, no backups. +- **Runbook §3** (`restore-procedure/drill.txt`): restore from ep0 → `hubdb-check` with the seal key from the Secret + (file → file) → `hosts=4 console_passwords_opened=4 failed=0`; a random key → `opened=0 failed=4` FAILED. Steps 4–5 + (into a live PVC) not run. `hubdb-check` red-proof R7 convicts. + +## Rows and register + +- **Closed:** R-519 (live-proven) → `CLOSED-ITEMS.md` in the same commit. +- **Narrowed:** R-173 (option A in force; left: runbook §3 steps 4–5), R-232 ((b) and (a) partly, for the hub DB). +- **R-231 not touched** (it is about `/opt/backup/scripts/`); the new units are versioned from day one. +- **Opened:** R-882 (Longhorn stale host PID), R-883 (8 workloads on a moving tag; zipline pinned), R-884 (`monitoring` + Prometheus Deployment OutOfSync), R-885 (scripts' Python tests not in CI), R-886 (Alertmanager cannot write + nflog/silences since the restart). +- **Register: 332 → 336.** `STATUS.md` updated (Tonight section; the two old "needs you" items marked decided/done). + +## Teardown — three layers + +- **Machines:** 9202 — bookstack removed through the product; its controller stays 0.296.0 (was 0.295.0); helper files + in the guest removed. DooPlex — scratch restores shredded; the `hubdb-check` binary shredded; the units, timers, + config and keys stay (they are the deliverable). +- **Hosts:** ep0 — the new user, tokens, namespace, prune job stay (the deliverable); nothing else changed. demo-hp — none. +- **Hub:** none provisioned. The hub runs v0.136.0. +- **Secrets:** a scan of 7,132 committed/working files against the 8 real secret values of this session → 0 hits. Scratch + copies (demo password, bookstack deploy values, the zipline pre-change dump, rule files, used signed envelopes) + shredded. + +## Unproven / say-so + +- `unproven.py --summary`: walked 20, partial 17, built 14, missing 4 — NOT WALKED 35 of 55; no number moved (the hub-DB capability is a new row in `00` §G, not one of the 55 walked claims). +- The `HubDBBackupStale` mail in the operator's inbox — seen only as Alertmanager's send counter. +- Runbook §3 steps 4–5. diff --git a/STATUS.md b/STATUS.md index d3384314..c50798f2 100644 --- a/STATUS.md +++ b/STATUS.md @@ -33,8 +33,14 @@ by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, an refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on DooPlex use "always the newest" (new row). -**Register:** 332 → 335 rows (1 closed: the cut-backup check; 4 opened: the Longhorn fault, the "always newest" apps, a -monitoring sync drift, script tests not in CI). +- **The alarm mail:** I hid the "copy sent" signal on purpose; after 30 minutes the alarm fired (14:22) and the mail + system sent it without error; a real send cleared it a minute later. **Please check your inbox** for + "HubDBBackupStale" around 16:22 — I cannot read your personal mailbox. +- **The mail system (Alertmanager) cannot save its own notes since the Longhorn restart.** Mail still goes out; but a + silence you set would be lost at its next restart (new row). + +**Register:** 332 → 336 rows (1 closed: the cut-backup check; 5 opened: the Longhorn fault, the "always newest" apps, a +monitoring sync drift, script tests not in CI, the mail system's notes). **Needs you:** 1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops. diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt index d9874a2b..cebf9ab6 100644 --- a/documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt @@ -3,3 +3,13 @@ backup_freshness.prom fan_metrics.prom felhom_hub_db_restore.prom node_housekeeping.prom +## 2026-10-05T14:23:16Z HubDBBackupStale FIRING (activeAt 13:52:06Z; fired 14:22); Alertmanager: 1 active alert, receiver email-notifications (to: the operator's address); email notifications_total=5 failed_total=0 since the pod started 13:20:48Z +## recovery push: unit Result=success +felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 112 min old +felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s) +felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 2 s +felhom-hub-db-backup: success signal written +felhom_hub_db_backup_last_success_timestamp_seconds 1791210207 +felhom_hub_db_backup_last_success_bytes 369807360 +## 2026-10-05T14:23:42Z HubDBBackupStale state=inactive +email notifications_total now: 6 diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index a71e3e4e..07dbcd69 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -325,7 +325,7 @@ stopping line that lies. | **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | -## Monitoring & notifications — 27 rows (P2 3, P3 16, P4 8) +## Monitoring & notifications — 28 rows (P2 3, P3 17, P4 8) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -355,6 +355,7 @@ stopping line that lies. | **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC | | **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator | | **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC | +| **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | | **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | ## Hub & operator — 25 rows (P2 1, P3 9, P4 15)