R-886 opened (336); alarm proven end to end; REPORT-hub-db-offsite-2026-10-05
gates / gates (push) Successful in 57s
gates / gates (push) Successful in 57s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+1
-1
@@ -25,7 +25,7 @@
|
||||
> (03:45, keep-daily 14, keep-weekly 8, max-depth 0). Success-only textfile metrics → `HubDBBackupStale` (26 h, critical)
|
||||
> and `HubDBRestoreTestStale` (8 d) in homelab-manifests `backup-freshness`. `hub/cmd/hubdb-check` opens a restored copy
|
||||
> with the seal key (runbook §3). Hub PVC 2 Gi + label `enabled`. Collateral: Longhorn instance-manager restart (R-882),
|
||||
> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Register 332 → 335.
|
||||
> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Alarm proven end to end (fired 14:22Z, Alertmanager email sent, 0 failed; cleared by a real push); Alertmanager nflog/silences unwritable since the restart (R-886). Register 332 → 336.
|
||||
> Report: `REPORT-hub-db-offsite-2026-10-05.md`.
|
||||
|
||||
> **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller
|
||||
|
||||
@@ -0,0 +1,130 @@
|
||||
# REPORT — the hub database off DooPlex (R-173 option A) and the torn-backup live test (R-519) — 2026-10-05 (evening)
|
||||
|
||||
| Part | Result | Releases / changes |
|
||||
|---|---|---|
|
||||
| **A** — PVC label + nightly `VACUUM INTO` snapshot | **done** — label `enabled`, volume 2 Gi, snapshot 353 MiB in 44 s, keep 2, mutex; 6 red-proofs | hub **v0.136.0** (one release) |
|
||||
| **B** — ep0: namespace, two tokens, prune job | **done** — `operator` ns, `dooplex-hub@pbs` with `!push` (DatastoreBackup) and `!restore` (DatastoreReader), `prune-operator-hubdb` (ns `operator` only); household jobs unchanged | ep0 config only |
|
||||
| **C** — DooPlex push / restore-test units, key, alarm | **done** — versioned scripts + units, 15 tests, 9 red-proofs; `enc.key` made, operator saved the paper key; two alarm rules, `promtool`-proven, 2 red-proofs | `scripts/hub-db-backup/`, homelab-manifests rules |
|
||||
| **D** — live proof | **done** — push, ep0 listing, restore test, token limits, paper-key restore, alarm fired + mailed + cleared; R-519 live on 9202; runbook §3 tested (steps 1–3) | `hub/cmd/hubdb-check` (tool, not in the image) |
|
||||
|
||||
**Side events (all with the operator's word):** the Longhorn instance-manager on DooPlex was restarted (77/77 volumes
|
||||
healthy in 110 s); my failed offline-grow attempt kept the hub down ~9.5 min (12:53–13:02Z); zipline came back on a
|
||||
newer release and was pinned to 4.7.0.
|
||||
|
||||
## Wrong claims (in the brief and the runbook)
|
||||
|
||||
1. **Runbook §1: "the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any
|
||||
sync."** The Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did not
|
||||
copy it down (`partA/step1-labels.txt`, before-state). Which one Longhorn reads was not measured; both say `enabled` now.
|
||||
2. **Brief + runbook: "keep 2 snapshots" on the 1 Gi volume.** Two 353 MiB snapshots plus the 370 MB live database did
|
||||
not fit (594 MB free). The operator chose to grow to 2 Gi.
|
||||
3. **Runbook Step 3 implied the snapshot is quick (0.63 s measured on a scratch copy).** On the live Longhorn volume: 44 s.
|
||||
4. **Runbook Step 2: `--schedule 'daily 03:45'`** — not a PBS calendar event (`unable to parse … 'daily'`); `03:45` is.
|
||||
5. **Runbook Step 2: `proxmox-backup-client namespace create … root@pam`** needs a password CC does not have;
|
||||
`proxmox-backup-debug api create …/namespace` works (the CLI then crashes printing the result — the namespace exists).
|
||||
**`user generate-token … --output-format json`** is refused; the API path is `…/access/users/<u>/token/<name>`.
|
||||
6. **Brief: the push token "DatastoreBackup only (no prune, no read)".** No prune: TRUE (`missing Datastore.Modify|
|
||||
Datastore.Prune`). **No read: FALSE** — the push token restored its own copy (`restore complete … 352.676 MiB`): PBS
|
||||
lets a backup's owner read it. Scope: ns `operator` only, and the content is encrypted with a key ep0 never sees.
|
||||
7. **Runbook Step 2: a token's ACL alone grants it** — PBS cuts a token's rights down by its user's, so the user needs
|
||||
both roles on the path too (done; the user has no password).
|
||||
8. **The 1 Gi → 2 Gi growth was assumed routine** (`allowVolumeExpansion=true`). It failed: the instance-manager called a
|
||||
vanished host PID (`nsenter: cannot open /host/proc/196610/ns/mnt`), and offline growth was blocked by the expansion's
|
||||
own attachment ticket (`partA/step1-expansion-failure.txt`, `step1-offline-expansion.txt`).
|
||||
9. **Runbook §3 as proposed: "a working reveal proves the key matches"** — it needs a running hub on the copy. A check
|
||||
that needs no hub now exists (`hubdb-check`), and the procedure says how to rebuild the key from the paper `data` field.
|
||||
10. **My own: "the hub is down ~2 minutes" for the offline grow** — it was ~9.5 minutes, and it did not work.
|
||||
|
||||
## Part A — evidence (`documentation/audits/hub-db-offsite-2026-10-05/partA/`)
|
||||
|
||||
- Architecture read first: `05-hub-architecture.md` §16 (new §16.3), `07-backup-architecture.md`, the runbook.
|
||||
- Hub baseline `cea8502f`-era main at v0.135.0; commits `d4be9f6` (code, docs, manifest), `5d8922e` (image tag 0.136.0).
|
||||
- `internal/store/snapshot.go` (`SnapshotInto` = `VACUUM INTO ?`), `internal/dbsnap` (`.tmp` → rename, 0600, keep 2,
|
||||
`ErrBusy`, `NeedsCatchUp`), `cmd/hub/main.go` (02:00 Budapest + start-up catch-up).
|
||||
- Tests: `dbsnap_test.go` — integrity_check `ok` and equal row counts in EVERY table vs. the live DB with rows still in
|
||||
the WAL (precondition asserted: a plain file copy held 0 of 150 hosts); keep 2; never two at once; a failed write
|
||||
leaves no file; catch-up age. `cmd/hub/r173_wiring_test.go` (AST). Red-proofs R1–R6 convict (`red-proof.txt`,
|
||||
`red-proof-run1.txt`; R2's first run convicted by HANGING — the test now fails in 2 s).
|
||||
- Live: `db snapshot written: hub-20261005T123037Z.db (369807360 bytes, 44.276s)`, mode `-rw-------`.
|
||||
- PVC/Volume labels both `enabled`; PVC capacity 2Gi; `/data` 2,028,392 KB, 37 % used (`step1-im-restart.txt`).
|
||||
- Full hub suite green (`go build/vet/test ./...`).
|
||||
|
||||
## Part B — evidence (`partB/`)
|
||||
|
||||
- Before (`ep0-before.txt`): prune jobs `prune-demo-hp`, `prune-demo-felhom` (ns each, keep-last 2, 03:30); GC
|
||||
`sun 04:30`; users `root@pam`, `felhom@pbs`; no `operator` ns.
|
||||
- After (`ep0-after.txt`): the same two jobs unchanged + `prune-operator-hubdb` (ns `operator`, max-depth 0, 03:45,
|
||||
keep-daily 14, keep-weekly 8); GC unchanged; user `dooplex-hub@pbs`; tokens `!push`, `!restore`; ACLs only on
|
||||
`/datastore/felhom-offsite/operator`.
|
||||
- Token secrets: ep0 → pipe → root 0600 files on DooPlex (36 bytes each), never on a command line or printed. No
|
||||
household namespace was listed or read (the `operator` check read only whether that one name exists).
|
||||
|
||||
## Part C — evidence (`partC/`)
|
||||
|
||||
- Tools present: `sqlite3` 3.46.1, `proxmox-backup-client` 4.2.3, `kubectl` (k3s) as root, textfile dir
|
||||
`/var/lib/node_exporter/textfile_collector` (already scraped), tunnel `felhom-ep0-pbs-tunnel` active on 127.0.0.1:18007.
|
||||
- `scripts/hub-db-backup/` (commit `cea8502f`), installed with `install.sh`; config
|
||||
`/etc/felhom-hub-backup/{env,token-push,token-restore,enc.key}` all root 0600. Timers enabled after the manual runs:
|
||||
next push Tue 02:31, next restore test Sun 2026-10-11 04:31.
|
||||
- `enc.key` created (`kdf none`, fingerprint `b2:19:bf:36:…`); **STOPPED for the paper key; the operator saved the `data`
|
||||
field** before the first push.
|
||||
- Tests: 15 (`test_hub_db_backup.py`, fakes for `kubectl` and `proxmox-backup-client`). Red-proofs P1–P9 all convict
|
||||
(`red-proof.txt`; P4 first did NOT convict — masked by integrity_check — and P5 errored; both tests were strengthened).
|
||||
**Not in CI** (R-885).
|
||||
- Alarm: `homelab-manifests` `ebc14b0`; `promtool check rules` (4 rules) and `promtool test rules` green; red: threshold
|
||||
260 h → FAILED; `absent()` removed → FAILED (the first attempt at that mutation did not apply — recorded)
|
||||
(`alarm-rule-test.txt`, `bf_test.yml`). Only `ConfigMap/prometheus-rules` was synced; `POST /-/reload` → 200; rules
|
||||
listed in `/api/v1/rules`.
|
||||
|
||||
## Part D — evidence (`partD/`)
|
||||
|
||||
- **Push** (`push-1.txt`): unit `Result=success`; `checked: 369807360 bytes, integrity ok, 4 host(s)`; `Encryption key
|
||||
fingerprint: b2:19:bf:36:…`; 352.676 MiB (17.242 MiB compressed) in 7.19 s; success timestamp written.
|
||||
- **ep0 listing with the read-only token:** `host/dooplex-hub/2026-10-05T13:50:18Z 352.677 MiB`.
|
||||
- **Restore test** (`restore-test-1.txt`): `checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable`;
|
||||
success timestamp written; no restored file left.
|
||||
- **Token limits + paper key** (`token-limits-and-paperkey.txt`): forget refused for both tokens; restore token cannot
|
||||
back up; neither can list the datastore root; push token CAN restore its own copy (claim 6); a key file rebuilt from
|
||||
the `data` field only → restore `rc=0`, integrity `ok`, 4 hosts; without a key → `missing key`.
|
||||
- **Alarm** (`alarm-drill.txt`): success file hidden 13:51:46Z → pending 13:52:06Z → **firing 14:22Z** → Alertmanager:
|
||||
1 active alert, receiver `email-notifications`, `notifications_total{email}` 5 and `failed_total` 0 → a real push
|
||||
14:23 (2 s) → inactive 14:23:42Z, email counter 6. **The mail itself was not read**: the Gmail connector reads the
|
||||
felhom.eu catch-all, and Alertmanager mails the operator's personal address — the operator is asked to confirm.
|
||||
- **R-519 on 9202** (`r519/`): 9202 raised to controller 0.296.0 by hand (`00-upgrade-9202.txt`); bookstack installed;
|
||||
run 1 complete (33 s); run 2 cut by `docker restart felhom-controller` at 14:05:59.67Z, 0.8 s after `Volume dump:
|
||||
bookstack/bookstack_bookstack_config` and before its database volume. After: the run record turned `interrupted`;
|
||||
the controller restarted bookstack itself; **/backups and /backups/apps both carry `data-interrupted-run`** ("A
|
||||
legutóbbi mentés (2026-10-05 16:05) megszakadt …"; negative control 0); **the restore point reads 14:04:44Z** — the
|
||||
run-1 database volume, its oldest part (config 14:05:58, SQL 14:05:54). Run 3 complete → notice gone from both pages,
|
||||
record clear, the torn `.tar.tmp` replaced. Bookstack removed through the product: 0 containers, 0 volumes, no
|
||||
folder, no backups.
|
||||
- **Runbook §3** (`restore-procedure/drill.txt`): restore from ep0 → `hubdb-check` with the seal key from the Secret
|
||||
(file → file) → `hosts=4 console_passwords_opened=4 failed=0`; a random key → `opened=0 failed=4` FAILED. Steps 4–5
|
||||
(into a live PVC) not run. `hubdb-check` red-proof R7 convicts.
|
||||
|
||||
## Rows and register
|
||||
|
||||
- **Closed:** R-519 (live-proven) → `CLOSED-ITEMS.md` in the same commit.
|
||||
- **Narrowed:** R-173 (option A in force; left: runbook §3 steps 4–5), R-232 ((b) and (a) partly, for the hub DB).
|
||||
- **R-231 not touched** (it is about `/opt/backup/scripts/`); the new units are versioned from day one.
|
||||
- **Opened:** R-882 (Longhorn stale host PID), R-883 (8 workloads on a moving tag; zipline pinned), R-884 (`monitoring`
|
||||
Prometheus Deployment OutOfSync), R-885 (scripts' Python tests not in CI), R-886 (Alertmanager cannot write
|
||||
nflog/silences since the restart).
|
||||
- **Register: 332 → 336.** `STATUS.md` updated (Tonight section; the two old "needs you" items marked decided/done).
|
||||
|
||||
## Teardown — three layers
|
||||
|
||||
- **Machines:** 9202 — bookstack removed through the product; its controller stays 0.296.0 (was 0.295.0); helper files
|
||||
in the guest removed. DooPlex — scratch restores shredded; the `hubdb-check` binary shredded; the units, timers,
|
||||
config and keys stay (they are the deliverable).
|
||||
- **Hosts:** ep0 — the new user, tokens, namespace, prune job stay (the deliverable); nothing else changed. demo-hp — none.
|
||||
- **Hub:** none provisioned. The hub runs v0.136.0.
|
||||
- **Secrets:** a scan of 7,132 committed/working files against the 8 real secret values of this session → 0 hits. Scratch
|
||||
copies (demo password, bookstack deploy values, the zipline pre-change dump, rule files, used signed envelopes)
|
||||
shredded.
|
||||
|
||||
## Unproven / say-so
|
||||
|
||||
- `unproven.py --summary`: walked 20, partial 17, built 14, missing 4 — NOT WALKED 35 of 55; no number moved (the hub-DB capability is a new row in `00` §G, not one of the 55 walked claims).
|
||||
- The `HubDBBackupStale` mail in the operator's inbox — seen only as Alertmanager's send counter.
|
||||
- Runbook §3 steps 4–5.
|
||||
@@ -33,8 +33,14 @@ by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, an
|
||||
refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on
|
||||
DooPlex use "always the newest" (new row).
|
||||
|
||||
**Register:** 332 → 335 rows (1 closed: the cut-backup check; 4 opened: the Longhorn fault, the "always newest" apps, a
|
||||
monitoring sync drift, script tests not in CI).
|
||||
- **The alarm mail:** I hid the "copy sent" signal on purpose; after 30 minutes the alarm fired (14:22) and the mail
|
||||
system sent it without error; a real send cleared it a minute later. **Please check your inbox** for
|
||||
"HubDBBackupStale" around 16:22 — I cannot read your personal mailbox.
|
||||
- **The mail system (Alertmanager) cannot save its own notes since the Longhorn restart.** Mail still goes out; but a
|
||||
silence you set would be lost at its next restart (new row).
|
||||
|
||||
**Register:** 332 → 336 rows (1 closed: the cut-backup check; 5 opened: the Longhorn fault, the "always newest" apps, a
|
||||
monitoring sync drift, script tests not in CI, the mail system's notes).
|
||||
|
||||
**Needs you:**
|
||||
1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops.
|
||||
|
||||
@@ -3,3 +3,13 @@ backup_freshness.prom
|
||||
fan_metrics.prom
|
||||
felhom_hub_db_restore.prom
|
||||
node_housekeeping.prom
|
||||
## 2026-10-05T14:23:16Z HubDBBackupStale FIRING (activeAt 13:52:06Z; fired 14:22); Alertmanager: 1 active alert, receiver email-notifications (to: the operator's address); email notifications_total=5 failed_total=0 since the pod started 13:20:48Z
|
||||
## recovery push: unit Result=success
|
||||
felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 112 min old
|
||||
felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s)
|
||||
felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 2 s
|
||||
felhom-hub-db-backup: success signal written
|
||||
felhom_hub_db_backup_last_success_timestamp_seconds 1791210207
|
||||
felhom_hub_db_backup_last_success_bytes 369807360
|
||||
## 2026-10-05T14:23:42Z HubDBBackupStale state=inactive
|
||||
email notifications_total now: 6
|
||||
|
||||
@@ -325,7 +325,7 @@ stopping line that lies.
|
||||
| **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
||||
|
||||
## Monitoring & notifications — 27 rows (P2 3, P3 16, P4 8)
|
||||
## Monitoring & notifications — 28 rows (P2 3, P3 17, P4 8)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -355,6 +355,7 @@ stopping line that lies.
|
||||
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
||||
| **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator |
|
||||
| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC |
|
||||
| **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator |
|
||||
| **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator |
|
||||
|
||||
## Hub & operator — 25 rows (P2 1, P3 9, P4 15)
|
||||
|
||||
Reference in New Issue
Block a user