R-232/R-173: hub restored into a throwaway k3s (runbook §3 steps 4–5 proven, corrected); R-173, R-861, R-518 closed; R-921, R-922 filed; STATUS, capability map, report
gates / gates (push) Successful in 5m23s
gates / gates (push) Successful in 5m23s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -26,6 +26,16 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-09 — can the business survive losing DooPlex: Gitea + secrets off-site, Gitea and the hub restored into throwaways
|
||||
|
||||
The full text of every row below: `git show 59f1ad2b86:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** (P2) | CLOSED 2026-10-09 — the last steps of the hub's recovery runbook (§3 steps 4–5) proven on a throwaway single-node k3s (the bench, an INTERNAL Docker network, no route out): the newest ep0 copy restored with the read-only token, the live hub image (same digest) deployed, scaled to 0, `hub.db` copied into the PVC by a helper pod (-wal/-shm removed), scaled to 1 → `console passwords sealed at rest (0 legacy …)`; customers 4/4 and hosts 4/4 equal live; 4/4 console passwords revealed (length only) and, second channel, `hubdb-check` 4/4 with the saved key, 0/4 with a random one. Found: a restored hub at once mails customers their pending notices (operator mail off does NOT stop them) — runbook §3 corrected. Deleted: the k3s container, its volumes, every copy (shredded). | `audits/dooplex-survival-2026-10-09/partD-*.txt`, `runbooks/RUNBOOK-hub-db-offsite-backup.md` §3 |
|
||||
| **R-861** | **The agent's sudoers lets the agent user reach root without the operator key.** (P2) | CLOSED 2026-10-09 — (a) proven: all three controller swaps of 0.304.0 (demo-hp, demo-felhom, Tester 1, ~05:05Z) ran `felhom-priv-apply controller-image 9201` (sudo log + the wrapper's `WROTE … 0.304.0` + the agent's "new controller healthy"), and the guests run 0.304.0 (Docker); the `tee` route 0 times that day. (b) B2 delivered + B3 accepted, (c) C2 accepted — `09` §3 decision 165. | `audits/dooplex-survival-2026-10-09/partE/R-861.txt` |
|
||||
| **R-518** | **„Mentés most" stopped every app for ~8 min.** (P2) | CLOSED 2026-10-09 — the two-tier night under one-stop-per-tier read back on demo-hp (2026-10-08): local tier 20:20–20:25Z, off-site tier 20:34–20:37Z, each with its own stop under ~2 min (agent journal + Proxmox task log; controller metrics.db; ep0 listing `2026-10-08T20:34:27Z`, 7-day cadence held). A third, useless stop when the off-site tier answered BUSY → new row **R-921**. | `audits/dooplex-survival-2026-10-09/partE/R-518.txt` |
|
||||
|
||||
## 2026-10-09 — the release day: D1–D4 proven live, the ep0-copy job installed
|
||||
|
||||
The full text of every row below: `git show b9073e8fb6:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user