R-173 option A in force (hub DB nightly to ep0, restore-tested, alarmed); R-519 proven live on 9202 and closed; R-173/R-232 narrowed; R-882..R-885 opened (332 -> 335); runbook §3 tested; hubdb-check
gates / gates (push) Successful in 59s
gates / gates (push) Successful in 59s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -3,9 +3,44 @@
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
|
||||
sent to it.**
|
||||
|
||||
**Updated 2026-10-05 (late afternoon, the hub-safety session): every box of ours healthy. The hub refuses forged form
|
||||
posts and keeps the console passwords locked; the System page shows boxes left behind; the agent can no longer make
|
||||
itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
|
||||
**Updated 2026-10-05 (evening, the hub-database session): every box of ours healthy. The hub database now leaves
|
||||
DooPlex every night, locked, to ep0, is test-restored every Sunday, and an alarm mails you if either stops. Report:
|
||||
`REPORT-hub-db-offsite-2026-10-05.md`.**
|
||||
|
||||
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box
|
||||
|
||||
**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three
|
||||
by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager.
|
||||
|
||||
**What works now (proven live):**
|
||||
- **The hub makes a clean copy of its database every night at 02:00** (hub 0.136.0). The first one: 353 MB, 44 s.
|
||||
- **DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30.** First send: 7 s.
|
||||
- **Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it**: it opens, it is whole, every console
|
||||
password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable.
|
||||
- **The two ep0 accounts can only do their one job:** the sending one cannot delete, the checking one cannot write,
|
||||
neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.)
|
||||
- **Your two saved keys work:** a key rebuilt from the paper copy you saved opened the copy, and your saved lock key
|
||||
opened all 4 console passwords in it (a wrong key opened none).
|
||||
- **The alarm:** proven by Prometheus' own rule test; the real alarm mail is below.
|
||||
- **A backup cut off by a restart is now said on the backup pages** (scratch box): the page said so, the restore point
|
||||
kept the older time of the part that was not redone, and the next full backup cleared the message.
|
||||
|
||||
**What broke, and what I did:**
|
||||
- **The hub's disk would not grow:** Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the
|
||||
disk could grow offline) did not work and kept the hub **down about 9.5 minutes**. You approved restarting the disk
|
||||
manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now.
|
||||
- **Zipline did not come back after that restart:** it is set to "always the newest", so it pulled a new release that
|
||||
refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on
|
||||
DooPlex use "always the newest" (new row).
|
||||
|
||||
**Register:** 332 → 335 rows (1 closed: the cut-backup check; 4 opened: the Longhorn fault, the "always newest" apps, a
|
||||
monitoring sync drift, script tests not in CI).
|
||||
|
||||
**Needs you:**
|
||||
1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops.
|
||||
2. **When convenient:** pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If
|
||||
you do nothing, any restart may upgrade one of them by surprise, as it did zipline.
|
||||
3. **Tester 2's one-time step** is unchanged (below).
|
||||
|
||||
## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights
|
||||
|
||||
@@ -38,7 +73,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
|
||||
minutes. Both figures are on the page now.
|
||||
|
||||
**Needs you:**
|
||||
1. **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
|
||||
1. **(DECIDED 2026-10-05 evening: A, done — see Tonight)** **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
|
||||
- **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an
|
||||
alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
|
||||
- **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A.
|
||||
@@ -47,7 +82,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
|
||||
`documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
|
||||
- **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy
|
||||
of the database cannot open the console passwords.
|
||||
2. **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
|
||||
2. **(DONE 2026-10-05 evening — see Tonight)** **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
|
||||
controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays
|
||||
proven by tests only.
|
||||
3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own
|
||||
|
||||
Reference in New Issue
Block a user