R-173 option A in force (hub DB nightly to ep0, restore-tested, alarmed); R-519 proven live on 9202 and closed; R-173/R-232 narrowed; R-882..R-885 opened (332 -> 335); runbook §3 tested; hubdb-check
gates / gates (push) Successful in 59s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 16:14:13 +02:00
parent cea8502f0b
commit afba622fb6
27 changed files with 611 additions and 44 deletions
+40 -5
View File
@@ -3,9 +3,44 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
sent to it.**
**Updated 2026-10-05 (late afternoon, the hub-safety session): every box of ours healthy. The hub refuses forged form
posts and keeps the console passwords locked; the System page shows boxes left behind; the agent can no longer make
itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
**Updated 2026-10-05 (evening, the hub-database session): every box of ours healthy. The hub database now leaves
DooPlex every night, locked, to ep0, is test-restored every Sunday, and an alarm mails you if either stops. Report:
`REPORT-hub-db-offsite-2026-10-05.md`.**
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box
**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three
by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager.
**What works now (proven live):**
- **The hub makes a clean copy of its database every night at 02:00** (hub 0.136.0). The first one: 353 MB, 44 s.
- **DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30.** First send: 7 s.
- **Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it**: it opens, it is whole, every console
password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable.
- **The two ep0 accounts can only do their one job:** the sending one cannot delete, the checking one cannot write,
neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.)
- **Your two saved keys work:** a key rebuilt from the paper copy you saved opened the copy, and your saved lock key
opened all 4 console passwords in it (a wrong key opened none).
- **The alarm:** proven by Prometheus' own rule test; the real alarm mail is below.
- **A backup cut off by a restart is now said on the backup pages** (scratch box): the page said so, the restore point
kept the older time of the part that was not redone, and the next full backup cleared the message.
**What broke, and what I did:**
- **The hub's disk would not grow:** Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the
disk could grow offline) did not work and kept the hub **down about 9.5 minutes**. You approved restarting the disk
manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now.
- **Zipline did not come back after that restart:** it is set to "always the newest", so it pulled a new release that
refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on
DooPlex use "always the newest" (new row).
**Register:** 332 → 335 rows (1 closed: the cut-backup check; 4 opened: the Longhorn fault, the "always newest" apps, a
monitoring sync drift, script tests not in CI).
**Needs you:**
1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops.
2. **When convenient:** pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If
you do nothing, any restart may upgrade one of them by surprise, as it did zipline.
3. **Tester 2's one-time step** is unchanged (below).
## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights
@@ -38,7 +73,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
minutes. Both figures are on the page now.
**Needs you:**
1. **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
1. **(DECIDED 2026-10-05 evening: A, done — see Tonight)** **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
- **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an
alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
- **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A.
@@ -47,7 +82,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
`documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
- **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy
of the database cannot open the console passwords.
2. **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
2. **(DONE 2026-10-05 evening — see Tonight)** **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays
proven by tests only.
3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own