catch-up session 2026-10-05: design 07 §6.1.1 (a box that is not always on), 08 §6.4; rulings 109-111, CC decisions 112-118; R-871/R-873..R-877 closed, R-872 narrowed (dated check), R-878 opened; live evidence; STATUS
gates / gates (push) Successful in 33s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 10:29:22 +02:00
parent ff25f1076d
commit 9bb45eaaa2
34 changed files with 790 additions and 18 deletions
+43 -8
View File
@@ -1,13 +1,48 @@
# STATUS — what works, what's broken, what's next
**Ready for the first real tester (Tester-2): yes. Tester 2 is a laptop that is switched off at night (your word,
2026-10-05) — it was offline all session; nothing was sent to it.**
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) stayed offline all day; nothing
was sent to it.**
**Updated 2026-10-05 (day, the night's fixes): every box of ours healthy. Fixed and proven live: the off-site clean-up
now deletes old copies (both demo boxes), a new box's first app install, the update's disk-space check, a killed
update's lost report. The power cut in the middle of an update was tested on demo-hp with your go: the box came back by
itself in 37 s, but the next update failed until I ran one command by hand — filed (R-876), fix next session.
One decision for you below (a box that is off at night). Report: `REPORT-night-fixes-2026-10-05.md`.**
**Updated 2026-10-05 (afternoon, the catch-up session): every box of ours healthy. Built and proven live: a box that was
off at its backup time makes the backups up once when it comes back; the household's banner; the OS update repairs
itself after a power cut (second crash on demo-hp, with your go). Report: `REPORT-catchup-2026-10-05.md`.**
## Today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut
**Decisions I took myself (you may reverse each — `09` decisions 112–118):**
- The make-up run starts 15 minutes after the box comes back; a backup due within 30 minutes is left to its normal time.
- A laptop that sleeps through the night: a nightly job that wakes up more than an hour late is skipped (otherwise app
updates would start at noon); the backups are made up instead. Tested, not measured (I may not suspend a box).
- The banner appears when the last backup is over 26 hours old; it suggests the latest evening hour the box is usually
on (5 of the last 7 days), and nothing when the box is usually on at its backup time.
- A box that is off at the 05:00 check now raises the missed-backup alarm after 2 nights without a database backup (3
without a whole-box backup) — not after one, so a box that broke last night gives only its "offline" alarm.
- The household hears "your server cannot be reached" at most once a week; you still hear every one.
- The restore-test's first check is 30 minutes after the agent starts (a box on for short times now gets tested).
- The update's repair step now also looks at dpkg's journal — the place the power cut left its mark.
**What works now (proven live):**
- **A missed night is made up once** (your choice A): the scratch box and demo-felhom were off across their backup time;
15 minutes after they came back, the missed backups ran by themselves (seconds). The household's timeline got one line:
"Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült." No mail.
- **The banner** (your idea) appeared on the scratch box, with the "change the backup time" button; "Close" kept it
closed. *No screenshot: there is no browser on DooPlex; I captured the page as the box served it.*
- **The power cut, again** (demo-hp, your go): back by itself in 38 seconds; **the next update repaired dpkg by itself
and finished — nobody typed anything.** No mail.
- **A restore-test 30 minutes after an agent start** ran and passed on demo-felhom.
- Controller 0.295.0, agent 0.145.0 (+ its root files) on demo-hp, demo-felhom and Tester 1; hub 0.134.0; new-install
image 0.295.0 baked and approved.
**Found today:**
- **My slip from this morning:** the Tester 1 test machine did not restart after the morning crash and stayed off for
1 h 17 min; my morning report said every box was healthy. It now starts by itself after a crash (proven by the second
crash).
- **What the household may notice from a make-up run:** the database backup stops an app with stored files for its copy —
1 second for opengist. The night does the same unseen; a big app may take longer, in the day (filed, small).
**Needs you:** nothing urgent. **Tester 2's one-time step** is unchanged (below). The new missed-backup alarm will be
checked at tomorrow's 05:00 run; if Tester 2 is still off, you will get its first real "backup missed" mail — that is
the fix working, not a new fault.
## Today (2026-10-05, day): the night's fixes, the power cut by day, a box that is off at night
@@ -38,7 +73,7 @@ One decision for you below (a box that is off at night). Report: `REPORT-night-f
**Needs you:**
1. **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its
1. **[DECIDED 2026-10-05 08:42 — option A, built the same afternoon; see above]** **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its
nightly database backups, second copy or off-site copy; the whole-box backup runs only about every 2 days; no
alarm says so; the household is mailed "your server cannot be reached" every night. No design document covers it
(R-871).