night fixes 2026-10-05: R-867/R-95/R-863..R-869 closed, R-870..R-876 opened (R-876 P2: repair misses dpkg's journal); Part E crash record; Part F spike; golden 0.294.0; rulings 100-103, CC decisions 104-108; STATUS
gates / gates (push) Successful in 32s
gates / gates (push) Successful in 32s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,11 +1,62 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 is installed, but offline since 2026-10-04 20:06 local.**
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 is a laptop that is switched off at night (your word,
|
||||
2026-10-05) — it was offline all session; nothing was sent to it.**
|
||||
|
||||
**Updated 2026-10-05 (morning, after the OS-update night): every box of ours healthy; the ruled undo proven (48 s);
|
||||
decision 78 proven on a fresh Tester 1 box; two P2 defects found (a new box's first app install can fail; the off-site
|
||||
clean-up refuses normal retention and mails an error weekly). One test (a power cut mid-update) not run — needs your
|
||||
permission. Morning note: `documentation/audits/night-2026-10-04/MORNING-NOTE.md`.**
|
||||
**Updated 2026-10-05 (day, the night's fixes): every box of ours healthy. Fixed and proven live: the off-site clean-up
|
||||
now deletes old copies (both demo boxes), a new box's first app install, the update's disk-space check, a killed
|
||||
update's lost report. The power cut in the middle of an update was tested on demo-hp with your go: the box came back by
|
||||
itself in 37 s, but the next update failed until I ran one command by hand — filed (R-876), fix next session.
|
||||
One decision for you below (a box that is off at night). Report: `REPORT-night-fixes-2026-10-05.md`.**
|
||||
|
||||
## Today (2026-10-05, day): the night's fixes, the power cut by day, a box that is off at night
|
||||
|
||||
**Decisions I took myself (you may reverse each — `09` decisions 104–108):**
|
||||
- The off-site clean-up's safety line is now "the last 7 calendar days" — the same number the clean-up keeps — not
|
||||
"8 days old". It can never block an honest clean-up again.
|
||||
- The image clean-up now waits while ANY app install, update, restore or undo is downloading — not only installs.
|
||||
- A killed update keeps its report on the box until the hub has it; the agent looks for such reports every 5 minutes.
|
||||
- **A second agent release today (0.144.1)**, against "one release per repo": the first fix for the lost report was
|
||||
proven NOT to work on demo-hp, and shipping it as it was would have been worse.
|
||||
|
||||
**What works now (proven live):**
|
||||
- **The weekly off-site clean-up really deletes old backups.** One clean-up each by hand: demo-felhom 16 → 14, demo-hp
|
||||
145 → 127 — exactly the backups I predicted. No error mail, the hub's count check quiet, the key files clean.
|
||||
This also closes R-95 (the box can no longer delete its own off-site history, and clean-up now works).
|
||||
- **A new box's first app install works the first time:** the image clean-up met an install at minute 3 on the scratch
|
||||
box, waited, and BookStack installed first try. A failed install now logs its real reason.
|
||||
- **The update's disk-space check counts the real download** (12.8 MB for 13 packages; it counted 0 before).
|
||||
- **A killed update still reports to the hub** (5 minutes later, once). The debug update runs with the hub away.
|
||||
- Controller 0.294.0 on demo-hp, demo-felhom and Tester 1 (floor per customer; Tester 2 not moved). Agent 0.144.1
|
||||
and its root files on all three. New-install image 0.294.0 baked and approved.
|
||||
|
||||
**Found today (filed, not fixed):**
|
||||
- **After a power cut in the middle of an update, every later update fails until someone runs one command on the box**
|
||||
(R-876, P2). The box itself comes back fine. Until the fix: `runbooks/crash-guard.md` has the command. I fix it next.
|
||||
- A box that is off at night (below): no catch-up, no missed-backup alarm (R-872), a "server cannot be reached" mail to
|
||||
the household every night (R-873), restore-tests never run (R-874).
|
||||
|
||||
**Needs you:**
|
||||
|
||||
1. **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its
|
||||
nightly database backups, second copy or off-site copy; the whole-box backup runs only about every 2 days; no
|
||||
alarm says so; the household is mailed "your server cannot be reached" every night. No design document covers it
|
||||
(R-871).
|
||||
- **A (my pick): a missed night runs once when the box comes back.** The database backups, the second copy and
|
||||
the off-site copy run a few minutes after the box is on again (they take seconds to minutes; whether
|
||||
any of them pauses an app is to be measured in the design — the whole-box backup, which does pause apps, already
|
||||
has its own catch-up). App and system updates still
|
||||
wait for a night. Costs: a design section and one controller release; the household may notice a busy disk for a
|
||||
few minutes after switching on.
|
||||
- **B: say plainly that the box must stay on at night.** The setup guide and the box's backup page say it; the
|
||||
missed-night alarm fires after 2 nights off. Costs: wording + one alarm; a laptop household gets an alarm it
|
||||
cannot fix except by changing habits.
|
||||
- **If you do nothing:** Tester 2 keeps having no database or off-site backup, nobody is told, and the household
|
||||
keeps getting the nightly "cannot be reached" mail. A restore-test alarm will fire around 2026-10-11.
|
||||
2. **Tester 2's one-time step** — unchanged from yesterday (below). If you wait, it keeps working; it just cannot get
|
||||
new root files.
|
||||
|
||||
**Recorded, not decisions:** Tester 1's Cloudflare tokens are NOT rotated (your ruling; R-870 has the steps).
|
||||
|
||||
## Today (2026-10-04, night): root files for installed boxes, test approvals, Docker self-repair
|
||||
|
||||
|
||||
Reference in New Issue
Block a user