9bb45eaaa2
gates / gates (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
79 lines
7.1 KiB
Markdown
79 lines
7.1 KiB
Markdown
# REPORT — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (2026-10-05, afternoon)
|
||
|
||
Brief: "a box that was off at night catches up when it comes back (decision A) …" (operator, 2026-10-05). Evidence:
|
||
`documentation/audits/catchup-2026-10-05/`. Architecture read before the claims: `07` §6.1, `08` (cool-downs, §6.2–6.3),
|
||
`09` §3 decision 11, `11` §5.4.1 and §8; the Part F spike (`audits/night-fixes-2026-10-05/partF/FINDINGS.md`).
|
||
|
||
## 1. The Part table
|
||
|
||
| Part | State | Note |
|
||
|---|---|---|
|
||
| §1 rulings recorded first | **done** | `09` decisions 109–111 |
|
||
| A.0 — the design | **done** | new `07` §6.1.1 (`[DESIGN — ruled 2026-10-05]` for 109–110; CC decisions 112–118 "operator may reverse"), `08` §6.4 |
|
||
| A — the catch-up (R-871) | **done** | spike on 9202 first (off across 09:05 → `db-dump scheduled for 2026-10-06 09:05`, nothing ran); built controller v0.295.0; 13 red-proofs; live on 9202 (twice — the second after a host crash interrupted the wait) and on demo-felhom |
|
||
| B — the banner (decision 110) | **done, changed** | rule + real-page tests + red-proofs; served live on 9202 and closed by its real route. **No screenshot: DooPlex has no browser or page renderer** (checked: no chromium/firefox/wkhtmltoimage/playwright); the page HTML as the box served it is the evidence. 9202's ledger was set by hand to a 3-day-old dump to make it show (scratch box, stated). |
|
||
| C — R-872, R-873, R-874, R-875 | **done; R-872 live pending** | hub v0.134.0 (R-872, R-873), agent v0.145.0 (R-874, R-875); each red-proofed. R-874 live on demo-felhom. R-872's first live 05:00 run: dated check 2026-10-06. R-873: tests only (no live occurrence — Tester 2 stayed off). |
|
||
| D — R-876 self-repair | **done** | agent v0.145.0; 3 red-proofs; live with the operator's go: crash mid-unpack, the next pass repaired dpkg by itself and finished; bundle signed to all three boxes |
|
||
| E — release, golden, records | **done** | controller v0.295.0, agent v0.145.0, hub v0.134.0 (one each); golden 0.295.0 baked + vouched (gate OK); floors for demo-hp, demo-felhom, tester-1 |
|
||
|
||
## 2. Claims in the brief that turned out wrong (or only partly true)
|
||
|
||
1. **"A controller start is the right trigger"** — not alone. A host **resume** is needed too: Go's timers run on
|
||
CLOCK_MONOTONIC, which does not count suspended time, so on a suspended laptop the 02:30 timer fires hours late — and
|
||
would start the app-update leg at noon. Built: a late daily timer (> 60 min) is skipped, and a resume watch triggers
|
||
the catch-up. *Reasoned and unit-tested; NOT measured — no suspend was allowed on a demo box.*
|
||
2. **"The whole-guest backup cannot collide with the catch-up"** — it CAN: its 48 h safety valve fires on the first
|
||
5-minute poll after a start, and the catch-up's database dump needs the apps up. Built: each waits for the other.
|
||
3. **"The box keeps a record of when it was on"** — TRUE, and it was not designed as one: the controller's system-metrics
|
||
table (one sample a minute, kept 30 days). Used for "off at 02:30" and the suggestion.
|
||
4. **"The repair step can check the journal without losing R-845's speed"** — TRUE: `--audit` and the journal are read
|
||
in ONE `sh -c` call; a clean pass still costs one call (pinned by a test).
|
||
5. **"Apps keep running during a catch-up"** (my morning STATUS said "to be measured") — the dump leg stops an app with a
|
||
volume for its copy: **measured 1 s** for opengist in the day (R-878).
|
||
6. R-873's brief offered "only when the box was off for more than 24 h" — rejected (a real outage would reach the
|
||
household a day late); the weekly rule was chosen (decision 116).
|
||
|
||
## 3. What was proven, with numbers
|
||
|
||
- **Part A (live).** 9202: W 09:35, off 07:33–07:38 UTC, start 07:38:09 → `missed [db-dump] … ONE catch-up in 15m0s` →
|
||
07:53:09 dump, done in 21 s. Host crash at 07:56 interrupted the next wait → new start 07:58:35 → ONE catch-up of all
|
||
three legs at 08:13:35, done in 22 s. demo-felhom: W 10:07, controller parked + stopped 08:05–08:10 (apps running) →
|
||
catch-up 08:25:02, dump in 1 s; hub event `backup_catchup_done` "Kimaradt mentés pótolva: a doboz ki volt kapcsolva
|
||
10:07-kor, a mentés most elkészült." — no mail. Window set back to 02:30.
|
||
- **Part B.** `nightchain.ComputeBanner` tests (shown / fresh / upgrade day / dismissed then back after a new miss / gone
|
||
after a success / off-site counts / no pattern / usually on); web tests through ServeHTTP and the real dismiss route.
|
||
- **Part C.** R-872 Tester 2 shape → both alarms; a dump 20 h ago → quiet; a new box → quiet. R-873: night 1 household
|
||
2 mails, night 2 household 0 / operator 2, after 8 days household again. R-874 live: demo-felhom agent start
|
||
07:38:46 → first check 08:08:46 → restore-test passed in 29 s.
|
||
- **Part D (live, operator's go).** 13 packages rolled back; crash 07:56:03.308 UTC during
|
||
`dpkg --force-confold … --unpack`; new boot 07:56:41 (38 s), guard armed, 1 unclean boot in window; at boot
|
||
`--audit` clean, journal 1 file, 1 package new / 12 old; next pass (nobody touched dpkg): `REPAIR configured=0
|
||
journal=1` → `PLAN upgrade=12` → `DONE rc=0 upgraded=12`, healthy; package list identical (279 lines). No mail.
|
||
|
||
## 4. Rows
|
||
|
||
Register before **340**, after **336**. Closed (6): R-871, R-873, R-874, R-875, R-876, R-877. Narrowed: R-872 (fixed,
|
||
dated check 2026-10-06). Opened (2): **R-877** (the Tester 1 VM had no start-on-boot; the morning crash left it off
|
||
1 h 17 min — filed and closed), **R-878** (P4, a daytime catch-up stops a volume app for its dump). STATUS updated.
|
||
|
||
## 5. Slips of mine, said plainly
|
||
|
||
- **This morning's report said every box was healthy; the Tester 1 VM had been off since the 06:14 crash** (R-877). Found
|
||
at 07:31 from its stale report time.
|
||
- **The hub image 0.134.0 was first built from a commit that was not yet pushed** (my unstaged doc edits blocked the
|
||
`git pull`; the push was then refused by the golden gate). I pushed after the golden bake and REBUILT the image from
|
||
the pushed commit `1b0678fa`; the deployed image is the rebuilt one (`sha256:f823b10e…`).
|
||
- Two of my red-proofs did not convict at first (a masked mutation, a test that matched another call); both tests were
|
||
strengthened and the red-proofs re-run — recorded in the red-proof files.
|
||
|
||
## 6. Teardown, three layers
|
||
|
||
- **Machines:** 9202: controller 0.295.0 (by hand, allowed there), window back to 02:30, its ledger was set by hand for
|
||
the banner and has since been overwritten by real catch-up runs. demo-felhom: window 02:30 again, controller unparked.
|
||
demo-hp: every package current, list identical. Tester 1: running, start-on-boot set. Bake VM: CT 9100 destroyed,
|
||
token shredded, `virgin`, qemu gone.
|
||
- **Hosts:** the park file on felhom-pve removed; nothing provisioned.
|
||
- **Hub:** floors 0.295.0 (MinAgent 0.131.0) for demo-hp, demo-felhom, tester-1; artifacts vouched (agent 0.145.0,
|
||
golden 0.295.0); 6 signed jobs (agent_update and agent_config_update for each of the three boxes); hub 0.134.0 deployed
|
||
(ArgoCD Synced/Healthy). Tester 2: read only, offline all session, nothing sent.
|