# REPORT — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (2026-10-05, afternoon) Brief: "a box that was off at night catches up when it comes back (decision A) …" (operator, 2026-10-05). Evidence: `documentation/audits/catchup-2026-10-05/`. Architecture read before the claims: `07` §6.1, `08` (cool-downs, §6.2–6.3), `09` §3 decision 11, `11` §5.4.1 and §8; the Part F spike (`audits/night-fixes-2026-10-05/partF/FINDINGS.md`). ## 1. The Part table | Part | State | Note | |---|---|---| | §1 rulings recorded first | **done** | `09` decisions 109–111 | | A.0 — the design | **done** | new `07` §6.1.1 (`[DESIGN — ruled 2026-10-05]` for 109–110; CC decisions 112–118 "operator may reverse"), `08` §6.4 | | A — the catch-up (R-871) | **done** | spike on 9202 first (off across 09:05 → `db-dump scheduled for 2026-10-06 09:05`, nothing ran); built controller v0.295.0; 13 red-proofs; live on 9202 (twice — the second after a host crash interrupted the wait) and on demo-felhom | | B — the banner (decision 110) | **done, changed** | rule + real-page tests + red-proofs; served live on 9202 and closed by its real route. **No screenshot: DooPlex has no browser or page renderer** (checked: no chromium/firefox/wkhtmltoimage/playwright); the page HTML as the box served it is the evidence. 9202's ledger was set by hand to a 3-day-old dump to make it show (scratch box, stated). | | C — R-872, R-873, R-874, R-875 | **done; R-872 live pending** | hub v0.134.0 (R-872, R-873), agent v0.145.0 (R-874, R-875); each red-proofed. R-874 live on demo-felhom. R-872's first live 05:00 run: dated check 2026-10-06. R-873: tests only (no live occurrence — Tester 2 stayed off). | | D — R-876 self-repair | **done** | agent v0.145.0; 3 red-proofs; live with the operator's go: crash mid-unpack, the next pass repaired dpkg by itself and finished; bundle signed to all three boxes | | E — release, golden, records | **done** | controller v0.295.0, agent v0.145.0, hub v0.134.0 (one each); golden 0.295.0 baked + vouched (gate OK); floors for demo-hp, demo-felhom, tester-1 | ## 2. Claims in the brief that turned out wrong (or only partly true) 1. **"A controller start is the right trigger"** — not alone. A host **resume** is needed too: Go's timers run on CLOCK_MONOTONIC, which does not count suspended time, so on a suspended laptop the 02:30 timer fires hours late — and would start the app-update leg at noon. Built: a late daily timer (> 60 min) is skipped, and a resume watch triggers the catch-up. *Reasoned and unit-tested; NOT measured — no suspend was allowed on a demo box.* 2. **"The whole-guest backup cannot collide with the catch-up"** — it CAN: its 48 h safety valve fires on the first 5-minute poll after a start, and the catch-up's database dump needs the apps up. Built: each waits for the other. 3. **"The box keeps a record of when it was on"** — TRUE, and it was not designed as one: the controller's system-metrics table (one sample a minute, kept 30 days). Used for "off at 02:30" and the suggestion. 4. **"The repair step can check the journal without losing R-845's speed"** — TRUE: `--audit` and the journal are read in ONE `sh -c` call; a clean pass still costs one call (pinned by a test). 5. **"Apps keep running during a catch-up"** (my morning STATUS said "to be measured") — the dump leg stops an app with a volume for its copy: **measured 1 s** for opengist in the day (R-878). 6. R-873's brief offered "only when the box was off for more than 24 h" — rejected (a real outage would reach the household a day late); the weekly rule was chosen (decision 116). ## 3. What was proven, with numbers - **Part A (live).** 9202: W 09:35, off 07:33–07:38 UTC, start 07:38:09 → `missed [db-dump] … ONE catch-up in 15m0s` → 07:53:09 dump, done in 21 s. Host crash at 07:56 interrupted the next wait → new start 07:58:35 → ONE catch-up of all three legs at 08:13:35, done in 22 s. demo-felhom: W 10:07, controller parked + stopped 08:05–08:10 (apps running) → catch-up 08:25:02, dump in 1 s; hub event `backup_catchup_done` "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült." — no mail. Window set back to 02:30. - **Part B.** `nightchain.ComputeBanner` tests (shown / fresh / upgrade day / dismissed then back after a new miss / gone after a success / off-site counts / no pattern / usually on); web tests through ServeHTTP and the real dismiss route. - **Part C.** R-872 Tester 2 shape → both alarms; a dump 20 h ago → quiet; a new box → quiet. R-873: night 1 household 2 mails, night 2 household 0 / operator 2, after 8 days household again. R-874 live: demo-felhom agent start 07:38:46 → first check 08:08:46 → restore-test passed in 29 s. - **Part D (live, operator's go).** 13 packages rolled back; crash 07:56:03.308 UTC during `dpkg --force-confold … --unpack`; new boot 07:56:41 (38 s), guard armed, 1 unclean boot in window; at boot `--audit` clean, journal 1 file, 1 package new / 12 old; next pass (nobody touched dpkg): `REPAIR configured=0 journal=1` → `PLAN upgrade=12` → `DONE rc=0 upgraded=12`, healthy; package list identical (279 lines). No mail. ## 4. Rows Register before **340**, after **336**. Closed (6): R-871, R-873, R-874, R-875, R-876, R-877. Narrowed: R-872 (fixed, dated check 2026-10-06). Opened (2): **R-877** (the Tester 1 VM had no start-on-boot; the morning crash left it off 1 h 17 min — filed and closed), **R-878** (P4, a daytime catch-up stops a volume app for its dump). STATUS updated. ## 5. Slips of mine, said plainly - **This morning's report said every box was healthy; the Tester 1 VM had been off since the 06:14 crash** (R-877). Found at 07:31 from its stale report time. - **The hub image 0.134.0 was first built from a commit that was not yet pushed** (my unstaged doc edits blocked the `git pull`; the push was then refused by the golden gate). I pushed after the golden bake and REBUILT the image from the pushed commit `1b0678fa`; the deployed image is the rebuilt one (`sha256:f823b10e…`). - Two of my red-proofs did not convict at first (a masked mutation, a test that matched another call); both tests were strengthened and the red-proofs re-run — recorded in the red-proof files. ## 6. Teardown, three layers - **Machines:** 9202: controller 0.295.0 (by hand, allowed there), window back to 02:30, its ledger was set by hand for the banner and has since been overwritten by real catch-up runs. demo-felhom: window 02:30 again, controller unparked. demo-hp: every package current, list identical. Tester 1: running, start-on-boot set. Bake VM: CT 9100 destroyed, token shredded, `virgin`, qemu gone. - **Hosts:** the park file on felhom-pve removed; nothing provisioned. - **Hub:** floors 0.295.0 (MinAgent 0.131.0) for demo-hp, demo-felhom, tester-1; artifacts vouched (agent 0.145.0, golden 0.295.0); 6 signed jobs (agent_update and agent_config_update for each of the three boxes); hub 0.134.0 deployed (ArgoCD Synced/Healthy). Tester 2: read only, offline all session, nothing sent.