Files
felhom.eu/REPORT-catchup-2026-10-05.md
T

7.1 KiB
Raw Blame History

REPORT — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (2026-10-05, afternoon)

Brief: "a box that was off at night catches up when it comes back (decision A) …" (operator, 2026-10-05). Evidence: documentation/audits/catchup-2026-10-05/. Architecture read before the claims: 07 §6.1, 08 (cool-downs, §6.2–6.3), 09 §3 decision 11, 11 §5.4.1 and §8; the Part F spike (audits/night-fixes-2026-10-05/partF/FINDINGS.md).

1. The Part table

Part State Note
§1 rulings recorded first done 09 decisions 109–111
A.0 — the design done new 07 §6.1.1 ([DESIGN — ruled 2026-10-05] for 109–110; CC decisions 112–118 "operator may reverse"), 08 §6.4
A — the catch-up (R-871) done spike on 9202 first (off across 09:05 → db-dump scheduled for 2026-10-06 09:05, nothing ran); built controller v0.295.0; 13 red-proofs; live on 9202 (twice — the second after a host crash interrupted the wait) and on demo-felhom
B — the banner (decision 110) done, changed rule + real-page tests + red-proofs; served live on 9202 and closed by its real route. No screenshot: DooPlex has no browser or page renderer (checked: no chromium/firefox/wkhtmltoimage/playwright); the page HTML as the box served it is the evidence. 9202's ledger was set by hand to a 3-day-old dump to make it show (scratch box, stated).
C — R-872, R-873, R-874, R-875 done; R-872 live pending hub v0.134.0 (R-872, R-873), agent v0.145.0 (R-874, R-875); each red-proofed. R-874 live on demo-felhom. R-872's first live 05:00 run: dated check 2026-10-06. R-873: tests only (no live occurrence — Tester 2 stayed off).
D — R-876 self-repair done agent v0.145.0; 3 red-proofs; live with the operator's go: crash mid-unpack, the next pass repaired dpkg by itself and finished; bundle signed to all three boxes
E — release, golden, records done controller v0.295.0, agent v0.145.0, hub v0.134.0 (one each); golden 0.295.0 baked + vouched (gate OK); floors for demo-hp, demo-felhom, tester-1

2. Claims in the brief that turned out wrong (or only partly true)

  1. "A controller start is the right trigger" — not alone. A host resume is needed too: Go's timers run on CLOCK_MONOTONIC, which does not count suspended time, so on a suspended laptop the 02:30 timer fires hours late — and would start the app-update leg at noon. Built: a late daily timer (> 60 min) is skipped, and a resume watch triggers the catch-up. Reasoned and unit-tested; NOT measured — no suspend was allowed on a demo box.
  2. "The whole-guest backup cannot collide with the catch-up" — it CAN: its 48 h safety valve fires on the first 5-minute poll after a start, and the catch-up's database dump needs the apps up. Built: each waits for the other.
  3. "The box keeps a record of when it was on" — TRUE, and it was not designed as one: the controller's system-metrics table (one sample a minute, kept 30 days). Used for "off at 02:30" and the suggestion.
  4. "The repair step can check the journal without losing R-845's speed" — TRUE: --audit and the journal are read in ONE sh -c call; a clean pass still costs one call (pinned by a test).
  5. "Apps keep running during a catch-up" (my morning STATUS said "to be measured") — the dump leg stops an app with a volume for its copy: measured 1 s for opengist in the day (R-878).
  6. R-873's brief offered "only when the box was off for more than 24 h" — rejected (a real outage would reach the household a day late); the weekly rule was chosen (decision 116).

3. What was proven, with numbers

  • Part A (live). 9202: W 09:35, off 07:33–07:38 UTC, start 07:38:09 → missed [db-dump] … ONE catch-up in 15m0s → 07:53:09 dump, done in 21 s. Host crash at 07:56 interrupted the next wait → new start 07:58:35 → ONE catch-up of all three legs at 08:13:35, done in 22 s. demo-felhom: W 10:07, controller parked + stopped 08:05–08:10 (apps running) → catch-up 08:25:02, dump in 1 s; hub event backup_catchup_done "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült." — no mail. Window set back to 02:30.
  • Part B. nightchain.ComputeBanner tests (shown / fresh / upgrade day / dismissed then back after a new miss / gone after a success / off-site counts / no pattern / usually on); web tests through ServeHTTP and the real dismiss route.
  • Part C. R-872 Tester 2 shape → both alarms; a dump 20 h ago → quiet; a new box → quiet. R-873: night 1 household 2 mails, night 2 household 0 / operator 2, after 8 days household again. R-874 live: demo-felhom agent start 07:38:46 → first check 08:08:46 → restore-test passed in 29 s.
  • Part D (live, operator's go). 13 packages rolled back; crash 07:56:03.308 UTC during dpkg --force-confold … --unpack; new boot 07:56:41 (38 s), guard armed, 1 unclean boot in window; at boot --audit clean, journal 1 file, 1 package new / 12 old; next pass (nobody touched dpkg): REPAIR configured=0 journal=1 → PLAN upgrade=12 → DONE rc=0 upgraded=12, healthy; package list identical (279 lines). No mail.

4. Rows

Register before 340, after 336. Closed (6): R-871, R-873, R-874, R-875, R-876, R-877. Narrowed: R-872 (fixed, dated check 2026-10-06). Opened (2): R-877 (the Tester 1 VM had no start-on-boot; the morning crash left it off 1 h 17 min — filed and closed), R-878 (P4, a daytime catch-up stops a volume app for its dump). STATUS updated.

5. Slips of mine, said plainly

  • This morning's report said every box was healthy; the Tester 1 VM had been off since the 06:14 crash (R-877). Found at 07:31 from its stale report time.
  • The hub image 0.134.0 was first built from a commit that was not yet pushed (my unstaged doc edits blocked the git pull; the push was then refused by the golden gate). I pushed after the golden bake and REBUILT the image from the pushed commit 1b0678fa; the deployed image is the rebuilt one (sha256:f823b10e…).
  • Two of my red-proofs did not convict at first (a masked mutation, a test that matched another call); both tests were strengthened and the red-proofs re-run — recorded in the red-proof files.

6. Teardown, three layers

  • Machines: 9202: controller 0.295.0 (by hand, allowed there), window back to 02:30, its ledger was set by hand for the banner and has since been overwritten by real catch-up runs. demo-felhom: window 02:30 again, controller unparked. demo-hp: every package current, list identical. Tester 1: running, start-on-boot set. Bake VM: CT 9100 destroyed, token shredded, virgin, qemu gone.
  • Hosts: the park file on felhom-pve removed; nothing provisioned.
  • Hub: floors 0.295.0 (MinAgent 0.131.0) for demo-hp, demo-felhom, tester-1; artifacts vouched (agent 0.145.0, golden 0.295.0); 6 signed jobs (agent_update and agent_config_update for each of the three boxes); hub 0.134.0 deployed (ArgoCD Synced/Healthy). Tester 2: read only, offline all session, nothing sent.