Files
felhom.eu/REPORT-catchup-2026-10-05.md
T

79 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (2026-10-05, afternoon)
Brief: "a box that was off at night catches up when it comes back (decision A) …" (operator, 2026-10-05). Evidence:
`documentation/audits/catchup-2026-10-05/`. Architecture read before the claims: `07` §6.1, `08` (cool-downs, §6.2–6.3),
`09` §3 decision 11, `11` §5.4.1 and §8; the Part F spike (`audits/night-fixes-2026-10-05/partF/FINDINGS.md`).
## 1. The Part table
| Part | State | Note |
|---|---|---|
| §1 rulings recorded first | **done** | `09` decisions 109–111 |
| A.0 — the design | **done** | new `07` §6.1.1 (`[DESIGN — ruled 2026-10-05]` for 109–110; CC decisions 112–118 "operator may reverse"), `08` §6.4 |
| A — the catch-up (R-871) | **done** | spike on 9202 first (off across 09:05 → `db-dump scheduled for 2026-10-06 09:05`, nothing ran); built controller v0.295.0; 13 red-proofs; live on 9202 (twice — the second after a host crash interrupted the wait) and on demo-felhom |
| B — the banner (decision 110) | **done, changed** | rule + real-page tests + red-proofs; served live on 9202 and closed by its real route. **No screenshot: DooPlex has no browser or page renderer** (checked: no chromium/firefox/wkhtmltoimage/playwright); the page HTML as the box served it is the evidence. 9202's ledger was set by hand to a 3-day-old dump to make it show (scratch box, stated). |
| C — R-872, R-873, R-874, R-875 | **done; R-872 live pending** | hub v0.134.0 (R-872, R-873), agent v0.145.0 (R-874, R-875); each red-proofed. R-874 live on demo-felhom. R-872's first live 05:00 run: dated check 2026-10-06. R-873: tests only (no live occurrence — Tester 2 stayed off). |
| D — R-876 self-repair | **done** | agent v0.145.0; 3 red-proofs; live with the operator's go: crash mid-unpack, the next pass repaired dpkg by itself and finished; bundle signed to all three boxes |
| E — release, golden, records | **done** | controller v0.295.0, agent v0.145.0, hub v0.134.0 (one each); golden 0.295.0 baked + vouched (gate OK); floors for demo-hp, demo-felhom, tester-1 |
## 2. Claims in the brief that turned out wrong (or only partly true)
1. **"A controller start is the right trigger"** — not alone. A host **resume** is needed too: Go's timers run on
CLOCK_MONOTONIC, which does not count suspended time, so on a suspended laptop the 02:30 timer fires hours late — and
would start the app-update leg at noon. Built: a late daily timer (> 60 min) is skipped, and a resume watch triggers
the catch-up. *Reasoned and unit-tested; NOT measured — no suspend was allowed on a demo box.*
2. **"The whole-guest backup cannot collide with the catch-up"** — it CAN: its 48 h safety valve fires on the first
5-minute poll after a start, and the catch-up's database dump needs the apps up. Built: each waits for the other.
3. **"The box keeps a record of when it was on"** — TRUE, and it was not designed as one: the controller's system-metrics
table (one sample a minute, kept 30 days). Used for "off at 02:30" and the suggestion.
4. **"The repair step can check the journal without losing R-845's speed"** — TRUE: `--audit` and the journal are read
in ONE `sh -c` call; a clean pass still costs one call (pinned by a test).
5. **"Apps keep running during a catch-up"** (my morning STATUS said "to be measured") — the dump leg stops an app with a
volume for its copy: **measured 1 s** for opengist in the day (R-878).
6. R-873's brief offered "only when the box was off for more than 24 h" — rejected (a real outage would reach the
household a day late); the weekly rule was chosen (decision 116).
## 3. What was proven, with numbers
- **Part A (live).** 9202: W 09:35, off 07:33–07:38 UTC, start 07:38:09 → `missed [db-dump] … ONE catch-up in 15m0s` →
07:53:09 dump, done in 21 s. Host crash at 07:56 interrupted the next wait → new start 07:58:35 → ONE catch-up of all
three legs at 08:13:35, done in 22 s. demo-felhom: W 10:07, controller parked + stopped 08:05–08:10 (apps running) →
catch-up 08:25:02, dump in 1 s; hub event `backup_catchup_done` "Kimaradt mentés pótolva: a doboz ki volt kapcsolva
10:07-kor, a mentés most elkészült." — no mail. Window set back to 02:30.
- **Part B.** `nightchain.ComputeBanner` tests (shown / fresh / upgrade day / dismissed then back after a new miss / gone
after a success / off-site counts / no pattern / usually on); web tests through ServeHTTP and the real dismiss route.
- **Part C.** R-872 Tester 2 shape → both alarms; a dump 20 h ago → quiet; a new box → quiet. R-873: night 1 household
2 mails, night 2 household 0 / operator 2, after 8 days household again. R-874 live: demo-felhom agent start
07:38:46 → first check 08:08:46 → restore-test passed in 29 s.
- **Part D (live, operator's go).** 13 packages rolled back; crash 07:56:03.308 UTC during
`dpkg --force-confold … --unpack`; new boot 07:56:41 (38 s), guard armed, 1 unclean boot in window; at boot
`--audit` clean, journal 1 file, 1 package new / 12 old; next pass (nobody touched dpkg): `REPAIR configured=0
journal=1` → `PLAN upgrade=12` → `DONE rc=0 upgraded=12`, healthy; package list identical (279 lines). No mail.
## 4. Rows
Register before **340**, after **336**. Closed (6): R-871, R-873, R-874, R-875, R-876, R-877. Narrowed: R-872 (fixed,
dated check 2026-10-06). Opened (2): **R-877** (the Tester 1 VM had no start-on-boot; the morning crash left it off
1 h 17 min — filed and closed), **R-878** (P4, a daytime catch-up stops a volume app for its dump). STATUS updated.
## 5. Slips of mine, said plainly
- **This morning's report said every box was healthy; the Tester 1 VM had been off since the 06:14 crash** (R-877). Found
at 07:31 from its stale report time.
- **The hub image 0.134.0 was first built from a commit that was not yet pushed** (my unstaged doc edits blocked the
`git pull`; the push was then refused by the golden gate). I pushed after the golden bake and REBUILT the image from
the pushed commit `1b0678fa`; the deployed image is the rebuilt one (`sha256:f823b10e…`).
- Two of my red-proofs did not convict at first (a masked mutation, a test that matched another call); both tests were
strengthened and the red-proofs re-run — recorded in the red-proof files.
## 6. Teardown, three layers
- **Machines:** 9202: controller 0.295.0 (by hand, allowed there), window back to 02:30, its ledger was set by hand for
the banner and has since been overwritten by real catch-up runs. demo-felhom: window 02:30 again, controller unparked.
demo-hp: every package current, list identical. Tester 1: running, start-on-boot set. Bake VM: CT 9100 destroyed,
token shredded, `virgin`, qemu gone.
- **Hosts:** the park file on felhom-pve removed; nothing provisioned.
- **Hub:** floors 0.295.0 (MinAgent 0.131.0) for demo-hp, demo-felhom, tester-1; artifacts vouched (agent 0.145.0,
golden 0.295.0); 6 signed jobs (agent_update and agent_config_update for each of the three boxes); hub 0.134.0 deployed
(ArgoCD Synced/Healthy). Tester 2: read only, offline all session, nothing sent.