From 9bb45eaaa2d595097de7ff60681d8eb38559a972 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 5 Oct 2026 10:29:22 +0200 Subject: [PATCH] =?UTF-8?q?catch-up=20session=202026-10-05:=20design=2007?= =?UTF-8?q?=20=C2=A76.1.1=20(a=20box=20that=20is=20not=20always=20on),=200?= =?UTF-8?q?8=20=C2=A76.4;=20rulings=20109-111,=20CC=20decisions=20112-118;?= =?UTF-8?q?=20R-871/R-873..R-877=20closed,=20R-872=20narrowed=20(dated=20c?= =?UTF-8?q?heck),=20R-878=20opened;=20live=20evidence;=20STATUS?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 10 ++ REPORT-catchup-2026-10-05.md | 78 ++++++++++++ STATUS.md | 51 ++++++-- .../architecture/00-capability-map.md | 3 +- .../architecture/07-backup-architecture.md | 74 +++++++++++ documentation/architecture/08-alarm-ladder.md | 20 +++ .../architecture/09-update-architecture.md | 38 ++++++ documentation/architecture/11-os-updates.md | 13 ++ .../audits/catchup-2026-10-05/final-state.txt | 9 ++ .../partA-spike/s1-set-window.txt | 12 ++ .../partA-spike/s2-off-across-W.txt | 9 ++ .../partA/hub-timeline-line.txt | 1 + .../partA/live-9202-1-deploy.txt | 12 ++ .../partA/live-9202-2-off-across-W.txt | 41 ++++++ .../live-9202-3-crash-interrupted-catchup.txt | 23 ++++ .../partA/live-demo-felhom.txt | 24 ++++ .../catchup-2026-10-05/partA/red-proofs.txt | 120 ++++++++++++++++++ .../partB/live-banner-9202.txt | 27 ++++ .../catchup-2026-10-05/partB/red-proofs.txt | 56 ++++++++ .../partC/r872-red-proofs.txt | 27 ++++ .../partC/r873-red-proof.txt | 8 ++ .../partC/r874-live-demo-felhom.txt | 11 ++ .../partC/r874-red-proof.txt | 9 ++ .../partC/r875-red-proof.txt | 8 ++ .../partD/d0-before-rollback.txt | 12 ++ .../catchup-2026-10-05/partD/d1-crash.txt | 5 + .../partD/d2-back-and-dpkg.txt | 25 ++++ .../partD/d3-next-pass-self-repair.txt | 25 ++++ .../partD/d4-events-mails.txt | 8 ++ .../partD/r876-red-proofs.txt | 17 +++ .../audits/catchup-2026-10-05/partE/vouch.txt | 3 + .../tester1/vm341-was-stopped.txt | 3 + documentation/backlog/CLOSED-ITEMS.md | 11 ++ documentation/backlog/OPEN-ITEMS.md | 15 +-- 34 files changed, 790 insertions(+), 18 deletions(-) create mode 100644 REPORT-catchup-2026-10-05.md create mode 100644 documentation/audits/catchup-2026-10-05/final-state.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA-spike/s1-set-window.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA-spike/s2-off-across-W.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA/hub-timeline-line.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA/live-9202-1-deploy.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA/live-9202-2-off-across-W.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA/live-9202-3-crash-interrupted-catchup.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA/live-demo-felhom.txt create mode 100644 documentation/audits/catchup-2026-10-05/partA/red-proofs.txt create mode 100644 documentation/audits/catchup-2026-10-05/partB/live-banner-9202.txt create mode 100644 documentation/audits/catchup-2026-10-05/partB/red-proofs.txt create mode 100644 documentation/audits/catchup-2026-10-05/partC/r872-red-proofs.txt create mode 100644 documentation/audits/catchup-2026-10-05/partC/r873-red-proof.txt create mode 100644 documentation/audits/catchup-2026-10-05/partC/r874-live-demo-felhom.txt create mode 100644 documentation/audits/catchup-2026-10-05/partC/r874-red-proof.txt create mode 100644 documentation/audits/catchup-2026-10-05/partC/r875-red-proof.txt create mode 100644 documentation/audits/catchup-2026-10-05/partD/d0-before-rollback.txt create mode 100644 documentation/audits/catchup-2026-10-05/partD/d1-crash.txt create mode 100644 documentation/audits/catchup-2026-10-05/partD/d2-back-and-dpkg.txt create mode 100644 documentation/audits/catchup-2026-10-05/partD/d3-next-pass-self-repair.txt create mode 100644 documentation/audits/catchup-2026-10-05/partD/d4-events-mails.txt create mode 100644 documentation/audits/catchup-2026-10-05/partD/r876-red-proofs.txt create mode 100644 documentation/audits/catchup-2026-10-05/partE/vouch.txt create mode 100644 documentation/audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt diff --git a/CONTEXT.md b/CONTEXT.md index 5a4dd1f9..30996970 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,16 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-05 (afternoon) — a box that is not always on (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0 +> vouched with agent 0.145.0, min_agent 0.131.0).** Rulings 109–111 (`09` §3: R-871 option A — a missed night runs once when the +> box comes back; the household's banner; Tester 2 read only). CC decisions 112–118, *operator may reverse*. Design `07` +> §6.1.1 (new): `internal/nightchain` ledger (`night-ledger.json`, an ATTEMPT record) + `CatchUp` (start + resume triggers, +> 15 min, backup legs only, `catchUpLegs`) + `scheduler.DailyLateLimit` (60 min) + `quiesce.SetCatchUpFn`; the banner +> (`nightchain.ComputeBanner`, 26 h, metrics-record suggestion, durable dismissal). Hub: R-872 down boxes judged at 48 h/72 h +> (open, dated check 2026-10-06), R-873 household liveness mail weekly, `backup_catchup_done` allowed. Agent: R-876 dpkg journal +> repair (live, second crash), R-874 first restore-test check 30 min after start (live), R-875. Closed R-871, R-873..R-877; +> opened R-877 (closed), R-878. Floors: demo-hp, demo-felhom, tester-1 at 0.295.0. Report: `REPORT-catchup-2026-10-05.md`. + > **2026-10-05 (day) — the night's fixes (controller v0.294.0, agent v0.144.0 + v0.144.1, golden 0.294.0 vouched with agent > 0.144.1, min_agent 0.131.0).** Rulings 100–103 (`09` §3: Tester 1's CF tokens NOT rotated → R-870; A1 by day on demo-hp > only, with the operator's go; Tester 2 is a laptop off at night; re-sign Tester 2 only if online — it never was). diff --git a/REPORT-catchup-2026-10-05.md b/REPORT-catchup-2026-10-05.md new file mode 100644 index 00000000..8a087386 --- /dev/null +++ b/REPORT-catchup-2026-10-05.md @@ -0,0 +1,78 @@ +# REPORT — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (2026-10-05, afternoon) + +Brief: "a box that was off at night catches up when it comes back (decision A) …" (operator, 2026-10-05). Evidence: +`documentation/audits/catchup-2026-10-05/`. Architecture read before the claims: `07` §6.1, `08` (cool-downs, §6.2–6.3), +`09` §3 decision 11, `11` §5.4.1 and §8; the Part F spike (`audits/night-fixes-2026-10-05/partF/FINDINGS.md`). + +## 1. The Part table + +| Part | State | Note | +|---|---|---| +| §1 rulings recorded first | **done** | `09` decisions 109–111 | +| A.0 — the design | **done** | new `07` §6.1.1 (`[DESIGN — ruled 2026-10-05]` for 109–110; CC decisions 112–118 "operator may reverse"), `08` §6.4 | +| A — the catch-up (R-871) | **done** | spike on 9202 first (off across 09:05 → `db-dump scheduled for 2026-10-06 09:05`, nothing ran); built controller v0.295.0; 13 red-proofs; live on 9202 (twice — the second after a host crash interrupted the wait) and on demo-felhom | +| B — the banner (decision 110) | **done, changed** | rule + real-page tests + red-proofs; served live on 9202 and closed by its real route. **No screenshot: DooPlex has no browser or page renderer** (checked: no chromium/firefox/wkhtmltoimage/playwright); the page HTML as the box served it is the evidence. 9202's ledger was set by hand to a 3-day-old dump to make it show (scratch box, stated). | +| C — R-872, R-873, R-874, R-875 | **done; R-872 live pending** | hub v0.134.0 (R-872, R-873), agent v0.145.0 (R-874, R-875); each red-proofed. R-874 live on demo-felhom. R-872's first live 05:00 run: dated check 2026-10-06. R-873: tests only (no live occurrence — Tester 2 stayed off). | +| D — R-876 self-repair | **done** | agent v0.145.0; 3 red-proofs; live with the operator's go: crash mid-unpack, the next pass repaired dpkg by itself and finished; bundle signed to all three boxes | +| E — release, golden, records | **done** | controller v0.295.0, agent v0.145.0, hub v0.134.0 (one each); golden 0.295.0 baked + vouched (gate OK); floors for demo-hp, demo-felhom, tester-1 | + +## 2. Claims in the brief that turned out wrong (or only partly true) + +1. **"A controller start is the right trigger"** — not alone. A host **resume** is needed too: Go's timers run on + CLOCK_MONOTONIC, which does not count suspended time, so on a suspended laptop the 02:30 timer fires hours late — and + would start the app-update leg at noon. Built: a late daily timer (> 60 min) is skipped, and a resume watch triggers + the catch-up. *Reasoned and unit-tested; NOT measured — no suspend was allowed on a demo box.* +2. **"The whole-guest backup cannot collide with the catch-up"** — it CAN: its 48 h safety valve fires on the first + 5-minute poll after a start, and the catch-up's database dump needs the apps up. Built: each waits for the other. +3. **"The box keeps a record of when it was on"** — TRUE, and it was not designed as one: the controller's system-metrics + table (one sample a minute, kept 30 days). Used for "off at 02:30" and the suggestion. +4. **"The repair step can check the journal without losing R-845's speed"** — TRUE: `--audit` and the journal are read + in ONE `sh -c` call; a clean pass still costs one call (pinned by a test). +5. **"Apps keep running during a catch-up"** (my morning STATUS said "to be measured") — the dump leg stops an app with a + volume for its copy: **measured 1 s** for opengist in the day (R-878). +6. R-873's brief offered "only when the box was off for more than 24 h" — rejected (a real outage would reach the + household a day late); the weekly rule was chosen (decision 116). + +## 3. What was proven, with numbers + +- **Part A (live).** 9202: W 09:35, off 07:33–07:38 UTC, start 07:38:09 → `missed [db-dump] … ONE catch-up in 15m0s` → + 07:53:09 dump, done in 21 s. Host crash at 07:56 interrupted the next wait → new start 07:58:35 → ONE catch-up of all + three legs at 08:13:35, done in 22 s. demo-felhom: W 10:07, controller parked + stopped 08:05–08:10 (apps running) → + catch-up 08:25:02, dump in 1 s; hub event `backup_catchup_done` "Kimaradt mentés pótolva: a doboz ki volt kapcsolva + 10:07-kor, a mentés most elkészült." — no mail. Window set back to 02:30. +- **Part B.** `nightchain.ComputeBanner` tests (shown / fresh / upgrade day / dismissed then back after a new miss / gone + after a success / off-site counts / no pattern / usually on); web tests through ServeHTTP and the real dismiss route. +- **Part C.** R-872 Tester 2 shape → both alarms; a dump 20 h ago → quiet; a new box → quiet. R-873: night 1 household + 2 mails, night 2 household 0 / operator 2, after 8 days household again. R-874 live: demo-felhom agent start + 07:38:46 → first check 08:08:46 → restore-test passed in 29 s. +- **Part D (live, operator's go).** 13 packages rolled back; crash 07:56:03.308 UTC during + `dpkg --force-confold … --unpack`; new boot 07:56:41 (38 s), guard armed, 1 unclean boot in window; at boot + `--audit` clean, journal 1 file, 1 package new / 12 old; next pass (nobody touched dpkg): `REPAIR configured=0 + journal=1` → `PLAN upgrade=12` → `DONE rc=0 upgraded=12`, healthy; package list identical (279 lines). No mail. + +## 4. Rows + +Register before **340**, after **336**. Closed (6): R-871, R-873, R-874, R-875, R-876, R-877. Narrowed: R-872 (fixed, +dated check 2026-10-06). Opened (2): **R-877** (the Tester 1 VM had no start-on-boot; the morning crash left it off +1 h 17 min — filed and closed), **R-878** (P4, a daytime catch-up stops a volume app for its dump). STATUS updated. + +## 5. Slips of mine, said plainly + +- **This morning's report said every box was healthy; the Tester 1 VM had been off since the 06:14 crash** (R-877). Found + at 07:31 from its stale report time. +- **The hub image 0.134.0 was first built from a commit that was not yet pushed** (my unstaged doc edits blocked the + `git pull`; the push was then refused by the golden gate). I pushed after the golden bake and REBUILT the image from + the pushed commit `1b0678fa`; the deployed image is the rebuilt one (`sha256:f823b10e…`). +- Two of my red-proofs did not convict at first (a masked mutation, a test that matched another call); both tests were + strengthened and the red-proofs re-run — recorded in the red-proof files. + +## 6. Teardown, three layers + +- **Machines:** 9202: controller 0.295.0 (by hand, allowed there), window back to 02:30, its ledger was set by hand for + the banner and has since been overwritten by real catch-up runs. demo-felhom: window 02:30 again, controller unparked. + demo-hp: every package current, list identical. Tester 1: running, start-on-boot set. Bake VM: CT 9100 destroyed, + token shredded, `virgin`, qemu gone. +- **Hosts:** the park file on felhom-pve removed; nothing provisioned. +- **Hub:** floors 0.295.0 (MinAgent 0.131.0) for demo-hp, demo-felhom, tester-1; artifacts vouched (agent 0.145.0, + golden 0.295.0); 6 signed jobs (agent_update and agent_config_update for each of the three boxes); hub 0.134.0 deployed + (ArgoCD Synced/Healthy). Tester 2: read only, offline all session, nothing sent. diff --git a/STATUS.md b/STATUS.md index f423ce6f..d4ee7fd6 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,13 +1,48 @@ # STATUS — what works, what's broken, what's next -**Ready for the first real tester (Tester-2): yes. Tester 2 is a laptop that is switched off at night (your word, -2026-10-05) — it was offline all session; nothing was sent to it.** +**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) stayed offline all day; nothing +was sent to it.** -**Updated 2026-10-05 (day, the night's fixes): every box of ours healthy. Fixed and proven live: the off-site clean-up -now deletes old copies (both demo boxes), a new box's first app install, the update's disk-space check, a killed -update's lost report. The power cut in the middle of an update was tested on demo-hp with your go: the box came back by -itself in 37 s, but the next update failed until I ran one command by hand — filed (R-876), fix next session. -One decision for you below (a box that is off at night). Report: `REPORT-night-fixes-2026-10-05.md`.** +**Updated 2026-10-05 (afternoon, the catch-up session): every box of ours healthy. Built and proven live: a box that was +off at its backup time makes the backups up once when it comes back; the household's banner; the OS update repairs +itself after a power cut (second crash on demo-hp, with your go). Report: `REPORT-catchup-2026-10-05.md`.** + +## Today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut + +**Decisions I took myself (you may reverse each — `09` decisions 112–118):** +- The make-up run starts 15 minutes after the box comes back; a backup due within 30 minutes is left to its normal time. +- A laptop that sleeps through the night: a nightly job that wakes up more than an hour late is skipped (otherwise app + updates would start at noon); the backups are made up instead. Tested, not measured (I may not suspend a box). +- The banner appears when the last backup is over 26 hours old; it suggests the latest evening hour the box is usually + on (5 of the last 7 days), and nothing when the box is usually on at its backup time. +- A box that is off at the 05:00 check now raises the missed-backup alarm after 2 nights without a database backup (3 + without a whole-box backup) — not after one, so a box that broke last night gives only its "offline" alarm. +- The household hears "your server cannot be reached" at most once a week; you still hear every one. +- The restore-test's first check is 30 minutes after the agent starts (a box on for short times now gets tested). +- The update's repair step now also looks at dpkg's journal — the place the power cut left its mark. + +**What works now (proven live):** +- **A missed night is made up once** (your choice A): the scratch box and demo-felhom were off across their backup time; + 15 minutes after they came back, the missed backups ran by themselves (seconds). The household's timeline got one line: + "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült." No mail. +- **The banner** (your idea) appeared on the scratch box, with the "change the backup time" button; "Close" kept it + closed. *No screenshot: there is no browser on DooPlex; I captured the page as the box served it.* +- **The power cut, again** (demo-hp, your go): back by itself in 38 seconds; **the next update repaired dpkg by itself + and finished — nobody typed anything.** No mail. +- **A restore-test 30 minutes after an agent start** ran and passed on demo-felhom. +- Controller 0.295.0, agent 0.145.0 (+ its root files) on demo-hp, demo-felhom and Tester 1; hub 0.134.0; new-install + image 0.295.0 baked and approved. + +**Found today:** +- **My slip from this morning:** the Tester 1 test machine did not restart after the morning crash and stayed off for + 1 h 17 min; my morning report said every box was healthy. It now starts by itself after a crash (proven by the second + crash). +- **What the household may notice from a make-up run:** the database backup stops an app with stored files for its copy — + 1 second for opengist. The night does the same unseen; a big app may take longer, in the day (filed, small). + +**Needs you:** nothing urgent. **Tester 2's one-time step** is unchanged (below). The new missed-backup alarm will be +checked at tomorrow's 05:00 run; if Tester 2 is still off, you will get its first real "backup missed" mail — that is +the fix working, not a new fault. ## Today (2026-10-05, day): the night's fixes, the power cut by day, a box that is off at night @@ -38,7 +73,7 @@ One decision for you below (a box that is off at night). Report: `REPORT-night-f **Needs you:** -1. **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its +1. **[DECIDED 2026-10-05 08:42 — option A, built the same afternoon; see above]** **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its nightly database backups, second copy or off-site copy; the whole-box backup runs only about every 2 days; no alarm says so; the household is mailed "your server cannot be reached" every night. No design document covers it (R-871). diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 0d5d5609..538cc62a 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -214,6 +214,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Always-on debug rings + on-demand log-bundle pulls with TTL/custody | controller v0.116, agent v0.83, hub v0.46 | **PROVEN-LIVE** | debug rings live-exercised `CAMPAIGN-3` fix-6 (1000-cap ring, ~55min horizon under load) | The **log-bundle-pull TTL/custody** half is changelog-only (no dedicated observability audit doc); ring persistence across restart is a known gap | | Operator alerting (Healthchecks → monitoring@felhom.eu) | k3s, Resend | **IMPLEMENTED** | operator infra, stated in production since 02-04; no corpus validation doc | | Backup-deadline alerting (`expected_backup_missed`) is ANCHORED — absence of signal is UNKNOWN, not failure | hub v0.75.0 | **IMPLEMENTED** | `audits/DIAG-backup-missed-2026-07-26.md` + red-proofs A/B/C + replay of the real 2026-07-26 03:00 reports (all three silent) | **No row status flips** — this signal had a FALSE-POSITIVE class (three instances: hub v0.12.0, v0.73.0, R-81), now anchored at first contact and read across retained host-report history. Still unit-proven only, not live-fired at a real deadline. The *underlying* PBS/offsite-DR tier gap it exposed is → R-82. | Per the status enum, no citation → not PROVEN-LIVE. Demoted pending an operator-cited live alert (candidate re-upgrade — see REPORT) | +| A box that is not always on (off at its backup time) makes the missed night up once when it comes back; the household sees a banner; the operator gets the missed-backup alarm, down or not | controller v0.295.0, hub v0.134.0, agent v0.145.0 | **PARTIAL — the catch-up and the banner PROVEN-LIVE (9202 + demo-felhom; banner on 9202 by a hand-set ledger); R-872's alarm proven by tests, live at the next 05:00 (dated check); a host SUSPEND reasoned + unit-tested, not measured** | `audits/catchup-2026-10-05/partA/`, `partB/`; design `07` §6.1.1 | R-872 live; a suspend never measured | ## G. Fleet & operator (hub) @@ -232,7 +233,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) | | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | -| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/` | +| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | | **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index b8e3bc7f..da44dfe1 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -396,6 +396,80 @@ warning that should precede such a failure is R-685. `pct config 9201`, both hosts). So the whole-guest tiers carry the guest and **none of the customer's data drives** — 916 GB on demo-felhom, 938 GB on demo-hp. +### 6.1.1 A box that is not always on — the catch-up, the banner, the alarms (R-871..R-874) + +> **Why this section exists (R-871):** until 2026-10-05 no architecture document covered a box that is OFF at its +> window W. Tester 2 is a laptop switched off at night (operator, 2026-10-05). Measured (`audits/night-fixes-2026-10-05/ +> partF/FINDINGS.md`, and live on 9202 `audits/catchup-2026-10-05/partA-spike/`): the controller's daily jobs always +> schedule the NEXT future time, so a missed 02:30/03:30/04:15 waited for the next night for ever; no alarm fired +> while the box was down at 05:00; the household was mailed "your server cannot be reached" every night. + +**[DESIGN — ruled 2026-10-05, `09` decision 109] A missed night runs ONCE when the box comes back.** Built controller +v0.295.0 (`internal/nightchain`): + +- **What a box does when it was off at W.** On a controller START and on a host RESUME, the controller asks its + night ledger (`/night-ledger.json`) which backup legs missed their last scheduled time. The ledger records when + each leg last RAN TO ITS END — an attempt record, used only to decide "missed", never as evidence a backup exists + (`R-100`'s rule). A leg that ran and failed was NOT missed: failures have their own alarms. +- **What the catch-up runs:** the three BACKUP legs, in the night's order — database dump, second copy, off-site copy + — through the SAME wrapped leg bodies the scheduled jobs run (`main.go` `withLeg`, `catchUpLegs`). +- **What it does not run:** the app-update leg and every Docker step (they restart apps; they wait for a real night). + The OS fast lane is the agent's and follows a whole-guest backup, not the catch-up (below). +- **When:** 15 minutes after the trigger (apps settle; a box switched on and off again at once does nothing). A leg + whose own next scheduled time is under 30 minutes away is left to its normal run. +- **Only once:** several missed nights = one catch-up (the question is about the LAST scheduled time). A normal night + followed by a daytime restart = none. A power cut in the middle of the chain = only the legs that did not end. +- **Never two at once:** every leg, scheduled or catch-up, holds one lock; a second trigger while one is pending does + nothing. +- **The whole-guest backup.** It keeps its own agent-side behaviour (inside [W+2h, W+6h), or the 48 h safety valve — + about one every 2 days for an evening-only box). **It CAN collide with a catch-up** (the valve fires on the first + 5-minute poll after a start), so each waits for the other: a scheduled quiesce defers while a catch-up runs + (`quiesce` `SetCatchUpFn`), and a catch-up waits up to 2 h while a quiesce holds the apps. +- **A suspended host** (a laptop lid): Go's timers run on CLOCK_MONOTONIC, which does not count suspended time, so + the 02:30 timer would fire hours late — and would start the app-update leg at noon. Two rules: a daily job whose + timer fires more than 60 minutes after its wall-clock time is SKIPPED (`scheduler.DailyLateLimit`), and a resume + watch (wall clock vs monotonic clock, checked every minute) triggers the catch-up. *Reasoned from the Go and Linux + clock semantics and unit-tested with injected clocks; not measured live (no suspend was allowed on a demo box).* +- **The first start on this release** seeds the ledger: nothing before it counts as missed (no surprise catch-up on + every box at the upgrade). +- **What the household sees:** one timeline line — "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 02:30-kor, a + mentés most elkészült." / "Missed backup made now: the box was off at 02:30, so the backup ran when it came back on." + (event `backup_catchup_done`, info: recorded, never mailed). + +**[DESIGN — ruled 2026-10-05, `09` decision 110] The banner (the operator's idea).** On every page of a logged-in +household: when the last daily backup (the database dump, and the off-site copy when configured) is over **26 h** old — +when the last one was, "the box was off at backup time (02:30)" when the box's own record says so, and a suggested +time. The record is the controller's system-metrics table (one sample a minute, kept 30 days): "on" at W = a sample +within 5 minutes of it; "usually on" in an hour = on in it on at least 5 of the last 7 days. It **never changes the +time** — a button opens the backup-time setting. The household can close it: it stays closed until the NEXT missed +backup time (durable, in the ledger) and disappears by itself after a successful night. + +*Decided by CC — operator may reverse* (`09` decisions 112–116): the 15-minute delay and 30-minute leave-to-normal +line; the 60-minute late-fire limit; 26 h as the banner's line (one night + the chain's two hours, = the hub's +`backupStaleAfter`); the suggestion = the LATEST hour H with H, H+1, H+2 usually on (a later evening disturbs least), +none when the current window is already usually on; the catch-up's line on the household's timeline but no mail. + +**The operator's side (hub v0.134.0, `08` §6.4).** R-872: a box DOWN at the 05:00 deadline is no longer skipped — it +is judged on longer lines (48 h without a dump, 72 h without a whole-guest backup), so a box that died last night +raises only its staleness alarm, and a box off at every deadline raises `expected_dbdump_missed` / +`expected_backup_missed`. R-873: the household hears "your server cannot be reached" at most once per 7 days (the +operator still gets every edge). R-874 (agent v0.145.0): the restore-test's first due-check runs 30 minutes after the +agent starts, so a box with short power-on sessions is still restore-tested. + +**[FACT, 2026-10-05] Proven live** (`audits/catchup-2026-10-05/partA/`): 9202 off across 09:35, on at 07:38 UTC → the +catch-up at 07:53:09 made the dump (21 s); a crash of its host mid-wait → at the next start ONE new catch-up made all +three legs (22 s); demo-felhom's controller off across 10:07 → the dump made at 08:25:03, 15 min after the start, and +the household's timeline line reached the hub. **What the household may notice:** the dump leg stops an app with a +volume for its copy — measured 1 s for opengist (182.5 KB), the same as at night; a large volume takes longer, in the +day (R-878). + +**Edge cases, stated.** A box on only in the day with W at night: the catch-up makes the backups every day it is +switched on (after 15 min); the banner suggests an evening time once a pattern exists. A box switched on during the +chain: legs already past are made up, legs still ahead run normally. A box whose controller restarts in the day after +a normal night: nothing. A box off for weeks: one catch-up when it returns; the hub's staleness alarm has been +running all along. An upgrade from an older release: the ledger is seeded, the first missed night after it is the +first one made up. + ### 6.2 Coverage per app class — and an unresolved count **[FACT]** Of 53 catalog templates, **52 keep data in Docker named volumes**; exactly **13** carry a diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index b86897ac..f13b7dae 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -340,6 +340,26 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by --- +## 6.4 A box that is not always on: the missed-backup deadline and the household's outage mail [DESIGN, hub v0.134.0, 2026-10-05] + +Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *operator may reverse*). + +- **R-872 — a box DOWN at the 05:00 deadline is judged, not skipped.** The check used to skip every customer whose node + is `down` ("they already have staleness events"), so a box down at EVERY deadline — a laptop off at night — was + never judged (measured 2026-10-05: `1 skipped (down)`). Now a down box is judged on longer lines: + `expected_dbdump_missed` after **48 h** without a `db_dump_completed`, `expected_backup_missed` after **72 h** + without a whole-guest backup in any retained host report — never for a box first seen less than 48 h ago. A box that + died last night still raises only its staleness alarm. A `disabled` box is still skipped (R-321). Pinned by + `monitor/r872_down_box_test.go`. +- **R-873 — "your server cannot be reached" reaches the HOUSEHOLD at most once per 7 days** (`node_stale`, + `node_down`, `host_stale`, `host_down`, read from the persisted notification log, so a hub restart does not reset + it). The operator still gets every edge. The recovery mail stays paired with a down mail the household actually + received (§6.2's pairing), so a held-back down mail also holds back its recovery. Pinned by + `notify/r873_liveness_weekly_test.go`. +- `backup_catchup_done` (info) is the box's "missed backup made now" line: recorded, never mailed. + +--- + ## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0] **"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.** diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index fd356bea..d8f2491b 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -815,6 +815,44 @@ its length, and both fixes cost something the household would notice — operato 1.30.0, agent 0.142.0 and the re-made golden; it lacks only agent 0.142.1's wrapper fix (R-858). *Operator ruling 2026-10-04 ~18:49.* +### 2026-10-05 (08:42) — three operator rulings (recorded before the work; the catch-up brief) + +109. **R-871, option A: a missed night runs once when the box comes back.** A box that was off during its backup time + makes the missed backups soon after it is switched on; updates that restart apps still wait for a real night. + **Rejected:** B — "the box must stay on at night", with an alarm only. *Operator ruling 2026-10-05.* +110. **The household is told on its dashboard** (the operator's idea): a banner the household can close, shown when a + daily backup was missed — when the last backup was, that the box was off during the backup time, and a suggested + different time. It never changes the time by itself. *Operator ruling 2026-10-05.* +111. **Tester 2: read only.** If it is online, CC re-signs its agent update (decision 97, act 1) and reads it back. + Nothing else. *Operator ruling 2026-10-05.* + +### 2026-10-05 (day, catch-up brief) — decided by CC — operator may reverse + +112. **When the catch-up runs (R-871).** Options: (a) at once after the box comes back — a box switched on and off + again starts a dump each time; (b) 15 minutes later, re-checking first. **Chosen (b)**, as the brief proposed; + a leg due within 30 minutes is left to its normal run. `07` §6.1.1. +113. **What a host suspend does to the night (R-871).** Options: (a) only a controller start triggers the catch-up — + a suspended laptop's 02:30 timer then fires hours late (Go timers run on CLOCK_MONOTONIC) and starts the + app-update leg at noon; (b) skip any daily job that fires over 60 minutes late and trigger the catch-up from a + resume watch. **Chosen (b).** Reasoned and unit-tested with injected clocks; not measured live. +114. **The banner's line and its suggestion (decision 110).** Options for the line: 24 h (a slow night flickers it), + 26 h (one night + the chain's two hours, = the hub's `backupStaleAfter`), 48 h (two missed nights before a word). + **Chosen 26 h.** Suggestion: the latest hour H with H, H+1, H+2 usually on (on ≥ 5 of the last 7 days); none + when the current window is already usually on (a new time would not help). +115. **R-872's lines for a box that is down at the deadline.** Options: (a) judge a down box like an up one — a box + that died last night raises a backup alarm on top of its staleness alarm; (b) 48 h without a dump / 72 h + without a whole-guest backup (the agent's own 48 h valve + a day). **Chosen (b).** hub v0.134.0, `08` §6.4. +116. **R-873's rule.** Options: (a) household liveness mail only after 24 h down — a real outage reaches the household + a day late; (b) at most once per 7 days to the household, the operator every edge. **Chosen (b)** (the brief's + first example): the first outage of a week still reaches the household at once. hub v0.134.0, `08` §6.4. +117. **R-874's first restore-test check after start.** Options: at start (a crash loop hammers a failing tier — the + earned restraint), 30 minutes after start, or keep one interval (6 h — a short-session box never reaches it). + **Chosen 30 minutes.** agent v0.145.0. +118. **R-876's repair trigger.** Options: (a) always run `dpkg --configure -a` (R-845's speed lost on every clean + pass); (b) read `--audit` AND the update journal in ONE `sh -c` call, repair when either shows something, and as + a belt repair + retry once when apt itself says "dpkg was interrupted". **Chosen (b)**: a clean pass still costs + one call (pinned by a test). agent v0.145.0. + ### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief) 100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) — diff --git a/documentation/architecture/11-os-updates.md b/documentation/architecture/11-os-updates.md index 32a9e95a..3802b72f 100644 --- a/documentation/architecture/11-os-updates.md +++ b/documentation/architecture/11-os-updates.md @@ -639,6 +639,19 @@ Agent v0.144.0 + v0.144.1 (wrapper and agent; `09` decisions 106–108). Evidenc (runbook `crash-guard.md`), then the next pass installed the 12. Operator mail `os_update_failed` (true); no household mail; the household's timeline showed "Controller elindult". +### 8.5 The self-repair after a power cut as BUILT (2026-10-05, agent v0.145.0) `[FACT]` + +- **R-876 fixed.** The wrapper reads `dpkg --audit` AND dpkg's update journal (`/var/lib/dpkg/updates/`) in ONE call + and repairs when either shows something; as a belt, when apt itself says "dpkg was interrupted", it repairs and + retries once (`09` decision 118). A clean pass still costs one call (R-845's speed, pinned by a test). +- **Proven live** (operator's go, demo-hp, `audits/catchup-2026-10-05/partD/`): 13 guest packages rolled back, crash + at 07:56:03 UTC during dpkg's unpack, back by itself (new boot 07:56:41, guard armed, 1 unclean boot in the window); + at boot `--audit` clean, journal 1 file, the same shape as §8.4. **The next pass, with nobody touching the box:** + `REPAIR configured=0 journal=1`, then `DONE rc=0 upgraded=12`, healthy; the guest's package list equals the one + before the rollback. No mail (the pass did not fail); the household's timeline: "Controller elindult". +- **R-874.** The agent's restore-test due-check runs 30 minutes after start (then every interval), so a box with short + power-on sessions is restore-tested (`09` decision 117). + ## 9. Where the rest lives - The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`). diff --git a/documentation/audits/catchup-2026-10-05/final-state.txt b/documentation/audits/catchup-2026-10-05/final-state.txt new file mode 100644 index 00000000..b9a7fdf8 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/final-state.txt @@ -0,0 +1,9 @@ +08:27:24 +('Tester-2-be8404', '0.142.0', '2026-10-04 18:05:48', None) +('demo-felhom-8363b5', '0.145.0', '2026-10-05 08:23:56', '0.145.0') +('demo-hp-bb76ea', '0.145.0', '2026-10-05 08:27:06', '0.145.0') +('tester-1-d70be4', '0.145.0', '2026-10-05 08:27:31', '0.145.0') +== demo-hp: agent felhom-agent 0.145.0 | controller 0.295.0 | healthy 20 of 21 | guard 1 +== felhom-pve: agent felhom-agent 0.145.0 | controller 0.295.0 | healthy 4 of 5 | guard 1 +Oct 05 10:08:58 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:58.009+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUN +9202: gitea.dooplex.hu/admin/felhom-controller:0.295.0 Up 29 minutes (healthy) diff --git a/documentation/audits/catchup-2026-10-05/partA-spike/s1-set-window.txt b/documentation/audits/catchup-2026-10-05/partA-spike/s1-set-window.txt new file mode 100644 index 00000000..2c6992fa --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA-spike/s1-set-window.txt @@ -0,0 +1,12 @@ +### 06:58:03 UTC: set 9202's window to 09:05 (Budapest) by POST /backups/window +HTTP 303 +2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 09:05 (next run 2026-10-05 09:05 CEST) +2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 10:05 (next run 2026-10-05 10:05 CEST) +2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 10:50 (next run 2026-10-05 10:50 CEST) +2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: rescheduled — recomputing next run +2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: rescheduled — recomputing next run +2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: next run at 2026-10-05 09:05:00 CEST (waiting 6m54s) +2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: next run at 2026-10-05 10:05:00 CEST (waiting 1h6m54s) +2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: rescheduled — recomputing next run +2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: next run at 2026-10-05 10:50:00 CEST (waiting 1h51m54s) +2026/10/05 06:58:05 backup_handlers.go:65: [INFO] [web] backup window set to 09:05 (legs 09:05/10:05/10:50) diff --git a/documentation/audits/catchup-2026-10-05/partA-spike/s2-off-across-W.txt b/documentation/audits/catchup-2026-10-05/partA-spike/s2-off-across-W.txt new file mode 100644 index 00000000..97cef227 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA-spike/s2-off-across-W.txt @@ -0,0 +1,9 @@ +### 07:03:01 pct shutdown 9202 +status: stopped +### 07:08:01 pct start 9202 +status: running +### 07:10:35 controller log after the start +2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:05 CEST +2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-10-05 10:05 CEST +2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-10-05 10:50 CEST +2026/10/05 07:08:08 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives diff --git a/documentation/audits/catchup-2026-10-05/partA/hub-timeline-line.txt b/documentation/audits/catchup-2026-10-05/partA/hub-timeline-line.txt new file mode 100644 index 00000000..98dfb57f --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA/hub-timeline-line.txt @@ -0,0 +1 @@ +('demo-felhom', 'backup_catchup_done', 'info', 'Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült.', 'controller', '2026-10-05 08:25:03') diff --git a/documentation/audits/catchup-2026-10-05/partA/live-9202-1-deploy.txt b/documentation/audits/catchup-2026-10-05/partA/live-9202-1-deploy.txt new file mode 100644 index 00000000..fef0d1b5 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA/live-9202-1-deploy.txt @@ -0,0 +1,12 @@ +### 07:29:47 9202: controller 0.295.0 by its bootstrap file +gitea.dooplex.hu/admin/felhom-controller:0.295.0 +gitea.dooplex.hu/admin/felhom-controller:0.295.0 started 2026-10-05T07:29:50.143207934Z +2026/10/05 07:29:50 scheduler.go:149: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:05 CEST +2026/10/05 07:29:50 nightchain.go:280: [INFO] [catch-up] controller start: no backup leg missed its last scheduled time — nothing to make up +{ + "seeded_at": "2026-10-05T07:29:50.484758959Z", + "ended": {}, + "db_dump_ok": "0001-01-01T00:00:00Z", + "last_catch_up": "0001-01-01T00:00:00Z", + "banner_dismissed_through": "0001-01-01T00:00:00Z" +} diff --git a/documentation/audits/catchup-2026-10-05/partA/live-9202-2-off-across-W.txt b/documentation/audits/catchup-2026-10-05/partA/live-9202-2-off-across-W.txt new file mode 100644 index 00000000..9c88f869 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA/live-9202-2-off-across-W.txt @@ -0,0 +1,41 @@ +### 07:33:00 W = 09:35 Budapest (07:35 UTC). pct shutdown 9202 +command 'lxc-stop -n 9202 --nokill --timeout 60' failed: exit code 1 +status: stopped +### 07:38:00 pct start 9202 +status: running +### 07:39:06 after the start +2026/10/05 07:38:09 scheduler.go:149: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:35 CEST +2026/10/05 07:38:09 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 09:35) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night) +### 07:53:51 the catch-up +2026/10/05 07:43:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health +2026/10/05 07:43:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives +2026/10/05 07:44:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan +2026/10/05 07:46:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan +2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: docker-socket-users +2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: offsite-credential-retry +2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan +2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: backup-cache +2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health +2026/10/05 07:48:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives +2026/10/05 07:50:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan +2026/10/05 07:52:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan +2026/10/05 07:53:09 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg +2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: offsite-credential-retry +2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: docker-socket-users +2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: backup-cache +2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health +2026/10/05 07:53:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives +2026/10/05 07:53:09 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (412.6 KB, 436ms, 72 tables) +2026/10/05 07:53:30 nightchain.go:340: [INFO] [catch-up] done: db-dump in 21s (missed at 2026-10-05 09:35) +{ + "seeded_at": "2026-10-05T07:29:50.484758959Z", + "ended": { + "db-dump": "2026-10-05T07:53:30.483916689Z" + }, + "db_dump_ok": "2026-10-05T07:53:30.484211296Z", + "last_catch_up": "2026-10-05T07:53:30.484384683Z", + "banner_dismissed_through": "0001-01-01T00:00:00Z" +} +2026/10/05 07:38:09 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 09:35) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night) +2026/10/05 07:53:09 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg +2026/10/05 07:53:30 nightchain.go:340: [INFO] [catch-up] done: db-dump in 21s (missed at 2026-10-05 09:35) diff --git a/documentation/audits/catchup-2026-10-05/partA/live-9202-3-crash-interrupted-catchup.txt b/documentation/audits/catchup-2026-10-05/partA/live-9202-3-crash-interrupted-catchup.txt new file mode 100644 index 00000000..39848676 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA/live-9202-3-crash-interrupted-catchup.txt @@ -0,0 +1,23 @@ +2026/10/05 07:58:35 main.go:340: [INFO] felhom-controller 0.295.0 starting (customer: demo-hp, domain: enkisfelhom.hu) +2026/10/05 07:58:35 scheduler.go:149: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-10-06 03:30 CEST +2026/10/05 07:58:35 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump tier2 offsite] (last scheduled 2026-10-05 04:15) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night) +2026/10/05 08:13:35 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg +2026/10/05 08:13:36 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (413.8 KB, 423ms, 72 tables) +2026/10/05 08:13:57 nightchain.go:332: [INFO] [catch-up] running the missed tier2 leg +2026/10/05 08:13:57 tier2.go:425: [INFO] [backup] Tier 2 copied paperless-ngx → /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx (82.5 MB, 1 leg(s), 0s) [SSD: state-only] +2026/10/05 08:13:57 tier2.go:425: [INFO] [backup] Tier 2 copied privatebin → /mnt/felhom-drives/scratch_hdd/backups/secondary/privatebin (20.3 KB, 0 leg(s), 0s) +2026/10/05 08:13:57 tier2.go:476: [INFO] [backup] Tier 2 run complete: 2 app(s) processed (incl. volume-only — F6) +2026/10/05 08:13:57 nightchain.go:332: [INFO] [catch-up] running the missed offsite leg +2026/10/05 08:13:57 main.go:1385: [INFO] [offbox] no scheduled off-site target on this box — the off-site leg does nothing +2026/10/05 08:13:57 nightchain.go:340: [INFO] [catch-up] done: db-dump, offsite, tier2 in 22s (missed at 2026-10-05 04:15) +{ + "seeded_at": "2026-09-30T07:54:19.875183Z", + "ended": { + "db-dump": "2026-10-05T08:13:57.406392075Z", + "offsite": "2026-10-05T08:13:57.69370068Z", + "tier2": "2026-10-05T08:13:57.693478189Z" + }, + "db_dump_ok": "2026-10-05T08:13:57.406683255Z", + "last_catch_up": "2026-10-05T08:13:57.693858077Z", + "banner_dismissed_through": "2026-10-05T00:30:00Z" +} diff --git a/documentation/audits/catchup-2026-10-05/partA/live-demo-felhom.txt b/documentation/audits/catchup-2026-10-05/partA/live-demo-felhom.txt new file mode 100644 index 00000000..511d5658 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA/live-demo-felhom.txt @@ -0,0 +1,24 @@ +### 08:05:01 W = 10:07 Budapest (08:07 UTC). park + stop the controller (apps keep running) +felhom-controller +4 +### 08:10:00 start the controller + unpark +felhom-controller +2026/10/05 08:10:02 [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 10:07 CEST +### 08:25:04 the catch-up +{ + "seeded_at": "2026-10-05T07:29:41.65020944Z", + "ended": { + "db-dump": "2026-10-05T08:25:03.369812125Z" + }, + "db_dump_ok": "2026-10-05T08:25:03.370056355Z", + "last_catch_up": "2026-10-05T08:25:03.370196251Z", + "banner_dismissed_through": "0001-01-01T00:00:00Z" +}### 08:25:19 the catch-up lines (this box logs without file:line) +2026/10/05 08:10:02 [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 10:07) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night) +2026/10/05 08:25:02 [INFO] [catch-up] running the missed db-dump leg +2026/10/05 08:25:03 [INFO] [catch-up] done: db-dump in 1s (missed at 2026-10-05 10:07) +### 08:25:20 window back to 02:30 +HTTP 303 +2026/10/05 08:25:22 [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15) +ls: cannot access '/var/lib/felhom-agent/guests/9201/controller-parked': No such file or directory +4 diff --git a/documentation/audits/catchup-2026-10-05/partA/red-proofs.txt b/documentation/audits/catchup-2026-10-05/partA/red-proofs.txt new file mode 100644 index 00000000..2b678ed9 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partA/red-proofs.txt @@ -0,0 +1,120 @@ +### RED-PROOF scheduler late-fire guard dropped + scheduler_test.go:128: a 6-hour-late fire ran the job (the app-update leg would start at noon) +--- FAIL: TestDaily_LateFireIsSkipped (0.00s) + --- PASS: TestDaily_LateFireIsSkipped/on_time (0.00s) + --- FAIL: TestDaily_LateFireIsSkipped/after_a_suspend (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.004s +FAIL +### restored +ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.456s +### RED-PROOF 1 Missed returns nothing (the old behaviour: no catch-up) + nightchain_test.go:99: no catch-up was scheduled for a box off across W (missed []) +--- FAIL: TestCatchUp_OffAtWThenOn_OneCatchUpAfter15Min (0.00s) + nightchain_test.go:135: the catch-up never finished +--- FAIL: TestCatchUp_TwoMissedNights_One (3.00s) + nightchain_test.go:151: the catch-up never finished +--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (3.00s) + nightchain_test.go:178: missed [], want tier2 + offsite of the night of the 4th +--- FAIL: TestCatchUp_LegAboutToRun_LeftToNormal (0.00s) + nightchain_test.go:189: the catch-up never finished +--- FAIL: TestCatchUp_NeverRunsAnythingButBackupLegs (3.00s) + nightchain_test.go:206: the catch-up never finished +--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (3.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 12.015s +FAIL +### RED-PROOF 2 a leg that ran is not checked (Ended ignored) + nightchain_test.go:121: a normal night was made up again: [db-dump tier2 offsite] +--- FAIL: TestCatchUp_NightRan_DaytimeRestart_None (0.00s) + nightchain_test.go:140: after the catch-up nothing is missed any more, got [db-dump tier2 offsite] +--- FAIL: TestCatchUp_TwoMissedNights_One (0.00s) + nightchain_test.go:153: ran [db-dump tier2 offsite], want only tier2 + offsite +--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s +FAIL +### RED-PROOF 3 the 15-minute delay dropped + nightchain_test.go:103: the catch-up must wait 15m0s first, slept [] +--- FAIL: TestCatchUp_OffAtWThenOn_OneCatchUpAfter15Min (0.00s) + nightchain_test.go:208: slept [1m0s 1m0s 1m0s] — want the 15 min delay then 3 one-minute waits for the quiesce +--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s +FAIL +### RED-PROOF 4 a pending catch-up does not block a second one + nightchain_test.go:133: a second catch-up was scheduled while one was pending: [db-dump tier2 offsite] +--- FAIL: TestCatchUp_TwoMissedNights_One (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s +FAIL +### RED-PROOF 5 the quiesce wait dropped + nightchain_test.go:208: slept [15m0s] — want the 15 min delay then 3 one-minute waits for the quiesce +--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s +FAIL +### RED-PROOF 6 the leave-to-normal rule dropped + nightchain_test.go:174: the dump due in 20 minutes was put in the catch-up: [db-dump tier2 offsite] +--- FAIL: TestCatchUp_LegAboutToRun_LeftToNormal (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s +FAIL +### RED-PROOF 8 a fresh ledger counts history as missed + nightchain_test.go:161: a freshly seeded ledger made up [db-dump tier2 offsite] +--- FAIL: TestCatchUp_FreshLedger_None (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.007s +FAIL +### restored +ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s +### RED-PROOF 7 the resume watch compares wall with wall (compiles) + nightchain_test.go:241: a 7-hour suspend was not seen +--- FAIL: TestResumeWatch_SeesASuspend (2.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 2.010s +FAIL +### RED-PROOF 9 the catch-up runs whatever Legs holds (not only the missed chain legs) + nightchain_test.go:153: ran [db-dump tier2 offsite], want only tier2 + offsite +--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (0.00s) + nightchain_test.go:191: an app update ran inside a catch-up +--- FAIL: TestCatchUp_NeverRunsAnythingButBackupLegs (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s +FAIL +### restored +ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s +### RED-PROOF 10 the whole-guest cycle does not wait for a running catch-up + r871_catchup_test.go:21: the backup started while a catch-up ran: start=1 stopped=[nextcloud] +--- FAIL: TestR871_ScheduledCycleWaitsForCatchUp (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.005s +FAIL +### restored +ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.005s +### RED-PROOF 12 the update leg inside the shared off-site body +--- PASS: TestR871_CatchUpWiring (0.01s) + r871_catchup_wiring_test.go:44: the update leg is inside the off-site body the catch-up runs +--- FAIL: TestR871_CatchUpRunsNoUpdateLeg (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.019s +FAIL +### restored +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.019s +### RED-PROOF 11 the start trigger not wired (test now keyed on the argument) + r871_catchup_wiring_test.go:37: the catch-up's START trigger (Evaluate(ctx, "controller start")) is called 0 times, want 1 +--- FAIL: TestR871_CatchUpWiring (0.02s) +--- PASS: TestR871_CatchUpRunsNoUpdateLeg (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.031s +FAIL +### restored +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.026s +### RED-PROOF hub: backup_catchup_done not allowed + r871_catchup_event_test.go:10: backup_catchup_done is not an allowed event type — the box's catch-up line would be dropped with a 400 +--- FAIL: TestR871_CatchUpEventIsAllowed (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/api 0.020s +FAIL +### restored +ok gitea.dooplex.hu/admin/felhom-hub/internal/api 0.019s diff --git a/documentation/audits/catchup-2026-10-05/partB/live-banner-9202.txt b/documentation/audits/catchup-2026-10-05/partB/live-banner-9202.txt new file mode 100644 index 00000000..444aa462 --- /dev/null +++ b/documentation/audits/catchup-2026-10-05/partB/live-banner-9202.txt @@ -0,0 +1,27 @@ +### 07:54:16 banner on 9202 — the ledger set by hand to a box whose last dump is 3 days old (scratch fixture, said so); window back to 02:30 +HTTP 303 +ledger set: last dump 2026-10-02T07:54:19.875183Z +2026/10/05 07:54:20 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump tier2 offsite] (last scheduled 2026-10-05 04:15) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night) +--- the banner as GET /launcher served it (9202, Hungarian household): +
+
+ + A legutóbbi mentés 3 nappal ezelőtt készült. Válassz egy olyan időpontot, amikor a doboz általában be van kapcsolva. + + Mentési idő módosítása +
+ + +
+
+
+
P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | -## Box system & updates — 20 rows (P2 4, P3 13, P4 3) +## Box system & updates — 18 rows (P2 3, P3 13, P4 2) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -330,10 +329,8 @@ stopping line that lies. | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | | **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | -| **R-875** | Box system & updates | P4 | **A kept OS report sent after the hub was away says "sent after the agent stopped mid-pass (R-868)" — wrong for that case: the agent did not stop, the hub was unreachable.** MEASURED 2026-10-05 05:53 UTC on demo-felhom (agent v0.144.1): the three reports of the R-866 hub-away pass reached the hub at the new agent's start with that reason text. The kept copy cannot tell the two causes apart. Fix direction: a neutral reason ("sent late — kept on the box until the hub could take it"), or the agent records which case it was. `audits/night-fixes-2026-10-05/partD/r866-kept-copies-sent-at-start.txt` | **READY — owner: CC** | — | — | CC | -| **R-876** | Box system & updates | P2 | **After a power cut in the middle of an OS update, every later OS pass FAILS until a person runs `dpkg --configure -a`: the wrapper's repair step runs only when `dpkg --audit` shows something, but a crash can leave dpkg's update journal (`/var/lib/dpkg/updates/`) non-empty with `--audit` clean — and that journal is exactly what apt refuses on.** MEASURED 2026-10-05 on demo-hp (Part E, the night's A1 by day, agent v0.144.1): crash at 06:13:55 UTC while dpkg ran; back by itself in 37 s (crash guard armed, 1 unclean boot); at boot `dpkg --audit` clean, all 13 packages `ii` (1 new, 12 old), but `/var/lib/dpkg/updates/` held 3 files; the next pass logged `REPAIR configured=0 fixed=0`, then `FAILED rc=100 step=install` with `E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a'`; operator mail `os_update_failed` (true). A by-hand `dpkg --configure -a` (rc 0) and the next pass finished (12 installed). Cause: R-845's speed-up skips the repair on a clean audit (`configs/felhom-os-apply` `repair()`). Without the fix a box that loses power mid-update fails its OS leg every night and mails the operator every night. Fix direction: also repair when `/var/lib/dpkg/updates/` is non-empty (one `ls`, no cost on a clean pass), pinned by a test with that exact state. `audits/night-fixes-2026-10-05/partE/` | **READY — owner: CC** (next agent release) | — | fix + the crash-state test | CC | -## Monitoring & notifications — 27 rows (P2 3, P3 17, P4 7) +## Monitoring & notifications — 26 rows (P2 3, P3 16, P4 7) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -362,8 +359,7 @@ stopping line that lies. | **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | | **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC | | **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator | -| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871 | — | CC | -| **R-873** | Monitoring & notifications | P3 | **A household whose box is off every night on purpose is mailed "Your server cannot be reached." every night and "reachable again" every morning; the operator gets about 6 mails a day for it.** FOUND 2026-10-05 (Part F spike): `node_down` after 90 min, mailed to the customer (Tester 2: 2026-10-04 19:36:42 UTC); the 6 h cooldown does not stop a daily repeat (`hub/internal/notify/dispatcher.go`). The R-285 gap (no notion of expected downtime) now reaching a real household. Fix direction depends on R-871: if B, a per-box "off at night" setting that quiets node_down within a nightly window; if A, the same plus the catch-up. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871, R-285 | — | CC | +| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC | ## Hub & operator — 23 rows (P2 1, P3 7, P4 15) @@ -506,4 +502,5 @@ stopping line that lies. the R-row. Duplicating them here would create the second source this design avoids. --> | item | due (UTC) | what to measure | |---|---|---| +| R-872 | 2026-10-06 | the first live 05:00 deadline run judges a down box on the longer lines (Tester 2, if still off): hub log + the two events (detail in the R-872 row) |