catch-up session 2026-10-05: design 07 §6.1.1 (a box that is not always on), 08 §6.4; rulings 109-111, CC decisions 112-118; R-871/R-873..R-877 closed, R-872 narrowed (dated check), R-878 opened; live evidence; STATUS
gates / gates (push) Successful in 33s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 10:29:22 +02:00
parent ff25f1076d
commit 9bb45eaaa2
34 changed files with 790 additions and 18 deletions
@@ -396,6 +396,80 @@ warning that should precede such a failure is R-685.
`pct config 9201`, both hosts). So the whole-guest tiers carry the guest and **none of the customer's
data drives** — 916 GB on demo-felhom, 938 GB on demo-hp.
### 6.1.1 A box that is not always on — the catch-up, the banner, the alarms (R-871..R-874)
> **Why this section exists (R-871):** until 2026-10-05 no architecture document covered a box that is OFF at its
> window W. Tester 2 is a laptop switched off at night (operator, 2026-10-05). Measured (`audits/night-fixes-2026-10-05/
> partF/FINDINGS.md`, and live on 9202 `audits/catchup-2026-10-05/partA-spike/`): the controller's daily jobs always
> schedule the NEXT future time, so a missed 02:30/03:30/04:15 waited for the next night for ever; no alarm fired
> while the box was down at 05:00; the household was mailed "your server cannot be reached" every night.
**[DESIGN — ruled 2026-10-05, `09` decision 109] A missed night runs ONCE when the box comes back.** Built controller
v0.295.0 (`internal/nightchain`):
- **What a box does when it was off at W.** On a controller START and on a host RESUME, the controller asks its
night ledger (`<data>/night-ledger.json`) which backup legs missed their last scheduled time. The ledger records when
each leg last RAN TO ITS END — an attempt record, used only to decide "missed", never as evidence a backup exists
(`R-100`'s rule). A leg that ran and failed was NOT missed: failures have their own alarms.
- **What the catch-up runs:** the three BACKUP legs, in the night's order — database dump, second copy, off-site copy
— through the SAME wrapped leg bodies the scheduled jobs run (`main.go` `withLeg`, `catchUpLegs`).
- **What it does not run:** the app-update leg and every Docker step (they restart apps; they wait for a real night).
The OS fast lane is the agent's and follows a whole-guest backup, not the catch-up (below).
- **When:** 15 minutes after the trigger (apps settle; a box switched on and off again at once does nothing). A leg
whose own next scheduled time is under 30 minutes away is left to its normal run.
- **Only once:** several missed nights = one catch-up (the question is about the LAST scheduled time). A normal night
followed by a daytime restart = none. A power cut in the middle of the chain = only the legs that did not end.
- **Never two at once:** every leg, scheduled or catch-up, holds one lock; a second trigger while one is pending does
nothing.
- **The whole-guest backup.** It keeps its own agent-side behaviour (inside [W+2h, W+6h), or the 48 h safety valve —
about one every 2 days for an evening-only box). **It CAN collide with a catch-up** (the valve fires on the first
5-minute poll after a start), so each waits for the other: a scheduled quiesce defers while a catch-up runs
(`quiesce` `SetCatchUpFn`), and a catch-up waits up to 2 h while a quiesce holds the apps.
- **A suspended host** (a laptop lid): Go's timers run on CLOCK_MONOTONIC, which does not count suspended time, so
the 02:30 timer would fire hours late — and would start the app-update leg at noon. Two rules: a daily job whose
timer fires more than 60 minutes after its wall-clock time is SKIPPED (`scheduler.DailyLateLimit`), and a resume
watch (wall clock vs monotonic clock, checked every minute) triggers the catch-up. *Reasoned from the Go and Linux
clock semantics and unit-tested with injected clocks; not measured live (no suspend was allowed on a demo box).*
- **The first start on this release** seeds the ledger: nothing before it counts as missed (no surprise catch-up on
every box at the upgrade).
- **What the household sees:** one timeline line — "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 02:30-kor, a
mentés most elkészült." / "Missed backup made now: the box was off at 02:30, so the backup ran when it came back on."
(event `backup_catchup_done`, info: recorded, never mailed).
**[DESIGN — ruled 2026-10-05, `09` decision 110] The banner (the operator's idea).** On every page of a logged-in
household: when the last daily backup (the database dump, and the off-site copy when configured) is over **26 h** old —
when the last one was, "the box was off at backup time (02:30)" when the box's own record says so, and a suggested
time. The record is the controller's system-metrics table (one sample a minute, kept 30 days): "on" at W = a sample
within 5 minutes of it; "usually on" in an hour = on in it on at least 5 of the last 7 days. It **never changes the
time** — a button opens the backup-time setting. The household can close it: it stays closed until the NEXT missed
backup time (durable, in the ledger) and disappears by itself after a successful night.
*Decided by CC — operator may reverse* (`09` decisions 112–116): the 15-minute delay and 30-minute leave-to-normal
line; the 60-minute late-fire limit; 26 h as the banner's line (one night + the chain's two hours, = the hub's
`backupStaleAfter`); the suggestion = the LATEST hour H with H, H+1, H+2 usually on (a later evening disturbs least),
none when the current window is already usually on; the catch-up's line on the household's timeline but no mail.
**The operator's side (hub v0.134.0, `08` §6.4).** R-872: a box DOWN at the 05:00 deadline is no longer skipped — it
is judged on longer lines (48 h without a dump, 72 h without a whole-guest backup), so a box that died last night
raises only its staleness alarm, and a box off at every deadline raises `expected_dbdump_missed` /
`expected_backup_missed`. R-873: the household hears "your server cannot be reached" at most once per 7 days (the
operator still gets every edge). R-874 (agent v0.145.0): the restore-test's first due-check runs 30 minutes after the
agent starts, so a box with short power-on sessions is still restore-tested.
**[FACT, 2026-10-05] Proven live** (`audits/catchup-2026-10-05/partA/`): 9202 off across 09:35, on at 07:38 UTC → the
catch-up at 07:53:09 made the dump (21 s); a crash of its host mid-wait → at the next start ONE new catch-up made all
three legs (22 s); demo-felhom's controller off across 10:07 → the dump made at 08:25:03, 15 min after the start, and
the household's timeline line reached the hub. **What the household may notice:** the dump leg stops an app with a
volume for its copy — measured 1 s for opengist (182.5 KB), the same as at night; a large volume takes longer, in the
day (R-878).
**Edge cases, stated.** A box on only in the day with W at night: the catch-up makes the backups every day it is
switched on (after 15 min); the banner suggests an evening time once a pattern exists. A box switched on during the
chain: legs already past are made up, legs still ahead run normally. A box whose controller restarts in the day after
a normal night: nothing. A box off for weeks: one catch-up when it returns; the hub's staleness alarm has been
running all along. An upgrade from an older release: the ledger is seeded, the first missed night after it is the
first one made up.
### 6.2 Coverage per app class — and an unresolved count
**[FACT]** Of 53 catalog templates, **52 keep data in Docker named volumes**; exactly **13** carry a