R-243: offsite_escrow_pending — an operator alarm when off-site is on and the escrow never done (7 days); 09 decisions 177-179; 07 R-899 note
gates / gates (push) Successful in 2m48s

Hub code unreleased; ships with tomorrow's hub release.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-08 08:00:08 +02:00
parent de4a8d20ba
commit b119301c6f
10 changed files with 382 additions and 0 deletions
@@ -405,6 +405,18 @@ read **only when the storage cannot be read**: a fresh saved copy → not due; a
unknown, as before. A storage that answers always wins, so a pruned archive still makes the tier due. Tests:
`TestBackupDue_R894_*` (felhom-agent `internal/localapi`).
**[DESIGN — operator ruling 2026-10-08, `09` §3 decision 177; built on controller and agent main the same day, ships
with their next releases — R-899] A daytime press never moves the night's backup. Every night takes its own.** The
agent's due-check reads the newest archive on the tier, and a „Mentés most" press is an archive like any other, so a
press at 08:49 made the 24 h tier due at 08:49 the next day — after the window closed — and that night had no
whole-guest backup, no OS leg and no kernel step (demo-hp, 2026-10-07 → 08). Now the controller keeps a ledger of the
last successful press and the last successful SCHEDULED run per tier (`whole-guest-ledger.json` beside the quiesce
marker); a tier the agent calls „not due" is still due on the scheduled path when a press came after the last scheduled
success and that success is older than tonight's gate opening. Such a tier waits for the window (it never fires the
safety valve); a scheduled success inside tonight's window ends it. And a press reaches the agent as
`trigger=manual`, after which the agent runs no OS leg (before, a press ran it at once, in the day). Pinned by
`TestR899_*` (controller) and `TestAfterPrimaryBackup` (agent).
**[FACT, 2026-10-04 — agent v0.140.0, `11-os-updates.md` §8.1] The OS leg closes the night.** After the
whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent
waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a
@@ -352,6 +352,7 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` |
| `offsite_escrow_pending` | warning | off-site is ON, the escrow is NOT done (`escrow_state` ≠ `escrowed`) and there has been no successful off-site run for **7 days** — counted from the last successful run, else from the first host report the hub received (an unknown anchor never fires). `offsite_stale` deliberately leaves this state out (pending is the designed onboarding state, `07` §6.1), so until hub main 2026-10-08 a box whose household never did the escrow step never backed up off-site and nothing fired (R-243). Re-sent at most once a week while true; the raise time is persisted, so a hub restart neither re-mails nor forgets (`09` §3 decision 179) | the escrow done, off-site off, or a successful run after the alarm → `offsite_escrow_pending_cleared` (info) | `monitor/offsite_escrow_pending_test.go` `TestR243_*`; `notify/r243_operator_only_test.go` |
| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` |
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
@@ -915,6 +915,24 @@ its length, and both fixes cost something the household would notice — operato
127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer.
*Operator ruling 2026-10-05.* (R-861)
### 2026-10-08 (07:12) — operator rulings on R-899 and the day's work (recorded before the work)
177. **R-899: a daytime whole-guest backup never moves the night's backup. Every night takes its own** (option A). A
household's „Mentés most" press keeps its own record but does not count for „has tonight's backup run". The 20-hour
rule was refused: it would not have caught the 2026-10-07 case (a press at 08:49, a window opening at 04:30).
*Operator ruling 2026-10-08 07:12.* Built the same day (controller `internal/quiesce/nightowed.go`; agent: no OS leg
after a press), ships with the next releases; `07` §6.1.
178. **Today (2026-10-08): progress with other work, without touching tonight's kernel night** (the second one, 8→9, on
demo-hp and demo-felhom). No change to those boxes, Tester 1 or the hub's running version; code is built, tested and
committed today and released tomorrow. *Operator ruling 2026-10-08 07:12.*
179. **R-243's line: `offsite_escrow_pending` fires after 7 days** with off-site ON, the escrow not done and no successful
off-site run. Options: 48 h (`offsite_stale`'s line — but that assumes runs are expected, and pending is the designed
onboarding state, so a slow household would be reported in its first days); 72 h (`expected_backup_missed` — the same
objection); **7 days** (`08` §6.3's line for „a box needing an action nobody took", `os_update_stale` /
`agent_behind`, and the off-site whole-guest cadence). Chosen 7 days: a household that does the step in its first week
is never reported, and one that does not is reported once a week. *Decided by CC (the brief asked CC to pick) —
operator may reverse.* `08` §6.3.
### 2026-10-07 (evening) — operator rulings on the first kernel night (recorded before the work)
174. **The household kernel mail text stays as written**, plus a reply address (a `Reply-To` that reaches the operator).