5fd2656021
gates / gates (push) Successful in 32s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
74 lines
4.7 KiB
Markdown
74 lines
4.7 KiB
Markdown
# Part F — a box that is OFF every night (Tester 2) — read-only spike, 2026-10-05
|
||
|
||
Collected by a read-only helper session (hub DB copy at ~05:11 UTC, non-secret columns only, copy shredded; hub
|
||
log; source at felhom.eu 7221ee5c, controller 69e9145, agent c8d12f1). Two claims re-checked by the main session
|
||
against source: the 05:00 deadline check skips `down` customers (`hub/internal/monitor/deadline.go` ~360), and the
|
||
controller's daily scheduler always picks the NEXT future time (`controller/internal/scheduler/scheduler.go`
|
||
`nextDailyRun`, ~389-408). Hub DB times are UTC; the hub log prints Budapest time.
|
||
|
||
## Q1 — Tester 2's record (customer `Tester-2`, host `Tester-2-be8404`)
|
||
|
||
Customer created 2026-09-30 (language en); host created 2026-10-04 16:07 UTC (agent 0.142.0, controller 0.292.0);
|
||
no apps installed. Online twice on 2026-10-04: ~15:54–17:13 UTC, and 18:03:42–18:06:24 UTC (last controller report).
|
||
Nothing since. One evening of data: no daytime pattern can be established.
|
||
|
||
| Leg | Result | Evidence |
|
||
|---|---|---|
|
||
| Whole-guest backup, local | ran once | 16:21:10 UTC, success |
|
||
| Whole-guest backup, PBS | ran once | 16:31:13 UTC, verify ok |
|
||
| OS update | ran once | `os_reports` 40/41, trigger `night`, 16:24/16:25 UTC, 49 + 106 packages, healthy |
|
||
| DB dump, tier-2 copy | never | — |
|
||
| Off-site copy | never | escrow `pending` |
|
||
| App updates | none | no apps |
|
||
| Restore-test | never | `restore_tests: []` in every report |
|
||
|
||
Events: 17:58 `host_stale`, 18:01 `node_stale`, 18:05–06 recovered; 18:51 stale again; 19:36 `node_down` +
|
||
`host_down` (error) — the household was mailed "Your server cannot be reached." at 19:36:42 UTC. Hub log
|
||
2026-10-05 05:00 Budapest: "Deadline check: … 0 backup missed … 1 skipped (down)".
|
||
|
||
**What fires if it stays off every night:**
|
||
- every evening `node_stale`/`host_stale` after 45 min; every night `node_down`/`host_down` after 90 min — the
|
||
household gets a `node_down` mail each night and a "reachable again" mail each morning (the 6 h cooldown does not
|
||
stop a daily repeat);
|
||
- **never** `expected_backup_missed` / `expected_dbdump_missed` while it is down at 05:00 (the deadline check skips
|
||
down customers);
|
||
- `restore_test_stale` ~2026-10-11 16:13 UTC (local tier, 7 d) and ~2026-10-16 (PBS, 12 d) — not skipped for down boxes;
|
||
- `os_update_stale` 2026-10-11 16:24 UTC unless a primary backup happens before;
|
||
- `offsite_stale` (48 h) once escrow is completed — the off-site leg runs only at night.
|
||
|
||
**What the household sees:** the backups page shows "Esedékes / Due" or "Naprakész / Up to date" and "No successful
|
||
backup yet" per tier. No text says the server must stay on at night or that a night was missed because it was off.
|
||
|
||
## Q2 — does anything catch up when the box comes back?
|
||
|
||
- **DB dump, tier-2, off-site, the chained app-update leg, the controller's 04:30 self-update:** NO. The daily
|
||
scheduler always picks the next future time; `LastRun` is in memory only and never consulted. (A hub floor still
|
||
reaches the box by day, after any report.)
|
||
- **Whole-guest backup:** PARTIAL. The controller asks every 5 min; due = newest archive ≥ 24 h old (read from
|
||
storage, survives a reboot); runs only in the window [W+2h, W+6h) (04:30–08:30 for W 02:30), OR by the safety valve
|
||
(never backed up, or last > 48 h). A box on only in the afternoon/evening gets about one every 2 days.
|
||
- **OS leg:** 90 s after any successful primary backup, ≥ 20 h apart — so it follows the valve, by day, under the
|
||
`night` label.
|
||
- **Restore-test:** checks every 6 h, the timer restarts at each agent start, no check at start — a box never on for
|
||
6 h in a row never runs one.
|
||
- Not verified: a SUSPENDED (not powered-off) laptop would probably run the daily jobs late rather than skip them
|
||
(reasoned from Go timers, not tested).
|
||
|
||
## Q3 — architecture
|
||
|
||
**No architecture document covers a box that is not always on.** Partial mentions: `11-os-updates.md` ("a box off for
|
||
a weekend does not alarm"); `07-backup-architecture.md` (the night chain and the [W+2h, W+6h) gate — no safety valve,
|
||
no catch-up); `08-alarm-ladder.md` (45/90 min); `where-felhom-stands.yaml` `fail.expected-downtime` (missing, R-285).
|
||
The intent "a box powered on only outside its window never starves" exists only as controller code comments
|
||
(`quiesce.go`).
|
||
|
||
## Q4 — register
|
||
|
||
R-285 (open, P4: no notion of expected downtime), R-30 (open, P3), R-195 / R-321 (closed: skips keyed off the
|
||
wrong fact), R-86 / R-100 (closed). **No row covers a box that is off every night.**
|
||
|
||
## Not determined
|
||
|
||
The daytime online pattern; Tester 2's actual W (on the box, which is off; 02:30 assumed); the type of signed job 53
|
||
(queued 18:35 UTC, after the box went off — the 2026-10-04 agent update, per the night's records).
|