Files
felhom.eu/documentation/audits/night-fixes-2026-10-05/partF/FINDINGS.md
T

74 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Part F — a box that is OFF every night (Tester 2) — read-only spike, 2026-10-05
Collected by a read-only helper session (hub DB copy at ~05:11 UTC, non-secret columns only, copy shredded; hub
log; source at felhom.eu 7221ee5c, controller 69e9145, agent c8d12f1). Two claims re-checked by the main session
against source: the 05:00 deadline check skips `down` customers (`hub/internal/monitor/deadline.go` ~360), and the
controller's daily scheduler always picks the NEXT future time (`controller/internal/scheduler/scheduler.go`
`nextDailyRun`, ~389-408). Hub DB times are UTC; the hub log prints Budapest time.
## Q1 — Tester 2's record (customer `Tester-2`, host `Tester-2-be8404`)
Customer created 2026-09-30 (language en); host created 2026-10-04 16:07 UTC (agent 0.142.0, controller 0.292.0);
no apps installed. Online twice on 2026-10-04: ~15:54–17:13 UTC, and 18:03:42–18:06:24 UTC (last controller report).
Nothing since. One evening of data: no daytime pattern can be established.
| Leg | Result | Evidence |
|---|---|---|
| Whole-guest backup, local | ran once | 16:21:10 UTC, success |
| Whole-guest backup, PBS | ran once | 16:31:13 UTC, verify ok |
| OS update | ran once | `os_reports` 40/41, trigger `night`, 16:24/16:25 UTC, 49 + 106 packages, healthy |
| DB dump, tier-2 copy | never | — |
| Off-site copy | never | escrow `pending` |
| App updates | none | no apps |
| Restore-test | never | `restore_tests: []` in every report |
Events: 17:58 `host_stale`, 18:01 `node_stale`, 18:05–06 recovered; 18:51 stale again; 19:36 `node_down` +
`host_down` (error) — the household was mailed "Your server cannot be reached." at 19:36:42 UTC. Hub log
2026-10-05 05:00 Budapest: "Deadline check: … 0 backup missed … 1 skipped (down)".
**What fires if it stays off every night:**
- every evening `node_stale`/`host_stale` after 45 min; every night `node_down`/`host_down` after 90 min — the
household gets a `node_down` mail each night and a "reachable again" mail each morning (the 6 h cooldown does not
stop a daily repeat);
- **never** `expected_backup_missed` / `expected_dbdump_missed` while it is down at 05:00 (the deadline check skips
down customers);
- `restore_test_stale` ~2026-10-11 16:13 UTC (local tier, 7 d) and ~2026-10-16 (PBS, 12 d) — not skipped for down boxes;
- `os_update_stale` 2026-10-11 16:24 UTC unless a primary backup happens before;
- `offsite_stale` (48 h) once escrow is completed — the off-site leg runs only at night.
**What the household sees:** the backups page shows "Esedékes / Due" or "Naprakész / Up to date" and "No successful
backup yet" per tier. No text says the server must stay on at night or that a night was missed because it was off.
## Q2 — does anything catch up when the box comes back?
- **DB dump, tier-2, off-site, the chained app-update leg, the controller's 04:30 self-update:** NO. The daily
scheduler always picks the next future time; `LastRun` is in memory only and never consulted. (A hub floor still
reaches the box by day, after any report.)
- **Whole-guest backup:** PARTIAL. The controller asks every 5 min; due = newest archive ≥ 24 h old (read from
storage, survives a reboot); runs only in the window [W+2h, W+6h) (04:30–08:30 for W 02:30), OR by the safety valve
(never backed up, or last > 48 h). A box on only in the afternoon/evening gets about one every 2 days.
- **OS leg:** 90 s after any successful primary backup, ≥ 20 h apart — so it follows the valve, by day, under the
`night` label.
- **Restore-test:** checks every 6 h, the timer restarts at each agent start, no check at start — a box never on for
6 h in a row never runs one.
- Not verified: a SUSPENDED (not powered-off) laptop would probably run the daily jobs late rather than skip them
(reasoned from Go timers, not tested).
## Q3 — architecture
**No architecture document covers a box that is not always on.** Partial mentions: `11-os-updates.md` ("a box off for
a weekend does not alarm"); `07-backup-architecture.md` (the night chain and the [W+2h, W+6h) gate — no safety valve,
no catch-up); `08-alarm-ladder.md` (45/90 min); `where-felhom-stands.yaml` `fail.expected-downtime` (missing, R-285).
The intent "a box powered on only outside its window never starves" exists only as controller code comments
(`quiesce.go`).
## Q4 — register
R-285 (open, P4: no notion of expected downtime), R-30 (open, P3), R-195 / R-321 (closed: skips keyed off the
wrong fact), R-86 / R-100 (closed). **No row covers a box that is off every night.**
## Not determined
The daytime online pattern; Tester 2's actual W (on the box, which is off; 02:30 assumed); the type of signed job 53
(queued 18:35 UTC, after the box went off — the 2026-10-04 agent update, per the night's records).