Files
felhom.eu/documentation/audits/night-fixes-2026-10-05/partF/FINDINGS.md
T

4.7 KiB
Raw Blame History

Part F — a box that is OFF every night (Tester 2) — read-only spike, 2026-10-05

Collected by a read-only helper session (hub DB copy at ~05:11 UTC, non-secret columns only, copy shredded; hub log; source at felhom.eu 7221ee5c, controller 69e9145, agent c8d12f1). Two claims re-checked by the main session against source: the 05:00 deadline check skips down customers (hub/internal/monitor/deadline.go ~360), and the controller's daily scheduler always picks the NEXT future time (controller/internal/scheduler/scheduler.go nextDailyRun, ~389-408). Hub DB times are UTC; the hub log prints Budapest time.

Q1 — Tester 2's record (customer Tester-2, host Tester-2-be8404)

Customer created 2026-09-30 (language en); host created 2026-10-04 16:07 UTC (agent 0.142.0, controller 0.292.0); no apps installed. Online twice on 2026-10-04: ~15:54–17:13 UTC, and 18:03:42–18:06:24 UTC (last controller report). Nothing since. One evening of data: no daytime pattern can be established.

Leg Result Evidence
Whole-guest backup, local ran once 16:21:10 UTC, success
Whole-guest backup, PBS ran once 16:31:13 UTC, verify ok
OS update ran once os_reports 40/41, trigger night, 16:24/16:25 UTC, 49 + 106 packages, healthy
DB dump, tier-2 copy never —
Off-site copy never escrow pending
App updates none no apps
Restore-test never restore_tests: [] in every report

Events: 17:58 host_stale, 18:01 node_stale, 18:05–06 recovered; 18:51 stale again; 19:36 node_down + host_down (error) — the household was mailed "Your server cannot be reached." at 19:36:42 UTC. Hub log 2026-10-05 05:00 Budapest: "Deadline check: … 0 backup missed … 1 skipped (down)".

What fires if it stays off every night:

  • every evening node_stale/host_stale after 45 min; every night node_down/host_down after 90 min — the household gets a node_down mail each night and a "reachable again" mail each morning (the 6 h cooldown does not stop a daily repeat);
  • never expected_backup_missed / expected_dbdump_missed while it is down at 05:00 (the deadline check skips down customers);
  • restore_test_stale ~2026-10-11 16:13 UTC (local tier, 7 d) and ~2026-10-16 (PBS, 12 d) — not skipped for down boxes;
  • os_update_stale 2026-10-11 16:24 UTC unless a primary backup happens before;
  • offsite_stale (48 h) once escrow is completed — the off-site leg runs only at night.

What the household sees: the backups page shows "Esedékes / Due" or "Naprakész / Up to date" and "No successful backup yet" per tier. No text says the server must stay on at night or that a night was missed because it was off.

Q2 — does anything catch up when the box comes back?

  • DB dump, tier-2, off-site, the chained app-update leg, the controller's 04:30 self-update: NO. The daily scheduler always picks the next future time; LastRun is in memory only and never consulted. (A hub floor still reaches the box by day, after any report.)
  • Whole-guest backup: PARTIAL. The controller asks every 5 min; due = newest archive ≥ 24 h old (read from storage, survives a reboot); runs only in the window [W+2h, W+6h) (04:30–08:30 for W 02:30), OR by the safety valve (never backed up, or last > 48 h). A box on only in the afternoon/evening gets about one every 2 days.
  • OS leg: 90 s after any successful primary backup, ≥ 20 h apart — so it follows the valve, by day, under the night label.
  • Restore-test: checks every 6 h, the timer restarts at each agent start, no check at start — a box never on for 6 h in a row never runs one.
  • Not verified: a SUSPENDED (not powered-off) laptop would probably run the daily jobs late rather than skip them (reasoned from Go timers, not tested).

Q3 — architecture

No architecture document covers a box that is not always on. Partial mentions: 11-os-updates.md ("a box off for a weekend does not alarm"); 07-backup-architecture.md (the night chain and the [W+2h, W+6h) gate — no safety valve, no catch-up); 08-alarm-ladder.md (45/90 min); where-felhom-stands.yaml fail.expected-downtime (missing, R-285). The intent "a box powered on only outside its window never starves" exists only as controller code comments (quiesce.go).

Q4 — register

R-285 (open, P4: no notion of expected downtime), R-30 (open, P3), R-195 / R-321 (closed: skips keyed off the wrong fact), R-86 / R-100 (closed). No row covers a box that is off every night.

Not determined

The daytime online pattern; Tester 2's actual W (on the box, which is off; 02:30 assumed); the type of signed job 53 (queued 18:35 UTC, after the box went off — the 2026-10-04 agent update, per the night's records).