catch-up session 2026-10-05: design 07 §6.1.1 (a box that is not always on), 08 §6.4; rulings 109-111, CC decisions 112-118; R-871/R-873..R-877 closed, R-872 narrowed (dated check), R-878 opened; live evidence; STATUS
gates / gates (push) Successful in 33s
gates / gates (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -214,6 +214,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Always-on debug rings + on-demand log-bundle pulls with TTL/custody | controller v0.116, agent v0.83, hub v0.46 | **PROVEN-LIVE** | debug rings live-exercised `CAMPAIGN-3` fix-6 (1000-cap ring, ~55min horizon under load) | The **log-bundle-pull TTL/custody** half is changelog-only (no dedicated observability audit doc); ring persistence across restart is a known gap |
|
||||
| Operator alerting (Healthchecks → monitoring@felhom.eu) | k3s, Resend | **IMPLEMENTED** | operator infra, stated in production since 02-04; no corpus validation doc |
|
||||
| Backup-deadline alerting (`expected_backup_missed`) is ANCHORED — absence of signal is UNKNOWN, not failure | hub v0.75.0 | **IMPLEMENTED** | `audits/DIAG-backup-missed-2026-07-26.md` + red-proofs A/B/C + replay of the real 2026-07-26 03:00 reports (all three silent) | **No row status flips** — this signal had a FALSE-POSITIVE class (three instances: hub v0.12.0, v0.73.0, R-81), now anchored at first contact and read across retained host-report history. Still unit-proven only, not live-fired at a real deadline. The *underlying* PBS/offsite-DR tier gap it exposed is → R-82. | Per the status enum, no citation → not PROVEN-LIVE. Demoted pending an operator-cited live alert (candidate re-upgrade — see REPORT) |
|
||||
| A box that is not always on (off at its backup time) makes the missed night up once when it comes back; the household sees a banner; the operator gets the missed-backup alarm, down or not | controller v0.295.0, hub v0.134.0, agent v0.145.0 | **PARTIAL — the catch-up and the banner PROVEN-LIVE (9202 + demo-felhom; banner on 9202 by a hand-set ledger); R-872's alarm proven by tests, live at the next 05:00 (dated check); a host SUSPEND reasoned + unit-tested, not measured** | `audits/catchup-2026-10-05/partA/`, `partB/`; design `07` §6.1.1 | R-872 live; a suspend never measured |
|
||||
|
||||
## G. Fleet & operator (hub)
|
||||
|
||||
@@ -232,7 +233,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/` |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
|
||||
|
||||
@@ -396,6 +396,80 @@ warning that should precede such a failure is R-685.
|
||||
`pct config 9201`, both hosts). So the whole-guest tiers carry the guest and **none of the customer's
|
||||
data drives** — 916 GB on demo-felhom, 938 GB on demo-hp.
|
||||
|
||||
### 6.1.1 A box that is not always on — the catch-up, the banner, the alarms (R-871..R-874)
|
||||
|
||||
> **Why this section exists (R-871):** until 2026-10-05 no architecture document covered a box that is OFF at its
|
||||
> window W. Tester 2 is a laptop switched off at night (operator, 2026-10-05). Measured (`audits/night-fixes-2026-10-05/
|
||||
> partF/FINDINGS.md`, and live on 9202 `audits/catchup-2026-10-05/partA-spike/`): the controller's daily jobs always
|
||||
> schedule the NEXT future time, so a missed 02:30/03:30/04:15 waited for the next night for ever; no alarm fired
|
||||
> while the box was down at 05:00; the household was mailed "your server cannot be reached" every night.
|
||||
|
||||
**[DESIGN — ruled 2026-10-05, `09` decision 109] A missed night runs ONCE when the box comes back.** Built controller
|
||||
v0.295.0 (`internal/nightchain`):
|
||||
|
||||
- **What a box does when it was off at W.** On a controller START and on a host RESUME, the controller asks its
|
||||
night ledger (`<data>/night-ledger.json`) which backup legs missed their last scheduled time. The ledger records when
|
||||
each leg last RAN TO ITS END — an attempt record, used only to decide "missed", never as evidence a backup exists
|
||||
(`R-100`'s rule). A leg that ran and failed was NOT missed: failures have their own alarms.
|
||||
- **What the catch-up runs:** the three BACKUP legs, in the night's order — database dump, second copy, off-site copy
|
||||
— through the SAME wrapped leg bodies the scheduled jobs run (`main.go` `withLeg`, `catchUpLegs`).
|
||||
- **What it does not run:** the app-update leg and every Docker step (they restart apps; they wait for a real night).
|
||||
The OS fast lane is the agent's and follows a whole-guest backup, not the catch-up (below).
|
||||
- **When:** 15 minutes after the trigger (apps settle; a box switched on and off again at once does nothing). A leg
|
||||
whose own next scheduled time is under 30 minutes away is left to its normal run.
|
||||
- **Only once:** several missed nights = one catch-up (the question is about the LAST scheduled time). A normal night
|
||||
followed by a daytime restart = none. A power cut in the middle of the chain = only the legs that did not end.
|
||||
- **Never two at once:** every leg, scheduled or catch-up, holds one lock; a second trigger while one is pending does
|
||||
nothing.
|
||||
- **The whole-guest backup.** It keeps its own agent-side behaviour (inside [W+2h, W+6h), or the 48 h safety valve —
|
||||
about one every 2 days for an evening-only box). **It CAN collide with a catch-up** (the valve fires on the first
|
||||
5-minute poll after a start), so each waits for the other: a scheduled quiesce defers while a catch-up runs
|
||||
(`quiesce` `SetCatchUpFn`), and a catch-up waits up to 2 h while a quiesce holds the apps.
|
||||
- **A suspended host** (a laptop lid): Go's timers run on CLOCK_MONOTONIC, which does not count suspended time, so
|
||||
the 02:30 timer would fire hours late — and would start the app-update leg at noon. Two rules: a daily job whose
|
||||
timer fires more than 60 minutes after its wall-clock time is SKIPPED (`scheduler.DailyLateLimit`), and a resume
|
||||
watch (wall clock vs monotonic clock, checked every minute) triggers the catch-up. *Reasoned from the Go and Linux
|
||||
clock semantics and unit-tested with injected clocks; not measured live (no suspend was allowed on a demo box).*
|
||||
- **The first start on this release** seeds the ledger: nothing before it counts as missed (no surprise catch-up on
|
||||
every box at the upgrade).
|
||||
- **What the household sees:** one timeline line — "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 02:30-kor, a
|
||||
mentés most elkészült." / "Missed backup made now: the box was off at 02:30, so the backup ran when it came back on."
|
||||
(event `backup_catchup_done`, info: recorded, never mailed).
|
||||
|
||||
**[DESIGN — ruled 2026-10-05, `09` decision 110] The banner (the operator's idea).** On every page of a logged-in
|
||||
household: when the last daily backup (the database dump, and the off-site copy when configured) is over **26 h** old —
|
||||
when the last one was, "the box was off at backup time (02:30)" when the box's own record says so, and a suggested
|
||||
time. The record is the controller's system-metrics table (one sample a minute, kept 30 days): "on" at W = a sample
|
||||
within 5 minutes of it; "usually on" in an hour = on in it on at least 5 of the last 7 days. It **never changes the
|
||||
time** — a button opens the backup-time setting. The household can close it: it stays closed until the NEXT missed
|
||||
backup time (durable, in the ledger) and disappears by itself after a successful night.
|
||||
|
||||
*Decided by CC — operator may reverse* (`09` decisions 112–116): the 15-minute delay and 30-minute leave-to-normal
|
||||
line; the 60-minute late-fire limit; 26 h as the banner's line (one night + the chain's two hours, = the hub's
|
||||
`backupStaleAfter`); the suggestion = the LATEST hour H with H, H+1, H+2 usually on (a later evening disturbs least),
|
||||
none when the current window is already usually on; the catch-up's line on the household's timeline but no mail.
|
||||
|
||||
**The operator's side (hub v0.134.0, `08` §6.4).** R-872: a box DOWN at the 05:00 deadline is no longer skipped — it
|
||||
is judged on longer lines (48 h without a dump, 72 h without a whole-guest backup), so a box that died last night
|
||||
raises only its staleness alarm, and a box off at every deadline raises `expected_dbdump_missed` /
|
||||
`expected_backup_missed`. R-873: the household hears "your server cannot be reached" at most once per 7 days (the
|
||||
operator still gets every edge). R-874 (agent v0.145.0): the restore-test's first due-check runs 30 minutes after the
|
||||
agent starts, so a box with short power-on sessions is still restore-tested.
|
||||
|
||||
**[FACT, 2026-10-05] Proven live** (`audits/catchup-2026-10-05/partA/`): 9202 off across 09:35, on at 07:38 UTC → the
|
||||
catch-up at 07:53:09 made the dump (21 s); a crash of its host mid-wait → at the next start ONE new catch-up made all
|
||||
three legs (22 s); demo-felhom's controller off across 10:07 → the dump made at 08:25:03, 15 min after the start, and
|
||||
the household's timeline line reached the hub. **What the household may notice:** the dump leg stops an app with a
|
||||
volume for its copy — measured 1 s for opengist (182.5 KB), the same as at night; a large volume takes longer, in the
|
||||
day (R-878).
|
||||
|
||||
**Edge cases, stated.** A box on only in the day with W at night: the catch-up makes the backups every day it is
|
||||
switched on (after 15 min); the banner suggests an evening time once a pattern exists. A box switched on during the
|
||||
chain: legs already past are made up, legs still ahead run normally. A box whose controller restarts in the day after
|
||||
a normal night: nothing. A box off for weeks: one catch-up when it returns; the hub's staleness alarm has been
|
||||
running all along. An upgrade from an older release: the ledger is seeded, the first missed night after it is the
|
||||
first one made up.
|
||||
|
||||
### 6.2 Coverage per app class — and an unresolved count
|
||||
|
||||
**[FACT]** Of 53 catalog templates, **52 keep data in Docker named volumes**; exactly **13** carry a
|
||||
|
||||
@@ -340,6 +340,26 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
|
||||
|
||||
---
|
||||
|
||||
## 6.4 A box that is not always on: the missed-backup deadline and the household's outage mail [DESIGN, hub v0.134.0, 2026-10-05]
|
||||
|
||||
Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *operator may reverse*).
|
||||
|
||||
- **R-872 — a box DOWN at the 05:00 deadline is judged, not skipped.** The check used to skip every customer whose node
|
||||
is `down` ("they already have staleness events"), so a box down at EVERY deadline — a laptop off at night — was
|
||||
never judged (measured 2026-10-05: `1 skipped (down)`). Now a down box is judged on longer lines:
|
||||
`expected_dbdump_missed` after **48 h** without a `db_dump_completed`, `expected_backup_missed` after **72 h**
|
||||
without a whole-guest backup in any retained host report — never for a box first seen less than 48 h ago. A box that
|
||||
died last night still raises only its staleness alarm. A `disabled` box is still skipped (R-321). Pinned by
|
||||
`monitor/r872_down_box_test.go`.
|
||||
- **R-873 — "your server cannot be reached" reaches the HOUSEHOLD at most once per 7 days** (`node_stale`,
|
||||
`node_down`, `host_stale`, `host_down`, read from the persisted notification log, so a hub restart does not reset
|
||||
it). The operator still gets every edge. The recovery mail stays paired with a down mail the household actually
|
||||
received (§6.2's pairing), so a held-back down mail also holds back its recovery. Pinned by
|
||||
`notify/r873_liveness_weekly_test.go`.
|
||||
- `backup_catchup_done` (info) is the box's "missed backup made now" line: recorded, never mailed.
|
||||
|
||||
---
|
||||
|
||||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||||
|
||||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||||
|
||||
@@ -815,6 +815,44 @@ its length, and both fixes cost something the household would notice — operato
|
||||
1.30.0, agent 0.142.0 and the re-made golden; it lacks only agent 0.142.1's wrapper fix (R-858). *Operator ruling
|
||||
2026-10-04 ~18:49.*
|
||||
|
||||
### 2026-10-05 (08:42) — three operator rulings (recorded before the work; the catch-up brief)
|
||||
|
||||
109. **R-871, option A: a missed night runs once when the box comes back.** A box that was off during its backup time
|
||||
makes the missed backups soon after it is switched on; updates that restart apps still wait for a real night.
|
||||
**Rejected:** B — "the box must stay on at night", with an alarm only. *Operator ruling 2026-10-05.*
|
||||
110. **The household is told on its dashboard** (the operator's idea): a banner the household can close, shown when a
|
||||
daily backup was missed — when the last backup was, that the box was off during the backup time, and a suggested
|
||||
different time. It never changes the time by itself. *Operator ruling 2026-10-05.*
|
||||
111. **Tester 2: read only.** If it is online, CC re-signs its agent update (decision 97, act 1) and reads it back.
|
||||
Nothing else. *Operator ruling 2026-10-05.*
|
||||
|
||||
### 2026-10-05 (day, catch-up brief) — decided by CC — operator may reverse
|
||||
|
||||
112. **When the catch-up runs (R-871).** Options: (a) at once after the box comes back — a box switched on and off
|
||||
again starts a dump each time; (b) 15 minutes later, re-checking first. **Chosen (b)**, as the brief proposed;
|
||||
a leg due within 30 minutes is left to its normal run. `07` §6.1.1.
|
||||
113. **What a host suspend does to the night (R-871).** Options: (a) only a controller start triggers the catch-up —
|
||||
a suspended laptop's 02:30 timer then fires hours late (Go timers run on CLOCK_MONOTONIC) and starts the
|
||||
app-update leg at noon; (b) skip any daily job that fires over 60 minutes late and trigger the catch-up from a
|
||||
resume watch. **Chosen (b).** Reasoned and unit-tested with injected clocks; not measured live.
|
||||
114. **The banner's line and its suggestion (decision 110).** Options for the line: 24 h (a slow night flickers it),
|
||||
26 h (one night + the chain's two hours, = the hub's `backupStaleAfter`), 48 h (two missed nights before a word).
|
||||
**Chosen 26 h.** Suggestion: the latest hour H with H, H+1, H+2 usually on (on ≥ 5 of the last 7 days); none
|
||||
when the current window is already usually on (a new time would not help).
|
||||
115. **R-872's lines for a box that is down at the deadline.** Options: (a) judge a down box like an up one — a box
|
||||
that died last night raises a backup alarm on top of its staleness alarm; (b) 48 h without a dump / 72 h
|
||||
without a whole-guest backup (the agent's own 48 h valve + a day). **Chosen (b).** hub v0.134.0, `08` §6.4.
|
||||
116. **R-873's rule.** Options: (a) household liveness mail only after 24 h down — a real outage reaches the household
|
||||
a day late; (b) at most once per 7 days to the household, the operator every edge. **Chosen (b)** (the brief's
|
||||
first example): the first outage of a week still reaches the household at once. hub v0.134.0, `08` §6.4.
|
||||
117. **R-874's first restore-test check after start.** Options: at start (a crash loop hammers a failing tier — the
|
||||
earned restraint), 30 minutes after start, or keep one interval (6 h — a short-session box never reaches it).
|
||||
**Chosen 30 minutes.** agent v0.145.0.
|
||||
118. **R-876's repair trigger.** Options: (a) always run `dpkg --configure -a` (R-845's speed lost on every clean
|
||||
pass); (b) read `--audit` AND the update journal in ONE `sh -c` call, repair when either shows something, and as
|
||||
a belt repair + retry once when apt itself says "dpkg was interrupted". **Chosen (b)**: a clean pass still costs
|
||||
one call (pinned by a test). agent v0.145.0.
|
||||
|
||||
### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief)
|
||||
|
||||
100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) —
|
||||
|
||||
@@ -639,6 +639,19 @@ Agent v0.144.0 + v0.144.1 (wrapper and agent; `09` decisions 106–108). Evidenc
|
||||
(runbook `crash-guard.md`), then the next pass installed the 12. Operator mail `os_update_failed` (true); no
|
||||
household mail; the household's timeline showed "Controller elindult".
|
||||
|
||||
### 8.5 The self-repair after a power cut as BUILT (2026-10-05, agent v0.145.0) `[FACT]`
|
||||
|
||||
- **R-876 fixed.** The wrapper reads `dpkg --audit` AND dpkg's update journal (`/var/lib/dpkg/updates/`) in ONE call
|
||||
and repairs when either shows something; as a belt, when apt itself says "dpkg was interrupted", it repairs and
|
||||
retries once (`09` decision 118). A clean pass still costs one call (R-845's speed, pinned by a test).
|
||||
- **Proven live** (operator's go, demo-hp, `audits/catchup-2026-10-05/partD/`): 13 guest packages rolled back, crash
|
||||
at 07:56:03 UTC during dpkg's unpack, back by itself (new boot 07:56:41, guard armed, 1 unclean boot in the window);
|
||||
at boot `--audit` clean, journal 1 file, the same shape as §8.4. **The next pass, with nobody touching the box:**
|
||||
`REPAIR configured=0 journal=1`, then `DONE rc=0 upgraded=12`, healthy; the guest's package list equals the one
|
||||
before the rollback. No mail (the pass did not fail); the household's timeline: "Controller elindult".
|
||||
- **R-874.** The agent's restore-test due-check runs 30 minutes after start (then every interval), so a box with short
|
||||
power-on sessions is restore-tested (`09` decision 117).
|
||||
|
||||
## 9. Where the rest lives
|
||||
|
||||
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
08:27:24
|
||||
('Tester-2-be8404', '0.142.0', '2026-10-04 18:05:48', None)
|
||||
('demo-felhom-8363b5', '0.145.0', '2026-10-05 08:23:56', '0.145.0')
|
||||
('demo-hp-bb76ea', '0.145.0', '2026-10-05 08:27:06', '0.145.0')
|
||||
('tester-1-d70be4', '0.145.0', '2026-10-05 08:27:31', '0.145.0')
|
||||
== demo-hp: agent felhom-agent 0.145.0 | controller 0.295.0 | healthy 20 of 21 | guard 1
|
||||
== felhom-pve: agent felhom-agent 0.145.0 | controller 0.295.0 | healthy 4 of 5 | guard 1
|
||||
Oct 05 10:08:58 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:58.009+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUN
|
||||
9202: gitea.dooplex.hu/admin/felhom-controller:0.295.0 Up 29 minutes (healthy)
|
||||
@@ -0,0 +1,12 @@
|
||||
### 06:58:03 UTC: set 9202's window to 09:05 (Budapest) by POST /backups/window
|
||||
HTTP 303
|
||||
2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 09:05 (next run 2026-10-05 09:05 CEST)
|
||||
2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 10:05 (next run 2026-10-05 10:05 CEST)
|
||||
2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 10:50 (next run 2026-10-05 10:50 CEST)
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: rescheduled — recomputing next run
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: rescheduled — recomputing next run
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: next run at 2026-10-05 09:05:00 CEST (waiting 6m54s)
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: next run at 2026-10-05 10:05:00 CEST (waiting 1h6m54s)
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: rescheduled — recomputing next run
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: next run at 2026-10-05 10:50:00 CEST (waiting 1h51m54s)
|
||||
2026/10/05 06:58:05 backup_handlers.go:65: [INFO] [web] backup window set to 09:05 (legs 09:05/10:05/10:50)
|
||||
@@ -0,0 +1,9 @@
|
||||
### 07:03:01 pct shutdown 9202
|
||||
status: stopped
|
||||
### 07:08:01 pct start 9202
|
||||
status: running
|
||||
### 07:10:35 controller log after the start
|
||||
2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:05 CEST
|
||||
2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-10-05 10:05 CEST
|
||||
2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-10-05 10:50 CEST
|
||||
2026/10/05 07:08:08 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
@@ -0,0 +1 @@
|
||||
('demo-felhom', 'backup_catchup_done', 'info', 'Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült.', 'controller', '2026-10-05 08:25:03')
|
||||
@@ -0,0 +1,12 @@
|
||||
### 07:29:47 9202: controller 0.295.0 by its bootstrap file
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.295.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.295.0 started 2026-10-05T07:29:50.143207934Z
|
||||
2026/10/05 07:29:50 scheduler.go:149: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:05 CEST
|
||||
2026/10/05 07:29:50 nightchain.go:280: [INFO] [catch-up] controller start: no backup leg missed its last scheduled time — nothing to make up
|
||||
{
|
||||
"seeded_at": "2026-10-05T07:29:50.484758959Z",
|
||||
"ended": {},
|
||||
"db_dump_ok": "0001-01-01T00:00:00Z",
|
||||
"last_catch_up": "0001-01-01T00:00:00Z",
|
||||
"banner_dismissed_through": "0001-01-01T00:00:00Z"
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
### 07:33:00 W = 09:35 Budapest (07:35 UTC). pct shutdown 9202
|
||||
command 'lxc-stop -n 9202 --nokill --timeout 60' failed: exit code 1
|
||||
status: stopped
|
||||
### 07:38:00 pct start 9202
|
||||
status: running
|
||||
### 07:39:06 after the start
|
||||
2026/10/05 07:38:09 scheduler.go:149: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:35 CEST
|
||||
2026/10/05 07:38:09 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 09:35) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
### 07:53:51 the catch-up
|
||||
2026/10/05 07:43:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health
|
||||
2026/10/05 07:43:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
2026/10/05 07:44:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:46:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: docker-socket-users
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: offsite-credential-retry
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: backup-cache
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health
|
||||
2026/10/05 07:48:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
2026/10/05 07:50:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:52:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:53:09 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: offsite-credential-retry
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: docker-socket-users
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: backup-cache
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health
|
||||
2026/10/05 07:53:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
2026/10/05 07:53:09 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (412.6 KB, 436ms, 72 tables)
|
||||
2026/10/05 07:53:30 nightchain.go:340: [INFO] [catch-up] done: db-dump in 21s (missed at 2026-10-05 09:35)
|
||||
{
|
||||
"seeded_at": "2026-10-05T07:29:50.484758959Z",
|
||||
"ended": {
|
||||
"db-dump": "2026-10-05T07:53:30.483916689Z"
|
||||
},
|
||||
"db_dump_ok": "2026-10-05T07:53:30.484211296Z",
|
||||
"last_catch_up": "2026-10-05T07:53:30.484384683Z",
|
||||
"banner_dismissed_through": "0001-01-01T00:00:00Z"
|
||||
}
|
||||
2026/10/05 07:38:09 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 09:35) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
2026/10/05 07:53:09 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 07:53:30 nightchain.go:340: [INFO] [catch-up] done: db-dump in 21s (missed at 2026-10-05 09:35)
|
||||
+23
@@ -0,0 +1,23 @@
|
||||
2026/10/05 07:58:35 main.go:340: [INFO] felhom-controller 0.295.0 starting (customer: demo-hp, domain: enkisfelhom.hu)
|
||||
2026/10/05 07:58:35 scheduler.go:149: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-10-06 03:30 CEST
|
||||
2026/10/05 07:58:35 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump tier2 offsite] (last scheduled 2026-10-05 04:15) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
2026/10/05 08:13:35 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 08:13:36 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (413.8 KB, 423ms, 72 tables)
|
||||
2026/10/05 08:13:57 nightchain.go:332: [INFO] [catch-up] running the missed tier2 leg
|
||||
2026/10/05 08:13:57 tier2.go:425: [INFO] [backup] Tier 2 copied paperless-ngx → /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx (82.5 MB, 1 leg(s), 0s) [SSD: state-only]
|
||||
2026/10/05 08:13:57 tier2.go:425: [INFO] [backup] Tier 2 copied privatebin → /mnt/felhom-drives/scratch_hdd/backups/secondary/privatebin (20.3 KB, 0 leg(s), 0s)
|
||||
2026/10/05 08:13:57 tier2.go:476: [INFO] [backup] Tier 2 run complete: 2 app(s) processed (incl. volume-only — F6)
|
||||
2026/10/05 08:13:57 nightchain.go:332: [INFO] [catch-up] running the missed offsite leg
|
||||
2026/10/05 08:13:57 main.go:1385: [INFO] [offbox] no scheduled off-site target on this box — the off-site leg does nothing
|
||||
2026/10/05 08:13:57 nightchain.go:340: [INFO] [catch-up] done: db-dump, offsite, tier2 in 22s (missed at 2026-10-05 04:15)
|
||||
{
|
||||
"seeded_at": "2026-09-30T07:54:19.875183Z",
|
||||
"ended": {
|
||||
"db-dump": "2026-10-05T08:13:57.406392075Z",
|
||||
"offsite": "2026-10-05T08:13:57.69370068Z",
|
||||
"tier2": "2026-10-05T08:13:57.693478189Z"
|
||||
},
|
||||
"db_dump_ok": "2026-10-05T08:13:57.406683255Z",
|
||||
"last_catch_up": "2026-10-05T08:13:57.693858077Z",
|
||||
"banner_dismissed_through": "2026-10-05T00:30:00Z"
|
||||
}
|
||||
@@ -0,0 +1,24 @@
|
||||
### 08:05:01 W = 10:07 Budapest (08:07 UTC). park + stop the controller (apps keep running)
|
||||
felhom-controller
|
||||
4
|
||||
### 08:10:00 start the controller + unpark
|
||||
felhom-controller
|
||||
2026/10/05 08:10:02 [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 10:07 CEST
|
||||
### 08:25:04 the catch-up
|
||||
{
|
||||
"seeded_at": "2026-10-05T07:29:41.65020944Z",
|
||||
"ended": {
|
||||
"db-dump": "2026-10-05T08:25:03.369812125Z"
|
||||
},
|
||||
"db_dump_ok": "2026-10-05T08:25:03.370056355Z",
|
||||
"last_catch_up": "2026-10-05T08:25:03.370196251Z",
|
||||
"banner_dismissed_through": "0001-01-01T00:00:00Z"
|
||||
}### 08:25:19 the catch-up lines (this box logs without file:line)
|
||||
2026/10/05 08:10:02 [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 10:07) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
2026/10/05 08:25:02 [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 08:25:03 [INFO] [catch-up] done: db-dump in 1s (missed at 2026-10-05 10:07)
|
||||
### 08:25:20 window back to 02:30
|
||||
HTTP 303
|
||||
2026/10/05 08:25:22 [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15)
|
||||
ls: cannot access '/var/lib/felhom-agent/guests/9201/controller-parked': No such file or directory
|
||||
4
|
||||
@@ -0,0 +1,120 @@
|
||||
### RED-PROOF scheduler late-fire guard dropped
|
||||
scheduler_test.go:128: a 6-hour-late fire ran the job (the app-update leg would start at noon)
|
||||
--- FAIL: TestDaily_LateFireIsSkipped (0.00s)
|
||||
--- PASS: TestDaily_LateFireIsSkipped/on_time (0.00s)
|
||||
--- FAIL: TestDaily_LateFireIsSkipped/after_a_suspend (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.004s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.456s
|
||||
### RED-PROOF 1 Missed returns nothing (the old behaviour: no catch-up)
|
||||
nightchain_test.go:99: no catch-up was scheduled for a box off across W (missed [])
|
||||
--- FAIL: TestCatchUp_OffAtWThenOn_OneCatchUpAfter15Min (0.00s)
|
||||
nightchain_test.go:135: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_TwoMissedNights_One (3.00s)
|
||||
nightchain_test.go:151: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (3.00s)
|
||||
nightchain_test.go:178: missed [], want tier2 + offsite of the night of the 4th
|
||||
--- FAIL: TestCatchUp_LegAboutToRun_LeftToNormal (0.00s)
|
||||
nightchain_test.go:189: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_NeverRunsAnythingButBackupLegs (3.00s)
|
||||
nightchain_test.go:206: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (3.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 12.015s
|
||||
FAIL
|
||||
### RED-PROOF 2 a leg that ran is not checked (Ended ignored)
|
||||
nightchain_test.go:121: a normal night was made up again: [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_NightRan_DaytimeRestart_None (0.00s)
|
||||
nightchain_test.go:140: after the catch-up nothing is missed any more, got [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_TwoMissedNights_One (0.00s)
|
||||
nightchain_test.go:153: ran [db-dump tier2 offsite], want only tier2 + offsite
|
||||
--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 3 the 15-minute delay dropped
|
||||
nightchain_test.go:103: the catch-up must wait 15m0s first, slept []
|
||||
--- FAIL: TestCatchUp_OffAtWThenOn_OneCatchUpAfter15Min (0.00s)
|
||||
nightchain_test.go:208: slept [1m0s 1m0s 1m0s] — want the 15 min delay then 3 one-minute waits for the quiesce
|
||||
--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s
|
||||
FAIL
|
||||
### RED-PROOF 4 a pending catch-up does not block a second one
|
||||
nightchain_test.go:133: a second catch-up was scheduled while one was pending: [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_TwoMissedNights_One (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 5 the quiesce wait dropped
|
||||
nightchain_test.go:208: slept [15m0s] — want the 15 min delay then 3 one-minute waits for the quiesce
|
||||
--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 6 the leave-to-normal rule dropped
|
||||
nightchain_test.go:174: the dump due in 20 minutes was put in the catch-up: [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_LegAboutToRun_LeftToNormal (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 8 a fresh ledger counts history as missed
|
||||
nightchain_test.go:161: a freshly seeded ledger made up [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_FreshLedger_None (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.007s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s
|
||||
### RED-PROOF 7 the resume watch compares wall with wall (compiles)
|
||||
nightchain_test.go:241: a 7-hour suspend was not seen
|
||||
--- FAIL: TestResumeWatch_SeesASuspend (2.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 2.010s
|
||||
FAIL
|
||||
### RED-PROOF 9 the catch-up runs whatever Legs holds (not only the missed chain legs)
|
||||
nightchain_test.go:153: ran [db-dump tier2 offsite], want only tier2 + offsite
|
||||
--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (0.00s)
|
||||
nightchain_test.go:191: an app update ran inside a catch-up
|
||||
--- FAIL: TestCatchUp_NeverRunsAnythingButBackupLegs (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
### RED-PROOF 10 the whole-guest cycle does not wait for a running catch-up
|
||||
r871_catchup_test.go:21: the backup started while a catch-up ran: start=1 stopped=[nextcloud]
|
||||
--- FAIL: TestR871_ScheduledCycleWaitsForCatchUp (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.005s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.005s
|
||||
### RED-PROOF 12 the update leg inside the shared off-site body
|
||||
--- PASS: TestR871_CatchUpWiring (0.01s)
|
||||
r871_catchup_wiring_test.go:44: the update leg is inside the off-site body the catch-up runs
|
||||
--- FAIL: TestR871_CatchUpRunsNoUpdateLeg (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.019s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.019s
|
||||
### RED-PROOF 11 the start trigger not wired (test now keyed on the argument)
|
||||
r871_catchup_wiring_test.go:37: the catch-up's START trigger (Evaluate(ctx, "controller start")) is called 0 times, want 1
|
||||
--- FAIL: TestR871_CatchUpWiring (0.02s)
|
||||
--- PASS: TestR871_CatchUpRunsNoUpdateLeg (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.031s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.026s
|
||||
### RED-PROOF hub: backup_catchup_done not allowed
|
||||
r871_catchup_event_test.go:10: backup_catchup_done is not an allowed event type — the box's catch-up line would be dropped with a 400
|
||||
--- FAIL: TestR871_CatchUpEventIsAllowed (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/api 0.020s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/api 0.019s
|
||||
@@ -0,0 +1,27 @@
|
||||
### 07:54:16 banner on 9202 — the ledger set by hand to a box whose last dump is 3 days old (scratch fixture, said so); window back to 02:30
|
||||
HTTP 303
|
||||
ledger set: last dump 2026-10-02T07:54:19.875183Z
|
||||
2026/10/05 07:54:20 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump tier2 offsite] (last scheduled 2026-10-05 04:15) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
--- the banner as GET /launcher served it (9202, Hungarian household):
|
||||
<div class="alerts-container">
|
||||
<div class="alert-banner alert-banner-warning" id="missed-backup-banner">
|
||||
<span class="alert-icon"><svg class="ico"><use href="#i-triangle-alert"/></svg></span>
|
||||
<span class="alert-message">A legutóbbi mentés 3 nappal ezelőtt készült. Válassz egy olyan időpontot, amikor a doboz általában be van kapcsolva.</span>
|
||||
<span class="alert-actions">
|
||||
<a href="/backups#window_start" class="btn btn-sm btn-primary">Mentési idő módosítása</a>
|
||||
<form method="POST" action="/backups/missed-banner/dismiss" style="display:inline">
|
||||
<input type="hidden" name="_csrf" value="(redacted)"><input type="hidden" name="back" value="/launcher">
|
||||
<button type="submit" class="btn btn-sm btn-outline">Bezárás</button>
|
||||
</form>
|
||||
</span>
|
||||
</div>
|
||||
</div
|
||||
### 07:54:54 POST /backups/missed-banner/dismiss (the Bezárás button)
|
||||
HTTP 302
|
||||
banner on /launcher after closing: 0
|
||||
2026/10/05 07:54:46 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:46 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:57 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:57 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:57 missed_backup_banner.go:100: [INFO] [web] missed-backup banner: closed by the household until the next missed backup time (after 2026-10-05T02:30:00+02:00)
|
||||
"banner_dismissed_through": "2026-10-05T00:30:00Z"
|
||||
@@ -0,0 +1,56 @@
|
||||
### RED-PROOF 1 stale is silent (the old behaviour)
|
||||
banner_test.go:32: not shown although the last backup is 2 days old
|
||||
--- FAIL: TestBanner_ShownWithReasonAndSuggestion (0.00s)
|
||||
banner_test.go:72: the banner did not come back after the next missed backup time
|
||||
--- FAIL: TestBanner_DismissedThenBackAfterANewMiss (0.00s)
|
||||
banner_test.go:92: banner = {Show:false LastBackup:2026-10-02 04:20:00 +0200 CEST DaysAgo:3 OffAt:02:30 Suggest:21:00 MissedAt:2026-10-05 02:30:00 +0200 CEST}, want shown with the off-site copy's
|
||||
--- FAIL: TestBanner_OffsiteTierCounts (0.00s)
|
||||
banner_test.go:107: banner = {Show:false LastBackup:2026-10-03 02:31:00 +0200 CEST DaysAgo:2 OffAt: Suggest: MissedAt:2026-10-05 02:30:00 +0200 CEST}, want shown, no suggestion, no 'off at'
|
||||
--- FAIL: TestBanner_NoPatternNoSuggestion (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### RED-PROOF 2 dismissal ignored
|
||||
banner_test.go:64: a closed banner came back without a new miss
|
||||
--- FAIL: TestBanner_DismissedThenBackAfterANewMiss (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### RED-PROOF 4 the seed ignored (banner on upgrade day)
|
||||
banner_test.go:53: shown on the day of the upgrade: {Show:true LastBackup:0001-01-01 00:00:00 +0000 UTC DaysAgo:0 OffAt:02:30 Suggest:21:00 MissedAt:2026-10-05 02:30:00 +0200 CEST}
|
||||
--- FAIL: TestBanner_NotShownBeforeTheRecordKnows (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### RED-PROOF 5 the off-site tier ignored
|
||||
banner_test.go:92: banner = {Show:false LastBackup:0001-01-01 00:00:00 +0000 UTC DaysAgo:0 OffAt: Suggest: MissedAt:0001-01-01 00:00:00 +0000 UTC}, want shown with the off-site copy's 3 days
|
||||
--- FAIL: TestBanner_OffsiteTierCounts (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.004s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
### RED-PROOF 3 a suggestion from noise (no day count) — test noise now evenings on 2 of 7 days
|
||||
banner_test.go:107: banner = {Show:true LastBackup:2026-10-03 02:31:00 +0200 CEST DaysAgo:2 OffAt: Suggest:21:00 MissedAt:2026-10-05 02:30:00 +0200 CEST}, want shown, no suggestion, no 'off at'
|
||||
--- FAIL: TestBanner_NoPatternNoSuggestion (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
### RED-PROOF 6 the banner not wired into executeTemplate
|
||||
r871_missed_backup_banner_test.go:37: not shown on /launcher although the last backup is 3 days old
|
||||
--- FAIL: TestR871_BannerShownThenClosedUntilTheNextMiss (0.06s)
|
||||
--- PASS: TestR871_BannerNotShownAfterASuccessfulNight (0.06s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.130s
|
||||
FAIL
|
||||
### RED-PROOF 7 the dismissal not recorded
|
||||
r871_missed_backup_banner_test.go:47: still shown after the household closed it
|
||||
--- FAIL: TestR871_BannerShownThenClosedUntilTheNextMiss (0.06s)
|
||||
--- PASS: TestR871_BannerNotShownAfterASuccessfulNight (0.07s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.145s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.148s
|
||||
@@ -0,0 +1,27 @@
|
||||
### RED-PROOF R-872a: the old skip of a down box
|
||||
r872_down_box_test.go:60: events map[] — a box off at every deadline must raise the missed-backup alarms, down or not
|
||||
--- FAIL: TestR872_DownEveryNightRaisesTheMissedAlarms (0.04s)
|
||||
--- PASS: TestR872_DownWithARecentDumpIsQuiet (0.04s)
|
||||
--- PASS: TestR872_NewDownBoxIsQuiet (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/monitor 0.118s
|
||||
FAIL
|
||||
### RED-PROOF R-872b: no grace for a new box
|
||||
--- PASS: TestR872_DownEveryNightRaisesTheMissedAlarms (0.04s)
|
||||
--- PASS: TestR872_DownWithARecentDumpIsQuiet (0.04s)
|
||||
r872_down_box_test.go:84: events map[expected_backup_missed:1 expected_dbdump_missed:1] for a box bound a day ago
|
||||
--- FAIL: TestR872_NewDownBoxIsQuiet (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/monitor 0.113s
|
||||
FAIL
|
||||
### RED-PROOF R-872c: the down dump line shortened to 12 h
|
||||
--- PASS: TestR872_DownEveryNightRaisesTheMissedAlarms (0.04s)
|
||||
r872_down_box_test.go:76: events map[expected_dbdump_missed:1] — a dump 20 h ago and a backup 30 h ago are inside the down box's lines
|
||||
--- FAIL: TestR872_DownWithARecentDumpIsQuiet (0.04s)
|
||||
r872_down_box_test.go:84: events map[expected_backup_missed:1 expected_dbdump_missed:1] for a box bound a day ago
|
||||
--- FAIL: TestR872_NewDownBoxIsQuiet (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/monitor 0.113s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/monitor 38.508s
|
||||
@@ -0,0 +1,8 @@
|
||||
### RED-PROOF R-873: the weekly household limit dropped
|
||||
r873_liveness_weekly_test.go:40: night 2: household 1 mails (want 0 — told once this week), operator 2 (want every edge)
|
||||
--- FAIL: TestR873_HouseholdHearsAnOutageAtMostWeekly (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/notify 0.043s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/notify 1.382s
|
||||
@@ -0,0 +1,11 @@
|
||||
Oct 05 09:38:46 demo-felhom felhom-agent[3840353]: time=2026-10-05T09:38:46.189+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled"
|
||||
Oct 05 09:38:46 demo-felhom felhom-agent[3935656]: time=2026-10-05T09:38:46.215+02:00 level=INFO msg="felhom-agent daemon starting" version=0.145.0 host_id=demo-felhom-8363b5 hub_url=https://hub.felhom.eu interval_s=900
|
||||
Oct 05 09:38:46 demo-felhom felhom-agent[3935656]: time=2026-10-05T09:38:46.836+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Oct 05 09:38:48 demo-felhom felhom-agent[3935656]: time=2026-10-05T09:38:48.047+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s
|
||||
Oct 05 10:08:46 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:46.837+02:00 level=INFO msg="backup: restore-test first evaluation after start (R-874)" after=30m0s
|
||||
Oct 05 10:08:47 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:47.156+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-backup archive=felhom-backup:backup/vzdump-lxc-9201-2026_10_04-07_49_00.tar.zst landed=2026-10-04T0
|
||||
Oct 05 10:08:47 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:47.948+02:00 level=INFO msg="restore-test: space preflight passed" storage=local-lvm required_bytes=8002555904 avail_bytes=365212749058
|
||||
Oct 05 10:08:47 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:47.972+02:00 level=INFO msg="restore-test: full-fidelity restore params derived from the archive config" scratch=990000 params=4
|
||||
Oct 05 10:08:48 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:48.047+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=restore-test
|
||||
Oct 05 10:09:16 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:09:16.524+02:00 level=INFO msg="restore-test: scratch guest torn down" vmid=990000
|
||||
Oct 05 10:09:16 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:09:16.524+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-backup:backup/vzdump-lxc-9201-2026_10_04-07_49_00.tar.zst duration_s=29.36536675
|
||||
@@ -0,0 +1,9 @@
|
||||
### RED-PROOF R-874: the first-evaluation timer dropped (back to the bare ticker)
|
||||
r874_first_eval_test.go:38: no evaluation within 3 s of start (FirstEval 30 ms) — a box with short sessions never gets a restore-test
|
||||
--- FAIL: TestR874_FirstEvaluationAfterStart (3.01s)
|
||||
--- PASS: TestR874_CrashLoopNeverEvaluates (0.10s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/backup 3.119s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/backup 0.151s
|
||||
@@ -0,0 +1,8 @@
|
||||
### RED-PROOF R-875: the v0.144.1 text back
|
||||
unsent_test.go:44: health_reason = "sent after the agent stopped mid-pass (R-868)", want the neutral "sent late …"
|
||||
--- FAIL: TestR868_KilledPassIsReportedAtStart (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.011s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 1.738s
|
||||
@@ -0,0 +1,12 @@
|
||||
### 07:55:21 before
|
||||
"armed": true,
|
||||
"last_boot_at": "2026-10-05T06:14:33Z",
|
||||
"tripped": false,
|
||||
"unclean_boots_in_window": 0,
|
||||
09:55:22 up 1:40, 1 user, load average: 2.60, 1.97, 1.49
|
||||
20
|
||||
279
|
||||
### 07:55:24 rollback (13 packages, simulated first)
|
||||
0 upgraded, 0 newly installed, 13 downgraded, 0 to remove and 0 not upgraded.
|
||||
0 upgraded, 0 newly installed, 13 downgraded, 0 to remove and 0 not upgraded.
|
||||
audit rc=0
|
||||
@@ -0,0 +1,5 @@
|
||||
07:55:41.640 debug pass started
|
||||
dpkg running at 07:56:02.802: 302936 /usr/bin/dpkg --force-confold --force-confdef --status-fd 24 --no-triggers --unpack --auto-deconfigure /var/cache
|
||||
CRASH at 07:56:03.308
|
||||
Timeout, server 192.168.0.104 not responding.
|
||||
ssh ended 07:56:10
|
||||
@@ -0,0 +1,25 @@
|
||||
ssh back 07:56:56
|
||||
09:56:56 up 0 min, 1 user, load average: 2.97, 0.65, 0.21
|
||||
"armed": true,
|
||||
"last_boot_at": "2026-10-05T07:56:41Z",
|
||||
"last_boot_unclean": true,
|
||||
"tripped": false,
|
||||
"unclean_boots_in_window": 1,
|
||||
kernel.panic = 10
|
||||
--- dpkg at boot
|
||||
audit rc=0
|
||||
1
|
||||
bind9-dnsutils 1:9.20.26-1~deb13u1 ii
|
||||
bind9-host 1:9.20.26-1~deb13u1 ii
|
||||
bind9-libs 1:9.20.26-1~deb13u1 ii
|
||||
libpcre2-8-0 10.46-1~deb13u2 ii
|
||||
libpython3.13-minimal 3.13.5-2+deb13u4 ii
|
||||
libpython3.13-stdlib 3.13.5-2+deb13u4 ii
|
||||
libssh2-1t64 1.11.1-1+deb13u1 ii
|
||||
libssl3t64 3.5.7-1~deb13u2 ii
|
||||
libxml2 2.12.7+dfsg+really2.9.14-2.1+deb13u1 ii
|
||||
openssl 3.5.7-1~deb13u2 ii
|
||||
openssl-provider-legacy 3.5.7-1~deb13u3 ii
|
||||
python3.13 3.13.5-2+deb13u4 ii
|
||||
python3.13-minimal 3.13.5-2+deb13u4 ii
|
||||
08:00:01 apps: 20 of 21 healthy
|
||||
@@ -0,0 +1,25 @@
|
||||
### 08:00:01 the next pass — no person touched dpkg
|
||||
=== felhom-agent 0.145.0 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true block=hub ===
|
||||
"healthy": true,
|
||||
"outcome": "applied",
|
||||
"healthy": true,
|
||||
"outcome": "nothing",
|
||||
"healthy": true,
|
||||
"outcome": "nothing",
|
||||
pass took 1m20.7s
|
||||
Oct 05 10:00:02 demo-hp felhom-os-apply[16364]: os-apply: START release=ring0-20261005T080001Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0
|
||||
Oct 05 10:00:13 demo-hp felhom-os-apply[16840]: os-apply: REPAIR configured=0 journal=1 fixed=0
|
||||
Oct 05 10:00:19 demo-hp felhom-os-apply[17215]: os-apply: PLAN upgrade=12 already=0 not-installed=0 from-snapshot=0
|
||||
Oct 05 10:00:34 demo-hp felhom-os-apply[18302]: os-apply: DONE rc=0 seconds=7.7 upgraded=12 restart-needed=cron,dbus-daemon,sshd,systemd,systemd-journal,systemd-logind,systemd-network reboot-needed=yes
|
||||
Oct 05 10:00:43 demo-hp felhom-os-apply[18833]: os-apply: START release=ring0-20261005T080001Z layer=host lane=fast mode=apply select=pending-fast packages=0
|
||||
Oct 05 10:00:47 demo-hp felhom-os-apply[19271]: os-apply: REPAIR configured=0 journal=0 fixed=0
|
||||
Oct 05 10:00:51 demo-hp felhom-os-apply[19388]: os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0
|
||||
Oct 05 10:00:51 demo-hp felhom-os-apply[19389]: os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)
|
||||
Oct 05 10:01:04 demo-hp felhom-os-apply[20758]: os-apply: START release=ring0-20261005T080001Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0
|
||||
Oct 05 10:01:09 demo-hp felhom-os-apply[20997]: os-apply: REPAIR configured=0 journal=0 fixed=0
|
||||
Oct 05 10:01:14 demo-hp felhom-os-apply[21238]: os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0
|
||||
Oct 05 10:01:14 demo-hp felhom-os-apply[21239]: os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)
|
||||
audit rc=0
|
||||
0
|
||||
0
|
||||
package list == before the rollback (279 lines)
|
||||
@@ -0,0 +1,8 @@
|
||||
--- hub events demo-hp since 07:55 UTC
|
||||
('controller_started', 'info', 'Controller elindult (0.295.0)', 'controller', '2026-10-05 07:58:50')
|
||||
('os_update_applied', 'info', 'System security fixes installed on the box (12 package(s)).', 'hub', '2026-10-05 08:00:42')
|
||||
--- notification_log since 07:55 UTC (all customers)
|
||||
--- os_reports demo-hp newest 3
|
||||
(73, '2026-10-05 08:01:22', 'debug', 'nothing', 1, 'docker')
|
||||
(72, '2026-10-05 08:01:00', 'debug', 'nothing', 1, 'host')
|
||||
(71, '2026-10-05 08:00:42', 'debug', 'applied', 1, 'guest')
|
||||
@@ -0,0 +1,17 @@
|
||||
### RED-PROOF R-876a: the journal does not count (v0.144.1)
|
||||
FAIL: test_the_next_pass_repairs_by_itself_and_finishes (test_felhom_os_apply.CrashLeftTheJournal.test_the_next_pass_repairs_by_itself_and_finishes)
|
||||
AssertionError: 18 not less than 16 : the repair must run before the install
|
||||
Ran 3 tests in 0.032s
|
||||
FAILED (failures=1)
|
||||
### RED-PROOF R-876b: no repair+retry when apt says interrupted
|
||||
FAIL: test_apt_interrupted_is_repaired_and_retried_once (test_felhom_os_apply.CrashLeftTheJournal.test_apt_interrupted_is_repaired_and_retried_once)
|
||||
AssertionError: 3 != 0 : {'failed': {'dpkg_audit': 'clean', 'rc': 100, 'tail': ["E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a' to correct the problem."]}, 'health_before': {'containers': {'app': {'health': 'healthy', 'id': 'bbb222', 'state': 'running'}, 'felhom-controller': {'health': 'healthy', 'id': 'aaa111', 'state': 'running'}}, 'controller': 'healthy', 'controller_docker_ok': True, 'docker_ok': True, 'network_ok': True}, 'layer': 'guest', 'mode': 'apply', 'pass_seconds': 0.0, 'plan': {'already': 0, 'from_snapshot': 0, 'not_installed': 0, 'upgrade': 2}, 'refused': None, 'release_id': 'os-t1', 'repair': {'clean_after': True, 'fixed': 0, 'half_configured_before': 0, 'journal_before': 0}, 'vmid': 9201}
|
||||
Ran 3 tests in 0.026s
|
||||
FAILED (failures=1)
|
||||
### RED-PROOF R-876c: repair always runs (R-845 speed lost)
|
||||
FAIL: test_a_clean_pass_still_costs_one_state_call (test_felhom_os_apply.CrashLeftTheJournal.test_a_clean_pass_still_costs_one_state_call)
|
||||
AssertionError: 2 != 1 : R-845: a clean pass reads dpkg's state ONCE and runs no repair
|
||||
Ran 3 tests in 0.035s
|
||||
FAILED (failures=1)
|
||||
### restored
|
||||
OK
|
||||
@@ -0,0 +1,3 @@
|
||||
artifacts POST HTTP 303
|
||||
Location: /configuration?flash=artifacts_set
|
||||
2026/10/05 09:48:14 [INFO] Artifact manifest set: agent=0.145.0 golden=0.295.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="78c00adce662d2d966b2ac50ebde46cde1ae225f0107c6a7c02b70ec8ce80c4f"
|
||||
@@ -0,0 +1,3 @@
|
||||
### 07:31:06 found: VM 341 (Tester 1) stopped since the 06:14 UTC demo-hp crash (no onboot)
|
||||
status: running
|
||||
onboot: 1
|
||||
@@ -26,6 +26,17 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-05 (afternoon) — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0; rulings 109–111, CC decisions 112–118)
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-871** | **No architecture covered a box that is not always on, and a missed night was never made up.** Decision 109 (option A) built: controller v0.295.0 `internal/nightchain` — a ledger of when each backup leg ran to its end; on a start or a host resume ONE catch-up 15 min later, backup legs only, never the update leg; a late daily timer after a suspend is skipped; the whole-guest backup and the catch-up wait for each other; decision 110's banner. Design `07` §6.1.1. Live: 9202 (dump made 15 min after the start; after a crash mid-wait, all three legs at the next start), demo-felhom (dump 15 min after the start; the household's timeline line reached the hub); banner served on 9202 and closed by its real route. 13 red-proofs. **Reasoning kept: the ledger records that a leg RAN, and is never read as evidence that a backup EXISTS.** Full text: `git show 1b0678fa:documentation/backlog/OPEN-ITEMS.md`. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partA/`, `partB/` |
|
||||
| **R-873** | **A household whose box is off every night was mailed "cannot be reached" every night.** hub v0.134.0: at most once per 7 days to the household (persisted), the operator every edge, the recovery mail stays paired (decision 116). Proven by test through the real dispatcher (red-proof); no live occurrence in the session (Tester 2 stayed off). | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r873-red-proof.txt` |
|
||||
| **R-874** | **A restore-test never ran on a box with short power-on sessions.** agent v0.145.0: first due-check 30 min after start (decision 117). Live on demo-felhom: start 07:38:46 UTC → `restore-test first evaluation after start (R-874)` at 08:08:46 → a due tier restored and passed in 29 s. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r874-*` |
|
||||
| **R-875** | **A kept report's reason said "the agent stopped mid-pass" for a hub-away pass.** agent v0.145.0: "sent late — kept on the box until the hub could take it". Test + red-proof. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r875-red-proof.txt` |
|
||||
| **R-876** | **After a power cut mid-update every later pass failed until a person ran `dpkg --configure -a`.** agent v0.145.0: dpkg's state = `--audit` AND the update journal in one call; repair on either; belt: repair + retry once when apt says "interrupted" (decision 118). Live (operator's go): crash at 07:56:03 UTC mid-unpack → back by itself → next pass `REPAIR configured=0 journal=1` → `DONE rc=0 upgraded=12`, no person, no mail; package list identical. **Reasoning kept: a check that reads one of two places dpkg keeps its state is a check that misses the other.** | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partD/` |
|
||||
| **R-877** | **The Tester 1 VM on demo-hp had no start-on-boot: the morning's demo-hp crash (06:14 UTC) left it off for 1 h 17 min, unnoticed** (the night-fixes report called every box healthy). Found 07:31 UTC; `qm set 341 --onboot 1`, started; the afternoon crash then brought it back by itself. Filed and closed in the same commit. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt` |
|
||||
|
||||
## 2026-10-05 (day) — the night's fixes: off-site clean-up guard, first-install image race, R8 download, a killed pass's report (controller v0.294.0, agent v0.144.0 + v0.144.1, golden 0.294.0; rulings 100–103, CC decisions 104–108)
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|
||||
@@ -191,7 +191,7 @@ stopping line that lies.
|
||||
| **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC |
|
||||
| **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator |
|
||||
|
||||
## Backup & restore — 53 rows (P2 9, P3 24, P4 20)
|
||||
## Backup & restore — 52 rows (P2 8, P3 23, P4 21)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -248,8 +248,7 @@ stopping line that lies.
|
||||
| **R-815** | Backup & restore | P4 | First-ever **GC** on `felhom-offsite` (armed today 13:11 UTC, never run) | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-815. WATCHING — no completion record found; schedule `sun 04:30` still present 2026-09-30.) — WATCHING | schedule | **Sun 2026-08-02 04:30 UTC** — confirm it completes | CC |
|
||||
| **R-816** | Backup & restore | P4 | **No off-site failure class has ever been seen live.** F-DIAG (controller v0.182.0, 2026-07-28) split off-site failures into six causes — quota, orphaned, no_repo, no_units, transport, unknown — each with its own Hungarian message. None of the six has been exercised by a real failure on a box; the recovery inventory records it only as a known limit (`documentation/architecture/_recovery-inventory-2026-07-28.md:955`). Filed 2026-10-03 from the F-DIAG row's residue when that row moved to `CLOSED-ITEMS.md`. | **READY — filed 2026-10-03 (triage); owner: CC.** Exercise each class once on a scratch guest (a full quota, a missing repository, a blocked transport) and read the message the household sees. | — | — | CC |
|
||||
| **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator |
|
||||
| **R-871** | Backup & restore | P2 | **No architecture document covers a box that is not always on, and the product has no catch-up for a missed night: a box that is OFF at its night window (Tester 2, a laptop switched off at night — operator 2026-10-05) never gets its database dumps, second copy, off-site copy or app updates.** FOUND 2026-10-05 (Part F spike, read only): the controller's daily jobs always schedule the NEXT future time (`controller/internal/scheduler/scheduler.go` `nextDailyRun`; `LastRun` in memory only, never consulted) — a missed 02:30/03:30/04:15 waits for the next night, for ever. Only the whole-guest backup catches up (outside its window only by the 48 h safety valve — about one every 2 days for an evening-only box), and the OS leg follows it by day under the `night` label. Nothing on the household's pages says the box must stay on at night. `07` §6.1 describes the night chain and the [W+2h, W+6h) gate but no catch-up and no safety valve; the intent lives only in controller comments (`quiesce.go`). **A promise to households — the operator's choice (STATUS, options A/B).** `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **WAITING-ON-OPERATOR — A (catch-up) or B (say plainly the box must be on at night)** | — | the decision, then a design section in `07` | operator |
|
||||
| **R-874** | Backup & restore | P3 | **A restore-test never runs on a box whose power-on sessions are all shorter than 6 hours.** FOUND 2026-10-05 (Part F spike, source): the agent checks every 6 h, the timer restarts at each agent start, and it does not check at start (`felhom-agent` restore-test schedule). Tester 2's sessions were ~1.5 h and ~5 min. `restore_test_stale` will then fire (~2026-10-11 for the local tier) with no visible cause for the household or the operator. Fix direction: check once shortly after start when the last test is overdue. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871 | — | CC |
|
||||
| **R-878** | Backup & restore | P4 | **A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer.** MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of `health: starting`; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. `audits/catchup-2026-10-05/partA/live-demo-felhom.txt` | **READY — owner: CC** | — | measure a large volume first | CC |
|
||||
|
||||
## Storage & devices — 12 rows (P3 7, P4 5)
|
||||
|
||||
@@ -305,7 +304,7 @@ stopping line that lies.
|
||||
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
|
||||
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
|
||||
|
||||
## Box system & updates — 20 rows (P2 4, P3 13, P4 3)
|
||||
## Box system & updates — 18 rows (P2 3, P3 13, P4 2)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -330,10 +329,8 @@ stopping line that lies.
|
||||
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |
|
||||
| **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
||||
| **R-875** | Box system & updates | P4 | **A kept OS report sent after the hub was away says "sent after the agent stopped mid-pass (R-868)" — wrong for that case: the agent did not stop, the hub was unreachable.** MEASURED 2026-10-05 05:53 UTC on demo-felhom (agent v0.144.1): the three reports of the R-866 hub-away pass reached the hub at the new agent's start with that reason text. The kept copy cannot tell the two causes apart. Fix direction: a neutral reason ("sent late — kept on the box until the hub could take it"), or the agent records which case it was. `audits/night-fixes-2026-10-05/partD/r866-kept-copies-sent-at-start.txt` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-876** | Box system & updates | P2 | **After a power cut in the middle of an OS update, every later OS pass FAILS until a person runs `dpkg --configure -a`: the wrapper's repair step runs only when `dpkg --audit` shows something, but a crash can leave dpkg's update journal (`/var/lib/dpkg/updates/`) non-empty with `--audit` clean — and that journal is exactly what apt refuses on.** MEASURED 2026-10-05 on demo-hp (Part E, the night's A1 by day, agent v0.144.1): crash at 06:13:55 UTC while dpkg ran; back by itself in 37 s (crash guard armed, 1 unclean boot); at boot `dpkg --audit` clean, all 13 packages `ii` (1 new, 12 old), but `/var/lib/dpkg/updates/` held 3 files; the next pass logged `REPAIR configured=0 fixed=0`, then `FAILED rc=100 step=install` with `E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a'`; operator mail `os_update_failed` (true). A by-hand `dpkg --configure -a` (rc 0) and the next pass finished (12 installed). Cause: R-845's speed-up skips the repair on a clean audit (`configs/felhom-os-apply` `repair()`). Without the fix a box that loses power mid-update fails its OS leg every night and mails the operator every night. Fix direction: also repair when `/var/lib/dpkg/updates/` is non-empty (one `ls`, no cost on a clean pass), pinned by a test with that exact state. `audits/night-fixes-2026-10-05/partE/` | **READY — owner: CC** (next agent release) | — | fix + the crash-state test | CC |
|
||||
|
||||
## Monitoring & notifications — 27 rows (P2 3, P3 17, P4 7)
|
||||
## Monitoring & notifications — 26 rows (P2 3, P3 16, P4 7)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -362,8 +359,7 @@ stopping line that lies.
|
||||
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
||||
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
||||
| **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator |
|
||||
| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871 | — | CC |
|
||||
| **R-873** | Monitoring & notifications | P3 | **A household whose box is off every night on purpose is mailed "Your server cannot be reached." every night and "reachable again" every morning; the operator gets about 6 mails a day for it.** FOUND 2026-10-05 (Part F spike): `node_down` after 90 min, mailed to the customer (Tester 2: 2026-10-04 19:36:42 UTC); the 6 h cooldown does not stop a daily repeat (`hub/internal/notify/dispatcher.go`). The R-285 gap (no notion of expected downtime) now reaching a real household. Fix direction depends on R-871: if B, a per-box "off at night" setting that quiets node_down within a nightly window; if A, the same plus the catch-up. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871, R-285 | — | CC |
|
||||
| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC |
|
||||
|
||||
## Hub & operator — 23 rows (P2 1, P3 7, P4 15)
|
||||
|
||||
@@ -506,4 +502,5 @@ stopping line that lies.
|
||||
the R-row. Duplicating them here would create the second source this design avoids. -->
|
||||
| item | due (UTC) | what to measure |
|
||||
|---|---|---|
|
||||
| R-872 | 2026-10-06 | the first live 05:00 deadline run judges a down box on the longer lines (Tester 2, if still off): hub log + the two events (detail in the R-872 row) |
|
||||
<!-- DUE-CHECKS-END -->
|
||||
|
||||
Reference in New Issue
Block a user