catch-up session 2026-10-05: design 07 §6.1.1 (a box that is not always on), 08 §6.4; rulings 109-111, CC decisions 112-118; R-871/R-873..R-877 closed, R-872 narrowed (dated check), R-878 opened; live evidence; STATUS
gates / gates (push) Successful in 33s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 10:29:22 +02:00
parent ff25f1076d
commit 9bb45eaaa2
34 changed files with 790 additions and 18 deletions
+11
View File
@@ -26,6 +26,17 @@
---
## 2026-10-05 (afternoon) — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0; rulings 109–111, CC decisions 112–118)
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-871** | **No architecture covered a box that is not always on, and a missed night was never made up.** Decision 109 (option A) built: controller v0.295.0 `internal/nightchain` — a ledger of when each backup leg ran to its end; on a start or a host resume ONE catch-up 15 min later, backup legs only, never the update leg; a late daily timer after a suspend is skipped; the whole-guest backup and the catch-up wait for each other; decision 110's banner. Design `07` §6.1.1. Live: 9202 (dump made 15 min after the start; after a crash mid-wait, all three legs at the next start), demo-felhom (dump 15 min after the start; the household's timeline line reached the hub); banner served on 9202 and closed by its real route. 13 red-proofs. **Reasoning kept: the ledger records that a leg RAN, and is never read as evidence that a backup EXISTS.** Full text: `git show 1b0678fa:documentation/backlog/OPEN-ITEMS.md`. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partA/`, `partB/` |
| **R-873** | **A household whose box is off every night was mailed "cannot be reached" every night.** hub v0.134.0: at most once per 7 days to the household (persisted), the operator every edge, the recovery mail stays paired (decision 116). Proven by test through the real dispatcher (red-proof); no live occurrence in the session (Tester 2 stayed off). | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r873-red-proof.txt` |
| **R-874** | **A restore-test never ran on a box with short power-on sessions.** agent v0.145.0: first due-check 30 min after start (decision 117). Live on demo-felhom: start 07:38:46 UTC → `restore-test first evaluation after start (R-874)` at 08:08:46 → a due tier restored and passed in 29 s. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r874-*` |
| **R-875** | **A kept report's reason said "the agent stopped mid-pass" for a hub-away pass.** agent v0.145.0: "sent late — kept on the box until the hub could take it". Test + red-proof. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partC/r875-red-proof.txt` |
| **R-876** | **After a power cut mid-update every later pass failed until a person ran `dpkg --configure -a`.** agent v0.145.0: dpkg's state = `--audit` AND the update journal in one call; repair on either; belt: repair + retry once when apt says "interrupted" (decision 118). Live (operator's go): crash at 07:56:03 UTC mid-unpack → back by itself → next pass `REPAIR configured=0 journal=1` → `DONE rc=0 upgraded=12`, no person, no mail; package list identical. **Reasoning kept: a check that reads one of two places dpkg keeps its state is a check that misses the other.** | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/partD/` |
| **R-877** | **The Tester 1 VM on demo-hp had no start-on-boot: the morning's demo-hp crash (06:14 UTC) left it off for 1 h 17 min, unnoticed** (the night-fixes report called every box healthy). Found 07:31 UTC; `qm set 341 --onboot 1`, started; the afternoon crash then brought it back by itself. Filed and closed in the same commit. | CLOSED 2026-10-05 — FIXED | `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt` |
## 2026-10-05 (day) — the night's fixes: off-site clean-up guard, first-install image race, R8 download, a killed pass's report (controller v0.294.0, agent v0.144.0 + v0.144.1, golden 0.294.0; rulings 100–103, CC decisions 104–108)
| Row | What | Closed | Evidence |
+6 -9
View File
@@ -191,7 +191,7 @@ stopping line that lies.
| **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC |
| **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator |
## Backup & restore — 53 rows (P2 9, P3 24, P4 20)
## Backup & restore — 52 rows (P2 8, P3 23, P4 21)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -248,8 +248,7 @@ stopping line that lies.
| **R-815** | Backup & restore | P4 | First-ever **GC** on `felhom-offsite` (armed today 13:11 UTC, never run) | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-815. WATCHING — no completion record found; schedule `sun 04:30` still present 2026-09-30.) — WATCHING | schedule | **Sun 2026-08-02 04:30 UTC** — confirm it completes | CC |
| **R-816** | Backup & restore | P4 | **No off-site failure class has ever been seen live.** F-DIAG (controller v0.182.0, 2026-07-28) split off-site failures into six causes — quota, orphaned, no_repo, no_units, transport, unknown — each with its own Hungarian message. None of the six has been exercised by a real failure on a box; the recovery inventory records it only as a known limit (`documentation/architecture/_recovery-inventory-2026-07-28.md:955`). Filed 2026-10-03 from the F-DIAG row's residue when that row moved to `CLOSED-ITEMS.md`. | **READY — filed 2026-10-03 (triage); owner: CC.** Exercise each class once on a scratch guest (a full quota, a missing repository, a blocked transport) and read the message the household sees. | — | — | CC |
| **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator |
| **R-871** | Backup & restore | P2 | **No architecture document covers a box that is not always on, and the product has no catch-up for a missed night: a box that is OFF at its night window (Tester 2, a laptop switched off at night — operator 2026-10-05) never gets its database dumps, second copy, off-site copy or app updates.** FOUND 2026-10-05 (Part F spike, read only): the controller's daily jobs always schedule the NEXT future time (`controller/internal/scheduler/scheduler.go` `nextDailyRun`; `LastRun` in memory only, never consulted) — a missed 02:30/03:30/04:15 waits for the next night, for ever. Only the whole-guest backup catches up (outside its window only by the 48 h safety valve — about one every 2 days for an evening-only box), and the OS leg follows it by day under the `night` label. Nothing on the household's pages says the box must stay on at night. `07` §6.1 describes the night chain and the [W+2h, W+6h) gate but no catch-up and no safety valve; the intent lives only in controller comments (`quiesce.go`). **A promise to households — the operator's choice (STATUS, options A/B).** `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **WAITING-ON-OPERATOR — A (catch-up) or B (say plainly the box must be on at night)** | — | the decision, then a design section in `07` | operator |
| **R-874** | Backup & restore | P3 | **A restore-test never runs on a box whose power-on sessions are all shorter than 6 hours.** FOUND 2026-10-05 (Part F spike, source): the agent checks every 6 h, the timer restarts at each agent start, and it does not check at start (`felhom-agent` restore-test schedule). Tester 2's sessions were ~1.5 h and ~5 min. `restore_test_stale` will then fire (~2026-10-11 for the local tier) with no visible cause for the household or the operator. Fix direction: check once shortly after start when the last test is overdue. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871 | — | CC |
| **R-878** | Backup & restore | P4 | **A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer.** MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of `health: starting`; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. `audits/catchup-2026-10-05/partA/live-demo-felhom.txt` | **READY — owner: CC** | — | measure a large volume first | CC |
## Storage & devices — 12 rows (P3 7, P4 5)
@@ -305,7 +304,7 @@ stopping line that lies.
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
## Box system & updates — 20 rows (P2 4, P3 13, P4 3)
## Box system & updates — 18 rows (P2 3, P3 13, P4 2)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -330,10 +329,8 @@ stopping line that lies.
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |
| **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC |
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
| **R-875** | Box system & updates | P4 | **A kept OS report sent after the hub was away says "sent after the agent stopped mid-pass (R-868)" — wrong for that case: the agent did not stop, the hub was unreachable.** MEASURED 2026-10-05 05:53 UTC on demo-felhom (agent v0.144.1): the three reports of the R-866 hub-away pass reached the hub at the new agent's start with that reason text. The kept copy cannot tell the two causes apart. Fix direction: a neutral reason ("sent late — kept on the box until the hub could take it"), or the agent records which case it was. `audits/night-fixes-2026-10-05/partD/r866-kept-copies-sent-at-start.txt` | **READY — owner: CC** | — | — | CC |
| **R-876** | Box system & updates | P2 | **After a power cut in the middle of an OS update, every later OS pass FAILS until a person runs `dpkg --configure -a`: the wrapper's repair step runs only when `dpkg --audit` shows something, but a crash can leave dpkg's update journal (`/var/lib/dpkg/updates/`) non-empty with `--audit` clean — and that journal is exactly what apt refuses on.** MEASURED 2026-10-05 on demo-hp (Part E, the night's A1 by day, agent v0.144.1): crash at 06:13:55 UTC while dpkg ran; back by itself in 37 s (crash guard armed, 1 unclean boot); at boot `dpkg --audit` clean, all 13 packages `ii` (1 new, 12 old), but `/var/lib/dpkg/updates/` held 3 files; the next pass logged `REPAIR configured=0 fixed=0`, then `FAILED rc=100 step=install` with `E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a'`; operator mail `os_update_failed` (true). A by-hand `dpkg --configure -a` (rc 0) and the next pass finished (12 installed). Cause: R-845's speed-up skips the repair on a clean audit (`configs/felhom-os-apply` `repair()`). Without the fix a box that loses power mid-update fails its OS leg every night and mails the operator every night. Fix direction: also repair when `/var/lib/dpkg/updates/` is non-empty (one `ls`, no cost on a clean pass), pinned by a test with that exact state. `audits/night-fixes-2026-10-05/partE/` | **READY — owner: CC** (next agent release) | — | fix + the crash-state test | CC |
## Monitoring & notifications — 27 rows (P2 3, P3 17, P4 7)
## Monitoring & notifications — 26 rows (P2 3, P3 16, P4 7)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -362,8 +359,7 @@ stopping line that lies.
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
| **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator |
| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871 | — | CC |
| **R-873** | Monitoring & notifications | P3 | **A household whose box is off every night on purpose is mailed "Your server cannot be reached." every night and "reachable again" every morning; the operator gets about 6 mails a day for it.** FOUND 2026-10-05 (Part F spike): `node_down` after 90 min, mailed to the customer (Tester 2: 2026-10-04 19:36:42 UTC); the 6 h cooldown does not stop a daily repeat (`hub/internal/notify/dispatcher.go`). The R-285 gap (no notion of expected downtime) now reaching a real household. Fix direction depends on R-871: if B, a per-box "off at night" setting that quiets node_down within a nightly window; if A, the same plus the catch-up. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **READY — owner: CC** (after R-871's decision) | R-871, R-285 | — | CC |
| **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC |
## Hub & operator — 23 rows (P2 1, P3 7, P4 15)
@@ -506,4 +502,5 @@ stopping line that lies.
the R-row. Duplicating them here would create the second source this design avoids. -->
| item | due (UTC) | what to measure |
|---|---|---|
| R-872 | 2026-10-06 | the first live 05:00 deadline run judges a down box on the longer lines (Tester 2, if still off): hub log + the two events (detail in the R-872 row) |
<!-- DUE-CHECKS-END -->