Files
felhom-controller/REPORT.md
T

5.7 KiB

REPORT — R-182: one operator email per backup run, and nothing dropped without a trace

Date: 2026-08-03 · Repo: felhom-controller · v0.193.1 → v0.194.0 · commit 88897a2 Baseline on arrival: db0d4b129d3b, v0.193.1 — matched §1.

The defect

Nine per-app recovery_unit_capture_failed events reached the hub on 2026-08-03 and two operator emails went out. The hub's operator cooldown key is customerID:eventType(+tier), and that event carries app but no tier — so the key held no app identifier. The first refused app took the hour's slot and every other app's failure was discarded before LogNotification, leaving no row on any channel.

The obvious fix was ruled against: putting app in the key produces one email per failing app, which on a full disk is a dozen.

What ships

internal/backup/runsummary.go — a per-run collector with exactly admissionSet's lifetime, fed by all three write legs, emitting backup_run_failures once at the end and only when something failed. The per-app event stays and becomes the record.

  • A refusal is noted ONCE, inside admitApp where the verdict is taken — not at the three legs that consult it. R-181's contract is one verdict per app per run; noting per leg listed a single refused app three times and produced "2 of 1 apps failed". Found by the digest's own test.
  • Deliberate skips are excluded — a disconnected or decommissioned drive has its own alert.
  • A manual run always reports: a unique run_id the hub's cooldown cannot collapse.

A gap the spec did not anticipate, and its fix. recovery_unit_capture_failed also fires from GetFullStatus's periodic sweep, outside any run. With it now record-only, those failures would have been recorded and never notified — a new silence created while closing one. The sweep therefore emits a digest too, deliberately with no run_id, so it stays under the ordinary hourly cooldown exactly as before while the mail now lists every failing app instead of whichever was first.

§3's safety property — confirmed, not assumed

hub/internal/monitor/deadline.go:396 (expected_backup_missed) and :417 (expected_dbdump_missed) are raised by the hub, from assessBackupFreshness(reportJSON,…) and GetEventsByType — i.e. from the box's report freshness and stored events, independently of any email the controller chooses to send. A digest's silence therefore still means "the run finished and found nothing wrong". This is what makes the whole design safe, and weakening that check re-opens a silent-failure path.

Files

internal/backup/runsummary.go (new) · runsummary_test.go (new) · admission.go · admission_test.go · backup.go · recovery_unit.go · internal/notify/notifier.go · internal/web/handler_debug.go · internal/api/router.go · cmd/controller/main.go · CHANGELOG.md · REUSE.md

Tests — go build && go vet && go test ./...28 packages ok, rc=0

7 new tests. controller_gates.py → all 8 OK.

Red-proofs

# Mutation Result
A all four noteFailure feeds removed (the pre-R-182 per-app-only path) RED — Scenario A and the refresh-sweep test both fail; the digest is never emitted
I (1st attempt) the main.go seam wiring commented out DID NOT FAIL — recorded as such. The AST test walked the backup package and not main.go, so the seam could be disconnected with the suite green. The test was fixed, not the result recorded
I (re-run) same mutation, against the fixed test RED — and the string SetRunSummaryNotify is still present in the file, which is why this is an AST walk and not strings.Contains

Live proof — demo-hp guest 9201

Method: endpoint-level — POST /api/debug/backup/dbdump, the exact endpoint the debug UI button calls, running the production RunDBDumps. No browser on DooPlex.

Filled for real with fallocate to 241 MB free / 100% used; the thin pool held 30.78 → 30.78 (instrument re-proven before use — demo-hp's pool is 53.93 GiB and a genuine fill would exhaust it).

[INFO] [backup] Run summary: 2 of 2 apps failed (manual run) — notifying the operator once
[INFO] Event pushed: backup_run_failures (error) — 2 of 2 apps failed to back up in this manual run: opengist, privatebin

Both apps refused, one digest naming both. Against the previous behaviour: two refusals, one email naming one app, one vanishing. Hub-side evidence is in felhom.eu/REPORT.md.

Scenario E: a second run in the same hour produced a second digest (11:50:45 and 11:51:28). Scenario B: after freeing the space, the run completed with 2 volume dumps and no digest.

Teardown

Fill file removed, guest helper scripts and the credential file shred-ed, pct fstrim 9201 returned 63.3 GiB, thin pool at 30.76. Disk back to 1.9G used / 64G free.

Observations — NOT acted on

  1. Only two apps are deployed on the demo box, so "several apps" was 2 of 2. That is exactly the measured case being closed (two refusals → one email), but a five-app demonstration would be stronger and needs a box with more apps.
  2. FormatOperatorEmail puts an emoji in every operator subject, including this digest. Part 3 asked for no emoji; the icon comes from the shared formatter, and changing it globally would alter every other operator mail and its tests. Left alone deliberately — the digest's own copy has none.
  3. The digest's leg name for a refusal is whole app (refused before any write) rather than a specific leg, because the reserve refuses all three at once. Accurate, but longer than the other leg names and it widens the column.