# REPORT — R-331: the operator Backup card said every customer had no backups **Hub v0.109.0 (with controller v0.225.0) · 2026-08-30** --- ## 1. What was wrong The hub customer page's **Backup** card read, for **every customer, indefinitely**: ``` Enabled Yes Snapshots 0 Repo Size 0 MB Integrity Unknown ``` Measured on `demo-hp` 2026-08-30, at which moment the truth was: | source | value | |---|---| | the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` | | that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` | | **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report | **A card that reads "no backups" over a working backup is worse than no card.** It is the R-88 direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an operator consults to answer "is this customer protected?". ## 2. Root cause The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and `integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice 8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment. The zeros were correct values for dead fields, rendered as if live. **The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker` **already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes were arriving. This is a **render fix over an existing feed**, not a new pipeline. ## 3. Why it was not a one-line template swap `snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered „Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's `StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have **moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards `stats_known`. ## 4. What changed `hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot keep: | report state | card shows | |---|---| | no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** | | `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" | | enabled, `stats_known:false` | **—**, plus "never been measured". Never `0` | | enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge | **A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every un-upgraded customer has zero backups. Pinned by a test. **The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten field would be a lie. `fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose entire defect was under-reporting a real backup reads as "nearly nothing". ## 5. Tests and the red-proof `r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a regression fails against the same numbers the defect was measured against. **The defect lived in the template's choice of source object, so a test one layer below it would have been green against the shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML. **RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests — `the rendered Backup card does not contain demo-hp's real snapshot count (67)`, `the card does not carry the real repository size (134.3 MB ...)`, `the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored immediately; `git diff` clean. **Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0. ## 6. Deployment PENDING at the time of writing — see the follow-up commit. The two halves ship independently and in either order: the hub renders "unknown" for any box still on controller 0.224.0, which is correct rather than wrong. ## 7. The push bypassed a gate, deliberately, and here is the declaration **`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden bake carries **0.223.0**, so a machine installed right now receives neither fix. **This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check: **it does not extend to a release that changes first-boot behaviour.** **A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change, `MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases wide rather than one. **The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement, five days overdue. It was **taken** during this session — see §8. ## 8. R-341's overdue check was taken, and its premise did not survive Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence: `documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. **Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still `MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`, which reads 03:54:54Z here — R-346's trap, avoided.) **Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0, CLOSE-WAIT 0. **The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1 upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent 0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the other way would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon. The question is now **moot**, and the row is closed as such. **What it does establish, which is worth more than the original question:** twelve days after the R-344 fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway concern retires with it. ## 9. Not done, and why - **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A second verdict over the same data is two things that can disagree — a shape this codebase has already paid for (`LastRun` vs `LastSuccess`, R-100). - **The dead `backup` fields were not removed from the controller's wire format.** Removing them would stop historical reports already in this hub's store from parsing, for no gain — nothing renders them now, and a controller-side test fails if anything starts producing them.