36f8630020
gates / gates (push) Failing after 16s
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the difference is stated rather than blurred. FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655, ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z here -- R-346's trap). Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0, CLOSE-WAIT 0. The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade, and reading it the other way would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon. Row closed as moot. What it DOES establish is worth more than the original question: twelve days after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits at baseline with zero established connections. R-336's ~323-day runway concern retires with it. BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and the newest golden bake carries 0.223.0, so a machine installed right now gets neither. The gate is RIGHT. This push therefore uses `git push --no-verify`, declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242. A BYPASS, not a waiver: the gate offers a waiver only for a release that DELIBERATELY needs no golden, and these need one. The operator was asked and ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box -- R-330 is a nightly false alarm about apps a new box has not installed yet, R-331 is a hub display over backups a new box has not taken yet -- and both arrive by self-update. That ground is recorded because it is what to re-check: it does NOT extend to a release changing first-boot behaviour. OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1, three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap is now two releases wide rather than one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
151 lines
8.6 KiB
Markdown
151 lines
8.6 KiB
Markdown
# REPORT — R-331: the operator Backup card said every customer had no backups
|
|
|
|
**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30**
|
|
|
|
---
|
|
|
|
## 1. What was wrong
|
|
|
|
The hub customer page's **Backup** card read, for **every customer, indefinitely**:
|
|
|
|
```
|
|
Enabled Yes Snapshots 0
|
|
Repo Size 0 MB Integrity Unknown
|
|
```
|
|
|
|
Measured on `demo-hp` 2026-08-30, at which moment the truth was:
|
|
|
|
| source | value |
|
|
|---|---|
|
|
| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` |
|
|
| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` |
|
|
| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report |
|
|
|
|
**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88
|
|
direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an
|
|
operator consults to answer "is this customer protected?".
|
|
|
|
## 2. Root cause
|
|
|
|
The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and
|
|
`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice
|
|
8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment.
|
|
The zeros were correct values for dead fields, rendered as if live.
|
|
|
|
**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this
|
|
package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker`
|
|
**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage
|
|
from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes
|
|
were arriving. This is a **render fix over an existing feed**, not a new pipeline.
|
|
|
|
## 3. Why it was not a one-line template swap
|
|
|
|
`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever
|
|
measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered
|
|
„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's
|
|
`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have
|
|
**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards
|
|
`stats_known`.
|
|
|
|
## 4. What changed
|
|
|
|
`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's
|
|
whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot
|
|
keep:
|
|
|
|
| report state | card shows |
|
|
|---|---|
|
|
| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** |
|
|
| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" |
|
|
| enabled, `stats_known:false` | **—**, plus "never been measured". Never `0` |
|
|
| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge |
|
|
|
|
**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is
|
|
the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every
|
|
un-upgraded customer has zero backups. Pinned by a test.
|
|
|
|
**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity
|
|
check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row
|
|
that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten
|
|
field would be a lie.
|
|
|
|
`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders
|
|
against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose
|
|
entire defect was under-reporting a real backup reads as "nearly nothing".
|
|
|
|
## 5. Tests and the red-proof
|
|
|
|
`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a
|
|
regression fails against the same numbers the defect was measured against. **The defect lived in the
|
|
template's choice of source object, so a test one layer below it would have been green against the
|
|
shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML.
|
|
|
|
**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests —
|
|
`the rendered Backup card does not contain demo-hp's real snapshot count (67)`,
|
|
`the card does not carry the real repository size (134.3 MB ...)`,
|
|
`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored
|
|
immediately; `git diff` clean.
|
|
|
|
**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0.
|
|
|
|
## 6. Deployment
|
|
|
|
PENDING at the time of writing — see the follow-up commit. The two halves ship independently and in
|
|
either order: the hub renders "unknown" for any box still on controller 0.224.0, which is correct
|
|
rather than wrong.
|
|
|
|
## 7. The push bypassed a gate, deliberately, and here is the declaration
|
|
|
|
**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was
|
|
CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden
|
|
bake carries **0.223.0**, so a machine installed right now receives neither fix.
|
|
|
|
**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs
|
|
no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated
|
|
ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not
|
|
installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by
|
|
self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check:
|
|
**it does not extend to a release that changes first-boot behaviour.**
|
|
|
|
**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change,
|
|
`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases
|
|
wide rather than one.
|
|
|
|
**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement,
|
|
five days overdue. It was **taken** during this session — see §8.
|
|
|
|
## 8. R-341's overdue check was taken, and its premise did not survive
|
|
|
|
Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence:
|
|
`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`.
|
|
|
|
**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still
|
|
`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation
|
|
as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`,
|
|
which reads 03:54:54Z here — R-346's trap, avoided.)
|
|
|
|
**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at
|
|
the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0,
|
|
CLOSE-WAIT 0.
|
|
|
|
**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1
|
|
upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent
|
|
0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the
|
|
other way would credit a changelog that was read in advance and found to contain no such mechanism. The
|
|
perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal
|
|
of the entire phenomenon. The question is now **moot**, and the row is closed as such.
|
|
|
|
**What it does establish, which is worth more than the original question:** twelve days after the R-344
|
|
fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero
|
|
established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway
|
|
concern retires with it.
|
|
|
|
## 9. Not done, and why
|
|
|
|
- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A
|
|
second verdict over the same data is two things that can disagree — a shape this codebase has already
|
|
paid for (`LastRun` vs `LastSuccess`, R-100).
|
|
- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would
|
|
stop historical reports already in this hub's store from parsing, for no gain — nothing renders them
|
|
now, and a controller-side test fails if anything starts producing them.
|