Files
felhom.eu/REPORT.md
T
admin 36f8630020
gates / gates (push) Failing after 16s
R-341 check taken, and the golden-currency bypass declared
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the
difference is stated rather than blurred.

FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on
ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655,
ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy
generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z
here -- R-346's trap).

Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented
baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0,
CLOSE-WAIT 0.

The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the
PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the
leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0
measures our fix, not the upgrade, and reading it the other way would credit a
changelog that was read in advance and found to contain no such mechanism. The
perturbation pre-registered for this window was Phase C at ~3%; the actual
perturbation was the removal of the entire phenomenon. Row closed as moot.

What it DOES establish is worth more than the original question: twelve days
after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits
at baseline with zero established connections. R-336's ~323-day runway concern
retires with it.

BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and
the newest golden bake carries 0.223.0, so a machine installed right now gets
neither. The gate is RIGHT. This push therefore uses `git push --no-verify`,
declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242.

A BYPASS, not a waiver: the gate offers a waiver only for a release that
DELIBERATELY needs no golden, and these need one. The operator was asked and
ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box --
R-330 is a nightly false alarm about apps a new box has not installed yet, R-331
is a hub display over backups a new box has not taken yet -- and both arrive by
self-update. That ground is recorded because it is what to re-check: it does NOT
extend to a release changing first-boot behaviour.

OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1,
three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap
is now two releases wide rather than one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:49:53 +02:00

151 lines
8.6 KiB
Markdown

# REPORT — R-331: the operator Backup card said every customer had no backups
**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30**
---
## 1. What was wrong
The hub customer page's **Backup** card read, for **every customer, indefinitely**:
```
Enabled Yes Snapshots 0
Repo Size 0 MB Integrity Unknown
```
Measured on `demo-hp` 2026-08-30, at which moment the truth was:
| source | value |
|---|---|
| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` |
| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` |
| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report |
**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88
direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an
operator consults to answer "is this customer protected?".
## 2. Root cause
The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and
`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice
8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment.
The zeros were correct values for dead fields, rendered as if live.
**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this
package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker`
**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage
from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes
were arriving. This is a **render fix over an existing feed**, not a new pipeline.
## 3. Why it was not a one-line template swap
`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever
measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered
„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's
`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have
**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards
`stats_known`.
## 4. What changed
`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's
whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot
keep:
| report state | card shows |
|---|---|
| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** |
| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" |
| enabled, `stats_known:false` | **&mdash;**, plus "never been measured". Never `0` |
| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge |
**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is
the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every
un-upgraded customer has zero backups. Pinned by a test.
**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity
check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row
that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten
field would be a lie.
`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders
against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose
entire defect was under-reporting a real backup reads as "nearly nothing".
## 5. Tests and the red-proof
`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a
regression fails against the same numbers the defect was measured against. **The defect lived in the
template's choice of source object, so a test one layer below it would have been green against the
shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML.
**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests —
`the rendered Backup card does not contain demo-hp's real snapshot count (67)`,
`the card does not carry the real repository size (134.3 MB ...)`,
`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored
immediately; `git diff` clean.
**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0.
## 6. Deployment
PENDING at the time of writing — see the follow-up commit. The two halves ship independently and in
either order: the hub renders "unknown" for any box still on controller 0.224.0, which is correct
rather than wrong.
## 7. The push bypassed a gate, deliberately, and here is the declaration
**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was
CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden
bake carries **0.223.0**, so a machine installed right now receives neither fix.
**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs
no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated
ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not
installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by
self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check:
**it does not extend to a release that changes first-boot behaviour.**
**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change,
`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases
wide rather than one.
**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement,
five days overdue. It was **taken** during this session — see §8.
## 8. R-341's overdue check was taken, and its premise did not survive
Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence:
`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`.
**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still
`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation
as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`,
which reads 03:54:54Z here — R-346's trap, avoided.)
**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at
the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0,
CLOSE-WAIT 0.
**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1
upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent
0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the
other way would credit a changelog that was read in advance and found to contain no such mechanism. The
perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal
of the entire phenomenon. The question is now **moot**, and the row is closed as such.
**What it does establish, which is worth more than the original question:** twelve days after the R-344
fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero
established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway
concern retires with it.
## 9. Not done, and why
- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A
second verdict over the same data is two things that can disagree — a shape this codebase has already
paid for (`LastRun` vs `LastSuccess`, R-100).
- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would
stop historical reports already in this hub's store from parsing, for no gain — nothing renders them
now, and a controller-side test fails if anything starts producing them.