Files
felhom.eu/REPORT.md
T
admin ac6ac037bc
gates / gates (push) Failing after 17s
R-331 live: the Backup card now reads 67 snapshots, not 0
Hub 0.109.0 + controller 0.225.0 deployed and verified from the live objects,
not from a rollout message (an ArgoCD "rolled out" can name the old image):
argocd sync=Synced rev==HEAD, deploy and pod both on felhom-hub:0.109.0, both
boxes on felhom-controller:0.225.0 (healthy).

Fetched from the live hub at the exact URL the operator's browser requests:
  demo-hp      67 snapshots / 134.3 MB / last success 15h ago / 50 GB quota
  demo-felhom  10 snapshots / 132.5 KB / last success 15h ago / 50 GB quota
Both read "Snapshots 0 / Repo Size 0 MB / Integrity Unknown" before this change.
Integrity row grep count is 0 on both pages.

Cross-checked against the SOURCE rather than against the card itself: the boxes'
own settings.json hold snapshot_count 67 / 10 and repo_size_bytes 140829678 /
135635, and 135635/1024 = 132.5 KB, matching the rendered value.

Stated rather than implied: the "never measured" branch was NOT verified live.
Both boxes report stats_known:true, so exercising it would have meant falsifying
a box's state. It is covered at render level by TestBackupCard_ThreeWayRuling and
TestBackupCard_OldControllerDegradesToUnknownNotEmpty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:53:10 +02:00

188 lines
10 KiB
Markdown

# REPORT — R-331: the operator Backup card said every customer had no backups
**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30**
---
## 1. What was wrong
The hub customer page's **Backup** card read, for **every customer, indefinitely**:
```
Enabled Yes Snapshots 0
Repo Size 0 MB Integrity Unknown
```
Measured on `demo-hp` 2026-08-30, at which moment the truth was:
| source | value |
|---|---|
| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` |
| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` |
| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report |
**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88
direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an
operator consults to answer "is this customer protected?".
## 2. Root cause
The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and
`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice
8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment.
The zeros were correct values for dead fields, rendered as if live.
**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this
package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker`
**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage
from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes
were arriving. This is a **render fix over an existing feed**, not a new pipeline.
## 3. Why it was not a one-line template swap
`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever
measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered
„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's
`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have
**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards
`stats_known`.
## 4. What changed
`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's
whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot
keep:
| report state | card shows |
|---|---|
| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** |
| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" |
| enabled, `stats_known:false` | **&mdash;**, plus "never been measured". Never `0` |
| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge |
**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is
the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every
un-upgraded customer has zero backups. Pinned by a test.
**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity
check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row
that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten
field would be a lie.
`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders
against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose
entire defect was under-reporting a real backup reads as "nearly nothing".
## 5. Tests and the red-proof
`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a
regression fails against the same numbers the defect was measured against. **The defect lived in the
template's choice of source object, so a test one layer below it would have been green against the
shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML.
**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests —
`the rendered Backup card does not contain demo-hp's real snapshot count (67)`,
`the card does not carry the real repository size (134.3 MB ...)`,
`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored
immediately; `git diff` clean.
**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0.
## 6. Deployment and live verification
Hub **0.109.0** built, pushed, manifest bumped, ArgoCD hard-refreshed and synced. Controller
**0.225.0** deployed to both demo boxes. Verified from the live objects, not from a rollout message
(an ArgoCD "rolled out" can name the old image):
```
argocd: sync=Synced rev=36f86300209e9ce914f2a97dac16ee9eeea249e2 (== HEAD)
deploy: gitea.dooplex.hu/admin/felhom-hub:0.109.0
pod: hub-6795879c4b-pf4lz ...felhom-hub:0.109.0 Running
boxes: ...felhom-controller:0.225.0 Up (healthy) on demo-felhom AND demo-hp
```
**The card, fetched from the live hub** (endpoint-level: the exact URL the operator's browser
requests; the residual is client-side rendering only — there is no browser on DooPlex):
| | demo-hp | demo-felhom |
|---|---|---|
| Off-site snapshots | **67** | **10** |
| Repo size | **134.3 MB** | **132.5 KB** |
| Last successful run | 15h ago | 15h ago |
| Soft quota | 50 GB | 50 GB |
| Integrity row | **absent** (grep count 0) | **absent** (grep count 0) |
Both read `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` before this change.
**Cross-checked against the source, not just against itself** — the numbers on the card are the
numbers on the boxes:
```
demo-hp settings.json offbox: snapshot_count 67, repo_size_bytes 140829678, stats_known true
demo-felhom settings.json offbox: snapshot_count 10, repo_size_bytes 135635, stats_known true
135635 / 1024 = 132.5 KB → matches the rendered value
```
**One honest detail worth keeping:** demo-felhom's "Last DB dump" reads `—`. That is correct, not a
regression — its only app (`opengist`) has no database, so the box has never taken a DB dump.
**Not verified live: the "never measured" branch.** Both boxes report `stats_known: true`, so the
degradation path could not be exercised on real hardware without falsifying a box's state. It is
covered by `TestBackupCard_ThreeWayRuling` and `TestBackupCard_OldControllerDegradesToUnknownNotEmpty`
at render level, and this is stated rather than implied.
## 7. The push bypassed a gate, deliberately, and here is the declaration
**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was
CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden
bake carries **0.223.0**, so a machine installed right now receives neither fix.
**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs
no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated
ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not
installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by
self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check:
**it does not extend to a release that changes first-boot behaviour.**
**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change,
`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases
wide rather than one.
**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement,
five days overdue. It was **taken** during this session — see §8.
## 8. R-341's overdue check was taken, and its premise did not survive
Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence:
`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`.
**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still
`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation
as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`,
which reads 03:54:54Z here — R-346's trap, avoided.)
**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at
the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0,
CLOSE-WAIT 0.
**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1
upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent
0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the
other way would credit a changelog that was read in advance and found to contain no such mechanism. The
perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal
of the entire phenomenon. The question is now **moot**, and the row is closed as such.
**What it does establish, which is worth more than the original question:** twelve days after the R-344
fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero
established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway
concern retires with it.
## 9. Not done, and why
- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A
second verdict over the same data is two things that can disagree — a shape this codebase has already
paid for (`LastRun` vs `LastSuccess`, R-100).
- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would
stop historical reports already in this hub's store from parsing, for no gain — nothing renders them
now, and a controller-side test fails if anything starts producing them.