Files
felhom.eu/REPORT.md
T
admin 36f8630020
gates / gates (push) Failing after 16s
R-341 check taken, and the golden-currency bypass declared
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the
difference is stated rather than blurred.

FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on
ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655,
ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy
generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z
here -- R-346's trap).

Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented
baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0,
CLOSE-WAIT 0.

The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the
PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the
leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0
measures our fix, not the upgrade, and reading it the other way would credit a
changelog that was read in advance and found to contain no such mechanism. The
perturbation pre-registered for this window was Phase C at ~3%; the actual
perturbation was the removal of the entire phenomenon. Row closed as moot.

What it DOES establish is worth more than the original question: twelve days
after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits
at baseline with zero established connections. R-336's ~323-day runway concern
retires with it.

BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and
the newest golden bake carries 0.223.0, so a machine installed right now gets
neither. The gate is RIGHT. This push therefore uses `git push --no-verify`,
declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242.

A BYPASS, not a waiver: the gate offers a waiver only for a release that
DELIBERATELY needs no golden, and these need one. The operator was asked and
ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box --
R-330 is a nightly false alarm about apps a new box has not installed yet, R-331
is a hub display over backups a new box has not taken yet -- and both arrive by
self-update. That ground is recorded because it is what to re-check: it does NOT
extend to a release changing first-boot behaviour.

OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1,
three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap
is now two releases wide rather than one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:49:53 +02:00

8.6 KiB

REPORT — R-331: the operator Backup card said every customer had no backups

Hub v0.109.0 (with controller v0.225.0) · 2026-08-30


1. What was wrong

The hub customer page's Backup card read, for every customer, indefinitely:

Enabled  Yes        Snapshots  0
Repo Size  0 MB     Integrity  Unknown

Measured on demo-hp 2026-08-30, at which moment the truth was:

source value
the box's own settings.json snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true
that night's controller log [offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s
this hub's own Offsite page 0.1 GB used of a 50 GB quota — read from the same stored report

A card that reads "no backups" over a working backup is worse than no card. It is the R-88 direction of failure — degrading to no backup rather than to unknown — on the one screen an operator consults to answer "is this customer protected?".

2. Root cause

The card rendered the report's backup object. Its snapshot_count, repo_size_mb and integrity_ok fields have had no producer since disk-tier restic moved to the host agent (slice 8C) — the controller's buildBackupReport leaves them zero deliberately and says so in a comment. The zeros were correct values for dead fields, rendered as if live.

The data was never missing. The live numbers ride in the report's offsite object, which this package already reads for the Offsite page (offsiteUsageBytes) and which monitor.OffsiteChecker already drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage from the same stored report, at the same moment the Backup card said 0 MB, is the proof the bytes were arriving. This is a render fix over an existing feed, not a new pipeline.

3. Why it was not a one-line template swap

snapshot_count: 0 means two opposite things — this repository holds nothing and nobody has ever measured this repository. R-225 measured that confusion one layer down: a rebuilt box rendered „Tarolo meret · 0 pillanatkep" over a store that really held snapshot f3d9cd67, and the controller's StatsKnown fixed it there. It was never on the wire, so rendering the count without it would have moved R-225 up to the hub instead of fixing anything. Controller v0.225.0 now forwards stats_known.

4. What changed

hub/internal/web/backup_card.go builds a typed backupCardView — resolved in Go, because the card's whole subject is a distinction a template {{if}} chain over map[string]interface{} float64s cannot keep:

report state card shows
no offsite object at all "No off-site data reported" — and says explicitly this is not the same as "no backups"
enabled:false + declared state the blocker by name (needs_credential) — a different operator action from "not enabled"
enabled, stats_known:false —, plus "never been measured". Never 0
enabled, stats_known:true the real count and size, including a real 0 — measured empty is knowledge

A pre-v0.225.0 controller sends no stats_known, which unmarshals to false → "unknown". That is the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every un-upgraded customer has zero backups. Pinned by a test.

The Integrity row is deleted, not re-sourced. Nothing produces it: the controller runs no integrity check, and NotifyIntegrityOK / NotifyIntegrityFailed exist and are called from nowhere. A row that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten field would be a lie.

fmtBytesAuto is new rather than reusing fmtBytesGB: that one is fixed at GB because it renders against GB quotas, and it turns demo-hp's real 140 829 678 bytes into 0.1 GB — which on a card whose entire defect was under-reporting a real backup reads as "nearly nothing".

5. Tests and the red-proof

r331_backup_card_test.go asserts the rendered page, using demo-hp's real reported values, so a regression fails against the same numbers the defect was measured against. The defect lived in the template's choice of source object, so a test one layer below it would have been green against the shipped bug — which is why these drive handleCustomerUnified and grep the HTML.

RED-PROOF (run 2026-08-30): restoring the pre-fix card markup fails all four tests — the rendered Backup card does not contain demo-hp's real snapshot count (67), the card does not carry the real repository size (134.3 MB ...), the card still shows an Integrity row, plus every branch of the three-way ruling. Restored immediately; git diff clean.

Green gate: go build ./... && go vet ./... && go test ./... in hub/ — 18 packages, rc 0.

6. Deployment

PENDING at the time of writing — see the follow-up commit. The two halves ship independently and in either order: the hub renders "unknown" for any box still on controller 0.224.0, which is correct rather than wrong.

7. The push bypassed a gate, deliberately, and here is the declaration

git push --no-verify was used for this change. repo_gates.py's golden-currency gate was CONVICTED and it was RIGHT: controller v0.224.0 and v0.225.0 are released and the newest golden bake carries 0.223.0, so a machine installed right now receives neither fix.

This is a BYPASS, not a waiver. The gate offers a waiver only for a release that deliberately needs no golden; these need one. The operator was asked and ruled bypass-now-bake-later, on the stated ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check: it does not extend to a release that changes first-boot behaviour.

A golden carrying 0.225.0 is OWED (RUNBOOK-manual-build.md §4.1; the vouch is a three-field change, MinAgent 0.129.0). This is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.

The other failing gate was fixed, not bypassed. due-checks was red on R-341's +7 d measurement, five days overdue. It was taken during this session — see §8.

8. R-341's overdue check was taken, and its premise did not survive

Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence: documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt.

Precondition passed, which is what makes the reading interpretable: proxmox-backup-proxy still MainPID 551655, ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0 — the same proxy generation as t0, so nothing restarted and re-based the count. (The anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z here — R-346's trap, avoided.)

Result: fd = 17. Not 17 more — seventeen total, exactly the documented baseline, against 405 at the first check on 2026-08-20. The socket histogram holds one LISTEN and nothing else: ESTAB 0, CLOSE-WAIT 0.

The verdict is "unanswerable", not "the upgrade fixed it". R-341 asks whether the PBS 4.2.5-1 upgrade changed the fd slope. Inside this interval we removed the leak ourselves (R-344, agent 0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the other way would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon. The question is now moot, and the row is closed as such.

What it does establish, which is worth more than the original question: twelve days after the R-344 fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway concern retires with it.

9. Not done, and why

  • No staleness verdict on the card. monitor.OffsiteChecker already owns that and alarms on it. A second verdict over the same data is two things that can disagree — a shape this codebase has already paid for (LastRun vs LastSuccess, R-100).
  • The dead backup fields were not removed from the controller's wire format. Removing them would stop historical reports already in this hub's store from parsing, for no gain — nothing renders them now, and a controller-side test fails if anything starts producing them.