10c223bdfe
gates / gates (push) Successful in 19s
The drill was to delete demo-hp's off-site history and get it back out of a Storage
Box snapshot, filling row 10's blank RTO. Phase 1 found there is nothing to get it
back from: no snapshot is reachable from a sub-account BY ANY NAME.
Measured, read-only, no delete verb issued against any live store:
- 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS, nine full days at
second granularity, plus 126 alternative shapes -> ZERO hits.
- The control is what makes that mean anything: the identical 600-name batch shape
with one real path appended returned it, 6 of 6.
- Structural cause: /home (u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev
0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a
different dataset than the one holding felhom-repo.
- Three tools agree with controls in the same run: SFTP, the port-23 shell,
rsync --list-only.
So yesterday's re-scope splits: clause (a) "the box cannot write into the snapshot
area" STANDS and is re-confirmed; clause (b) "the rest is recoverable file by file"
is NOT SUPPORTED. STOPPED before Phase 2 on the operator's ruling — with no recovery
leg the deletion would have destroyed real history to buy only an alarm test that
could not fire at the specified size. Store verified untouched at 69 snapshots.
R-432 ANSWERED (negatively; its panel-read next step withdrawn as unnecessary).
R-433 no snapshot reachable by any name — decides R-95's remedy and its rank.
R-434 the drop alarm's text promises a file-by-file recovery that cannot be performed.
R-435 the drop detector is blind to a single-app deletion (>50% of 69 needed, ~9 given).
R-436 LEAD: the provider offers `rclone serve restic --stdio` and restic 0.14.0 speaks
`rclone:` (measured, controlled) — real prevention may need no new machine, IF
the vendor pins --append-only. Ask before building.
07 §8 row 10: text corrected, status NOT moved, RTO still blank.
No code, no version bump, no image, no golden. REPORT.md's only copy of the R-331
report preserved as REPORT-r331-backup-card.md before overwrite.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
188 lines
10 KiB
Markdown
188 lines
10 KiB
Markdown
# REPORT — R-331: the operator Backup card said every customer had no backups
|
|
|
|
**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30**
|
|
|
|
---
|
|
|
|
## 1. What was wrong
|
|
|
|
The hub customer page's **Backup** card read, for **every customer, indefinitely**:
|
|
|
|
```
|
|
Enabled Yes Snapshots 0
|
|
Repo Size 0 MB Integrity Unknown
|
|
```
|
|
|
|
Measured on `demo-hp` 2026-08-30, at which moment the truth was:
|
|
|
|
| source | value |
|
|
|---|---|
|
|
| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` |
|
|
| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` |
|
|
| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report |
|
|
|
|
**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88
|
|
direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an
|
|
operator consults to answer "is this customer protected?".
|
|
|
|
## 2. Root cause
|
|
|
|
The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and
|
|
`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice
|
|
8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment.
|
|
The zeros were correct values for dead fields, rendered as if live.
|
|
|
|
**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this
|
|
package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker`
|
|
**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage
|
|
from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes
|
|
were arriving. This is a **render fix over an existing feed**, not a new pipeline.
|
|
|
|
## 3. Why it was not a one-line template swap
|
|
|
|
`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever
|
|
measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered
|
|
„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's
|
|
`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have
|
|
**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards
|
|
`stats_known`.
|
|
|
|
## 4. What changed
|
|
|
|
`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's
|
|
whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot
|
|
keep:
|
|
|
|
| report state | card shows |
|
|
|---|---|
|
|
| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** |
|
|
| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" |
|
|
| enabled, `stats_known:false` | **—**, plus "never been measured". Never `0` |
|
|
| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge |
|
|
|
|
**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is
|
|
the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every
|
|
un-upgraded customer has zero backups. Pinned by a test.
|
|
|
|
**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity
|
|
check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row
|
|
that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten
|
|
field would be a lie.
|
|
|
|
`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders
|
|
against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose
|
|
entire defect was under-reporting a real backup reads as "nearly nothing".
|
|
|
|
## 5. Tests and the red-proof
|
|
|
|
`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a
|
|
regression fails against the same numbers the defect was measured against. **The defect lived in the
|
|
template's choice of source object, so a test one layer below it would have been green against the
|
|
shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML.
|
|
|
|
**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests —
|
|
`the rendered Backup card does not contain demo-hp's real snapshot count (67)`,
|
|
`the card does not carry the real repository size (134.3 MB ...)`,
|
|
`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored
|
|
immediately; `git diff` clean.
|
|
|
|
**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0.
|
|
|
|
## 6. Deployment and live verification
|
|
|
|
Hub **0.109.0** built, pushed, manifest bumped, ArgoCD hard-refreshed and synced. Controller
|
|
**0.225.0** deployed to both demo boxes. Verified from the live objects, not from a rollout message
|
|
(an ArgoCD "rolled out" can name the old image):
|
|
|
|
```
|
|
argocd: sync=Synced rev=36f86300209e9ce914f2a97dac16ee9eeea249e2 (== HEAD)
|
|
deploy: gitea.dooplex.hu/admin/felhom-hub:0.109.0
|
|
pod: hub-6795879c4b-pf4lz ...felhom-hub:0.109.0 Running
|
|
boxes: ...felhom-controller:0.225.0 Up (healthy) on demo-felhom AND demo-hp
|
|
```
|
|
|
|
**The card, fetched from the live hub** (endpoint-level: the exact URL the operator's browser
|
|
requests; the residual is client-side rendering only — there is no browser on DooPlex):
|
|
|
|
| | demo-hp | demo-felhom |
|
|
|---|---|---|
|
|
| Off-site snapshots | **67** | **10** |
|
|
| Repo size | **134.3 MB** | **132.5 KB** |
|
|
| Last successful run | 15h ago | 15h ago |
|
|
| Soft quota | 50 GB | 50 GB |
|
|
| Integrity row | **absent** (grep count 0) | **absent** (grep count 0) |
|
|
|
|
Both read `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` before this change.
|
|
|
|
**Cross-checked against the source, not just against itself** — the numbers on the card are the
|
|
numbers on the boxes:
|
|
|
|
```
|
|
demo-hp settings.json offbox: snapshot_count 67, repo_size_bytes 140829678, stats_known true
|
|
demo-felhom settings.json offbox: snapshot_count 10, repo_size_bytes 135635, stats_known true
|
|
135635 / 1024 = 132.5 KB → matches the rendered value
|
|
```
|
|
|
|
**One honest detail worth keeping:** demo-felhom's "Last DB dump" reads `—`. That is correct, not a
|
|
regression — its only app (`opengist`) has no database, so the box has never taken a DB dump.
|
|
|
|
**Not verified live: the "never measured" branch.** Both boxes report `stats_known: true`, so the
|
|
degradation path could not be exercised on real hardware without falsifying a box's state. It is
|
|
covered by `TestBackupCard_ThreeWayRuling` and `TestBackupCard_OldControllerDegradesToUnknownNotEmpty`
|
|
at render level, and this is stated rather than implied.
|
|
|
|
## 7. The push bypassed a gate, deliberately, and here is the declaration
|
|
|
|
**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was
|
|
CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden
|
|
bake carries **0.223.0**, so a machine installed right now receives neither fix.
|
|
|
|
**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs
|
|
no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated
|
|
ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not
|
|
installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by
|
|
self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check:
|
|
**it does not extend to a release that changes first-boot behaviour.**
|
|
|
|
**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change,
|
|
`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases
|
|
wide rather than one.
|
|
|
|
**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement,
|
|
five days overdue. It was **taken** during this session — see §8.
|
|
|
|
## 8. R-341's overdue check was taken, and its premise did not survive
|
|
|
|
Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence:
|
|
`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`.
|
|
|
|
**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still
|
|
`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation
|
|
as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`,
|
|
which reads 03:54:54Z here — R-346's trap, avoided.)
|
|
|
|
**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at
|
|
the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0,
|
|
CLOSE-WAIT 0.
|
|
|
|
**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1
|
|
upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent
|
|
0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the
|
|
other way would credit a changelog that was read in advance and found to contain no such mechanism. The
|
|
perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal
|
|
of the entire phenomenon. The question is now **moot**, and the row is closed as such.
|
|
|
|
**What it does establish, which is worth more than the original question:** twelve days after the R-344
|
|
fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero
|
|
established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway
|
|
concern retires with it.
|
|
|
|
## 9. Not done, and why
|
|
|
|
- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A
|
|
second verdict over the same data is two things that can disagree — a shape this codebase has already
|
|
paid for (`LastRun` vs `LastSuccess`, R-100).
|
|
- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would
|
|
stop historical reports already in this hub's store from parsing, for no gain — nothing renders them
|
|
now, and a controller-side test fails if anything starts producing them.
|