3401fcdc1c
Seam sweep: TieredBackend was the FIRST, not the only one. BackupArchiveLister has the identical silent-degrade shape and a worse blast radius (it degrades to the pre-R-84 in-memory-only behaviour), and no compile-time witness existed in production code anywhere in either repo. No defect found, so no version bump and no deploy — the witnesses are guards, proven by breaking a signature and watching go build fail where it previously passed. Live outage: age_state=unknown captured on real hardware for the first time, with demo-felhom's local tier genuinely due throughout — the controller deferred and zero app stacks were stopped. The R-88 breaker did NOT arm and no whole_guest_backup_failed travelled, because felhom-pbs was not due; recorded as conditions-did-not-arise rather than claimed as coverage. Post-boot: the volume changed device name (sdb->sda) across the reboot and the mount survived only because fstab uses by-id. That was never tested before.
148 lines
8.1 KiB
Markdown
148 lines
8.1 KiB
Markdown
# REPORT — closing the backup arc: seam sweep + live outage validation (2026-07-27)
|
|
|
|
**Overwritten** per the standing rule. **No version shipped** — the sweep found no defect, so the
|
|
witnesses are compile-time-only source changes riding the next release. Agent `9842c52`→`0236553`,
|
|
controller `ca013c8`→`8f46495`. Live fleet unchanged: hub v0.79.0, agent v0.105.0, controller v0.178.0.
|
|
|
|
## Part 1 — the seam table
|
|
|
|
Every optional-capability type assertion across all four repos. `StackDataProvider`, `BackupReporter`
|
|
and `RestoreTestReporter` are **not** in this table: they are interface-typed struct *fields*, which
|
|
the compiler already enforces at assignment. The dangerous shape is `x.(Interface)`.
|
|
|
|
| Interface | Asserted at | On failure | Silent? | Witness now |
|
|
|---|---|---|---|---|
|
|
| `quiesce.TieredBackend` | `tiers.go:109` (`resolveDueTiers`) | falls back to the untargeted single-tier path | **SILENT** | ✅ moved to production |
|
|
| `quiesce.TieredBackend` | `quiesce.go:375` (`allTiersForManualRun`) | manual run covers only the untargeted tier | **SILENT** | ✅ same witness |
|
|
| `quiesce.TieredBackend` | `tiers.go:191`, `:199` | **unchecked — would panic** | no (loud) | ✅ same witness |
|
|
| `localapi.BackupArchiveLister` | `server.go:1155` (`newestArchiveOn`) | degrades to `archiveAbsent` = pre-R-84 in-memory-only | **SILENT** | ✅ added |
|
|
| `agentapi.AgentVersionReporter` | `features.go:157` | version gate → live probe | silent, benign (both decide correctly) | ✅ added |
|
|
| anonymous `GuestMemory` iface | `features.go:86` | returns `errNoMemoryProbe` → `SupportUnknown` | **no** — distinct error | n/a (anonymous) |
|
|
|
|
**Was `TieredBackend` the only one, or the first one? It was the first one.**
|
|
`BackupArchiveLister` has the identical shape and a worse blast radius: a silent degrade to
|
|
`archiveAbsent` is the R-84 bug returning — a cold store after a restart reading as "no backup ever" —
|
|
with nothing in any log. It had **no witness of any kind**, and in fact **zero compile-time witnesses
|
|
existed in production code anywhere in either repo** before today.
|
|
|
|
**No defect found.** Every implementation satisfies its interface today, so per the task's rule there
|
|
is no version bump and no deploy. What was missing was the *guard*, not the behaviour.
|
|
|
|
**The guard is proven, not assumed.** Breaking `BackupRunner.NewestArchiveTime`'s signature now fails
|
|
`go build`:
|
|
|
|
```
|
|
cmd/felhom-agent/seam_witnesses.go:28:38: *backup.BackupRunner does not implement
|
|
localapi.BackupArchiveLister (wrong type for method NewestArchiveTime)
|
|
```
|
|
|
|
Before the witness, that same break compiled and vetted clean. The controller's witness also **moved
|
|
out of a `_test.go` file** — a test-file witness fires on `go test`/`go vet` but not on `go build`
|
|
alone, which is exactly the path the R-88 Part 2 near-miss would have taken.
|
|
|
|
## Part 2 — the live outage
|
|
|
|
Pre-flight cleared: 0 tasks in flight on ep0, 0 backup/restore processes on either box (counted by
|
|
`COMM`, since `pgrep -f` self-matches and gave a false `1` first), 0 scratch guests, no quiesce marker.
|
|
|
|
### The headline — `age_state=unknown` on real hardware, first ever
|
|
|
|
```json
|
|
{"vmid":9201,"due":true,
|
|
"reason":"backup storage unreadable and no in-memory record — age UNKNOWN, treating as due",
|
|
"target":"felhom-pbs","age_state":"unknown"}
|
|
```
|
|
|
|
Identical on both boxes. The transition is timestamped on demo-felhom:
|
|
|
|
```
|
|
18:35:30 felhom-pbs/content 200 ← before power-off
|
|
18:39:26 felhom-pbs/content 500 ← after
|
|
18:39:11 local/content 200 ← local tier, same window, unaffected
|
|
```
|
|
|
|
Agent journal, both boxes: `could not read the backup storage for the due-check … target=felhom-pbs`
|
|
— the degrade path genuinely executed. **Positive observables throughout**, per standing rule 3: an
|
|
empty quiesce log alone would have been equally consistent with a dead loop.
|
|
|
|
### The behaviour it gates
|
|
|
|
demo-felhom's `local` tier was **genuinely due** (25.8 h) for the whole window, so the due set was
|
|
`[local(known), felhom-pbs(unknown)]` → `oldestAge` nil → `valveLicensed` false → **defer**. This is
|
|
the 09:02 incident's exact condition replayed with the fix in place.
|
|
|
|
| | demo-felhom | demo-hp |
|
|
|---|---|---|
|
|
| App containers running | **10**, unchanged | **3**, unchanged |
|
|
| Quiesce cycles | **0** | **0** |
|
|
| Quiesce marker | absent | absent |
|
|
| vzdump attempted | none since 07:13 | none since 09:37 |
|
|
|
|
**Not one app stack was stopped.** Per-tier isolation demonstrated on the same box in the same
|
|
minute: `local` → 200 with a real age, `felhom-pbs` → 500 → unknown.
|
|
|
|
### Item 4 — conditions did not arise, and that is confirmed, not assumed
|
|
|
|
`felhom-pbs` was not due on either box (28 h / 7 h against a 168 h cadence), so no backup ran, nothing
|
|
failed, **the R-88 breaker did not arm**, and **no `whole_guest_backup_failed` travelled**. I did not
|
|
force one. Confirmed positively: the newest hub event is `16:10:33`, *before* the window opened at
|
|
16:24, and 0 events exist in the window.
|
|
|
|
## Part 2.2 — post-boot
|
|
|
|
| Check | Result |
|
|
|---|---|
|
|
| Mount | `/mnt/pbs-datastore ext4 rw,relatime,discard` — **relatime present, `noatime` ABSENT** ✅ |
|
|
| Swap | `/swapfile 4G` back — **survived its first reboot** (it was added and verified without one) ✅ |
|
|
| PBS services | `proxmox-backup` + `proxmox-backup-proxy` both active ✅ |
|
|
| Datastore | `felhom-offsite → /mnt/pbs-datastore`, online ✅ |
|
|
| Snapshots | demo-felhom **2**, demo-hp **3** — **exact match**, same timestamps ✅ |
|
|
| Schedules | `gc-schedule sun 04:30`, `verify-new 1`, 2 prune jobs ✅ |
|
|
| Filesystem | `EXT4-fs (sda): recovery complete` — journal replayed after the unclean power-off, then mounted r/w; **0** I/O or EXT4 errors ✅ |
|
|
|
|
### The unplanned finding: the volume changed device name
|
|
|
|
```
|
|
before: volume=/dev/sdb root=/dev/sda1
|
|
after: volume=/dev/sda root=/dev/sdb1
|
|
```
|
|
|
|
The kernel enumerated the disks in the opposite order across the reboot. The mount survived only
|
|
because fstab uses `/dev/disk/by-id/scsi-0HC_Volume_106469259`. **A `/dev/sdb` entry would have failed
|
|
to mount — or mounted the root disk's device node.** That choice was load-bearing and had never been
|
|
tested; this reboot tested it.
|
|
|
|
### Recovery
|
|
|
|
Both boxes returned to `age_state=known` on both tiers within ~60 s of boot; `felhom-pbs` on
|
|
demo-felhom back to `due=false, within cadence window, age_seconds=102102`. No breaker was armed, so
|
|
none needed clearing; no failure event fired, so no recovery event was expected — and none appeared.
|
|
|
|
## Part 3 — CX33
|
|
|
|
**Not available even with the server powered off**, so the powered-on theory is settled: it is the
|
|
Cost-Optimized line's "Limited availability", not the power state. No rescale happened. R-90 stays
|
|
`BLOCKED`/`WATCHING` with the 4 GiB swap as the interim — now a *proven* interim, since it survived a
|
|
reboot. Escape hatch if 8 GB becomes urgent: the **Regular Performance (CPX)** / **General Purpose
|
|
(CCX)** lines carry no availability warning, at higher cost. Operator's call.
|
|
|
|
## Comes off the never-live-validated list
|
|
|
|
- **`age_state=unknown` observed on real hardware** — the path R-88 Part 2 exists for.
|
|
- **The window gate refusing an unknown-driven valve, with a genuinely due tier present** — zero app
|
|
stacks stopped.
|
|
- **Per-tier isolation under a real offsite outage.**
|
|
- **The by-id mount surviving a reboot** (and surviving a device rename, which was not anticipated).
|
|
- **The 4 GiB swap surviving a reboot.**
|
|
- **Datastore integrity across an unclean power-off** — journal replay, exact snapshot match.
|
|
|
|
## Honestly cannot come off it
|
|
|
|
- **The R-88 breaker has still never armed on real hardware.** No backup was due, so none failed.
|
|
- **`whole_guest_backup_failed` has never travelled from a REAL failure** — only from the induced
|
|
POST on 2026-07-27 15:04:57. The quiesce→notify link remains unit-proven.
|
|
- **R-97b's app-alarm suppression has never fired live** — it needs a real quiesce cycle.
|
|
- **`age_state=absent` has never been observed live** — it needs a genuinely fresh tier.
|
|
- **The legacy-agent path (Scenario C) is now unobservable on this fleet** — both agents are v0.105.0.
|
|
- **The first in-window scheduled backup** lands 2026-07-28 at the earliest.
|