3401fcdc1c
Seam sweep: TieredBackend was the FIRST, not the only one. BackupArchiveLister has the identical silent-degrade shape and a worse blast radius (it degrades to the pre-R-84 in-memory-only behaviour), and no compile-time witness existed in production code anywhere in either repo. No defect found, so no version bump and no deploy — the witnesses are guards, proven by breaking a signature and watching go build fail where it previously passed. Live outage: age_state=unknown captured on real hardware for the first time, with demo-felhom's local tier genuinely due throughout — the controller deferred and zero app stacks were stopped. The R-88 breaker did NOT arm and no whole_guest_backup_failed travelled, because felhom-pbs was not due; recorded as conditions-did-not-arise rather than claimed as coverage. Post-boot: the volume changed device name (sdb->sda) across the reboot and the mount survived only because fstab uses by-id. That was never tested before.
4.6 KiB
4.6 KiB
OPEN-ITEMS — the single source of truth for open work
Rebuilt 2026-07-27 by read-only triage. ROADMAP.md keeps the full history and reasoning; this
page keeps only what is open, and it is the file to read first. REPORT.md is per-session and
overwritten — nothing durable may live only there.
State: BLOCKED · READY · WAITING-ON-OPERATOR · WATCHING. Every row has an owner.
| ID | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|
| R-88a | SHIPPED (controller v0.176.0, 2026-07-27) | — | Live on both boxes; breaker 15m→4h, per-tier, never permanent | — | |
| R-88b | /backup/due cannot say unknown |
SHIPPED + PROVEN-LIVE (agent v0.105.0 + controller v0.178.0, 2026-07-27) | — | age_state=unknown captured on real hardware during a deliberate ep0 outage; controller deferred, zero app stacks stopped |
— |
| R-95 | restic offsite credential can delete (readonly=False, forget --prune runs from the box); SFTP cannot express append-only |
READY #1 | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST --append-only |
CC |
| R-94 | Hub hands out host-install 1.19.0; 1.20.0 is what carries R-82's backup default |
READY #2 | — | Bump configs.go:28, and stop hand-syncing a version constant across repos |
CC |
| R-86 | Restore-tests are interval-scheduled, not backup-aligned | READY #3 | R-90 (ep0 headroom) informs cadence | Trigger a tier ~24 h after its own newest archive | CC |
| R-87 | The restic tier is never restore-tested | READY #4 | — | Design a controller-side test (no scratch-guest analogue transfers) | CC |
| — | Storage Box snapshots on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but 0 taken yet |
WATCHING | first run tonight 00:00 | Confirm size_snapshots > 0 tomorrow; until then the mitigation is armed, not proven |
CC |
| — | PBS-storage-1 (u629193, box 611421) still status=active, 19.9 MB |
WAITING-ON-OPERATOR | operator console | Delete the box | operator |
| R-90 | ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged | BLOCKED (interim proven) | Hetzner CX33 availability — confirmed unavailable even powered OFF, so it is the Cost-Optimized "Limited availability", not the power state | Re-check CX33; escape hatch if urgent = CPX/CCX lines (no availability warning, higher cost) | operator |
| R-91 | Old 13 GB datastore copy at /srv/pbs-felhom on ep0's root disk |
WATCHING | demo-felhom's first post-migration PBS backup | Delete once it lands; fix CONTEXT.md:1018 same commit |
CC |
| — | First-ever GC on felhom-offsite (armed today 13:11 UTC, never run) |
WATCHING | schedule | Sun 2026-08-02 04:30 UTC — confirm it completes | CC |
| — | demo-felhom's next weekly PBS backup (newest is 2026-07-26) | WATCHING | schedule | ~2026-08-02; also releases R-91 | CC |
| — | demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC) | WATCHING | schedule | ~2026-07-30 18:38 UTC | CC |
| R-97 | SHIPPED (controller v0.177.0 + hub v0.78.0/v0.79.0, 2026-07-27) | — | v0.79.0 (R-97c) replaced a FALSE operator-only comment with a real operatorOnlyEvents register |
— | |
| R-89 | Retention as a per-customer commercial policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC |
| R-92 | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC |
| R-93 | drill-r50 is both a blocked customer and the only drift fixture |
READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
Why the READY rows rank this way
- R-95 — the largest data exposure: the tier holding the customer's documents and photos is the
one whose credential can delete. The snapshot mitigation is now armed (daily 00:00, keep 7),
but it has taken zero snapshots so far and it does not touch the root cause — the box can still
forget --pruneits own repo. - R-94 — a one-line constant, but until it moves every hub-driven install gets the pre-R-82 backup default. Cheapest high-consequence fix on the list.
- R-86 — an operator ruling already exists; it only waits on knowing what load ep0 can take.
- R-87 — real and unbuilt, but needs its own design, so it should not jump work that is specified.