Files
felhom.eu/documentation/backlog/OPEN-ITEMS.md
T
admin 65409aecd1 docs: R-88 Part 1 shipped; Phase 0 root cause; R-97 minted
R-88 split: Part 1 (the failure breaker) SHIPPED in controller v0.176.0 and live
on both boxes; Part 2 (unknown != never) stays OPEN and is agent-side.

Phase 0 established the root cause at source: newestArchiveOn's (time.Time, bool)
signature cannot represent 'unknown', so a storage read ERROR collapses into a
positive 'no successful backup recorded yet'. The errored and genuine-never paths
are byte-identical on the wire, which is why Part 2 cannot be done controller-side.

R-97: the whole-guest backup tier has no failure signal to the hub at all —
internal/quiesce never imports internal/notify, so three failed backups and three
app-stack outages produced zero backup_failed events. Its only trace was a
customer-tier Hungarian app_start_failed for an app the backup itself had stopped.
2026-07-27 16:27:31 +02:00

40 lines
4.5 KiB
Markdown

# OPEN-ITEMS — the single source of truth for open work
**Rebuilt 2026-07-27 by read-only triage.** `ROADMAP.md` keeps the full history and reasoning; this
page keeps only what is **open**, and it is the file to read first. `REPORT.md` is per-session and
**overwritten** — nothing durable may live only there.
State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row has an owner.
| ID | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|
| **R-88a** | ~~Failing backup re-quiesces every 5 min, no backoff~~ | **SHIPPED** (controller v0.176.0, 2026-07-27) | — | Live on both boxes; breaker 15m→4h, per-tier, never permanent | — |
| **R-88b** | `/backup/due` cannot say *unknown* — "read errored" and "never backed up" are byte-identical, so nil still bypasses the window gate | **READY #1** | — | Agent wire change: give unknown its own representation; compat rule both ways + MinAgent floor | CC |
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY #2** | — | Snapshot plan as the stopgap (row below), then split prune off-box or move to REST `--append-only` | CC |
| **R-94** | Hub hands out host-install `1.19.0`; `1.20.0` is what carries R-82's backup default | **READY #3** | — | Bump `configs.go:28`, and stop hand-syncing a version constant across repos | CC |
| **R-86** | Restore-tests are interval-scheduled, not backup-aligned | **READY #4** | R-90 (ep0 headroom) informs cadence | Trigger a tier ~24 h after **its own** newest archive | CC |
| **R-87** | The restic tier is never restore-tested | **READY #5** | — | Design a controller-side test (no scratch-guest analogue transfers) | CC |
| — | Enable Hetzner Storage Box **snapshots** on `storage-box-pool-1``snapshot_plan=null`, 0/10 used, server-side so SFTP cannot delete them | WAITING-ON-OPERATOR | operator ruling | One console/API call; immediate immutability for the restic tier | operator |
| — | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | WAITING-ON-OPERATOR | operator console | Delete the box | operator |
| **R-90** | ep0: 3.8 GB, **no swap**, OOM'd 2026-07-27 killing PBS for ~15 min | **BLOCKED** | Hetzner CX33 availability | Rescale; or add a swapfile as an interim (needs no console) | operator |
| **R-91** | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC |
| — | First-ever **GC** on `felhom-offsite` (armed today 13:11 UTC, never run) | WATCHING | schedule | **Sun 2026-08-02 04:30 UTC** — confirm it completes | CC |
| — | demo-felhom's next weekly PBS backup (newest is 2026-07-26) | WATCHING | schedule | ~2026-08-02; also releases R-91 | CC |
| — | demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC) | WATCHING | schedule | ~2026-07-30 18:38 UTC | CC |
| **R-89** | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC |
| **R-96** | Two standing rules agreed in chat, never committed | READY (XS) | — | Add both beside `CONTEXT.md` S-1/S-2 | CC |
| **R-92** | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC |
| **R-93** | `drill-r50` is both a blocked customer and the only drift fixture | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
## Why the READY rows rank this way
1. **R-88b** — the loop itself is fixed (R-88a, v0.176.0), so the active harm is gone; what remains is
that an unknown still fires the safety valve, so a read failure can take **one** out-of-window
quiesce. Bounded now, but it is the fourth appearance of this class and the only one still open.
2. **R-95** — the largest *data* exposure: the tier holding the customer's documents and photos is the
one whose credential can delete, and the mitigation is a console click nobody has made.
3. **R-94** — a one-line constant, but until it moves every hub-driven install gets the pre-R-82
backup default. Cheapest high-consequence fix on the list.
4. **R-86** — an operator ruling already exists; it only waits on knowing what load ep0 can take.
5. **R-87** — real and unbuilt, but needs its own design, so it should not jump work that is specified.