Files
felhom.eu/documentation/backlog/OPEN-ITEMS.md
T

10 KiB

OPEN-ITEMS — the single source of truth for open work

Rebuilt 2026-07-27 by read-only triage. ROADMAP.md keeps the full history and reasoning; this page keeps only what is open, and it is the file to read first. REPORT.md is per-session and overwritten — nothing durable may live only there.

State: BLOCKED · READY · WAITING-ON-OPERATOR · WATCHING. Every row has an owner.

ID What State Blocked on Next action Owner
R-88a Failing backup re-quiesces every 5 min, no backoff SHIPPED (controller v0.176.0, 2026-07-27) Live on both boxes; breaker 15m→4h, per-tier, never permanent
R-88b /backup/due cannot say unknown SHIPPED + PROVEN-LIVE (agent v0.105.0 + controller v0.178.0, 2026-07-27) age_state=unknown captured on real hardware during a deliberate ep0 outage; controller deferred, zero app stacks stopped
R-95 restic offsite credential can delete (readonly=False, forget --prune runs from the box); SFTP cannot express append-only READY #1 Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST --append-only CC
R-94 Hub hands out host-install 1.19.0; 1.20.0 is what carries R-82's backup default READY #2 Bump configs.go:28, and stop hand-syncing a version constant across repos CC
R-86 Restore-tests are interval-scheduled, not backup-aligned READY #3 R-90 (ep0 headroom) informs cadence Trigger a tier ~24 h after its own newest archive CC
R-87 The restic tier is never restore-tested READY #4 Design a controller-side test (no scratch-guest analogue transfers) CC
Storage Box snapshots on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but 0 taken yet WATCHING first run tonight 00:00 Confirm size_snapshots > 0 tomorrow; until then the mitigation is armed, not proven CC
PBS-storage-1 (u629193, box 611421) still status=active, 19.9 MB WAITING-ON-OPERATOR operator console Delete the box operator
R-90 ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged BLOCKED (interim proven) Hetzner CX33 availability — confirmed unavailable even powered OFF, so it is the Cost-Optimized "Limited availability", not the power state Re-check CX33; escape hatch if urgent = CPX/CCX lines (no availability warning, higher cost) operator
R-91 Old 13 GB datastore copy at /srv/pbs-felhom on ep0's root disk WATCHING demo-felhom's first post-migration PBS backup Delete once it lands; fix CONTEXT.md:1018 same commit CC
First-ever GC on felhom-offsite (armed today 13:11 UTC, never run) WATCHING schedule Sun 2026-08-02 04:30 UTC — confirm it completes CC
demo-felhom's next weekly PBS backup (newest is 2026-07-26) WATCHING schedule ~2026-08-02; also releases R-91 CC
demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC) WATCHING schedule ~2026-07-30 18:38 UTC CC
R-97 Whole-guest backup tier had no hub signal; quiesce blamed the apps SHIPPED (controller v0.177.0 + hub v0.78.0/v0.79.0, 2026-07-27) v0.79.0 (R-97c) replaced a FALSE operator-only comment with a real operatorOnlyEvents register
F-CRIT-2 A failed offsite backup left a phantom snapshot (1 B, manifest-less, NEWEST) that RESET the tier's freshness clock — 7 days silent on the real 168h cadence, invisible to both the R-88 breaker and the hub deadline monitor SHIPPED + PROVEN-LIVE (agent v0.106.0, 2026-07-28) NewestArchiveTime now counts only plausibly-complete entries (measured 1 MiB floor; undecidable ⇒ not counted). Campaign fault 2 replayed on demo-hp: phantom rejected + logged once, tier correctly DUE and backed up, and no thrash on the inverse
R-99 Server-side prune never removes a phantom snapshot. Confirmed it does NOT count them toward keep-last (dry-run kept 2 real + the phantom) so there is no retention/data-loss bug — but one accumulates per aborted upload, forever READY (S) Decide a cleanup path. Deletion on a customer datastore is a separate ruling — detection shipped, removal deliberately not automated CC
F-CRIT-1 An app that fails to restart after a quiesce never alarms on any channel — restartAll discarded the error AND StateStopped was whitelisted on invariant I1, which the quiesce path had made false SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28) Both causes fixed. Live on demo-hp: alarmed 9s after grace expiry, banner shows (stopped); a deliberate user stop stayed silent through 9 dead-app scans
F-A1 A restore-test in flight made a healthy backup report as FAILED (HTTP 409 read as a tier failure): breaker armed + operator emailed, on both boxes SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28) 409 → contention: tier stays DUE, dropped before anything stops (15m), and BLOCKED alarm if contention outlives the agent's 120m ceiling (3h). Hub DB: 409 → 0 operator emails, real failure → 1
R-100 A restic offsite tier that fails every night never goes stale on the hub — isStale counted from LastRun, which the controller writes unconditionally on failure SHIPPED + PROVEN-LIVE (controller v0.181.0 + hub v0.80.0, 2026-07-28) Anchored on a new last_success. Severity corrected during Phase 0: this was NOT a silencebackup_failed does fire nightly and reaches the operator (live DB: 5 sends). The real defect is defeated defence in depth: the hub-side pull net was anchored on a field the failing controller keeps refreshing, so it could not compensate for a lost push (cf. F-HUB). Live on demo-hp: induced failure → last_run advanced 11:25:48Z, last_success held 11:24:20Z; demo-felhom healthy → anchor advanced. Legacy degrade logged once per customer, live
R-101 Tier-2 (cross-drive) LastRun is also written on failure (recordTier2Failure), and three customer surfaces render it without a status: the Tier2DestDisconnected and Tier2DestInactive branches of backups_apps.html, and the restore-confirm dialog (Legutóbbi másolat: {{.Tier2LastRun}}) — which presents a possibly-failed run's timestamp at the moment the customer decides whether to restore READY (S) Found by R-100's P0.3/P0.4 sweep; filed not fixed (out of that task's path). No hub verdict reads it, so this is a UI-honesty issue, not an alarm one. The two degraded branches do carry a tag-warn; the restore-confirm is the sharpest instance. Fix: pair the timestamp with the status everywhere, as the main Tier-2 branch already does CC
F-REBOOT A guest rebooted during its backup does not come back — shutdown completes, start never happens, no self-heal; 9m47s total appliance outage with every alarm silent SHIPPED + PROVEN-LIVE (agent v0.107.0, 2026-07-28) 60 s guest-power watchdog; onboot is the deliberate-stop discriminator (already the stale-lock path's, and what pve-guests consults), retry bounded 3x/1m-2m-4m then escalates once. Live on demo-hp: 120 s unattended vs the incident's 587 s with a human; Scenario B proven (an onboot:0 guest left stopped)
F-LEAK A failed restore-test cannot destroy its own scratch guest (403 VM.Allocate); the 10-slot VMID band shrinks silently SHIPPED + PROVEN-LIVE (agent v0.110.0 + host-install v1.21.0, 2026-07-28) Three attempts, two refuted live. (1) Pool adoption: PUT /pools/{pool} also needs VM.Allocate on the VM — membership cannot bootstrap its own authority. (2) Per-path /vms/990000..990009 ACLs: work, but PVE's destroy calls remove_vm_access (LXC.pm:906) which deletes every ACL at /vms/<vmid>consumed by the op it authorises, one use per slot. (3) SHIPPED: 4th root-fenced exception, band enforced in sudoers literally (pct destroy 99000[0-9] --purge) + in code + at the caller; API destroy still tried first. Live: band PERMITTED, 9201/9100/9999/990010/1 REFUSED, and pct start 990000 REFUSED too
F-OBS deadapp-check leaves NO positive observable on a default (info-level) box — "no alarms" was indistinguishable from "never ran" SHIPPED + PROVEN-LIVE (controller v0.180.0 + agent v0.109.0, 2026-07-28) INFO summary every 20th scan carrying scans/evaluated/down. Agent v0.109.0 fixes the same shape in the guest-power watchdog shipped hours earlier in v0.107.0 — it logged only at startup and when it acted, so its health could be read only from absence
R-89 Retention as a per-customer commercial policy on the hub READY (increment 2) Policy object + reconciler → ep0 prune job; keep box tokens write-only CC
R-92 Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable READY (XS) Widen precision when retention becomes customer-visible CC
R-93 drill-r50 is both a blocked customer and the only drift fixture READY (XS) Retire it for a synthetic fixture, or unblock + silence per-customer CC

Why the READY rows rank this way

  1. R-95 — the largest data exposure: the tier holding the customer's documents and photos is the one whose credential can delete. The snapshot mitigation is now armed (daily 00:00, keep 7), but it has taken zero snapshots so far and it does not touch the root cause — the box can still forget --prune its own repo.
  2. R-94 — a one-line constant, but until it moves every hub-driven install gets the pre-R-82 backup default. Cheapest high-consequence fix on the list.
  3. R-86 — an operator ruling already exists; it only waits on knowing what load ep0 can take.
  4. R-87 — real and unbuilt, but needs its own design, so it should not jump work that is specified.