Files
felhom.eu/documentation/backlog/OPEN-ITEMS.md
T
admin bcbe2707d6 E-2 complete: wrapper, installer Case A/B, offer flow, degraded banner
Live: hub 0.81.0, agent 0.113.0, controller 0.185.1 on both demo boxes;
host-install 1.22.0 (script; no reinstall performed).

E-2a wrapper proven live as root on demo-hp: F-1 subdirectory refused, F-2
unmounted path refused, root device refused, idempotent re-apply is a no-op,
repointing refused -- 0 stray storages. The agent PVE role was NOT widened.

Scenario E proven live on BOTH boxes: healthy renders nothing, no message key.

Records three defects I introduced and caught: unreachable routes (mounted
outside /api/storage/, caught by the first live call), a hollow test exposed by
its own red-proof, and another gofmt-realignment no-op.

Not live-proven: the degraded banner and offer acceptance (both boxes healthy),
backup_target_absent end-to-end, Case A/B on a real install, drive-loss recovery.
2026-07-29 09:16:59 +02:00

19 KiB
Raw Blame History

OPEN-ITEMS — the single source of truth for open work

Rebuilt 2026-07-27 by read-only triage. ROADMAP.md keeps the full history and reasoning; this page keeps only what is open, and it is the file to read first. REPORT.md is per-session and overwritten — nothing durable may live only there.

State: BLOCKED · READY · WAITING-ON-OPERATOR · WATCHING. Every row has an owner.

ID What State Blocked on Next action Owner
R-88a Failing backup re-quiesces every 5 min, no backoff SHIPPED (controller v0.176.0, 2026-07-27) Live on both boxes; breaker 15m→4h, per-tier, never permanent
R-88b /backup/due cannot say unknown SHIPPED + PROVEN-LIVE (agent v0.105.0 + controller v0.178.0, 2026-07-27) age_state=unknown captured on real hardware during a deliberate ep0 outage; controller deferred, zero app stacks stopped
R-95 restic offsite credential can delete (readonly=False, forget --prune runs from the box); SFTP cannot express append-only READY #1 Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST --append-only CC
R-94 Hub hands out host-install 1.19.0; 1.20.0 is what carries R-82's backup default READY #2 Bump configs.go:28, and stop hand-syncing a version constant across repos CC
R-86 Restore-tests are interval-scheduled, not backup-aligned READY #3 R-90 (ep0 headroom) informs cadence Trigger a tier ~24 h after its own newest archive CC
R-87 The restic tier is never restore-tested READY #4 Design a controller-side test (no scratch-guest analogue transfers) CC
Storage Box snapshots on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but 0 taken yet WATCHING first run tonight 00:00 Confirm size_snapshots > 0 tomorrow; until then the mitigation is armed, not proven CC
PBS-storage-1 (u629193, box 611421) still status=active, 19.9 MB WAITING-ON-OPERATOR operator console Delete the box operator
R-90 ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged BLOCKED (interim proven) Hetzner CX33 availability — confirmed unavailable even powered OFF, so it is the Cost-Optimized "Limited availability", not the power state Re-check CX33; escape hatch if urgent = CPX/CCX lines (no availability warning, higher cost) operator
R-91 Old 13 GB datastore copy at /srv/pbs-felhom on ep0's root disk WATCHING demo-felhom's first post-migration PBS backup Delete once it lands; fix CONTEXT.md:1018 same commit CC
First-ever GC on felhom-offsite (armed today 13:11 UTC, never run) WATCHING schedule Sun 2026-08-02 04:30 UTC — confirm it completes CC
demo-felhom's next weekly PBS backup (newest is 2026-07-26) WATCHING schedule ~2026-08-02; also releases R-91 CC
demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC) WATCHING schedule ~2026-07-30 18:38 UTC CC
R-97 Whole-guest backup tier had no hub signal; quiesce blamed the apps SHIPPED (controller v0.177.0 + hub v0.78.0/v0.79.0, 2026-07-27) v0.79.0 (R-97c) replaced a FALSE operator-only comment with a real operatorOnlyEvents register
F-CRIT-2 A failed offsite backup left a phantom snapshot (1 B, manifest-less, NEWEST) that RESET the tier's freshness clock — 7 days silent on the real 168h cadence, invisible to both the R-88 breaker and the hub deadline monitor SHIPPED + PROVEN-LIVE (agent v0.106.0, 2026-07-28) NewestArchiveTime now counts only plausibly-complete entries (measured 1 MiB floor; undecidable ⇒ not counted). Campaign fault 2 replayed on demo-hp: phantom rejected + logged once, tier correctly DUE and backed up, and no thrash on the inverse
R-99 Server-side prune never removes a phantom snapshot. Confirmed it does NOT count them toward keep-last (dry-run kept 2 real + the phantom) so there is no retention/data-loss bug — but one accumulates per aborted upload, forever READY (S) Decide a cleanup path. Deletion on a customer datastore is a separate ruling — detection shipped, removal deliberately not automated CC
F-CRIT-1 An app that fails to restart after a quiesce never alarms on any channel — restartAll discarded the error AND StateStopped was whitelisted on invariant I1, which the quiesce path had made false SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28) Both causes fixed. Live on demo-hp: alarmed 9s after grace expiry, banner shows (stopped); a deliberate user stop stayed silent through 9 dead-app scans
F-A1 A restore-test in flight made a healthy backup report as FAILED (HTTP 409 read as a tier failure): breaker armed + operator emailed, on both boxes SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28) 409 → contention: tier stays DUE, dropped before anything stops (15m), and BLOCKED alarm if contention outlives the agent's 120m ceiling (3h). Hub DB: 409 → 0 operator emails, real failure → 1
R-100 A restic offsite tier that fails every night never goes stale on the hub — isStale counted from LastRun, which the controller writes unconditionally on failure SHIPPED + PROVEN-LIVE (controller v0.181.0 + hub v0.80.0, 2026-07-28) Anchored on a new last_success. Severity corrected during Phase 0: this was NOT a silencebackup_failed does fire nightly and reaches the operator (live DB: 5 sends). The real defect is defeated defence in depth: the hub-side pull net was anchored on a field the failing controller keeps refreshing, so it could not compensate for a lost push (cf. F-HUB). Live on demo-hp: induced failure → last_run advanced 11:25:48Z, last_success held 11:24:20Z; demo-felhom healthy → anchor advanced. Legacy degrade logged once per customer, live
R-101 Tier-2 LastRun is written on failure and rendered to the customer as „Legutóbbi másolat" — including in the restore confirm dialog SHIPPED + PROVEN-LIVE (controller v0.182.0, 2026-07-28) CrossDriveBackup.LastSuccess + SuccessTracked; the dialog names the last successful copy and discloses a failed newest attempt. Legacy rows migrate truthfully on first touch (an ok row adopts its time; an error row seeds nothing) — without the marker all 7 fleet rows would have flipped to „Még nincs sikeres másolat" on deploy. Part 2: the three record* sites rebuilt the whole struct; replaced by tier2Update (copy-and-overlay, safe by construction) — the naive fix would have had recordTier2Failure CLEAR the anchor. Live on demo-hp, rendered dialog read in both states
C9-F1 Tier-2 „Fájlok visszaállítása" is offered for apps whose copy has no restorable file leg; stops the app, restores 0 files, reports „Nincs hiányzó fájl — minden fájl megvan a helyén." SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28) Phase 0 sized it: 43 of 53 catalog apps read NOTHING, 9 read file legs but never their DB/volumes, 1 stateless. Honesty half shipped: Tier2RestoreCoverage refuses UP FRONT without stopping the app and NAMES the working action; a run that proceeds claims only what it examined and discloses that the database and volumes are not covered. Live on demo-felhom: bookstack refused, uptime stayed „Up About an hour" (was „Up 25 seconds"); paperless A1 re-run still byte-identical, 16/16 docs clean
C9-F2 An app in a Docker crash loop never alarms on any channel; StateRestarting is in no down-set SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28) StateRestarting deliberately NOT added to IsDownState (that alarms on every deploy fleet-wide); a SUSTAINED run becomes down after crashLoopAfter=5m, set above the 120s deploy timeout, Mealie's 60s start_period and R-97b's 180s grace. Dashboard counter uses the same predicate so it no longer contradicts the alarm. Red-proof that matters: the naive IsDownState change fails the brief-restart test
C9-F3R-104 An interrupted offsite run leaves an exclusive restic lock the self-heal cannot reach: resticStep (offbox.go:634-648) has unlock --remove-all, but ensureOffboxRepo's probe fails first, classifyResticProbe (offbox.go:77-93) has no lock case → "other" → fail-fast. Tier dead until a human unlocks; ClassifyOffsiteFailure likewise has no lock case so the operator is told „A távoli mentés ismeretlen okból nem sikerült" for a precisely-known, self-healable condition READY (MEDIUM) Add a lock case to both classifiers and let the probe path escalate to unlock --remove-all. Answers Phase C item 8: the repo is NOT usable after a killed run. Cleared manually this run; tier proven working again (ok, 1m35s). Reachable by any interruption — container restart, OOM, host reboot mid-backup CC
C9-F1bR-103 Tier-2's restore cannot cover 43 of 53 apps; the action that CAN is the keep-side unit restore (POST /backup/restoreRestoreFromRecoveryUnit, replays volume tars + DB dumps). v0.183.0 NAMES it in the refusal text but does not route to it READY Put the working action in the card the customer already opened. Deliberately its own task: it places a DESTRUCTIVE operation (overwrites live data with the backup state) behind a button reached via a NON-destructive one, so the confirm copy must carry that difference — the reason it was not folded into v0.183.0 CC
C9-F4R-102 Nothing reads the Tier-2 copy's recovery-unit/ mirror. It is written by EVERY Tier-2 run (tier2.go:369, „Unit leg (always)") and read by no code path: RecoveryUnitPath resolves to backups/**primary**/ (appbackup/paths.go:46-48), and the only reader of the secondary tree is tier2_restore.go:79, which reads hdd/+userdata/ only READY (potentially > C9-F1) Tier-2 exists for the case where the PRIMARY drive is lost — and in exactly that case the primary unit is gone while this mirror survives on the second drive, unreachable by any customer action, leaving offsite as the only route. Verified by enumeration: 6 references to "secondary" in the tree, one writer, one reader, one wipe-warning lister CC
R-108 Network storage can host an app's namespace, and FileBrowser binds a network share at its ROOT. Local drives are userdata-scoped (web/handlers.go:2450-2460); network paths are bound at the share root (:2432) and served with download: true. No IsNetwork() filter guards the deploy dropdown (settings.go:904-914), the per-app migrate targets (handlers.go:674-679), or handleStorageMigrateApp (storage_handlers.go:410-424 — its whole-namespace sibling DOES refuse, :397) READY — BLOCKS an architectural target This is why D5 was not adopted in the 2026-07-28 07-backup-architecture.md rewrite: D5 moves app secrets into the local recovery unit so Tier-1/Tier-2 restore stop needing the guest, and that is safe only if no browsing surface can reach the backup tree. Every other surface was verified clean (SMB both namespace shapes, FileBrowser for drives, .fab import + download, /api/debug/*, all three ServeFile sites, storage-path add) — 07 §10.1 has the full sweep. Not a leak today (the unit's app.yaml is secret-stripped). Verified LIVE in demo-hp's generated FileBrowser compose CC
F-DIAG Four distinct offsite failure causes collapse into two operator-visible strings SHIPPED (controller v0.182.0, 2026-07-28) ClassifyOffsiteFailure → quota / orphaned / no_repo / no_units / transport / unknown, each with its own Hungarian message. Unclassifiable says so rather than being folded into a neighbour. Secrets: the old message was a raw err.Error() passthrough carrying sftp:<user>@<host>:<path>; redaction is now by the target's actual host/user/path (a first regex-only attempt leaked on a bare hostname and its own test caught it). Unit-proven; not yet exercised by a live offsite failure of each class
F-OPS A manual pct restore inherits the source guest's bind mounts — during a real DR, on a different host, under pressure DOCUMENTED (2026-07-28) documentation/runbooks/RUNBOOK-manual-guest-restore.md: which mpN are volumes vs host binds, the mp9 source-VMID trap (it can bind another guest's bootstrap credentials), strip-and-re-add before first boot, and a positive pre-start verification. Docs only by design — the agent already neutralises binds on its own restore paths, and a second implementation would drift
F-REBOOT A guest rebooted during its backup does not come back — shutdown completes, start never happens, no self-heal; 9m47s total appliance outage with every alarm silent SHIPPED + PROVEN-LIVE (agent v0.107.0, 2026-07-28) 60 s guest-power watchdog; onboot is the deliberate-stop discriminator (already the stale-lock path's, and what pve-guests consults), retry bounded 3x/1m-2m-4m then escalates once. Live on demo-hp: 120 s unattended vs the incident's 587 s with a human; Scenario B proven (an onboot:0 guest left stopped)
F-LEAK A failed restore-test cannot destroy its own scratch guest (403 VM.Allocate); the 10-slot VMID band shrinks silently SHIPPED + PROVEN-LIVE (agent v0.110.0 + host-install v1.21.0, 2026-07-28) Three attempts, two refuted live. (1) Pool adoption: PUT /pools/{pool} also needs VM.Allocate on the VM — membership cannot bootstrap its own authority. (2) Per-path /vms/990000..990009 ACLs: work, but PVE's destroy calls remove_vm_access (LXC.pm:906) which deletes every ACL at /vms/<vmid>consumed by the op it authorises, one use per slot. (3) SHIPPED: 4th root-fenced exception, band enforced in sudoers literally (pct destroy 99000[0-9] --purge) + in code + at the caller; API destroy still tried first. Live: band PERMITTED, 9201/9100/9999/990010/1 REFUSED, and pct start 990000 REFUSED too
F-OBS deadapp-check leaves NO positive observable on a default (info-level) box — "no alarms" was indistinguishable from "never ran" SHIPPED + PROVEN-LIVE (controller v0.180.0 + agent v0.109.0, 2026-07-28) INFO summary every 20th scan carrying scans/evaluated/down. Agent v0.109.0 fixes the same shape in the guest-power watchdog shipped hours earlier in v0.107.0 — it logged only at startup and when it acted, so its health could be read only from absence
E-2 Drive-role machinery around the moved vzdump target SHIPPED (hub 0.81.0, agent 0.113.0, controller 0.185.1, host-install 1.22.0 — 2026-07-29) Parts 15 complete. Role model + offer-only assignment + Hungarian degraded banner + absent-target signal + installer Case A/B. NOT yet live-proven: the DEGRADED banner and the offer acceptance (both demo boxes are healthy, so neither state occurs naturally) and backup_target_absent end-to-end. Installer is installer-logic-tested, not install-tested — no reinstall was performed CC
E-2a The target move needs a root-fenced wrapper — the agent cannot do it SHIPPED + PROVEN-LIVE (agent v0.113.0 + host-install v1.22.0, 2026-07-29) felhom-backup-target-apply behind a literal FELHOM_BACKUPTARGET sudoers alias; the agent's PVE role was NOT widened. Enforces F-1 (mountpoint -q) and F-2 (is_mountpoint 1 hardcoded), refuses a root-device target, has NO storage-removal path (grep-assertable), is idempotent and refuses to repoint. All five laws proven live as root on demo-hp with 0 stray storages
E-2b NotifyStorageDisconnected/Reconnected defined and called NOWHERE — a drive going absent emitted no event on any channel SHIPPED + PROVEN-LIVE (controller v0.184.1 + agent v0.112.0 + hub v0.81.0, 2026-07-29) Seam wired in ReconcileDriveGates; a target drive raises the specific backup_target_absent instead. A keying bug was caught before deploy: a.Path is the registered GUEST path, not the agent's host MountPath, so the target branch was unreachable — every absent drive, target included, fell through to the generic event (v0.184.1). Tests observe the WIRE (httptest hub), not a mock
E-2c E-1 put the whole-guest backups on a drive POST /disks/eject would eject SHIPPED + PROVEN-LIVE (agent v0.112.0, 2026-07-29) Eject + decommission refuse 409 on the backup-target mount, naming the storage and the remedy. Live on BOTH boxes: demo-hp /mnt/nvme-1tb and demo-felhom /mnt/hdd_1 both refused, drives unmoved. NOT a role reclassification — RoleForStorage untouched, because on both boxes that drive is ALSO the enrolled user-data drive; TestEjectStillAllowedOnANonTargetDrive pins the non-over-correction and /var/lib/vz is still refused by the PRE-EXISTING role gate, not this one
PETI peti-felhom deliberately NOT migrated. Its whole-guest backup still shares a device with its guest, so a drive failure there is offsite-only recovery ACCEPTED RISK — parked operator's next visit (tester reinstalling from scratch) Accepted until the reinstall; re-evaluate if that slips past ~2026-09-01. Do not migrate, do not touch operator
R-109 The DR recipe records no backup target. It lists every storage's name/type/content but never which one holds the local archives — and each demo box now carries TWO content=backup dir storages, felhom-backup (live) and local (frozen 2026-07-28 archives) READY (XS) Add the resolved BackupTarget() to the host-half. Third recipe-completeness defect beside R-105/R-106 CC
R-89 Retention as a per-customer commercial policy on the hub READY (increment 2) Policy object + reconciler → ep0 prune job; keep box tokens write-only CC
R-92 Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable READY (XS) Widen precision when retention becomes customer-visible CC
R-93 drill-r50 is both a blocked customer and the only drift fixture READY (XS) Retire it for a synthetic fixture, or unblock + silence per-customer CC

Why the READY rows rank this way

  1. R-95 — the largest data exposure: the tier holding the customer's documents and photos is the one whose credential can delete. The snapshot mitigation is now armed (daily 00:00, keep 7), but it has taken zero snapshots so far and it does not touch the root cause — the box can still forget --prune its own repo.
  2. R-94 — a one-line constant, but until it moves every hub-driven install gets the pre-R-82 backup default. Cheapest high-consequence fix on the list.
  3. R-86 — an operator ruling already exists; it only waits on knowing what load ep0 can take.
  4. R-87 — real and unbuilt, but needs its own design, so it should not jump work that is specified.