20 KiB
ROADMAP — future features & open work
What this is: the prioritized decision log of planned/open work. Items are intentions, not claims about live behavior — the capability map (
architecture/00-capability-map.md) is the only place that states what the platform does today.Lifecycle: idea → spiked → spec'd → in-progress → shipped (item collapses to a one-liner with the version, and the corresponding capability-map row changes status with evidence). Items can also be killed (keep the one-liner + why — decisions are worth remembering).
Coupling rule: every item names the capability-map row(s) it flips. Every map gap row points back here by ID. Neither file duplicates the other's content.
Priorities: P1 = closed-alpha blocker · P2 = close during alpha · P3 = post-alpha. Existing loose notes in this folder (
FOLLOWUP-*,FIX-M*) are absorbed as references below.
P1 — closed-alpha blockers
| ID | Item | Size | Status | Notes / map rows flipped |
|---|---|---|---|---|
| R-1 | Peti convergence: clean-slate proxmox2 reinstall (spec'd 07-15), first live auto-confirm, supervised escrow ceremony, execute parked publish trains (agent 0.81→0.88, controller → 0.137) | L | spec'd | Flips: publish train PARTIAL→PROVEN-LIVE; appliance/BYO/day-0 "real customer" notes; escrow ceremony. The single biggest unproven surface — an alpha where fixes can't ship remotely is dead. Reinstall arc SHIPPED hub v0.57.0 (2026-07-16): the clean-slate reinstall-of-existing-customer path is now first-class — claim re-issue (F2), offsite re-issue (F3), escrow-honesty-on-re-issue (2.3) all auto-fire on re-enrollment. Peti's proxmox2 clean-slate now walks a supported path |
| R-2 | ~hub/internal/notify/, store.go, hub/internal/claim/) |
S | killed (2026-07-16) | Not a real issue: the "foreign WIP" was in-flight code from a concurrent CC session on the customer-claim arc, snapshotted before it committed. All of it landed cleanly — notify/+claim/engine.go in 6b40eb8 (v0.50.0), store.go in a1d0450 (v0.54.0), plus follow-up e205a2d; v0.55.0 shipped. Working tree is clean, no stashes. Lesson already codified: never run two writing sessions on one felhom.eu clone (CLAUDE.md §git add -A) |
| R-3 | Friend-alpha onboarding runbook (generalized from pilot/RUNBOOK-peti-return-2026-07-13): hardware prep → golden → install → claim → ceremony → "first restore by the customer" scripted step |
M | idea | Flips: "customer performs a restore" MISSING row; produces the tester-agreement sibling of PETI-tester-agreement.md. Next from-scratch rehearsal to include customer DELETE + re-create — the ESCROW cascade is now DEFINED (hub v0.60.1): host delete DEMOTES escrow to retained custody (never destroys), customer Danger-zone delete PURGES it (the one true purge point). S6b (manual stale-host delete before re-enroll) is OBSOLETE — re-enrollment upserts the existing host row cleanly (store.UpsertHost ON CONFLICT DO UPDATE; handleAdminCreateHost no duplicate refusal) + the v0.57.0 arc auto-fires the re-issues; the rehearsal live-confirms it. NON-escrow offboarding NOW ANSWERED by the middle-tier Customer RESET (hub v0.61.0, LIVE): one operator action deprovisions the Hetzner sub-account/box (repo data destroyed), destroys the PBS namespace + backup groups + token, clears the DR recipe / one-time secret / claim state / retained escrow custody (separate ack) — identity + basic config survive. WG peer release rides host delete (peers are host-scoped, gone before RESET runs — RESET refuses while any host row exists). Remaining consistency gap: the customer Danger-zone DELETE still leaves host rows and does NOT run the offsite/PBS teardown (RESET is the teardown path; DELETE is escrow-purge + config-drop). Decide whether DELETE should require a prior RESET (or subsume it) — new item R-25b |
| R-4 | Claim-code deliverability: test-send to gmail.com / freemail.hu / citromail.hu / t-online.hu; tighten DMARC p=none → p=quarantine (pending since email.md 02-04) |
S | idea | A claim code in spam bricks onboarding at step 1. Cheap, do before first invite |
P2 — during alpha
| ID | Item | Size | Status | Notes |
|---|---|---|---|---|
| R-5 | Hub: Storage Box box-level aggregate — total BX11 fill, Σ(customer soft quotas), oversubscription ratio, operator alert threshold | M | idea | Per-customer OffsiteChecker exists (hub v0.41). Source: Hetzner Storage Box API (verify Robot-legacy vs. new API endpoint for our box before speccing — 2025 migration); NOT ssh du over restic repos. Token → k8s Secret ref. Flips map row G/MISSING |
| R-6 | Spike: LAN service discovery from the guest — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | idea | Shared prerequisite for R-7 + R-8. Spike-first: the traps are networking, not Samba/DLNA |
| R-7 | SMB server share (gated on R-6): samba + wsdd, LAN-only binding (never tunnel), user model (single household user first), which roots are shared (dedicated shares/ vs app userdata — every SMB-writable path needs a backup class), uid-1000 convention, paperless consume flow |
L | idea | Flips map row E/MISSING. Alpha-relevant: "my box is a NAS" is a core household expectation |
| R-8 | DLNA (gated on R-6): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea | Don't add catalog weight before proving the cheap path |
| R-9 | Uninstaller trio (from 07-15 Peti session): cluster-aware felhom_guests guard (node-local pct list deletes cluster-wide pveum objects); saferemove detection + time estimate + opt-in --quick-remove (never mutate storage.cfg); smarter restore_storage default for BYO clusters (shared storage, not local-lvm) |
M | idea | Second item's rejected alternative (temp-disable-and-restore) stays rejected — crash window silently downgrades cluster wipe policy |
| R-10 | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) | XS | idea | One-line hardening; batch with the next controller task |
| R-11 | Tester-facing one-pager: what the box does, known limitations, how to report (channel decision: Messenger group?) | S | idea | Pairs with R-3 |
| R-16 | Operator hygiene: campaign6 autofs orphan (clears on host reboot) + tied-CreatedAt flash duplicates (audiobookshelf/komga/romm) | XS | open (doc-drift bit CLOSED) | Viktor's own action items from 6D/6E. Doc-drift leftover CLOSED (host-install v1.17.0, 2026-07-17): the R-20-noted stale "EMPTY by default" operator-key comment corrected (keys are PINNED). Remaining = the two operator items above |
| R-22 | PBS-DR pre-check self-grant (F4). On a non-default storage id the token-auth GET /storage/<id> pre-check 403s (no ACL yet) and used to abort before the root-run grant that creates it. |
S | SHIPPED + PROVEN-LIVE agent v0.89.0 (2026-07-17) | On a 403 the reconcile self-grants via the root wrapper + re-reads, then converges. Red-proof TestSelfGrant_PreCheck403DoesNotAbortBeforeGrant; live-reproduced on the demo (marker aside + ACLs revoked → self-grant → converged state=adopted in ~3 s, ACLs restored, offsite active). Origin tests/VALIDATION-n100-baremetal-2026-07-16.md F4. |
| R-17 | Old-box archive (u629193-sub1) retirement decision — 9/9 byte-identical restores verified | XS | awaiting-decision | Viktor ruling |
| R-19 | Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE |
| R-20 | XS | closed (2026-07-16) | Confirmed against scripts/felhom-host-install.sh source (not changelog): keys resolve at L1181–1219 (script constants OPERATOR_KEY_*, populated, --operator-pubkey-file override), pinned automatically by step_agent_config() "STEP 6/8" (L2044; python builds authz.signers L2146–2156, reinstall preserves existing), verified at L2332–2337 ("authz signers: N … operator-signed self-update armed"). No interactive prompt or post-install hand-edit — fully automatic. Doc-drift note: the L193–197 "EMPTY by default" comment is stale vs the now-populated constants (→ R-16 hygiene) |
|
| R-23 | Immediate-sync Direction-2 follow-ups (hub v0.58 / controller v0.140, shipped 2026-07-16): (a) live-validate the operator-UI save→apply round-trip end-to-end — needs an operator login, so fold into a Peti/alpha supervised session (one config save → box wakes in seconds → self-restart → startup report; also demonstrates the restart single-fire once a bump has advanced the generation past 0); (b) cosmetic: the controller Waiter's "recovered" INFO logs on the next hold completion (pollOnce blocks ~240 s), not at reconnect — surface recovery at connect time |
S | idea | Flips the new map row "config/state change round-trips in seconds" PARTIAL→PROVEN-LIVE. Transport + mechanism already proven live (SPIKE-immediate-sync-transport-2026-07-16, v0.58/v0.140 REPORTs: 240 s no-annotation hold, 0.047 s wake, restart 1-WARN/0-storm); only the login-gated UI-triggered bump leg is unexercised |
P3 — post-alpha
| ID | Item | Size | Status | Notes |
|---|---|---|---|---|
| R-26 | Guided old-history recovery via a retained superseded escrow + the recovery code. Enabled by hub v0.60.0 (Part B) which now RETAINS superseded escrow blobs (host_escrow_superseded, ListSupersededEscrow). Build the flow that, given the customer's recovery code, unwraps a retained old blob → recovers the old repo passphrase → mounts/reads the moved-aside .orphaned-<date> repo for restore. |
M | idea (enabled by v0.60.0) | Turns "history recoverable in principle" into a real customer-drivable path; pairs with the controller v0.142.0 orphaned-repo move-aside. Origin DIAGNOSE-offbox-repo-orphaned-2026-07-17 |
| R-27 | Customer-facing self-bind page (R-21 slice C follow-on). Today an unclaimed appliance is bound by the OPERATOR on the Hosts page (hub v0.62.0). Build the customer-facing flow so a customer can claim/bind their own freshly-installed box (e.g. enter a claim code / the appliance's displayed pairing id → the hub binds it to their account → delivery proceeds). Turns "operator binds every box" into true self-service onboarding. | M | idea | Origin: hub v0.62.0 slice C (operator-bind ruling; self-bind deliberately deferred). Reuses the appliance_registrations + one-shot delivery machinery; adds a customer-auth surface + a pairing-id/claim-code channel. Pairs with the claim engine |
| R-25b | Customer DELETE ↔ RESET consistency. The middle-tier Customer RESET (hub v0.61.0) runs the full external teardown (Hetzner sub-account/box + PBS namespace/groups/token) and refuses while any host row exists. The Danger-zone DELETE still (a) leaves host rows and (b) does NOT run that teardown — it purges escrow custody + drops the config only. Decide the model: DELETE requires a prior RESET, or DELETE subsumes RESET's teardown, or they stay orthogonal (RESET = recycle-in-place, DELETE = escrow-purge). | S | idea | Origin: hub v0.61.0 RESET ship. Flips a future "customer fully offboarded (external resources released)" map row. Cheap once the model is chosen |
| R-25 | Device-node TOCTOU hardening (drive init). Graduate the controller v0.141.0 Observation: the format → resolveEnrollUUID(path) → AssignDisk(uuid) sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. |
S | idea | From the v0.141.0 F6 commit's security-review finding (felhom-controller REPORT). Low real risk (single-operator, agent-guarded), but cheap to close |
| R-24 | Guest resources as hub desired-state (live resize). F5 (host-install v1.17.0) auto-sizes RAM/cores at INSTALL only. Make guest cores/RAM a per-host pbs_dr-sibling descriptor field the agent reconciles (pct set -memory/-cores), so the operator can right-size a running box from the hub — and land it in seconds via the agent-plane poke (R-13). |
M | idea (F5 follow-on) | Follows F5 (VALIDATION-n100 — appliance auto-size shipped); the live-resize path reuses the desired-state + poke machinery (agent v0.89 / hub v0.59). Would flip a new map row "operator right-sizes a running guest from the hub" |
| R-12 | Cluster mode: agent-follows-guest, bind-mount reconciliation on HA migration | XL | idea | Scoped 07-15; interim = HA-group pin to one node. Driven by Peti's two-node cluster |
| R-13 | OOB management arc: dual-use existing WireGuard + hub desired-state channel as mutual-repair | L | first slice PROVEN-LIVE (poke channel) | FIRST SLICE PROVEN-LIVE — the agent-plane poke channel (Direction-2a), agent v0.89.0 + hub v0.59.0 (2026-07-17): the ep0-relayed contentless poke (hub→ep0 felhom-poke forced-cmd→UDP→box WG /32:51822, peer-confined, zero ep0/box infra change) reaches the agent and fires an immediate desired-state cycle. Full path live: real operator manifest save → sync-poke delivered to 10.77.0.2; box → poke received → immediate desired-state cycle (~31 ms ep0→box, save→tick ≈ ~0.45 s). This is ONLY the listener+sender; the rest of the mutual-repair arc (self-heal actions over the channel) stays open. Per SPIKE-immediate-sync-transport-2026-07-16 P4. The controller-plane Direction-2 wait channel (hub v0.58 / controller v0.140) shipped the config-puller leg separately |
| R-28 | Agent fast-tick-until-first-convergence (immediacy SECONDARY). The hub v0.63.0 system-initiated pokes cannot reach a box on the ONE leg that matters most at onboarding: the WG-registration auto-provision (the box's tunnel does not exist until it fetches the WG block, so a poke is undeliverable by construction — the register bump is deliberately poke-free). Close it from the agent side: while ANY desired-state item is still unapplied (e.g. a pbs_dr descriptor freshly minted, guests not yet at desired run-state), the agent ticks on a fast 30 s state-based cadence instead of the 15-min backbone, and self-disarms the instant it converges. State-based, not a fixed burst — no timer to leak, no fleet-wide load once converged. |
M | idea | Rides the agent v0.90.0 train (agent-repo task; hub already mints the state). Names the coupling: flips the capability-map immediacy row's remaining "WG-registration leg unfired live" caveat → PROVEN once a real onboarding converges in seconds without a poke. Candidate train-passenger: the Guests-0/0 onboarding observation (a box reporting 0/0 guests briefly at first boot) — the same fast-tick shortens that window. Pairs with R-13 (poke channel) + R-23 (UI leg) as the third immediacy leg |
| R-14 | Headscale/WireGuard spike: Minecraft/gaming port connectivity (CGNAT-proof, sovereign DERP fallback) | M | idea | |
| R-15 | Multi-user dashboard accounts (household members, roles) | L | idea | Single password is a stated alpha limitation (R-11) |
| R-21 | Bare-metal Felhom ISO — per-PVE-release auto-install ISO for blank customer hardware → first-boot wrapper (invokes felhom-host-install.sh) → universal secret-free / operator-bind (option C) |
XL | SHIPPED (slices A+B+C, 2026-07-17) — physical N100 boot + the live boot→bind→day-0 composition fold into the supervised rehearsal (R-1) | PHYSICAL RUN 2026-07-16 (tests/VALIDATION-n100-baremetal-2026-07-16.md): demo N100 reinstalled clean-slate from a pipeline ISO → chain reached rc-0 first try on real hardware (closes slice A's operator-gated boundary), serial-filter safety proven on metal, PBS-DR reconciler self-healed on the reused peer, DMI verdict = key on MAC+UUID. F1 (HIGH, slice-B input): this cheap AMI AN3PLUS 0.01 firmware won't UEFI-boot the ISO's GRUB from USB (relocation 0x0) — SB-off/shim-bypass don't help; worked around live with a grub-mkimage loader built from the box's own GRUB. Pipeline must ship a firmware-compatible loader / PXE path. Reused-customer edges (F2 claim re-issue, F3 offsite re-issue, F4 non-default-storage-id ACL 403) feed R-1/Peti. UX: F6 drive-init doesn't mount+attach, F5 guest-RAM not configurable, F7 back-route. — Slice A (build pipeline + first-boot bootstrap) DONE + validated on VM 310: build gate/red-proof, disk-filter fail-safe, stub→retry-unit→real public-channel host-install fetch+invoke→retry, resume-decision, exactly-once, no-net retry+recovery all GREEN. Operator-gated remainder: host-install rc-0 terminal success (drill customer needs the password-gated create-UI). Slice B — SHIPPED (scripts v1.18.0, 2026-07-17): the F1 firmware fix is now a first-class pipeline mode `build-felhom-iso.sh --loader shim |
Absorbed / superseded notes in this folder
FOLLOWUP-nas-automount-guest-reboot-reassert.md— shipped (agent v0.84/v0.85, CAMPAIGN-3); keep for historyFOLLOWUP-golden-default-controller-tag.md— verify against current golden flow; close or promote to an itemFIX-M18-NOTES.md,FIX-M19-NOTES.md,DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md— historical diagnoses; superseded by shipped fixes