Files
felhom.eu/documentation/architecture/00-capability-map.md
T
admin 070b0ce2dc 07-backup-architecture: rewrite as the recovery model (D1-D6); D5 BLOCKED by R-108
Replaces the 2026-07-14 DRAFT entirely - it was verified against controller v0.132.0,
51 versions stale, while being cited as authoritative. Now verified against controller
0.183.0 / agent 0.110.0 / hub 0.80.0 / catalog 4252121. NOT ratified.

Every statement is labelled [DESIGN] (a decision from the architecture discussion) or
[FACT] (file:line, live output, or a citation to _recovery-inventory-2026-07-28.md).

Phase 0 verified D5's premise and it FAILED. Ten read surfaces were checked; nine are
clean (SMB under both namespace shapes, FileBrowser for local drives, .fab import and
download, /api/debug/*, all three ServeFile sites, registering the backup dir as a
drive). One is not: FileBrowser binds a NETWORK share at its ROOT while local drives
are userdata-scoped, and no IsNetwork() filter guards the deploy dropdown, the per-app
migrate targets, or handleStorageMigrateApp. So an app namespace - and under D5 its
plaintext secrets - can land somewhere a browsing surface serves with download:true.
Not a leak today (the unit's app.yaml is secret-stripped); it becomes one under D5.
D5 is therefore written as a BLOCKED target, not as settled. -> R-108

Core artifact is the failure -> recovery matrix (15 rows). RTO cells carry only
measured durations; 12 are blank and each blank is listed explicitly as a finding.
RPO cells carry the configured cadence that bounds RPO, labelled as such - no RPO has
ever been measured from an incident and none is invented.

The 7/53 vs 9/43/1 Tier-2 coverage counts are both recorded, unresolved; neither adopted.

Coupling (S-1): capability map section C now cites the matrix instead of restating
routes; ROADMAP gains R-102..R-108; OPEN-ITEMS gains only the in-flight R-108 plus
C9-F1b/F3/F4 -> R-n cross-references; CONTEXT records D1-D6 as standing ruling S-3.

No code, no version bump, no deploy, no box touched.
2026-07-28 20:50:41 +02:00

70 KiB
Raw Blame History

00 — Felhom Capability Map

What this is: the single cross-component truth table of what the Felhom platform can do today, at what confidence level, with verifiable evidence. Rows are scenarios (user- or operator-visible outcomes), not modules — a scenario spans agent + controller + hub + catalog, and this is the only doc that shows that view.

What this is NOT: a roadmap. Planned work lives in documentation/backlog/ROADMAP.md and is referenced from gap rows by ID (→ R-n). A row here never claims future behavior.

Status enum (strict):

Status Meaning
PROVEN-LIVE Exercised end-to-end on real infrastructure; MUST cite a campaign/drill/validation doc in documentation/audits/ or documentation/tests/. No citation → not PROVEN-LIVE.
IMPLEMENTED Shipped + unit/red-proof tested, but the real flow has not been exercised live (or not on the surface that matters — noted per row).
PARTIAL Some legs live, some missing/unvalidated — the note says which.
MISSING Does not exist. Present-tense fact; if planned, the row points at a roadmap ID.

Update rule (end-of-session checklist item): if a task changed any capability's status, update the row in the same session — with the new evidence citation. A PROVEN-LIVE claim is subject to the cardinal rule like any other claim.

Verified 2026-07-16 against evidence corpus @ felhom.eu tip 4b18cc5 by CC (capability-map audit); see REPORT.md for the per-row verdict table.


A. Provisioning & day-0

Scenario Components Status Evidence Gap / roadmap
Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box installer, agent, hub, golden PROVEN-LIVE (nested VM) DRILL-day0-vm-2026-07-12, DRILL-day0-take2-2026-07-12 First firing on real customer hardware pending → R-1
BYO install: --mode byo, mandatory caps, host-mutation disclosure, coexistence guards installer v1.15+, agent PARTIAL DRILL-GL6-2026-07-08 (demo box); GL-8 coexistence fixes Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1
Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot host-install); selectable UEFI loader; universal secret-free / operator-bind mode scripts v1.19.0 (scripts/iso/) + hub v0.62.0 + assistant container PROVEN-LIVE on TWO different boards (N100 2026-07-18; HP t740 2026-07-21) tests/VALIDATION-n100-rehearsal-2026-07-18.md — the full chain on real metal in a single pass: the generic reusable pairing ISO (v1.20.0, --loader mkimage, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted (appliance_registrations id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. F1 is closed on physical hardware. Prior nested legs: slice A SPIKE-baremetal-iso-2026-07-16 (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; --loader mkimage boots+installs SB-off; mkimage SB-enforcing FAILS Access Denied; surgery byte-identical); slice C (2026-07-17): the GENERIC secret-free ISO — box self-registers as an unclaimed appliance (POST /api/v1/appliance/register, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) F1 loader caveat: --loader mkimage fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → Secure Boot must be OFF; default shim keeps SB. Slice C bind is operator-password-gated (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17) — see the dedicated self-bind row
Box survives a wrong-NIC install: hub-unreachable first boot → legible Hungarian console screen (NIC table + remedy) + NIC sweep self-heal (bounded DHCP + hub probe per NIC, success-only persist), and the baked root password is operator-knowable (<iso>.rootpw.txt) scripts v1.24.0 (scripts/iso/felhom-bootstrap.sh network_gate/sweep_nics, build-felhom-iso.sh rootpw emission) PROVEN-LIVE (nested drill — nested ≠ metal: metal proof rides the next real multi-NIC install) audits/SPIKE-firstboot-nic-sweep-2026-07-22.md — dead-NIC install from the virgin v1.24.0 ISO baked the 192.168.100.2 fallback (WITH a dead default gateway), the R-59 screen painted on the console (screendump captured), and after the cable move the box swept to the working NIC, re-leased and self-registered at the hub unaided in under a minute; the drill also caught + fixed the stale-fallback-route trap (flush before the bounded dhclient) and verified the emitted rootpw against the installed box's shadow hash R-59 ships as a first-boot gate, not an install-time abort (recorded deviation — the fallback is the auto-installer's own, initrd hook out of scope); sweep is structurally first-boot-only (state.json gate + unit done-flag condition); a box past install-start gets the screen but its interfaces are never touched
Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) controller v0.122, hub v0.50 PROVEN-LIVE (drill VM) DRILL-day0-vm-2026-07-12 §10/F-4 (gate ON via real edge; claimed, code consumed) Never executed by a non-Viktor human → R-3. Deliverability (R-4), gmail half DONE 2026-07-18: the rehearsal's claim email was the first sent under the tightened DMARC p=quarantine and landed in the gmail Inbox, not spam (tests/VALIDATION-n100-rehearsal-2026-07-18.md). freemail.hu remains Viktor's open half. (Dropped mis-cited CAMPAIGN-4 F-C — that is the escrow-claim 502, not password claim)
Customer binds their own appliance (self-service): operator-sent 7-day tokenized capability link → public two-factor /bind/<token> (console pairing code + retrieval passphrase) → hub stages the bind, no operator hub v0.66.0 + ISO scripts v1.20.0 PROVEN-LIVE (real customer-zero bind on metal, 2026-07-18) tests/VALIDATION-n100-rehearsal-2026-07-18.md: operator minted + emailed the link 16:28:55 (7-day TTL, expiry 2026-07-25 recorded); the customer bound their own box at 16:29:55 with attempts=0, locked=0appliance_bound carries source customer_selfbind, and the credential was delivered 26 s later with no operator action. Hub-side lifecycle in hub-state.txt (selfbind_tokens mint→email→consume). Prior unit evidence: hub v0.66.0 (web/selfbind.go, store/selfbind.go; Scenarios AF + F1/F2; 4 red-proofs verified red — THE TRAP /bind/ exemption, no-oracle, lockout, single-active); GC verdict §3 (no appliance GC → TTL stands alone) R-27 slice 1. No appliance list ever rendered; wrong code == wrong passphrase (one generic failure); 5-attempt lockout → call support; expiry falls back to operator-bind. Live first-run DONE 2026-07-18 (rehearsal; the console banner rendered on the real ISO). R-27b (controller second-box dismissable prompt) deferred; multi-box-per-link = repeated operator sends
Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge controller v0.127, agent v0.88/0.89 PROVEN-LIVE (drill VM, endpoint-exact) agent v0.88.0 REPORT (ceremony ~4s, one-shot claim 200→410, R absent from every payload); SPIKE-controller-escrow-2026-07-13 Customer-facing browser wizard FIRST LIVE FIRING 2026-07-18 (tests/VALIDATION-n100-rehearsal-2026-07-18.md, S6): customer zero drove the wizard on the reborn box — ceremony started 16:56:29, recovery code claimed one-shot 16:56:39 (absent from logs by design), hub-verified and EscrowState auto-confirmed 16:56:41, offsite runs enabled 12 s after the ceremony began; the v0.138.0 „megerősítésre vár, legfeljebb 15 perc" awaiting card rendered and flipped on the ACK (operator screenshots: Viktor's set). Honest caveat: at a 12-second confirm the awaiting window is so short that catching both states on screen is luck, not procedure. Prior: endpoints driven on the drill VM. agent v0.89.0: /escrow/preflight pbs_storage_id row now live-reloads (reads current agent.json) — a pbsdr convergence that seeds the id flips it green with NO service restart. hub v0.60.0 (data-first retention): a re-escrow with a DIFFERENT sealed passphrase no longer destroys the old blob — the hub RETAINS it (host_escrow_superseded), so a previous passphrase stays recoverable with its recovery code (turns the reinstall-orphan incident from "history destroyed" into "history recoverable"). Guided-recovery flow = R-26. Red-proof TestSaveHostEscrow_RetainsSuperseded. hub v0.60.1 — custody survives the host lifecycle: host deletion (with the escrow ack) DEMOTES the current blob to retained custody (moved into host_escrow_superseded, never destroyed; existing superseded rows spared); the customer Danger-zone Delete is the one true purge point (cascades both escrow tables incl. already-deleted hosts). No operator path through host lifecycle can lose a blob. Red-proofs TestDeleteHost_DemotesEscrowNeverDestroys + TestDeleteCustomer_PurgesEscrowCustody. agent v0.93.0 (2026-07-21) — recovery codes can no longer contain a hyphenated word. The EFF large list holds exactly four entries containing the hyphen the words are joined with (drop-down, felt-tip, t-shirt, yo-yo); drawing one produced a code that reads as 11 words instead of 10 — ambiguous to transcribe in exactly the situation R exists for. They are now excluded from GENERATION only: the draw space goes 7776 → 7772 and a 10-word code 129.248 → 129.241 bits, still well clear of the 128-bit floor. Every code already issued remains valid — R is verified as a whole passphrase by the PBS scrypt KDF and is never re-split, so no customer needs to re-run a ceremony. This also retired the long-standing ~1/5 TestGenerateRecoveryCode_EntropyAndFormat flake, which was this defect and not a flaky test
DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation installer v1.15, agent v0.86, hub v0.51 PROVEN-LIVE (2026-07-21) DRILL-day0-take2-2026-07-12 §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) ⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39). On the reborn N100 the descriptor auto-provisioned and the agent reported converged state=applied (16:45:53), yet the storage is dead: pvesm statusfelhom-pbs: error fetching datastores - 401 Unauthorized / inactive, and a direct probe with the stored credential returns 401 on every endpoint including /version while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and consumed_at is still NULL; the converged state machine will not re-apply, and the agent's 15-minute verify loop cannot even read the credential to notice (open /etc/pve/priv/storage/felhom-pbs.pw: permission denied — non-root agent reading a file it writes through a root wrapper). A tier that reports applied while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See tests/VALIDATION-n100-rehearsal-2026-07-18.md F2 and pbs-dr-state.txt. agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17: the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick pbsdr: pre-check 403 … self-granting … (R-22)converged state=adopted in ~3 s, ACLs self-restored, pvesm status felhom-offsite=active, zero operator action. No more one-shot pveum grant 2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end. The three defects that let a box be applied and dead simultaneously are each addressed: the hub stamps a monotonic secret_generation into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow read verb so the non-root agent can read the credential it writes (it never could — /etc/pve/priv is 0700 root:www-data, which made the verify loop blind by construction); and pbs.ProbeAuth turns a 401 into a loud auth_failed that the existing pbsdrheal damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says applied, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (rc=0) and probed successfully (credential probe OK storage=felhom-pbs). STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged. The operator pressed Re-issue PBS credentials; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): 08:39:31Z hub mints a fresh secret, generation 0 → 1, and the descriptor gains "secret_generation": 1 — with token_id and fingerprint byte-identical, i.e. exactly the re-key shape that used to be invisible → 10:39:34 the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → 10:39:38 ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD previous_state=applied (leg c: the exact R-39 failure state, detected out loud for the first time ever) → 10:39:45 one-time token secret consumed secret_len=36 (leg a: NO short-circuit — this is the line that never appeared on 2026-07-18) → 10:39:45 felhom-pbs-apply reconcile (the set-only wrapper, no --server) → 10:39:47 pbsdr: converged state=applied. Corroboration: the agent marker hash moved to afbb3b41… (it was byte-identical to the pre-reissue marker in the failure); consumed_at stamped 08:39:45Z; the on-disk secret's mtime moved 2026-07-18 20:28:522026-07-21 10:39:45; a live probe with the NEW credential returns 200; three consecutive hub reports trace the whole state machine applied → auth_failed → applied; and zero pbsdr_selfheal escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no consumed-failed.json. Row upgraded to PROVEN-LIVE (2026-07-21). Evidence: felhom-agent/REPORT.md (2026-07-21).
Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE — local daily + offsite weekly as scheduled work agent v0.970.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 PROVEN-LIVE (2026-07-26) audits/SPIKE-r82-phase0-2026-07-26.md; per-repo CHANGELOGs/REPORTs. Restore round-trip on demo-hp: --selftest=restore-test against felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Zpass:true, verified:"boot+running", mount_parity:"ok" (mp0=/var/lib/docker 50G, mp1=/mnt/sys_drive 20G, mp8/mp9 throwaway stand-ins for the archived binds), source_tier:"pbs", 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL. That tier was PROVEN-LIVE as applied since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held zero, ever — "applied and empty", the R-39 shape one level quieter. What earns PROVEN-LIVE here: (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it restores into a bootable, mount-complete guestmount_parity is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) the multi-tier quiesce ran through the real UI endpoint (POST /api/guest-backup/trigger, authed+CSRF) and produced exactly ONE stop/start pair with BOTH backups inside itquiescing 1 stack(s) 17:01:39 → local done 17:02:56 "next tier may start (app still quiesced)" → felhom-pbs snapshotted 17:03:06 → unquiescing 17:03:06. App downtime 1m27s for both tiers, and the app came back healthy. Known gaps, recorded not hidden: the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → R-82
Restore-proof is UNATTENDED — the scheduler covers EVERY tier and a failure is heard agent v0.104.0, hub v0.77.0 IMPLEMENTED (not PROVEN-LIVE) per-repo CHANGELOGs; backlog/SPEC-r85-phase4-5-2026-07-26.md. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. The row above is earned by a MANUAL --selftest=restore-test; this one is about the SCHEDULED path, and the distinction is the whole point. Before R-85 the scheduler could only ever see cfg.Backup.BackupTarget(), so the offsite tier was never a candidate — and a failed restore-test was a [WARN] line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — restore_test_failed (broken now) and restore_test_stale (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. Why NOT PROVEN-LIVE: rotation has not been observed selecting both tiers across consecutive UNATTENDED cadences. At a 24h cadence that is a multi-day observation window, and it had not elapsed when this shipped. A single passing run proves the code path, not the rotation. → R-85
Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive hub v0.61.0, felhom-tenantsync v1.1.0 PROVEN-LIVE (external teardown, incl. two real firings) tests/VALIDATION-n100-rehearsal-2026-07-18.md — two live firings, both host-delete-first, on two different customers (demo-vm-felhom 15:49:57, demo-felhom 16:08:51): every leg ok (claim, db_purge, descriptor, hetzner, pbs), escrow acked separately, each completing in 89 s (hub-state.txt customer_resets). The Hetzner sub-account destruction is now verified against the live pool box — and produced the run's sharpest lesson: a sub-account is an access-control object, not a data object. Deleting it left its /home intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (a finding by S7's own criterion) and why RESET now needs a base-dir purge → R-32. Prior: hub v0.61.0 REPORT; ep0 live drill 2026-07-17 (throwaway drill-reset-01 with a real backup: deprovision deleted:true destroyed the namespace + backup group + token, idempotent re-run deleted:false, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. Live-clicked 2026-07-18 (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. R-25b CLOSED (hub v0.69.0, 2026-07-21): the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below
Customer DELETE cascade (top lifecycle tier): one guided operator action → hosts → RESET → residue → purge; host rows deleted (custody DEMOTED), full external teardown (Hetzner repo destroyed, PBS namespace/token revoked, tunnel + zone removed), then the customer record and ALL escrow ciphertext purged hub v0.69.0 UNIT-PROVEN; live leg PENDING hub/internal/web/customer_delete_test.go — leg ORDER observed from inside leg 2 (hosts already gone, customer row still present, custody still retained); 9 fail-closed gate cases each asserting zero mutations + zero external calls + no journal row; resume-after-external-failure converges; purgeEscrow custody semantics; preview leaks no secret. 5 red-proofs (ack gate, stale-preview gate, ONLINE-host gate, leg order inverted, purgeEscrow=true) Three acknowledgements + typed customer-id + stale-preview check + ONLINE-host refusal, ALL before any write. Ruling-3 preserved BY CONSTRUCTION (leg 2 never sees a host row); custody purged EXACTLY ONCE, in leg 3. Coupling: hub-only — no agent/controller/catalog change; the cascade calls the same service paths as manual host-delete and standalone RESET, so their rules move together. v0.70.0 (2026-07-21): added the residue leg — GetCustomers() is REPORT-derived, so before it a fully deleted customer stayed on the Customers list and its report stream kept the staleness/offsite checkers alerting (live: demo-vm-felhom deleted 07-18, still emailing offsite_stale on 07-21). The leg also purges the credential-bearing appliance_registrations + selfbind_tokens. Ghost customers (config row already gone) are now deletable — 404 means "nothing here", not "no config row"; the Hetzner/descriptor legs record skipped_no_config. Gap: the end-to-end live leg on a scratch customer (external Hetzner teardown observed from outside) is not yet run
Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling installer PARTIAL DRILL-GL6-2026-07-08 Phase 1/5 (KEPT-vs-WIPED printed verbatim; drive data intact ×3); GL-4 code Secret purge (GL6-F1 .bak residue) fixed v1.12.0; enrolled-drive mnt-*.mount units survive (GL6-F2, open); cluster-aware felhom_guests guard + saferemove cost warning missing → R-9

B. Apps & catalog

Scenario Components Status Evidence Gap / roadmap
Deploy an app from the catalog (env config, memory guard, health-aware progress) controller, catalog (~52 apps, images pinned) PROVEN-LIVE CAMPAIGN-2 T-DEPLOY-SET (7 apps, env config, health-aware); RERUN-p1p3 (×4 PASS) Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only
App lifecycle: start/stop/restart/update/logs/remove/redeploy controller PROVEN-LIVE CAMPAIGN-2 T-LIFECYCLE (stop/start/restart/update/logs); remove live in CAMPAIGN-3 Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical
Protected infra stacks can't be stopped/removed from UI controller PROVEN-LIVE CAMPAIGN-nomercy + RERUN-p1p3 T-SEC-PROTECTED (refuse stop/remove, stay Up) (Cited CAMPAIGN-2 T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal)
Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad backup: block degrades to legacy, loudly) controller v0.132, catalog PROVEN-LIVE CAMPAIGN-2 T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs
Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) agent v0.94.0→v0.95.0, controller v0.169.0→v0.171.0, hub v0.73.1 PROVEN-LIVE (healthy path + delivery) 2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix): the card on guest 9201 now shows BOTH real disks with real verdicts + human model labels„AirDisk 512GB SSD" → Rendben (34°C) (the system SSD, via LVM/dm resolution) and „TOSHIBA MQ04ABF100" → Rendben (30°C) (the USB, via union-path SMART). /disks carries smart.health=PASSED + model_name for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (SPIKE-smart-coverage-2026-07-25.md had proven both disks answer smartctl -a -j PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. Notification pipeline PROVEN-LIVE 2026-07-24 — a disk_health_degraded POST (the exact notify.PushEvent wire call) was 400-rejected by hub v0.73.0 and 200-accepted + „Operator email sent" by hub v0.73.1 No new smartctl load; feature-detect by payload presence → MinAgent floor unchanged; no sudoers/-d sat change. No global banner (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin local on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + model_name capture. A live FAILING/Figyelmeztetés verdict still needs a genuinely degrading disk (the demo disks are healthy) — that arm stays unit+red-proof covered. SMART history/trending (hub-side) PARKED (ROADMAP R-73)
App crashes → customer notified (one event per transition, no flapping spam) controller v0.120, hub v0.48 IMPLEMENTED controller v0.120.0 (dead-app alerting, app_start_failed, one-event-per-transition red-proofs); CAMPAIGN-3 F11 surfaced the gap End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven
Post-deploy optional config (API keys etc.) with restart controller, catalog .felhom.yml IMPLEMENTED feature long-standing; config page renders (CAMPAIGN-2 T-PAGE-ALL is GET-only) The config-save+restart flow is exercised in no campaign (CAMPAIGN-3 explicitly skipped interactive app config). Demoted: T-PAGE-ALL is a page-render smoke test, not this flow
Backup classification: 13 bind-bearing apps carry mandatory/optional/excluded classes catalog, controller v0.132133 PROVEN-LIVE SPIKE-backup-classification-2026-07-14, CAMPAIGN-6D/6E Remaining ~39 apps are legacy-class by design (unit-only offsite)

C. Protection & recovery (the product promise)

Coupling (2026-07-28, S-1). The failure → recovery matrix in 07-backup-architecture.md §8 is authoritative for which failure has which recovery route, who can invoke it, and what its measured RTO is. This section stays authoritative for per-capability status. Neither restates the other — rows below carry a → 07 §8 row n pointer instead of repeating the route. Where a row's status and the matrix's status differ in wording, the matrix is about the failure and the row is about the mechanism; that is not a contradiction, and both cite the same evidence.

The matrix's blank RTO/RPO cells are deliberate: no number is estimated anywhere. Two counts of Tier-2 app coverage disagree (9/43/1 vs 7/45/1) and are both recorded there, unresolved — do not adopt either from this page.

Scenario Components Status Evidence Gap / roadmap
Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes controller v0.118 PROVEN-LIVE CAMPAIGN-2 T-BAK-FULL (pg+mariadb autodiscovered); atomicity CAMPAIGN-6B P4 + CAMPAIGN-6E B1/B2 (SIGKILL mid-write → only .tar.tmp touched, last-good byte-unchanged); DB restore CAMPAIGN-6D P-FAB (Cited CAMPAIGN-3 F7 is the finding of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 DB replay route → 07-backup-architecture.md §8 row 3
Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary controller v0.135 PROVEN-LIVE CAMPAIGN-6E-2026-07-15 (P-TIER2 deep-4 PASS), CAMPAIGN-6C Route + RTO → 07-backup-architecture.md §8 rows 1, 2, 4, 5. The matrix records that the copy's recovery-unit/ mirror is read by no path (→ R-102)
Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping controller v0.134, agent, hub PROVEN-LIVE CAMPAIGN-6D-2026-07-15 (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); VALIDATION-offbox-storagebox-2026-07-09 (byte-perfect round-trip) Raw-data quota (SP-1) + retention regrouping (SP-2) are SPIKE-restic-snapshot-shape dry-run verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. Reinstall-continuity (controller v0.142.0, 2026-07-17): a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (wrong password or no key found) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. The live leg FIRED on its own during the 2026-07-18 rehearsal (tests/VALIDATION-n100-rehearsal-2026-07-18.md, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard classified it, pushed offbox_repo_orphaned, skipped the run and showed the card (16:58:14) rather than nightly-spamming a raw restic error; the operator-confirmed reset then moved the repo aside (never deleted) to .orphaned-20260718 and re-initialised (16:59:26→16:59:32), and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — the finding is that it had to fire at all (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). DIAGNOSE-offbox-repo-orphaned-2026-07-17 Route + RTO → 07-backup-architecture.md §8 rows 4, 10, 12, 15 (incl. the R-95 delete exposure and the R-104 stale-lock defect)
Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live controller v0.134/134.1/135 PROVEN-LIVE (2026-07-20) CAMPAIGN-6D accept legs (immich end-to-end from offsite alone) 2026-07-19: audits/DIAG-immich-restore-2026-07-19.md finds no offsite path loads a DB dump — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "immich end-to-end from offsite alone" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume (immich_postgres_data is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" overclaimed scope: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → PARTIAL, scope-corrected. Evidence: audits/DIAG-immich-restore-2026-07-19.md (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). 2026-07-19, controller v0.148.0: the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but round 2 found it aborts against a running app (audits/DIAG-immich-restore-round2-2026-07-19.md, H4: the replay races immich's own schema repair; clip_index recreated by the app 2 s before the dump's CREATE INDEX). 2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47) — both restore paths now replay into a DB-ONLY window (StartStackServices brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. (The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard BackupStatus fix. R-47 shipped in v0.153.0.) 2026-07-20: the clean run HAPPENED — endpoint-level supervised reconstitute of immich from snapshot 49e7cb46 (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no already exists, operation reported SUCCESS, immich's own DatabaseService logged No schema drift detected twice, 11 assets active, 4/4 containers healthy, 231 public indexes. Operator confirmed the immich timeline renders correctly after the reconstitute (screenshot held, 2026-07-20). Evidence: felhom-controller/REPORT.md §4b. 2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE. The operator deleted the photos in immich own UI and emptied the trash (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. 40 file(s) placed against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets active, No schema drift detected, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: felhom-controller/REPORT.md 4e Route + RTO → 07-backup-architecture.md §8 rows 3, 4 — the matrix also records that no offsite action unpacks the named-volume tars it captures (→ R-107)
Manual .fab export/import: class-scoped capture, browser up/download, tunnel-proof chunking controller v0.125/128/130/136 PROVEN-LIVE CAMPAIGN-6D P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking CAMPAIGN-6B P2 (100 MiB via real CF edge, 120 MiB→413) Chunking proven at the real CF edge via curl --resolve; the rendered browser file-picker upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B finding; fix verified in 6D
Guest-loss DR: PBS restore with full-fidelity layout from archive, restore-test verification agent v0.75/0.76, PBS PROVEN-LIVE CAMPAIGN-2 T-P9-DESTROY-RESTORE (whole-guest pct restore of 9201 → running+healthy) + T-PBS-VERIFY (verify_state: ok, 13 snapshots); DRILL-GL6-2026-07-08 Phase 0d (restore-test mount_parity: ok) (Cited VALIDATION-newbox-restore is offbox restic file-restore, wrong tier — corrected.) Real offsite guest-loss round-trip still R1-blocked → S5 DR drill Route + RTO → 07-backup-architecture.md §8 rows 6, 8, 9 — measured 84112 s local / 1101 s PBS into a scratch guest; a restore to a DIFFERENT host is unmeasured
PBS-DR secret self-heal on reused-peer re-provision hub v0.56 IMPLEMENTED hub v0.56.0 (pbsdrheal/reconciler.go, RestageHostPBSSecret, all §10 red-proofs); SPIKE-pbsdr-selfheal-2026-07-15 (root cause) Reconciler is scoped to one host (PBSDRHEAL_ONLY_HOST), not fleet-wide; already fired live hands-free on drill qm300 (07-15) — real-customer firing + fleet-wide widening pending
Box survives an unattended app or guest-network failure (a dead app member, a boot-orphaned app, a dead DHCP client) — it is noticed, and where safe it is repaired controller v0.156.0, agent v0.92.1 PROVEN-LIVE (2026-07-21) All three legs exercised on the live demo box, operator-present, in one session — felhom-controller/REPORT.md + felhom-agent/REPORT.md (2026-07-21). Dead primary: docker stop immich-server 12:50:40 CEST → degraded 13 s later → exactly one app_start_failed + dashboard banner → restart → banner self-cleared (the 2026-07-20 shape that was silent for 18 h). Boot orphan: pct reboot 9201[bootrecon] 1 boot-orphaned app(s) found: [bookstack] → started in 1 attempt of 2, zero alerts (success inside the boot grace is silent); StartedAt proves Docker's unless-stopped did NOT resurrect it — only the sweep did, which also answers P1 and confirms the F5 hypothesis. Dead DHCP client: deliberate replay of the incident — kill -9 12:43:18 → detected on process liveness 57 s later while the lease was still live → healed 12:45:18 with the incident's verbatim invocation; the tunnel never dropped (cloudflared Up 29 hours), i.e. the outage was prevented rather than merely observed The gap this validation surfaced → R-55, now FIXED (controller v0.157.0, 2026-07-21). For a drive-backed app a customer's deliberate Stop did NOT survive a reboot — the boot bind gate recreated and started every deployed drive-backed app unconditionally. Pre-existing, not introduced by R-52 (whose own gate was observed correct). The gate now also requires the app to still HAVE containers, which is R-52's own existing-Exited vs absent predicate: a UI Stop is compose down and removes them. So "a stopped app stays stopped" now holds for drive-backed apps too — PROVEN LIVE 2026-07-21 (TASK-F Part 3, operator-present). immich was stopped through the real UI endpoint (compose down → 0 containers), calibre-web and bookstack left running, then pct reboot 9201: the gate recreated calibre-web and logged 1 drive-backed app(s) left stopped — zero containers means the customer stopped them on purpose; immich came back stopped, where the identical fixture had brought it back running hours earlier. Zero alerts, ~15 s to steady state. The static-guest half of the network leg stays deliberately out of scope → R-50
Crash/power-loss mid-backup/mid-migration → self-heal on next run controller, agent PROVEN-LIVE CAMPAIGN-6D P5-REST (SIGKILL mid-offbox → auto-restart ~15s, run marked failed not false-success, no stale lock); CAMPAIGN-6E B1-B3 (Cited CAMPAIGN-2 T-RBT-* legs were empty / auth-hollow — corrected.) Live mid-migration crash→self-heal is the weakest sub-claim (P5-REST is mid-backup)
An app can be withdrawn from the catalog without orphaning the customers running it (available / hidden / abandoned) controller v0.158.1, catalog metadata PROVEN-LIVE (2026-07-21) TASK-F Part 1. Verified on 9201 through the real endpoints: lifecycle: abandoned arrived via the normal catalog sync; plant-it renders 0 times on the Alkalmazások page (control app renders 10); a direct POST /api/stacks/plant-it/deployHTTP 409 "Ez az alkalmazás jelenleg nem telepíthető."; the app page carries the permanent notice and offers no Telepítés button. felhom-controller/REPORT.md (2026-07-21) Deployed instances keep FULL function in every state — lifecycle governs what is offered, never what runs. Orphan detection deliberately never sees the field (red-proofed): a withdrawn template stays in the catalog tree, or every deployed instance would read Elavult and be offered deletion. Unknown values fail OPEN; the deploy gate fails CLOSED. R-57
Box survives a site/network change (relocation, different subnet, DHCP re-lease) with the control plane intact agent v0.96.0 (island NIC), host-install v1.19.0, controller (unchanged), bootstrap PROVEN-LIVE (2026-07-25) R-50 SHIPPED and deployed to the whole fleet. The control plane now rides a host-internal, portless island bridge (vmbr9, 169.254.253.1/30.2/30) with a fixed private address that no LAN/DHCP/site move can invalidate. Proven end-to-end: the spike's F1 replay (renumber the LAN → agent stays bound on the island, control plane HTTP 200; the LAN-literal contrast reproduces the original bind: cannot assign requested address daemon-death) + cold-reboot survival (SPIKE-island-bridge-2026-07-25.md), the migration runbook run verbatim (RUNBOOK-island-migration.md), a fresh provision auto-attaching the island net1 (A4), and the live migration of both demo boxes (demo-hp + demo-felhom, 2026-07-25) — island /storage HTTP 200, LAN DNS pinned to the LAN IP (Finding-1), apps served throughout (0 container restarts), hub reporting 0.96.0. Origin: audits/AUDIT-vacation-remote-ops-2026-07-20.md — the real relocation where the agent's LAN-literal bind took storage/PBS/quiesce/restore-test/DR down silently; that is now structurally impossible on a migrated box Fleet: DONE. Remaining: R-74 — bring the island to Peti's 2-node cluster (SDN vnet / bridge parity), its own supervised runbook. Related historical: R-51 (dead-primary alerting), R-52 (boot desired-state reconciliation), both shipped
Soft-quota: usage bar, pre-push enlargement block, customer notification controller v0.109/134, hub v0.41/55 PROVEN-LIVE 6D/6E; hub OffsiteChecker
A customer (not the operator) performs a restore via UI alone all MISSING (as evidence) Alpha will produce this; script it into R-3. 2026-07-19: the C6 evidence attempt ran and found a product gap instead of evidenceaudits/DIAG-immich-restore-2026-07-19.md. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING by finding, not by absence of attempt — the rehearsal system working, not failing. 2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0 (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. 2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped (audits/DIAG-immich-restore-round2-2026-07-19.md): a deliberate run from snapshot 49e7cb46 did recover all 11 assets (status=active, files resolve), but the operation reported failure and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. 2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED. (The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.) The C6 drill ran end-to-end through the UI: photos deleted, trash emptied, the full files+database restore pressed on /backups/restore, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: felhom-controller/REPORT.md 4e. Residual: the run was performed by the OPERATOR, not by a customer — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means trash, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash and verify the app's content, not the file count Lane split → 07-backup-architecture.md §3: this row is Lane 1 (customer, unassisted). §8 rows 15 are the routes it would exercise

D. Storage & devices

Scenario Components Status Evidence Gap / roadmap
Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts controller, agent v0.87 PROVEN-LIVE DISPOSITION-ia-finding2-systemdisks-2026-07-13 (legacy EFI+LVM host, root not offered, byte-identical); enroll/format live in storage-lifecycle-acceptance-2026-06-15 (E10 re-enroll, data intact); agent fence self-test refuses /dev/sda (Cited CAMPAIGN-2 T-STG-ENROLL/SEC-FORMAT were auth-hollow CSRF-403.) Fresh-USB wizard enroll+format through the customer UI PROVEN-LIVE (controller v0.141.0, 2026-07-17): a 64 GB scratch USB driven through the real /api/storage/init endpoints (login+CSRF) → confirm → detached format (~27 s mkfs) → mount → register → mounted+registered at /mnt/felhom-drives/scratch1. F6 (initialize-to-usable) now covered: the wizard runs the chain as a detached, disconnect-safe, pollable job (3-step progress) with an agent format-status poll for a slow mkfs 2026-07-26 — a SILENT failure class on the channel every agent-backed capability depends on (this row, data migration, USB enrollment, guest RAM, quiesce/PBS) is now DETECTED (controller v0.173.0, R-77). No row status flips. controller.yaml and bootstrap.json could disagree on local_api.endpoint indefinitely with no signal: the R-50 island migration rewrote the latter, the fleet kept dialling the former, and for 17.5 h the only alert was a generic "agent unreachable" that read as an infrastructure blip. Drift now raises its own event type (local_api_endpoint_drift) naming both values. It is DETECTION ONLY — the authority ruling is R-78 — so the class is now loud, not prevented. Evidence: audits/DIAG-agent-channel-2026-07-26.md.
Data migration between drives (all / per-app), crash-safe controller PROVEN-LIVE CAMPAIGN-6C 4P-5 (scope=app round-trip, byte-identical); storage-lifecycle-acceptance-2026-06-15 (two migrate-all runs via dashboard UI, sha256 byte-identical) (Cited CAMPAIGN-2 T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS
NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated controller v0.113117, agent v0.81/84/85 PROVEN-LIVE SPIKE-nas-verify-2026-07-11, SPIKE-nas-dsm-2026-07-11, CAMPAIGN-3-2026-07-11 (boot/reassert fixes)
USB drive enrollment + unplug detection + recommission controller, agent PROVEN-LIVE storage-lifecycle-acceptance-2026-06-15 E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); CAMPAIGN-4/6A (3 USB re-establish across device-letter reshuffle) (Cited RUNBOOK-usb could NOT complete a wizard enrollment; CAMPAIGN-2 legs were auth-hollow.) Fresh-USB wizard enrollment specifically still unproven
Decommission (migrate-first and anyway-paths), eject agent, controller PROVEN-LIVE storage-lifecycle-acceptance-2026-06-15 E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) (Cited CAMPAIGN-2 T-STG-DECOM-* were auth-hollow; SPIKE-decommission was report-only, button still vestigial.)
Boot ordering: automount + networking survive reboot; appliance self-heal watchdog agent v0.85 PROVEN-LIVE CAMPAIGN-4-2026-07-13 (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + CAMPAIGN-6A-2026-07-14 1D (re-arm reboot-survival across 9 guest + 1 host reboots) (CAMPAIGN-3 F10/F11/F12 were the CRITICAL/HIGH failures; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: skip-active on pct reboot carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A)

E. Access, networking & household use

Scenario Components Status Evidence Gap / roadmap
Remote access via Cloudflare Tunnel + Traefik (per-app subdomains) cloudflared, traefik PROVEN-LIVE CAMPAIGN-2 T-FLT-CF Per-customer zone-scoped CF tokens (blast-radius ruling)
LAN access when internet is down (lan_resolver) agent IMPLEMENTED Never drilled as a customer experience ("net down — can I reach my photos?") → R-19
Phone photo backup immich (classified) PROVEN-LIVE 6D end-to-end restore proof
Documents/OCR paperless-ngx (classified) PROVEN-LIVE CAMPAIGN-6C 4P-1 (deploy paperless-ngx, ingest 3 docs via consume flow, OCR + PDF/A ~90s) + 4P-2/3/5 Consume-folder ingestion awkward without SMB → R-7
Files from Windows Explorer / Mac Finder (SMB server) controller v0.145.0 + felhom-samba:1.0.0 PROVEN-LIVE felhom-controller REPORT.md (v0.144.0) + controller/sharing.md; transport verdict audits/SPIKE-lan-discovery-2026-07-18.md „Megosztás" page: enable + one household password + shares (new folder or picked existing, per-share read-only). Fourth protected infra stack (host-net, smbd+nmbd+wsdd). Live on demo: 445 reachable, NetBIOS FELHOM resolves, write/read byte-compare PASS, write to a read-only share REFUSED, SMB writes land as uid 1000. Explorer leg PASSED 2026-07-18 (Viktor): Network → FELHOM → both shares open; a real Explorer save into dokumentumok landed owned uid 1000, and a write into the read-only filmek was refused by Windows with the folder left untouched. Share data RIDES BOTH BACKUP TIERS (R-7b, controller v0.145.0, Model B sibling shares source): tier-2 cross-drive legs + an offsite _shares restic snapshot carrying the share definitions and the credential copy, with a „Megosztások" restore. All four legs PROVEN-LIVE on demo 2026-07-18 — tier-2 tree md5-verified; offsite snapshots e0b9d723 (Viktor 12:18:16Z) and 4e2b15ec both carrying manifest + passdb.tar; restore round-trip returned a deleted probe file byte-identical and a deleted share DEFINITION with its original flags without overwriting live files; samba liveness → hub-accepted health_critical. Remaining human leg: SMB positive auth with the real household password
Media to TV via DLNA MISSING Jellyfin app exists; DLNA/SSDP unvalidated → R-6, R-8
File access via browser FileBrowser (infra app, auto-mount sync) IMPLEMENTED FileBrowser runs healthy + userdata-bound (storage-lifecycle-acceptance-2026-06-15, CAMPAIGN-3) Actual browse/download through FileBrowser is exercised in no doc. (Cited CAMPAIGN-2 T-PAGE-ALL renders only the controller dashboard pages, not FileBrowser.) Demoted. 2026-07-26, controller v0.172.0 (R-75) — status DELIBERATELY UNCHANGED. The canonical drop-zone now has its own FileBrowser source („Beolvasás" → /srv/beolvasas, a separate bind of <system namespace>/userdata/import) and the app page carries a per-app deep link into it. Verified live on demo-hp: the source and bind are in the generated config, the app page renders https://files.enkisfelhom.hu/files/Beolvas%C3%A1s/paperless, and a file written through FileBrowser's OWN mount was consumed and deleted by paperless in ~30 s. That is still not a browse. Nothing in this arc drove the FileBrowser HTTP UI — no browser exists on DooPlex — so the row's standing caveat survives intact and the upgrade to PROVEN-LIVE remains unearned. What it would take: a human click-through, or an authenticated /api/resources round-trip against the live instance. See controller/import-and-data-paths.md
Indítópult (app launcher) — one-tap grid of the household's openable apps controller v0.163.0 IMPLEMENTED New FIRST sidebar page /launcher: colored tiles (deterministic slug color or .felhom.yml brand_color) + white glyph/monogram, one per openable app (tile ⟺ „Megnyitás" — subdomain presence is the single criterion; controller excluded). Operational → <a target=_blank> to the public URL; stopped → greyed + state badge, no link. / stays the Vezérlőpult. Endpoint-level + render-test verified; felhom-controller/REPORT.md (2026-07-24) Live operator click-through of a real tile → app pending (browser automation not available on DooPlex). Follow-up: curate brand_color for top catalog apps (R-72). Sharing the launcher outside the household is now the capability-URL guest link — see the row below
Indítópult megosztás (vendég link) — capability URL /s/<token> serves a standalone read-only guest launcher (no account, no admin session); optional per-share password; QR controller v0.165.0 IMPLEMENTED 160-bit crypto/rand token, constant-time match (empty stored = disabled = byte-identical to the mux default 404); guest headers noindex/no-referrer/no-store; optional SEPARATE bcrypt share password + its own per-IP attempt map; signed cookie HMAC(token|passwordHash) keyed with session_secret (rotate-token OR change-password invalidates all cookies); token redacted in logs (/s/<redacted>). Groups AG (14 tests) + 3 red-proofs; §13 endpoint-level live validation on 9201 all-pass (felhom-controller/REPORT.md 2026-07-24). Design ruling: member accounts SUPERSEDED by this capability-URL model; per-member tile visibility parked under the SSO arc (R-15). Full operator browser click-through + a validation doc pending → then PROVEN-LIVE. Accepted residuals: link-preview crawlers fetch once (noindex prevents indexing); reverse-proxy/CF access logs hold the path (ops-tier); the modal link carries the request Host (LAN-IP admin ⇒ LAN-IP link)
Forgot dashboard password → instant reset code controller v0.123, hub PROVEN-LIVE DRILL-day0-take2-2026-07-12 F-15 (live re-run of the exact failure path: hash applied 1s after request, code accepted first try)
Multiple household users / per-person accounts MISSING Single dashboard password; acceptable for alpha → R-15
WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) agent v0.72, hub v0.35 IMPLEMENTED SPIKE-oob-wg-operator-peer-2026-07-05, SPIKE-felhom-sshd-2026-07-05 Mutual-repair desired-state arc not built → R-13
Break-glass management-plane recovery agent v0.71, hub v0.34 IMPLEMENTED runbooks/break-glass.md

F. Notifications & monitoring

Scenario Components Status Evidence Gap / roadmap
Health-degradation email (edge-triggered, cooldowns, Hungarian) via hub → Resend controller, hub IMPLEMENTED delivery pipeline live-proven for the enlarge-block trigger (CAMPAIGN-6D P3-DELIVERY, op+customer "Kedves Ügyfél!"); NotifyHealthChange ok→warn/fail edge-trigger implemented The health-degradation trigger specifically has never fired an email live in any doc. Demoted (pipeline proven for a different event). Deliverability to HU freemail → R-4
Event catalog: app_start_failed, dead-app, offbox_enlarge_blocked, claim/reset codes, critical severity controller, hub v0.31/48/50/55 PROVEN-LIVE live-delivered: CAMPAIGN-6D P3-DELIVERY (enlarge-block, op+customer); DRILL-day0-vm F-4 (claim code); DRILL-day0-take2 F-15 (reset code) app_start_failed/dead-app delivery is unit-only (6C inconclusive) — the pipeline + 3 event families are live, those two are not
Prefs safety: empty-email wipe guard controller v0.137 + hub v0.71.0 IMPLEMENTED controller leg red-proofed 07-15; hub-side no-clobber belt (handleSavePreferences preserves a stored non-empty address on an empty-email push) red-proofed 07-22 Born from a live incident; controller 0.160.0 guards both its push legs, so the hub belt covers older/rogue boxes
Paired recovery notifications + prefs seeding at claim + priority headers (power-outage audit F11/F12/F14-light) hub v0.71.0 IMPLEMENTED (recovery leg PARTIAL until a live staleness cycle fires it) hub/CHANGELOG.md v0.71.0; 17 tests + 4 red-proofs (REPORT.md 2026-07-22); Resend headers mechanism probed live (HTTP 200) pre-implementation; operator+customer test rows live-fired via the controller's own test endpoint Recovery = explicit eventType branch, severity semantics frozen; customer gate = PAIRING (notification_log evidence), not enabled_events. Live legs pending: a natural *_recovered mail (next real staleness cycle or the reboot-drill arc — never fabricated by blocking reports) and seed-at-claim on a real claim (Peti Friday reinstall). F14-full (operator push channel, ntfy/Telegram) stays open → R-69
System + container metrics (SQLite, Chart.js, 30-day downsampling) controller IMPLEMENTED metrics collection + /monitoring render present (page 200) The cited CAMPAIGN-2 T-RES-CPU/T-SOAK-LOOP are H1/H2 harness artifacts (auth-302), not metrics tests; SQLite/Chart.js/30-day downsampling validated in no campaign. Demoted
Always-on debug rings + on-demand log-bundle pulls with TTL/custody controller v0.116, agent v0.83, hub v0.46 PROVEN-LIVE debug rings live-exercised CAMPAIGN-3 fix-6 (1000-cap ring, ~55min horizon under load) The log-bundle-pull TTL/custody half is changelog-only (no dedicated observability audit doc); ring persistence across restart is a known gap
Operator alerting (Healthchecks → monitoring@felhom.eu) k3s, Resend IMPLEMENTED operator infra, stated in production since 02-04; no corpus validation doc
Backup-deadline alerting (expected_backup_missed) is ANCHORED — absence of signal is UNKNOWN, not failure hub v0.75.0 IMPLEMENTED audits/DIAG-backup-missed-2026-07-26.md + red-proofs A/B/C + replay of the real 2026-07-26 03:00 reports (all three silent) No row status flips — this signal had a FALSE-POSITIVE class (three instances: hub v0.12.0, v0.73.0, R-81), now anchored at first contact and read across retained host-report history. Still unit-proven only, not live-fired at a real deadline. The underlying PBS/offsite-DR tier gap it exposed is → R-82.

G. Fleet & operator (hub)

Scenario Components Status Evidence Gap / roadmap
Customer/host management: 8-tab detail, scoped auto-refresh, safe stale-host deletion, capability chips hub v0.470.53 PROVEN-LIVE hub v0.53.0 dead-host roll-up live on the Peti cluster (proxmox1 down 23h); CAMPAIGN-4-2026-07-13 (operator UI driven live); DRILL-day0-take2 F-16 (offsite/freeze buttons live) 8-tab render + capability chips are render-test-validated (hub UI is password-gated; CC cannot log in). (Cited "daily operator use" was a no-doc citation; AUDIT-hub-gui-2026-06-30 predates these features at hub v0.25)
Config/state change round-trips in seconds (hub↔box immediacy; 15-min cycle stays the backbone): box→hub out-of-cycle report (Dir 1) + hub→box GET /api/v1/wait long-poll wake (Dir 2) controller v0.139/140, hub v0.58/0.63 PROVEN-LIVE (2026-07-21) Transport proven live through the real DNS-only ingress: SPIKE-immediate-sync-transport-2026-07-16 + hub v0.58.0 / controller v0.140.0 REPORTs — 240 s no-annotation hold (25 s heartbeat defeats nginx's 60 s proxy_read_timeout, no ingress change), 0.047 s wake-on-change, hub rollout restart = 1 WARN + 0-storm reconnect; Dir-1 2 s box→hub round-trip live in controller v0.139.0 The operator-UI save→apply round-trip is not fired end-to-end live (hub UI password-gated; CC can't log in) → R-23; the wake transport and the ACK→config_version→ConfigRefresher delivery chain are each proven, only the UI-triggered bump leg is unexercised. Agent-plane (host-domain desired-state) poke first slice PROVEN-LIVE (Direction-2a, agent v0.89.0 + hub v0.59.0, 2026-07-17): contentless ep0-relayed UDP poke → agent immediate desired-state cycle, per SPIKE-immediate-sync-transport-2026-07-16 P4. Full path live-proven: a real operator manifest save fired poke: sync-poke delivered to 10.77.0.2; the box (0.89.0) received it and logged poke received → triggering an immediate desired-state cycleout-of-band report triggered~31 ms ep0→box, sub-ms to the report cycle (WG-confined, from 10.77.0.1 to the 10.77.0.2-bound socket); save→tick ≈ ~0.45 s (SSH-dominated), well under ≤23 s. R-13 first slice (listener+sender only; the rest of the mutual-repair arc stays open). System-initiated immediacy wired (hub v0.63.0, this REPORT): the mutation sites that only OPERATOR actions used to notify now fire the correct plane's notifier when the hub itself mints state — agent-plane pokes at PBSDRAutoProvision (the observed slice-C lag), ReissuePBSDR (also the pbsdrheal escalation), handlePBSDRReissue, and the two admin desired-state api writers; controller-plane bump at reissueOnReenroll. Unit-tested + red-proofed, not yet fired on a real system event (folds into the rehearsal bind sequence). Still PARTIAL: the R-23 operator-UI save→apply leg and the agent fast-tick-until-first-convergence SECONDARY (the WG-registration leg a poke can't reach pre-tunnel) remain unfired live. Fast-tick SHIPPED (agent v0.90.0, R-28): while any desired-state item is unapplied — incl. the pre-tunnel window a poke can't reach — the agent pulses the out-of-band trigger every 30 s and self-disarms on convergence (state-based; four cached sources; LOUD states excluded). LIVE on both demo agents (the fast-tick armed: 30s … startup line verified). REAL-ONBOARDING PROOF DONE — tests/VALIDATION-n100-rehearsal-2026-07-18.md (ledger 8, S5): on a genuine first onboarding on metal, every post-bind leg landed seconds apart with no ~15-minute stall anywhere — bind 16:29:55 → credential delivered 16:30:21 (26 s) → agent 0.90.0 up 16:30:49 → WG registered + tunnel applied 16:30:51 (~2 s) → poke listener 16:30:54 → controller 16:32:28 → floor-lifted and running current 16:32:39. Bind → running-current = 2 min 44 s. The PBS-DR descriptor auto-provisioned on the same cadence (agent converged state=applied 16:45:53) — though see the DR-tier row: the descriptor converged while the credential behind it was already stale (R-39). The pre-tunnel fast-tick window is therefore proven in its real setting; the remaining PARTIAL is the R-23 operator-UI save→apply leg alone 2026-07-21 — THE LAST PARTIAL LEG IS CLOSED (R-23(a) restart leg). The operator saved the global floor to a version the box did NOT run (0.153.0 → v0.154.0) and the managed self-update fired exactly once: 06:57:13Z UpdateState pending (initiated_by=auto-floor) → 06:57:17Z agent controller swap requested06:57:21Z container restarted → 06:57:29Z new controller healthy. Save → healthy on the new version = 16 s. Over a 39-minute window: swap requests 1, agent-driven bootstrap restarts 1, rollbacks 0, container RestartCount 0; VerifyStartup confirmed on the next boot and the following periodic check logged Current version 0.154.0 is up to date (the at/above-floor branch doing nothing, as designed). The 2026-07-20 attempt could not prove this because it targeted an already-running version. Evidence: felhom-controller/REPORT.md §6.
Customer right-sizes guest RAM from the controller (agent-enforced bounds, live cgroup apply, no reboot) agent v0.90.0 + controller v0.143.0 (R-24) PROVEN-LIVE (grow and shrink on metal, 2026-07-18) tests/VALIDATION-n100-rehearsal-2026-07-18.md (ledger 9) — the apply is now proven in both directions on a normal-sized box: customer zero shrank 11675 → 8192 MB at 16:50:22 and grew 8192 → 12288 MB at 17:02:17, each a live cgroup apply with no reboot (local-api: guest-memory resized in the agent journal, [web] memory resized in the controller log), and the new total rippled into the deploy page's memory math at 17:05:15 (total=12288MB). F5 auto-sizing had landed the guest at 11675 MB. Controller-direct (R-24's hub-desired-state framing SUPERSEDED, Viktor 2026-07-17). Agent GET/POST /guest/memory enforces every bound FRESH (min 2048 / max host_total2048 / shrink floor max(2048, usage+512)) + verify-after-apply; PVE SetConfig hot-applies (Phase-0 PROVEN on the nested box: maxmem moves with the guest running, /proc/meminfo ripples via lxcfs, no reboot). Controller "Szerver memória (RAM)" card + code→Hungarian map, gated on FeatureGuestMemoryResize (MinAgent 0.90.0). LIVE-validated end-to-end through the real endpoint on the demo (above_max + below_min refusals render the Hungarian, agent English never leaks; SupportYes via the version header) Row complete as of the 2026-07-18 rehearsal — the refusals were proven on the nested demo, the applies on the N100. Cores stay observation
Publish train: MinAgent floors, gated auto-Reissue, version channels, floor-field-LAST rules hub v0.45/0.53, agent PARTIAL runbooks/publish-train-rules.md; demo-fleet updates proven Box-side floor lift PROVEN-LIVE on a fresh install (tests/VALIDATION-n100-rehearsal-2026-07-18.md): the day-0 golden deployed controller 0.143.0 at 16:32:28 and the managed floor lifted it to 0.145.0 by 16:32:34 — a 5-second, fully unattended update inside the first minute of controller life, update-state.json recording initiated_by: auto-floor with controller_updated pushed to the hub. So the mechanism is no longer nested-only. Still never proven on a real REMOTE customer — parked trains RUNBOOK-publish-0.79/0.81/0.85-* await Peti → R-1. Action before first invite: rebuild the golden to 0.145.x now that this evidence is banked, so fresh boxes don't sit two versions stale
Agent self-update: A/B slots, crash-loop auto-rollback, operator-signed agent v0.70+ PROVEN-LIVE (demo) SPIKE-agent-selfupdate-2026-07-05 Remote-customer proof pending → R-1
Controller self-update: anonymous registry, no credentials in guest controller v0.112 PROVEN-LIVE (demo) 07-10 arc
Offsite provisioning: Hetzner API, sub-account per customer, host-key pinning, credential re-issue hub v0.370.39 PROVEN-LIVE VALIDATION-offsite-provisioning-e2e-2026-07-09, SPIKE-hetzner-api-provisioning-2026-07-09
Per-customer offsite fill + staleness + freeze lever hub v0.41 IMPLEMENTED OffsiteChecker (hub/internal/monitor/offsite.go): fill 90/95% vs soft quota, staleness >48h No live-fired leg: CAMPAIGN-offsite-overnight-2026-07-10 recorded no quota/fill/staleness emails, and the freeze write-block was inconclusive (only the Hetzner readonly:true API op succeeded). Demoted
Box-level Storage Box aggregate (total fill, Σ quotas, oversubscription alert) hub v0.64.0 (R-5) IMPLEMENTED (data pipeline PROVEN-LIVE) monitor.OffsiteBoxChecker — fetch-throttled Hetzner GET (1/15 min), fill (used/storage_box_type.size, 80/90%) + oversubscription (Σ shared+enabled ConfigJSON quotas / capacity, 2.0×), escalation-only operator alert on the customer-less "pool-box" scope; Offsite-tab panel + dashboard tile. Phase-0-pinned live shape (box 611714) + live-computed in-cluster: 0.2% full (2.6 GB of 1.00 TB), Σ shared quota 150 GB, oversub 0.15x. Tests + 4 red-proofs; hub v0.64.0 REPORT Two open legs: the UI render is unit-verified only (hub UI password-gated → no screenshot); the alert emails are unit + red-proof verified but NOT fired live (real pool nominal — a live-fire emails Viktor). Thresholds pending Viktor's ruling (named config keys). READ-ONLY (GET)
Operator sees PBS DR datastore fill at a glance (Offsite "PBS DR" tab + dashboard gauge) hub v0.65.0 + tenantsync v1.2.0 (R-5) IMPLEMENTED (data pipeline PROVEN-LIVE) The PBS DR datastore (felhom-offsite on ep0) fill — NOT a Hetzner box. Option A: a read-only usage op on the felhom-tenantsync ep0 forced command (twin of fingerprint; df on the datastore path — no customer_id, no admin token, NO mutation), polled by monitor.PBSDRBoxChecker (OffsiteBoxChecker clone; 15-min throttle; states ok/unavailable/degraded; fill 80/90% on the "pbsdr-box" operator scope). /offsite split into Restic + PBS DR tabs; two dashboard gauges. Graceful: hub deploy ⟂ ep0 update (ep0 ≤ v1.1.0 → gauge "n/a" until updated). Phase-0-pinned (df on ep0 PBS 4.2.3) + live-computed in-cluster (ep0 updated to v1.2.0 this session): 19.1% full (7.1 GB of 37.2 GB). 10 Go tests + a bash harness + 3 red-proofs; hub v0.65.0 REPORT Open legs: UI render unit-verified only (hub UI password-gated); the fill alert email is unit + red-proof verified, NOT fired live (datastore nominal at 19%). Separate PBS threshold keys (default 80/90); no oversubscription (namespaces, not quotas). READ-ONLY
Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store hub v0.53, conventions IMPLEMENTED 07-13 closing bundle
Operator login password changeable from UI hub v0.54 IMPLEMENTED 07-13