Files
felhom.eu/documentation/architecture/00-capability-map.md
T
admin 54a4644721 docs: R-39 fleet fix SHIPPED (hub 0.68.0 + agent 0.91.2); R-50b(a) SHIPPED; (b)/(c) open
R-39's three legs are closed and deployed: the hub stamps a monotonic secret_generation
so a re-key finally moves the descriptor hash; the wrapper gains a narrow read verb so
the non-root agent can read the credential it writes; and ProbeAuth turns a 401 into a
loud auth_failed the existing damper escalates to a fresh mint. Plus a consumed_at
honesty gauge for the applied-but-never-consumed disagreement.

Recorded in the R-39 row, because both are the kind of thing a future reader needs:

- A load-bearing fact the spec did not flag, checked rather than trusted: Apply bails out
  if the storage status probe ERRORS and adopt converges without consuming when the
  storage reads active, so the fix depended on PVE's 401 behaviour. PVE's storage_info
  wraps activation in eval{} and leaves active=0, so a 401 returns HTTP 200 with
  active:0 — never an API error. The chain is sound by proof, not inference.

- A defect I shipped and caught: v0.91.0 built the probe seam and main.go never wired it,
  so the leg was inert while every test passed. Same class as controller v0.154.0 the day
  before. Fixed in v0.91.1 (artifact superseded, not overwritten); v0.91.2 made a healthy
  probe observable so "no auth_failed" can never again be confused with "never probed".

The DR-tier capability row is deliberately NOT upgraded to PROVEN-LIVE: the decisive
evidence is STOP-2, the operator pressing Re-issue and the box converging where the
identical click did nothing on 2026-07-18.

R-50b(a) shipped — wrapper sha256 in the manifest + agent reporting + host drift surface,
with unknown-on-either-side reading as quiet rather than drift. (b)/(c) remain open: the
wrapper is still fetched unversioned from raw/branch/main.
2026-07-21 10:24:11 +02:00

47 KiB
Raw Blame History

00 — Felhom Capability Map

What this is: the single cross-component truth table of what the Felhom platform can do today, at what confidence level, with verifiable evidence. Rows are scenarios (user- or operator-visible outcomes), not modules — a scenario spans agent + controller + hub + catalog, and this is the only doc that shows that view.

What this is NOT: a roadmap. Planned work lives in documentation/backlog/ROADMAP.md and is referenced from gap rows by ID (→ R-n). A row here never claims future behavior.

Status enum (strict):

Status Meaning
PROVEN-LIVE Exercised end-to-end on real infrastructure; MUST cite a campaign/drill/validation doc in documentation/audits/ or documentation/tests/. No citation → not PROVEN-LIVE.
IMPLEMENTED Shipped + unit/red-proof tested, but the real flow has not been exercised live (or not on the surface that matters — noted per row).
PARTIAL Some legs live, some missing/unvalidated — the note says which.
MISSING Does not exist. Present-tense fact; if planned, the row points at a roadmap ID.

Update rule (end-of-session checklist item): if a task changed any capability's status, update the row in the same session — with the new evidence citation. A PROVEN-LIVE claim is subject to the cardinal rule like any other claim.

Verified 2026-07-16 against evidence corpus @ felhom.eu tip 4b18cc5 by CC (capability-map audit); see REPORT.md for the per-row verdict table.


A. Provisioning & day-0

Scenario Components Status Evidence Gap / roadmap
Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box installer, agent, hub, golden PROVEN-LIVE (nested VM) DRILL-day0-vm-2026-07-12, DRILL-day0-take2-2026-07-12 First firing on real customer hardware pending → R-1
BYO install: --mode byo, mandatory caps, host-mutation disclosure, coexistence guards installer v1.15+, agent PARTIAL DRILL-GL6-2026-07-08 (demo box); GL-8 coexistence fixes Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1
Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot host-install); selectable UEFI loader; universal secret-free / operator-bind mode scripts v1.19.0 (scripts/iso/) + hub v0.62.0 + assistant container PROVEN-LIVE (physical N100, one pass, 2026-07-18) tests/VALIDATION-n100-rehearsal-2026-07-18.md — the full chain on real metal in a single pass: the generic reusable pairing ISO (v1.20.0, --loader mkimage, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted (appliance_registrations id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. F1 is closed on physical hardware. Prior nested legs: slice A SPIKE-baremetal-iso-2026-07-16 (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; --loader mkimage boots+installs SB-off; mkimage SB-enforcing FAILS Access Denied; surgery byte-identical); slice C (2026-07-17): the GENERIC secret-free ISO — box self-registers as an unclaimed appliance (POST /api/v1/appliance/register, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) F1 loader caveat: --loader mkimage fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → Secure Boot must be OFF; default shim keeps SB. Slice C bind is operator-password-gated (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17) — see the dedicated self-bind row
Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) controller v0.122, hub v0.50 PROVEN-LIVE (drill VM) DRILL-day0-vm-2026-07-12 §10/F-4 (gate ON via real edge; claimed, code consumed) Never executed by a non-Viktor human → R-3. Deliverability (R-4), gmail half DONE 2026-07-18: the rehearsal's claim email was the first sent under the tightened DMARC p=quarantine and landed in the gmail Inbox, not spam (tests/VALIDATION-n100-rehearsal-2026-07-18.md). freemail.hu remains Viktor's open half. (Dropped mis-cited CAMPAIGN-4 F-C — that is the escrow-claim 502, not password claim)
Customer binds their own appliance (self-service): operator-sent 7-day tokenized capability link → public two-factor /bind/<token> (console pairing code + retrieval passphrase) → hub stages the bind, no operator hub v0.66.0 + ISO scripts v1.20.0 PROVEN-LIVE (real customer-zero bind on metal, 2026-07-18) tests/VALIDATION-n100-rehearsal-2026-07-18.md: operator minted + emailed the link 16:28:55 (7-day TTL, expiry 2026-07-25 recorded); the customer bound their own box at 16:29:55 with attempts=0, locked=0appliance_bound carries source customer_selfbind, and the credential was delivered 26 s later with no operator action. Hub-side lifecycle in hub-state.txt (selfbind_tokens mint→email→consume). Prior unit evidence: hub v0.66.0 (web/selfbind.go, store/selfbind.go; Scenarios AF + F1/F2; 4 red-proofs verified red — THE TRAP /bind/ exemption, no-oracle, lockout, single-active); GC verdict §3 (no appliance GC → TTL stands alone) R-27 slice 1. No appliance list ever rendered; wrong code == wrong passphrase (one generic failure); 5-attempt lockout → call support; expiry falls back to operator-bind. Live first-run DONE 2026-07-18 (rehearsal; the console banner rendered on the real ISO). R-27b (controller second-box dismissable prompt) deferred; multi-box-per-link = repeated operator sends
Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge controller v0.127, agent v0.88/0.89 PROVEN-LIVE (drill VM, endpoint-exact) agent v0.88.0 REPORT (ceremony ~4s, one-shot claim 200→410, R absent from every payload); SPIKE-controller-escrow-2026-07-13 Customer-facing browser wizard FIRST LIVE FIRING 2026-07-18 (tests/VALIDATION-n100-rehearsal-2026-07-18.md, S6): customer zero drove the wizard on the reborn box — ceremony started 16:56:29, recovery code claimed one-shot 16:56:39 (absent from logs by design), hub-verified and EscrowState auto-confirmed 16:56:41, offsite runs enabled 12 s after the ceremony began; the v0.138.0 „megerősítésre vár, legfeljebb 15 perc" awaiting card rendered and flipped on the ACK (operator screenshots: Viktor's set). Honest caveat: at a 12-second confirm the awaiting window is so short that catching both states on screen is luck, not procedure. Prior: endpoints driven on the drill VM. agent v0.89.0: /escrow/preflight pbs_storage_id row now live-reloads (reads current agent.json) — a pbsdr convergence that seeds the id flips it green with NO service restart. hub v0.60.0 (data-first retention): a re-escrow with a DIFFERENT sealed passphrase no longer destroys the old blob — the hub RETAINS it (host_escrow_superseded), so a previous passphrase stays recoverable with its recovery code (turns the reinstall-orphan incident from "history destroyed" into "history recoverable"). Guided-recovery flow = R-26. Red-proof TestSaveHostEscrow_RetainsSuperseded. hub v0.60.1 — custody survives the host lifecycle: host deletion (with the escrow ack) DEMOTES the current blob to retained custody (moved into host_escrow_superseded, never destroyed; existing superseded rows spared); the customer Danger-zone Delete is the one true purge point (cascades both escrow tables incl. already-deleted hosts). No operator path through host lifecycle can lose a blob. Red-proofs TestDeleteHost_DemotesEscrowNeverDestroys + TestDeleteCustomer_PurgesEscrowCustody
DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation installer v1.15, agent v0.86, hub v0.51 IMPLEMENTED DRILL-day0-take2-2026-07-12 §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) ⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39). On the reborn N100 the descriptor auto-provisioned and the agent reported converged state=applied (16:45:53), yet the storage is dead: pvesm statusfelhom-pbs: error fetching datastores - 401 Unauthorized / inactive, and a direct probe with the stored credential returns 401 on every endpoint including /version while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and consumed_at is still NULL; the converged state machine will not re-apply, and the agent's 15-minute verify loop cannot even read the credential to notice (open /etc/pve/priv/storage/felhom-pbs.pw: permission denied — non-root agent reading a file it writes through a root wrapper). A tier that reports applied while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See tests/VALIDATION-n100-rehearsal-2026-07-18.md F2 and pbs-dr-state.txt. agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17: the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick pbsdr: pre-check 403 … self-granting … (R-22)converged state=adopted in ~3 s, ACLs self-restored, pvesm status felhom-offsite=active, zero operator action. No more one-shot pveum grant 2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end. The three defects that let a box be applied and dead simultaneously are each addressed: the hub stamps a monotonic secret_generation into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow read verb so the non-root agent can read the credential it writes (it never could — /etc/pve/priv is 0700 root:www-data, which made the verify loop blind by construction); and pbs.ProbeAuth turns a 401 into a loud auth_failed that the existing pbsdrheal damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says applied, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (rc=0) and probed successfully (credential probe OK storage=felhom-pbs). This row is NOT upgraded to PROVEN-LIVE yet — the decisive evidence is STOP-2, the operator pressing Re-issue and the box converging where the identical click did nothing on 2026-07-18. Until that runs, the capability is SHIPPED but the destroy-then-recover proof this row asks for is not in hand. Evidence: felhom-agent/REPORT.md (2026-07-21).
Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive hub v0.61.0, felhom-tenantsync v1.1.0 PROVEN-LIVE (external teardown, incl. two real firings) tests/VALIDATION-n100-rehearsal-2026-07-18.md — two live firings, both host-delete-first, on two different customers (demo-vm-felhom 15:49:57, demo-felhom 16:08:51): every leg ok (claim, db_purge, descriptor, hetzner, pbs), escrow acked separately, each completing in 89 s (hub-state.txt customer_resets). The Hetzner sub-account destruction is now verified against the live pool box — and produced the run's sharpest lesson: a sub-account is an access-control object, not a data object. Deleting it left its /home intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (a finding by S7's own criterion) and why RESET now needs a base-dir purge → R-32. Prior: hub v0.61.0 REPORT; ep0 live drill 2026-07-17 (throwaway drill-reset-01 with a real backup: deprovision deleted:true destroyed the namespace + backup group + token, idempotent re-run deleted:false, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. Live-clicked 2026-07-18 (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. Consistency gap → R-25b: the Danger-zone DELETE leaves host rows and doesn't run this teardown
Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling installer PARTIAL DRILL-GL6-2026-07-08 Phase 1/5 (KEPT-vs-WIPED printed verbatim; drive data intact ×3); GL-4 code Secret purge (GL6-F1 .bak residue) fixed v1.12.0; enrolled-drive mnt-*.mount units survive (GL6-F2, open); cluster-aware felhom_guests guard + saferemove cost warning missing → R-9

B. Apps & catalog

Scenario Components Status Evidence Gap / roadmap
Deploy an app from the catalog (env config, memory guard, health-aware progress) controller, catalog (~52 apps, images pinned) PROVEN-LIVE CAMPAIGN-2 T-DEPLOY-SET (7 apps, env config, health-aware); RERUN-p1p3 (×4 PASS) Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only
App lifecycle: start/stop/restart/update/logs/remove/redeploy controller PROVEN-LIVE CAMPAIGN-2 T-LIFECYCLE (stop/start/restart/update/logs); remove live in CAMPAIGN-3 Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical
Protected infra stacks can't be stopped/removed from UI controller PROVEN-LIVE CAMPAIGN-nomercy + RERUN-p1p3 T-SEC-PROTECTED (refuse stop/remove, stay Up) (Cited CAMPAIGN-2 T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal)
Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad backup: block degrades to legacy, loudly) controller v0.132, catalog PROVEN-LIVE CAMPAIGN-2 T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs
App crashes → customer notified (one event per transition, no flapping spam) controller v0.120, hub v0.48 IMPLEMENTED controller v0.120.0 (dead-app alerting, app_start_failed, one-event-per-transition red-proofs); CAMPAIGN-3 F11 surfaced the gap End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven
Post-deploy optional config (API keys etc.) with restart controller, catalog .felhom.yml IMPLEMENTED feature long-standing; config page renders (CAMPAIGN-2 T-PAGE-ALL is GET-only) The config-save+restart flow is exercised in no campaign (CAMPAIGN-3 explicitly skipped interactive app config). Demoted: T-PAGE-ALL is a page-render smoke test, not this flow
Backup classification: 13 bind-bearing apps carry mandatory/optional/excluded classes catalog, controller v0.132133 PROVEN-LIVE SPIKE-backup-classification-2026-07-14, CAMPAIGN-6D/6E Remaining ~39 apps are legacy-class by design (unit-only offsite)

C. Protection & recovery (the product promise)

Scenario Components Status Evidence Gap / roadmap
Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes controller v0.118 PROVEN-LIVE CAMPAIGN-2 T-BAK-FULL (pg+mariadb autodiscovered); atomicity CAMPAIGN-6B P4 + CAMPAIGN-6E B1/B2 (SIGKILL mid-write → only .tar.tmp touched, last-good byte-unchanged); DB restore CAMPAIGN-6D P-FAB (Cited CAMPAIGN-3 F7 is the finding of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10
Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary controller v0.135 PROVEN-LIVE CAMPAIGN-6E-2026-07-15 (P-TIER2 deep-4 PASS), CAMPAIGN-6C
Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping controller v0.134, agent, hub PROVEN-LIVE CAMPAIGN-6D-2026-07-15 (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); VALIDATION-offbox-storagebox-2026-07-09 (byte-perfect round-trip) Raw-data quota (SP-1) + retention regrouping (SP-2) are SPIKE-restic-snapshot-shape dry-run verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. Reinstall-continuity (controller v0.142.0, 2026-07-17): a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (wrong password or no key found) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. The live leg FIRED on its own during the 2026-07-18 rehearsal (tests/VALIDATION-n100-rehearsal-2026-07-18.md, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard classified it, pushed offbox_repo_orphaned, skipped the run and showed the card (16:58:14) rather than nightly-spamming a raw restic error; the operator-confirmed reset then moved the repo aside (never deleted) to .orphaned-20260718 and re-initialised (16:59:26→16:59:32), and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — the finding is that it had to fire at all (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). DIAGNOSE-offbox-repo-orphaned-2026-07-17
Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live controller v0.134/134.1/135 PROVEN-LIVE (2026-07-20) CAMPAIGN-6D accept legs (immich end-to-end from offsite alone) 2026-07-19: audits/DIAG-immich-restore-2026-07-19.md finds no offsite path loads a DB dump — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "immich end-to-end from offsite alone" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume (immich_postgres_data is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" overclaimed scope: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → PARTIAL, scope-corrected. Evidence: audits/DIAG-immich-restore-2026-07-19.md (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). 2026-07-19, controller v0.148.0: the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but round 2 found it aborts against a running app (audits/DIAG-immich-restore-round2-2026-07-19.md, H4: the replay races immich's own schema repair; clip_index recreated by the app 2 s before the dump's CREATE INDEX). 2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47) — both restore paths now replay into a DB-ONLY window (StartStackServices brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. (The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard BackupStatus fix. R-47 shipped in v0.153.0.) 2026-07-20: the clean run HAPPENED — endpoint-level supervised reconstitute of immich from snapshot 49e7cb46 (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no already exists, operation reported SUCCESS, immich's own DatabaseService logged No schema drift detected twice, 11 assets active, 4/4 containers healthy, 231 public indexes. Operator confirmed the immich timeline renders correctly after the reconstitute (screenshot held, 2026-07-20). Evidence: felhom-controller/REPORT.md §4b. 2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE. The operator deleted the photos in immich own UI and emptied the trash (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. 40 file(s) placed against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets active, No schema drift detected, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: felhom-controller/REPORT.md 4e
Manual .fab export/import: class-scoped capture, browser up/download, tunnel-proof chunking controller v0.125/128/130/136 PROVEN-LIVE CAMPAIGN-6D P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking CAMPAIGN-6B P2 (100 MiB via real CF edge, 120 MiB→413) Chunking proven at the real CF edge via curl --resolve; the rendered browser file-picker upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B finding; fix verified in 6D
Guest-loss DR: PBS restore with full-fidelity layout from archive, restore-test verification agent v0.75/0.76, PBS PROVEN-LIVE CAMPAIGN-2 T-P9-DESTROY-RESTORE (whole-guest pct restore of 9201 → running+healthy) + T-PBS-VERIFY (verify_state: ok, 13 snapshots); DRILL-GL6-2026-07-08 Phase 0d (restore-test mount_parity: ok) (Cited VALIDATION-newbox-restore is offbox restic file-restore, wrong tier — corrected.) Real offsite guest-loss round-trip still R1-blocked → S5 DR drill
PBS-DR secret self-heal on reused-peer re-provision hub v0.56 IMPLEMENTED hub v0.56.0 (pbsdrheal/reconciler.go, RestageHostPBSSecret, all §10 red-proofs); SPIKE-pbsdr-selfheal-2026-07-15 (root cause) Reconciler is scoped to one host (PBSDRHEAL_ONLY_HOST), not fleet-wide; already fired live hands-free on drill qm300 (07-15) — real-customer firing + fleet-wide widening pending
Crash/power-loss mid-backup/mid-migration → self-heal on next run controller, agent PROVEN-LIVE CAMPAIGN-6D P5-REST (SIGKILL mid-offbox → auto-restart ~15s, run marked failed not false-success, no stale lock); CAMPAIGN-6E B1-B3 (Cited CAMPAIGN-2 T-RBT-* legs were empty / auth-hollow — corrected.) Live mid-migration crash→self-heal is the weakest sub-claim (P5-REST is mid-backup)
Box survives a site/network change (relocation, different subnet, DHCP re-lease) with the control plane intact agent, controller, bootstrap PARTIAL audits/AUDIT-vacation-remote-ops-2026-07-20.md — a real relocation of the demo box: guest + hub telemetry + WG/PBS + Cloudflare tunnel all survived untouched, but the controller↔agent control plane did not (agent binds a LAN literal → bind: cannot assign requested address → storage/PBS-backup/quiesce/restore-test/DR down until fixed). Mitigated for the window by pinning vmbr0 static R-50 (island-bridge control plane, spike-first) is the durable fix. Related: R-51 (dead-primary alerting) and R-52 (boot desired-state reconciliation) — the same event left two apps Exited with no alarm and no recovery
Soft-quota: usage bar, pre-push enlargement block, customer notification controller v0.109/134, hub v0.41/55 PROVEN-LIVE 6D/6E; hub OffsiteChecker
A customer (not the operator) performs a restore via UI alone all MISSING (as evidence) Alpha will produce this; script it into R-3. 2026-07-19: the C6 evidence attempt ran and found a product gap instead of evidenceaudits/DIAG-immich-restore-2026-07-19.md. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING by finding, not by absence of attempt — the rehearsal system working, not failing. 2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0 (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. 2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped (audits/DIAG-immich-restore-round2-2026-07-19.md): a deliberate run from snapshot 49e7cb46 did recover all 11 assets (status=active, files resolve), but the operation reported failure and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. 2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED. (The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.) The C6 drill ran end-to-end through the UI: photos deleted, trash emptied, the full files+database restore pressed on /backups/restore, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: felhom-controller/REPORT.md 4e. Residual: the run was performed by the OPERATOR, not by a customer — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means trash, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash and verify the app's content, not the file count

D. Storage & devices

Scenario Components Status Evidence Gap / roadmap
Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts controller, agent v0.87 PROVEN-LIVE DISPOSITION-ia-finding2-systemdisks-2026-07-13 (legacy EFI+LVM host, root not offered, byte-identical); enroll/format live in storage-lifecycle-acceptance-2026-06-15 (E10 re-enroll, data intact); agent fence self-test refuses /dev/sda (Cited CAMPAIGN-2 T-STG-ENROLL/SEC-FORMAT were auth-hollow CSRF-403.) Fresh-USB wizard enroll+format through the customer UI PROVEN-LIVE (controller v0.141.0, 2026-07-17): a 64 GB scratch USB driven through the real /api/storage/init endpoints (login+CSRF) → confirm → detached format (~27 s mkfs) → mount → register → mounted+registered at /mnt/felhom-drives/scratch1. F6 (initialize-to-usable) now covered: the wizard runs the chain as a detached, disconnect-safe, pollable job (3-step progress) with an agent format-status poll for a slow mkfs
Data migration between drives (all / per-app), crash-safe controller PROVEN-LIVE CAMPAIGN-6C 4P-5 (scope=app round-trip, byte-identical); storage-lifecycle-acceptance-2026-06-15 (two migrate-all runs via dashboard UI, sha256 byte-identical) (Cited CAMPAIGN-2 T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS
NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated controller v0.113117, agent v0.81/84/85 PROVEN-LIVE SPIKE-nas-verify-2026-07-11, SPIKE-nas-dsm-2026-07-11, CAMPAIGN-3-2026-07-11 (boot/reassert fixes)
USB drive enrollment + unplug detection + recommission controller, agent PROVEN-LIVE storage-lifecycle-acceptance-2026-06-15 E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); CAMPAIGN-4/6A (3 USB re-establish across device-letter reshuffle) (Cited RUNBOOK-usb could NOT complete a wizard enrollment; CAMPAIGN-2 legs were auth-hollow.) Fresh-USB wizard enrollment specifically still unproven
Decommission (migrate-first and anyway-paths), eject agent, controller PROVEN-LIVE storage-lifecycle-acceptance-2026-06-15 E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) (Cited CAMPAIGN-2 T-STG-DECOM-* were auth-hollow; SPIKE-decommission was report-only, button still vestigial.)
Boot ordering: automount + networking survive reboot; appliance self-heal watchdog agent v0.85 PROVEN-LIVE CAMPAIGN-4-2026-07-13 (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + CAMPAIGN-6A-2026-07-14 1D (re-arm reboot-survival across 9 guest + 1 host reboots) (CAMPAIGN-3 F10/F11/F12 were the CRITICAL/HIGH failures; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: skip-active on pct reboot carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A)

E. Access, networking & household use

Scenario Components Status Evidence Gap / roadmap
Remote access via Cloudflare Tunnel + Traefik (per-app subdomains) cloudflared, traefik PROVEN-LIVE CAMPAIGN-2 T-FLT-CF Per-customer zone-scoped CF tokens (blast-radius ruling)
LAN access when internet is down (lan_resolver) agent IMPLEMENTED Never drilled as a customer experience ("net down — can I reach my photos?") → R-19
Phone photo backup immich (classified) PROVEN-LIVE 6D end-to-end restore proof
Documents/OCR paperless-ngx (classified) PROVEN-LIVE CAMPAIGN-6C 4P-1 (deploy paperless-ngx, ingest 3 docs via consume flow, OCR + PDF/A ~90s) + 4P-2/3/5 Consume-folder ingestion awkward without SMB → R-7
Files from Windows Explorer / Mac Finder (SMB server) controller v0.145.0 + felhom-samba:1.0.0 PROVEN-LIVE felhom-controller REPORT.md (v0.144.0) + controller/sharing.md; transport verdict audits/SPIKE-lan-discovery-2026-07-18.md „Megosztás" page: enable + one household password + shares (new folder or picked existing, per-share read-only). Fourth protected infra stack (host-net, smbd+nmbd+wsdd). Live on demo: 445 reachable, NetBIOS FELHOM resolves, write/read byte-compare PASS, write to a read-only share REFUSED, SMB writes land as uid 1000. Explorer leg PASSED 2026-07-18 (Viktor): Network → FELHOM → both shares open; a real Explorer save into dokumentumok landed owned uid 1000, and a write into the read-only filmek was refused by Windows with the folder left untouched. Share data RIDES BOTH BACKUP TIERS (R-7b, controller v0.145.0, Model B sibling shares source): tier-2 cross-drive legs + an offsite _shares restic snapshot carrying the share definitions and the credential copy, with a „Megosztások" restore. All four legs PROVEN-LIVE on demo 2026-07-18 — tier-2 tree md5-verified; offsite snapshots e0b9d723 (Viktor 12:18:16Z) and 4e2b15ec both carrying manifest + passdb.tar; restore round-trip returned a deleted probe file byte-identical and a deleted share DEFINITION with its original flags without overwriting live files; samba liveness → hub-accepted health_critical. Remaining human leg: SMB positive auth with the real household password
Media to TV via DLNA MISSING Jellyfin app exists; DLNA/SSDP unvalidated → R-6, R-8
File access via browser FileBrowser (infra app, auto-mount sync) IMPLEMENTED FileBrowser runs healthy + userdata-bound (storage-lifecycle-acceptance-2026-06-15, CAMPAIGN-3) Actual browse/download through FileBrowser is exercised in no doc. (Cited CAMPAIGN-2 T-PAGE-ALL renders only the controller dashboard pages, not FileBrowser.) Demoted
Forgot dashboard password → instant reset code controller v0.123, hub PROVEN-LIVE DRILL-day0-take2-2026-07-12 F-15 (live re-run of the exact failure path: hash applied 1s after request, code accepted first try)
Multiple household users / per-person accounts MISSING Single dashboard password; acceptable for alpha → R-15
WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) agent v0.72, hub v0.35 IMPLEMENTED SPIKE-oob-wg-operator-peer-2026-07-05, SPIKE-felhom-sshd-2026-07-05 Mutual-repair desired-state arc not built → R-13
Break-glass management-plane recovery agent v0.71, hub v0.34 IMPLEMENTED runbooks/break-glass.md

F. Notifications & monitoring

Scenario Components Status Evidence Gap / roadmap
Health-degradation email (edge-triggered, cooldowns, Hungarian) via hub → Resend controller, hub IMPLEMENTED delivery pipeline live-proven for the enlarge-block trigger (CAMPAIGN-6D P3-DELIVERY, op+customer "Kedves Ügyfél!"); NotifyHealthChange ok→warn/fail edge-trigger implemented The health-degradation trigger specifically has never fired an email live in any doc. Demoted (pipeline proven for a different event). Deliverability to HU freemail → R-4
Event catalog: app_start_failed, dead-app, offbox_enlarge_blocked, claim/reset codes, critical severity controller, hub v0.31/48/50/55 PROVEN-LIVE live-delivered: CAMPAIGN-6D P3-DELIVERY (enlarge-block, op+customer); DRILL-day0-vm F-4 (claim code); DRILL-day0-take2 F-15 (reset code) app_start_failed/dead-app delivery is unit-only (6C inconclusive) — the pipeline + 3 event families are live, those two are not
Prefs safety: empty-email wipe guard controller v0.137 IMPLEMENTED red-proofed 07-15 Born from a live incident; guard itself unit-proven
System + container metrics (SQLite, Chart.js, 30-day downsampling) controller IMPLEMENTED metrics collection + /monitoring render present (page 200) The cited CAMPAIGN-2 T-RES-CPU/T-SOAK-LOOP are H1/H2 harness artifacts (auth-302), not metrics tests; SQLite/Chart.js/30-day downsampling validated in no campaign. Demoted
Always-on debug rings + on-demand log-bundle pulls with TTL/custody controller v0.116, agent v0.83, hub v0.46 PROVEN-LIVE debug rings live-exercised CAMPAIGN-3 fix-6 (1000-cap ring, ~55min horizon under load) The log-bundle-pull TTL/custody half is changelog-only (no dedicated observability audit doc); ring persistence across restart is a known gap
Operator alerting (Healthchecks → monitoring@felhom.eu) k3s, Resend IMPLEMENTED operator infra, stated in production since 02-04; no corpus validation doc Per the status enum, no citation → not PROVEN-LIVE. Demoted pending an operator-cited live alert (candidate re-upgrade — see REPORT)

G. Fleet & operator (hub)

Scenario Components Status Evidence Gap / roadmap
Customer/host management: 8-tab detail, scoped auto-refresh, safe stale-host deletion, capability chips hub v0.470.53 PROVEN-LIVE hub v0.53.0 dead-host roll-up live on the Peti cluster (proxmox1 down 23h); CAMPAIGN-4-2026-07-13 (operator UI driven live); DRILL-day0-take2 F-16 (offsite/freeze buttons live) 8-tab render + capability chips are render-test-validated (hub UI is password-gated; CC cannot log in). (Cited "daily operator use" was a no-doc citation; AUDIT-hub-gui-2026-06-30 predates these features at hub v0.25)
Config/state change round-trips in seconds (hub↔box immediacy; 15-min cycle stays the backbone): box→hub out-of-cycle report (Dir 1) + hub→box GET /api/v1/wait long-poll wake (Dir 2) controller v0.139/140, hub v0.58/0.63 PROVEN-LIVE (2026-07-21) Transport proven live through the real DNS-only ingress: SPIKE-immediate-sync-transport-2026-07-16 + hub v0.58.0 / controller v0.140.0 REPORTs — 240 s no-annotation hold (25 s heartbeat defeats nginx's 60 s proxy_read_timeout, no ingress change), 0.047 s wake-on-change, hub rollout restart = 1 WARN + 0-storm reconnect; Dir-1 2 s box→hub round-trip live in controller v0.139.0 The operator-UI save→apply round-trip is not fired end-to-end live (hub UI password-gated; CC can't log in) → R-23; the wake transport and the ACK→config_version→ConfigRefresher delivery chain are each proven, only the UI-triggered bump leg is unexercised. Agent-plane (host-domain desired-state) poke first slice PROVEN-LIVE (Direction-2a, agent v0.89.0 + hub v0.59.0, 2026-07-17): contentless ep0-relayed UDP poke → agent immediate desired-state cycle, per SPIKE-immediate-sync-transport-2026-07-16 P4. Full path live-proven: a real operator manifest save fired poke: sync-poke delivered to 10.77.0.2; the box (0.89.0) received it and logged poke received → triggering an immediate desired-state cycleout-of-band report triggered~31 ms ep0→box, sub-ms to the report cycle (WG-confined, from 10.77.0.1 to the 10.77.0.2-bound socket); save→tick ≈ ~0.45 s (SSH-dominated), well under ≤23 s. R-13 first slice (listener+sender only; the rest of the mutual-repair arc stays open). System-initiated immediacy wired (hub v0.63.0, this REPORT): the mutation sites that only OPERATOR actions used to notify now fire the correct plane's notifier when the hub itself mints state — agent-plane pokes at PBSDRAutoProvision (the observed slice-C lag), ReissuePBSDR (also the pbsdrheal escalation), handlePBSDRReissue, and the two admin desired-state api writers; controller-plane bump at reissueOnReenroll. Unit-tested + red-proofed, not yet fired on a real system event (folds into the rehearsal bind sequence). Still PARTIAL: the R-23 operator-UI save→apply leg and the agent fast-tick-until-first-convergence SECONDARY (the WG-registration leg a poke can't reach pre-tunnel) remain unfired live. Fast-tick SHIPPED (agent v0.90.0, R-28): while any desired-state item is unapplied — incl. the pre-tunnel window a poke can't reach — the agent pulses the out-of-band trigger every 30 s and self-disarms on convergence (state-based; four cached sources; LOUD states excluded). LIVE on both demo agents (the fast-tick armed: 30s … startup line verified). REAL-ONBOARDING PROOF DONE — tests/VALIDATION-n100-rehearsal-2026-07-18.md (ledger 8, S5): on a genuine first onboarding on metal, every post-bind leg landed seconds apart with no ~15-minute stall anywhere — bind 16:29:55 → credential delivered 16:30:21 (26 s) → agent 0.90.0 up 16:30:49 → WG registered + tunnel applied 16:30:51 (~2 s) → poke listener 16:30:54 → controller 16:32:28 → floor-lifted and running current 16:32:39. Bind → running-current = 2 min 44 s. The PBS-DR descriptor auto-provisioned on the same cadence (agent converged state=applied 16:45:53) — though see the DR-tier row: the descriptor converged while the credential behind it was already stale (R-39). The pre-tunnel fast-tick window is therefore proven in its real setting; the remaining PARTIAL is the R-23 operator-UI save→apply leg alone 2026-07-21 — THE LAST PARTIAL LEG IS CLOSED (R-23(a) restart leg). The operator saved the global floor to a version the box did NOT run (0.153.0 → v0.154.0) and the managed self-update fired exactly once: 06:57:13Z UpdateState pending (initiated_by=auto-floor) → 06:57:17Z agent controller swap requested06:57:21Z container restarted → 06:57:29Z new controller healthy. Save → healthy on the new version = 16 s. Over a 39-minute window: swap requests 1, agent-driven bootstrap restarts 1, rollbacks 0, container RestartCount 0; VerifyStartup confirmed on the next boot and the following periodic check logged Current version 0.154.0 is up to date (the at/above-floor branch doing nothing, as designed). The 2026-07-20 attempt could not prove this because it targeted an already-running version. Evidence: felhom-controller/REPORT.md §6.
Customer right-sizes guest RAM from the controller (agent-enforced bounds, live cgroup apply, no reboot) agent v0.90.0 + controller v0.143.0 (R-24) PROVEN-LIVE (grow and shrink on metal, 2026-07-18) tests/VALIDATION-n100-rehearsal-2026-07-18.md (ledger 9) — the apply is now proven in both directions on a normal-sized box: customer zero shrank 11675 → 8192 MB at 16:50:22 and grew 8192 → 12288 MB at 17:02:17, each a live cgroup apply with no reboot (local-api: guest-memory resized in the agent journal, [web] memory resized in the controller log), and the new total rippled into the deploy page's memory math at 17:05:15 (total=12288MB). F5 auto-sizing had landed the guest at 11675 MB. Controller-direct (R-24's hub-desired-state framing SUPERSEDED, Viktor 2026-07-17). Agent GET/POST /guest/memory enforces every bound FRESH (min 2048 / max host_total2048 / shrink floor max(2048, usage+512)) + verify-after-apply; PVE SetConfig hot-applies (Phase-0 PROVEN on the nested box: maxmem moves with the guest running, /proc/meminfo ripples via lxcfs, no reboot). Controller "Szerver memória (RAM)" card + code→Hungarian map, gated on FeatureGuestMemoryResize (MinAgent 0.90.0). LIVE-validated end-to-end through the real endpoint on the demo (above_max + below_min refusals render the Hungarian, agent English never leaks; SupportYes via the version header) Row complete as of the 2026-07-18 rehearsal — the refusals were proven on the nested demo, the applies on the N100. Cores stay observation
Publish train: MinAgent floors, gated auto-Reissue, version channels, floor-field-LAST rules hub v0.45/0.53, agent PARTIAL runbooks/publish-train-rules.md; demo-fleet updates proven Box-side floor lift PROVEN-LIVE on a fresh install (tests/VALIDATION-n100-rehearsal-2026-07-18.md): the day-0 golden deployed controller 0.143.0 at 16:32:28 and the managed floor lifted it to 0.145.0 by 16:32:34 — a 5-second, fully unattended update inside the first minute of controller life, update-state.json recording initiated_by: auto-floor with controller_updated pushed to the hub. So the mechanism is no longer nested-only. Still never proven on a real REMOTE customer — parked trains RUNBOOK-publish-0.79/0.81/0.85-* await Peti → R-1. Action before first invite: rebuild the golden to 0.145.x now that this evidence is banked, so fresh boxes don't sit two versions stale
Agent self-update: A/B slots, crash-loop auto-rollback, operator-signed agent v0.70+ PROVEN-LIVE (demo) SPIKE-agent-selfupdate-2026-07-05 Remote-customer proof pending → R-1
Controller self-update: anonymous registry, no credentials in guest controller v0.112 PROVEN-LIVE (demo) 07-10 arc
Offsite provisioning: Hetzner API, sub-account per customer, host-key pinning, credential re-issue hub v0.370.39 PROVEN-LIVE VALIDATION-offsite-provisioning-e2e-2026-07-09, SPIKE-hetzner-api-provisioning-2026-07-09
Per-customer offsite fill + staleness + freeze lever hub v0.41 IMPLEMENTED OffsiteChecker (hub/internal/monitor/offsite.go): fill 90/95% vs soft quota, staleness >48h No live-fired leg: CAMPAIGN-offsite-overnight-2026-07-10 recorded no quota/fill/staleness emails, and the freeze write-block was inconclusive (only the Hetzner readonly:true API op succeeded). Demoted
Box-level Storage Box aggregate (total fill, Σ quotas, oversubscription alert) hub v0.64.0 (R-5) IMPLEMENTED (data pipeline PROVEN-LIVE) monitor.OffsiteBoxChecker — fetch-throttled Hetzner GET (1/15 min), fill (used/storage_box_type.size, 80/90%) + oversubscription (Σ shared+enabled ConfigJSON quotas / capacity, 2.0×), escalation-only operator alert on the customer-less "pool-box" scope; Offsite-tab panel + dashboard tile. Phase-0-pinned live shape (box 611714) + live-computed in-cluster: 0.2% full (2.6 GB of 1.00 TB), Σ shared quota 150 GB, oversub 0.15x. Tests + 4 red-proofs; hub v0.64.0 REPORT Two open legs: the UI render is unit-verified only (hub UI password-gated → no screenshot); the alert emails are unit + red-proof verified but NOT fired live (real pool nominal — a live-fire emails Viktor). Thresholds pending Viktor's ruling (named config keys). READ-ONLY (GET)
Operator sees PBS DR datastore fill at a glance (Offsite "PBS DR" tab + dashboard gauge) hub v0.65.0 + tenantsync v1.2.0 (R-5) IMPLEMENTED (data pipeline PROVEN-LIVE) The PBS DR datastore (felhom-offsite on ep0) fill — NOT a Hetzner box. Option A: a read-only usage op on the felhom-tenantsync ep0 forced command (twin of fingerprint; df on the datastore path — no customer_id, no admin token, NO mutation), polled by monitor.PBSDRBoxChecker (OffsiteBoxChecker clone; 15-min throttle; states ok/unavailable/degraded; fill 80/90% on the "pbsdr-box" operator scope). /offsite split into Restic + PBS DR tabs; two dashboard gauges. Graceful: hub deploy ⟂ ep0 update (ep0 ≤ v1.1.0 → gauge "n/a" until updated). Phase-0-pinned (df on ep0 PBS 4.2.3) + live-computed in-cluster (ep0 updated to v1.2.0 this session): 19.1% full (7.1 GB of 37.2 GB). 10 Go tests + a bash harness + 3 red-proofs; hub v0.65.0 REPORT Open legs: UI render unit-verified only (hub UI password-gated); the fill alert email is unit + red-proof verified, NOT fired live (datastore nominal at 19%). Separate PBS threshold keys (default 80/90); no oversubscription (namespaces, not quotas). READ-ONLY
Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store hub v0.53, conventions IMPLEMENTED 07-13 closing bundle
Operator login password changeable from UI hub v0.54 IMPLEMENTED 07-13