MEASURED 2026-09-15 (BIGNIGHT Phase 6): privatebin updated 2.0.5 -> 2.0.6, catalog reverted to 2.0.5, and the box read „Frissítés elérhető — ma" over an Update that would have moved the pin BACKWARDS onto a possibly-migrated datadir. - stacks.CatalogOrder: the comparison gains a fourth answer (Ahead) and moves out of web, so the badge and UpdatePreflight cannot drift apart. - The badge: ahead reads „Naprakész"/"Up to date", tag-ok, with a title saying why. - The refusal: UpdatePreflight returns `downgrade` (409), born as a bundle key; the API now renders update refusals through errText so it reaches English households. - Ahead is narrow: every differing service must be orderable AND newer, else Behind. - Ordering is util.Version.Compare behind a tag normaliser — no second comparator. - Three red-proofs, each seen to fail. R-589 was already fixed in v0.258.0; only its register row was stale. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
344 KiB
felhom-controller
Central management container for Felhom home servers. The in-guest controller: one per
customer LXC, Docker-only, Hungarian web dashboard for managing the customer's app stacks, app-data
backups, monitoring and notifications. All Proxmox/disk operations are delegated to the host agent
(felhom-agent).
Current version: v0.59.0
Documentation has moved (and is now code-verified)
The authoritative architecture & feature documentation now lives in the central docs home:
felhom.eu/documentation/controller/— module map, deploy & stack lifecycle, backup architecture, storage/monitoring/metrics, and auth/hub/sync/integrations. Those docs are grounded in current source (v0.59.0).Quick build & deploy is in the "Build & Deploy" section below and in the repo
CLAUDE.md(authoritative for the workflow). NOTE the demo controller runs in an LXC guest (9201) under the bootstrap-managed mechanism (/etc/felhom-controller-image+felhom-controller-bootstrap.service), not the bare-metal/opt/dockercompose path some sections below still describe.The prose below this banner is retained legacy reference and may lag the central docs — when they disagree, the central docs win. (Some sections still mention restic / pre-8C disk handling that has since moved to the host agent.)
Table of Contents
- Architecture
- Features
- App Management
- App Export/Import
- Backup System
- Storage Management
- Monitoring & Health
- Notifications
- Update Management
- Authentication & Settings
- Central Hub
- Setup Wizard
- Disaster Recovery
- Asset Sync
- Debug Mode
- Geo-Restriction
- App-to-App Integrations
- Network File Sharing — SMB
- Async-job feedback
- Dashboard language (i18n)
- Repository Layout
- Configuration
- REST API
- Build & Deploy
- Roadmap
Architecture
┌─────────────────────────────────────────────────────────────────┐
│ Customer Hardware (N100 mini PC / Raspberry Pi) │
│ │
│ ┌──────────┐ ┌────────────────────────────────────────────┐ │
│ │ Traefik │ │ felhom-controller (privileged container) │ │
│ │ (reverse │──▶│ │ │
│ │ proxy) │ │ ┌──────────┐ ┌─────────────────────────┐│ │
│ └──────────┘ │ │ Web UI │ │ Stack Manager ││ │
│ │ │ (HU dash │ │ (compose ops, git sync, ││ │
│ ┌──────────┐ │ │ board) │ │ deploy, delete, update) ││ │
│ │cloudflared│ │ └──────────┘ └─────────────────────────┘│ │
│ │ (tunnel) │ │ ┌──────────┐ ┌─────────────────────────┐│ │
│ └──────────┘ │ │ Backup │ │ Storage Manager ││ │
│ │ │ (3-layer │ │ (disk scan, format, ││ │
│ ┌──────────┐ │ │ restic) │ │ mount, migrate) ││ │
│ │ App │ │ └──────────┘ └─────────────────────────┘│ │
│ │ stacks │ │ ┌──────────┐ ┌─────────────────────────┐│ │
│ │ (docker │ │ │Scheduler │ │ Monitor & Metrics ││ │
│ │ compose) │ │ │(cron-like│ │ (health, SQLite ││ │
│ └──────────┘ │ │ jobs) │ │ time-series, Chart.js) ││ │
│ │ └──────────┘ └─────────────────────────┘│ │
│ │ ┌──────────┐ ┌─────────────────────────┐│ │
│ │ │ Notify │ │ REST API + Hub Reporter ││ │
│ │ │ (events) │ │ (JSON push + events) ││ │
│ │ └──────────┘ └─────────────────────────┘│ │
│ │ ┌──────────┐ │ │
│ │ │ Assets │ │ │
│ │ │ (Hub │ │ │
│ │ │ sync) │ │ │
│ │ └──────────┘ │ │
│ └────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
│ events + reports │ git pull │ asset sync
▼ ▼ ▼
hub.felhom.eu gitea.dooplex.hu hub.felhom.eu
(central dashboard) (stack definitions) (logos, screenshots)
Key Architecture Decisions
- Pure Go, no frameworks — stdlib
net/http+html/template. Only external deps:bcrypt,yaml.v3,modernc.org/sqlite(pure Go, no CGO). - Privileged container — Required for disk operations (format, mount, fstab),
/devaccess, and Docker socket control. /host-devindirection — Docker overrides/devwith a tmpfs. The host's/devis mounted at/host-devto access block devices.StackDataProviderinterface — Breaks circular import between the backup packages and stacks. Defined ininternal/appbackup(and re-exposed via a type alias ininternal/backup). Implemented bystackAdapterinmain.go. ProvidesGetStackHDDPath()for per-drive backup routing.- Atomic file writes — All persistent state (
settings.json,app.yaml) written to.tmpthenos.Renamefor crash safety. go:embedtemplates — All HTML/CSS/JS compiled into the binary. No runtime file dependencies.- Design system v2 (v0.96.0, TASK-D0) — The whole UI renders in the Felhom v2 language (canonical
reference:
felhom.eu/documentation/design/design-system.md): navy token palette, single 2px radius, no shadows, exception-based status color (nominal = blue/neutral; amber/red only on deviation —stateColoremitsrun/progress/warn/neutral/off,usageColor/tempColoremitnominal/warn/crit; a stopped app is neutral, NOT red). Capacity bars are the hairline.metercomponent; state chips are.tag; informational pills are.metarow. Fonts (Plus Jakarta Sans + JetBrains Mono, variable woff2, latin+latin-ext) and a 30-icon Lucide sprite are vendored in the binary (internal/web/static/fonts/served at/static/fonts/;templates/icons.html) — no CDN, no emoji. The setup wizard serves the same embedded stylesheet viaweb.StyleCSS(). - Guest RAM resize (v0.143.0, R-24; MinAgent 0.90.0) — the Rendszer page's "Szerver memória (RAM)"
card shows the guest's current/used memory + the allowed range and lets the customer resize it. The
controller only proxies + maps the agent's machine
codeto Hungarian; the agent (felhom-agent v0.90.0) enforces every bound and applies the change live via PVE SetConfig — no reboot (min 2048 / max host_total−2048 / a shrink is refused below max(2048, usage+512)).systemPageDatacallsmemoryCardData(a 2 s agent GET behind theFeatureGuestMemoryResizegate);POST /api/system/memory/resize(ServeSystemAPI,internal/web/system_memory_handlers.go) →agentapi.ResizeMemory→ the code→Hungarian map. A JS confirm fires only on a shrink; an outdated agent hides the control; an unreachable agent falls back to the guest's own/proc/meminfo. Memory only (cores stay observation). The lxcfs ripple means the deploy-page memory math follows a resize for free. - „Hálózat" card (v0.159.0, R-66) — Beállítások → Rendszer, between „Verzió és frissítés" and
„Szerver memória": Helyi cím (LAN), Hálózati név (
\\<SMBServerName>— rendered ONLY while Megosztás is enabled; the NetBIOS name exists only while samba runs), Átjáró, plus a muted footer ("read this page aloud during remote troubleshooting"). Every value is live-computed per render and stored NOWHERE (S-5): the reads go through the samba-container netns door (internal/stacks/guestnet.go— the controller's own netns is the docker bridge, so in-process answers like/proc/net/routewould report 172.x, the S-2 trap); with Megosztás off the door is closed and rows render „—" (an address-less row beats a wrong address). Companions: the Debug system dump gains anetworksection (interfaces without veth*/docker*/br-* plumbing, default route + gateway + source interface, DNS from the guest's resolv.conf, and the SAMElan_addressthe card shows for cross-checking), and the NAS add form names the NetBIOS trap — helper text under Szerver, plus a purely lexical hint appended to anunreachablefailure when the submitted server is a single-label non-IP name (looksLikeFlatNetworkName; no NetBIOS/mDNS resolution is ever attempted). - Page IA (v0.97.0, TASK-D1) — the settings monolith is split into four pages, each with its own
data builder (
systemPageData/storagePageData/notificationsPageData/securityPageData, sharingsettingsBaseData) and template. Routes:/settings(Rendszer),/settings/notifications(GET→page, POST→save on the same path),/settings/security,/storage(main-nav Tárhely — Meghajtók), and/storage/network(Hálózati tárhely / NAS; v0.98.0 split — nested sub-links under Tárhely). The enrollment wizards live at/storage/{init,attach};/settings/storage/{init,attach}301 to them; all storage action successes redirect to/storage?storage_msg=…./storageshows a unified drive view: server-rendered registry cards enriched in place from the agent/api/disks(role/durable-id/actions, joined on mount path), plus read-only Rendszermeghajtók and Nem regisztrált meghajtók groups; agent-down degrades to a single warn note. Consequential actions use an in-page.confirm-overlay(openDialog) or the LIGHT inline two-step (felhomConfirmin layout.html, v0.123.0: the trigger swaps in place to "kérdés + Igen/Mégse"; form buttons opt in viadata-confirm="…") — never nativeconfirm()/prompt()(OS-modals freeze browser automation; drill F-11). Shared app-list row (v0.126.0):templates/app_row.html(app_list_row/app_list_row_end) is the ONE row grammar for app lists — icon + name (+ optional secondary line) left, caller action block right; used by the dashboard installed-apps list, the Távoli mentés toggle list and the Visszaállítás restore-to-verify/.fab lists; the backups-apps expander header is ALIGNED to the same grammar (own markup — it carries the toggle). Protected infra stacks The off-site restore list is keyed on the STORE (v0.204.0, R-237):offsite_restore_list.gobuilds it fromOffsiteInventoryList(the repository's own snapshot tags), NOT from deployed + offsite-toggled apps. A rebuilt box has neither and used to be shown nothing to restore while its snapshots sat in the repository. Installed-ness is a property OF a row (it changes what restoring implies), never a filter on it; an unreadable store renders as UNKNOWN and keeps the action; thefelhom-offboxand_sharesmarker tags are never offered as apps. Its two preconditions now name a reason AND a route (v0.207.0, R-252/R-253): a rebuilt box keeps its drives but loses their REGISTRATION, so the page renders a notice — „Előbb csatold vissza az adatmeghajtót", linking to/storage— wheneverHasRestoreDestination()is false, asked through the backup manager's own predicate so page and resolver read the sameGetSchedulableStoragePaths(). And the not-installed row no longer promises „a visszaállítás előbb újratelepíti": reconstitution writes to the app's ownGetStackHDDPath, which exists only after the customer picks a drive at deploy time, so the copy says to install it first and links to/stacks/<app>/deploy. Both notices are conditional — a healthy box renders exactly as before, pinned by a test that fails if either becomes unconditional. A run that skipped a selected app isincomplete(v0.205.0, R-234): the off-site verdict now countsmissingUnprotectedbesidemandatoryGaps— an app the customer selected that is DEPLOYED but has no recovery unit is not protected, so the run is notok. A selected-but-UNDEPLOYED app is named with what to do and does NOT move the verdict (a permanently amber box is a status nobody reads); a disconnected/decommissioned drive has its own signal. The manual „Távoli mentés most” also refuses SYNCHRONOUSLY when a run is already in flight, instead of answering „elindult” and leaving the previous run's verdict on the card. The box does not mint a repository key over a sealed package (v0.206.0, R-241):WriteOffboxSecretsauto-generated on ONE input — does the password file exist — while its two neighbours in the same file both consultedGetHubEscrowIdentityPresent(). A rebuilt box's credential self-heal therefore wrote a fresh key over the package the hub was holding for it, and the recovery screen then correctly reported that nothing was recoverable under the key the box held. The guard is a CONJUNCTION (a package held AND no key present), so a first-time box mints exactly as before; the refusal is a HOLDING state that still writes the transport and declaresoffsite.state=awaiting_recovery_key, so/recoverycan bring the tier up the moment the key arrives.--abandon-status/--abandon-extend=N/--abandon-stopare the operator levers. The recovery offer has a third shape (v0.206.0, R-241): shape (c) — the hub holds a package for a key OTHER than the one we are using. Shapes (a) "no key at all" and (b) "a run proved the repo will not open" are proxies for that question and have each been wrong in opposite directions. The comparison was already computed on every report ACK and persisted nowhere; it now lives insettings.HubEscrowKeySHA256. A KNOWN difference offers however old the reading; a hash never learned falls back to (a)/(b). Abandoning the old history is a finishable thing (v0.206.0, R-241): the customer's confirmed set-aside starts a 14-day countdown, visible on Távoli mentés and reversible by recovering with the code, at the end of which the set-aside store AND the hub's retained sealed package are removed TOGETHER (two-phase commit; the hub half isPurgeSupersededEscrowForCustomer, hub v0.98.0). Afterwards shape (c) has nothing to compare and the offer ends because the state is right. The recovery screen appears once per ENTRY into the offered state (an epoch), the reminder bar is per-visit, the durable opt-out silences the bar only, and none of the three removes the entry point on Távoli mentés. While a recovery is outstanding, „Helyreállítási kód létrehozása" is UNAVAILABLE — creating a new code would put the earlier history beyond every shipped path. A code for an EARLIER package is now recognised as correct (v0.214.0, R-311; MinAgent 0.129.0). When the supplied code does not open the package the hub currently holds, the agent tries the RETAINED ones and answers 422 if one opens; the screen then says the code is correct, names the supersession date, says the earlier package is kept and that the CURRENT backups are unaffected, and routes to support. It deliberately promises no restore — there is no in-product route to a set-aside store (R-312), and the retained package may itself predate the repository-password field. On an agent older than 0.129.0 the screen keeps the earlier hedged wording, which was honest: until something looked, a correct-but-earlier code and a mistype really were indistinguishable. (traefik/cloudflared/filebrowser) get curated Hungarian display identity from theinframeta.gomap (name + description + generic/static/infra-logo.svgfallback icon); filebrowser is the only infra stack with a customer link (files.<domain>). Universal app placeholder (v0.163.0):app_list_rownow DEFAULTS its fallback icon to the embeddedAppPlaceholderSVG(a 2×2 app-grid glyph, served at/static/app-placeholder.svg), so a logo-less app shows a placeholder on every list surface instead of a hidden icon; infra rows still override with the server glyph. The felhom brand mark is never an app placeholder (brand = platform identity only). Enforcement — one entry point (2026-08-02): runpython3 scripts/controller_gates.pyfromcontroller/after any template change. It is THE runner and invokes every gate:template_id_gate.py(JS element-ID integrity),emoji_gate.py(no emoji),native_confirm_gate.py(zero native confirm/prompt),app_row_dedup_gate.py(row markup single-sourced),mojibake_gate.py(no double-encoded UTF-8 in templates/Go sources),docker_run_volume_path_gate.py(everydocker … -vmount reviewed), andreuse_refs_check.pyon the repo root. It exits non-zero if any gate does, and a missing gate script is a FAILURE, not a skip.--fast(what.githooks/pre-pushruns) selects the gates that touch no network and no container runtime — today all of them. The GoTestNoEmojiInTemplatesmirrors the emoji gate. Why a runner: of this project's gates, only the ones named by aCLAUDE.mdentry point ever got run — the 2026-08-02 census found the two unnamed ones red, one for nineteen days. - Europe/Budapest timezone — All scheduled jobs, timestamps, and UI labels use Hungarian timezone.
Module Map
| Module | Path | Responsibility |
|---|---|---|
| Config | internal/config/ |
YAML loader, validation, FELHOM_* env overrides |
| Settings | internal/settings/ |
Runtime-mutable settings.json (passwords, backup prefs, storage paths, notifications) |
| Stacks | internal/stacks/ |
Compose operations, scanning, .felhom.yml metadata, deploy/delete flow; base-infra bring-up (infra.go — EnsureBaseStack) |
| Infra | internal/infra/ |
Pure renderers (embedded text/template) for the base-infra stacks (traefik/cloudflared/filebrowser); pinned image tags as the single source of truth (web filebrowser sync delegates here) |
| Crypto | internal/crypto/ |
AES-256-GCM encryption for sensitive app.yaml values (passwords, secrets), key management |
| Sync | internal/sync/ |
Git-based app catalog sync (clone/pull, content-hash copy) |
| AppBackup | internal/appbackup/ |
Self-contained app-data backup primitives: DB dump discovery/execution (DiscoverDatabases, DumpOne), Docker-volume/app-data discovery (StackDataProvider, DiscoverAppData), keep-side path helpers (AppDBDumpPath, AppVolumeDumpPath, AppDataDir). DiscoverDatabases reads each container's com.docker.compose.project label and prefers it as the stack name (v0.218.0, R-355) — the label is the stack name by construction, since compose runs with cmd.Dir set to the stack directory and no -p; the deployed-stack set (M19, v0.62.0) remains the fallback for containers not started by compose, and an attribution that resolves to no known stack now WARNs instead of being returned silently. Before this, paperless-ngx (container paperless-postgres) was attributed to a non-existent stack paperless, so its 72-table PostgreSQL dump landed outside its recovery unit, never reached the off-site copy, and was never restored — and the same value reaching writeSafetyDump meant a destructive restore of that app took no undo copy at all. ListDumpFiles takes an optional cached(name,size,mod) lookup so an unchanged dump isn't re-validated (line-scan) every ~5-min cycle (M18, v0.62.0). No dependency on restic/cross-drive/drive-mount. Imported directly by appexport and storage. |
| Backup | internal/backup/ |
Per-drive 3-layer backup: DB dumps → restic snapshots → cross-drive copies, restore. Re-exposes the appbackup primitives via aliases/forwarders (appbackup_bridge.go) for the disk/host-side code and the web/api/report consumers. |
| Storage | internal/storage/ |
Disk scanning (lsblk), partitioning (sfdisk), formatting (mkfs.ext4), mounting, data migration (rsync) |
| System | internal/system/ |
System info (/proc), CPU collector, mount points, disk usage, FS info |
| Monitor | internal/monitor/ |
System health checks, storage watchdog, legacy Healthchecks pinger (deprecated) |
| Metrics | internal/metrics/ |
SQLite time-series store, system + container metric collection |
| Scheduler | internal/scheduler/ |
Central job scheduler (periodic + daily, skip-if-running, panic recovery). UpdateDaily reschedules a daily job at runtime (no restart) via a per-job reschedule signal (v0.168.0). |
| Backupwindow | internal/backupwindow/ |
Pure time math for the customer-configurable backup window (v0.168.0): ParseHHMM/FmtHHMM, LegTimes (W / W+60m / W+105m, wrap-safe), GateWindow ([W+2h, W+6h)), EffectiveWindow (settings > yaml > "02:30"). Offsets are constants — derived, never stored. |
| SelfUpdate | internal/selfupdate/ |
Version checking (registry), update trigger, state persistence, startup verification |
| Notify | internal/notify/ |
Email notifications via hub relay, preference sync, per-event cooldowns |
| Report | internal/report/ |
Hub report builder + HTTP pusher (system, stacks, backup, health) |
| Assets | internal/assets/ |
Hub-managed asset syncer: downloads logos/screenshots with SHA-256 change detection |
| SelfTest | internal/selftest/ |
Startup self-test: 9 diagnostic checks (Docker, dirs, storage, hub, restic, metrics) |
| Util | internal/util/ |
Shared utilities: TruncateStr for debug log output truncation |
| AppExport | internal/appexport/ |
Per-app export/import via .fab bundles (config + DB + user data), optional AES-256 encryption |
| API | internal/api/ |
REST JSON endpoints, diagnostic dump (/api/debug/dump) |
| Web | internal/web/ |
Hungarian dashboard, auth, page handlers, template functions, alerts |
Features
Disk health — "Lemezek állapota" card + degradation alert (v0.169.0; severity ladder v0.215.0)
The dashboard shows a per-physical-disk health card driven by the agent's SMART summary (serialized
into /disks from agent v0.94.0 — no new smartctl load; the controller only reads it). One pure
verdict function (agentapi.DiskVerdictFor) is the shared truth for the card chip and the check.
Why it does not trust smart_status.passed: attributes 187/197/198 all carry thresh: 0 and a
normalized SMART value floors at 1, so a drive's own overall verdict is structurally incapable of
failing on unreadable sectors — the real failing drive read PASSED at 352 pending sectors. The ladder
reads the raw counters instead. Evidence:
felhom.eu/documentation/audits/DIAG-smart-passed-trap-2026-08-14.md.
Four labels, and there will not be a fifth — a predicted failure is „Hiba", the same word a self-reported failure gets:
- Rendben — clean.
- Figyelmeztetés — a first sighting of unreadable sectors below the bar, reallocated sectors, NVMe
media errors,
percentage_used ≥ 90, ortemperature_c ≥ 55. - Hiba — the drive reports FAILING, or (v0.215.0) unreadable sectors are sustained across
two consecutive checks, unreadable sectors sit alongside reallocation, the count reaches 64,
temperature_c ≥ 60, NVMe'scritical_warningis set, orpercentage_used ≥ 100. - Nincs adat — nil/UNKNOWN/old agent. Never alarms, and never erases a disk's stored history.
DiskVerdictFor(smart, prior) takes an agentapi.DiskPrior; a zero prior is the fail-safe, so a first
sighting can only reach Figyelmeztetés. The full 14-row truth table, evaluated top-down, is documented
in internal/agentapi/diskverdict.go together with the provenance of 64 / 55 / 60.
- The card fetches
/disksthrough a 60 s TTL cache (dashboard refresh-spam can't smartctl-storm the host); an unreachable agent renders "Nincs adat" and the page still loads. The card replays the prior that produced the stored verdict (diskRecord.PriorSawUncorrectable) so the chip and the email cannot disagree. - An hourly
disk-health-check(v0.215.0; was 6-hourly — the real drive's benign excursion lasted about one hour, and/disksmeasures ~0.82 s) emitsdisk_health_degradedagainst persisted per-disk state indisk-health-state.jsonundercfg.Paths.DataDir. It fires on an escalation against the last alerted verdict, and re-alerts a disk already at Hiba once it has both doubled its unreadable-sector count and waited out a 24 h cooldown. A disk's first verdict baselines silently; recovery and UNKNOWN never notify; a disappeared disk is forgotten. A restart no longer re-baselines — that is why the state is persisted. Every cycle logs a positive observable (disk-health check complete: N disk(s) evaluated, M alert(s)). - Severity is a wire contract: the alert emits
warning/critical— the hub coerces anything outside{info, warning, error, critical}toinfoand never emails it.DiskAlertKind.Severity()is the single mapping. The hub allowlist must includedisk_health_degraded. - Five message shapes (
notify.DiskAlert), because the customer's action differs: Figyelmeztetés, drive-reported FAILING, Hiba-from-sector-count, Hiba-from-heat, and the still-worsening re-alert. - No global banner (deliberate): the card + email carry it.
1. App Management
The controller manages Docker Compose stacks through a complete lifecycle: catalog sync, first-time deployment, runtime operations, and deletion.
Git Sync (internal/sync/)
The app catalog lives in a separate Git repository. The controller:
- Shallow-clones the catalog on startup
- Periodically fetches updates (configurable, default 15 min)
- Copies only
docker-compose.ymland.felhom.ymlto the stacks directory - Never overwrites
app.yaml(user secrets are safe) - Uses SHA-256 content hashing — only writes files that actually changed
- Triggers stack rescan after sync so the dashboard updates immediately
- Post-sync hook: auto-injects missing deploy fields (new secrets, domains) into existing
app.yamlfor stacks whose templates were updated (see Missing Field Injection below) - Manual sync via "Sablonok frissitese" button or
POST /api/sync
First-Time Deploy Flow
- Customer sees app card with "Telepites" button
- Deploy page pre-generates and displays all auto-values before the user clicks deploy:
domainfields: shown as readonly text input with the customer's configured base domainsubdomainfields: editable text input pre-filled with the default from.felhom.yml, shown with.base-domainsuffix. Validated for DNS-safe format, reserved names, and uniqueness across deployed stacks. Locked after deploy — changing requires Remove + Redeploysecretfields: pre-generated and shown as masked password inputs with a "Megjelenítés" reveal button — user can see/copy all DB passwords and keys before deploying- User-configurable inputs (admin password, language, storage path) remain editable
- Section header prompts the user to note down any passwords they need
checkBeforeDeploy()JS guard fetches live state first (prevents double-deploy from another tab)- Memory validation (F1, v0.61.0): the controller runs as a Docker container inside an LXC, where
/proc/meminfoshows the Proxmox host's RAM (no lxcfs in the container) and the container's own cgroup is unlimited (the guest cap lives on the LXC ancestor). So the guest cap is read from the Docker daemon (system.GuestMemTotalMB()→docker infoMemTotal — the daemon runs in the LXC and reports the guest's lxcfs-backed RAM; the cgroup limit is preferred when present, e.g. non-nested):usable_memory = guest_cap - reserved_memory_mb(default 384MB reserved)- Hard block if
committed_used + new_request > usable_memory, wherecommitted_used=CommittedMemory()(sum of running apps' mem requests) — the guest-wide RSS is not observable from the container, so the controller's own committed accounting is the accurate, cheap "used". /api/system/inforeports the guest cap as total and committed memory as used. 4b. Docker-data volume reserved-buffer gate (v0.58.0, storage-split prevention layer): the OS rootfs and Docker data are split onto separate volumes; infra (controller/traefik/cloudflared/filebrowser) shares the one Docker data-root (/var/lib/docker) and is protected by prevention, not placement.system.GetDockerVolumeHeadroom()measures the Docker-data volume viastatfs("/")(the controller container's overlay root is the upperdir on that volume — true with the golden's overlay2 driver) and reservesmax(5 GB, 10%).deployStackrefuses a new deploy (HTTP 507) when free space is at/under the buffer; the deploy page shows the warning + disables the button. Fail-open on a statfs error. The runtime disk monitor (healthcheck.go, warn 80% / crit 90%) watches the same volume and trips above the buffer. (Assumes the split guest's large data volume; the golden bakes overlay2 + log rotation so images+volumes live on the data volume, not/var/lib/containerd.)
- Pre-generated secret values are submitted as hidden form inputs so the same values the user saw are saved to
app.yaml(no silent re-generation on submit). Controller savesapp.yaml, sets in-memoryDeployed+Deployingflags, then runsdocker compose up -dasynchronously in a goroutine — API returns immediately so the UI switches to the progress panel without waiting for image pulls. On failure the goroutine reverts both disk and in-memory state and setsDeployError. - 3-step progress panel polls
GET /api/stacks/{name}every 3s: config saved →deploying(pulling images) → containers starting → health check passed. NewStateDeployingstate shown while compose-up is in progress (no containers yet). - Post-deploy: locked fields (DB_PASSWORD, etc.) become read-only; the "Automatikusan generált értékek" section continues to show the saved values on the settings page
- The deploy/settings page includes start/stop/restart buttons for deployed apps, plus a "Megnyitás ↗" link to the app's subdomain URL (only visible when running)
Catch-All Page for Stopped Apps
When a user visits a stopped or undeployed app's subdomain (e.g., travel.demo-felhom.eu), the controller serves a branded error page instead of Traefik's raw 404:
- Traefik catch-all router: The controller's
docker-compose.ymlregisters a second router (catchall) withpriority=1(lowest) andHostRegexp(.+). Running apps always win; only requests with no matching container reach the controller. CatchAllMiddlewareinserver.gointercepts requests whereHost≠felhom.DOMAIN, serves the catch-all page without auth (user has no session on the app subdomain).findStackBySubdomain()identifies the app by matching the subdomain against deployedapp.yamlSUBDOMAINenv or metadata fallback.catchall.html— standalone template (no layout, inline CSS) showing the app name, status ("leállítva" / "nincs telepítve" / "nem található"), and links to the controller dashboard or the app's detail page.- Subdomain links on the Alkalmazások page are only shown for deployed apps (non-deployed apps have no guaranteed subdomain yet).
Indítópult (app launcher page) (v0.163.0)
/launcher (launcherHandler + templates/launcher.html) is the first sidebar item, above
Vezérlőpult; / still lands on the Vezérlőpult. It renders a grid of large tappable tiles — one per
openable deployed app. Openability has a single criterion, shared with the „Megnyitás" button: the
stack has a subdomain (env SUBDOMAIN > .felhom.yml subdomain > protectedStackSubdomains),
resolved through the extracted Server.subdomainMap helper (the dashboard and Alkalmazások pages use
the same helper). The controller's own stack is excluded by name.
Each tile is a colored rounded square: funcmap.tileColor(slug, brand) returns a validated
.felhom.yml brand_color (#rgb/#rrggbb, Metadata.BrandColor) or, when absent/invalid, a
deterministic FNV-1a-of-slug → HSL color (fixed S/L, hue varies per app). The white monochrome logo
renders on top of a monogram initial (funcmap.initial, multibyte-safe); if the logo fails to
load the monogram shows through (the launcher does NOT use the app-placeholder here). Operational apps
are <a target="_blank" rel="noopener"> links; stopped/degraded apps render greyed + unclickable with
the Hungarian state badge. Empty state links to /stacks.
Indítópult megosztása — guest launcher via capability URL (v0.165.0)
The admin launcher's "Indítópult megosztása" button mints a capability URL —
https://<host>/s/<token>, where token is a 160-bit crypto/rand value
(newShareToken, base64.RawURLEncoding, 27 chars) — that serves a standalone, read-only guest
launcher with no account and no admin session. The link grants information only, zero
control: app names + public URLs; every privilege stays behind each app's own auth and the
controller admin password. The tile visual is shared with the admin launcher via the launch_tile
template partial; the app slice comes from the extracted Server.launcherApps() helper.
- Routing (
internal/web/share.go,share_handlers.go):/s/<token>joins the RequireAuth pre-auth allowlist after the claim-gate block (an unclaimed box never serves the guest page — the claim gate stays supreme) and is exempted from session CSRF (the guest password POST carries a pre-auth HMAC CSRF,validShareCSRF, mirroring the claim POST). Token match issubtle.ConstantTimeCompare; an empty stored token (= sharing OFF, there is no separate flag) matches nothing, so a wrong/disabled token returns a byte-identical mux-default 404 (share404). Guest responses setX-Robots-Tag: noindex, nofollow/Referrer-Policy: no-referrer/Cache-Control: no-store. The token is a secret: the ServeHTTP debug line and the 404 WARN redact/s/paths to/s/<redacted>. - Optional per-share password (
settings.LauncherSharePasswordHash): a SEPARATE bcrypt credential (never the adminPasswordHash), guarded by its OWN per-IP 5/1-min attempt map (shareAttempts, never the adminloginAttempts). A correct password mints a signed gate cookie = HMAC-SHA256(token|passwordHash) keyed with the persisted, box-scopedweb.session_secret— so rotating the token OR changing the password invalidates every outstanding cookie with no bookkeeping. - Guest state labels ride the v0.164.0 ruling and never expose internal vocabulary: clickable ⇔
isOperationalState && !routeUnpublished(operational AND route actually published, so a tap never dead-ends);StateStopped⇒ "A tulajdonos leállította"; any other non-clickable state ⇒ "Átmenetileg nem elérhető". Empty ⇒ "Jelenleg nincs elérhető alkalmazás." (buildGuestAppsis the pure, tested mapping; templateslauncher_shared.html+launcher_share_password.html). - Admin modal (in
launcher.html): current link + copy button, QR code (GET /launcher/share/qr.png, ~256px PNG viagithub.com/skip2/go-qrcode, admin-authed,no-store), set/clear share password, "Új link készítése" (rotate), "Megosztás kikapcsolása" (clears token AND password). The management POSTs live under/launcher/share/*and ride the normal admin session + session CSRF; rotate/disable use the inlinedata-confirm(felhomConfirm) affordance. A feature-detected "Megosztás…" button (v0.165.1) opens the OS share sheet vianavigator.share(title + text + URL only — never the QR as afiles:attachment); hidden unless the browser supports it, with "Link másolása" as the universal fallback (the non-cancel rejection path falls back to it too).
Design ruling: member accounts are superseded by this capability-URL model; per-member tile visibility is parked under the SSO arc.
Dashboard "Megnyitás" Button
Running apps on the Vezérlőpult now show a "Megnyitás ↗" button that opens the app's subdomain in a new tab. The Subdomains map is built in dashboardHandler from app.yaml env or metadata fallback.
The optional .felhom.yml open_path field (Metadata.OpenPath) is appended to that URL for apps whose UI isn't at / — e.g. gokapi → /admin, ghost → /ghost/. Empty = bare root. Rendered via .Meta.OpenPath in dashboard.html, deploy.html, and app_info.html (the same field flows through all three open-link sites; no handler change needed). Must start with /; cosmetic only (does not affect routing).
App lifecycle — withdrawing an app without orphaning anyone (v0.158.0)
.felhom.yml carries an optional top-level lifecycle: (Metadata.Lifecycle), the catalog's
answer to "stop offering this app" that does not punish the customers already running it.
| value | offered for new installs? | shown to someone already running it |
|---|---|---|
available (default; absent/empty ≡ this) |
yes | nothing special |
hidden |
no | nothing — "we stopped offering this" is not their problem |
abandoned |
no | „Nem karbantartott" badge + a notice on the app page that updates and security fixes will no longer arrive |
A deployed instance keeps full function in every state. Lifecycle governs what is OFFERED, never
what runs — deleting a template instead would mark every deployed instance Elavult and offer a
Törlés button for working software.
Three predicates on Metadata are the single interpretation of the field — every surface goes
through them: EffectiveLifecycle(), CanInstall(), IsAbandoned().
- Listing —
web.visibleCatalogStacksdrops a template that is not installable AND not deployed here (Deployed || Protected || CanInstall()). - Deploy gate —
api.deployStackrefuses server-side before any mutation with „Ez az alkalmazás jelenleg nem telepíthető." (409).stacks.DeployStackrepeats the check for any caller that does not route through the API. Hiding the button is not a gate. - Unknown values fail OPEN (→
available+ one WARN), deliberately opposite to the gate's fail-closed posture: a typo, or a state from a newer catalog than this controller, must never pull a working app out of every customer's catalog. Both read the sameEffectiveLifecycle, so they cannot disagree. - Orphan detection must never see this field.
getCatalogTemplateSlugskeys on directory + compose presence only; withdrawn templates stay in the catalog tree. Asserted byTestCatalogTemplateSlugs_IgnoresLifecyclewith a red-proof. - Badges are generic plumbing:
web.MetaBadge+ themeta_badgetemplate partial + thelifecycleBadgefuncmap entry. R-56's difficulty labels are intended as a sibling funcmap function returning the same*MetaBadge— no new markup or CSS.
What is installed, and is it current? (v0.233.0 — update arc slices 1 & 2)
Two additions, and neither changes how an update behaves. Slice 1 is a record; slice 2 is a label.
Slice 1 — app.yaml gains installed_images. After every successful compose up from
StartStack, RestartStack, the guarded update (v0.237.0; UpdateStack before it) and the deploy path, Manager.recordInstalledImages
(internal/stacks/installed.go) reads what each container is ACTUALLY running and writes it down,
keyed by compose SERVICE name:
installed_images:
web:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:aaaa… # the only identifier that cannot move; "" if never pulled
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
- Read from the CONTAINER, never from
docker-compose.yml. That file is the value that has already moved: the catalog syncer overwrites a deployed app's compose file on a 15-minute cycle with no deployed check, and file and container can disagree indefinitely (measured live,SPIKE-app-update-2026-09-01§3).Manager.checkLocalImagesis a line scan of that file and is deliberately NOT reused. - A failed write NEVER refuses the action — the deliberate opposite of
SetDesiredState.desired_stateis the customer's INTENT, so an act whose intent could not be recorded is refused;installed_imagesis an OBSERVATION, and refusing to start an app because a note could not be written would trade a real outage for a bookkeeping gap. Logged at ERROR and the app stays up. - NOT called from
StartStackServices— that path starts only the database service for the R-47 restore window, and a partial record would overwrite a complete one. - Absent means UNKNOWN and never means current. Every
app.yamlpredating v0.233.0 has no entry. - Seeded at startup for apps nobody touches (v0.234.0).
Manager.BackfillInstalledImagesruns once at boot, beside the desired-state backfill, and records what every deployed app is ALREADY on. It only READS containers — it starts nothing and writes no compose file. It never overwrites an existing record, and it refuses to seed a partial observation:updateBadgereads a service-count mismatch as BEHIND, so a degraded app seeded from what is visible would show „Frissítés elérhető" while being perfectly current. Without this, v0.233.0's label never appeared on an app that simply ran (found on demo-felhom, 2026-09-03). - Its own docker seam (
Manager.installedExecFn) carries a context and a 30 s timeout, whichcomposeExecCustomEnv/execCommanddo not — a bookkeeping read must not be able to wedge a lifecycle action.
Slice 2 — one badge, and no version number. .felhom.yml gains optional
catalog_since: "YYYY-MM-DD" (Metadata.CatalogSince + CatalogSinceAge), the date the catalog last
moved that app's pins. web.updateBadge compares the recorded reference for each service against what
the current template pins and returns a *MetaBadge rendered by the existing meta_badge partial —
no new markup, no new CSS, which is exactly what metabadge.go's comment asks of its second user.
| state | badge |
|---|---|
| every service matches the template | „Naprakész" (tag-ok) |
| every differing service is provably NEWER than the catalog (v0.260.0) | „Naprakész" (tag-ok), with a title saying the app is ahead |
any service differs, catalog_since usable |
„Frissítés elérhető — N napja" (tag-warn) |
any service differs, catalog_since absent/malformed/future |
„Frissítés elérhető" |
| no record, or the template cannot be read | nothing is rendered |
The fourth row is v0.260.0 (R-524), and it is also a REFUSAL. The comparison lives in
stacks.CatalogOrder (internal/stacks/updateorder.go) — Unknown / Current / Behind / Ahead —
and web.compareInstalledToTemplate is a thin wrapper over it, because the badge and
Manager.UpdatePreflight must reach the same verdict. A box AHEAD of the catalog (the catalog was
reverted under it) reads „Naprakész" and its Update is refused with reason downgrade:
„Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére."
Ahead is narrow on purpose: EVERY differing service must be orderable and newer, or the answer
falls back to Behind. Ordering is util.Version.Compare (the one comparator) behind a tag
normaliser — X.Y/X.Y.Z, optional leading v, and a trailing suffix that must be IDENTICAL on
both sides, so 31.0.14-apache → 31.0.15-apache orders while postgres:16-alpine,
26.05.2-ls310 → -ls311, apache-2.57.0, a date stamp and a digest pin do not.
- No version string is shown to the customer anywhere — operator ruling, 2026-09-02: a household
cannot act on
26.05.2, only on "you are behind, and by this long". Versions stay in the logs, the API and the hub. - No registry is queried. A customer's box must not need eight upstream registries to render a
page. Known limitation: for the 23 floating pins (
postgres:16-alpine,mariadb:11.6, …) the reference can be identical while the image behind it has moved, so those apps can read „Naprakész" when they may not be. Digest-level comparison needs a registry query and is deferred. - Information only. The badge is wired to no action. (What the
Frissítésbutton itself does changed in v0.237.0 — see "The guarded update" below.)
Reasoning and the seven-slice plan: felhom.eu/documentation/architecture/09-update-architecture.md.
Freeze the version, keep the fixes flowing (v0.235.0 — update arc slice 3)
Operator ruling, 2026-09-06. An app's version is frozen to what the customer has; only a deliberate Update moves it. Everything else in a template — health checks, memory limits, new deploy fields — still arrives on the 15-minute cycle, and a broken definition still repairs itself.
app.yamlgainspinned_images(service → ref): what the app is SUPPOSED to run. Notinstalled_images, which is an observation. Written only by the deploy path, the guarded update, a restore, and the one-timeAdoptPins. Absent = unpinned = pre-v0.235.0 behaviour.applied-compose.ymlin the stack dir stores the exact definition the pin came from. The syncer copies onlydocker-compose.ymland.felhom.yml, so that name is safe.Syncerrenders instead of copying, via the nil-safeSetRenderPlanFnseam. Catalog images equal the pin → copy verbatim (fixes flow, self-healing works). They differ → write the stored definition, whole — never a substitution of refs into a newer template (wger 2.6)..felhom.ymlalways flows, even to a frozen app: it holds no image and carriescatalog_since. Known limitation, R-458.- Nothing was added to the thirteen
compose up -dcall sites. Most are repairs; a repair that refuses to repair leaves an app down. - The badge reads
Stack.CatalogImages, neverTemplateImages. After the freeze the live compose file is the frozen one, so comparing against it would answer „Naprakész" on apps that are behind.
Reasoning: felhom.eu/documentation/architecture/09-update-architecture.md §3, §5.
The guarded update (v0.237.0 — update arc slice 4)
POST /api/stacks/{name}/update no longer updates on the spot. It refuses what it must, starts a
job, and answers 202. The page polls GET /api/stacks/{name}.
| field | meaning |
|---|---|
updating |
a guarded update is in progress |
update_phase / update_phase_label |
checking, backing-up, safety-dump, pinning, pulling, starting, verifying, done, failed — and the Hungarian label for each |
update_error |
the customer sentence when the update did not complete |
hold_reason |
the hold's sentence while the app is held (failed update OR failed restore) |
Refused with 409 before anything moves: held; a backup, restore, app-data op or quiesce holding
the app; a migration; already updating; deploying; not enough memory for the NEW template's request
(the deploy's own memoryVerdict, releasing the app's current request); less than 2 GB free on
the Docker data root (a fixed floor — image sizes are not known without a registry query); and — since
v0.239.0 — no copy on any backup tier AND no way to take one now (drive unresolvable, disconnected,
or a migration running).
Any backup tier counts (v0.239.0, R-475). The update leans on the first FRESH copy in the order
second drive (Tier 2), the app's own recovery unit (Tier 1, „helyi"), off-site (Tier 3, looked up with
a 15 s bound — unreachable counts as absent, with a WARN; since v0.240.0 it runs no per-app stats). update.backup_max_age applies to whichever
tier is chosen. An app with no copy anywhere is backed up first. Tier 2 is required nowhere in the
update path; the backups page's „Teljes visszaállítás" still reads the Tier-2 predicate alone.
Since v0.241.0 (R-479) an app whose data is bind-mounted files walks second drive → off-site → own
unit (its unit holds settings and database dumps, not the files), and the hold sentence ends with
what the chosen copy holds („… — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a
fájlokat nem." / „… a beállításokat, az adatbázist és a fájlokat tartalmazza.").
The sequence. With no fresh copy on any tier the app is backed up first (RunAppBackupNow: this
app's DB dump, volume dump, unit capture, then a Tier-2 copy whose failure is only a WARN — the unit
just captured is marked proven current, so a quiet app's own unit counts as fresh). Then a database
safety dump, then the pin moves, then pull, up, and the health wait (.felhom.yml check, or 60 s of
every container running for an app with none; bounded by update.health_timeout). A failed pull
puts the pin back. An app that does not become healthy is stopped and HELD — the pin stays on the
new version, and the hold sentence names the tier and the date of the copy it can be restored from
(„második meghajtó" / „saját meghajtó" / „távoli mentés"). A successful unit restore — and, since
v0.239.0, a successful off-site restore — lifts an update hold, and since v0.240.0 the held update's
sentence leaves the card with it. A copy older than the app's current deployed_at does not count
(v0.240.0) — it belongs to a previous install. The off-site lookup is one restic snapshots call.
Config (controller.yaml):
update:
backup_max_age: 24h # the chosen copy (any tier) must be younger than this, or the app is backed up first
health_timeout: 5m # how long the new version has to become healthy before the app is held
Crash safety: <data>/update-journal.json is written before every phase; RecoverUpdates (before
the boot sweep) puts a pin back or marks an interrupted update for ResumeInterruptedUpdates.
Every unattended start path honours a hold: the boot sweep and the app-stop guard (as before), and since v0.237.0 the drive-return gate and the nightly volume dump. The nightly capture and Tier-2 run skip a held app so its restore point is not overwritten — and, since v0.238.1, an app whose update is still in progress (the periodic capture overwrote a primary unit during a health wait, found live).
The page (v0.238.0). Frissítés follows the job — the button shows the phase label and the page
reloads when the update ends. An updating card offers no lifecycle button; a held card shows the hold
sentence with a Mentések link and nothing that would start it; a failed update that held nothing
shows its sentence. These checks run BEFORE isOperational, which counts restarting as operational.
Not done, deliberately: the old version is never put back automatically — whether that works is
per-app and was measured unpredictable. Reasoning: felhom.eu/documentation/architecture/09-update-architecture.md §6.
App Info Pages
Each app can define rich metadata in .felhom.yml:
app_info: tagline, use_cases, first_steps, prerequisites, default_creds, docs_urloptional_config: groups of post-deploy configurable env vars (e.g., API keys for metadata providers)resources: mem_request, mem_limit, pi_compatible, needs_hdd, hungarian_uiinitial_credentials(v0.84.0): for apps that auto-generate a first-login password into a file at first boot (vs taking it from a deploy field) — e.g. Crafty →/crafty/app/config/default-creds.txt. Fields:file(path inside the container),format(json|regex|plain), optionalcontainer(defaults to the stack's main container),username_key/password_key(json),username_pattern/password_pattern(regex, first capture group),note.ReadInitialCredentials(internal/stacks/initialcreds.go) reads the file live viadocker exec … catand parses it with the pureparseInitialCreds; the value is never persisted toapp.yaml.appDetailHandlersurfaces it on/apps/{slug}as a "Kezdeti belépési adatok" card (masked password + reveal/copy), labelled as the initial credential (stays valid only until the customer changes it in-app). Hidden when the container is down / file missing / parse fails. Reuse for any future self-seeding app.data_paths(v0.172.0, R-75): the customer-facing folder annotation —{path, root, role, label}whereroot ∈ {import, userdata, hdd}androle ∈ {import, library, export}. It ANNOTATES paths that must already exist as compose binds and can never declare one (so no new filesystem-write primitive comes from catalog data). Validation is deliberately ASYMMETRIC: a malformed path is a whole-block reject (reusingappbackup.ValidateRelPath, the same refusal set asbackup:), an unknown role fails OPEN with one WARN (theLifecycleprecedent — presentation, not data handling). Rendered on/apps/{slug}as „Hova tegyem a fájlokat?" for DEPLOYED apps only, each row a FileBrowser deep link plus a class-driven consequence line. Full contract:felhom.eu/documentation/controller/import-and-data-paths.md.backup(v0.132.0; three lists since v0.172.0): the referential-coupling classification block (Task 2). Optional listsuserdata:(relative to${USERDATA_PATH}),hdd:(relative to${HDD_PATH}) andimport:(relative to${IMPORT_PATH}, R-75), each of{path, class}whereclass ∈ {mandatory, optional, excluded}(COUPLED / DECOUPLED-precious / DECOUPLED-bulk).LoadMetadatavalidates the block against the app's compose binds and rejects the WHOLE block (→ nil + one[ERROR], app behaves as legacy) on any defect. Semantics (appbackup.ClassifyBinds): an explicit entry wins over the:rodefault; an unlisted writable bind defaultsmandatory, an unlisted:robind defaultsexcluded; no block at all = legacy behavior per tier. INERT as of v0.132.0 — the schema/parser/classifier + theManager.ClassifiedBinds/StackDataProvider.GetStackClassifiedBindsseam exist, but no backup tier consumes them yet (Task 3 = tier policy engine, Task 4 = manual.fabUI). Seefelhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md.- Capture-set computation (v0.133.0, Task 3-core, INERT):
appbackup.ComputeCaptureSet(binds, hasClassification, tier, hddPath)(internal/appbackup/captureset.go) is the pure path-algebra that turns classified binds into a tier-filtered, structurally-guarded, containment-deduped absoluteCaptureSet{HasClassification, Paths, Skipped}(slash algebra, nofilepath/FS/log).TierOffsite= mandatory only;TierSecondary= mandatory + optional; excluded dropped; legacy short-circuits to unit-only. Structural guards (traversal / bare HDD drive-root / reservedbackups/zone) move refused would-be captures intoSkippedfor the engines to log. CompanionCrossAppOverlaps(pure; WARN wiring deferred to 3a/3b). Seefelhom.eu/documentation/architecture/07-backup-architecture.md§3. - Offsite tier engine (v0.134.0, Task 3a — consumes the above):
internal/backup/offbox.gooffbox_capture.go+offbox_restore.go. Each toggled app's push is ONE multi-path restic snapshot = recovery unit + itsTierOffsitemandatory set (offboxCaptureSet); legacy/undeployed stay unit-only. Skipped/missing mandatory paths are loud gaps (English log + HungarianLastWarning) because restic 0.14.0 silently skips a missing source path (SP-3.4). Quota readsstats --mode raw-data(real repo bytes, SP-1); a pre-push gate blocks an ENLARGEMENT that would cross the soft quota (unit-only push continues;OffboxTarget.EnlargedBlocked; edge-triggered notify). Retentionforget --group-by host,tags(SP-2). Restore (RestoreOffboxScratch) scratches to a data drive off the rootfs (F-A1) behind a headroom gate; unit-only default via--includethe absolute unit path;PlaceOffsiteRestoremerges a full scratch into live viarsync --ignore-existing(never--delete), refusing on the puremapOffsiteRestorePathsguards. It restores FILES ONLY — no database, no stack restart — and the UI now says so.
- Truthful hub-managed empty state (v0.161.0, R-70): when
controller.yaml'soffsite.enabledis true but nooffboxtarget exists yet (the pre-apply window — or a burned one-time credential, DIAG-f10), the Távoli mentés status card AND the target empty-state line say „Felhom offsite tárhely kiépítve — a beállítás automatikus, folyamatban…" instead of „igényelhető szolgáltatás" / „Még nincs beállítva…". Data keyOffsiteHubEnabledfrombackupsOffboxData; own-NAS setup form unchanged. The hub side (v0.72.0) watches the same window from its end (delivery-state detector, stuck event, R-71c credential self-heal). - Coherent snapshot pairs (v0.148.0, R-44): every offsite run — manual AND nightly — refreshes
the DB/volume dumps and recovery units (
offsitePreDump→runDBDumpsInternal) BEFORE the restic capture, so each snapshot is an internally coherent{DB@T, files@T}bundle and retention is a history of restorable points. Order is the mechanism: the gap can only add files the DB does not reference yet, never remove one it does. Each manifest carriesoffsite_run_iddumps_at; a manifest without them is a pre-v0.148 pair of unknown skew, surfaced at restore time. The periodic refresh carries the prior stamp forward and never invents one. A dump-leg failure is a loud WARN that does NOT abort the push (data-first: a degraded backup beats none).
- A scratch restore NAMES ITS SCOPE (v0.198.0, R-204 item 3 —
restoreScratchOutcomeMsg).mode=unit(the default) restores the recovery unit only — the app's definition, configuration and DB dumps — becauseRestoreOffboxScratchpasses--include <unit path>; the userdata that is in the SAME snapshot is excluded by it. The outcome flash used to be one sentence for both modes and named neither scope, so on the last step of a disaster recovery the customer was told „visszaállítva" after the thing they were looking for had not been. The unit outcome now states what came back, that the customer's own files did NOT, and the step that gets them; the full outcome states that the files came with it (an absence is not a statement). The wizard's intent card 1 carries the same scope BEFORE the choice. Themode=fulltwo-step size gate is untouched, and the default staysunit— all three wizard forms setmodeexplicitly, so changing it would alter nothing the customer sees while silently changing a mode-less POST. - Offsite reconstitution (v0.148.0, R-43 —
offbox_reconstitute.go): the leg that was missing.ReconstituteFromOffsite(/backup/offbox/reconstitute, „Teljes visszaállítás (fájlok + adatbázis)") makes the live app equal to the chosen snapshot: safety dump → stop → files overwritten (rsyncRestoreOverwrite: no--ignore-existing, no--delete) → the snapshot's NAMED VOLUMES replayed (restoreDockerVolumesFrom, reading the SCRATCH unit — v0.218.0, R-354) → the DATABASE SERVICE ONLY started (StartStackServices, v0.153.0) → the snapshot's dump replayed (reimportDBDumpsFrom, reading the SCRATCH unit) → the full stack started → health wait. The volume leg did not exist before v0.218.0 — the archives live inside the recovery unit, whose placement is (correctly) skipped, so the off-site restore returned files and a database and silently nothing else. For the 40 of 53 catalogue apps that declare no data drive, that archive is the entire dataset. Volumes replay BEFORE the database, so a logical dump still wins over a volume-tar copy of the same database, and inside the stopped window because Docker will not replace a volume in use.restoreDockerVolumesFromis the local restore path's own replay with an explicit directory — one implementation, two callers. Two invariants: nothing is ever deleted (post-snapshot files survive as extras), and thepre-restore-safety dump is verified on disk BEFORE anything is stopped or overwritten — if it cannot be taken the operation refuses with zero changes. Safety dumps appear inListDumpFiles(they are the undo). The live recovery unit is still never overwritten, which is why the replay source is the scratch. Honesty surfaces (OffsiteScratchPair): dump age, an unstamped-pair warning, and the R-44 empty-dump sniff — all warn-level, none of them gates. - Where the pre-restore undo copy is written (v0.221.0–.1, R-361).
DumpOneTotakes an EXPLICIT final path;DumpOnekeeps its signature and calls it with the canonical<stack>-<dbtype>.sql. The safety dump asks forpre-restore-<stamp>-…directly — it used to dump to the canonical name and rename afterwards, which destroyed the app's own backup on every restore. The.tmpderives from the final path, so a nightly dump and a safety dump in the same directory cannot share a scratch file.db_dumpsin the manifest lists the app's own dumps only; the undo copies stay on disk and stay visible, capped at 3 per app, pruned from the capture side above the already-current early return. - What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380). A ladder, and every
rung is observable: replay → rollback → hold. The pre-restore undo copy has always been
taken; since v0.220.0 it is also put back when the replay fails — the whole set for this run,
matched on the run's own stamp (never on the
pre-restore-prefix, and never just the first file), re-applied with the DB service still up and before any restart, into a re-discovered container (the DB-only start re-creates it, so the captured id is dead by then — R-379, v0.220.1). The app then starts and the message says both that the restore failed and that the data is back. If the rollback ALSO fails the app is held stopped — the operator's ruling — the hold is persisted inSettings.RestoreHolds, every start path refuses it (customer button, app-stopRecover(), boot sweep, viadriveStartGateabove its driveless early return), the app-stop marker is ended so nothing auto-restarts it, and the row goes red. Cleared with--clear-restore-hold <app>, which requires a controller restart.--single-transactionon the Postgres import is a belt only; MariaDB DDL is not transactional, which is why the rollback is the fix. - Where an off-site restore puts the data (v0.219.0, R-356).
ReconstituteFromOffsiteandPlaceOffsiteRestoreresolve the destination withManager.GetAppDrivePath— the same resolverCaptureRecoveryUnitwrote the snapshot with: the app'sHDD_PATHif it declares one,systemDataPathotherwise. Before v0.219.0 both used the RAWGetStackHDDPathand read an empty answer as "the app is not installed", which refused 40 of the 53 catalogue apps permanently — those apps are correctly driveless — while telling the customer to reinstall a running app "in the same place" they are never offered. Three separate refusals now, three separate sentences: not deployed (isStackDeployed, asked ofListDeployedStacksand failing CLOSED on a nil provider); deployed but no resolvable data root; and the R-351 placement mismatch, which still requiresack_placement=1. The 13needs_hddapps are unchanged. The capture-side rawGetStackHDDPath(offbox_capture.go) is fenced and must keep no fallback — it resolves declareduserdata/importfile legs, which do not live on the system disk. - The restore wizard (v0.154.0, R-48 —
web/restore_wizard.go). The offsite restore controls used to render as up to five inline forms per app row, two of which — the missing-only merge and the true reconstitution — were sibling buttons whose difference is whether the data comes back. That mis-selection caused the round-2 incident. Each row now carries ONE entry linking to a per-app wizard; the three intents are cards with consequence sentences, and the dangerous one keeps the R-43 double-confirm verbatim.deriveWizardStepis pure — (op running, size-gate flash, scratch ready) → step + which intents unlock — and a running op outranks a stale?full_prep=, so no commit button survives into a restore. While ANY op runs, every mutation form is suppressed server-side rather than offered and refused. No new endpoint, no job registry (that stays R-45), and the page works with JavaScript disabled. v0.155.0 fix: the "is an op running" read must come fromRestoreStatus()(theopRunningdisplay flag, set synchronously byBeginRestoreOp), NOTManager.IsRunning()(the concurrency single-flight, whichRestoreOffboxScratchnever acquires — so v0.154.0's execution step was unreachable for the verification restore). The strip's highlight is its own derivedPhase, so a finished restore reads „Eredmény" while the intent step is available again; the outcome card is window-bounded and app-bound. - The DB-only replay window (v0.153.0, R-47). Until v0.153.0 the whole stack was started before
the replay, so the application's own schema management raced the dump: measured live on
2026-07-19 (H4), immich-server rebuilt
clip_indextwo seconds before the dump'sCREATE INDEXand the replay abortedalready existsunderON_ERROR_STOP=1. The DB service is now brought up alone (appbackup.DBServiceNamesreads the LIVE compose'sservices:map to name it), the dump is replayed with the app still down, and only then does the full start run. Fail-closed: a dump with no identifiable DB service refuses before the first mutation. Every exit from the window — replay failure, DB-only start failure — still does a best-effort full start, so a failed restore never leaves the box with a database and no application.
- Capture-set computation (v0.133.0, Task 3-core, INERT):
The /apps/{slug} page renders hero section, screenshots, setup guide, and optional config form.
Stack Operations
| Operation | What it does |
|---|---|
| Start | docker compose up -d — pre-start memory check rejects with 409 if insufficient RAM |
| Stop | docker compose stop (blocked for protected stacks) |
| Restart | docker compose restart |
| Update | docker compose pull + docker compose up -d |
| Remove | docker compose down --volumes + remove app.yaml + optional HDD/backup cleanup; template preserved for redeploy |
| Delete | docker compose down --rmi local --volumes + optional HDD data cleanup (orphaned stacks only) |
Remove vs Delete: "Eltávolítás" (Remove) is for deployed catalog stacks — it reverts the stack to "Nincs telepítve" state while keeping the template for easy redeployment. "Törlés" (Delete) is for orphaned stacks — it removes the entire stack directory including templates. Both require stopping the stack first.
Remove modal shows three sections: (1) always-removed items (Docker volumes, app.yaml, cross-drive schedule), (2) optional HDD data deletion with reimport warning, (3) optional backup data deletion (DB dumps + cross-drive rsync) with restic retention note.
Protected stacks (traefik, cloudflared, felhom-controller) cannot be stopped, removed, or deleted from the UI. Restart is allowed.
Orphan detection: Deployed stacks with no matching catalog template are marked as orphaned with an "Elavult" badge and can be safely deleted.
Base-infrastructure bring-up (stacks/infra.go + internal/infra/, v0.41.0)
Canonical import root (v0.172.0, R-75). ${IMPORT_PATH} = <system namespace root>/userdata/import
— ONE drop-zone per box, on the always-available system drive, never per data drive (each drop-zone app
has exactly one ingest bind, so a per-drive import/ would put a dead lookalike on every other drive,
and import/* is class: excluded so files stranded there are unbacked too). Injected at BOTH
compose-env builders; no per-drive fallback — unresolvable leaves it unset so compose fails loudly.
The system drive is deliberately NOT a registered StoragePath, so the FileBrowser bind
(/srv/beolvasas, sidebar „Beolvasás"), the skeleton and the system-owned beolvasas SMB share each
reach it explicitly. The userdata skeleton is catalog-derived (DeriveUserdataDirs +
UserdataSkeletonCarry, sorted — the sort is load-bearing, see REUSE.md) and can only ever ADD.
The controller stands up its own base stack — traefik (reverse proxy), cloudflared (external tunnel), filebrowser — instead of relying on the bare-metal scripts/docker-setup.sh (which a Proxmox-provisioned guest never runs). internal/infra renders the compose + config files from controller.yaml via embedded text/templates (lifted from docker-setup.sh); image tags are pinned constants there (TraefikImage/CloudflaredImage/FileBrowserImage) and the web FileBrowser sync path delegates to the same renderers, so the pinned versions can never diverge.
Manager.EnsureBaseStack() creates the traefik-public network, then deploys traefik → cloudflared → filebrowser under ${stacks_dir}/<name>. It is:
- single-flight (a
TryLockguard — it's called from both first boot and every health tick, so overlapping runs must not race on the same stack dir), - idempotent (skips a stack whose container is already running; never overwrites an existing filebrowser compose, preserving the storage mounts
SyncFileBrowserMountsmanages), - non-fatal (logs, never crashes the controller).
cloudflared is only deployed when a tunnel token is configured. Triggers: a first-boot goroutine (after stack init) and an unconditional call on every system-health tick (self-heal — cheap when healthy thanks to the idempotency). monitor.EffectiveProtected mirrors the cloudflared condition so a LAN-only node (no tunnel token) doesn't report a perpetual "protected container not running" FAIL.
Backend transports — self-signed HTTPS backends (ensureServersTransports → RenderServersTransports, v0.83.0). Traefik talks HTTP to app backends by default, which is correct for every catalog app that serves plain HTTP. The exception is an app that serves its own self-signed TLS on the internal docker bridge (the first is Crafty, HTTPS-only on :8443): Traefik must speak https to it and skip verifying a per-container self-signed cert (no CA to verify against; the hop never leaves the host). insecureSkipVerify is not settable via Docker labels in traefik v3 — it must live in static/file config — so EnsureBaseStack writes a file-provider dynamic file dynamic/serverstransports.yml defining a named transport insecure-skip-verify (write-if-changed; hot-loaded by the file watcher). An app opts in per-service via two catalog labels — loadbalancer.server.scheme=https + loadbalancer.serverstransport=insecure-skip-verify@file (the @file suffix is the cross-provider reference). Backend verification stays the default (ON) for every other service — there is deliberately no global insecureSkipVerify in traefik.yml. This write runs outside ensureTraefik (which early-returns when traefik is already up) so an established node still materializes the file on a self-heal tick.
Mount prerequisite (Section-G): the controller writes these stacks under
/opt/docker/stacksinside its container, butdocker compose upruns on the guest Docker daemon. The golden's controller-bootstrap (felhom-agentbuild-golden.sh) therefore bind-mounts that path same-path (-v /opt/docker/stacks:/opt/docker/stacks) so the daemon resolves every relative bind source — without it, all bind-mounted stacks (base infra and customer apps) silently break.
FileBrowser admin password (v0.243.0, R-513). FileBrowser used to accept admin/admin on every box. After ensureFileBrowser, every tick until decided, Manager.EnsureFileBrowserAdminPassword (internal/stacks/filebrowser_password.go) probes http://filebrowser:80/api/auth/login with admin/admin: 200 → a generated password:16 is set through FileBrowser's API (never the auth.adminPassword config key — measured to overwrite a hand-set password on every start), verified both ways, and stored AES-encrypted in settings.json (filebrowser_admin_state: generated); 401 → operator, never touched. The FileBrowser app page shows user admin and the password behind the R-254 reveal (POST /apps/filebrowser/initial-credentials/reveal), or „az üzemeltető állította be".
Controller routing + the wildcard cert anchor (wireController → RenderControllerRoute, v0.41.1 / v0.42.1). filebrowser self-registers with traefik via Docker labels + traefik-public membership baked into its compose; the controller can't (it's started by the golden bootstrap before traefik-public exists, and the v2 bootstrap.json carries no domain — that comes from the hub pull). So EnsureBaseStack wires the controller post-pull: it docker network connect traefik-public felhom-controller and writes a traefik file-provider route dynamic/controller.yml (Host(felhom.<domain>) → http://felhom-controller:8080, write-if-changed). When DNS-01 ACME is configured, that route is also the wildcard-cert anchor: its router-level tls.domains: *.<domain> makes traefik proactively obtain the wildcard *.<domain> + apex via Cloudflare DNS-01 at startup (an entrypoint-level http.tls.domains does not trigger issuance in traefik v3 — only a router-level tls.domains does). Every other router then serves that one real wildcard cert by SNI — no per-app certresolver labels. This is what lets a LAN client reach the box directly at *.<domain> with the real cert (the felhom-agent split-horizon resolver depends on it).
Missing Field Injection (deploy.go)
When app templates are updated (e.g., a new APP_KEY secret is added to .felhom.yml), existing deployed apps need the new field in their app.yaml. The controller handles this automatically:
- On startup:
InjectMissingFields()runs for all deployed stacks - After sync: the post-sync hook runs for stacks whose templates were updated
- For each deployed stack, compares
.felhom.ymldeploy_fieldsagainstapp.yamlenv vars - Missing
secretfields: auto-generated using the field's generator spec (password:N,hex:N,base64key:N) - Missing
domainfields: filled with the customer's configured domain - Missing
subdomainfields: filled with the field's default value or the.felhom.ymlsubdomain:metadata - Other field types (e.g.,
text,select): logged as warning for manual configuration - Locked fields are added to the locked list automatically
Generator types: password:N (alphanumeric), hex:N (hex-encoded random bytes), base64key:N (base64: + N random bytes base64-encoded, for Laravel APP_KEY etc.), static:VALUE (literal value).
Container State Display
| State | Color | Label | Meaning |
|---|---|---|---|
| Running + healthy | Green | "Fut" | All containers running and healthy |
| Running + starting | Orange | "Indulas..." | Healthcheck not yet passed |
| Deploying | Orange | "Telepítés..." | Compose up in progress (image pull, container creation) |
| Running + unhealthy | Yellow | "Nem egeszseges" | Docker or controller-side healthcheck failing |
| Degraded | Red | "Leallitva" (counts with stopped) | A SUPERVISED member of a multi-container app is dead — e.g. the app's database — while other members are still up. Counts as DOWN: raises the dead-app banner and app_start_failed (R-51, R-384) |
| Stopped/exited | Red | "Leallitva" | All containers stopped |
| Restarting | Yellow | "Ujrainditas..." | Restart loop; becomes down only after 5 minutes sustained (crash loop) |
| Not deployed | Gray | "Nincs telepitve" | Compose file exists, not deployed |
The order these are decided in is load-bearing (v0.222.0, R-384). degraded is evaluated
before unhealthy/starting/restarting. An app whose database dies drags its own front end
unhealthy seconds later — so if unhealthy were decided first (as it was until v0.222.0), the
symptom would mask the fault and the app would be silently down. A down member whose restart policy
is no/on-failure is a finished one-shot init/migrate container and stays benign. The full ladder,
including which states deliberately do NOT alarm and why, is documented in
felhom.eu/documentation/architecture/08-alarm-ladder.md.
Route-unpublished indicator (F5, v0.61.0). Traefik's Docker provider only publishes a route to a
container that is healthy (or has no healthcheck), so an unhealthy/restarting deployed app returns a
hard 404 at its URL even though the container is running. The routeUnpublished template helper
(funcmap.go) drives a distinct "URL nem elérhető – útvonal nincs publikálva" indicator on the dashboard
and stacks cards for such apps, so a dead URL isn't mistaken for a merely-degraded-but-reachable one.
Controller-side Health Probes (internal/stacks/healthprobe.go)
For apps that declare a healthcheck: section in .felhom.yml, the controller probes the container directly over the Docker network (both are on traefik-public). This complements Docker-level healthchecks and is the only health mechanism for distroless/scratch images that lack shell utilities.
Three probe types are supported:
http— Any HTTP response (even 4xx/5xx) = service is alive. Only connection refused/timeout = unhealthy.api— HTTP request with response validation (expected status code, body content). Fails if expectations aren't met.tcp— Simple port reachability check vianet.Dial.
Multiple checks per app are supported (all must pass). The probe scheduler runs every 10 seconds; per-app intervals default to 5 minutes and are configurable via healthcheck.interval in .felhom.yml. Probe results are stored in Stack.HealthProbe and exposed via the API. Failed probes override the stack state to StateUnhealthy; the override clears automatically when the next probe passes.
Fast initial probing: On start/restart, stale health probe results are cleared (so the stack doesn't immediately appear "unhealthy" from a previous result). Until the first healthy probe, the controller checks every 10 seconds instead of the normal 5-minute interval, giving fast feedback on whether the app came up successfully.
2. App Export/Import (.fab bundles)
Per-app export creates a self-contained .fab file (tar.gz, optionally encrypted) that can be stored externally or used to restore the app on the same server. Distinct from the automatic backup system — user-initiated, per-app, produces a single portable file.
Bundle contents: manifest.json + config/ (compose, .felhom.yml, app.yaml with plaintext secrets) + database/ (gzipped SQL dump) + data/ (HDD bind mount tars AND Docker named volume tars — additive since v0.130.0; a needs_hdd app bundles both).
Encryption: Optional AES-256-CTR + HMAC-SHA256 with scrypt key derivation (N=32768). Format: "FABE" magic header + salt + IV + encrypted tar.gz + HMAC tag. Streaming for multi-GB files.
Export flow: Estimate size → check free space → optionally stop app → copy config → dump DB → tar user data (docker volumes stream via docker cp through a stopped helper container — v0.125.0; never a docker run -v host mount, which strands data on containerized controllers) → fail-loud bundle assertion (every manifest-claimed tar must exist non-empty, AND a needs_hdd bundle claiming NO data at all is refused — v0.130.0 C6B-F1) → create tar.gz → optionally encrypt → atomic rename. App restarts automatically after export if it was stopped.
Mount discovery (v0.130.0, C6B-F1): stacks.ExportDataMounts — the ${HDD_PATH} binds unioned with the ${USERDATA_PATH} root (single userdata entry) when the compose uses the standard userdata convention; pre-fix the adapter was ${HDD_PATH}-only, so 12/13 needs_hdd catalog apps exported hollow (config-only) bundles. The root (not per-bind) keying is what round-trips through the import's basename→<HDD_PATH>/<subdir> mapping. A basename collision between mounts fails the export loudly. The share-removal endpoint also refuses while a deployed app's HDD_PATH is on the share (C6B-F2 guard).
Class-scoped export (v0.136.0, Task 4 — the SQ6 fix): for a classified app the userdata root tar is exclude-scoped — it keeps only dirs that are an ancestor-or-descendant of a SELECTED bind relpath (mandatory ∪ checked-optional ∪ opted-in-excluded; R1-C, the tier2Reconcile keep-rule), so sibling apps' content no longer rides along. No selected userdata bind ⇒ the root tar is skipped entirely (radarr → state-only). Non-selected HDD bind mounts are skipped; a mount matching no classified bind is kept (fail toward capture). Mechanics unchanged: ONE userdata tar, per-mount skip, manifest v1 + import untouched. The plan is pure (appexport/fabplan.go computeFabPlan over appbackup.ComputeFabBuckets); tarDirectoryExcluding prunes excluded subtrees in the walk. Legacy (no-block) apps export byte-identically to v0.130.0. Mandatory paths are a server-side floor (a client cannot deselect them). The export page shows the class selection UI (locked mandatory, pre-selected optional checkboxes, opt-in excluded behind the two-number warning + FileBrowser pointer); the estimate carries an additive class split (ExportEstimate.MandatoryItems/OptionalItems/ExcludedItems + BaseBytes).
Import flow: Decrypt if needed → extract → validate-before-destroy (v0.125.0: every manifest-claimed data tar must be present non-empty BEFORE the app is stopped or any volume removed — hollow bundles from containerized ≤0.124.0 exporters are refused with the app untouched) → prepare stack dir (create new or compose down --volumes for existing) → restore config (re-encrypt app.yaml with current server key) → restore user data (HDD tars; volumes populate via docker cp streaming) → restore DB (start DB service, wait for ready, import dump) → start full stack → refresh UI.
Architecture: internal/appexport/ package with ExportStackProvider adapter interface (same pattern as backup.StackDataProvider). exportAdapter in main.go bridges stacks.Manager to the provider.
API endpoints: /api/export/estimate, /api/export/start, /api/export/status, /api/export/bundles, /api/export/manifest, /api/export/import, /api/export/import/status; browser download (v0.124.0): /api/export/download/estimate, /api/export/download/start, GET /api/export/download?file=<basename>; browser upload (v0.128.0): POST /api/export/upload/{init,chunk,finalize,abort}.
Browser download (v0.124.0 — portability, NOT a backup tier): the same export pipeline runs with dest = <DataDir>/fab-downloads/ (same producer → byte-identical bundle), then streams via a guarded endpoint (basename-shape + dir-containment guard; Content-Disposition: attachment; io.Copy; the staged bundle is removed after the stream and a 1h TTL sweep runs on startup + each start). Estimate is shown BEFORE starting; the batch UI downloads apps one at a time (no combined archive). handler_export_download.go.
Browser upload (v0.128.0 — the download's return leg): the /import page uploads a .fab straight from the browser. Chunked because the Cloudflare tunnel caps request bodies at ~100 MB (probed live 2026-07-13: 120 MiB → edge 413, 80 MiB → origin): JS slices the file into 64 MiB chunks (strictly sequential offsets; one retry per chunk re-synced from the 409 received_bytes echo), the server streams each to a .part-<random> file in the DEFAULT drive's exports dir (io.Copy, 96 MiB per-request cap, free-space gate = size + 1 GiB), finalize checks the exact declared size, fsyncs and atomically renames (collision → lowest-free "name (N).fab"). Single-flight; no client-side hash (the .fab format self-validates at import); in-memory state — startup GC sweeps *.part-*, 15-min idle timeout aborts server-side. The scan + validation + import pipeline are untouched. handler_export_upload.go.
UI: Export button on app info page, the "Hordozható mentéscsomag (.fab)" section on /backups/restore (per-app download + batch), standalone import page at /import (upload zone + drive-scanned bundle list).
3. Backup System
The backup system implements a 3-2-1 backup architecture. Each tier is a complete, self-sufficient backup — any single tier can fully restore an app.
The restore carries the customer's own previous answers (v0.217.0, R-351).
internal/backup/offbox_placement.go. Every recovery unit's manifest.json records drive and
namespace_root, and its compose/app.yaml records SUBDOMAIN/DOMAIN. Until v0.217.0 nothing read
them back, so a restore into a destination different from the recorded one succeeded silently.
CheckPlacementcompares the recorded drive against where the restore is about to write, before the safety dump and before the first byte. A difference is named — both values — and refused. The customer may proceed deliberately withack_placement, a separate form field fromconfirm=1: one click must not carry two decisions.- An UNKNOWN recording is never a mismatch. A pre-field or unreadable manifest falls back to the previous behaviour rather than blocking, and is never rendered as an empty value.
- The not-installed refusal names where the data belonged, read from the prepared scratch.
- The deploy page prefills the address and data folder from the app's own backup
(
RecordedUnitForStack), labelled as coming from the backup and still editable — a memory, not a lock.RecordedAddress.Known()requires both halves: an absentSUBDOMAINwould otherwise surface the catalog's default as though the customer had chosen it. - Where the app's data will live is stated on the deploy page before the button is pressed (R-352). Measured 2026-08-21: 13 of 53 catalogue templates declare a storage field; the other 40 have none and their data goes to the system drive. Visibility only — no placement changed.
- Starting a restore is gated by
restoreOpBlocked(), which reads the display flag as well as the concurrency flag. Before v0.217.0 a second press started a second run and was told it had. - The off-site listing's per-app size calls run concurrently, bounded to 4. Measured before the change: 2605 ms + 5 × 2697 ms ≈ 16 s. The bound protects the Storage Box's session cap; a refused size call returns 0, which under-reports rather than fails visibly.
The reserve — per-app backup admission (v0.192.0 decision B2, widened by v0.193.0 / R-181).
internal/backup/admission.go. Since the mp1→mp0 merge (R-165) local backups and Docker's
data-root share one filesystem, so an unbounded backup write is a stopped box rather than a slow one.
Before any of an app's three local write legs runs — DB dump, volume dump, recovery-unit capture —
admitApp takes one verdict for that app for that run and the other two legs reuse it. A refused
app writes nothing at all, is not stopped, keeps its previous unit byte-identical, and produces
exactly one operator alert (recovery_unit_capture_failed, operator-tier).
- The verdict is lazy, not run-wide. It is taken at the app's first write, because app A's dump can put app B under the reserve; a verdict taken at run start would read a disk that no longer exists by the time B writes.
- It is never re-decided between an app's own legs, and the memo is reset per run.
- Two questions, both against two thresholds (97% used / 1 GiB free). Headroom: is the
filesystem already below the reserve? Size: would this app's own write take it below? The size
estimate is the app's previous
.sql+.taralready on disk. No history → headroom-only, deliberately — otherwise the first backup is the one that can never happen — and the alert says so. - The thresholds sit beyond
fillwatch's critical band (95% / 2 GiB), so the customer is always warned before a refusal is possible. - It refuses; it never deletes. Nothing here is generational — one unit per app at one fixed path — so "prune the oldest" could only destroy a different app's only local copy.
Sidebar behaviour (v0.146.0). Groups that own sub-pages — Tárhely, Biztonsági mentés,
Megosztás — render as accordions: the header is a real <button class="nav-group-toggle">
(keyboard- and AT-reachable for free) carrying a chevron, and exactly one group is open at a
time. The group containing the active page is rendered open server-side (.is-open in
layout.html), so the right group is already open before any JS runs and stays open if JS never
does; the vanilla listener in layout.html only handles clicks. Groups without sub-items
(Vezérlőpult, Alkalmazások, Rendszermonitor, Debug) are plain links, unchanged. Making the header a
button cost no reachability because every group's own landing page is also its first sub-item
(/storage → Meghajtók, /backups → Áttekintés, /sharing → Hálózati megosztás). Collapse uses
grid-template-rows: 0fr → 1fr rather than max-height, so it animates to the content's real
height with no magic number to drift as item counts change.
Page map (v0.124.0 IA split — sidebar children under Biztonsági mentés):
| Route | Page | Sections |
|---|---|---|
/backups |
Áttekintés | storage overview, whole-guest Rendszermentés (per tier since v0.243.0, R-517: newest success, a failed attempt under it, „nincs beállítva" for absent storage; „Naprakész" and the remote tick from successes only), status stat cards, single-copy warning, backup-target banner + offer (v0.186.0) |
/backups/remote |
Távoli mentés | Felhom-offsite status card (3 states, display-only), tier-3 status block + quota, participation toggles (+ zero-toggle hint; the persisted zero-toggle run-warning is DISPLAY-replaced by a "kijelölés módosult" note once ≥1 app is toggled — offboxWarningDisplay, v0.126.0), manual-target form (#offbox-section) |
/backups/apps |
Alkalmazások | schedule, Adatbázisok table, per-app 1./2./3. tier rows (tier-2 config entry; tier-3 actions deep-link to /backups/remote#offbox-section); since v0.242.0 a REMOVED app whose recovery unit was kept is listed after the deployed rows („Eltávolítva — visszaállítható") with one action, the unit restore that reinstalls it (R-487) — the list is keyed on the drives, not on what is deployed |
/backups/restore |
Visszaállítás | restore panel, offsite restore list (one „Visszaállítás…" entry per app since v0.154.0), existing verification copies, .fab download/import loop |
/backups/restore/app?name=<app> |
Visszaállítás — | R-48 per-app offsite restore wizard (v0.154.0). GET-only; three described intent cards (ellenőrzés / hiányzó fájlok / teljes visszaállítás), a visible phase strip, and a server-derived step. Adds NO mutation endpoint — every card posts to the pre-existing /backup/offbox/{restore,place,reconstitute} |
Backup-target banner + offer (E-2 · v0.186.0, R-114 + R-112). The /backups page renders the
whole-system backup-target state server-side, from the AGENT's view (never from our own intent flag).
Four outcomes, three of which the customer sees nothing for or one thing for:
| State | Renders |
|---|---|
| healthy — a real drive holds the target | nothing (no badge, no reassurance: a working box must look normal) |
degraded, never configured (local/unset) |
the system-disk copy + an offer control that POSTs /api/storage/backup-target/assign |
configured, drive absent (TargetAbsent) |
the absent-drive copy, no offer — the remedy is to reconnect that drive |
| unknown (agent unreachable / pre-R-82) | nothing — absence of an answer is not degradation |
degradedMessageFor is the single decision point for customer copy; backupTargetView returns nil
for the two silent states. The absent copy is verbatim the hub's backup_target_absent email so the
banner and the mail agree. The offer never auto-submits, and restart_required from assign is shown
rather than papered over with a self-restart (the agent deliberately does not restart itself).
Shared data builders: backupsCommonData (chrome + full-status + flash) + backupsOffboxData
(offbox target/toggles) in handlers.go; shared partials in templates/backups_shared.html.
(The v0.124.0 split was MOVE-only, gated one-shot by backups_split_move_check.py; the gate was
retired in v0.126.0 when the moved blocks were legitimately rewritten onto the shared row partial.)
| Tier | Contents | Location | Can fully restore? |
|---|---|---|---|
| 1. Nightly restic | DB + Config + User data | Same drive as app | Yes (not against drive failure) |
| 2. Cross-drive | DB + Config + User data | Different physical device | Yes |
| 3. Remote | Everything | Cloud / remote server | Future |
Key principles:
- User data backup is mandatory — every app with HDD bind mounts is included automatically. There is no per-app toggle.
- Each tier includes everything needed to restore: DB dumps, config, and user data. No tier depends on another tier's data.
- Tier 2 is configurable for ALL apps — not just apps with HDD data. Non-HDD apps
back up config + DB dumps to the secondary drive (small but protects against drive failure).
(v0.118.0, F6): the nightly tier-2 RUN (
RunAllTier2) now actually includes volume-only apps too — it used to skip every non-HDD app, leaving them a single controller-level copy onsys_drive. A sys_drive app's restore-point drive label is now clear ("Belső SSD (rendszer)"), never blank. - 3-2-1 on single vs multi drive (v0.118.0, F6): on a box with a second physical drive, tier-2 is
the off-drive copy. On a single-drive box (no off-drive target at all) there is genuinely only
ONE local copy —
FullBackupStatus.SingleCopyWarningsurfaces an honest Hungarian notice on the backup page ("Csak egy másolat készül…") instead of implying a 3-2-1 guarantee the box cannot keep. - The
AppBackupPrefs.Enabledfield in settings.json is legacy and not read by any code. - v2 layout + class-driven legs (v0.135.0, Task 3b):
backups/secondary/<stack>/is the v2 relpath-mirroring layout —.felhom-tier2-layoutmarker (written LAST) +recovery-unit/+hdd/<relpath>/+userdata/<relpath>/. For a classified app the appdata leg is theTierSecondarycapture set (per-bind mandatory + optional; excluded drops out —tier2_capture.go); legacy apps keep a byte-identical resolver set in the same layout. N>1 appdata dirs + nested binds are native (the old flat-appdata N>1 refusal is gone). First v2 run per app = delete-and-rebuild of the old flatappdata/+ a reconcile pass that prunes dest dirs a bind no longer covers; all removals go throughtier2SafeRemove(refuses anything outsidebackups/secondary/). The SSD fallback is a state-only tier (unit + mandatory; optional skipped, honest reason). NETWORK (NAS) storage is never a tier-2 target — pinned or auto (F-6C-1: rsync-ogunder root_squash → wrong-owner restore). Restore reads v2 behind the marker gate; a pre-v2 copy is refused.
Per-app Tier 2 contents by app type:
| App type | Tier 2 contents | Example |
|---|---|---|
| HDD + DB | Config + DB + User data | Immich, Paperless-ngx |
| HDD, no DB | Config + User data | — |
| Docker volumes + DB | Config + DB + Volume data | Tandoor |
| Docker volumes, no DB | Config + Volume data | Mealie (SQLite) |
| DB, no HDD/volumes | Config + DB | Vikunja |
| Config only | Config | Gokapi, Homepage |
Tier 1: Nightly Backup (mandatory, same drive)
The nightly backup has two phases that run sequentially. All paths are per-drive — each physical drive gets its own restic repo and per-app DB dump directories.
Customer-configurable backup window (v0.168.0). ONE setting on the backup page — "Mentési időablak kezdete" (start W, default "02:30") — drives every leg at FIXED, never-stored offsets so they can never be misordered: DB dump at W, tier-2 mirror at W+60m, off-box at W+105m (wrap-safe). The whole-guest (agent PBS/vzdump) cycle is gated to [W+2h, W+6h) with a safety valve (runs regardless once the last successful backup is older than cadence+24h, so a box only ever on outside its window never starves); manual "Mentés most" is never gated. A saved window fans out to the three daily legs via
scheduler.UpdateDailyand takes effect without a restart. Precedence: settings > controller.yamldb_dump_schedule> "02:30". Seeinternal/backupwindow.
Absent-storage tier skip (v0.243.0, R-518). A tier whose storage the agent (≥ 0.131.0) reports
absentis dropped from the manual and the scheduled run before anything is stopped (quiesce.skipAbsentTiers), logged, and reported once asbackup_tier_skipped.unknownor a legacy agent is never skipped.
OOM visibility (v0.243.0, R-514). The 30 s dead-app check also reads
State.OOMKilledfor running app containers (Manager.ScanOOMKilled); the dashboard shows „Memória elfogyott" and the hub getsapp_oomonce per container run. The controller does not restart the app.
Multi-tier whole-guest backup (v0.174.0, R-82 Slice B). The agent can serve SEVERAL whole-guest backup tiers with independent cadences — "local daily + PBS weekly" (agent >= v0.97.0,
GET /backup/tiers). The controller owns quiescing, so it reconciles them: it collects EVERY due tier up front and runs them inside ONE quiesce window — one stop, N sequential backups (vzdump holds a guest lock), one resume. Two cycles on the weekly night would mean two app outages for one night's work. The app stays quiesced until the LAST tier snapshots, so every tier is app-consistent; the consequence is that both-due-night downtime is (first tier's full backup) + (last tier's snapshot), which is why tiers run fast-first (the agent advertises primary/local first). A manual "Mentés most" covers every tier, due-ness ignored. The window gate's safety valve evaluates the OLDEST due tier, so a stale DR tier cannot be starved by a fresher local one. Against a pre-R-82 agent (/backup/tiers404s) the loop degrades to the single untargeted tier, logs it once, and still takes the backup — MinAgent is unchanged. Seeinternal/quiesce(tiers.go) andinternal/agentapi/backup_tiers.go.
Atomic dump writes (v0.118.0, CAMPAIGN-3 F7). BOTH dump paths are crash-safe: the DB dump (
dbdump.goDumpOne) and the Docker-volume dump (DumpAppVolumes) write to a.tmpsibling, fsync, thenos.Renameover the restore point ONLY on success. A mid-write failure (a NFS cut mid-tar, an EIO, a timeout) removes only the.tmpand leaves the last good.sql/.tarbyte-untouched — a tier-1 restore is replace-semantics, so an in-place write that got truncated to 0 bytes used to destroy the only restore point..tar.tmpfiles are invisible to the restore-point/stale scans and orphans are swept on the next run.
Stale-primary sweep (v0.118.0, F5). After each cycle,
pruneStalePrimaryDirsremoves an orphanedbackups/primary/<app>dir left on an OLD drive when an app'sHDD_PATHmoved to another drive. Guarded: only for a DEPLOYED app whose CURRENT drive differs from the dir's drive; never the current-drive dir (the live restore point) or an undeployed app's dir; strictly underbackups/primary/.
NAS backup locality (v0.118.0, CAMPAIGN-3 Part 4 — decision A)— SUPERSEDED by R-108 (v0.187.0, 2026-07-30). Decision A said a NAS-resident app's tier-1 artifacts live on the NAS itself (nas-media/backups/primary/<app>), beside the data, with the tier-2 cross-drive copy as the off-NAS mitigation. That locality is exactly what made abackups/tree reachable through FileBrowser's share-ROOT bind (download: true), and it is why architectural target D5 — app secrets in the local recovery unit — could not be adopted.An app's data namespace may no longer live on network storage at all (operator ruling 2026-07-30), so the case decision A described can no longer arise: no app on a NAS ⇒ no
backups/primary/on a NAS.settings.RefuseAsAppNamespaceis the single predicate; every placement surface consults it (deploy POST, per-app migrate, decommission-with-migrate). The NAS keeps its browse capability unchanged — the share-root:rslavebind is load-bearing for automount wake (R-67) and was deliberately NOT narrowed; scoping it is undefinable anyway, since apps on a share store at<share>/<app>and creating auserdata/layer would write Felhom convention onto a customer's own NAS.The NAS-outage window decision A warned about is therefore also gone: an app's tier-1 artifacts are always on a local drive now, because the app itself always is.
Drive layout (v0.26.0):
<drive>/
├── felhom-data/ ← all controller-managed data (namespace, v0.26.0+)
│ ├── appdata/<app>/ ← app user data
│ └── backups/
│ ├── primary/
│ │ ├── restic/ ← one restic repo per drive (all apps on this drive)
│ │ └── <app>/
│ │ ├── db-dumps/ ← per-app DB dump files
│ │ └── volume-dumps/ ← per-app Docker volume tars (v0.33.0)
│ └── secondary/
│ ├── restic/ ← secondary restic repo (cross-drive)
│ ├── _infra/ ← infra config mirror
│ └── <app>/rsync/ ← per-app rsync data
├── .felhom-infra-backup/ ← DR marker (stays at drive root for scanner)
├── Dokumentumok/ ← user files (not controller-managed)
└── media/ ← user files (not controller-managed)
Note (Model A — corrected in v0.52.0):
HDD_PATHinapp.yamlis the in-guest mount point (e.g.,/mnt/felhom-usb). Under slice-10 Model A the host agent binds<drive>/felhom-datadirectly onto that mount, so the in-guest mount already is thefelhom-datanamespace root. Neither the compose templates nor the path helpers add afelhom-datasegment for a drive-resident app: app data is${HDD_PATH}/appdata/<app>and backups${HDD_PATH}/backups/..., single-nested. Only the SSD-only system-data fallback (a bare root,inGuestDrive=false) appendsfelhom-data. SeeNamespaceRoot(drivePath, inGuestDrive)ininternal/appbackup/paths.go. Earlier catalog templates used${HDD_PATH}/felhom-data/appdata/<app>, which double-nested to.../felhom-data/felhom-data/...on a Model-A drive; v0.52.0 dropped that segment in the catalog and locks deploy↔backup path agreement withinternal/stacks/hddpath_agreement_test.go.
Path computation is centralized in backup/paths.go via the FelhomDataDir = "felhom-data" constant:
PrimaryResticRepoPath(drivePath)→<drive>/felhom-data/backups/primary/restic/AppDBDumpPath(drivePath, stackName)→<drive>/felhom-data/backups/primary/<stack>/db-dumps/AppVolumeDumpPath(drivePath, stackName)→<drive>/felhom-data/backups/primary/<stack>/volume-dumps/AppDataDir(drivePath, name)→<drive>/felhom-data/appdata/<name>/(final segment is the app's real appdata dir NAME, resolved viaAppDataDirNamesfrom compose binds — NOT always the stack name; F-S2)SecondaryResticRepoPath(drivePath)→<drive>/felhom-data/backups/secondary/restic/AppSecondaryRsyncPath(drivePath, stackName)→<drive>/felhom-data/backups/secondary/<stack>/rsync/SecondaryInfraPath(drivePath)→<drive>/felhom-data/backups/secondary/_infra/InfraBackupDir(mountPath)→<drive>/.felhom-infra-backup/(unchanged — stays at drive root for DR scanner)
⚠️ Stale: the restic/secondary helpers above (
PrimaryResticRepoPath,SecondaryResticRepoPath,AppSecondaryRsyncPath,SecondaryInfraPath) describe the pre-strip layout — restic/cross-drive was removed in slice 8C. This section is rewritten when Tier 2 (Phase 3) lands.
Per-app recovery unit (Phase 2, v0.53.x; secret model rewritten by D5, v0.188.0)
Each app's backups/primary/<app>/ is a self-contained, recreatable recovery unit:
backups/primary/<app>/
├── compose/ docker-compose.yml + .felhom.yml + app.yaml (0600 — CARRIES the portable secrets)
├── db-dumps/ app-consistent DB dump(s)
├── volume-dumps/ named-volume tars
└── manifest.json image pins, secret NAMES, data_key names, portable NAMES, checksums, secret_source
- The secret split (D5, schema 2, operator ruling 2026-07-30). The unit was secret-free until
v0.188.0, and that made "restore from the drive alone" false: the fast, local, customer-doable
Tier-1/2 restore secretly depended on the slow, operator-driven whole-guest restore, because a
data-encrypting key or a DB password absent from the guest cannot be regenerated without leaving the
restored data unreachable. Tier-1/2 now needs the drive and nothing else. What travels is decided
in ONE place,
stacks.PortableSecretEnvVars:- TRAVELS — every
type: secretfield (45 of 53 across the catalog): the declareddata_keys, the 18 DB/root passwords, and the internal signing/encryption secrets. Each of these either decrypts data sitting on the SAME drive or authenticates to a container on an internal compose network with no external listener, so possessing it adds nothing to possessing the drive — which is exactly D2's argument for keeping the DATA plaintext. Written into the unit's app.yaml at 0600, plaintext, like the data beside it. - WITHHELD — every
type: passwordfield (7 admin/UI logins) plus thenonPortableSecretsregister (vaultwarden/ADMIN_TOKEN, whose/adminpanel is on the app's public web port). These authenticate against published services, so their blast radius is NOT bounded by the drive. They stay in the guest and are regenerated on restore (O4). Excluding this class is what licenses the plaintext ruling — the two are coupled and must not be relaxed independently. - The register is code, not a catalog flag, deliberately: a security boundary a catalog push can
silently move is not a boundary (cf. R-97a). Adding an app whose
type: secretfield gates an internet-reachable login means adding a row there.
- TRAVELS — every
- Fail-closed is unchanged. A
data_keymissing from both the unit and the guest still refuses the restore outright (never generated). D5 makes the key normally present; "normally" is not a reason to soften the gate. - Precedence: the UNIT WINS over the guest when both hold a value. Not "newest wins" — the unit's secrets are captured in the same run as the dumps beside them, so the unit's value is the one that MATCHES THE DATA BEING RESTORED, while the guest's is merely the most recent. A rotated data key does not decrypt data encrypted with the old one, and a rotated DB password does not match the hash inside the restored data directory. Pinned in both directions.
- Resettable secrets (O4, v0.99.0) — now the rare path, since the portable class comes from the
unit. An unrecoverable withheld secret gets a generated replacement from its catalog
generatespec (stacks.GenerateSecretForFieldvia thebackup.SetSecretGeneratorseam) rather than redeploying blank; the value persists encrypted throughRecreateStackDefinitionFromUnit→SaveAppConfig. ⚠️ R-127: a regenerated database password is NOT harmless —POSTGRES_PASSWORDis ignored once PGDATA is non-empty, so the restored data dir keeps the old role hash and the app cannot authenticate against its own rows, while the dump replay (local trust socket) still reports success. The WARN says so. - Helpers:
RecoveryUnitPath/RecoveryUnitComposePath/RecoveryUnitManifestPath(internal/appbackup/paths.go). Capture:Manager.CaptureRecoveryUnit(internal/backup/recovery_unit.go), run from the daily DB dump and the periodicRefreshCache(idempotent checksum-skip); the split itself is inbuildUnitAppYaml. The env + portable values come fromStackDataProvider.GetStackRecoveryInfo, which keepsNonSecretEnvand the secret set disjoint by construction.data_keyfields are marked in.felhom.yml(DeployField.DataKey). - A schema-1 (pre-D5) unit carries no secrets and still restores from the guest — the restore degrades rather than failing, and the next capture rewrites the unit (the app.yaml checksum changes).
- Consequence for the other tiers: the unit is copied by Tier 2 (another customer drive, plaintext,
same reasoning) and pushed offsite by restic (
offbox.go— encrypted at rest under the customer-owned repo password). Neither tier's code changed; the secrets simply travel with the unit they already carried. - Restore replays the DB dump (F17, v0.61.0; re-sequenced v0.153.0, R-47).
RestoreFromRecoveryUnit(and theRestoreAppfallback) stops the app → restores named-volume tars → recreates the compose definition and persists the recovered env (RecreateStackDefinitionFromUnit— starts nothing) → starts the DATABASE SERVICE ONLY (StartStackServices, named from the unit's compose) → replays eachdb-dumps/*.sqlinto that DB → starts the full stack → health wait. Before v0.153.0RecreateStackFromUnitended in a fullcompose up -d, so this path carried the same H4 race as the offsite one (see the reconstitution section above), with the same fail-closed rule and the same guarantee that every exit still brings the app back up. The replay itself usesbackup.reimportDBDumps→appbackup.ImportDump(psql / mariadb client, using the live container's own discovered credentials). The DB replay runs AFTER the volume restore, so the logical SQL dump wins over any volume-tar copy of the database (the dumps use DROP/CREATE —pg_dump --clean --if-exists,mariadb-dumpdefault--add-drop-table— so replay is idempotent). Volume-restore and DB-import failures now surface (restore returns an error) instead of a swallowed WARN. Prior to v0.61.0 the per-app restore never replayed the.sql, so DB-resident data did not come back. - The restore now STATES what came back (v0.226.0, R-353).
RestoreFromRecoveryUnitreturns(UnitRestoreResult, error)— volumes replayed, DBs replayed, and what the manifest LISTED — andunitRestoreOutcomeMsg(internal/web/handlers.go) turns that into the customer's sentence. It replaces<app> visszaállítva (<snapshot>)., which was equally true of a run that returned an entire dataset and one that returned nothing; on 2026-08-21 it was printed over a unit holding onlymanifest.jsonandcompose/. Three cases, three sentences: data returned (named and counted); nothing returned and the unit listed nothing („ez a mentés csak a beállításokat tartalmazta"); nothing returned though the unit listed dumps („…de egyik sem állt vissza. Az adataid változatlanok maradtak."). Every one is a claim about the BACKUP, never about the app — see CONTEXT.md's ruling and 07-backup-architecture §6.3.
Off-site content proof (v0.231.0, R-87)
What it is: every night the box restores ONE app's newest off-site snapshot into a throwaway folder, asks whether that backup still contains the app's actual data, records which snapshot it proved, and deletes the copy.
The question it answers, and why the integrity check beside it cannot. restic check proves the
stored bytes are the bytes we stored. It cannot tell us we stored the wrong thing. A hollow
recovery unit — no database dump, no volume tar — backs up cleanly, checks cleanly at 100 % depth,
restores cleanly, and gives the customer nothing back. Measured on demo-hp 2026-08-31 (R-403):
120 082 104 B became 7 036 B in one nightly run, recorded as a success.
WHAT IT DOES NOT PROVE, said plainly because the green tick invites the other reading: it does not prove a restore puts data back into a running app. It restores to a throwaway folder, looks, and deletes. It never touches a live app. Putting data back is drill work, and
07-backup-architecture.md§8 matrix row 4 does not move on the strength of this job.
| job | offsite-proof, sched.Daily at 05:30 |
| cadence | per SNAPSHOT, not per clock (R-86's model) — an app is due when its newest snapshot's ID differs from the ID last proved for it. One app per run; eight apps are covered in eight nights, and a NEW backup makes an app due again immediately |
| acceptance rule | two parts, and both are needed. (1) everything the manifest declares is present in the restored unit, AND (2) the manifest declares what the app is supposed to have. Part 1 alone passes a hollow unit, which is the shape this exists to catch |
| where the expectation comes from | the unit's own compose/docker-compose.yml, never the live box — the snapshot may predate the app's current shape. Database: DBServiceNames (the same discriminator the restore path uses). Volumes: ParseComposeNamedVolumes, as an existence check, not a name match |
| outcomes | three: pass, fail (readable and empty), and cannot judge. An app that legitimately has no database and no named volumes passes |
| repository writes | none. --no-lock, no unlockStale, and the exec seam rather than resticStep, so the unlock --remove-all escalation is unreachable. Asserted on the argv as a non-effect |
| guard | takes the single-writer flag itself and SKIPS rather than waits. Since v0.232.0 every off-site entry point takes it — RestoreOffboxScratch, OffboxRestorePrepareFull, RestoreSharesScratch and RestoreOffbox were all missing it (R-411/R-408), and the invariant is now pinned by an AST walk rather than asserted in a comment |
| scratch | backups/offsite-proof/<app> — a separate root from the customer's backups/offsite-restore/, so the nightly delete can never reach a copy the customer made, and a proof copy can never be offered for placement. On a box with NO registered data drive it falls back to the system data path (v0.232.0, R-414), because that is where a driveless app's unit already lives; the customer's FULL restore does not fall back and still refuses |
| if it cannot run at all | it records cannot_run, not silence (v0.232.0, R-414). A proof that never started is a standing property of the machine, not a transient failure, so it is written where the hub can read it — last_proof_result is never ABSENT, because absent already means a controller too old to have the feature. Per-snapshot due-ness is NOT advanced, so the app is retried once a drive is registered |
| timeout | 10 min (proofRestoreTimeout) — ~150x the slowest single app measured |
| cost, measured on demo-hp | one app 2.3–4.0 s, all eight back to back 25 s, peak scratch = that app's logical size (213 MB largest). The weekly check beside it takes 40.3 s |
| result | persisted on settings.OffboxTarget (proved_snapshots, last_proof_*) and published on OffboxReportStatus. last_proof_result absent = NOT RECORDED (a pre-0.231.0 controller), never "failed" |
Notifications. A failure emits one offsite_proof_empty, severity error, operator-only. A
pass, a skip, a "cannot judge" and a restore error emit nothing — a nightly success mail is how
people stop reading their alerts, and alarming on our own blind spot trains the operator to discount
the one alarm that matters.
⚠ IT IS DELIBERATELY NOT
backup_integrity_failed. That type means the store is damaged and carries a hub-side Hungarian template saying so. Here the store is sound and the CONTENT is absent — a different cause and a different action. The message says the backup is readable and does not contain the app's data, and explicitly that the store is not damaged.
restic restore --verifyis NOT used as a correctness check and must not be. Measured ondemo-hp2026-08-31: a byte changed in place in a restored 160 MB tar, with size and mtime preserved, passed clean; verify took 131 ms on a 213 MB tree, which cannot be hashing. It is a size-and-mtime reconciliation.
Off-site integrity check (v0.227.0, R-359/R-397)
What it is: a restic check against the off-site repository, run by the controller itself. Until
v0.227.0 nothing verified that the off-site copies were readable — the whole-guest tier had verify
jobs, the tier holding the customer's documents and photos had none.
| job | offsite-integrity, sched.Daily at 06:00 |
| cadence | due-ness, not a weekday — runs when the last SUCCESSFUL check is older than monitoring.integrity.max_age_days (default 7). A box switched off on its check day is checked the next day it is on |
| depth | --read-data-subset=100% by default since v0.228.0 (R-399) — the check re-reads and re-hashes every stored byte, not just the catalogue. monitoring.integrity.read_data_subset: absent or empty = the default; off (any case) = structure and index only; any form restic accepts (10%, 1/7, 50M) = itself. A malformed value WARNs and falls back to the default, never to off — a typo must not quietly remove the protection |
| why full depth | the structure check does not detect a size-preserving pack corruption. Measured on demo-hp 2026-08-30: a pack was damaged without changing its size, plain restic check reported no errors were found and exited clean, every read-data form caught it. Cost on that 134 MB store: 35.0 s structure vs 39.2 s at 100% |
| slow notice | a completed check over integritySlowNoticeThreshold (5 min) logs a WARN naming the duration, the depth and R-401. Operator log only — no hub event, no customer alarm, and it never changes the depth by itself. A skip or an unreachable store never warns: neither has a duration to judge |
| guard | takes the single-writer flag and SKIPS rather than waits |
| timeout | 30 min (integrityCheckTimeout) — bounds a hung repository so it cannot pin the flag |
| by hand | POST /api/debug/backup/integrity — same code path, due-ness ignored, every other guard intact |
| result | persisted on settings.OffboxTarget (last_integrity_check, last_integrity_ok, last_integrity_depth — v0.228.0) and published on OffboxReportStatus. Depth empty = NOT RECORDED (a pre-0.228.0 controller), never "structure" |
Three outcomes, not two. Skipped (a sibling operation held the flag), Unreachable (the repo
could not be opened, or the check timed out) and failed are different facts. Only a failure notifies;
a skip and an unreachable repository do not advance due-ness, so tomorrow tries again. A failure
does advance it — re-checking a broken store nightly is load with no new information.
Notifications. backup_integrity_ok is severity info, which severityNotifies drops — it mails
nobody, by design. backup_integrity_failed is error and reaches the operator; the customer leg is
switchable and OFF by default. The customer gets a sentence; restic's output goes to the log,
truncated.
⚠ THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION, and this is the thing to know before trusting it. Measured on
demo-hp2026-08-30 against a throwaway repo whose pack was corrupted without changing its size:restic checkreturnedno errors were found, exit 0; every--read-data*form returnedPack ID does not match …and exit 1. The structure check verifies the index, the pack inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots — but it does not re-hash pack contents. Choosing the read-data depth is R-399, and the cost curve is measured: structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s on a 134.3 MB / 67-snapshot store. Those figures do not extrapolate: the structure check's cost tracks the index, read-data's tracks the data.
The restore record survives a restart (v0.246.0, R-550)
The restore op-status (internal/backup/opstatus.go, served at GET /api/backup/restore-status) was
in memory only. Chaos night round 10 hard-reset a box four seconds into a restore; afterwards the
status was the Go zero value and nothing told the household whether the restore finished. Operator
ruling 2026-09-17 reversed the in-memory choice for the restore record only (notification cooldowns
stay in memory).
restore-status.jsoninDataDir(besidesettings.json), written atomically at both ends of an op.- At startup (
loadRestoreRecordAtStartup, before any page is served) a record still marked running becomes a failed result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — and a per-app notice that stays until that app's next restore. restore_interrupted(warning, for the household; hub v0.117.0) is pushed once per interruption, after the notifier exists (reportInterruptedRestore). Best-effort like everyPushEvent(3 attempts, 3 s apart); the page notice does not depend on it./backups/restoreshows a „Megszakadt visszaállítás" card per interrupted app; the off-site wizard's „Eredmény" card falls back to the interrupted record („Észlelve: …").
Restore refusals (v0.226.0)
Three guards added on the off-site restore surface, all server-side:
| guard | where | refuses when |
|---|---|---|
| Free space, R-357 | ReconstituteFromOffsite, before mapOffsiteRestorePaths / writeSafetyDump / StopStack |
the live namespace has less free than the scratch's size. Same wording as the two non-destructive gates (offsiteNoSpaceMsgFmt). No headroom multiplier — this copies a measured tree, not a predicted download. Fail-closed when either probe reads ≤ 0. The app is never stopped for a refused restore. |
| Incomplete scratch, R-358 | offboxPlaceHandler AND offboxReconstituteHandler |
OffboxFullScratchReady is false — no .felhom-restore-complete.json, unreadable, wrong schema, or full:false. „A visszaállítási másolat nem teljes…". The wizard flags control a button; these control the operation. |
| Restore in flight, R-360 | offboxVerifyCopyDeleteHandler |
any backup or restore op is running (restoreOpBlocked(), not IsRunning()). Previously it refused only during a backup, so the copy a restore was writing into could be deleted from the UI. |
The scratch completion marker (.felhom-restore-complete.json, 0600, atomic) is written by
RestoreOffboxScratch only after restic returns nil, and any stale one is cleared before restic starts.
It carries full, so a unit-only verification restore — which writes the same directory — can no
longer unlock the full-restore actions (R-396).
Tier 2 — off-drive copy (Phase 3, v0.55.x)
For every HDD app, Tier 2 (internal/backup/tier2.go) rsync-mirrors the recovery unit
(backups/primary/<app>/) + the app's resolved appdata/<name>/ to <target>/backups/secondary/<app>/
on a different physical disk — the only off-drive protection bind-mounted HDD app data can get (PBS
can't reach bind mounts). The appdata dir NAME is derived from the app's compose ${HDD_PATH} binds,
not assumed to be the stack name (F-S2, v0.131.0: paperless-ngx writes appdata/paperless;
tier2AppDataName → appbackup.AppDataDirNames); an app resolving to >1 distinct appdata dir is
refused loudly. This copies the recovery unit + appdata/<name> ONLY — not the browsable
userdata/ tree (F-S1, owned by the backup-classification redesign) and not the namespace
wholesale. Auto-targeted: prefer another registered user-data drive (off-disk via
system.SamePhysicalDevice); else the internal SSD for small units only, behind a size-aware
rootfs-headroom guard (tier2FitsHeadroom) that refuses rather than fills the ~8 GB guest rootfs
(reserve = max(2 GB, 20%)), recording an honest "needs a 2nd HDD" status. Status persists via
settings.CrossDriveBackup and drives the "2. mentés" card. Runs daily (tier2-backup, 03:30) or via
POST /api/backup/tier2. restic is not used — a plain browsable mirror.
In-place file restore from the Tier-2 copy (C2, v0.100.0 — closes drill finding F2) —
POST /backup/tier2/restore (backup.RestoreTier2Files, internal/backup/tier2_restore.go) + the
"Fájlok visszaállítása" button on the healthy Tier-2 layer row. Additive-only semantics
(rsyncRestoreMissing: rsync -a --ignore-existing): files missing from the live resolved
appdata/<name> dir (F-S2 — compose-derived, not the stack name)
are copied back from the RECORDED Tier-2 copy; existing live files are never overwritten (a
customer edit after the last copy wins) and nothing is ever deleted — this exactly serves the
"I deleted my files" scenario with zero risk to newer data. Source = the recorded
CrossDriveBackup.DestinationPath (never a fresh target selection). Single-flight with
backup/restore; refusals (no copy / never ran / copy dir gone / either drive disconnected /
decommissioned) happen before the app is stopped, with customer-readable Hungarian reasons;
stop → copy → start → health-wait. Out of scope by design: overwrite/point-in-time restore (offbox +
operator paths) and per-file selection. Apps that index their data dir (e.g. Nextcloud) may need a
rescan (occ files:scan) before restored files appear in their own UI.
COVERAGE — read this before assuming an app is protected by this button (C9-F1, v0.183.0). This restore reads
hdd/anduserdata/only. It does not readrecovery-unit/, which every Tier-2 run also writes and which holds the app's DB dumps and named-volume tarballs. Counted at catalogue459766cb1639by running the production rule over all 53 templates (v0.229.0): 45 apps have no readable subtree at all (their data is entirely in named volumes — BookStack, Docmost, Vaultwarden, Gitea, …), 7 have file legs but never their database or volumes, and 1 (bentopdf) is stateless. So this button is a guaranteed no-op for 45 of 53 apps and only ever partial for the rest.Since v0.183.0 it is HONEST about that instead of silently reporting success:
Tier2RestoreCoverageis consulted before anything starts, an app with no readable subtree is refused without being stopped and told which action does work, and a run that does proceed claims only what it examined („Minden vizsgált fájl megvan a helyén.") plus a disclosure that the database and internal volumes are not part of this restore.Since v0.229.0 the action that covers those apps is on the SAME row — see below. C9-F1b / R-103 and C9-F4 / R-102 are CLOSED. The refusal no longer sends anyone to another page: where the copy holds an openable unit it names „Teljes visszaállítás a másolatból", the button beside it.
Full restore FROM THE SECOND DRIVE's mirror (R-102 + R-103, v0.229.0) —
POST /backup/tier2/unit-restore (backup.RestoreTier2Unit → RestoreFromRecoveryUnitAt,
internal/backup/tier2_restore.go) + the "Teljes visszaállítás a másolatból" button on the Tier-2
layer row, in btn-danger-outline beside the additive one.
Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on
every run. Until v0.229.0 nothing read it, because every reader of a unit could only name a path
under backups/primary/ — so in the one failure Tier-2 exists for (the primary drive is lost, and the
primary unit with it) the surviving copy was unreachable. RestoreFromRecoveryUnitAt takes the unit
DIRECTORY, so the same restore that always worked from the primary now works from anywhere.
- The source moves; the destination does not. Data lands in the live Docker volumes and the live database container exactly as before; only the read path changes.
- It OVERWRITES, unlike the additive button beside it. The two are separate buttons because they are separate promises, and the confirm carries the difference in words and names the copy's date — differently when that date is only an ATTEMPT and not a proven copy (R-101).
- Fail-closed: the mirror must carry a parseable
manifest.json. Arecovery-unit/directory that exists is not a package, and a restore armed over one would stop the app and replay nothing. - Two predicates, not one wider one:
Tier2Coverage.CanRestoreUnit()gates this action;CanRestore()still gates only the file restore. Merging them would be R-356 again. - Refusals (no recorded copy, drive disconnected, pre-v2 layout, no openable unit) all happen before
the app is stopped; a second press is refused by
restoreOpBlocked()(R-351b). - Proven live on
demo-hp2026-08-31 with the primary unit moved aside — 3 volumes of 3, 1 database of 1, 28.65 s, an accented filename byte-identical, andsecrets recovered=2/2with the guest'sapp.yamlalso moved aside:felhom.eu/documentation/audits/DRILL-r102-tier2-unit-2026-08-31/. - Since v0.230.0 it also REFILLS the app's own drive before returning — see the R-403 note below.
The nightly copy refuses to replace a complete package with an empty one (R-403, v0.230.0) —
internal/backup/r403_hollow.go + the precondition in RunTier2.
The measurement, because this was run before it was fixed. On the shipped v0.229.0, on demo-hp:
an app's Tier-2 copy went from 120 082 104 B (4 database dumps + 3 volume tars) to 7 036 B (none of
either) in one nightly run, reported as a success. RunTier2 guarded the unit leg with os.Stat
alone, rsyncMirror is rsync -a --delete, and nothing compared the two sides — and an empty
recovery unit is a folder that exists. Evidence:
felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/.
- The predicate asks the MANIFEST, never the byte size.
unitCarriesDatais true when the unit's manifest lists a database dump or a volume tar. Absent or unparseable manifest ⇒ hollow, fail closed. - The refusal is one shape only: source hollow AND destination not. complete→complete,
complete→hollow and hollow→hollow all mirror as before.
--deletestays and the data legs are untouched — §8 row 5's derived-copy rebuild is a design decision and a copy that legitimately shrinks still shrinks. - The other legs still run and the run is not failed; a preserved package must not cost the customer their file legs or raise a red alarm on a healthy box.
- What the surfaces say. A preserved package is older than the run that preserved it, so the
per-app card carries „A másolat adatcsomagja régebbi, mint a legutóbbi mentés…", and the
„Teljes visszaállítás a másolatból" confirm names the package's own date (from the mirrored
manifest's
created_at) plus aFIGYELEMclause saying why. The restore outcome names the same date. Only an app whose leg was actually preserved shows any of it — a warning that fires on everything costs the same as the comforting lie it replaces. - The cause is closed too:
RestoreTier2Unitrefills an absent or hollow primary unit from the mirror inside the call, because the hollow manifest was written two seconds later by the 5-minute capture job. Never over a complete primary, never after a failed restore, and the capture itself is not guarded — it describes reality, and with the primary refilled there is nothing hollow left to describe.
Per-app Tier-2 config panel (v0.57.0) — GET/POST /stacks/{name}/backup
(internal/web/tier2_config_handler.go + templates/tier2_config.html). The "2. mentés" row's
Beállítás button links here (was the dead-end deploy page). Shows the effective off-drive target
(pinned or auto), whether it's the size-limited internal SSD, the last-run reason, and lets the customer
pin a registered drive (off physical disk) or toggle Tier 2 off. Always visible — single-SSD apps
get the "csak DB/konfiguráció" note, non-HDD apps the "already in the PBS whole-guest snapshot" context.
Two preference fields on CrossDriveBackup — UserDisabled + PreferredTarget (set via
Settings.SetTier2Preference) — are preserved across the runner's status writes (withTier2Prefs):
selectTier2Target honors a valid pin before auto-picking; RunTier2 skips a disabled app. The runner
re-validates the pin off-disk at run time. Manager.Tier2Info(stackName) is the read-only panel view
(effective target + eligible alternative drives).
Phase 1 — Database Dumps (internal/backup/dbdump.go, scheduled 02:30)
- Auto-discovery of PostgreSQL and MariaDB containers via
docker ps+docker inspect - Dumps via
docker exec pg_dump/docker exec mariadb-dumpwith 5-minute timeout - Dumps are written to the app's home drive:
AppDBDumpPath(appDrive, stackName) - Atomic writes (
.tmp→.sql) to prevent corruption - Validation after each dump: checks file size, header presence, counts
CREATE TABLE - Results cached in
settings.jsonsurviving container restarts
Phase 1b — Docker Volume Dumps (internal/backup/backup.go runVolumeDumps, part of the same run)
- F3 (v0.99.0): re-wired into the nightly/manual app-data backup run (
runDBDumpsInternal) — after the restic removalDumpAppVolumesSafehad no caller, sovolume-dumps/was never produced. Runs AFTER the DB dumps and BEFOREcaptureAllRecoveryUnitsso the manifests enumerate fresh tars. - Gate order (load-bearing): protected-stack (
cfg.IsProtectedStack) and has-volumes (GetDockerVolumes()) checks come BEFOREDumpAppVolumesSafe— the Safe variant stops the stack before its own volume check, so unconditional calls would bounce every volume-less app nightly. Disconnected/decommissioned drives skip with the same summary style as the DB loop. - Each volume-bearing stack is stopped before dump, restarted after (
DumpAppVolumesSafe()) — prevents inconsistent tars of live databases. - For each volume:
docker run --rm -v <vol>:/vol:ro -v <dumpDir>:/out alpine tar cf /out/<vol>.tar -C /vol . - 10-minute timeout per volume; a per-stack failure lands in the run summary as
FAIL <app> volumes:, flips the run's Success flag and fails the run (no silent partial) — other stacks still proceed - Stale tars cleaned up (volumes that no longer exist)
- Volume names resolved with project prefix via
ResolveDockerVolumeNames()(e.g.,mealie_mealie_data) - Dumps written to
AppVolumeDumpPath(nsRoot, stackName)
Phase 2 — Restic Snapshot (internal/backup/restic.go, scheduled 03:00)
- Apps are grouped by drive via
groupStacksByDrive()— each drive's apps are backed up to that drive's restic repo - App drive resolution:
GetStackHDDPath()(fromStackDataProvider) → falls back toSystemDataPath - Auto-generated repository password (32 random bytes, base64url), shared across all repos, synced to hub
- Paths included in each per-drive snapshot (v0.34.0: per-app scoped):
- Per-app DB dump dirs on that drive
- Per-app Docker volume dump dirs (
volume-dumps/*.tar) - Per-app HDD mount paths (user data)
- Per-app stack config dir (
<StacksDir>/<stackName>/— only for stacks on this drive) controller.yaml— only on the system drive (not duplicated across all drives)
- Auto-detects and unlocks stale locks (restic repo lock)
- Weekly prune on Sundays with configurable retention (keep-daily, keep-weekly, keep-monthly)
- Weekly integrity check (
restic check) on Sunday 04:00 — checks all primary repos
Protects against: accidental deletion, data corruption, point-in-time rollback. Does NOT protect against drive failure (backup is on the same physical drive).
Tier 2: Cross-Drive Backup (opt-in, different device) (internal/backup/crossdrive.go)
Complete backup to a different physical drive. Available for all apps — apps with HDD data back up config + DB + user data + Docker volumes; apps without HDD back up config + DB dumps + Docker volumes.
-
Auto-enable for small apps (v0.14.1): Apps without HDD mounts (config-only, DB-only) are automatically configured for daily rsync Tier 2 when ≥2 storage paths are registered.
AutoEnableSmallApps()runs at the start of each nightly backup cycle. Never overwrites existing user-configured cross-drive settings (even disabled ones). -
Infrastructure config backup (v0.14.1):
syncInfraConfig()rsyncs the stacks directory andcontroller.yamlto<dest>/backups/secondary/_infra/on every secondary destination drive. Runs before per-app backups. Cross-drive restic also includes infra paths. -
Two methods:
- rsync — Simple mirror with
--delete(fast, no versioning, browsable on disk) - restic — Versioned, deduplicated, encrypted (shared repo across apps, not browsable)
- rsync — Simple mirror with
-
Per-app configuration in settings.json: destination path, method, schedule (daily/weekly/manual)
-
Pre-backup DB dump:
DumpStackDB()runs fresh pg_dump/mariadb-dump before each cross-drive backup; non-fatal on failure (wired viaDBDumperinterface to avoid circular imports) -
Pre-backup volume dump (v0.33.0, safe stop/start v0.34.0):
DumpAppVolumesSafe()stops the stack, exports Docker named volumes to tar, restarts — wired viaVolumeDumperinterface -
Empty mounts allowed:
RunAppBackupaccepts apps with no HDD mounts — the rsync mount loop simply doesn't execute, but DB + config copy still runs -
Drive-type-aware validation (
ValidateDestination):Destination type Space checks External mount (different device than /)Block if <100 MB free System drive (same device as /)Require ≥10 GB free AND <90% used; logged warning -
Secondary drive layout (v0.14.1):
<dest-drive>/backups/secondary/ ├── _infra/ ← infrastructure config mirror (v0.14.1) │ ├── controller.yaml │ └── stacks/ ← full stacks dir (all app configs) ├── <app>/rsync/ ← per-app rsync mirror │ ├── _db/ ← DB dump files │ ├── _config/ ← compose.yml, app.yaml, .felhom.yml │ ├── _volumes/ ← Docker volume tars (v0.33.0) │ └── <user data> ← HDD mount contents (if app has HDD data) └── restic/ ← shared restic repo (all cross-drive apps)- DB dump files read from per-app home drive path (
AppDBDumpPath) _prefix directories prevent collision with user data- For non-HDD apps, only
_db/,_config/, and_volumes/(if applicable) are present (no user data directory)
- DB dump files read from per-app home drive path (
-
Restic backup paths: includes HDD mounts (if any) + config dir + per-app DB dump dir from home drive + stacks dir + controller.yaml (infra, v0.14.1)
-
Safety guards: destination ≠ source, path-overlap check (HDD mounts only), writable check
-
Chained execution: runs immediately after nightly restic — daily apps every night, weekly apps on Sundays
-
Hub reporting after manual triggers (v0.27.2):
OnCrossDriveCompletecallback on Router pushes infra backup snapshot to Hub + writes local infra backup after both single-app and run-all manual triggers complete (previously only automatic scheduled runs reported) -
Per-app concurrency lock prevents overlapping runs
-
Status (last_run, duration, size, error) persisted to settings.json
Protects against: primary drive failure, drive theft/damage.
Tier 3: Off-box (NAS) Backup — live
The off-site "1" of 3-2-1: each off-box-toggled app's recovery unit + DB dumps + volume tars are
backed up to the customer's NAS / Felhom offsite as an encrypted restic repo over SFTP (see the
off-box section below and internal/backup/offbox.go). The per-app "3. mentés" row on the backups
page renders one of four real states via the pure tier3State helper (internal/web/backup_page_state.go):
unconfigured (no target) / off (app not toggled) / escrow_pending (fork-4 key-escrow gate holds —
never a false success) / active (status badge + restic → <host> + relative last-run). Off-box run
status is repo-global (one LastRun); no per-app run time is fabricated.
The household is ASKED for the recovery code (v0.245.0, R-543). While the off-site tier is
configured and its escrow state is not escrowed, every authenticated page carries a reminder bar —
„A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking /backup/escrow.
It is the R-241 bar, second instance (internal/web/escrow_banner.go): same session-cookie
dismissal, back at the next visit, gone for good when the state is escrowed. It is added in
executeTemplate, the single render choke point, so it reaches every page; the login and claim
pages bypass that function, and a session check keeps it off the public guest share page. The pause
itself is UNCHANGED — it is the zero-knowledge escrow design, not a defect.
…but only once the box can do it (v0.246.0, R-546). For the first minutes after a bind the
agent's escrow preflight is not ok (typically no PBS storage yet — ~17 min measured). The bar now
consults the agent's own ok (escrow_readiness.go, cached 60 s, probed only while paused) and is
held back while the agent says not ready; /backup/escrow shows a waiting card („A doboz még készül
— … pár perc múlva …") that polls the preflight and reloads itself, instead of a red checklist; and
POST /api/escrow/start refuses 409 with the same sentence before staging or starting. Readiness
unknown (agent unreachable) keeps the bar — the fail-loud rule of R-543. A re-ceremony on an
escrowed box keeps the checklist, where a red row is a real fault.
Restore (internal/backup/restore.go)
Both Tier 1 (restic) and Tier 2 (rsync) restores are supported. All deployed apps appear in the restore dropdown with per-app snapshot filtering.
| App type | Config restored | DB restored | User data restored | Docker volumes restored |
|---|---|---|---|---|
| Has HDD data | Yes | Yes | Yes (always) | Yes (if present) |
| Docker volumes, no HDD | Yes | Yes | n/a | Yes |
| DB only, no HDD/volumes | Yes | Yes | n/a | n/a |
| Config only | Yes | — | n/a | n/a |
Snapshot API (GET /api/backup/snapshots?stack=<name> — F1, v0.99.0):
- Backed by
backup.Manager.ListRestorePoints(internal/backup/restore_points.go). The keep-side restore has exactly one restore point per app — the current recovery unit — so the endpoint returns at most one entry:time= newest artifact mtime (manifest / db-dumps / volume-dumps),short_id:"helyi",tier:1,drive_labelfrom the storage registry (empty on the SSD fallback) - Never emits tier-2 entries: Tier-2 copies are not restorable via
POST /backup/restore(it only reads the primary unit) — listing them would silently restore tier-1 data while claiming tier-2 - Guards: empty/traversal stack name → 400 (
validStackParam), unknown stack → 404, known stack with no unit yet →ok:true, data:[](the UI shows "Nincs elérhető mentés") - History: the route was a restic-era leftover fetched by the template but unregistered — the dropdown could never populate and the restore button never enabled (drill finding F1)
Restore type info shown per-app when selected in dropdown (Hungarian banners):
- Has HDD or Docker volumes: "Teljes visszaallitas: adatbazis + konfiguracio + felhasznaloi adatok"
- Has DB, no user data: "Adatbazis es konfiguracio visszaallitasa"
- Config only: "Csak konfiguracio visszaallitasa"
Tier 1 restore (RestoreApp):
- Stop app → resolve app's home drive →
restic restore <id> --target / --include <path>...→ populate Docker volumes from restored tars → restart app → health check - Restore paths: config dir, DB dump dir, volume dump dir, HDD mounts
- Docker volumes restored via
restoreDockerVolumes():docker volume rm -f→docker volume create→docker run alpine tar xf
Tier 2 restore (RestoreAppFromTier2):
- Stop app → rsync config from
_config/→ rsync HDD data (single/multi-mount) → copy DB dumps from_db/(streamingcopyFile) → restore Docker volumes from_volumes/tars → restart app → health check - Uses rsync
--deletefor config and HDD data to ensure exact mirror state - Single-mount apps: data directly in rsync dir (excluding
_*); multi-mount: per-leaf subdirectories
Common:
- v0.34.0: Post-restore health check (
waitForHealthy) polls container state withdocker psrefresh every 5s for up to 90s. Warning logged if app doesn't reach running state; restore still returns success (data is restored regardless). - Running flag prevents concurrent backup/restore operations
- Snapshot ID validated (8-64 lowercase hex, or special
tier2-rsync) - Import from
.fabbundle link shown in restore section for cross-system migration
Backup Page UI (internal/web/templates/backups.html)
Unified per-app status table with expandable rows showing per-tier backup status:
Status dot per app:
| Dot color | Meaning |
|---|---|
| Green | 2+ tiers configured with successful backups + destination healthy |
| Yellow | Only 1 tier, or Tier 2 failing, or Tier 2 configured but never run, or destination disconnected/inactive |
| Red | Tier 2 destination blocked or inaccessible |
Every app starts as yellow (1 tier only). Green requires Tier 2 configured with successful backup.
Per-app backup tiers (3 rows per app):
- 1. mentes (Tier 1, always present) — Auto badge + "helyi" + last run + contents. The contents label is PER TIER (v0.244.0, R-537): it describes what that tier actually captured, not what the app is shaped like.
- 2. mentes (Tier 2, configurable for ALL apps) — one of:
- Configured: method (rsync/restic) + destination + schedule + last run + status + contents + browsable indicator (folder icon for rsync) + action buttons
- Not configured: "1. mentes auto" + "Nincs 2. masolat" + settings link
- 3. mentes (Tier 3, off-box/NAS — live) — one of four
tier3Statestates:unconfigured("Nincs beallitva" + Beallitas link),off("Kikapcsolva" + Bekapcsolas link),escrow_pending("Kulcsletetre var"),active(status badge + "restic -> " + relative last-run)
Backup contents per app (shown per tier):
- Apps whose files live on the data drive (class A: calibre-web, immich, nextcloud, paperless-ngx): Tier 1 reads "DB + Konfig" — a Tier-1 unit has no file-copy step, so it does not hold them — and a sentence under the row says where the files ARE protected. That sentence renders by tier-3 STATE, not by the app's shape (v0.245.0, R-543) (
driveFilesNoteFor):active→ „…védi";escrow_pending→ „…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel" + the route; no off-site and no second drive → „Az alkalmazás fájljairól jelenleg nincs másolat…" + both ways out. It takestier3State's own vocabulary, so the row and the sentence cannot disagree. Tier 2/3 read "DB + Konfig + Adatok", because those tiers DO carry the file legs. - Apps whose data is entirely in Docker volumes (45 of 53 templates): "Konfig + DB + Adatok" or "Konfig + Adatok" on every tier — the unit really does hold their data.
- Apps with DB only: "DB + Konfig"
- The restore refuses rather than lying (v0.244.0, R-538): „Visszaállítás indítása" on a unit that cannot return an app's drive-side files is REFUSED before anything is stopped, naming the route that can (the off-site „Teljes visszaállítás (fájlok + adatbázis)", or the second drive's „Fájlok visszaállítása"), and saying plainly when no copy exists. A database-and-settings-only restore is a separately-worded second step.
- Apps with neither: "Konfig"
Deploy page shows cross-drive (Tier 2) configuration form for all deployed apps, not just those with HDD data. Non-HDD apps can configure destination, method, and schedule.
Other sections:
- Schedule overview with next run times for DB dump, restic, prune
- Snapshot history table (last 20 snapshots aggregated from all per-drive repos, sorted by time)
- Storage overview card (total size across repos, snapshot count, DB dump count/size, encryption key with show/copy)
- Restore section: app dropdown → per-app snapshot dropdown (Tier 1 + Tier 2 grouped) → restore type info → confirmation checkbox → execute → import from
.fabbundle link
4. Storage Management
Off-box (NAS) backup — restic-over-SFTP (v0.93.0, Part B). The "1 off-site" leg of 3-2-1 for the app-data tier: each off-box-toggled app's recovery unit + DB dumps + volume tars are backed up to the customer's NAS as an encrypted restic repo over SFTP (
internal/backup/offbox.go). No kernel mount — restic talks SFTP directly; the NAS sees only ciphertext. Distinct from the local cross-drive rsync copy and the agent's PBS whole-CT DR.
- Fail-fast (load-bearing): every restic call uses
-o sftp.command="ssh … -oConnectTimeout=10 … -s sftp"so a dead NAS errors in ~10 s, never hangs the backup runner; a failure raises thebackup_failedoperator alert.-oStrictHostKeyChecking=yes+ a pinnedknown_hosts(no blind TOFU).- init-if-absent (idempotent),
forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune, single-flight (sharesm.running) + migration-guard, restic's own exit code checked, restore to a scratch dir (non-destructive).- Unit discovery (v0.104.0) — durable, deployment-independent. Each toggled app's recovery unit is discovered by scanning the durable storage registry — every registered schedulable storage path (
GetSchedulableStoragePaths) ∪ the system-data fallback drive — forbackups/primary/<app>, rather than inferring the drive from the app's liveapp.yamlHDD_PATH(which silently fell back tosystemDataPathfor a toggled-but-undeployed app → offbox looked on the wrong drive, backed up nothing, and reportedok/0). If the same app's unit exists on two drives (drive churn), the newest by manifestCreatedAtis used and the stale one is WARN-logged. The WRITE side (CaptureRecoveryUnit/ dumps) is unchanged — this only changes offbox's read/discovery path. Boundary: decommissioned or non-schedulable drives are not searched.- No silent success (v0.104.0): a run where ≥1 app is toggled but 0 were backed up (no unit found anywhere) is a hard error →
LastStatus="error"+ operator alert (was a misleadingok/0 snapshots). A partial run (some units missing) staysokbut sets a HungarianLastWarningnaming the skipped apps, shown on/backups.- Crash-lock self-heal (v0.110.0). A crash mid-prune leaves a restic EXCLUSIVE lock that plain
restic unlockcan't clear (the recreated container's new hostname stops restic proving the dead PID stale for ~30 min). Every backup/prune/restore step runs throughresticStep, which on a lock error escalates tounlock --remove-all+ one retry — safe because the repo has a SINGLE legitimate writer (per-customer sub-account isolation + the in-process single-flight mutex). Boundary: a DR-cloned SECOND controller writing the same repo would defeat this premise — operator-supervised territory, out of scope for the auto-heal. A crash mid-run also flips a persistedLastStatus="running"to a truthful error on the next startup (self-corrects on the next successful run).- Secrets (SSH key + auto-gen repo password) are 0600 files in the data dir — never logged/committed.
- Password custody + atomicity (v0.105.0, fork-4; pairs with agent v0.77.0). The repo password is the irreplaceable DATA key for the offsite tier, so it rides the customer-recovery-code (R) escrow (age-under-R, in the agent's
IdentityBundle— custody spike validated a recovered password opens the real repo). Atomicity gate: enabling offsite pushes the password to the agent (StageEscrowSecret→POST /escrow/stage-secret) and marksEscrowState="pending"; no offsite backup runs until the escrow is confirmed (OffboxRunnable()), so an un-recoverable offsite copy can never exist. Ceremony — PRIMARY path (v0.127.0, agent ≥ v0.88.0): the customer wizard at/backup/escrow(web/escrow_handlers.go+templates/backups_escrow.html): preflight (agentGET /escrow/preflight+ version gate) → warnings (re-ceremony adds the supersede copy) → password re-auth (login rate limiter) → re-stage-first (offbox configured →PushOffboxPasswordForEscrow; a staging failure ABORTS the start — no hash-less blob, ever) → agent job (POST /escrow/ceremony, pollGET /escrow/ceremony/status@2 s) → one-shot R reveal (POST /escrow/ceremony/claim,Cache-Control: no-store; R exists only in the page's JS scope; 10-min unclaimed TTL → void, re-run supersedes) → typed-back (two random words, client-side) → finish. R is never templated/logged/persisted on either side. Operator fallback (CLI, unchanged text mode): enable offsite (→ pending) → runfelhom-agent --selftest=escrow-create --upload(the staged restic password auto-injected, then wiped) → hand the customer the fresh R (once; supersedes any prior code) → the auto-confirm flips escrowed hands-free. (The manual "Letét megerősítése" button is GONE from the card; the deprecated endpoint remains for legacy hash-less blobs.) DR: recover R → the escrow yields the password →POST /backup/offbox/inject-password {password}pre-places it 0600 → configure offbox → restore. The SFTP access key is regenerated at DR (a fresh sub-account key), NOT escrowed; the DR recipe carries only the non-secretoffsite_resticcoordinates (DRResticCoord).- Hub-verified auto-confirm (v0.108.0, SLICE 3; pairs with agent v0.79.0 + hub v0.40.0). The report ACK carries
escrow:{identity_blob_present, restic_pw_sha256, created_at}(the hash is recorded at ceremony time — non-reversible sha256 of the sealed password, safe to serve).report.EscrowAutoConfirmerflipsEscrowStatepending→escrowed ONLY whensha256(local repo_password)matches — i.e. the stored escrow provably covers the CURRENT key; blob-presence alone never confirms (a stale blob would re-open the fork-4 gap). Mismatch → stays pending + a loud warn naming the ceremony (deduped per hash); no hash / no row / no local file → silently pending; never un-confirms. On flip it wipes the agent-staged secret. The canonical hasher (backup.HashResticPassword, trimmed-string sha256) is pinned by a cross-repo test vector against the agent's. The manualPOST /backup/offbox/confirm-escrowis a deprecated fallback for legacy hash-less blobs (e.g. the demo's). Stale-blob re-check (v0.127.0, Scenario F): an ESCROWED box re-compares the ACK hash every cycle — mismatch OR a present blob with an EMPTY hash (a superseding ceremony that missed the staged secret) sets an in-memory stale flag (EscrowAutoConfirmer.StaleBlob→ the Távoli mentés card's warning + re-ceremony CTA) + one warn per distinct hub hash. State never flips, runs never block; a covering blob (or a fresh auto-confirm) clears it. Awaiting-confirmation card (v0.138.0): the flip above lands on the next report ACK, so a just-finished ceremony sits pending for up to ~15 min. To avoid re-showing the yellow "Helyreállítási kód szükséges" card during that gap, a successful recovery-code claim stampsOffboxTarget.CeremonyCompletedAt(RFC3339, persisted; zeroed on the flip). While stamped and pending,/backups/remoteshows an info "megerősítésre vár, legfeljebb 15 perc" card (offboxCeremonyWaitState); pastescrowCeremonyGraceWindow(35 min = 2 cycles + slack) it degrades to a warn "a megerősítés nem érkezett meg" + re-ceremony CTA — never an indefinite wait. The wizard's final step shows a matching "Mi történik ezután?" note. Display-only: no state change, no run-gate effect. (Phase-0 A: the wait itself is correct; only the feedback was missing.)- Injection guard (
ValidateOffboxTarget): host/user/repo must not start with-(ssh option-injection) or carry metacharacters/traversal;OffboxConfiguredfails closed on an invalid target. Image:restic+openssh-client(re-added; restic's sftp backend shells out tossh).- UI: the "Külső (NAS) mentés" section on the backups page (configure target, per-app toggles, run-now, restore, status). Config:
settings.OffboxTarget+ per-appAppBackupPrefs.Offbox. Daily at 04:15.- Soft quota (v0.109.x, SLICE 4; pairs with hub v0.41.0). The shared-model quota (
quota_gb) rides the descriptor intoOffboxTarget.QuotaGB(0 = no limit — dedicated boxes are Hetzner-enforced; the hash includes it, so a hub-side quota change re-applies via key-auth-first, no password consumed).RepoSizeBytespersists fromrestic stats(last-known on failure — stale-but-safe). Pre-run gate: ≥100% refuses NEW backup runs (Hungarian error + operator alert) but the prune step still runs (the only way back under quota) and restore is never gated; ≥80% sets a Hungarian usage warning./backupsshows a usage bar when quota>0. The hub report carriesoffsite:{enabled, escrow_state, last_run, last_status, snapshot_count, repo_size_bytes, quota_gb}— the hub's OffsiteChecker alerts on fill (90/95%) and staleness (escrowed + no run >48h); the Hetznerreadonlyfreeze is an OPERATOR lever on the hub (never automatic). A RE-apply preserves the existing target's EscrowState + runtime status (v0.109.1 — custody tracks the preserved repo password).- Hub-driven provisioning apply-bridge (v0.106.x, SLICE 2; pairs with hub v0.38.x — validated live 2026-07-09). When the hub provisions the offsite tier (a Hetzner Storage Box sub-account or dedicated box), the served
controller.yamlgains anoffsite:section (host/user/port/repo_path/quota_gb +host_fingerprint) andinternal/offsiteapply.Bridgereconciles it at startup: scan + verify the box host key againsthost_fingerprint(no blind TOFU) → generate an ed25519 keypair → consume the hub's one-time transient password (POST /api/v1/offsite/consume-password/{id}, single-use) → install the pubkey (sshpass -e ssh-copy-id -p 23 -s -f, pinned known_hosts,StrictHostKeyChecking=yes; the installer ensures~/.sshexists — ssh-copy-id's SFTP mode needs it) → verify key-only auth → configure the offbox target →EscrowState="pending"(the fork-4 gate above still holds) → persist a descriptor-hash marker LAST. Idempotent (marker → no re-consume) and fail-safe (any failure → nothing persisted, retried on the next config refresh/restart; a consumed-but-failed install logs a loud "password is spent — reset on the hub"). The 15-min hub-report → config-refresh cycle is the trigger, so descriptor-to-applied latency is ≤ ~15 min.- Key-auth-first (v0.107.0): on a descriptor change, if the ALREADY-INSTALLED key still authenticates (
SFTPKeyAuthProber, pinned to the freshly-verified host key — the probe never bypasses the fingerprint verify), the bridge re-pins + reconfigures WITHOUT consuming a one-time password. Fresh guests fall through to the full consume+install path. The hub's "Re-issue offsite credentials" (v0.39.0) is the recovery for a genuinely-spent password on a fresh guest.- Staged-secret wipe (v0.107.0): confirm-escrow calls the agent's
DELETE /escrow/stage-secret(agent ≥ v0.78.0) wheneverEscrowStateflips toescrowed— best-effort, loud-logged on failure.- Settle-gate (R-71a, v0.162.0) — the day-0 race removed. The apply-bridge runs BEHIND a settle-gate (
Bridge.AwaitSettle→ReconcileWhenSettled): before the consume/install path it polls the self-updater's own state via theSettleProviderseam (aSettleFuncadapter overupdater.GetFloor()/IsUpdateRunning()in main.go — no second floor-fetch path). While a managed update is running OR the box is below the operator floor (an auto-floor update is imminent), the gate WAITS rather than consume the single-use password — the update's restart would otherwise kill the bridge mid-install and burn it (the F10 day-0 shape). At/above floor with no update in flight, it GOes on the first poll with zero added latency (B′). Bounds: 10 s poll, 90 s floor-knowledge sub-bound (sized to the ~5–10 s report-ACK floor latency; the floor is in-memory, not persisted, so it is unknown until the first ACK on any restart), 5 min overall — both bounds GO+WARN and lean on the R-71c hub self-heal as the belt (a hub that cannot serve a floor cannot serve a consume, so proceeding never burns a password). The gate is wired only when a self-updater exists (no updater → no floor-update to race → reconcile immediately). Ordering-only: the consume/install/persist internals and the 404-no-oracle contract are untouched. Three-layer defense: the v1.25.0 golden≥floor build gate PREVENTS the trigger, (a) DEFERS it, R-71c HEALS a burn.
NAS network storage (v0.92.0, Part A2; pairs with agent v0.50.0). A customer NAS share (NFS or SMB) is a distinct storage KIND from a physical drive (
StoragePath.Kind == "network"), for bulk media. The controller is a thin proxy over the agent's/netstorage/*(A1) + the local registry — it holds no mount authority and never persists the SMB password (it passes the credential to the agent's add request, which writes the 0600 file). EndpointsPOST /api/storage/netstorage/{add,remove}+GET /api/storage/netstorage(merges the agent's per-share liveness:ok|idle(benign idle-unmount) |unreachable|unknown). A network share registers as/mnt/felhom-drives/<name>, is selectable as a media app'sHDD_PATH(Schedulable), and shows in the "Hálózati tárhely (NAS)" settings section.
- NOT a drive: the drive lifecycle (eject/decommission/migrate/wipe/SMART) is refused on a network path server-side (
refuseNetworkLifecycle). The drive-absent gate (planDriveGates) and the missing-storage surface skip network paths, so anunreachableNAS is a recoverable warning (networkStorageWarnings→ a distinct app-card badge), never the drive "missing → stopped" cascade.- Verify-before-commit (v0.113.0, pairs with agent v0.81.0; SPIKE-nas-verify-2026-07-11): add no longer registers blind.
POST /api/storage/netstorage/addreturns{started:true}and a DETACHED single-flight orchestration job (internal/web/netstorage_job.go, migrate.go shape; polled onGET /api/storage/netstorage/add/status) drivesagent_add → verifying → probing → registering: the agent installs the units + runs its own detached mount-verify (journal-classified, agent-side auto-rollback), then the controller re-execs itself as--netprobe <dir>at uid/gid 1000 (SysProcAttr.Credential,netprobe*.go) to prove a media app can WRITE through the share (the squash trap), and only then registers. ANY failure = full rollback (nothing registered, nothing installed, no creds file); categorized Hungarian errors (netAddMessage, §3.2 map — notenfs_exportmerges not-found/not-permitted: NFSv4 returns identical strings). An agent restart mid-verify (verify-statusphasenone) ⇒ controller rollback; any agent-configured share NOT in the registry surfaces as a remove-only "Árva megosztás" orphan row. The page (storage_network.html) is on the canonical form pattern with staged poll progress and the protocol-honest NAS guidance (SMB-first; NFS map-all-users vs full-fidelityanonuid=<uid+100000>). Authoritative doc:felhom.eu/documentation/controller/network-storage-nas.md.- Agent-capability gate (v0.114.0): the verify-before-commit add is COUPLED to agent ≥ v0.81.0, so the add entry point now probes the agent first (
internal/agentapi/features.go,Supports(FeatureNetstorageVerify)— a route probe onGET /netstorage/verify-status: 2xx ⇒ supported, typed 404 ⇒ older agent, transport/5xx ⇒ indeterminate, cached 5 min both polarities). On an older agent the add is refused SYNCHRONOUSLY (HTTP 412, machine codeagent_outdated, honest Hungarian message) BEFORE the single-flight claim — never a misleading mid-pipeline rollback; indeterminate NEVER refuses (a down agent speaks through the existing error paths). The settings page swaps the add form for a banner onSupportNo(share list + remove stay usable in every state). Convention: every future coupled feature adds afeatureProbesrow + a gate call at its entry point and declaresMinAgentin its CHANGELOG header — seefelhom.eu/documentation/runbooks/publish-train-rules.md.- Consuming-namespace verification (v0.117.0, RCA AUDIT-nas-cwa-rca-2026-07-11 fix 2): a guest reboot silently replaces an idle NAS trigger with a plain local STUB dir in the app namespace while host-side (agent) health stays green — so the controller now verifies where the apps consume:
internal/system/fsclass*.goclassifies a path by statfs f_type in THIS process's namespace (networknfs/cifs/smb2 |autofsidle trigger = HEALTHY, never force-mounted |stub|unknown= fail-open). Three consumers: (1) the--netprobechild now REQUIRES a mounted network fs after its create (exit 5 → categorynot_network_fs, full rollback — a writable stub can never verify); (2)POST /api/stacks/{name}/deployrefuses (409) a registered networkHDD_PATHthat classifies as a stub (refuseNetworkStubDeploy; idle autofs deploys fine — first app access mounts it); (3) the dashboard/stacks app cards gain a distinct stub badge ("Hálózati tárhely hibás — az alkalmazás nem a NAS-t látja",NetworkStubs) that WINS over the recoverableunreachablebadge. The agent pair (v0.84.0ReassertNetworkMounts) re-arms triggers on guest start; this layer is the detection net beneath it.- One classification, two surfaces (v0.119.0, CAMPAIGN-3 F8): the share row on the
/storage/networkpage now reads the SAME classifier. Its health used to come only from the agent's server-level TCP dial (server:2049/445), which stays green when a single export isexportfs -u'd — so the row showed benign "Készenlét" while the stacks cards showed the stub.fuseNetHealth(netstorage_handlers.go) fuses the agent view withclassifyFSPath(Where): a newstubhealth (badge "Hibás — az alkalmazások nem a NAS-t látják") overrides idle/ok when the namespace sees local disk; a whole-serverunreachablestill wins over stub; autofs-healthy / network /unknownleave the agent health intact (never force-mount an idle trigger). The row and the stacks badge derive from ONE classifier and can never contradict.- mapped_uid validated at the door (v0.119.0, F8's sibling F4):
handleNetStorageAddrange-checks the container uid/gid (1..65533; the guest maps<uid>→<uid>+100000, so 65534=nobody and a host-side value like 101000 must not be entered) after the<=0default — out of range → a friendly Hungarian 400, nothing installed (previously a rawagent_errorfrom the agent).- Deploy view truth (v0.117.0, RCA fix 4): a deployed app's read-only storage select now marks
selectedby the app's STOREDHDD_PATH(extra disabled<path> (nem elérhető)option when the stored path left the schedulable list) —IsDefaultselects only for NEW deploys. Pre-fix the view showed the default drive regardless of app.yaml (the RCA's S-C symptom).- Limits (v1): a share's
+100000uid mapping is fixed at add-time (one app / same-uid apps); for write apps on a soft NFS mount, an in-flight file can truncate if the NAS vanishes mid-write (prefer atomic-write apps / SSD-staging).
⚠️ INTERMEDIARY-MOUNT model (v0.67.x, pairs with agent v0.35.x). External data drives are now visible in the guest at a STABLE path
/mnt/felhom-drives/<name>, NOT the raw/mnt/<name>. The agent keeps a single permanent parent bind/mnt/felhom-drivesin the guest and swaps each drive'sfelhom-datanamespace underneath it host-side (mount --bind), so the guest sees attach/detach live with no reboot (mount propagation), the bind source never disappears (C1-immune), and onlyfelhom-datacrosses in (confinement). The per-drivepct set -mpNbind is deprecated.
- The registered storage path + every app's
HDD_PATH+ the FileBrowser source = the stable/mnt/felhom-drives/<name>. The controller maps it back to the raw/mnt/<name>(agentWhere()) only for agent calls (assign/attach/eject/decommission).GET /api/diskscarriesguest_path+bound_under_parent(the guest-visible signal).- Drive-absent gate (
internal/web/intermediary.go,driveGateLoop30s): an absent/mnt/felhom-drives/<name>drive stops + blocks its apps (StoppedStacks= the gate-stopped set); a returned drive re-attaches + auto-restarts them;actionStackrefuses to start an app whose drive is absent ("tárhely nem elérhető"). SSD/system paths are never gated.- H1 endpoints
POST /api/storage/{disconnect,reconnect,restart-apps}drive the host-side eject/reconnect (no guest reboot).- Lifecycle (v0.68.x): a "Leszerelés" button decommissions a drive (migrate-then-decommission OR decommission-anyway type-to-confirm — non-destructive, never touches the parent mp); a "Visszacsatlakoztatás" button one-click re-enrolls a decommissioned/ejected drive (clears the marker, re-binds under the parent, restarts gate-stopped apps). Decommissioning the DEFAULT auto-promotes another schedulable drive (
defaultPromotionTarget), or BLOCKS if it's the only one. Eject and decommission keep the RAW drive mounted (logical retire) so re-enroll re-binds it.- Guest-reboot convergence is DETERMINISTIC via the agent's
guest_boot_id: the controller persistsLastGuestBootIDand, when it changes, recreates EVERY deployed drive-backed app onto the re-propagated drive (processGuestBootChange— no fragile container-uptime sampling). v0.71.0 — boot-race fix: on a guestpct reboot, in-guest dockerd auto-starts the apps ~18s BEFORE the agent re-binds the drive, so their volume bind fails at create-time (mkdir …/userdata: permission denied,RestartCount=0→ never retried → stuckExited). The old recovery sampled the agent'sBoundUnderParentONCE, raced that rebind, recreated nothing, and burned its boot-id one-shot.processGuestBootChangenow gates on the REAL live in-guest bind (driveBindLive: is/mnt/felhom-drives/<drive>an actual mountpoint in the controller's own/mntrslave/proc/self/mountinfo?) and waits for it (pollLiveBinds, bounded ~120s) before recreating — including apps stuckExitedwith a create-time mount failure (shouldRecreateOnBootis state-independent). The guest-only reboot path (which the host-reboot sweep never exercised) is now covered; drives that never go live in the window are left to the drive-absent gate.processGuestBootChangealso runs on every periodicdriveGateLooptick now (idempotent, boot-id gated) so a momentarily-unreachable agent right after a guest reboot no longer permanently strands recovery. v0.157.0 — R-55, the gate now honours a customer's Stop. Until this version the recreate keyed onDeployed && HDD_PATH && drive-presentalone, so a drive-backed app the customer had deliberately Stopped was silently restarted on every guest reboot (proven live: immich, stopped from the UI seconds earlier, came back running).shouldRecreateOnBootnow also requires the app to still HAVE containers (len(Stack.Containers) > 0, fromdocker ps -a, soExitedones count) — R-52'sexisting-Exited vs absentdistinction (bootrecon.isBootOrphan) translated to this gate. A UI Stop iscompose down, which REMOVES the containers; a guest that went down under a running app leaves them present. State is still NOT a filter — that part of the original design is load-bearing and unchanged;hasContainersanswers a different question ("does docker still have records of it") which, unlike liveness, survives a reboot as a statement of intent. The evidence is sampled BEFORE any recreate, because recreate's ownStopStackerases it. Apps stopped by the drive-absent gate are also at zero containers and are likewise left alone here — they are restored byReconcileDriveGates'Returnbranch fromStoragePath.StoppedStacks, on the same loop tick. Honoured Stops are logged at INFO (left stopped …), counted separately from the "no live bind" skips so an intended outcome never fires a WARN. Agent-path prerequisite (also v0.71.0): the whole drive gate needscfg.LocalAPI.Endpoint(the per-guest agent local API).bootstrap.MaybeIngestnow callsensureLocalAPIon the already-configured path — merginglocal_apifrombootstrap.jsoninto an existing controller.yaml that lacks it (seeded beforelocal_apiexisted) — because without itagentClient()returns "agent not configured" and the entire gate + boot recovery silently die. v0.72.0 — FileBrowser convergence on boot-recreate: FileBrowser is base-infra (it binds each drive'suserdatabut has noHDD_PATH, so it is NOT in the drive-backed recreate set) — after a host reboot its mounts could be stale (the early first-boot bring-up bound them before the drives went live).processGuestBootChangenow, afterpollLiveBindsconfirms the binds and the apps are recreated, triggersgo s.SyncFileBrowserMounts()so FileBrowser converges against the now-live drives. The sync runs unconditionally (FileBrowser reflects the current bind state even if no app needed recreating). The recreate loop is a purerecreateDriveBackedApps(stacks, present, recreate, syncFB)that callssyncFBexactly once, after every recreate. Live-accepted over two realfelhom-pvereboots (FileBrowser non-stale, all drive-backed apps recovered, agent tolerated a/dev/sdXswap by UUID).⚠️ Rebuilt on the agent-delegated disk model (v0.43.0), made ROLE-AWARE in v0.44.0, UX-polished in v0.45.0. After the 8C de-privileging, the controller holds no Proxmox/disk credentials and no destructive authority — disk execution + the gate live entirely in the host agent. The drive UI is driven by the agent's authoritative role (
system|backup|user-data, fromGET /api/disks): the appliance's own system storage and the backup safety-net are visibly protected (lock badge, NO destructive controls); the customer manages their own user-data drives with informed consent. The agent re-enforces role at wipe time — the UI lockout is defense-in-depth, not the gate.
Overview (
settings.html←GET /api/disks): styled cards (not a table) — name, mono device/mount, badges for class (gyors/lassú), data (Adatot tartalmaz), role (🔒 Rendszer / 🔒 Biztonsági mentés — védett / Felhasználói adat) and registered state, plus a capacity bar (the monitoringsystem-bar, from the agent'stotal_bytes/used_bytes). Eject/Wipe render only for user-data drives mounted under/mnt.
- (v0.45.0) Deterministic order —
agentDisksListHandlersorts the list server-side (sortDisksForView): user-data → system → backup (then unrecognized), alpha by name within a tier, so it no longer reorders on each reload (the agent's view iterates an unordered Go map).- (v0.45.0) Purpose + app-backing clarity (B4) —
localandlocal-lvmare both shown (not collapsed); each card carries a plain-Hungarian purpose description keyed on the agent's role/type, the app-backing storages are tagged (local-lvm→ "Alkalmazás-rendszer"; user-data → "Alkalmazás-adatok"), and a one-line tiering note above the list answers "which storage do the apps use?". Pure presentation — role/type stay authoritative from the agent.- (v0.45.0) Register shortcut (B3) — a mounted, unregistered user-data drive offers Regisztrálás as its PRIMARY action:
POST /api/storage/register→registerStoragePathrecords the existing mount (no format, no eject) + FileBrowser-syncs. Leválasztás/Törlés stay secondary.Customer wipe/eject — a type-to-confirm modal that names the deployed apps that break (
GET /api/storage/impact→appsUsingPath) and disables the destructive button until the mount name is typed exactly. Wipe (POST /api/storage/wipe): eject (unmount + deregister) → server-side two-step customer-confirmed format (learn the agent's durable id, then re-submitconfirmed:truebound to it). The agent refuses a protected device regardless of what the controller sends.Guided init (
/storage/init,POST /api/storage/init,web/storage_handlers.go): format → resolve the new fs UUID →assign→ register. The selector lists only user-data targets. A data-bearing user-data device now uses the customer-confirmation flow (type-to-confirm → re-submitconfirmed:true+ durable id), NOT thefelhom-opsigncommand. The opsign surface remains a fallback only if a protected device somehow reaches init. F6 (v0.141.0): the chain runs as a DETACHED single-flight job (web/storage_init_job.go,context.Background()) the wizard polls viaGET /api/storage/init/status(3-step progress: formatting → mounting → registering) — a closed tab / lost connection no longer aborts the post-mkfs mount+register.registeris the LAST step (marker-last crash-safety). A slow mkfs that outruns the agentapi client's 15 s timeout is followed by polling the agent'sGET /disks/format/status(agentapi.Client.FormatStatus→awaitAgentFormat) before continuing. Live-validated on a 64 GB USB (mkfs ~27 s → done, mounted+registered at/mnt/felhom-drives/scratch1).Guided attach (
/settings/storage/attach,POST /api/storage/attach): non-destructive — resolve the existing fs UUID →assign→ register. Selector restyled to cards (user-data only).Eject (
POST /api/storage/eject): benign unmount + deregister, with the agent's dependent-guest warning + the affected-app list (parity with wipe). The eject is ROLE-GATED at the agent (felhom- agent v0.24.0):POST /disks/ejectrefuses to unmount a system/backup mount — the UI hiding the button is defense-in-depth, not the control. Only user-data mounts are ejectable.
agentapi(internal/agentapi) is the pinned client to the agent local API:Disks/AssignDisk/EjectDisk/FormatDisk(…, confirmed, durableID);DiskInfo.role+capacity;FormatResult.{role,needs_confirmation,durable_id};ErrNeedsConfirmation(user-data) vsErrFormatRefused(system/backup).FormatResult.PendingOp.OpsignCommand()for the operator path.
- (v0.74.0) Client lifecycle — ONE shared client, reused.
Server.agentClient()builds theagentapi.Clientonce (memoized viasync.Once) and returns the same instance to all ~19 call sites; thehttp.Transportis bounded + expiring (MaxIdleConnsPerHost:2,IdleConnTimeout:90s). This replaced a per-callagentapi.New(...)that leaked one idle ESTABLISHED socket per call and exhausted the ephemeral source-port range to the agent's:8443after ~5 days (EADDRNOTAVAIL). Safe becausecfg.LocalAPIis static per process (a config-apply triggers a graceful self-restart).- (v0.90.0) Channel health-check —
internal/channelhealth. A ~60s scheduler job (agent-channel-health) probes the channel viaServer.ProbeAgentChannel(the SAME memoized client +GET /storage— not a fresh client) and classifies failures intoup | down:<reason>(pin_mismatch / unauthorized / unreachable / timeout / misconfigured / construction_error). Transient reasons (refused/timeout) are debounced (N≥2 consecutive) so a clean agent restart's ~1s blip doesn't page; pin/401/DNS/construction alert on the first down. On a transition it fires an English operator-only event (Notifier.NotifyAgentChannelDown/Recovered, hub 1h cooldown) and sets a Hungarian dashboard banner (AlertManager.SetAgentChannelAlert). The first observation seeds state silently. This is the controller half of the self-health story (the agent watches its own privileged capabilities; the controller watches its link to the agent). A construction error (agentClient()can't build — a latching config fault) is surfaced distinctly. Detection/surfacing only — it never touches the pin, transport, or gate.The
StoragePathregistry (settings.go:AddStoragePath/default/schedulable/label) is unchanged.
- (v0.64.0)
AutoDiscoverStoragePathsis now ADDITIVE — it no longer bails when the registry is non-empty; instead it registers only deployed-app paths missing from the registry. It never removes or mutates an existing entry, never re-adds or reactivates a path already present in ANY state (incl.Decommissioned), and never flipsIsDefault(a new path defaults only if the registry has no default at all). This is NOT auto-register-on-attach (that recommendation was rejected — manual enrollment is by design); it only picks up drives that deployed apps already reference.- (v0.64.0)
InferStorageLabeldisambiguates the internal SSD — a path whose basename is thefelhom-datanamespace dir (the internal system volume, e.g./mnt/sys_drive/felhom-data) now reads "Belső SSD (rendszer)" instead of the colliding "Tárhely (felhom-data)". Model-A user drives register their MOUNT ROOT (e.g./mnt/felhom-usb), never.../felhom-data, so this can't mislabel a user drive. Still overridable viaSetStorageLabel.(v0.65.0) Data migration (
internal/stacks/migrate.go) — move app data between drives, in-process over the controller's/mnt:/mnt:rslaveRW mount; crash-safe + resumable via a journal (<dataDir>/migration.json).MigrateAll(source,target)moves the whole felhom-data namespace (every app + a conflict-merge walk for non-app/customer content);MigrateApp(app,target)moves one app's subtree (drive→drive AND SSD→drive). Pipeline: validate → stop → copy (rsync -a --checksum, additive, no--delete) → verify (rsync -ani --checksum, zero pending) → flip+redeploy (RedeployFromEnv) → cleanup. CLEANUP — the only destructive step — is gated on every unit verified AND every app redeployed. Conflict-merge: skip-identical (checksum vs target + its(N)siblings), rename-on-differ to lowest-free<base>(N)<ext>, never overwrite; idempotent. Single-flight + mutually exclusive with the backup orchestrator. UI:POST /api/storage/migrate{,-app}+GET /api/storage/migrate/status(poll); migrate-all on the settings page, per-app on the app-info page, shared Hungarian progress panel.(v0.65.0) Self-serve decommission (
handleStorageDecommission,POST /api/storage/decommission) — retire a drive, non-destructive (never formats). Two choices (no partial): migrate-all-then-decommission (runsMigrateAll; the migration done-hook soft-marks the source + calls the agent on success), or decommission-anyway (type-to-confirm; stops the apps, keeps theirHDD_PATHso they show "Hiányzó tárhely"). Both end atsettings.SetDecommissioned(soft marker retained — blocks A1 resurrection) +agentapi.Decommission(agent v0.32.0:POST /disks/decommission, role-gated user-data, intent + bind-prune + unmount). Re-enrolling a decommissioned drive (registerStoragePath) clears the marker + restoresSchedulable. A deployed app whose drive is decommissioned/disconnected/absent shows the "Hiányzó tárhely" warning badge on the dashboard/stacks/app card.(v0.66.0) Userdata layout + shared-storage ownership convention (
internal/appbackup/userdata.go). Each drive's felhom-data namespace gains a customer-facinguserdata/tree (sibling ofappdata//backups/) — the ONLY thing FileBrowser mounts. Skeleton:media/{movies,tv,music,audiobooks,books,comics,photos},downloads,import/{paperless,calibre},roms,documents. Convention: every userdata dir is group-owned bySharedContentGID(1000), mode 2775 (setgid + group-rwx) —EnsureUserdataDir= MkdirAll → explicitChmod(ModeSetgid|0775)(MkdirAll's mode is umask-masked AND drops setgid) → chown group 1000. Setgid makes new files inherit group 1000 so FileBrowser (uid 1000) + the content apps collaborate without permission collisions.
USERDATA_PATH=<HDD_PATH>/userdatais injected into the compose env (the sharedwithUserdataPath, used by BOTHstackEnvandcomposeExecWithEnv— the initial deploy builds env from values, not app.yaml). Catalog media mounts use${USERDATA_PATH}/....- Pre-create: the full skeleton is created with the convention on
registerStoragePath+syncFileBrowserMounts(system + additional drives); a deploy belt incomposeExecCustomEnv(onup) pre-creates every${USERDATA_PATH}/...bind source the app declares, so Docker never auto-creates a userdata dir as guest-root (covers apps not in the skeleton).- NAS shares browse too (v0.160.0, R-67): a registered network storage binds its share ROOT into FileBrowser (
/mnt/felhom-drives/<name>:/srv/<name>:rslave) — no userdata scoping and NO skeleton (Felhom convention dirs are never written onto a customer's own NAS; the sync is read-only toward the share). The gate differs from drives: an idle autofs trigger is HEALTHY and included (first access wakes it — Phase-0-probed through an rslave bind on demo-hp), while astubclassifier verdict (this namespace sees a local dir, not the NAS) EXCLUDES the share from both the mounts and the source list that pass — uploading into a stub would be silently shadowed by the real mount later.unknownfails open. NAS add-success and remove trigger the same debouncedSyncFileBrowserMounts; pure assembly lives inbuildFileBrowserPaths(handlers.go) with every edge seamed.- FileBrowser mounts
<drive>/userdata(wasappdata) and runs as uid 1000 → it can create folders + upload into the 2775 setgid tree (fixes the old permission-denied); app internals (appdata/) are no longer browsable. (v0.66.2) its entrypoint is wrapped["sh","-c","umask 002; exec /home/filebrowser/filebrowser"]so folders the customer creates come out 2775 (group-writable) — the gtstef image is a single Go binary that ignores aUMASKenv (verified), so the wrapper is the mechanism. Without it, customer folders were 2755 (group-read only) and apps in group 1000 couldn't write into them.- Run-identity: PUID/PGID-1000 apps (radarr/sonarr/calibre with
UMASK=002) write group- writable, so FileBrowser can fully manage their output. Root-only apps (jellyfin :ro; komga + audiobookshelf RW after theuser:1000fallback) write group-readable via setgid (FileBrowser browses/reads, the app manages its own files).- Migration-safe:
migrate.go's merge walk preserves the source dir's setgid + group (andcopyFilethe full file mode + group), so the convention survives a whole-driveMigrateAll.The privileged controller-side disk subsections below are historical (the
internal/storage/*scan/ format code was removed in 8C — execution is the agent's now).
The storage subsystem handles the full lifecycle of external storage: detection, initialization, path registration, and data migration.
CURRENT (post-de-privileging + Impl-2b, v0.95.0): the in-guest storage code below (
internal/storage/scan.go,format.go,attach.go) is retired — all disk ops are delegated to the host agent viainternal/agentapi. The two enrollment wizards (/settings/storage/init,/settings/storage/attach) now populate candidates from the agent's raw-device scan (GET /api/disks/candidates→ agent Impl-2a, proxied byagentDiskCandidatesHandler):initialize= every unclaimed disk (blank or data-bearing),attach= the mountable-FS subset. The agent's unclaimed-disk filter (Impl-1claim.go) is authoritative + fail-safe (never offers OS/enrolled/claimed disks), so the controller does NO client- or server-side filtering. Enrollment posts to the unchanged/api/storage/init(format via the agent's Impl-1 guardedmkfs→ mount → bind → intent) or/api/storage/attach(mount + bind, no format). The legacy text below is kept for historical context.
Disk Scanning (internal/storage/scan.go)
ScanDisks()useslsblk -J -bfor block device enumeration- System disk detection via host fstab parsing (
/host-fstab) + UUID resolution viablkid - Partitions enriched with filesystem type, UUID, and label from direct
blkidprobing (Docker containers have incomplete udev cache) - Returns
AvailableDisks(non-system, non-loop, non-CDROM),SystemDisks, andFormatablePartitions(empty partitions on system disks that are safe to format) - Handles NVMe (
nvme0n1p1), SCSI (sdb1), and eMMC (mmcblk0p1) naming
Disk Initialization Wizard (internal/storage/format.go)
A step-by-step UI at /settings/storage/init:
- Scan — Lists available disks with model, size, partition info
- Select — User picks a disk and enters a mount name (e.g.,
hdd_1) - Confirm — User types "FORMAZAS" to confirm destructive operation
- Format pipeline:
wipefs→sfdisk(GPT) →mkfs.ext4→blkidUUID → backup fstab → append UUID-based fstab entry → mount →findmntverification →chown 1000:1000→ createfelhom-data/andDokumentumok/subdirectories - Auto-registers new storage path in settings.json
- Smart partition detection: skips repartitioning for existing empty partitions
Safety guards: system disk detection, mount path conflict check, confirmation required, progress channel for real-time UI feedback.
System-disk partition formatting: When the system disk has an empty partition (no filesystem, not mounted, not used for /, /boot, /boot/efi, or swap), the init wizard detects it via FormatablePartitions in the scan result and offers to format just that partition. Uses IsSystemPartition() (granular per-partition check via fstab) instead of IsSystemDisk() (whole-disk block), so sda1 can be formatted while sda3 (root) remains protected.
Attach Existing Drive Wizard (internal/storage/attach.go)
A step-by-step UI at /settings/storage/attach for drives that already have a filesystem (e.g., a previously used ext4 drive). Unlike the init wizard, this does not format the drive — existing data is preserved.
Problem solved: Mounting a whole drive at /mnt/<name> would mix existing user data with the controller's directory structure (felhom-data/, Dokumentumok/, etc.). The bind-mount approach isolates the controller's working directory from other data on the drive.
- Scan — Lists available disks, filtered to partitions that have an existing filesystem (FSType != "")
- Mount raw — Partition is mounted read-only at a hidden staging path (
/mnt/.felhom-raw/<label>) - Browse — Directory browser shows the drive's contents. User can navigate and create a new folder (e.g.,
felhom_data) - Configure — User enters a mount name and display label. Warning: mount path is immutable until detached
- Finalize — Bind-mounts the selected subfolder at
/mnt/<name>. Two fstab entries are created (both withnofail):- Raw mount:
UUID=<uuid> /mnt/.felhom-raw/<x> <fstype> defaults,nofail,noatime 0 2 - Bind mount:
/mnt/.felhom-raw/<x>/<subfolder> /mnt/<name> none bind,nofail 0 0
- Raw mount:
- Sets permissions (
chown 1000:1000), createsfelhom-data/andDokumentumok/subdirectories - Auto-registers the storage path in settings.json + syncs FileBrowser mounts
Cancel at any point cleans up the temporary raw mount. The bind mount path (/mnt/<name>) is a real mount point, so all existing code (disk usage, IsMountPoint checks, etc.) works unchanged.
Storage Path Registry (internal/settings/settings.go)
Multiple external storage paths supported with:
- Label: Human-readable name (editable inline)
- Default flag: New deploys use this path by default
- Schedulable flag: Path appears in deploy dropdown
- Disconnected state:
Disconnected,DisconnectedAt,StoppedStacks— set by watchdog or safe-disconnect API, cleared on reconnect - Auto-discovery: On startup, scans deployed apps'
HDD_PATHvalues and registers unknown paths - Thread-safe CRUD: Add, Remove, SetDefault, SetSchedulable, SetLabel, SetDisconnected, ClearDisconnected
Data Migration (internal/storage/migrate.go)
Move app data between storage paths (e.g., SSD → HDD, HDD → new HDD):
- Validate: stack exists, deployed, has HDD data, target differs from source
- Estimate total size, check free space on target
- Stop the application
rsync -a --info=progress2per mount path with real-time progress parsing- Update
app.yamlHDD_PATH to new location - Start the application
- Rollback on failure: reverts config, restarts on old storage
Progress UI at /stacks/{name}/migrate with byte counter and percentage.
Stale Data Cleanup
After migration, the deploy page detects leftover data on previous storage paths:
- Shows path, size, and a delete button
- Two-step confirmation required
- Protected paths (
felhom-data/,felhom-data/appdata/,felhom-data/backups/,media/,Dokumentumok/) cannot be deleted
FileBrowser Mount Sync
When storage paths are added or removed, syncFileBrowserMounts() auto-regenerates FileBrowser's docker-compose.yml with volume mounts for all registered paths. It then recreates the container only when the generated config.yaml or compose actually changed (v0.82.0, F2) — gated by the pure helper fbNeedsRecreate(oldCfg,newCfg,oldCompose,newCompose), which compares the on-disk content captured before the writes against the final content read after them (so the integrations' ReapplyConfigForTarget edits count). When nothing changed (a controller restart, a no-op sync) it issues a plain up -d --remove-orphans that does not bounce the running FileBrowser. The restore-mode DB reset (down -v) still forces a recreate.
Storage Watchdog (internal/monitor/watchdog.go)
Continuously monitors registered storage paths for disconnection/reconnection (primarily USB drives):
- Probe loop:
ProbeStoragePath()callssyscall.Statfs()with 3-second timeout in a goroutine. Runs every 5s per connected path, 30s per disconnected path. - Debouncing: 3 consecutive probe failures required before declaring a drive disconnected (prevents false positives from transient I/O).
- Disconnect reaction (automatic, ~15s detection):
- Stops all deployed stacks whose
HDD_PATHis under the disconnected drive (skips protected stacks) - Persists
Disconnected,DisconnectedAt,StoppedStackstosettings.json - Lazy-unmounts stale VFS entries (
umount -l) — for attach-wizard drives, unmounts bind first, then raw - Fires alert refresh (red banner on all pages), notification (
storage_disconnected), and immediate hub report push
- Stops all deployed stacks whose
- Auto-reconnect (for UUID-based fstab entries):
- Checks
/host-dev/disk/by-uuid/<uuid>for device reappearance - Cleans stale mounts, then
mount -T /host-fstab <path>(raw + bind for attach-wizard drives) - Verifies with a post-mount probe
- Runs
restic unlockif stale lock files exist - Validates
StoppedStacks(filters to actually-stopped stacks), clearsDisconnectedflag - Fires alert refresh, notification (
storage_reconnected), hub report push
- Checks
Safe disconnect UI (manual, Settings page):
- "Leválasztás" button shown for USB drives (detected via sysfs symlink path containing
/usb) - Confirmation dialog lists affected apps
- Flow: stop apps →
sync→umount(fallbackumount -l) → mark disconnected → notification - Disconnected card: dashed border, red badge, timestamp, stopped apps list, "Csatlakoztatás" (reconnect) button
- After reconnect: "Alkalmazások indítása" button to restart auto-stopped stacks
USB detection (system.IsUSBDevice): Reads /host/sys/block/<disk> symlink — if target path contains /usb, it's a USB device. The removable sysfs flag is unreliable for USB HDDs (returns 0). USB drives show an orange "USB" badge on their storage card alongside Aktív/Alapértelmezett badges (v0.27.2). Handles findmnt bind-mount suffix stripping (/dev/sdb1[/subdir] → /dev/sdb1) for attach-wizard drives (v0.32.5).
Backup guards: Nightly DB dumps, restic snapshots, and cross-drive backups all skip disconnected, removed, and inactive drives with WARN log (not treated as failures). Cross-drive RunAppBackup() returns nil (not error) for unavailable destinations — prevents noisy error aggregation in scheduled runs (v0.32.5).
Tier2 destination unavailable (v0.32.5): When a Tier2 backup destination drive is disconnected, removed from storage, or deactivated (Inaktív), the backup page shows:
- Yellow status dot with "2. mentés szünetel" tooltip (not red)
- Warning badge: "Cél meghajtó leválasztva" (disconnected/removed) or "Cél meghajtó inaktív" (deactivated)
- Grayed-out last-run info and backup contents
- Hidden "Futtatás most" button (prevents futile manual triggers)
- "Beállítás" link preserved for reconfiguration
- Tier2 config persists — backups auto-resume when drive returns/reactivates
- Detection:
IsStoragePathKnown()catches removed paths,IsStoragePathSchedulable()catches inactive/disconnected/decommissioned
UI integration: Disconnected drives show with hatched red bars on dashboard, monitoring, and backup pages. Per-app backup rows show "Meghajtó leválasztva" badge. Health check emits warnings for disconnected paths.
5. Monitoring & Health
System Health Checks (internal/monitor/healthcheck.go)
RunHealthCheck() evaluates multiple subsystems and returns a HealthReport with status (ok/warn/fail):
| Check | Warning | Critical |
|---|---|---|
| Disk usage (SSD/HDD) | >= 90% | >= 95% |
| Memory | available < 512MB | available < 256MB |
| CPU temperature | >= 75C | >= 85C |
| Docker daemon | — | unreachable |
| Protected containers | — | not running |
| Storage paths | not a mount point (data on SSD), drive disconnected | path inaccessible, disk >= 95% |
Backup destination validation (CheckBackupDestination) has tiered checks:
- Path doesn't exist → critical/blocked
- Not writable → critical/blocked
- Same block device as root → warning (data on system drive)
- Disk >95% full → critical/blocked
- Disk >90% full → warning
Healthchecks.io Integration (deprecated)
Legacy pinger (internal/monitor/pinger.go) still runs for backward compatibility but is no longer the primary monitoring mechanism. Monitoring is now handled by the Hub event system (see Notifications). A deprecation log is emitted on startup if ping UUIDs are configured.
Metrics Store (internal/metrics/)
- SQLite with WAL mode for concurrent reads during collection
- System metrics: CPU%, memory (total/used/available), temperature, load average — collected every 60 seconds
- Container metrics: CPU%, memory, network I/O, block I/O per container
- Downsampled queries for chart time ranges (1h, 6h, 24h, 7d, 30d)
- 30-day auto-prune via daily scheduler job
Monitoring Page
Full-page system monitor at /monitoring:
- System Overview: hostname, OS, kernel, CPU model/cores, uptime
- System Metrics Charts: 4 line charts (CPU, Memory, Temperature, Load) in 2x2 grid
- Memory Distribution Bar: stacked bar showing per-container memory usage, OS/system overhead, and free memory (real-time from
/proc/meminfo+ container stats) - Container Resources: horizontal bar charts (CPU% and Memory per container)
- Per-container Detail: click-to-expand historical charts
- Hub Connection Status: shows Hub URL, customer ID, connection state (connected/unreachable), last successful push, last error
Chart.js 4.4.7 embedded locally (works in offline environments), dark theme matching site design.
Host (Proxmox box) Health — agent-proxied (slice 9, internal/agentapi + agent_host_metrics_handler.go)
The de-privileged controller (slice 8C) sees only its own cgroup and cannot read the host. The top card of /monitoring ("Szerver állapota (gazdagép)") instead shows the real Proxmox box, proxied from the host agent's GET /host/metrics:
- Host block: CPU% + load average, memory used/total, CPU/chassis temperature (or "n/a" when the hardware exposes no sensor — graceful-null), uptime.
- Per-storage capacity: a used/total bar per host storage target, with thin-pool fill (a full lvmthin pool corrupts every guest on it) and disk SMART temperature/wear.
Path: GET /api/host-metrics → Client.HostMetrics() (leaf-pinned, per-guest-token agentapi client) → agent GET /host/metrics. Host-wide and token-authed (assumption: one customer per host — the home-server model). It is a live fetch (a fresh agent collect, not the 15-minute hub snapshot), so the page polls it every 8 s while open. When the agent is unconfigured/unreachable the card shows a "nem elérhető" banner; the controller's own metric charts are unaffected.
Storage-bar ordering + labels (v0.57.0): the agent enumerates storages via pvesm in a
non-deterministic order, so the per-storage capacity list (#host-storage-bars) reordered on every poll.
enrichHostStorageTargets (agent_host_metrics_handler.go) sorts the response server-side —
user-data (usb/local-dir) → system+apps (lvmthin/lvm) → builtin local → backup
(pbs/nfs/cifs) → other, alphabetical by id within a tier — and attaches a friendly Hungarian
label + one-line purpose per entry (rendered by monitoring.html, with the raw PVE id shown muted).
Display labels only — the PVE storage ids are never renamed (vzdump/PBS configs reference them by
name). This is distinct from the server-rendered, user-data-only buildStorageBars "Tárhely" list.
Alert System (internal/web/alerts.go)
State-based alerts displayed on all pages:
- Sources: health issues, Hub connection status, backup disabled, storage disconnected, update available
- Hub alerts:
hub-disabled(warning) when Hub not enabled,hub-unreachable(error) when last push failed and no success in 30 min - Sorted by severity (error > warning > info), capped at 5 visible
- Refreshed every 5 min + on startup + on storage state changes
6. Notifications
Hub Event System (internal/notify/notifier.go)
The controller pushes structured events to the Hub's /api/v1/event endpoint. The Hub handles notification dispatch, cooldown management, and dead man's switch detection.
Core method: PushEvent(eventType, severity, message, details) — non-blocking goroutine, 2 retries with 3s backoff, never blocks the caller.
⚠ THE SEVERITY VOCABULARY IS THE HUB'S, AND IT IS EXACT:
{info, warning, error, critical}. The hub coerces anything else toinfoat ingest, andinfois dropped byseverityNotifiesbefore both delivery legs. So a severity outside that set means the event is stored, the POST returns200, the dashboard shows it — and it is e-mailed to nobody.This shipped twice:
DiskAlertKind.Severitysentwarnuntil v0.215.0, andapp_start_failedsent it until v0.223.0 — 91 of those events were stored and not one was ever delivered. It is now pinned by an AST walk over the whole controller (TestR329_EveryEmittedSeverityIsInTheHubVocabulary); the six call sites that pass a variable are registered by name, so a new one fails the test. Since v0.107.0 the hub also logs aWARNnaming any severity it had to rewrite. Full contract:felhom.eu/documentation/architecture/08-alarm-ladder.md§6.1.Who receives what.
processOperatorconsults only the operator switch, the address and a one-hour cooldown — never customer preferences — so a valid severity always reaches the operator. The customer leg additionally consultsoperatorOnlyEventsand the customer's own enabled events.
Event Types
| Event Type | Severity | Trigger |
|---|---|---|
backup_failed |
error | Nightly restic backup fails |
db_dump_completed |
info | Nightly database dumps succeed |
db_dump_failed |
error | Nightly database dumps fail |
backup_integrity_ok |
info | Weekly restic check passes |
backup_integrity_failed |
error | Weekly restic check fails |
crossdrive_completed |
info | Cross-drive secondary backup succeeds |
crossdrive_failed |
error | Cross-drive secondary backup fails |
health_degraded |
warning | Health status degrades (ok→warn) |
health_critical |
error | Health status critical (any→fail) |
health_recovered |
info | Health status recovers (fail/warn→ok) |
disk_warning |
warning | Disk usage crosses 90% |
disk_critical |
critical | Disk usage crosses 95% (this row read error until v0.223.0; the emitter is fillwatch.Band.Severity() and it has always sent critical) |
storage_disconnected |
error | Storage drive physically removed |
storage_reconnected |
info | Storage drive reconnected |
controller_started |
info | Controller process starts |
controller_updated |
info/error | Self-update success or failure |
app_deployed |
info | New app deployed via API |
app_removed |
info | App removed via API |
app_start_failed |
warning | A DEPLOYED app is not running (fix-3) — fired ONCE per running→down transition. Customer-switchable („Alkalmazás nem fut"), OFF by default; the OPERATOR is e-mailed regardless. Was warn until v0.223.0 — see the severity note below |
disaster_recovery_started |
warning | DR restore begins |
disaster_recovery_completed |
info/error | DR restore finishes (success/partial) |
Each event carries typed detail structs (e.g., BackupDetails, DiskDetails, HealthDetails) serialized as JSON.
Deployed-app-down alerting (fix-3, v0.120.0, CAMPAIGN-3). A deadapp-check scheduler job (every
30 s, after a 90 s boot grace) scans stackMgr.GetStacks(): a DEPLOYED app whose containers are
exited/degraded (stacks.IsDownState minus the stopped exclusion added in v0.164.0 — see below;
a Docker created/dead container, the F11 dead-at-boot case, resolves to exited) gets a
state-based WARN dashboard banner ("Telepített alkalmazás nem
fut: ", grouped above 3 so a reboot storm doesn't wall the dashboard) that self-clears when the
app runs again, AND an app_start_failed hub event fired once per running→down transition
(Notifier.NotifyAppStartFailures tracks per-app state; down→down cycles are silent — the hub owns
the real cooldown, the controller adds no timer). The boot grace prevents false alarms during the
controller's own startup while STILL firing for an app that never came up. This closes the campaign's
4-hour silent CWA death.
Dead-primary alerting (R-51, v0.156.0). fix-3 above only ever saw stacks that were entirely
down. A multi-container app whose MAIN container died while its helpers kept running aggregated to
StateRunning ("partial") and therefore alerted on nothing — immich-server was Exited for 18 h,
the app 100 % unreachable, with no banner and no event (F4, AUDIT-vacation-remote-ops-2026-07-20).
aggregateState's mixed branch now inspects each DOWN member's docker restart policy: always /
unless-stopped means docker was supposed to be keeping it up, so the stack becomes
StateDegraded — a down state, so the existing banner and the existing app_start_failed event
fire unchanged. no / on-failure is a finished one-shot init/migrate container and stays benign.
An unreadable policy counts as supervised (fail-closed: a member is known dead, only the excuse is
missing). The unhealthy / restarting / paused / unknown exclusions are untouched — folding
unhealthy into down is precisely the flapping fix-3 avoided. UI: „Részlegesen leállt", warn
colour, counted with the stopped apps, URL flagged unpublished (Traefik withholds the route when the
routed member is the dead one). Policy reads are one docker inspect per down member of a mixed
stack, cached per container+state.
Deliberate stops are silent (v0.164.0). Stopping an app from the UI (Leállítás → StopStack →
docker compose down → zero containers → the deployed stack aggregates to StateStopped) is the
user's own action, not a fault, and must not raise the banner OR the app_start_failed email. The
scan's pure core was extracted to classifyRunStates([]stacks.Stack) and its down predicate is now
stacks.IsDownState(st.State) && st.State != stacks.StateStopped — the SINGLE fix-3 derivation point,
so StateStopped is dropped from both the banner dead-list and the notifier Down-set at once (the
launcher tile still shows greyed „Leállítva"; the monitoring page and dashboard counters are factual
display, not alarms, and are unchanged). This rests on two invariants: I1 — a UI stop always ends
at StateStopped (compose down removes the containers); I2 — the P2 restart-policy census
(53 templates / 78 services, all unless-stopped) means a crashing app never comes to rest at
stopped, so faults still surface as exited/degraded/restarting/unhealthy. If either
invariant changes, revisit the suppression.
C9-F2 (v0.183.0) — the
restartinghalf of that sentence was a wish, not a fact.restartingwas named above as a state through which faults "still surface", but it was in no down set at all:IsDownStateexcludes it, so a crash-looping app raised no banner, noapp_start_failed, no email and no hub event — andunless-stoppedmeans Docker retries forever, so the silence was permanent. Campaign 9 watched docmost loop for nine minutes while the F-OBS heartbeat printed „180 scans since boot, 4 deployed app(s) evaluated, 0 currently down".The fix does not add
StateRestartingtoIsDownState— that alarms on every deploy and update fleet-wide. A SUSTAINED restarting run becomes down afterstacks.crashLoopAfter(5 min), chosen above the deploy flow's 120 s health timeout, Mealie's 60 sstart_periodand R-97b's 180 s quiesce grace, so the suppression windows compose into one bounded delay rather than leaving a gap. Carried byStack.RestartingSince(stamped inrefreshStatusLocked, cleared on any other state, not persisted) and read viaStack.CrashLooping(now)— used by BOTH the alarm and the dashboard "how many of my apps work" counter, which previously countedrestartingas running and so contradicted the alarm on the same screen. Pinned bycrashloop_classify_test.go; the test that a brief restart stays silent is the one that fails against the naive fix.
IsDownState itself is deliberately UNCHANGED (other
callers rely on stopped counting as down). An out-of-band docker compose stop leaves the containers
present → StateExited → still alerts, which is correct (out-of-band tampering is reportable).
R-330 (v0.224.0) — the alarm must be told by EVERY mechanism that stops an app, and until now it was told by one. The controller stops a customer's app on purpose in two quite separate places: the quiesce loop, for the whole-guest (vzdump/PBS) backup, and the
AppStopGuardpaths — the nightly volume dump, an off-site reconstitution and a.fabexport. R-97b built the suppression window for the first and it works. The second registered with nothing, soclassifyRunStatesnever knew, and the nightly backup e-mailed the customer „Telepített alkalmazás nem fut" about apps it was holding down itself.Measured on
demo-hp2026-08-30 (controller 0.223.0):DumpAppVolumesSafeholds each stack down ~13 s whiledeadapp-checkscans every 30 s, so the scan caught whichever stack was mid-cycle — 3 events at the 02:30 CESTdb-dumpleg, 2 more at the 04:15offbox-backupleg, every night, 61 e-mails, while every other scan that day logged0 currently down.
classifyRunStates'squiescedargument is now the union of both sets (unionSuppressedinscanDeployedAppRunStates). A third way to stop an app means a third set here — that omission is the whole of this defect. The window cannot latch: a restart that was attempted and broke callsReleaseFailedand alarms on the next scan,Beginreplaces the previous operation's set, and a 6 h backstop covers a hold nothing released. Suppression is not persisted — after a crash the guard holds nothing,Recover()either brings the app back or leaves it genuinely down, and a down app must alarm.
Boot desired-state reconciliation (R-52, v0.156.0, internal/bootrecon; rebuilt on recorded intent
in R-166, v0.189.0). A deployed: true app that missed its boot start used to stay down until a
human noticed — the same shutdown that produced F4 left immich and calibre-web Exited while ten
sibling containers came back, and they were still down 18 h later (F5). At startup (5 s after the
quiesce and app-stop recoveries, so the three never race) the controller performs one bounded
sweep: every deployed, non-protected, not-mid-deploy stack that is down gets StartStack, at most
2 attempts 30 s apart, then it stops and the alarm owns the problem. Never a restart loop. The
whole sweep fits inside the 90 s boot grace, so a successful recovery is silent and a failed one
still alerts. Outcome is logged per attempt at INFO; no new hub event (the existing alarm is the
escalation).
What "down" means here changed in v0.189.0. Until then the sweep required the stack to still HAVE
containers, because the UI's Stop is compose down (which removes them) and "zero containers" was
read as a deliberate stop. That inference was wrong in two silent ways: a power cut mid-compose
and an interrupted deploy also leave zero containers, and both were skipped as "the customer
stopped it" and left down indefinitely. Since R-166 the sweep reads the customer's recorded
intent (desired_state in app.yaml) instead:
desired_state |
containers | result |
|---|---|---|
stopped |
any | never started — the customer said so, and no observation overrides it |
running |
0 | recovered — the power-cut / interrupted-deploy case, invisible before v0.189.0 |
running |
>0 and down | recovered (unchanged) |
running |
>0 and up | left alone |
| absent (legacy) | 0 | not started — byte-identical to the pre-v0.189.0 behaviour |
| absent (legacy) | >0 and down | recovered — byte-identical to the pre-v0.189.0 behaviour |
Absent means UNKNOWN, never "running". Every app.yaml written before v0.189.0 lacks the field,
so absent is what an upgraded box reads for every app that has not been started or stopped since;
reading it as "running" would start every deliberately-stopped app on the first boot after the
upgrade. Where intent is unknown the sweep falls back to the old inference rather than inventing an
answer, and a running-only startup backfill converges the unambiguous cases (deployed and
observed up) without waiting for a button press. stopped is never backfilled from any signal.
The sweep no longer looks only once (R-157 mechanism A, v0.190.0). It used to sample the fleet at
T+5 s and return; at that moment docker is still restoring containers after a hard reset, so an app
that had not yet settled into a down state was never a candidate — measured failing on three of six
hard resets. It is now a settle-then-sweep window: the fleet (name, state, container count) is
sampled every 5 s, called settled after 3 identical samples, and swept once, at the end, on a
settled fleet. The window ends on whichever comes first — settled, or a 50 s budget — and the log
says which. Sampling is read-only and there is still exactly one sweep with its per-app attempt bound
intact: this widens a bounded window, it does not remove the bound. settle (5 s) + budget (50 s) + one 30 s retry = 85 s stays inside the 90 s deadAppBootGrace, which is what keeps a successful
recovery silent; a window that genuinely overruns emits a LATE RECOVERY WARN naming the apps rather
than the grace being widened to hide it.
The sweep asks before it starts (R-171, v0.190.0). Three things legitimately hold an app down,
and starting it would be the wrong repair for all three: its data drive is absent (compose would
create the bind sources on the guest rootfs — the hazard the drive gate exists to prevent), a
quiesce is holding it for a whole-guest backup, or an app-data operation (volume dump, offsite
restore, .fab export) is holding it. All three are refused through one seam, reusing the signals
their owners already publish. Fail-safe: a drive whose liveness cannot be determined is treated as
absent. Held apps are reported separately from StillDown — they are not a fault the sweep failed
to fix, and reporting them as one is a false alarm. This closed a regression v0.189.0 introduced:
before it, a drive-gate-stopped app read as running + zero containers, so the sweep started it,
burned both attempts and handed it to the dead-app alarm.
Both boot gates read intent (R-170, v0.190.0). There are two: the R-52 sweep above, and the
drive-backed boot recreate gate (shouldRecreateOnBoot, internal/web/intermediary.go), which
re-creates a drive-backed app onto its re-propagated drive after a guest reboot. Until v0.190.0 the
second still ended in && hasContainers, so the two disagreed about the same question. It now uses
the identical three-way table — stopped → never, running → recreate whatever the container count,
absent → exactly the old hasContainers behaviour. Its drive-presence term is untouched and still
load-bearing: an app whose drive is absent is never recreated there either.
Desired state — who owns it (R-166, v0.189.0). app.yaml gains desired_state, a tri-state
"" / running / stopped. It is written by the customer's own action and nothing else: the
/api/stacks/{name}/{action} switch (start/restart/update → running, stop → stopped),
DeployStack, UpdateOptionalConfig's redeploy branch, and the .fab import. StartStack and
StopStack are deliberately not writers — a census found 14 callers of which only 2 are the
customer, and recording intent in the primitive would make a nightly backup indistinguishable from
the customer pressing Stop, which is the confusion the feature exists to end. Intent is written
before the act, and an action whose intent cannot be recorded is refused.
Interrupted app-data operations (R-166, v0.189.0, backup.AppStopGuard). A volume dump, an
off-site reconstitution and a .fab export all stop an app, work on its data, and start it again.
A controller killed inside that window left the app down with nothing on disk recording why or that
it was owed a restart. A persisted marker (<data_dir>/appstop-state.json — its own file, never
quiesce's, so one file has one writer) is now written before the stop and cleared only after a
restart that succeeded; a failed restart deliberately keeps it. At startup Recover() restarts the
recorded apps, clears the marker, and its outcome is reported to the operator on the existing
backup_failed event — an interrupted operation means the backup did not complete. The defer in
those functions is not the mechanism: a SIGKILL runs no deferred function (Campaign 8 fault 10, on
live hardware), which is exactly what the marker covers.
Default Enabled Events
Events the customer receives notifications for (configurable in settings):
backup_failed, db_dump_failed, disk_warning, disk_critical, storage_disconnected, node_down, health_critical, expected_backup_missed, expected_dbdump_missed
Preference Sync
Notification preferences (email, enabled events, cooldown hours) are:
- Stored locally in
settings.json - Synced to Hub on save and on controller startup via
POST /api/v1/preferences - Hub sync failure doesn't block local save
Empty-email save guard (v0.137.0): the save handler REFUSES a submit with a blank e-mail box while any event is still enabled (it would store an empty address AND push it to the hub, wiping the provisioning-seeded alert delivery). The form re-renders with a Hungarian error and the customer's ticked events preserved; no save, no sync. Clearing the e-mail with zero events enabled is allowed (an intentional turn-everything-off).
7. Update Management
App Catalog Sync
- Periodic
git fetch+git reset --hardof the app catalog repo - Content-hash comparison prevents unnecessary file writes
- Post-sync stack rescan detects new/changed apps immediately
- Stale lock recovery: automatically removes
.git/index.lock,.git/shallow.lock, and.git/HEAD.lockbefore each fetch — prevents permanent sync failures after interrupted operations (e.g. container restart mid-sync)
Planned Update Classifications
| Marker | Behavior |
|---|---|
| No marker | Optional — shown on dashboard, customer clicks "Update" |
UPDATE_REQUIRED=true |
Mandatory — auto-applied during next update window |
UPDATE_SECURITY=true |
Critical — applied immediately |
Controller Self-Update (internal/selfupdate/)
The controller can update itself to the latest registry version with a one-click Settings button. In the
LXC architecture there is no in-container compose to drive (the old docker compose -f /opt/docker/felhom-controller/docker-compose.yml up -d path does not exist in the guest — it produced
"docker-compose.yml nem elérhető"). Instead (Phase 1, v0.85.0) the controller pulls the target image
in-guest and delegates the container swap to the host agent, which owns the restart + health-verify +
rollback.
How It Works
1. Check Gitea Docker Registry V2 API for the latest semver tag (queryRegistry, BasicAuth).
2. If newer than current (ldflags Version): docker login --password-stdin → docker pull <image>
→ docker logout — IN-GUEST over the shared docker socket (token via stdin, never argv).
3. Delegate to the host agent: agentapi.SwapController → POST /controller/swap {image} (202). The agent
(external to this container) rewrites /etc/felhom-controller-image, restarts
felhom-controller-bootstrap.service, polls the new controller to healthy, and ROLLS BACK to the
previous image if it doesn't come up. The controller never docker-rm/recreates itself.
4. On startup the new container reads update-state.json → VerifyStartup marks success (current==target)
or failure (rollback → version mismatch). The Settings button polls /api/health and reloads.
The button is latest-only (no version picker) and opt-in. No host agent wired (un-provisioned guest) → self-update unavailable.
Phase 2 — managed updates: the version FLOOR (v0.86.0)
On top of the opt-in button, the controller now honors an operator-enforced minimum version (FLOOR).
The hub returns the customer's effective floor (per-customer override else a global default) on the
report ACK (min_controller_version, alongside latest_version). The controller's report pusher
(internal/report/pusher.go, OnPushResponse) hands the floor to the updater (SetFloor) and calls
MaybeAutoUpdate() — on the existing report cycle, no new timer/endpoint:
- If the box is below the floor it auto-updates to the floor (not latest) by reusing the Phase 1
flow above (
performUpdate,initiatedBy="auto-floor") — same pull → agent swap → rollback. No customer click. - At/above the floor: nothing (it does not chase latest — that's the button's job).
- Guards: dev build / no agent / backup running → skip; floor must be pullable (floor ≤ latest
available; floor > latest → warn + do nothing); one attempt per below-floor condition (in-memory flag
- persisted
update-state.json) → no flapping/storm.
- persisted
- Settings UI shows "Minimális verzió (üzemeltető): X" and, during an auto-update, the same restart-poll panel as the button.
The floor is the auto-target (the operator raises it for a controlled fleet rollout); latest stays
the customer's manual opt-in. Floor source + operator UI are hub-side (felhom-hub v0.15.0). No agent
change — Phase 2 reuses the Phase 1 POST /controller/swap.
Pull-based config-refresh (v0.94.0)
The same report ACK also carries a per-customer config_version (a hub-side stored counter,
bumped on every config save; felhom-hub v0.26.0). This is how an operator config edit reaches a running
box — the hub never connects into the box (it replaced the retired inbound "Push Config"). Wired in
OnPushResponse beside the floor reconcile (internal/report/config_refresh.go, ConfigRefresher.Reconcile):
- Changed vs. the last-applied version (
settings.applied_config_version) →bootstrap.RefreshConfigre-pullscontroller.yamlfrom the hub and rewrites it, re-merging the per-guestlocal_apifrombootstrap.json(reuses the first-boot pull machinery; overwritescontroller.yamlsince the hub is its source of truth; never touchessettings.json) → record the new version → graceful self-restart (api.GracefulSelfRestart→os.Exit(0)→ Dockerrestart: unless-stoppedre-reads the new config). - First-ever ACK (nothing recorded) → record the baseline without restarting (the first-boot pull already has the current config).
- Unchanged version → no-op (so no restart storm — after a refresh
applied == ACK). - Failed pull → keep the current config, do not record/restart, retry next report cycle.
- Record-before-restart so the restarted process sees the version applied and doesn't loop. Only the
felhom-controllercontainer restarts; customer app stacks are untouched. (A config apply rotatesweb.session_secret, so dashboard sessions are invalidated — same as the old Push-Config path.)
Design Philosophy
- No automatic rollback — follows the Watchtower pattern (24k+ GitHub stars, no rollback). Docker's
restart: unless-stoppedpolicy is the crash safety net. The Hub's dead man's switch detects when the controller goes down. - Audit state file —
update-state.jsonin the data volume records every update attempt (previous version, target version, initiator, result). Operators can SSH in and revert usingPreviousImagefrom this file. - Backup-aware — refuses to start an update while a backup is in progress (
backupRunning()guard).
Package Structure
| File | Purpose |
|---|---|
version.go |
ParseVersion("X.Y.Z") → Version{Major,Minor,Patch}, Compare() returns -1/0/1. Hand-rolled, no external deps. Rejects "dev" and "latest". |
state.go |
UpdateState struct persisted as JSON. LoadState(), SaveState() (atomic: .tmp + rename), ClearState(). Status values: "pending", "success", "failed". |
updater.go |
Core Updater struct. Registry check via HTTP GET to gitea.dooplex.hu/v2/admin/felhom-controller/tags/list with Basic Auth (git username/token). Update trigger: docker pull → compose file regex replace → docker compose up -d. Thread-safe with sync.Mutex. |
Update Trigger Flow
- Guard checks: concurrent update lock, dev version check, backup running check, compose file accessible
- Write
update-state.jsonwith status"pending"(audit trail) docker pull <image>:<targetVersion>- Read compose file → replace image tag via regexp → atomic write (
.tmp+ rename) docker compose -f /opt/docker/felhom-controller/docker-compose.yml -p felhom-controller up -d- Docker kills the current container, starts the new one
Startup Verification
Called once from main.go before the scheduler starts:
- Load
update-state.json— if missing or status !="pending", nothing to do - Compare running
Versionwithstate.TargetVersion - Match → mark
"success", notify via hub - Mismatch → mark
"failed", notify via hub - No rollback attempt — operator reverts manually if needed
Auto-Update Scheduling
Two separate scheduler jobs prevent interference with backups:
| Job | Type | Default | Purpose |
|---|---|---|---|
selfupdate-check |
sched.Every |
6h | Check registry, cache result (for UI). Never triggers update. |
selfupdate-auto |
sched.Daily |
04:30 | If auto-update enabled + update available + backup not running → trigger. |
The auto-update time (config.SelfUpdate.AutoUpdateTime, default "04:30") is deliberately separate from the backup window (02:30-~04:00) to avoid collisions. The backupRunning() guard is the hard safety check — if backups run long past 04:30, the update is skipped and retried the next day.
An initial version check fires 30s after startup so the Settings page shows version info quickly.
Compose File Access
The controller needs write access to its own docker-compose.yml. This is achieved via Docker volume mount ordering:
volumes:
# 1. Directory mount — gives access to compose file + config
- /opt/docker/felhom-controller:/opt/docker/felhom-controller
# 2. Read-only override — prevents accidental config writes
- /opt/docker/felhom-controller/controller.yaml:/opt/docker/felhom-controller/controller.yaml:ro
# 3. Named volume override — persistent data in Docker-managed volume
- controller-data:/opt/docker/felhom-controller/data
API Endpoints
| Method | Path | Auth | Description |
|---|---|---|---|
| GET | /api/selfupdate/status |
Session or API key | Current status (cached, no network call) |
| POST | /api/selfupdate/check |
Session or API key | Force registry check, return result |
| POST | /api/selfupdate/update |
Session or API key | Trigger update (async, returns immediately) |
Self-update endpoints accept either session auth (for UI) or hub API key as bearer token (for external triggering from build scripts or hub). This enables the post-v0.16.0 deploy workflow:
# After building + pushing new image:
curl -s -X POST https://felhom.demo-felhom.eu/api/selfupdate/update \
-H "Authorization: Bearer <HUB_API_KEY>"
Settings Page UI
The "Verzió és frissítés" card on the Settings page (/settings) shows:
- Current version and latest available version
- "Frissítés elérhető" (update available) badge
- Last check time and any errors
- Registry mode line (v0.112.0): "Registry: nyilvános (hitelesítés nélkül)" vs "Registry: hitelesített" — credential-less is a supported mode, not an error state
- Auto-update status with configured time
- Last update result (success/failed/pending)
- Buttons: "Frissítés keresése" (check) + "Frissítés telepítése" (apply)
Registry access modes (v0.112.0)
Git Sync credentials (git.username/git.token) are optional — for private catalogs only; version
discovery and self-update work without them:
- Anonymous (both empty):
queryRegistryperforms the Docker Registry v2 anonymous token dance — plain GET → 401 withWWW-Authenticate→ token fetched from the ADVERTISED realm (parsed from the header, never hardcoded — registry-agnostic) with no credentials → Bearer retry.pullImageskipsdocker loginentirely (docker's native anonymous flow covers public packages). - Authenticated (both set): the previous BasicAuth + login/pull/logout path, unchanged.
- Half-configured (only one set): loud incomplete-credentials error — never a silent anonymous downgrade.
- A registry that genuinely denies anonymous access surfaces "registry denied anonymous access — a private registry requires Git Sync credentials".
After triggering an update, the page polls /api/health every 3s and reloads when the new container responds.
A global info-level alert ("Új controller verzió elérhető") appears on all pages when an update is available, linking to the Settings page.
Configuration
self_update:
enabled: true
check_interval: "6h" # How often to check registry
image: "gitea.dooplex.hu/admin/felhom-controller" # Default
auto_update: false # Set true for unattended updates
auto_update_time: "04:30" # When to auto-apply (after backups)
health_timeout_seconds: 60 # Reserved for future use
Edge Cases
| Scenario | Behavior |
|---|---|
Version == "dev" |
ParseVersion returns error → no updates reported, trigger refused |
| Registry unreachable | Log warning, return error in check result. No crash. |
| No registry credentials | Return error "Registry hitelesítő adatok hiányoznak" |
| Compose file not writable | Refuse update before doing anything |
| Backup running | Refuse with "Mentés fut, próbálja később" |
| Concurrent update | Mutex prevents duplicates: "Frissítés már folyamatban" |
| Bad update (crash loop) | Docker restarts container. State file stays "pending". Operator SSH-reverts using PreviousImage. |
| Corrupt state file | Treated as "no pending update", logged, deleted |
8. Authentication & Settings
Customer-claim gate (internal/web/claim.go, v0.122.0 — closes DRILL-day0-vm F-4/F-5)
The dashboard password is customer-owned, set through a one-time claim code the hub emails to
the registered address (no operator-set path, no open-until-set window). This closes the fresh-box
race where a new felhom.<domain> cert appears in CT logs minutes before any password exists.
- States (precedence): a SET password (settings→config) always wins — the gate never shows.
Else a delivered claim-code hash + not-yet-claimed → GATED: every route serves the claim
page (
302 → /claim) or401JSON (API); only/claim*,/static/*,/api/healthpass. Else (no password, no hash) → legacy-open with a red transition banner until the hub delivers a hash (transitional only, never the fresh-box state). - Claim/reset flow:
GET /claim(code + new password ×2, min 12) →POST /claimverifies the code (bcrypt match AND generation not yet consumed AND ≤ 72 h old), sets the customer's password viasettings.SetPasswordHash, marksClaimed(set-only), consumes the generation (single-use), invalidates sessions, issues a fresh one.POST /claim/request-new-code(the "Új kód kérése" / login-page "Elfelejtett jelszó") forwards to the hub, which emails a fresh code to the registered address only. Reset rides the same page (a claimed box reaches/claimpre-auth). - Anti-brute-force: per-source + global counter, 5 failures → 15-minute lockout (both scopes),
raising the allowlisted
claim_lockoutevent. Pre-auth CSRF is an HMAC overweb.session_secret(fixes the CTRL-007 bare-double-submit weakness). The per-source key is the client IP resolved by the sharedclientIP(r)helper — XFF first-hop, elseRemoteAddrwith the ephemeral port stripped (v0.129.0 F-B; keying on the rawRemoteAddrlet distinct direct connections evade the counter). The login form and the escrow wizard re-auth share the same helper/key. - Delivery: the hub bakes
web.claim_code_{hash,generation,issued_at}into the Day-0 controller.yaml (gate-from-first-boot) and serves the freshest state in the report ACK (report/claim_sync.gocaches it idempotently by generation — newer advances, same/older/nil never rewrites, a hub outage never clears). The report carriesclaimed(hub ingests set-only). - Escape hatch:
felhom-controller --print-reset-codeprints a one-time local code (generation above cached/baked/consumed); the same gate consumes it. Root-gated bydocker execreachability. It runs as a SEPARATE PROCESS, which is whyeffectiveClaimCodeREADS THROUGH to the persisted state (v0.198.0, R-204 item 1). Until then the hatch persisted a new code while the running server kept the old one cached, so the minted code was refused until the controller was restarted — and nothing said so. The read-through (settings.ReloadClaimCode) is on the claim path only, and only while the box carries no password (claimGateActivereturns onauthEnabled()first). Deliberately not a watcher, a signal handler or a TTL: a TTL would leave a window in which a SUPERSEDED code still works, which is worse than the bug. Refreshes hash/generation/issuedAt only — neverClaimConsumedGeneration, which this process alone writes and must stay monotonic.effectiveClaimCodereturns an error and every caller fails closed (the gate stays UP, the claim is refused); an ABSENT settings file is not an error, since a pre-first-save box falls back to the controller.yaml bake.
Session Auth (internal/web/auth.go)
- bcrypt password verification with configurable source priority:
settings.json→controller.yaml→ no auth (open access) - 7-day session duration with random 32-byte hex tokens
?next=redirect after login preserves the page the user was visiting- Session cleanup every 15 minutes
- All sessions invalidated on password change
- Conditional logout link (hidden when auth is disabled)
- Each session stores a dedicated CSRF token (separate 32-byte random value) alongside the session token
CSRF Protection (internal/web/csrf.go)
Synchronizer-token CSRF protection on all browser-facing state-mutating endpoints.
How it works:
CsrfProtectmiddleware wraps all route handlers inmain.go- Safe methods (GET, HEAD, OPTIONS) pass through without validation
- For POST/DELETE/PATCH: reads token from
_csrfform field orX-CSRF-Tokenrequest header; constant-time compares against the session's stored CSRF token - On rejection: JSON
{"ok":false,"error":"CSRF token missing or invalid"}for/api/paths; HTTP 403 text page for UI routes - Logs:
[WARN] CSRF rejected: METHOD /path from addr (reason)
Exempt paths (no CSRF check):
- Requests with
Authorization: Bearer ...header — hub→controller API calls (selfupdate, config/apply). Browsers cannot auto-send Bearer headers, so cross-site requests are impossible on these endpoints. - Auth-disabled mode (
authEnabled() == false) — CSRF is meaningless when there is no session.
Token delivery to templates:
executeTemplate(w, r, name, data)wrapper inserver.goauto-injectsCSRFField(template.HTMLhidden<input>) andCSRFToken(raw string) into every page's data maplayout.htmlemits<meta name="csrf-token" content="{{.CSRFToken}}">and definescsrfHeaders()JS function in<head>(before page scripts)- Forms:
{{.CSRFField}}(or{{$.CSRFField}}inside{{range}}loops — outer scope required) - JS
fetch()calls:headers: csrfHeaders()— returns{'X-CSRF-Token': metaContent} - Dynamically-created JS forms: read token from
document.querySelector('meta[name="csrf-token"]').content navigator.sendBeacon()replaced withfetch(..., {keepalive: true})where used —sendBeaconcannot send custom headers
Settings Persistence (internal/settings/settings.go)
Runtime-mutable settings in settings.json (separate from infrastructure config):
| Section | Contents |
|---|---|
password_hash |
bcrypt hash override |
notifications |
email, enabled events, cooldown hours |
db_validations |
per-DB dump validation results (survives restarts) |
app_backup |
per-app map: enabled flag, cross-drive config (method, dest, schedule, runtime status) |
storage_paths |
registered paths with label, default flag, schedulable flag, disconnected state |
cross_drive_restic_password |
auto-generated restic password for cross-drive repos |
All public methods use sync.RWMutex. File writes are atomic (.tmp + rename).
Settings Page (/settings)
Five sections:
- System config — read-only display of
controller.yamlvalues - Version & update — current/latest version, check/update buttons, auto-update status, last update result
- Storage paths — add/remove, edit labels, set default, toggle schedulable, per-path app list with sizes, safe disconnect/reconnect for USB drives
- Password change — current + new + confirm, min 8 chars
- Notifications — email, event checkboxes, cooldown hours, test email button
9. Central Hub Reporting
Report Push (internal/report/)
Periodic JSON push (default every 15 min) to the central felhom-hub service:
- System: hostname, OS, CPU, memory, disk usage, uptime
- Containers: running/stopped counts, per-container CPU/memory
- Backup: last DB-dump run, success (disk-tier backup is the host agent's; restic password is no longer reported — removed v0.69.0)
- Health: current status, issues, warnings
- Stacks: deployed apps with versions and states
- Config hash: SHA256 of
controller.yamlfor Hub-side config comparison - Geo-restriction (always present, v0.70.0):
geo_restrictionis always populated —Enabled=falsewith an empty country list when never configured — so the Hub always renders the geo section ("Inaktív" when off) instead of hiding it.buildGeoRestrictionReportininternal/report/builder.go. - App telemetry (v0.28.0+): Per-stack memory (current/avg/peak) and CPU averages from the last 15 minutes of metrics data, plus log scan results (error/warning counts with deduplicated issues). Only non-protected, deployed stacks are included. Backward-compatible: old Hub versions silently ignore this field.
- Controller telemetry (v0.32.4+): The controller's own container (
felhom-controller) is included as a special entry in theapp_telemetryarray. Its memory/CPU metrics come from the same metrics collector, and its log warnings/errors are scanned viadocker logsusing the same pipeline as app containers. This reuses all existing Hub telemetry infrastructure (memory trend charts, known issues, fleet aggregation) with zero Hub-side changes. - DR recipe — customer + apps half (v0.73.0):
dr_recipeis the controller half of the secret-free reconstruction recipe (SPIKE-dr-recipe-2026-06-16.md) that complements escrow (keys) + PBS/restic (bytes).BuildDRRecipeAppHalf(internal/report/dr_recipe.go) emits{recipe_version, customer{id,display,domain}, apps[]}; each deployed, non-protected app contributesAppRecipe{catalog_ref, enabled, storage_bindings}where bindings are parsed from the compose (${HDD_PATH}/${USERDATA_PATH}volume binds →{container_path, drive, subpath}, e.g. romm →felhom-flash:userdata/roms). THE BOUNDARY: the emitter is the enforcement point — it ships an explicit allowlist of those three fields and reads NOTHING fromAppConfig.Env, so noENC:/token/password can leak (allowlist, not denylist → new fields excluded by default). The load-bearingTestBuildAppRecipe_NoSecrets+ its red-proof companion live here. The hub assembles this half with the agent's storage/guest/PBS half into one customer recipe.recipe_version=1, ignore-unknown on read.
Bearer token authentication, 3-attempt retry with 5-second backoff. Push status tracked via PushStatus struct (LastAttempt, LastSuccess, LastError, consecutive failures) — used by the monitoring page and alert system to show Hub connection health.
Immediate out-of-cycle report on user actions (v0.139.0, generalizing the v0.70.0 geo push): besides the periodic cycle, user actions with hub-side effects fire a debounced, coalescing out-of-cycle report push (report.Trigger in internal/report/trigger.go: buffered-1 signal channel + single worker; quiet window 2 s, min spacing 15 s, trailing-edge — a burst coalesces to ≤ 1 + ceil(burst/15 s) pushes and the LAST state always reaches the Hub). One canonical fire closure in main.go does the full BuildReport+Claimed+Push; the trigger adds NO retry of its own (the Pusher owns retries) and every failure degrades to the 15-min cycle, which stays the reconciliation backbone. Wired call sites: geo settings save/manual sync + app deploy/remove/delete (api.Router.reportPushNow), and via the web.Server.SetReportTrigger seam (reportTriggerNow, fired only AFTER a successful local commit): escrow recovery-code claim (the ACK hash-match flips pending→escrowed in seconds), notification-prefs save, app-email toggle, offsite target config + per-app offsite toggle, customer claim completion. hub.enabled: false → the seams stay nil (strict no-op).
Direction 2 — hub→box wait channel (v0.140.0): the reverse immediacy path, so an OPERATOR action on the hub reaches the box in seconds. report.Waiter (internal/report/waiter.go) holds a hanging authenticated GET {hub}/api/v1/wait?gen=N (same hub URL + key as the pusher; no new config keys) against hub ≥ v0.58.0's in-memory operator-intent generation counter. The hub completes the hold the instant any operator intent bumps that customer's generation (config save/delete, claim resend, offsite re-issue/freeze, floor, block/unblock, log-pull); on a generation change the Waiter fires the same Direction-1 report.Trigger — and nothing else, so the immediate report's ACK delivers everything through the unchanged config-refresh/escrow/claim/floor machinery (the box pulls even the wake-up; the hub never connects inbound). Its http.Client has no overall timeout (a held GET must stay open for the hub's ~240 s hold, which streams a 25 s heartbeat newline to defeat the nginx 60 s read-timeout — no ingress change needed); a per-request context bounds a dead connection. First-observation records-not-fires (no restart echo); a same-generation timeout fires nothing; any error (transport, a 404 from a pre-v0.58.0 hub, malformed body) backs off 5 s→5 min and the 15-min cycle keeps reconciling. Constructed beside the trigger under the same hubPusher != nil && cfg.Hub.Enabled gate. Grounding: felhom.eu/documentation/audits/SPIKE-immediate-sync-transport-2026-07-16.md.
Config apply + self-restart (internal/api/router.go, internal/api/selfrestart.go)
POST /api/config/apply (Hub-authed) writes a new controller.yaml, but the new config only takes effect on restart — singletons such as the Cloudflare client are built once at startup (so a rotated CF API token would otherwise keep failing). Behaviour (v0.70.0):
- No-op guard: if the pushed body is byte-identical to the current file, do nothing — no rewrite, no restart (the Hub may re-push idempotently; never flap on a no-op).
- Otherwise: write
0600, respond200(flushed), then gracefully self-restart —gracefulSelfRestartwaits ~500 ms for the response to flush, thenos.Exit(0). The container runsrestart: unless-stopped, so Docker brings it back with the fresh config, and a startup report is pushed. The exit sits behind an injectable seam (Router.restart/SetRestarter) for unit testing.
Manual restart (v0.70.0): POST /api/selfrestart (session auth + CSRF via the /api/ mount) runs the same helper — surfaced as the "Vezérlő újraindítása" button on the settings page (confirm → POST → poll GET / every 2 s → reload), so a customer can recover the controller without rebooting the whole guest.
Full-server restart (v0.81.0): POST /api/server/reboot (session auth + CSRF) reboots the whole guest via the host agent's GuestReboot primitive (HandleServerReboot in internal/web/storage_handlers.go, delegating to the testable serverReboot core; agent reboots detached + returns 202). Surfaced as a separate "Kiszolgáló újraindítása" settings card alongside the controller-only restart, reusing the same pollRestart() reload loop. It replaces the retired drive-activation banner (v0.81.0): in the intermediary-mount model an enrolled drive binds live into the running guest, so storage no longer needs a reboot to activate — this button is purely a deliberate full-system restart.
App Telemetry (internal/metrics/telemetry.go, internal/metrics/logscanner.go, internal/report/telemetry.go)
Each report push now includes per-app telemetry data:
Metrics collection (telemetry.go):
MetricsStore.GetContainerTelemetry(since)aggregates container-level memory (avg, peak, current) and CPU averages from thecontainer_metricsSQLite table for the last 15 minutes.
Log scanning (logscanner.go):
ScanContainerLogs(containerNames, since, logger)runsdocker logs --since=15m --tail=1000sequentially on all non-protected deployed containers.- Classifies lines by keyword match (errors:
error,fatal,panic,crit,oom,killed,exception,traceback; warnings:warn,warning) on the first 5 words (case-insensitive). - Deduplicates via fingerprinting: strips ANSI escape codes, ISO timestamps (with timezone offsets), and syslog timestamps (including mid-line); replaces 6+ digit numbers with
<N>, 8+ char hex with<HEX>, UUIDs with<UUID>. Groups identical fingerprints, keeps top 10 per container. - Returns
[]ContainerLogSummarywithErrorCount,WarnCount,RecentIssues []LogIssue. - Error context (v0.111.0): each error-severity issue carries
Context []string— up to ±5 raw lines around its FIRST occurrence in the scrape window (≤11 lines, ≤400 chars/line with…, ANSI-stripped, redacted). Warns never carry context. The classification loop is the pureanalyzeLogLines()(unit-tested with synthetic windows). - Redaction (
redact.go, v0.111.0):RedactLine()maskspassword|passwd|secret|token|api[_-]?key|authorization|bearervalues ([REDACTED], incl.Authorization: Bearer <tok>) and 64-hex strings ([REDACTED-HEX64]) on every context and log-tail line before it leaves the box — controller-side, authoritative.
Report integration (report/telemetry.go):
buildAppTelemetrySection()calls both, thenbuildAppTelemetry()aggregates by stack — summing container metrics, merging issues, capping at 10 per app. Additionally,buildControllerTelemetry()creates a special entry for the controller container itself (app_name: "felhom-controller").- Results stored as
[]AppTelemetryin theReportstruct fieldapp_telemetry. - Context budget (v0.111.0):
enforceContextBudget()caps the per-report total of all issue-context bytes at 16KB, dropping context from the lowest-count issues first.
On-demand log tails (report/logtail.go, v0.111.0 — pull-based, same ACK-flag pattern as escrow/config-refresh):
- The report ACK (
PushResponse) gainslog_tail_requests: [app…]— apps the operator requested logs for on the hub.main.go'sOnPushResponsehands them toreport.SetPendingLogTails(); the NEXTBuildReportdrains them (consume-once) and shipslog_tails: [{app, collected_at, lines[]}]. - Collection:
stacks.GetLogs(app, 200)(compose logs, ordered as emitted) for stacks;metrics.FetchContainerLogTail()(docker logs --tail=200) for thefelhom-controllercontainer. Caps: ≤400 chars/line, ≤64KB/app head-truncated (newest lines kept), every line throughRedactLine(). - Fail-safe: a failed fetch or failed push leaves the hub's request pending — the next ACK re-arms it. The hub clears the request when a tail arrives. NO hub→controller push channel exists; the guest listens to no one.
Infrastructure Backup to Hub — RETIRED (2026-06-16)
Removed. The controller no longer pushes any infra-backup to the Hub, and the Hub no longer accepts or stores one (hub v0.12.0). The builder (
internal/report/infra_backup.go) and local mirror (internal/backup/local_infra.go) were deleted back in slice 8C; the last caller-less stub (Pusher.PushInfraBackup) and thebackup_completedevent were removed in controller v0.69.0. DR now rests on the agent's PBS whole-CT snapshot + the Hub-generated controller.yaml. The text below is historical and describes the removed mechanism — much of this section (and thelocal_infra.go/setup/scanner.go/PullRecovery/restore_drivesreferences elsewhere in this README) is stale slice-8C debt. Seefelhom.eu/documentation/audits/SPIKE-infra-backup-2026-06-15.md.
After each backup cycle (including manual Tier 2 triggers via OnCrossDriveComplete callback), the controller pushed a full infrastructure snapshot to the Hub for disaster recovery. This snapshot included:
controller.yaml(base64-encoded, full config including secrets)settings.json(base64-encoded, backup prefs, storage paths, cross-drive configs)- Disk layout (UUIDs, labels, mount points, fstab options, bind-mount topology)
- Deployed stacks manifest (app names, HDD paths) with actual config files:
docker-compose.yml,app.yaml,.felhom.yml(base64-encoded per stack, v0.34.0) - Restic passwords (primary + cross-drive, base64-encoded)
This enables fully automated recovery when the system drive is replaced — the new controller pulls the snapshot from the Hub, auto-mounts surviving drives by UUID, and restores all applications.
Hub Dashboard
The hub service (separate Go app in the felhom.eu repo) provides:
- Multi-customer overview table with status indicators and event count badges
- Customer detail page with system/storage/containers/backup/health/events sections
- Event timeline: last 50 events with severity filter, colored badges, source tracking
- Dead man's switch: staleness detection (30min stale, 60min down), missed backup detection (daily at 05:00)
- Notification dispatch: operator (English) + customer (Hungarian) emails via Resend with per-event cooldowns
- Infra backup status per customer (last sync, stack count, disk count)
- Color coding: green (<30min), yellow (30-60min), red (>60min since last report)
- 90-day report + event retention with daily prune at 04:30 Budapest time
10. First-Run Setup Wizard
When the controller starts with no valid customer configuration (customer.id empty), it enters setup mode — a web-based wizard that handles all initial configuration. This replaces the old interactive shell wizard in docker-setup.sh.
Setup Mode Detection (internal/setup/setup.go)
NeedsSetup(cfg) returns true when customer.id is empty or a .needs-setup marker file exists. In setup mode, the controller skips normal startup (no scheduler, no backup, no stacks) and serves only the wizard UI on two listeners:
:8080— behind Traefik (accessible via domain, e.g.https://felhom.example.com):8081— direct HTTP (accessible via LAN IP, e.g.http://192.168.0.100:8081)
Wizard Flow
┌──────────────────────────────────┐
│ 1. Welcome │
│ Choose: Restore / Fresh install │
└─────────┬───────────┬────────────┘
│ │
┌─────▼─────┐ ┌──▼───────────────┐
│ 2a. Scan │ │ 2b. Hub download │
│ drives for│ │ (customer ID + │
│ local │ │ password) │
│ backups │ │ │
└─────┬─────┘ └──────┬────────────┘
│ │
┌─────▼─────┐ │
│ 2a.2 Hub │ │
│ recovery │ │
│ (fallback)│ │
└─────┬─────┘ │
│ │
┌─────▼─────┐ ┌──────▼───────────┐
│ Execute │ │ Execute fresh │
│ restore │ │ install │
└─────┬─────┘ └──────┬───────────┘
│ │
└───────┬───────┘
▼
os.Exit(0) → Docker restarts
→ normal mode
Hub Pre-Seeding
When docker-setup.sh is run with --hub-customer / --hub-password, the controller receives
pre-seeded credentials via environment variables:
| Env var | Purpose |
|---|---|
FELHOM_SETUP_CUSTOMER_ID |
Pre-fills customer ID in wizard forms |
FELHOM_SETUP_PASSWORD |
Pre-fills retrieval password for auto-processing |
In hub mode, the welcome page shows three cards instead of two:
- "Visszaállítás a Hub-ról" — auto-calls
PullRecovery(), shows infra backup details - "Visszaállítás helyi meghajtóról" — standard drive scan
- "Friss telepítés" — auto-calls
PullConfig(), downloads config only
Both hub paths auto-process when credentials are pre-seeded (no form entry needed). On error, the wizard falls back to the manual form with the error displayed.
Key Components
| File | Purpose |
|---|---|
setup/setup.go |
NeedsSetup() detection, SetupState persistence to setup-state.json |
setup/handlers.go |
HTTP handlers for each wizard step (welcome, scan, hub-restore, fresh, manual) |
setup/scanner.go |
Scans all block devices for .felhom-infra-backup/ directories (current + history/) via lsblk + temp mounts; returns rich info (app names, disk count) |
setup/hub.go |
Hub recovery pull (GET /api/v1/recovery/{id}) and config download |
setup/csrf.go |
Lightweight CSRF protection (cookie + hidden field, SameSite=Strict) |
setup/network.go |
Detects local IPs for LAN access URL display |
setup/templates/ |
8 embedded HTML templates (Hungarian, dark theme matching main UI) — includes setup_hub_versions.html for Hub backup version picker |
Local Infra Backup (internal/backup/local_infra.go)
The controller writes infrastructure snapshots to every connected drive after each backup cycle and on startup. Location: <drive>/.felhom-infra-backup/. Files:
backup.json— full infra backup (config, settings, disk layout, passwords, stacks)metadata.json— schema version, timestamp, customer ID, controller version, SHA256 checksumhistory/— previous backup versions (last 5), rotated automatically before each write{timestamp}-backup.json+{timestamp}-metadata.jsonpairs (timestamp format:20060102T150405Z)- Oldest entries pruned when count exceeds 5
During setup wizard drive scan, both current and historical backups are discovered, integrity-verified, and offered for one-click restore. The scan results table shows app names/count, disk count, and a "korábbi" badge for historical versions.
Recovery Info (internal/recovery/info.go)
Generates recovery-info.txt on the system data partition with customer ID, Hub URL, retrieval password, and recovery instructions in Hungarian. Updated on startup and after config changes. Also displayed on the Settings page in a "Vészhelyzeti információk" section.
No secret is rendered into a page's response body (v0.207.0 + v0.208.0, R-249/R-254). Three endpoints implement one rule — the page carries a BOOLEAN, the value comes from an explicit authenticated POST with Cache-Control: no-store, and the reveal is LOGGED (reading a value off markup left no trace at all):
| Secret | Endpoint | Read from |
|---|---|---|
| retrieval passphrase | POST /settings/retrieval-password/reveal |
settings |
| an app's generated first-login password | POST /apps/<slug>/initial-credentials/reveal |
live from the container — never a cached copy |
| an already-deployed app's auto-generated secret field | POST /stacks/<name>/auto-field/reveal |
the decrypted app.yaml; authorised by requiring a type: secret auto-field of that stack |
They are deliberately per-secret, not one generic "reveal any named secret" endpoint — that would turn three narrow exposures into one lever with a parameter. The PRE-DEPLOY hidden input is untouched and deliberate: a form must carry what it submits (see §318 below). scripts/secret_in_markup_gate.py (in controller_gates.py) enforces the rule over all templates; its measured blind spot — a secret arriving under a neutral page-data key — is in its docstring.
The retrieval passphrase is NOT rendered into that page (v0.207.0, R-249). securityPageData passes only HasRetrievalPassword (a boolean), and the value is fetched by an explicit act: POST /settings/retrieval-password/reveal → {"ok":true,"data":{"password":"…"}}, behind RequireAuth + CsrfProtect like every other POST, Cache-Control: no-store, and logged (retrieval passphrase revealed via the security page from <ip> — the value is never logged). Until v0.207.0 the page rendered the plaintext into a display:none span, so any fetch of the page returned it; the toggle was cosmetic. This follows the rule escrow_handlers.go already states for the recovery code: a secret is revealed by an XHR, never templated server-side into HTML.
11. Disaster Recovery
When a system drive fails and is replaced, the recovery flow uses the setup wizard:
1. docker-setup.sh deploys fresh controller with minimal config
- With --hub-customer: credentials pre-seeded via env vars
- Without: user enters credentials manually in wizard
2. Controller detects empty customer.id → enters setup mode
3. User opens wizard at http://<LAN-IP>:8081
4. Hub mode: welcome page shows Hub restore / local scan / fresh install
Non-hub mode: welcome page shows restore / fresh install
5. Hub restore: auto-connects to Hub, shows version picker if multiple versions
Local restore: scans all drives for .felhom-infra-backup/ directories (current + history/)
6. User selects backup version → restore: config, settings, passwords, disk layout
7. Controller restarts into normal mode with full config
8. Controller auto-mounts surviving drives by UUID from disk layout
9. Dashboard shows "Visszaállítás" (Restore) page for app-level recovery
10. User confirms → sequential restore: rsync first, restic fallback, DB import
Backup sources (priority order):
- Local infra backup (
.felhom-infra-backup/on surviving drives) — fastest, no network needed - Hub recovery endpoint (
GET /api/v1/recovery/{id}) — requires retrieval password, supports?version=IDfor specific versions; Hub retains ~14 versions via GFS pruning (7 daily / 4 weekly / 3 monthly) - Manual config (wizard form) — enter all details manually as last resort
Hub verification: After setup, the controller periodically verifies customer standing via the Hub report push response (customer_blocked field). If blocked or Hub unreachable for >7 days, the controller enters limited mode (no new deployments).
12. Asset Sync
App assets (logos, screenshots) are managed centrally by the Hub and downloaded to each controller via a daily sync process. This decouples asset updates from controller image rebuilds — new app icons only require a Hub redeploy.
How It Works (internal/assets/syncer.go)
1. Fetch manifest from Hub: GET /api/v1/assets/manifest (Bearer auth)
2. Compare SHA-256 checksums with local cache (<dataDir>/assets/)
3. Download changed/new files: GET /api/v1/assets/file/{filename}
4. Remove local files not in Hub manifest (stale cleanup)
5. Save local manifest copy for next comparison
Asset Resolution (two-tier)
| Priority | Path | Source |
|---|---|---|
| 1 | <dataDir>/assets/ |
Downloaded from Hub (synced cache) |
| 2 | /usr/share/felhom/assets/ |
Baked into Docker image (fallback) |
The Resolve(filename) method checks the synced cache first, then falls back to the baked-in directory. This ensures assets are always available even before the first sync.
The Felhom logo (/static/felhom-logo.svg) also uses this two-tier resolution: the logo handler checks synced assets first, then falls back to the embedded SVG constant. This allows logo updates via Hub without a controller rebuild. The logo is also used as an SVG favicon.
Configuration
assets:
sync_enabled: true # Opt-in: download assets from Hub API
sync_schedule: "05:00" # Daily sync time (HH:MM, Budapest timezone)
Asset sync requires hub.enabled: true with valid hub.url and hub.api_key. The initial sync runs 10 seconds after startup (to let subsystems initialize), then daily at the configured time.
Sync Status
The syncer tracks status (last sync time, result, file count, total bytes) accessible via GET /api/assets/status. On-demand sync can be triggered via POST /api/assets/sync.
File Types
The Hub serves three asset types per app:
{slug}-logo.svg— primary SVG logo{slug}-logo.png— PNG fallback{slug}-screenshot-{N}.webp— app screenshots
Key Design Decisions
- Opt-in via
sync_enabled— backward compatible, baked-in assets still work without Hub - SHA-256 change detection — only downloads files that actually changed (bandwidth efficient)
- Atomic file writes — downloads to
.tmpthenos.Renamefor crash safety - Stale file cleanup — removes local files not in the Hub manifest (e.g., deleted apps)
- Non-blocking initial sync — runs in a goroutine with 10s delay, doesn't block startup
13. Debug Mode
When logging.level: "debug" is set in controller.yaml, the controller exposes a full diagnostic dashboard at /debug. All debug endpoints are gated — at info level, the sidebar link disappears and all /api/debug/* routes return 404.
R-400 (v0.228.0): the table below is now MECHANICALLY pinned to the dispatcher. controller/scripts/debug_route_gate.py compares every /api/debug/... reference in debug.html against every subpath == case in handler_debug.go and fails on either difference. Before it existed the page referenced 24 addresses and 17 were answered; three of the seven dead ones fetched on page LOAD, so whole panels had been permanently blank. Six controls were deleted and one (backup/crossdrive) implemented — the "Tárhely teszt" section went entirely, which is why the section numbers below skip 4.
Debug Page Sections
| # | Section | Endpoints | Description |
|---|---|---|---|
| 1 | Rendszer diagnosztika | GET /api/debug/dump |
Full state dump: controller info, storage, stacks, network (guest-netns interfaces/route/DNS via the samba door, R-66; best-effort per item), scheduler, health, alerts. JSON download. |
| 2 | Értesítés teszt | POST /api/debug/event/test, GET /api/debug/event/history |
Send test events with configurable type/severity, view event history ring buffer. |
| 3 | Mentés teszt | POST /api/debug/backup/dbdump · POST /api/debug/backup/crossdrive · POST /api/debug/backup/integrity |
Trigger a DB dump, run the Tier-2 (cross-drive) sweep over every deployed HDD-backed app, or run an off-site integrity check by hand. crossdrive is asynchronous and answers with the app list it started for; integrity is synchronous and answers with the verdict. backup/infra was DELETED (R-400) — the disk-tier infra backup moved to the host agent in slice 8C and nothing in this repo backs it. |
| 5 | Hub & Kapcsolatok | POST /api/debug/hub/{push,test-connectivity,preferences-sync}, POST /api/debug/gitea/test-connectivity |
Test Hub/Gitea connectivity with latency. Push reports and sync preferences. hub/infra-push was DELETED (R-400) — Pusher.PushInfraBackup was removed 2026-06-16. |
| — | Telemetria teszt | GET /api/debug/telemetry |
Run the full telemetry collection pipeline on-demand (metrics query + log scan). Returns per-app table: container list, memory current/avg/peak, CPU avg, catalog limit, log error/warning counts, and top issues. Useful for verifying container→stack mapping and testing log scanner patterns without waiting for the 15-minute report cycle. |
| 6 | Önfrissítés teszt | POST /api/debug/selfupdate/dry-run |
Dry-run update check: current vs new image lines, compose writability, backup state. |
| 7 | DR / Telepítő varázsló | POST /api/debug/dr/trigger-setup |
Trigger setup mode via marker file (requires "RESET"). dr/infra-status and its panel were DELETED (R-400) — it rendered the two retired infra-backup mechanisms above, and it fetched on page LOAD, so the panel had been permanently blank. |
| 8 | Naplóviewer | GET /api/debug/logs?level=&limit=&after=, GET /api/debug/agent-logs |
In-memory log viewer (last 5000 entries, spill-persisted across restart — fix-6), level filter, 2s auto-refresh, color-coded entries. Two tabs (v0.116.0): Vezérlő (own ring) and Ügynök (the agent's always-DEBUG ring proxied over the local API; a pre-0.83 agent renders the "available after the agent's next update" notice). |
Key Implementation Details
- Log buffer (
internal/web/logbuffer.go): Ring buffer implementingio.Writer. Since v0.116.0 it ALWAYS exists (anylogging.level) and captures every line INCLUDING[DEBUG]: the logger isio.MultiWriter(LevelFilterWriter(os.Stdout, logging.level), logBuffer)— stdout/docker-logs keep respectinglogging.level, the ring holds the full detail for remote diagnostics.logBuffer.Lines(maxBytes)renders the newest-kept plain-text tail (the reportcontroller_log_tailsource). New leveled lines useinternal/logx(Debugf/Infof/Warnf/Errorf); legacyisDebug()call sites are unchanged.- Ring sizing, spill persistence, periodic-noise policy (fix-6, v0.120.0, CAMPAIGN-3). The
campaign measured the 1000-entry ring wrapping in ~6.5 min under load and dying on every restart —
the exact post-incident window was the first thing lost. Three changes: (a) cap 1000→5000
(
Entries/the debug handler display cap raised to match — a larger ring is useless if the viewer can't request more than 1000 of it; the Naplóviewer default pull is 1000). (b) periodic-noise policy: a periodic job's ROUTINE success is not ring-worthy — the every-cycle scheduler "job finished" line andrefreshStatusLockedper-cycle enumeration are logged at a new[TRACE]level that the ring DROPS at write-time (levelPriority("TRACE") < DEBUG). Failures and state changes are never TRACE, so nothing load-bearing is lost; this was the biggest ring filler. (c) spill persistence:LogBuffer.SpillTo/LoadFromatomically (tmp+rename, JSON-lines) spill the ring to<DataDir>/debug-ring.logon the SSD state dir (the persistent data volume that survives container recreation — NEVER a NAS/HDD path) every 30 s and on clean shutdown, and load it back on boot so a restart / recreate preserves the pre-restart window. Corruption-safe: a truncated/partial line is skipped on load, never fatal.
- Ring sizing, spill persistence, periodic-noise policy (fix-6, v0.120.0, CAMPAIGN-3). The
campaign measured the 1000-entry ring wrapping in ~6.5 min under load and dying on every restart —
the exact post-incident window was the first thing lost. Three changes: (a) cap 1000→5000
(
- Controller self-log pull (
internal/report/selftail.go): the hub's report ACK may carrycontroller_log_requested— the NEXT report shipscontroller_log_tail(ring, 128 KB cap, consume-once, the v0.111.0 app-tail pattern; additive fields, app-tail wire unchanged). Serving a pull logs the customer-visibleoperator log pull servedINFO line. - Storage simulation:
simulatedPathsmap in watchdog prevents the watchdog from re-probing simulated-disconnected paths. Disconnect runs all real steps exceptlazyUnmount(drive stays physically mounted). - DR trigger safety: Uses marker file (
data/.needs-setup) instead of modifying controller.yaml. Pre-checks that infra backup exists on at least one drive. - Routing:
/api/debug/carved out in HTTP mux (same pattern as/api/storage/), routed to web server with auth + CSRF. - DebugCallbacks: 7 closures wired from main.go for operations needing modules not on Server struct (hub push, infra backup, connectivity tests, telemetry preview).
- Telemetry debug:
GetTelemetryPreviewcallback callsreport.BuildAppTelemetryForDebug()(exported wrapper around the privatebuildAppTelemetrySection()). Result renders as a table with collapsible raw JSON. Available regardless of hub configuration.
Per-Module Logging
All modules emit structured log lines at [INFO], [WARN], and [ERROR] levels for operational events (state changes, completions, failures). When logging.level: "debug", additional detailed [DEBUG] [module] prefixed log lines are emitted. Each module with stateful debug (struct-based) exposes a SetDebug(bool) method, wired from main.go. Modules without a struct use package-level DebugLogger variables (e.g., system.DebugLogger).
Standard-level logging (always active):
[INFO]— Operational events: stack deploy/start/stop, backup completion, config changes, disk operations, sync results[WARN]— Degraded states: health threshold breaches, unsafe backup destinations, retryable failures, best-effort operation failures[ERROR]— Hard failures: data restore errors, integration apply failures, compose file update errors, disk format failures
| Module | Debug Field | Prefix | Key Areas |
|---|---|---|---|
stacks |
cfg.Logging.Level |
[DEBUG] [stacks] |
Stack CRUD, compose commands, env vars, HDD mounts, encryption migration, health probes |
backup |
ResticManager.debug |
[DEBUG] [restic] / [DEBUG] [backup] |
Restic commands, snapshot operations, restore scanning, drive mounting |
cloudflare |
Client.debug + GeoSyncManager.debug |
[CF-DEBUG] / [DEBUG] [cloudflare] |
API requests/responses, WAF rule CRUD, zone resolution, geo sync diff |
integrations |
Manager.debug |
[DEBUG] [integrations] |
Toggle apply/revoke timing, lifecycle hooks, config reapply |
system |
DebugLogger |
[DEBUG] [system] |
Memory/disk/CPU/load/temp collection, mount probing, USB detection |
monitor |
Pinger.debug |
[DEBUG] [pinger] |
Health ping URLs, retry attempts, response codes |
settings |
Settings.debug |
[DEBUG] [settings] |
Load/save sizes, storage path ops, geo/integration state changes |
scheduler |
Scheduler.debug |
[DEBUG] [sched] |
Job registration, execution timing, daily schedule calculations |
web |
cfg.Logging.Level |
[DEBUG] [web] |
HTTP requests, auth decisions, session management, storage API ops |
api |
Router.debug |
[DEBUG] [api] |
API routing, handler entry points, request details |
selfupdate |
Updater.debug |
[DEBUG] [selfupdate] |
Version checks, update preconditions, docker pull timing |
assets |
Syncer.debug |
[DEBUG] [assets] |
Manifest fetch, hash comparison, file download timing |
storage |
logger-based | [DEBUG] [storage] |
Disk scanning, formatting, attach, drive migration |
metrics |
logger-based | [DEBUG] [metrics] |
Per-container log scanning, error/warning counts |
appexport |
Exporter.debug |
[DEBUG] [appexport] |
Export/import steps, crypto operations, bundle scanning |
14. Geo-Restriction
Country-based access control via Cloudflare WAF Custom Rules. The controller manages WAF rules in the http_request_firewall_custom phase to block requests from non-allowed countries. Rules are identified by a [felhom-geo] description prefix — other WAF rules are never touched.
Prerequisites
The existing cf_api_token (used for DNS-01 ACME) needs Zone WAF:Edit permission added. No new token is needed — just expanded permissions on the same token. The settings UI only appears when a CF API token is configured.
Architecture
┌─────────────┐ ┌──────────────────┐ ┌──────────────────────┐
│ Settings UI │────▶│ GeoSyncManager │────▶│ Cloudflare WAF API │
│ (settings. │ │ (geosync.go) │ │ /zones/{id}/ │
│ html) │ │ diff & apply │ │ rulesets/{id}/rules │
└─────────────┘ └──────────────────┘ └──────────────────────┘
│ ▲
│ POST /api/geo/* │ Scheduler (6h)
▼ │ + deploy/remove hooks
┌─────────────┐ │
│ API layer │──────────────┘
│ (geo.go) │
└─────────────┘
Rule structure:
- Global rule:
(not ip.src.country in {"HU"})→ block (withhttp.host neexclusions for apps that have per-app overrides) - Per-app rule:
(http.host eq "app.example.com" and not ip.src.country in {"HU" "US"})→ block - Block response: HTTP 403 with Hungarian message
Local network access is inherently unaffected — traffic from the LAN goes directly to the server, bypassing Cloudflare entirely.
Cloudflare API Client (internal/cloudflare/)
| File | Purpose |
|---|---|
client.go |
HTTP client with Bearer token auth, 15s timeout, generic do() helper |
zone.go |
Zone ID resolution — tries exact domain, then parent domains progressively |
waf.go |
WAF rule CRUD, expression builders (BuildGlobalExpression, BuildAppExpression) |
countries.go |
~250 ISO 3166-1 alpha-2 codes with Hungarian names |
geosync.go |
Sync orchestrator — diffs desired vs existing rules, creates/updates/deletes |
GeoSyncManager uses a StackLister interface (implemented by geoStackAdapter in main.go) to get deployed app hostnames without circular imports.
Settings Model
Stored in settings.json (runtime-modifiable):
type GeoRestriction struct {
Enabled bool `json:"enabled"`
AllowedCountries []string `json:"allowed_countries"`
AppOverrides map[string]AppGeoOverride `json:"app_overrides,omitempty"`
LastSync string `json:"last_sync,omitempty"`
LastSyncError string `json:"last_sync_error,omitempty"`
ZoneID string `json:"zone_id,omitempty"`
RulesetID string `json:"ruleset_id,omitempty"`
}
Thread-safe access via GetGeoRestriction(), SetGeoRestriction(), SetGeoAppOverride(), RemoveGeoAppOverride(), SetGeoSyncState().
API Endpoints
| Method | Path | Description |
|---|---|---|
| GET | /api/geo/status |
Current geo settings + sync state |
| POST | /api/geo/settings |
Update global settings (enable/disable, countries) |
| POST | /api/geo/sync |
Trigger manual sync |
| GET | /api/geo/countries |
Full country list for search UI |
| POST | /api/stacks/{name}/geo/override |
Set per-app country override |
| DELETE | /api/stacks/{name}/geo/override |
Remove per-app override |
All mutating endpoints trigger an async Cloudflare sync. The /api/geo/ path accepts both session auth and Hub Bearer token auth (via selfUpdateAuthMiddleware), enabling Hub-side geo-disable for lockout recovery.
Sync Triggers
- Settings change — user saves geo settings or per-app override
- Deploy/remove — app deployment or removal changes the hostname list
- Scheduler — periodic verification every 6 hours
- Startup — delayed initial sync 15s after boot
- Manual — "Szinkronizálás" button on settings page
UI
Settings page ("Beállítások" → "Földrajzi korlátozás"):
- Enable/disable toggle
- Searchable country autocomplete with tag-based selection
- Hungary pinned with
confirm()warning on removal - Per-app overrides summary with add/edit/remove
- Sync status display (last sync time, errors)
App detail page (per-app override, shown when geo is globally enabled):
- Toggle for custom country restriction
- Independent country selector
15. App-to-App Integrations
Generic framework for connecting deployed applications to each other. Provider apps declare available integrations in .felhom.yml, and users enable/disable them via toggle switches on the provider's deploy/settings page ("Beállítások").
Architecture (internal/integrations/)
integrations.go— Core types:Handlerinterface (Apply/Revoke),ApplyContext(carries domain, decrypted env vars, provider metadata, stacks dir, logger, restart func),StatusInfo(UI data),IntegrationKey()/ParseIntegrationKey()key helpersmanager.go—Managercoordinates toggle operations, builds apply contexts from decrypted app.yaml env vars. UsesStackProviderinterface (GetStack, GetStacks, RestartStack) to break circular imports with stacks package — adapted viaintegrationStackAdapterin main.go. Key methods:Toggle(ctx, provider, target, enable)— Validates both apps deployed+running, calls Apply/Revoke, persists stateListForProvider(slug)— Returns[]StatusInfofor UI with target deployment/running statusReapplyConfigForTarget(name)— Re-applies all active integrations targeting a stack (config-only, no restart). Used bySyncFileBrowserMountsafter config regeneration
lifecycle.go— Lifecycle hooks called from API router goroutines:OnStackStop— Revokes active integrations, sets"provider_stopped"/"target_unavailable"(keepsenabled=true)OnStackStart— Re-applies enabled integrations after 5s delay (waits for stack state refresh). Accepts bothStateRunningandStateStartingviaisStackUp()helperOnStackRemove— Revokes and permanently deletes integration state
- Handler implementations — One file per integration pair (e.g.
onlyoffice_filebrowser.go,onlyoffice_nextcloud.go)
Integration State
Stored in settings.json under integrations map (key: "provider:target"):
enabled— User intent (survives stop/restart)status— Current state:"active","error","disabled","provider_stopped","target_unavailable"last_error— Most recent error messageenabled_at— RFC3339 timestamp
CRUD methods in settings.go: GetIntegrationState, SetIntegrationState, RemoveIntegrationState, GetIntegrationsForProvider, GetIntegrationsForTarget (all use existing RWMutex + atomic write pattern).
Lifecycle
- Enable: User toggles on → validates both apps deployed+running → calls
Handler.Apply()→ persists state as"active" - Disable: User toggles off → calls
Handler.Revoke()→ persists state as"disabled" - Provider/target stops:
OnStackStop→ callsHandler.Revoke()→ sets status to"provider_stopped"or"target_unavailable"(keepsenabled=true) - Provider/target starts:
OnStackStart(5s delay) → finds enabled integrations with non-active status → re-applies if both sides running/starting - Provider/target removed:
OnStackRemove→ revokes and deletes integration state permanently - FileBrowser config regen:
SyncFileBrowserMountsregeneratesconfig.yamlfrom scratch →ReapplyConfigForTarget("filebrowser")patches integration config synchronously → recreates the container only when the finalconfig.yaml/compose differ from the pre-sync content (fbNeedsRecreategate, v0.82.0)
Important: SyncFileBrowserMounts uses --force-recreate (rather than a plain up -d) when something changed, because config.yaml is a bind mount — without --force-recreate, docker compose up -d won't recreate the container when only the config file changes (compose only detects compose-file changes). The recreate is now gated on an actual change (v0.82.0, F2): a controller restart or no-op sync where the generated config+compose are byte-identical issues a plain up -d --remove-orphans and does not bounce the customer's file UI. ReapplyConfigForTarget calls each handler's Apply with a no-op RestartStack since the caller handles the restart.
Built-in Handlers
OnlyOffice → FileBrowser (onlyoffice_filebrowser.go):
- Apply: Reads
JWT_SECRET+SUBDOMAINfrom OnlyOffice app.yaml (decrypted), strips any existingintegrations:block from FileBrowserconfig.yamlviaremoveIntegrationsSection(), appends new block withurl(public HTTPS),internalUrl(http://onlyoffice:80),secret,viewOnly: false. Atomic write (.tmp+ rename). Restarts FileBrowser - Revoke: Strips
integrations:block from config.yaml, restarts FileBrowser
OnlyOffice → Nextcloud (onlyoffice_nextcloud.go):
- Apply: Runs
docker exec -u www-data nextcloud php occcommands:app:install onlyoffice(tolerates "already installed")app:enable onlyofficeconfig:app:set onlyoffice DocumentServerUrl --value=https://{subdomain}.{domain}config:app:set onlyoffice DocumentServerInternalUrl --value=http://onlyoffice:80config:app:set onlyoffice jwt_secret --value={JWT_SECRET}config:app:set onlyoffice StorageUrl --value=http://nextcloud(internal callback URL)
- Revoke: Runs
occ app:disable onlyoffice(tolerates container not running / app not enabled)
OnlyOffice compose template notes: Requires Traefik middleware X-Forwarded-Proto=https in labels so the Document Server generates HTTPS URLs for editor resources (prevents mixed content errors in browser).
Metadata (.felhom.yml)
Provider apps declare integrations in their .felhom.yml. Parsed into IntegrationDef struct in metadata.go, with HasIntegrations() helper.
integrations:
- target: filebrowser
label: "FileBrowser integráció"
description: "Dokumentumok szerkesztése a fájlkezelőben"
- target: nextcloud
label: "Nextcloud integráció"
description: "Dokumentumok szerkesztése a Nextcloudban"
API Endpoints
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/integrations/{provider} |
List integrations for a provider app (status, target availability) |
| POST | /api/integrations/{provider}/{target} |
Enable/disable integration ({"enabled": true/false}) |
Routes registered before hasSuffix-based stack routes in router.go (see router bug pattern).
UI
Toggle switches on the provider's deploy/settings page ("Integrációk" section, within deploy.html). Data wired in deployHandler() for deployed apps only. Each integration shows:
- Label and description from
.felhom.ymlmetadata - Status badge: "Aktív", "Nincs telepítve", "Célalkalmazás leállítva", "Hiba"
- Toggle checkbox (disabled when target not deployed/running)
- JS
toggleIntegration()→ POST to API → reload on success
Wiring (main.go)
integrationStackAdaptertype implementsintegrations.StackProvider(same pattern asstackAdapter,geoStackAdapter)integrations.NewManager(sett, adapter, domain, stacksDir, encKey, logger)— registers built-in handlers- Wired into API router via
SetIntegrationManager()and web server viaSetIntegrationManager()
16. Network File Sharing — SMB („Megosztás") (v0.144.0, R-7 slice 1; connect card + status contract v0.151.0)
LAN file sharing so the box behaves like a NAS: the customer enables sharing, sets ONE household SMB
password, and exports folders that appear in Windows Explorer / Mac Finder as \\FELHOM\<share>.
Architecturally this is an EMBEDDED CONTROLLER FEATURE, not a catalog app — three reasons:
it requires network_mode: host (the R-6 spike proved the default docker bridge is deaf to the LAN
multicast that WSD/mDNS discovery needs), its configuration is a dynamic share list rendered into
smb.conf (not env vars), and its share roots must ride the backup classification. It is therefore
the fourth protected infra stack: traefik / cloudflared / filebrowser / samba.
The image (controller/infra-images/samba/, felhom-samba:1.1.0)
Our own pinned image (alpine 3.21 by digest) running four daemons (v1.1.0) — smbd (445), nmbd
(NetBIOS flat-name resolution), wsdd (WS-Discovery), under tini. nmbd is not optional: the R-6
spike proved wsdd alone makes the box visible in Explorer while the double-click still fails
0x80070035, because WSD supplies an icon, not a name→IP mapping. The image is deliberately dumb —
/etc/samba/smb.conf is bind-mounted READ-ONLY by the controller, nothing is templated inside, no
name or password is baked, and the passdb lives on a named volume so the household password survives
container recreation. Built by controller/scripts/build-samba-image.sh (never :latest).
Data model (internal/settings/smb.go)
SMBSettings{Enabled, ServerName, UserSet}—ServerNameis the NetBIOS name (≤15, NetBIOS-safe).SMBShare{Name, Path, ReadOnly, Offsite, CreatedAt}— the share registry.- The SMB password is NEVER persisted. Only
UserSet(a boolean) is stored; the secret lives in the container's passdb, applied viasmbpasswdon STDIN.
Rendering + lifecycle (internal/infra/samba.go, internal/stacks/samba.go)
Pure renderers produce a hardened smb.conf (server min protocol = SMB2, bind interfaces only
on lo eth0, disable netbios = no, map to guest = never, per-share force user/group = felhom
so every SMB write lands as uid:gid 1000) and a compose file (network_mode: host, pinned image,
config :ro, passdb volume, one bind per share — :ro for read-only shares as defence in depth).
ensureSamba joins EnsureBaseStack after filebrowser, gated on SMB.Enabled (the cloudflared
conditional-deploy precedent); ReconcileSamba runs after every mutation. Both are idempotent —
unchanged config plus a running container performs zero compose calls. Config writes are atomic
(tmp+fsync+rename). A share whose drive is disconnected/decommissioned is rendered ABSENT from
smb.conf (never export a dead mountpoint) while its configuration is retained.
Nothing in this feature deletes or moves customer files. Disabling sharing is compose down
(passdb volume kept); deleting a share is a config-only edit. The only os.Mkdir* is the guarded
new-share-folder create.
UI + the picker guard (internal/web/sharing_handlers.go, templates/sharing.html)
Top-nav category „Megosztás" → „Hálózati megosztás": enable/server-name card, household
password, shares table (Név · Mappa · Írásvédett · Felhőmentés · Törlés), and a create flow — either a
NEW folder under <storage>/shares/ or an EXISTING folder chosen in a browse modal.
sharingResolvePath is the security gate for every customer-supplied path: absolute →
EvalSymlinks (before containment, so a planted symlink cannot escape) → must live inside a
registered, live storage root → must not be in a deny-listed system subtree → must be a directory.
Refusals are uniform (Ez a mappa nem osztható meg.) so the picker can never act as a filesystem
oracle. The deny-list is DERIVED from stacks.SharingDeniedRoots, itself provably a subset of
ProtectedHDDPaths — it can only shrink relative to the delete guard, never drift into a stale second
list. The drive root is an exact-match denial (a whole drive is never shareable) while user-data
folders under it stay shareable. sharingResolveStorageRoot is a separate, strictly tighter check
used only as the new-folder parent. The picker endpoint is /api/sharing/browse, registered on the
main mux behind RequireAuth+CsrfProtect (the /api/ subtree is routed there, not in the web
ServeHTTP switch).
Discovery is per-platform, and the two halves do not overlap (v1.1.0)
| Client | Working form | Served by |
|---|---|---|
| Windows | \\<NÉV> |
nmbd (flat-name resolution) + wsdd (Network view) |
| macOS | smb://<NÉV>.local |
avahi/mDNS |
| Any | smb://<IP> |
direct — always works |
A Mac cannot use the bare smb://<NÉV>, and nothing we ship can change that. Captured live
2026-07-20: macOS broadcasts a correct NBNS query for <NÉV><20>, nmbd answers in 140 µs with a
textbook positive response (flags 0x8580, RCODE=0, right address), and macOS never opens a TCP
connection — NetBIOS there feeds legacy browsing, not smb:// URL resolution. avahi templates its
config and _smb._tcp service file from FELHOM_SERVER_NAME at entrypoint, so a rename
re-advertises; both it and dbus are non-fatal on failure, because a discovery gap must never become
a sharing outage. Automatic Finder-sidebar appearance is NOT claimed — the record is published
and answers browse queries, but was not observed working on the test Mac.
„Csatlakozás a megosztáshoz" card (v0.151.0). Shown only while sharing is enabled: the Windows
form (\\<NÉV>), the Mac form (smb://<NÉV>), and — when derivable — the direct address
smb://<IP> as the fallback for networks that do not resolve the name. The address comes from
stacks.SambaLANAddress(), which reads the guest's netns through the samba container
(network_mode: host); the controller itself is on a docker bridge and would answer 172.x. It is
derived per render and stored nowhere — the guest holds the address by DHCP, so a persisted copy
eventually misdirects customers — and an underivable address simply omits the line, because a page
without an address beats a page with a wrong one.
/sharing/status — two channels, one envelope (v0.147.0 card, v0.151.0 contract)
The bring-up poll target reports phase (the ensure JOB, which the page answers with a one-shot
location.reload() when it turns terminal-running) and running (the service LEVEL, straight from
the liveness probe). Keeping them apart is load-bearing: v0.147.0 coerced idle→running on the
PHASE channel, so every steady-state page load saw a fresh success edge and reloaded, forever. Since
v0.151.0 the level never reaches the phase channel, and a terminal running is served exactly
once (consumeIfRunning) so a real bring-up cannot re-arm the reload on the page it just caused.
failed, needs_password and in-flight phases are never consumed — their client path shows a card
and stops, with no reload.
Backup classification (internal/stacks/samba_classify.go)
ClassifiedBinds("samba") resolves from the shares registry rather than catalog metadata (samba has
no .felhom.yml and its binds are absolute share paths). Per-share Felhőmentés ON → mandatory
(offsite + tier-2); OFF → optional (tier-2 only). smb.conf/passdb are never classified.
Share backup EXECUTION — the sibling shares source (R-7b, v0.145.0)
The earlier KNOWN GAP is closed: share data is in both live tiers. It did not get there through
GetStackClassifiedBinds — the engines are recovery-unit shaped and Model B′ deliberately left every
per-app path byte-identical. Instead internal/backup runs a sibling shares source off the same
registry, applying the same per-share class rule:
| Tier | Entry point | Shape |
|---|---|---|
| 2 (cross-drive) | RunSharesTier2 — after the per-stack loop in RunAllTier2 |
legs grouped by SOURCE DRIVE → backups/secondary/_shares/<driveKey>/<share> + _payload/, layout marker LAST |
| 3 (offsite) | runOffboxSharesLeg — after the per-app loop, before retention |
ONE restic backup --tag felhom-offbox --tag _shares = manifest staging dir + every MANDATORY share |
| restore | RestoreSharesScratch → PlaceSharesRestore |
scratch first, then a missing-only merge, each destination prefix-asserted against LIVE storage roots |
The payload (shares_payload.go) is what makes a restore give back a working feature rather than
loose files: a byte-deterministic _shares-manifest.json of the definitions plus a best-effort,
secret-bearing passdb.tar. Definitions protection is the floor — a quota-blocked offsite push
degrades to the manifest alone, never to nothing.
_shares is a reserved key (restic tag, dest root, status record). ValidateSMBShareName refuses
a leading underscore, and both run loops skip a _shares stack loudly. It never reaches a customer
surface: backup.DisplayStackName maps it to „Megosztások" at the notification and prose boundaries,
while the persisted set, the tag and the paths keep the raw key.
Liveness: monitor.EffectiveProtected adds infra.SambaContainerName exactly while sharing is
on, so a dead sharing service raises the standard protected-container issue → alert → degradation
e-mail. Note the container name is NOT the stack name (samba vs felhom-samba).
17. Async-job feedback (v0.147.x, feedback slice 1)
Three long operations that used to be silent now report what they are doing. These are three targeted cards on the two existing patterns (the deploy 3-step panel and the storage-init status poll), NOT a framework — a unified async-job layer is ROADMAP R-45.
Verification-restore visibility (backup/offbox_verify_copies.go, web/offbox_handlers.go,
templates/backups_restore.html). The offsite verification-restore flash now names the full path
it wrote to, and /backups/restore lists existing verification copies (app · size · date · path)
with a per-copy delete.
offsiteRestoreRootFor(drivePath)is THE placebackups/offsite-restoreis spelled;offboxRestoreScratchDirbuilds on it so listing and delete resolve byte-identical paths to what the restore wrote.ListOffsiteRestoreCopies()sweeps every candidate drive in the same preference order the restore path uses to choose one, so it can never miss a copy the restore was capable of creating.DeleteOffsiteRestoreCopy(stack)takes a stack name, never a path — the customer cannot hand the controller a directory to remove. Guarded byisSafeStackNameplus a containment assertion on the resolved path.POST /backup/offbox/verify-copy/deleteadditionally requiresconfirm=1, is double-confirmed in the UI, and refuses while any backup/restore op is running.
SMB bring-up progress (web/samba_ensure_job.go, templates/sharing.html). /sharing/enable
and /sharing/password no longer run ReconcileSamba() inside the POST — they start a detached
single-flight job (the storage_init_job.go shape) and the page polls GET /sharing/status.
- The opening phase is decided before the work starts, from
stacks.SambaImagePresent():pulling(„képfájl letöltése") when the pinned image is not in local Docker storage, elsestarting. Afterwards the image is always present, so the distinction is unrecoverable later. - Terminal success is probed via
stacks.SambaRunning()—compose up -dexits 0 on a crash-loop. Anilreconcile withUserSet == falsereportsneeds_password, notrunning. /sharing/statuslets live container state win over a stale/absent job, so a page loaded after a restart still tells the truth.
Offsite backup progress (backup/offbox_progress.go, templates/backups_remote.html).
RunOffboxBackupWithProgress (the manual trigger only; the nightly RunOffboxBackup is
unchanged and stays silent) installs a progress sink, adds --json to the app-backup leg and scans
restic's stdout line-by-line through offboxStreamRunner — a streaming sibling of the existing
offboxRunner seam, injectable for tests. Published on the existing GET /backup/offbox/status
under progress.
- Restic's status object is only partly usable in practice, and the fallbacks matter more than the
percentage: on an incremental run restic transfers no bytes (
bytes_doneisomitempty, so absent) andpercent_donestays 0 for the whole run; and because restic 0.14 counts a file only when it completes, an app dominated by one large archive freezes the file counters too. The card therefore degrades: bytes -> files -> current file + elapsed seconds. - A run is not only the per-app loop.
Phase(shares,retention) names the post-app stages and clears app-scoped counters, so the card never shows the last app's finished numbers against work that is no longer about that app.
Infra image pins (internal/infra). infra.Images() returns every controller-managed infra
image, derived from the existing pinned consts. felhom-controller --print-infra-images prints it
(config-free by design — no controller.yaml, data dir or settings are touched) so the golden bake
(felhom-agent configs/build-golden.sh) can ask the controller image it is about to bake instead of
keeping its own list, which had already drifted. A go/ast test fails if a *Image const is added
without reaching Images().
18. Dashboard language (i18n) (v0.247.0–v0.259.0)
Design: felhom.eu/documentation/architecture/10-localisation.md. Inventory:
felhom.eu/documentation/audits/I18N-INVENTORY-2026-09-17.md.
What a household sees: Hungarian by default, and since v0.250.0 a Magyar / English switch in the sidebar footer of every dashboard page. Every template is converted (v0.247.0 three pages, v0.248.0 apps and settings, v0.249.0 backups, v0.250.0 storage, sharing, sign-in, claim, guest share, catch-all, debug). Since v0.252.0 the sentences the program BUILDS follow the language too — flash lines, page data, the JSON the page's script reads, the alert banners, the country names, and the four page titles built around an app name. Since v0.253.0 every error message carries its key too. Since v0.254.0 the notes a background run SAVES follow the language as well, and the switch is a globe — on the dashboard and on the sign-in, claim and recovery pages. Since v0.257.0 the app catalog's own copy follows it, and since v0.258.0/v0.259.0 so do the last Go-composed sentences: the update and lifecycle badges, the data-folder promise, the channel banners, the claim page's nine answers (R-596) and the Backup page's protection warnings and tier names (R-598). Still Hungarian: the one R-570 producer, the 18 formal „ön" forms (counted and ratcheted, R-516), everything the operator reads (by design), and the apps' own UIs (not ours).
THE ONE FAILURE MODE THIS SECTION EXISTS TO WARN ABOUT: a composed sentence handed to a renderer as page DATA. Five instances — R-566, R-573, R-590, R-596, R-598. The template renders
{{.Error}}correctly in both languages; only the VALUE is Hungarian. No structural check sees it: not the parity fixtures (they render the field faithfully whatever it holds), notTestI18nEnglishPages(it reads templates, not structs), noti18n_missing_gate(there is no key to be missing), noti18n_go_parity(it proves the Hungarian is unchanged, not which language reached the page). Only a live English page or a handler-level render test finds one — and the render test must assert the Hungarian sentence is GONE, not merely that the English is present.Probing it live: the
felhom_langcookie works ONLY on anonymous pages (sign-in, claim, recovery, bind).langForstep 2 means a request carrying a session reads the household's saved setting and deliberately ignores the cookie — so/backupswithfelhom_lang=enreturns Hungarian, which reads exactly like an unfixed defect. Use?lang=behind auth (R-602).And an English value containing an apostrophe never matches on the page —
html/templateescapes'to', and the assertion fails saying the sentence never arrived (R-603).
- Bundles:
internal/i18n/locales/hu.json(authoritative, every key) anden.json, embedded. Flatkey → text; a value may carry template actions ({{.RecoveryAbandonDate}}) and inline markup (<strong>), which are the message's parameters. English plurals arekey.one/key.other(Bundle.Plural); Hungarian has one form. - Templates: a converted template carries
{{T "key"}}markers.loadTemplatesparses one template set per language, expanding the markers textually beforehtml/templateparses the file — so the Hungarian set is parsed from byte-identical source in identical escaping contexts. An undefined key is left in place and the set fails to load ("function T not defined"). - Fallback: a key missing from English shows the Hungarian text and is counted (startup INFO line
i18n: en template set shows Hungarian for N markers); never a key, never a blank. - Language of a request:
?lang=hu|en(testing override, never persisted) →settings.jsonlanguage→hu.executeTemplaterenders from that language's set. Pages outside the dashboard chrome (login, claim, recovery, both guest share pages, the catch-all) render throughexecuteTemplateLang: the household language, but NO session CSRF fields and NO escrow reminder (TestI18nDirectRenderPagesHaveNoAdminChrome,TestDirectRenderHandlersFollowLanguage). - The switch:
POST /settings/language(lang,back; session CSRF like every form) saves the language and redirects tobackwithout its query. Shown in the sidebar footer of every dashboard page, in both languages, since v0.250.0 (until then only on non-Hungarian pages or with?lang=, while English was partial). - Go-side copy: a handler names its title key (
data["TitleKey"], pinned equal to its Hungarian literal byTestHandlerTitleKeysMatchHungarianTitle); non-Hungarian sets override the copy-producing template funcs (stateLabel,timeAgo,timeAgoStr,nextRunLabel,statusText,infraMeta) from the bundle. The Hungarian funcs are untouched. - Go-side messages (v0.252.0):
Bundle.Msgffills a message's own printf verbs; English reorders with Go's explicit argument indexes (%[2]s), so the Hungarian value stays the format string the code always had. Handlers uses.msg(r, key, …); background and view-model builders take alangand uses.msgLang; the API has the same pair (internal/api/i18n_api.go).s.bundle()falls back to the embeddedi18n.Shared(), so a Server built withoutloadTemplatesstill renders sentences. - Flash lines (v0.252.0): a flash travels in the redirect URL and is rendered by the NEXT request,
so it carries a bundle KEY plus repeated
faparameters (?flash=flash.share.enabled&fa=…), built withflashQueryand resolved withs.flashFrom. A value the bundle does not know — a link minted by an older controller — is shown verbatim, never as a key and never dropped. - Alert banners (v0.252.0): an alert is built by a background health cycle and read minutes later,
so
AlertcarriesMessageKey+MessageArgs(andLinkTextKey) andGetAlerts(lang)renders on the way out. An alert with no key rendersMessageverbatim — which is how the two exceptions work. - Country names (v0.252.0):
country.<ISO2>keys, translated at DISPLAY ininternal/api/geo.goand re-sorted per language. Theinternal/cloudflaretable is untouched: it validates codes, and the report puts codes (never names) on the wire. - NOT translated — it is not ours: the report's
health.warnings/health.issuesand everynotifyeventmessage. The hub composes the household's e-mail from those and falls back to the controller's own sentence when it has no entry (FormatCustomerEmail), so translating them here would change an e-mail nobody asked to change. They follow the language in slice 3 (R-558). Pinned by wire goldens ininternal/monitorandinternal/notify. - Errors (v0.253.0):
util.MsgError(key, args…)/util.MsgErrorf(kind, key, args…)carry the bundle key from the package that MAKES the error to the handler that PRINTS it.Error()returns the Hungarian byte for byte (so an un-converted printer is unchanged);Unwrap() []errorreturns the kind and any wrapped cause, soerrors.Isanswers for both; an error ARGUMENT renders recursively, and a foreign error (restic, docker, ssh, stdlib) prints verbatim. Display ends calls.errText(r, err)/r.errText(req, err)—TestNoErrErrorInPageOutputconvicts any sink that does not. - Plurals (v0.253.0): a key that carries
.one/.otherin a language is a plural key and its FIRST parameter is the count (i18n.Bundle.form). Hungarian has one form at every count..one/.otherare RESERVED suffixes —TestNoOrdinaryKeyEndsInAPluralSuffix. - Saved notes (v0.254.0): a background run has no request and no reader, so a note it SAVES is
written in the BOX's language at write time —
util.Text(lang, key, …)viam.note/s.note(m.noteErr/s.noteErrfor an error). Consequence, by operator ruling (slice 2 §16 option 1): a household that switches sees the previous run's note in the old language until the next run. TRAP, and it cost a 25-minute hang:UpdateOffboxStatusand its siblings hold the settings WRITE lock while running their callback, andboxLang()wants the READ lock —sync.RWMutexis not reentrant, so a note rendered inside such a callback DEADLOCKS while holding the settings lock. Resolve the language BEFORE the callback;TestNoteHelpersAreNotCalledUnderTheSettingsLockguards it. - The switch is a globe (v0.254.0):
templates/lang_globe.html, a<details>/<summary>menu with no script, drawn inline (the sprite lives only inlayout.html, so<use>would render nothing on the shells). Language names are shown in their own language and never translated. - A visitor's language (v0.254.0): the sign-in, claim and recovery pages carry the same globe, posting
to
POST /lang, which sets the display-onlyfelhom_langcookie.langFor's order:?lang=→ household setting when a session exists → cookie when none → setting →hu; a signed-in household never reads the cookie./langis CSRF-exempt because its only achievable effect is the language of the page the victim's own browser shows them — if it gains a second effect it needs CSRF that day.safeBackPathrefuses a protocol-relative//hostas well as an absolute URL. A successful CLAIM carries the cookie into the household setting (§16). No globe on the guest share pages or the not-found page — R-577. - Report: the hub report carries
"language"(always present). No hub release reads it yet. - Tools and gates:
scripts/i18n_extract.pyconverts a template (moves each Hungarian run into hu.json, leaves a marker);scripts/i18n_missing_gate.py(incontroller_gates.py) checks keys exist, no orphans, the English gap and the formal „ön" count only shrink (ratchets), no pleading English.scripts/i18n_go_parity.py(v0.252.0, incontroller_gates.py) freezes every Go string literal at slice 2's base commit (scripts/i18n_go_base.json) and refuses a key whose Hungarian is not that text byte for byte;scripts/i18n_go_keys.jsonrecords what each key replaced. Three decoys. The emoji, native-confirm, secret-in-markup and retrieval-promise gates read templates expanded throughscripts/i18n_bundle.py; mojibake scans the bundles. - Coverage:
TestI18nParityCoversEveryMarkerfails when any marker occurrence is rendered by no parity case. Fixtures are written only when missing (-update-i18n-golden,-i18n-golden-only). - Release gate:
TestI18nParitycompares the Hungarian render of every fixture state with HTML captured from the unconverted templates (internal/web/testdata/i18n_parity/). Fixtures are never regenerated to make a conversion pass.
Catalog copy in a second language (v0.257.0, R-560 slice 5)
The message bundle above covers text the CONTROLLER owns. An app's own words — its one-line
description, tagline, use_cases, first_steps, prerequisites, default_creds, the deploy-field
labels/descriptions/placeholders and select-option labels, optional_config group and field text,
integration labels, data_paths labels and the initial-credentials note — come from the catalog's
.felhom.yml and are Hungarian. They are NOT bundle keys: they belong to the catalog, change on a
catalog push, and a box may run a catalog the controller has never seen.
-
Format: a sibling block in the SAME file, so a reviewer sees one app whole:
description: "Titkosított jegyzet és szöveg megosztás" i18n: en: description: "Encrypted note and text sharing"Metadata.I18n map[string]MetadataOverlay(internal/stacks/metadata_i18n.go),json:"-"— the untranslated half never travels to a page or an API answer; the merged view does. -
Read path:
Metadata.For(lang), reached only throughstacks.LocalizeStacks(a list),LocalizeStackPtr(one stack) orStack.MetaFor.For("hu")is the parsed struct withI18ncleared and nothing else —TestMetaForHuIsIdentitydeep-compares it over all 53 real catalog files ininternal/stacks/testdata/catalog/.LocalizeStacksreturns the INPUT SLICE for Hungarian rather than a copy. -
Field-by-field fallback: an English field that is absent — or blank — shows the Hungarian one. A half-translated app is a legal state, which is what lets the catalog be translated in batches.
-
The three prose lists replace WHOLE (
use_cases,first_steps,prerequisites): merging by index produces a list nobody wrote. Every other list is matched by its own key —deploy_fieldsbyenv_var, options byvalue,optional_configgroups bymatch_group(the Hungariangroupvalue they translate; a group has no other identity), its fields byenv_var, integrations bytarget,data_pathsbypath. Position matching mistranslates silently. -
Never writes through. The metadata is the stack manager's and is shared by concurrent requests; an in-place merge would put one household's language on another's page.
-
Pages: the dashboard rows, the Apps list, the logs page, the deploy page and the app page.
TestNoDirectMetaCopyReadOnPagesAST-parsesinternal/weband fails on anyx.Meta.<copy>read outside a named, reasoned allow-list — so a NEW page that reads catalog copy off the manager fails the suite instead of quietly rendering Hungarian to an English household. -
Two producers have no request and therefore no language — the integration rows (
internal/integrations) and the initial-credentials note (internal/stacks/initialcreds.go). Both are re-taken from the localised metadata in the handler. -
Older controllers are unaffected:
LoadMetadatauses non-strictyaml.Unmarshal(this repo constructs noyaml.Decoder), so a pre-0.257.0 box drops the whole block. Proven live on a 0.255.0 box: pages byte-identical before and after the block arrived, no parse warning in a live log. -
The catalog side carries the freeze and the gate:
app-catalog-felhom.eu/scripts/copy_freeze/hu.json(every Hungarian copy string, captured before any translation) andscripts/check-copy-i18n.py(freeze + structure + language + credential tokens + a coverage ratchet), fifth row of that repo'scatalog_gates.py.
Repository Layout
controller/
├── cmd/controller/main.go # Entry point, wires all 17 modules (setup mode branch + normal startup)
├── internal/
│ ├── config/config.go # YAML loader, validation, env overrides
│ ├── crypto/crypto.go # AES-256-GCM encryption for app.yaml secrets, key management
│ ├── settings/settings.go # Runtime settings (JSON, atomic writes, RWMutex)
│ ├── stacks/
│ │ ├── manager.go # Stack scanning, compose ops, container status
│ │ ├── metadata.go # Parse .felhom.yml app metadata
│ │ ├── deploy.go # First-deploy: secret gen, app.yaml, compose up; missing field injection
│ │ └── delete.go # Stack deletion/removal + HDD/backup data cleanup
│ ├── sync/sync.go # Git sync: clone/pull app catalog, content-hash copy
│ ├── storage/
│ │ ├── scan.go, scan_linux.go # Disk detection via lsblk + blkid
│ │ ├── format.go, format_linux.go # Partition, format, mount pipeline
│ │ ├── attach.go, attach_linux.go # Attach existing FS drive (raw mount + bind mount)
│ │ ├── safety.go, safety_linux.go # System disk detection, mount guards, fstab ops
│ │ ├── migrate.go # App data migration (rsync with progress)
│ │ └── *_other.go # Non-Linux stubs for cross-compilation
│ ├── backup/
│ │ ├── backup.go # Orchestrator (per-drive dumps + restic + cross-drive chain)
│ │ ├── paths.go # Per-drive path helpers (FelhomDataDir constant, PrimaryResticRepoPath, AppDataDir, InfraBackupDir, etc.)
│ │ ├── local_infra.go # Local infra backup to all drives (.felhom-infra-backup/)
│ │ ├── dbdump.go # DB auto-discovery + dump (pg_dump, mariadb-dump)
│ │ ├── restic.go # Restic operations (init, snapshot, prune, check) — repoPath as param
│ │ ├── appdata.go # StackDataProvider interface, app data discovery
│ │ ├── crossdrive.go # Per-app backup to secondary storage (rsync/restic)
│ │ ├── restore.go # Per-app restore from per-drive repo
│ │ ├── restore_scan.go # DR: scan drives for backup data, build restore plan
│ │ ├── restore_app_linux.go # DR: per-app restore (rsync config/data + docker compose up)
│ │ └── restore_drives_linux.go # DR: auto-mount drives by UUID from Hub infra backup
│ ├── cloudflare/
│ │ ├── client.go # CF API client (Bearer auth, generic JSON helper)
│ │ ├── zone.go # Zone ID resolution (domain → zone)
│ │ ├── waf.go # WAF rule CRUD + expression builders
│ │ ├── countries.go # ISO 3166-1 country codes + Hungarian names
│ │ └── geosync.go # Geo sync orchestrator (diff & apply rules)
│ ├── integrations/
│ │ ├── integrations.go # Core types: Handler interface, ApplyContext, StatusInfo
│ │ ├── manager.go # Manager: Toggle, ListForProvider, StackProvider interface
│ │ ├── lifecycle.go # OnStackStop, OnStackStart, OnStackRemove hooks
│ │ ├── onlyoffice_filebrowser.go # OnlyOffice → FileBrowser handler (config.yaml patch)
│ │ └── onlyoffice_nextcloud.go # OnlyOffice → Nextcloud handler (occ commands)
│ ├── assets/syncer.go # Hub asset sync (download, SHA-256 compare, resolve)
│ ├── api/
│ │ ├── router.go # REST API endpoints (~36 routes)
│ │ └── geo.go # Geo-restriction API handlers
│ ├── scheduler/scheduler.go # Central job scheduler (Every, Daily)
│ ├── system/
│ │ ├── info.go, info_linux.go # RAM, disk, CPU, temperature, load average
│ │ ├── cpu_linux.go # Background /proc/stat sampling
│ │ └── mounts_linux.go # Mount points, disk usage, FS info, backup dest checks, storage probing, USB detection
│ ├── monitor/
│ │ ├── pinger.go # Healthchecks.io HTTP ping client
│ │ ├── healthcheck.go # System health checks (disk, mem, CPU, temp, Docker)
│ │ └── watchdog.go # Storage watchdog (probe, disconnect/reconnect, safe eject)
│ ├── metrics/
│ │ ├── store.go # SQLite time-series (WAL mode, downsampled queries)
│ │ ├── collector.go # Background collector (60s, system + docker stats)
│ │ └── sysinfo.go # Static system info (/proc, /etc)
│ ├── selfupdate/
│ │ ├── version.go # Semver parsing + comparison (hand-rolled)
│ │ ├── state.go # Update audit state (JSON, atomic writes)
│ │ └── updater.go # Registry check, update trigger, startup verify
│ ├── notify/notifier.go # Email relay to hub, preference sync, cooldowns
│ ├── report/
│ │ ├── builder.go # Hub report builder (all subsystems → JSON)
│ │ ├── pusher.go # HTTP POST to hub (retry, Bearer auth, parses customer_blocked)
│ │ └── infra_pull.go # DR: pull recovery/config from Hub (retrieval password auth)
│ ├── setup/ # First-run setup wizard (web-based, replaces docker-setup.sh wizard)
│ │ ├── setup.go # NeedsSetup() detection, state persistence
│ │ ├── handlers.go # HTTP handlers for all wizard steps
│ │ ├── scanner.go # Drive scanner for local infra backups
│ │ ├── csrf.go # Lightweight CSRF (cookie + hidden field)
│ │ ├── network.go # Local IP detection for LAN access URLs
│ │ └── templates/ # 7 wizard HTML templates (Hungarian)
│ ├── recovery/info.go # Recovery info file generator (recovery-info.txt)
│ └── web/
│ ├── server.go # HTTP server, routing, static files, catch-all middleware, executeTemplate wrapper
│ ├── auth.go # Session auth + per-session CSRF token, login/logout, session cleanup
│ ├── csrf.go # CsrfProtect middleware, csrfToken/csrfField helpers
│ ├── handlers.go # Page handlers (dashboard, stacks, deploy, backups, etc.)
│ ├── handler_restore.go # DR: restore page handler + APIs (scan, restore all, skip)
│ ├── handler_debug.go # Debug page handler + 20 debug API endpoints (debug-mode only)
│ ├── logbuffer.go # Ring buffer (io.Writer) for in-memory log capture
│ ├── storage_handlers.go # Storage API handlers (scan, format, attach, migrate, cleanup, disconnect/reconnect)
│ ├── alerts.go # State-based alert generation
│ ├── funcmap.go # Template functions (state colors, Hungarian formatting)
│ ├── embed.go # go:embed for templates + Chart.js
│ └── templates/ # 21 HTML files (4-page settings split) + style.css + icons.html sprite (Hungarian UI, design system v2)
├── configs/
│ ├── controller.yaml.example # Full config reference
│ └── example-felhom-metadata.yml # .felhom.yml format reference
├── Dockerfile # Multi-stage: Go 1.24 builder + debian-slim runtime
├── docker-compose.yml # Controller's own compose (privileged, /mnt rshared)
└── go.mod # Go 1.24, deps: bcrypt, yaml.v3, modernc.org/sqlite
Configuration
18.5 The e-mails the hub sends (v0.256.0, slice 3)
The hub writes a household's e-mails in the household's language from hub v0.118.0. It learns
which language from the box: every report carries "language" (since v0.247.0), and the hub prefers
the last reported one over the operator's creation-time default. Nothing here needs a setting.
The box sends about a third of those sentences itself, and the hub cannot translate one — it
arrives as finished text naming a drive, an app or a number. So POST /api/v1/event carries
message_customer: the same sentence in the household's language, beside the unchanged Hungarian
message.
messageis always Hungarian. It is what the operator is mailed, what the hub logs and whatnotification_logrecords; a household switching language moves nothing an operator reads.- A Hungarian household sends no second copy at all (
omitempty), so its payload is byte-for-byte what every box has always sent. - 19 producers render both from ONE bundle key (
event.*): health ×3, controller started/updated/ update-failed, storage disconnected/reconnected, backup target absent/restored, app deployed/ deploy-started/deploy-failed/removed, crossdrive ×2, disaster recovery ×2, db-dump completed. - An untranslatable tail (a docker error, a validator's sentence) is appended to both renderings.
- Not converted, deliberately: the operator-tier types (the hub keeps them off the customer channel) and the producers whose sentence is composed in another package and reaches the notifier already finished. Those still send Hungarian only, so an English household can still see one Hungarian line in some mails. Tracked as a row.
A new box starts in the language the operator picked. The hub renders customer.language into
controller.yaml; GetLanguage prefers the household's stored choice, then that, then Hungarian.
The config value is never written into settings.json — persisting it would record a choice the
household never made.
Controller config (controller.yaml)
Single YAML file per customer, infrastructure-only. Does not contain app-specific config.
Key sections:
customer:
name: "Demo Felhom"
id: "demo-felhom"
paths:
stacks_dir: "/opt/docker/stacks"
data_dir: "/opt/docker/felhom-controller/data"
system_data_path: "/mnt/sys_drive" # NVMe/system drive — fallback for apps without HDD
git:
repo_url: "https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git"
sync_interval: "15m"
# Per-drive backup paths are computed automatically:
# <drive>/backups/primary/restic/ — restic repo per drive
# <drive>/backups/primary/<app>/db-dumps/ — DB dumps per app
# <drive>/backups/secondary/ — cross-drive rsync + restic
backup:
enabled: true
restic_password_file: "/opt/docker/felhom-controller/data/restic-password"
db_dump_schedule: "02:30"
restic_schedule: "03:00"
retention: { keep_daily: 7, keep_weekly: 4, keep_monthly: 6 }
monitoring:
health_interval: "5m"
ping_uuids:
heartbeat: "uuid-here"
system_health: "uuid-here"
db_dump: "uuid-here"
backup: "uuid-here"
backup_integrity: "uuid-here"
web:
listen: ":8080"
setup_listen: ":8081" # Plain HTTP for setup wizard LAN access
hub:
enabled: true
url: "https://hub.felhom.eu"
api_key: "bearer-token-here"
assets:
sync_enabled: true # Download app assets (logos, screenshots) from Hub API
sync_schedule: "05:00" # Daily sync time (HH:MM, Budapest timezone)
system:
reserved_memory_mb: 384 # RAM reserved for OS + controller
Environment variable overrides: FELHOM_LOGGING_LEVEL=debug, FELHOM_HUB_ENABLED=false, etc.
Runtime settings (settings.json)
Auto-managed by the controller. Contains password hash overrides, notification preferences, per-app backup configs, storage path registry, DB validation cache, Hub verification state (hub_verified, hub_verified_at), retrieval password for disaster recovery, and pending event queue. All writes are atomic (write .tmp, rename).
Per-app config (app.yaml)
Auto-generated during deployment. Contains env vars, locked fields list, deploy timestamp. Secret fields are locked (read-only after first deploy). Missing fields from updated templates are auto-injected on startup and after sync (see Missing Field Injection).
Encryption at rest: Sensitive env values (type: password and type: secret from .felhom.yml metadata) are stored encrypted as ENC:base64(nonce+ciphertext) using AES-256-GCM. The 32-byte encryption key is stored at {dataDir}/encryption.key (generated on first run, 0600 permissions). Values are decrypted transparently when passed to docker-compose or displayed in the UI. The key is included in infra backups (Hub + local drives) and restored during disaster recovery. On upgrade, existing plaintext values are migrated automatically on startup.
Scheduler Jobs
| Job | Type | When | Purpose |
|---|---|---|---|
| status-refresh | periodic | 30s | Refresh container states |
| stack-scan | periodic | 2m | Rescan stacks directory |
| heartbeat | periodic | 5m | Legacy Healthchecks ping (deprecated — Hub handles via event system) |
| system-health | periodic | configurable | Health checks + alert refresh |
| backup-cache | periodic | 5m | Refresh backup status cache |
| hub-report | periodic | 15m | Push report to central hub |
| db-dump | daily | 02:30 | Database dumps |
| backup | daily | 03:00 | Restic backup → cross-drive chain |
| backup-integrity | daily | Sun 04:00 | Restic check |
| metrics-prune | daily | 04:00 | Delete metrics older than 30 days |
| selfupdate-check | periodic | 6h | Check registry for new version (cache for UI) |
| selfupdate-auto | daily | 04:30 | Auto-update if enabled + backup not running |
| asset-sync | daily | 05:00 | Download changed app assets from Hub |
All daily jobs use Europe/Budapest timezone. Skip-if-running prevents concurrent execution. Panic recovery in all jobs.
REST API
Stack Operations
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/health |
Health check (no auth) |
| GET | /api/stacks |
List all stacks |
| GET | /api/stacks/{name} |
Stack details |
| POST | /api/stacks/{name}/deploy |
First-time deploy |
| POST | /api/stacks/{name}/start |
Start stack (409 if insufficient memory) |
| POST | /api/stacks/{name}/stop |
Stop stack |
| POST | /api/stacks/{name}/restart |
Restart stack |
| POST | /api/stacks/{name}/update |
Pull + recreate |
| POST | /api/stacks/{name}/optional-config |
Update optional env vars |
| GET | /api/stacks/{name}/logs |
Container logs (?raw=1 for plain text) |
| GET | /api/stacks/{name}/hdd-data |
HDD data paths + sizes — resolved from the app's OWN app.yaml HDD_PATH (v0.236.0, R-442), never the global config |
| GET | /api/stacks/{name}/backup-data |
Backup data paths + sizes (DB dumps, cross-drive rsync) |
| POST | /api/stacks/{name}/remove |
Remove deployed stack (revert to "not deployed"). remove_hdd_data: true deletes the app's folders under its recorded HDD_PATH and lists them; 409 + a Hungarian sentence when the data was asked for but its location cannot be resolved or the drive is absent — nothing is touched, the app is kept (v0.236.0, R-442). hdd_paths_removed is [] for an SSD app (never null); hdd_paths_missing, hdd_note, backup_paths_refused state what was not found / not removed. remove_backups: true (v0.240.0, R-474) deletes the app's whole recovery unit, its Tier-2 mirror(s) on any registered drive and its backup preferences (backup_paths_removed lists them); without it the backups AND the Tier-2 record are kept, so the removed app can still be restored from the second drive (R-486). Off-site snapshots are never touched by removal. Since v0.242.0 a removal also clears the app's update hold (R-491) and volumes_removed lists the named volumes actually removed, [] when none (R-489 — a volume recreated by a unit restore carries no compose label and is not listed yet, measured 2026-09-13; the row stays open) |
| DELETE | /api/stacks/{name} |
Delete orphaned stack — same R-442 resolution and refusal shape as /remove |
| POST | /api/sync |
Trigger catalog sync |
| GET | /api/system/info |
System info + sync status |
Backup & Restore
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/backup/status |
Full backup status |
| POST | /api/backup/run |
Trigger manual backup |
| GET | /api/backup/snapshots |
List snapshots (?stack={name} for filtering). Since v0.242.0 it also answers for a REMOVED app whose unit is on a connected drive (R-487) — 404 only when no unit exists anywhere |
| POST | /api/stacks/{name}/cross-backup |
Save cross-drive config |
| POST | /api/stacks/{name}/cross-backup/run |
Trigger cross-drive backup |
| GET | /api/stacks/{name}/cross-backup/status |
Cross-drive status |
| POST | /api/backup/cross-drive/run-all |
Run all scheduled cross-drive backups |
| GET | /backup/offbox/status |
Offsite run status + live progress for a manual run (v0.147.x) |
| POST | /backup/offbox/verify-copy/delete |
Delete ONE verification copy (stack name + confirm=1; v0.147.0) |
| GET | /sharing/status |
SMB bring-up phase + live container state (v0.147.0) |
Storage
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/storage/scan |
Scan available disks |
| POST | /api/storage/init |
Format and mount a disk |
| GET | /api/storage/init/status |
Format progress |
| POST | /api/storage/attach/mount-raw |
Temp-mount partition for browsing |
| GET | /api/storage/attach/browse?path= |
List directories on raw mount |
| POST | /api/storage/attach/mkdir |
Create folder on raw mount |
| POST | /api/storage/attach |
Finalize attach (bind mount + fstab) |
| GET | /api/storage/attach/status |
Attach progress |
| POST | /api/storage/attach/cancel |
Cleanup temp raw mount |
| POST | /api/storage/migrate |
Start app data migration |
| GET | /api/storage/migrate/status |
Migration progress |
| POST | /api/storage/disconnect |
Safe disconnect (stop apps, unmount) |
| POST | /api/storage/reconnect |
Reconnect disconnected drive |
| POST | /api/storage/restart-apps |
Restart auto-stopped apps |
| GET | /api/storage/status |
All storage paths with connection state |
Self-Update
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/selfupdate/status |
Update status (cached check result + last state) |
| POST | /api/selfupdate/check |
Force registry check |
| POST | /api/selfupdate/update |
Trigger self-update (async) |
Self-update endpoints accept session auth OR Authorization: Bearer <hub_api_key> for external triggering.
Config Management
| Method | Endpoint | Description |
|---|---|---|
| POST | /api/config/apply |
Apply new controller.yaml from Hub (atomic write) |
| GET | /api/config/hash |
Get SHA256 hash of current controller.yaml |
| GET | /api/config |
Get raw controller.yaml content (text/yaml) for live diff and pull |
Config endpoints accept session auth OR Authorization: Bearer <hub_api_key> (same as self-update). The /api/config/apply endpoint:
- Accepts raw YAML body (the generated config from Hub)
- Validates YAML is parseable before writing
- Atomic write: writes to
.tmpthenos.Renamefor crash safety - Does NOT reload config — restart required to apply changes
- Returns
{"ok": true, "message": "Config applied. Restart controller to apply changes."}
Metrics
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/metrics/system |
System metrics time-series (`?range=1h |
| GET | /api/metrics/containers/summary |
Current container stats |
| GET | /api/metrics/containers/{name} |
Per-container time-series |
| GET | /api/metrics/sysinfo |
Static system info |
Assets
| Method | Endpoint | Description |
|---|---|---|
| POST | /api/assets/sync |
Trigger on-demand asset sync from Hub (async) |
| GET | /api/assets/status |
Asset sync status (last sync, file count, total bytes) |
Integrations
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/integrations/{provider} |
List integrations for provider app (status, target availability) |
| POST | /api/integrations/{provider}/{target} |
Enable/disable integration ({"enabled": true/false}) |
Debug (debug mode only)
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/debug/dump |
Full diagnostic JSON dump (controller state, storage, stacks, backup, hub, scheduler, health, alerts). Returns 404 when logging.level is not "debug". |
| GET | /api/debug/telemetry |
Run telemetry collection on-demand; returns per-app metrics + log summary with latency. Response: {latency_ms, app_count, total_errors, total_warnings, app_telemetry[]}. |
Response format: {"ok": true/false, "data": ..., "error": "...", "message": "..."}
App-email relay (internal/mailrelay/)
Gives deployed apps outbound email (password resets, invites, confirmations) through one managed path —
app → in-process SMTP shim → hub → Resend — with the Resend key staying hub-side (never on the box).
Architecture Shape 1: the shim runs in-process inside the controller, reusing the existing hub client.
Implements felhom.eu/documentation/audits/SPIKE-smtp-app-relay-2026-06-28.md.
- The shim (
internal/mailrelay/) is ago-smtpserver with two listeners —:2525plaintext+STARTTLS and:2465implicit-TLS (self-signed cert at boot). It advertises AUTH PLAIN+LOGIN and accepts any credentials, ignoring them (apps send none; some require the offer).Datareads the raw message, enforces the From-header domain allowlist (reject5xxbefore any hub call), then forwards the raw MIME to the hubPOST /api/v1/mailwith the controller's hub Bearer key — single-shot (no retry, no spool in v1). The hub HTTP status maps to an SMTP reply (2xx→250, 4xx→451, 5xx→554). Listeners bind to the app Docker network only (the controller container joinstraefik-public); apps reach the shim byfelhom-controller. - Lifecycle (
lifecycle.go): the shim starts/stops at runtime to match the global app-email toggle (no controller restart). Wired inmain.go, gated on a configured hub + themail_relaykill-switch. - Toggles + injection: a global toggle (
settings.AppEmail, Settings page) and a per-app toggle (AppConfig.EmailEnabled, on the app's config page, shown only for apps with ansmtp_mapping). When both are on and the app declares.felhom.ymlsmtp_mapping,stackEnvinjects the relay env at compose time (host=shim, port=2525, security/from per the mapping, From=<app>@felhom.eu) — derived each compose, never persisted toapp.yaml. Config knobs:mail_relay(listeners,shim_host,from_domains, kill-switch). - What the box never holds: the Resend key, or any durable mail queue. v2 (deferred) = a separate
felhom-smtp-shimcontainer + accept-and-spool retry + aResend-Idempotency-Key.
Build & Deploy
Build
# On build server (192.168.0.180)
cd ~/build/felhom-controller
git -C ~/git/felhom-controller pull
./build.sh v0.20.0 --push
Deploy on customer node
Option A: Self-Update API (v0.16.0+)
After building and pushing the new image, trigger the controller's self-update endpoint:
curl -s -X POST https://felhom.demo-felhom.eu/api/selfupdate/update \
-H "Authorization: Bearer <HUB_API_KEY>"
The controller pulls the new image, updates its own compose file, and runs docker compose up -d to replace itself. The Settings page also has a "Frissítés telepítése" button for manual triggering.
Option B: Manual SSH (pre-v0.16.0 or fallback)
# On customer node (e.g., 192.168.0.162)
cd /opt/docker/felhom-controller
sudo docker pull gitea.dooplex.hu/admin/felhom-controller:<VERSION>
sudo sed -i 's|image: gitea.dooplex.hu/admin/felhom-controller:.*|image: gitea.dooplex.hu/admin/felhom-controller:<VERSION>|' docker-compose.yml
sudo docker compose up -d
Important: Always use docker compose up -d, NOT docker compose restart — restart doesn't pick up new images.
Docker Requirements
The controller container needs:
privileged: true(disk operations)- Docker socket mount (
/var/run/docker.sock) /mntmount withpropagation: rshared(container mounts visible to host)/devmounted as/host-dev(block device access)/etc/fstabmounted as/host-fstab(persistent mount config)
See docker-compose.yml for the full volume configuration.
Roadmap
Completed
- Stack management with deploy flow and memory validation
- Git-based app catalog sync
- Central job scheduler
- System monitoring with SQLite metrics and Chart.js charts
- Healthchecks.io integration (5 ping types)
- 3-layer backup system (DB dumps + restic + cross-drive)
- Per-app backup restore with auto stop/restart
- Storage management (scan, format, mount, registry)
- Attach existing drive wizard (v0.15.0) — bind-mount subfolder from pre-formatted drive, directory browser
- App data migration between storage paths
- Storage watchdog (v0.17.0) — USB disconnect detection (~15s), auto-stop apps, auto-remount on reconnect, safe eject UI
- Central hub reporting
- Email notifications via hub relay
- Settings persistence and password management
- Dashboard alert system
- Per-drive backup architecture (v0.14.0) — per-drive restic repos, per-app DB dumps, path helpers
- Cross-drive restic pruning (v0.14.0)
- Auto Tier 2 for small apps (v0.14.1) — auto-enable daily rsync for non-HDD apps when ≥2 drives
- Infrastructure config in cross-drive backup (v0.14.1) — stacks dir + controller.yaml in
_infra/+ restic - Disaster recovery (v0.15.5) — Hub-based infra backup, auto-mount by UUID, restore UI with full-page takeover
- Controller self-update (v0.16.0) — Watchtower-style pull + restart, Settings page UI, API key auth, auto-update scheduling
- Hub-managed config (v0.20.0) — Config apply endpoint (
POST /api/config/apply), config hash in reports for sync comparison - Config content endpoint (v0.21.1) —
GET /api/configreturns raw YAML for Hub live diff and pull operations - First-run setup wizard (v0.22.0) — Web-based wizard replaces shell scripts, drive scan for local backups, Hub recovery, fresh install flow
- Setup wizard logo fix (v0.22.2) — Use embedded SVG instead of filesystem path
- Hub-managed asset sync (v0.22.3) — Download app logos/screenshots from Hub API with SHA-256 change detection, daily sync schedule
In Progress / Planned
- Update classification and auto-apply (optional/required/security markers)
- Docker volume backup + Tier 2 restore (v0.33.0)
- Raspberry Pi testing (pi-customer-1)
- CSRF protection on POST endpoints (v0.23.0)
- Verbose debug logging across all modules (v0.24.0)
- Diagnostic dump endpoint
/api/debug/dump(v0.24.0) - Startup self-test with 9 subsystem checks (v0.24.0)
- Login rate limiting
Test Environments
| Node | Hardware | Domain | Status |
|---|---|---|---|
| demo-felhom | Acemagic GK3PLUS N100, 16G RAM, 512G SSD + 1TB HDD | demo-felhom.eu | Active |
| felhotest | Proxmox VM (4-16G RAM, 8 vCPU, 200G + 100G SCSI) | router.abonet.hu:33022 | Active |
| pi-customer-1 | Raspberry Pi 3B+, 1G RAM, 32G SD | pi-customer-1.local | Not yet tested |
Related Repositories
| Repository | Purpose |
|---|---|
| felhom-controller | This repo — controller + deploy scripts |
| app-catalog-felhom.eu | Docker Compose templates + .felhom.yml metadata |
| felhom.eu | Website + app assets + felhom-hub service |