New documentation/controller/ subtree (module map + deploy/stack-lifecycle, backup, storage/monitoring/metrics, auth/hub/sync/integrations) grounded in current source; top-level documentation/README.md index across controller/agent/platform/hub/audits; REORG-NOTES with the verification ledger + flagged doc-gaps. Supersedes (keeps) the v0.33 controller planning map. Additive only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
15 KiB
Controller: deploy & stack lifecycle
Source of truth: felhom-controller
internal/stacks/,internal/infra,internal/syncat v0.59.0.
This document describes how the in-guest controller represents application stacks, deploys them, runs their lifecycle (update/stop/start/restart/remove/delete), and brings up the base infrastructure (traefik / cloudflared / filebrowser) that everything else depends on. Every claim below is verified against the current source.
1. The stack model
A stack is one customer app, represented on disk as a directory under
/opt/docker/stacks/<name>/ (config key paths.stacks_dir, default /opt/docker/stacks)
containing:
docker-compose.yml— the compose template (synced from the catalog, see §7)..felhom.yml— app metadata (Metadata): display name, deploy fields, resource requests/limits, optional config, healthcheck spec, integrations.app.yaml— the per-app deployment record (AppConfig), written only after the app is first deployed.
In memory each stack is a Stack struct (internal/stacks/manager.go:62):
| Field | Meaning |
|---|---|
State (ContainerState) |
aggregated container state |
Deployed |
has an app.yaml with deployed: true |
Deploying |
a compose up is in progress (image pull, etc.) |
Protected |
a system stack that must never be stopped/removed (§4) |
Orphaned |
deployed but no longer present in the synced catalog |
DeployError |
last async deploy error, surfaced to the UI |
HealthProbe |
latest controller-side probe result (§6) |
ContainerState values (manager.go:23): running, starting (health: starting),
unhealthy, stopped, restarting, exited, paused, unknown, not_deployed,
deploying, orphaned. The effective state is derived by resolveContainerState
(combines Docker State + the health hint in Status) and then aggregateState across
all of a stack's containers, with priority `unhealthy > starting > restarting > all-running
stopped
(manager.go:437,:468`).
ScanStacks (manager.go:250) discovers stacks by walking stacks_dir, reading metadata
and app.yaml, marking Protected via cfg.IsProtectedStack, and computing Orphaned
against the synced catalog cache (getCatalogTemplateSlugs, manager.go:1078). While an
async deploy is in flight it deliberately does not overwrite Deployed/AppConfig —
the deploy goroutine owns those (H3 fix, manager.go:293).
AppConfig (app.yaml)
type AppConfig struct {
Deployed bool `yaml:"deployed"`
DeployedAt string `yaml:"deployed_at"`
Env map[string]string `yaml:"env"`
LockedFields []string `yaml:"locked_fields"`
}
(deploy.go:98). Env holds the resolved deploy-field values; sensitive ones are stored
encrypted (§3). LockedFields lists env vars marked locked_after_deploy that
UpdateStackConfig refuses to change (deploy.go:434).
2. Deploy flow
Entry point: DeployStack(req) (deploy.go:121), driven by the deploy form ("Telepítés").
-
Atomic check-and-set of
Deploying(H1). Under the manager lock, the stack must exist, must not already beDeploying, and must not already beDeployed; thenDeployingis set true (deploy.go:122). This single critical section prevents two concurrent deploys of the same app. Any later validation failure clears the flag viaclearDeploying(). -
Memory gate.
system.GetMemoryMB()vs the app'smem_requestand the configuredreserved_memory_mb. A hard block (real used + new request > usable) aborts with a Hungarian error; a soft over-commit case returns a non-fatal warning string (deploy.go:159). -
Field resolution. For each
DeployField:domainis auto-filled fromcustomer.domain;subdomainis validated (DNS-safe, not reserved, not already in use);secretuses the value previewed to the user or generates one;passwordis required (never silently generated — the user must know it);pathmust exist on the host. Required-but-empty fails (deploy.go:206–292). -
app.yamlis persisted withdeployed: falsefirst (CTRL-T2-1). The in-memoryAppConfigcarriesDeployed: true(for UX), but a clone withDeployed: falseis whatSaveAppConfigwrites to disk at this point (deploy.go:302–314). The durable record is intentionally not marked deployed yet. -
In-memory state flips to deployed (
s.Deployed = true,s.AppConfig = appCfg), so the UI immediately stops showing a stale "Telepítés" button during the pull (deploy.go:331). -
Async compose.
go m.runComposeDeploy(...)runsdocker compose up -doff the request path, so the HTTP handler returns at once and the UI polls for progress (an image pull can take 30–60s). The compose command injects the resolved env plusDOMAIN(deploy.go:338,composeExecWithEnvmanager.go:505).
CTRL-T2-1 crash-safety (the key current behaviour)
The on-disk app.yaml records deployed: true only after docker compose up -d
actually succeeds. runComposeDeploy (deploy.go:346):
- On compose failure: reverts in-memory state (
Deployed=false,Deploying=false, recordsDeployError,AppConfig=nil) and re-saves the revertedAppConfig(stilldeployed:false) to disk (deploy.go:350). - On compose success: re-runs
SaveAppConfigwithappCfg.Deployed == true, flipping the durable record to deployed (deploy.go:371). If that save fails, it reverts in memory so the customer can cleanly redeploy rather than be stuck half-recorded (deploy.go:375). - Only then is
Deployingcleared andRefreshStatus()run.
Why it matters: a crash or power loss during the image-pull window leaves app.yaml
at deployed: false. On restart, ScanStacks therefore sees the app as not deployed
(no ghost-deployed stack with no containers), and DeployStack will cleanly redeploy
instead of refusing with "already deployed".
The UI's three-step progress (form submit → "pulling/deploying" via the deploying
state and DeployError → settled running/unhealthy) is driven entirely by this
in-memory state, which the front-end polls.
3. SaveAppConfig and secret encryption (fail-closed, H10)
SaveAppConfig(stackDir, cfg, encKey, sensitiveVars) (deploy.go:669) clones the env and,
for each var named in sensitiveVars (the secret/password deploy fields, computed by
SensitiveEnvVars, deploy.go:741), encrypts the value with the AES-256 key
(crypto.Encrypt) unless it is empty or already encrypted.
Fail-closed (H10, v0.59.0): if encryption of a sensitive value errors, SaveAppConfig
returns the error and writes nothing — it never falls through to a plaintext write
(deploy.go:683–691). Earlier code logged a WARN and persisted plaintext, which leaked
the secret to disk; that path is gone. Callers (DeployStack, UpdateStackConfig, etc.)
already propagate this error, so the deploy fails cleanly with no secret on disk.
The write itself is atomic: write to app.yaml.tmp then rename (H04, deploy.go:711),
mode 0600. Decryption for compose happens lazily via LoadAppConfigDecrypted /
stackEnv (deploy.go:727, manager.go:809), which injects DOMAIN plus the decrypted
env into the compose process environment. A one-time startup MigrateEncryption
(manager.go:128) re-saves any deployed app still holding plaintext sensitive values.
4. Lifecycle: update / stop / start / restart / remove / delete
All operations resolve the stack dir, build env via stackEnv (decrypted), and shell out
to compose. Notably start/restart use up -d, not bare restart, so env changes and
template changes (new images, healthchecks) are picked up (manager.go:709, :643).
- Start (
StartStack,manager.go:643):compose up -d; clears the stale health probe so the next probe runs fresh. - Stop (
StopStack,manager.go:682): refuses protected stacks, thencompose down. - Restart (
RestartStack,manager.go:709):compose up -d; clears health probe. - Update (
UpdateStack,manager.go:746):compose pullthencompose up -d --remove-orphans. - Config update (
UpdateStackConfig,deploy.go:405): rejects locked fields, re-savesapp.yaml(encrypting secrets), thenup -dwith the decrypted env. - Remove (
RemoveStack,delete.go:283): for a deployed (non-orphaned) stack —compose down --volumes(keeps images for redeploy), optional HDD-data and backup-path cleanup, then deletes onlyapp.yaml, reverting the stack to "not deployed". Compose template files are preserved so the user can redeploy. - Delete (
DeleteStack,delete.go:80): for an orphaned stack only —compose down --rmi local --volumes, optional HDD-data removal, then removes the whole stack directory.
Protected-stack enforcement (server-side, defence in depth)
The protected set is configured under stacks.protected (default
traefik, cloudflared, felhom-controller, filebrowser — see
configs/controller.yaml.example:55 and the setup writer internal/setup/handlers.go:476).
config.IsProtectedStack matches case-insensitively (config.go:347). Enforcement exists
at three layers:
- Router/API:
actionStackblocks any action other thanrestarton a protected stack with HTTP 403 (internal/api/router.go:412). - Manager:
StopStack(manager.go:683),RemoveStack(delete.go:289) andDeleteStack(delete.go:86) each independently refuse protected stacks, so the guard holds even if a caller bypasses the router. - The quiesce loop's
RunningAppStacks(manager.go:232) also excludes protected stacks, so the controller never stops its own tunnel/proxy or itself for a backup.
Deploying-guard on remove/delete (H2)
Both RemoveStack (delete.go:309) and DeleteStack (delete.go:106) refuse while the
stack is Deploying, and both also require the stack to be stopped (not
running/starting/restarting). HDD-path deletion is gated by ProtectedHDDPaths
(delete.go:58), which refuses to wipe top-level namespace directories (appdata/,
backups/, media/, etc.), and ParseComposeHDDMounts cleans paths before the prefix
check to block traversal (C10, delete.go:555).
5. Base-infra bring-up: EnsureBaseStack
EnsureBaseStack (internal/stacks/infra.go:27) renders and deploys the routing/access
infrastructure. Properties (verified):
- Single-flight: guarded by
infraMu.TryLock()(infra.go:28). It is fired both at first boot and on everysystem-healthtick (self-heal); a second concurrent invocation returns immediately rather than racing acompose upon the same dir. - Idempotent: each component is skipped when its container is already running
(
containerRunning,infra.go:256), so the healthy-state re-run is a cheap trio ofdocker inspectcalls. - Non-fatal by contract: per-component failures are collected into a joined error for
the caller to log — it must never crash the controller (
infra.go:67).
Deploy order is load-bearing (the composes declare traefik-public as
external: true, so it must exist first):
ensureTraefikNetwork— creates the externaltraefik-publicdocker network if absent, tolerating a create/inspect race (infra.go:211). A failure here is fatal to the run (every stackupwould fail without it).ensureTraefik— preparesdynamic/,certs/, a0600acme.json, renders the traefik config fromcustomer.email+infrastructure.cf_api_token, writes it, andcompose up -d(infra.go:73).wireController— writes the file-provider routeHost(felhom.<domain>) → http://felhom-controller:8080and joins thefelhom-controllercontainer totraefik-public. Both steps are idempotent (route rewritten only when content changed, so the traefik file watcher doesn't reload every tick; network-connect skipped when already attached). A missing domain is a logged no-op, not an error (infra.go:160).ensureCloudflared— only wheninfrastructure.cf_tunnel_tokenis set; a LAN-only node legitimately runs without it, andmonitor.EffectiveProtectedmirrors this condition so such a node doesn't report cloudflared as a perpetually-missing protected container (infra.go:55,monitor/healthcheck.go:252).ensureFileBrowser— preserves an existing compose. Iffilebrowser/docker-compose.ymlalready exists, itcompose up -dwithout regenerating, so the storage mounts thatweb.SyncFileBrowserMountsmanages are kept intact. Only on first provision does it render the initial compose +config.yamlwith no mounts (infra.go:122).
6. Health probes
RunHealthProbes (internal/stacks/healthprobe.go:23) is a scheduler-driven, controller-side
probe (independent of Docker's own healthcheck). It only probes stacks in running/unhealthy
state that declare a healthcheck in metadata and whose interval has elapsed. The interval is
adaptive: a fast 10s retry while unhealthy, otherwise the configured interval (default 5m);
when HealthProbe is nil (just started) it probes immediately (healthprobe.go:41).
Targets are collected under a read-lock, probed concurrently with the lock released
(healthprobe.go:91), and results applied under the write-lock. Check types: tcp
(dial), http (any response = healthy), and api (validate expected status / body
substring) — runSingleCheck, probeTCP, probeHTTP (healthprobe.go:169+). A failed
probe overrides a Docker-running stack to unhealthy; a passing probe clears that
override back to running. refreshStatusLocked re-applies the last probe result so a
status refresh never resurrects a stale healthy state (manager.go:418).
7. How stacks get their templates (catalog git-sync)
The compose templates and metadata come from the app catalog via internal/sync
(Syncer, sync.go:22). On a configurable interval (git.sync_interval, default 15m,
with an initial sync at startup) it shallow-clones/pulls git.repo_url (branch
git.branch) into <data_dir>/catalog-cache via git fetch --depth 1 +
reset --hard origin/<branch> (sync.go:247). It then copies only docker-compose.yml
and .felhom.yml from each templates/<app>/ into stacks_dir/<app>/, hashing content
to skip unchanged files and never overwriting app.yaml (copyTemplates,
sync.go:311; copyIfChanged, sync.go:395).
After a sync that changed anything, it triggers a rescan (ScanStacks) and, for updated
stacks, a post-sync hook that runs InjectMissingFields — auto-generating values for any
new secret/domain/subdomain deploy fields that aren't yet in a deployed app's
app.yaml (sync.go:207, deploy.go:801). TriggerSync (manual) debounces to one run
per 30s. Sync is disabled (manual mode) when git.repo_url is empty (sync.go:79). The
catalog cache also feeds orphan detection (getCatalogTemplateSlugs): a deployed stack
whose slug is no longer in the cache is flagged Orphaned.