docs: centralize controller documentation + top-level index (code-verified, v0.59.0)
New documentation/controller/ subtree (module map + deploy/stack-lifecycle, backup, storage/monitoring/metrics, auth/hub/sync/integrations) grounded in current source; top-level documentation/README.md index across controller/agent/platform/hub/audits; REORG-NOTES with the verification ledger + flagged doc-gaps. Supersedes (keeps) the v0.33 controller planning map. Additive only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,256 @@
|
||||
# Controller: deploy & stack lifecycle
|
||||
|
||||
> Source of truth: felhom-controller `internal/stacks/`, `internal/infra`, `internal/sync` at v0.59.0.
|
||||
|
||||
This document describes how the in-guest controller represents application stacks, deploys
|
||||
them, runs their lifecycle (update/stop/start/restart/remove/delete), and brings up the base
|
||||
infrastructure (traefik / cloudflared / filebrowser) that everything else depends on. Every
|
||||
claim below is verified against the current source.
|
||||
|
||||
## 1. The stack model
|
||||
|
||||
A *stack* is one customer app, represented on disk as a directory under
|
||||
`/opt/docker/stacks/<name>/` (config key `paths.stacks_dir`, default `/opt/docker/stacks`)
|
||||
containing:
|
||||
|
||||
- `docker-compose.yml` — the compose template (synced from the catalog, see §7).
|
||||
- `.felhom.yml` — app metadata (`Metadata`): display name, deploy fields, resource
|
||||
requests/limits, optional config, healthcheck spec, integrations.
|
||||
- `app.yaml` — the per-app deployment record (`AppConfig`), written **only after** the app
|
||||
is first deployed.
|
||||
|
||||
In memory each stack is a `Stack` struct (`internal/stacks/manager.go:62`):
|
||||
|
||||
| Field | Meaning |
|
||||
|---|---|
|
||||
| `State` (`ContainerState`) | aggregated container state |
|
||||
| `Deployed` | has an `app.yaml` with `deployed: true` |
|
||||
| `Deploying` | a `compose up` is in progress (image pull, etc.) |
|
||||
| `Protected` | a system stack that must never be stopped/removed (§4) |
|
||||
| `Orphaned` | deployed but no longer present in the synced catalog |
|
||||
| `DeployError` | last async deploy error, surfaced to the UI |
|
||||
| `HealthProbe` | latest controller-side probe result (§6) |
|
||||
|
||||
`ContainerState` values (`manager.go:23`): `running`, `starting` (`health: starting`),
|
||||
`unhealthy`, `stopped`, `restarting`, `exited`, `paused`, `unknown`, `not_deployed`,
|
||||
`deploying`, `orphaned`. The effective state is derived by `resolveContainerState`
|
||||
(combines Docker `State` + the health hint in `Status`) and then `aggregateState` across
|
||||
all of a stack's containers, with priority `unhealthy > starting > restarting > all-running
|
||||
> stopped` (`manager.go:437`, `:468`).
|
||||
|
||||
`ScanStacks` (`manager.go:250`) discovers stacks by walking `stacks_dir`, reading metadata
|
||||
and `app.yaml`, marking `Protected` via `cfg.IsProtectedStack`, and computing `Orphaned`
|
||||
against the synced catalog cache (`getCatalogTemplateSlugs`, `manager.go:1078`). While an
|
||||
async deploy is in flight it deliberately does **not** overwrite `Deployed`/`AppConfig` —
|
||||
the deploy goroutine owns those (H3 fix, `manager.go:293`).
|
||||
|
||||
### `AppConfig` (`app.yaml`)
|
||||
|
||||
```go
|
||||
type AppConfig struct {
|
||||
Deployed bool `yaml:"deployed"`
|
||||
DeployedAt string `yaml:"deployed_at"`
|
||||
Env map[string]string `yaml:"env"`
|
||||
LockedFields []string `yaml:"locked_fields"`
|
||||
}
|
||||
```
|
||||
(`deploy.go:98`). `Env` holds the resolved deploy-field values; sensitive ones are stored
|
||||
**encrypted** (§3). `LockedFields` lists env vars marked `locked_after_deploy` that
|
||||
`UpdateStackConfig` refuses to change (`deploy.go:434`).
|
||||
|
||||
## 2. Deploy flow
|
||||
|
||||
Entry point: `DeployStack(req)` (`deploy.go:121`), driven by the deploy form ("Telepítés").
|
||||
|
||||
1. **Atomic check-and-set of `Deploying` (H1).** Under the manager lock, the stack must
|
||||
exist, must not already be `Deploying`, and must not already be `Deployed`; then
|
||||
`Deploying` is set true (`deploy.go:122`). This single critical section prevents two
|
||||
concurrent deploys of the same app. Any later validation failure clears the flag via
|
||||
`clearDeploying()`.
|
||||
|
||||
2. **Memory gate.** `system.GetMemoryMB()` vs the app's `mem_request` and the configured
|
||||
`reserved_memory_mb`. A hard block (real used + new request > usable) aborts with a
|
||||
Hungarian error; a soft over-commit case returns a non-fatal warning string
|
||||
(`deploy.go:159`).
|
||||
|
||||
3. **Field resolution.** For each `DeployField`: `domain` is auto-filled from
|
||||
`customer.domain`; `subdomain` is validated (DNS-safe, not reserved, not already in
|
||||
use); `secret` uses the value previewed to the user or generates one; `password` is
|
||||
**required** (never silently generated — the user must know it); `path` must exist on
|
||||
the host. Required-but-empty fails (`deploy.go:206`–`292`).
|
||||
|
||||
4. **`app.yaml` is persisted with `deployed: false` first (CTRL-T2-1).** The in-memory
|
||||
`AppConfig` carries `Deployed: true` (for UX), but a clone with `Deployed: false` is what
|
||||
`SaveAppConfig` writes to disk at this point (`deploy.go:302`–`314`). The durable record
|
||||
is intentionally *not* marked deployed yet.
|
||||
|
||||
5. **In-memory state flips to deployed** (`s.Deployed = true`, `s.AppConfig = appCfg`),
|
||||
so the UI immediately stops showing a stale "Telepítés" button during the pull
|
||||
(`deploy.go:331`).
|
||||
|
||||
6. **Async compose.** `go m.runComposeDeploy(...)` runs `docker compose up -d` off the
|
||||
request path, so the HTTP handler returns at once and the UI polls for progress (an
|
||||
image pull can take 30–60s). The compose command injects the resolved env plus `DOMAIN`
|
||||
(`deploy.go:338`, `composeExecWithEnv` `manager.go:505`).
|
||||
|
||||
### CTRL-T2-1 crash-safety (the key current behaviour)
|
||||
|
||||
The on-disk `app.yaml` records `deployed: true` **only after** `docker compose up -d`
|
||||
actually succeeds. `runComposeDeploy` (`deploy.go:346`):
|
||||
|
||||
- On compose **failure**: reverts in-memory state (`Deployed=false`, `Deploying=false`,
|
||||
records `DeployError`, `AppConfig=nil`) and re-saves the reverted `AppConfig`
|
||||
(still `deployed:false`) to disk (`deploy.go:350`).
|
||||
- On compose **success**: re-runs `SaveAppConfig` with `appCfg.Deployed == true`, flipping
|
||||
the durable record to deployed (`deploy.go:371`). If *that* save fails, it reverts in
|
||||
memory so the customer can cleanly redeploy rather than be stuck half-recorded
|
||||
(`deploy.go:375`).
|
||||
- Only then is `Deploying` cleared and `RefreshStatus()` run.
|
||||
|
||||
Why it matters: a crash or power loss **during the image-pull window** leaves `app.yaml`
|
||||
at `deployed: false`. On restart, `ScanStacks` therefore sees the app as *not deployed*
|
||||
(no ghost-deployed stack with no containers), and `DeployStack` will cleanly redeploy
|
||||
instead of refusing with "already deployed".
|
||||
|
||||
The UI's three-step progress (form submit → "pulling/deploying" via the `deploying`
|
||||
state and `DeployError` → settled `running`/`unhealthy`) is driven entirely by this
|
||||
in-memory state, which the front-end polls.
|
||||
|
||||
## 3. `SaveAppConfig` and secret encryption (fail-closed, H10)
|
||||
|
||||
`SaveAppConfig(stackDir, cfg, encKey, sensitiveVars)` (`deploy.go:669`) clones the env and,
|
||||
for each var named in `sensitiveVars` (the `secret`/`password` deploy fields, computed by
|
||||
`SensitiveEnvVars`, `deploy.go:741`), encrypts the value with the AES-256 key
|
||||
(`crypto.Encrypt`) unless it is empty or already encrypted.
|
||||
|
||||
**Fail-closed (H10, v0.59.0):** if encryption of a sensitive value errors, `SaveAppConfig`
|
||||
**returns the error and writes nothing** — it never falls through to a plaintext write
|
||||
(`deploy.go:683`–`691`). Earlier code logged a WARN and persisted plaintext, which leaked
|
||||
the secret to disk; that path is gone. Callers (`DeployStack`, `UpdateStackConfig`, etc.)
|
||||
already propagate this error, so the deploy fails cleanly with no secret on disk.
|
||||
|
||||
The write itself is atomic: write to `app.yaml.tmp` then `rename` (H04, `deploy.go:711`),
|
||||
mode `0600`. Decryption for compose happens lazily via `LoadAppConfigDecrypted` /
|
||||
`stackEnv` (`deploy.go:727`, `manager.go:809`), which injects `DOMAIN` plus the decrypted
|
||||
env into the compose process environment. A one-time startup `MigrateEncryption`
|
||||
(`manager.go:128`) re-saves any deployed app still holding plaintext sensitive values.
|
||||
|
||||
## 4. Lifecycle: update / stop / start / restart / remove / delete
|
||||
|
||||
All operations resolve the stack dir, build env via `stackEnv` (decrypted), and shell out
|
||||
to compose. Notably **start/restart use `up -d`, not bare `restart`**, so env changes and
|
||||
template changes (new images, healthchecks) are picked up (`manager.go:709`, `:643`).
|
||||
|
||||
- **Start** (`StartStack`, `manager.go:643`): `compose up -d`; clears the stale health
|
||||
probe so the next probe runs fresh.
|
||||
- **Stop** (`StopStack`, `manager.go:682`): refuses protected stacks, then `compose down`.
|
||||
- **Restart** (`RestartStack`, `manager.go:709`): `compose up -d`; clears health probe.
|
||||
- **Update** (`UpdateStack`, `manager.go:746`): `compose pull` then
|
||||
`compose up -d --remove-orphans`.
|
||||
- **Config update** (`UpdateStackConfig`, `deploy.go:405`): rejects locked fields, re-saves
|
||||
`app.yaml` (encrypting secrets), then `up -d` with the decrypted env.
|
||||
- **Remove** (`RemoveStack`, `delete.go:283`): for a *deployed* (non-orphaned) stack —
|
||||
`compose down --volumes` (keeps images for redeploy), optional HDD-data and backup-path
|
||||
cleanup, then deletes **only** `app.yaml`, reverting the stack to "not deployed". Compose
|
||||
template files are preserved so the user can redeploy.
|
||||
- **Delete** (`DeleteStack`, `delete.go:80`): for an *orphaned* stack only —
|
||||
`compose down --rmi local --volumes`, optional HDD-data removal, then removes the whole
|
||||
stack directory.
|
||||
|
||||
### Protected-stack enforcement (server-side, defence in depth)
|
||||
|
||||
The protected set is configured under `stacks.protected` (default
|
||||
`traefik`, `cloudflared`, `felhom-controller`, `filebrowser` — see
|
||||
`configs/controller.yaml.example:55` and the setup writer `internal/setup/handlers.go:476`).
|
||||
`config.IsProtectedStack` matches case-insensitively (`config.go:347`). Enforcement exists
|
||||
at three layers:
|
||||
|
||||
- **Router/API:** `actionStack` blocks any action other than `restart` on a protected
|
||||
stack with HTTP 403 (`internal/api/router.go:412`).
|
||||
- **Manager:** `StopStack` (`manager.go:683`), `RemoveStack` (`delete.go:289`) and
|
||||
`DeleteStack` (`delete.go:86`) each independently refuse protected stacks, so the guard
|
||||
holds even if a caller bypasses the router.
|
||||
- The quiesce loop's `RunningAppStacks` (`manager.go:232`) also excludes protected stacks,
|
||||
so the controller never stops its own tunnel/proxy or itself for a backup.
|
||||
|
||||
### Deploying-guard on remove/delete (H2)
|
||||
|
||||
Both `RemoveStack` (`delete.go:309`) and `DeleteStack` (`delete.go:106`) refuse while the
|
||||
stack is `Deploying`, and both also require the stack to be stopped (not
|
||||
running/starting/restarting). HDD-path deletion is gated by `ProtectedHDDPaths`
|
||||
(`delete.go:58`), which refuses to wipe top-level namespace directories (`appdata/`,
|
||||
`backups/`, `media/`, etc.), and `ParseComposeHDDMounts` cleans paths before the prefix
|
||||
check to block traversal (C10, `delete.go:555`).
|
||||
|
||||
## 5. Base-infra bring-up: `EnsureBaseStack`
|
||||
|
||||
`EnsureBaseStack` (`internal/stacks/infra.go:27`) renders and deploys the routing/access
|
||||
infrastructure. Properties (verified):
|
||||
|
||||
- **Single-flight:** guarded by `infraMu.TryLock()` (`infra.go:28`). It is fired both at
|
||||
first boot and on every `system-health` tick (self-heal); a second concurrent invocation
|
||||
returns immediately rather than racing a `compose up` on the same dir.
|
||||
- **Idempotent:** each component is skipped when its container is already running
|
||||
(`containerRunning`, `infra.go:256`), so the healthy-state re-run is a cheap trio of
|
||||
`docker inspect` calls.
|
||||
- **Non-fatal by contract:** per-component failures are collected into a joined error for
|
||||
the caller to *log* — it must never crash the controller (`infra.go:67`).
|
||||
|
||||
Deploy order is load-bearing (the composes declare `traefik-public` as
|
||||
`external: true`, so it must exist first):
|
||||
|
||||
1. `ensureTraefikNetwork` — creates the external `traefik-public` docker network if absent,
|
||||
tolerating a create/inspect race (`infra.go:211`). A failure here is fatal to the run
|
||||
(every stack `up` would fail without it).
|
||||
2. `ensureTraefik` — prepares `dynamic/`, `certs/`, a `0600` `acme.json`, renders the
|
||||
traefik config from `customer.email` + `infrastructure.cf_api_token`, writes it, and
|
||||
`compose up -d` (`infra.go:73`).
|
||||
3. `wireController` — writes the file-provider route `Host(felhom.<domain>) →
|
||||
http://felhom-controller:8080` and joins the `felhom-controller` container to
|
||||
`traefik-public`. Both steps are idempotent (route rewritten only when content changed,
|
||||
so the traefik file watcher doesn't reload every tick; network-connect skipped when
|
||||
already attached). A missing domain is a logged no-op, not an error (`infra.go:160`).
|
||||
4. `ensureCloudflared` — **only when `infrastructure.cf_tunnel_token` is set**; a LAN-only
|
||||
node legitimately runs without it, and `monitor.EffectiveProtected` mirrors this
|
||||
condition so such a node doesn't report cloudflared as a perpetually-missing protected
|
||||
container (`infra.go:55`, `monitor/healthcheck.go:252`).
|
||||
5. `ensureFileBrowser` — **preserves an existing compose.** If
|
||||
`filebrowser/docker-compose.yml` already exists, it `compose up -d` *without
|
||||
regenerating*, so the storage mounts that `web.SyncFileBrowserMounts` manages are kept
|
||||
intact. Only on first provision does it render the initial compose + `config.yaml` with
|
||||
no mounts (`infra.go:122`).
|
||||
|
||||
## 6. Health probes
|
||||
|
||||
`RunHealthProbes` (`internal/stacks/healthprobe.go:23`) is a scheduler-driven, controller-side
|
||||
probe (independent of Docker's own healthcheck). It only probes stacks in `running`/`unhealthy`
|
||||
state that declare a `healthcheck` in metadata and whose interval has elapsed. The interval is
|
||||
adaptive: a fast **10s** retry while unhealthy, otherwise the configured interval (default 5m);
|
||||
when `HealthProbe` is nil (just started) it probes immediately (`healthprobe.go:41`).
|
||||
|
||||
Targets are collected under a read-lock, probed **concurrently** with the lock released
|
||||
(`healthprobe.go:91`), and results applied under the write-lock. Check types: `tcp`
|
||||
(dial), `http` (any response = healthy), and `api` (validate expected status / body
|
||||
substring) — `runSingleCheck`, `probeTCP`, `probeHTTP` (`healthprobe.go:169`+). A failed
|
||||
probe overrides a Docker-`running` stack to `unhealthy`; a passing probe clears that
|
||||
override back to `running`. `refreshStatusLocked` re-applies the last probe result so a
|
||||
status refresh never resurrects a stale healthy state (`manager.go:418`).
|
||||
|
||||
## 7. How stacks get their templates (catalog git-sync)
|
||||
|
||||
The compose templates and metadata come from the app catalog via `internal/sync`
|
||||
(`Syncer`, `sync.go:22`). On a configurable interval (`git.sync_interval`, default 15m,
|
||||
with an initial sync at startup) it shallow-clones/pulls `git.repo_url` (branch
|
||||
`git.branch`) into `<data_dir>/catalog-cache` via `git fetch --depth 1` +
|
||||
`reset --hard origin/<branch>` (`sync.go:247`). It then copies **only** `docker-compose.yml`
|
||||
and `.felhom.yml` from each `templates/<app>/` into `stacks_dir/<app>/`, hashing content
|
||||
to skip unchanged files and **never** overwriting `app.yaml` (`copyTemplates`,
|
||||
`sync.go:311`; `copyIfChanged`, `sync.go:395`).
|
||||
|
||||
After a sync that changed anything, it triggers a rescan (`ScanStacks`) and, for *updated*
|
||||
stacks, a post-sync hook that runs `InjectMissingFields` — auto-generating values for any
|
||||
new `secret`/`domain`/`subdomain` deploy fields that aren't yet in a deployed app's
|
||||
`app.yaml` (`sync.go:207`, `deploy.go:801`). `TriggerSync` (manual) debounces to one run
|
||||
per 30s. Sync is disabled (manual mode) when `git.repo_url` is empty (`sync.go:79`). The
|
||||
catalog cache also feeds orphan detection (`getCatalogTemplateSlugs`): a deployed stack
|
||||
whose slug is no longer in the cache is flagged `Orphaned`.
|
||||
Reference in New Issue
Block a user