09-update-architecture.md gains the fourth dated operator ruling (2026-09-06, Option 1) and its section 5 is rewritten from a proposed shape into the shipped one: the pin, the stored definition, the render table, the four writers, the startup ordering, and the trap this slice set for slice 2 - the live compose file is now the frozen one, so a badge comparing against it would answer Naprakesz on exactly the apps that are behind. 02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being true today, so it is corrected, and the two sections describing the old seam now carry a banner saying they describe v0.234.0 and below - kept because every box under v0.235.0 still behaves that way and because they are the measured account of why it changed. R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458 opened for the .felhom.yml asymmetry, with what would settle it by measurement. Live evidence: two real catalog pushes travelling the real 15-minute cycle, both reverted, the tree byte-identical afterwards. The restart that used to take 18.3 seconds and pull a new image now takes 0.1 seconds and pulls nothing.
41 KiB
Felhom Controller Architecture — Part 2: Controller Module Map
How to read this document. Two kinds of statement appear, and where this document marks them it marks them like this — the same wording as
07-backup-architecture.md:11-17, carried here on 2026-08-22 (R-376) so a reader meets one convention and not eight:
- [DESIGN] — a decision taken. Not derived from code; the code may not implement it yet.
- [FACT] — an observed property, carrying a
file:line, a live command output or a citation.Statements in this document are NOT yet all marked. Marking them wholesale is a large judgement exercise and a wrong mark is worse than none, so only what a session touches is marked (R-376). An unmarked statement therefore means "not yet classified", never "observed". That ambiguity is exactly what cost this project three sessions in August 2026: the hot/bulk placement decision sat unmarked beside a marked
[FACT], and was read as an observation and reported as a defect.
EXECUTED (slice 8C, 2026-06-10 — controller v0.37.0). This map's target state is now realized: the disk-execution subsystem (
storage/*, restic, cross-drive, drive-restore,disk_layout,local_infra,infra_backup,setup/scanner,monitor/watchdog+pinger, the storage UI) is deleted (~12.3k LOC);backup.Manageris split to app-data only; disk management is rewired to the host agent's local API (web/agent_disk_handlers.go→ agent/disks); and the container is de-privileged (noprivileged,/dev,/etc/fstab, rshared). The in-guest controller is now Docker-only with no disk/Proxmox privileges, as designed. See doc 03 §6/§9.
Status: audit (keep / port / delete / modify / add), grounded in the v0.33 source.
Subject: the v0.33 controller in felhom-controller/controller/ (110 .go files,
~40 K LOC) audited against 01-topology-and-trust.md and
../proxmox-platform.md.
This is a planning map, not the port. No controller code was changed. Source citations use
controller/internal/...:line(a different repo, so links are not clickable). Classifications reflect the target model: the in-guest controller is Docker-only and holds no Proxmox credentials; everything host/disk/Proxmox moves to a new host agent (out of scope here); the controller reaches the agent through a constrained local API.
Classification scheme
KEEP (host-agnostic, ~unchanged) · PORT (survives, needs rework) · DELETE (→agent) (responsibility moves to the host agent) · DELETE (obsolete) (no longer needed) · MODIFY (stays, materially changes) · NEW (no v0.33 equivalent). Risk tags: clean · needs-rework · hazard (entangles a delete-target with a keep/port target).
0. Executive summary
- The app domain is largely intact and portable: stack lifecycle (
stacks/), catalog git-sync (sync/), app-to-app integrations (integrations/),.fabexport/import (appexport/), the scheduler, crypto, asset sync, the hub report/notify channels, and most of the web UI KEEP/PORT cleanly. - The disk/storage/host half deletes wholesale to the agent: all of
storage/,monitor/watchdog.go, the restic/cross-drive/disk-layout/drive-mount parts ofbackup/,report/infra_backup*+infra_pull, and the host-physical parts ofsystem/. - The setup wizard (
setup/) is obsolete — the agent provisions the controller. - The single biggest hazard is
backup/: the keep side (DB dumps, Docker-volume archive, per-app restore — needed byappexport/and the backup UI) and the delete side (restic, cross-drive, drive-mount) are interleaved inside the same files (backup.go,restore.go,paths.go), not cleanly file-separated. Extracting the app-data-backup subset into a clean retained package is the critical refactor. - Intent-vs-reality corrections (vs the task's provisional split):
monitor/pinger.gois already dead (legacy Healthchecks.io, "deprecated… now handled by Hub" permain.go) → DELETE(obsolete), not keep.backup.go/restore.go/paths.godo not split on file boundaries — they split within the file.settings/is not pure app domain — it stores disk/disconnect/decommission state.system/is genuinely mixed-per-function, not per-file.
0a. App state: desired / in-flight / observed — S-1 CONTRACT (2026-08-02, decision D-b, R-166)
This section is a live contract, not migration history — the rest of this document is the v0.33 keep/port/delete inventory. Read this before touching
stacks/,backup/orbootrecon/. Shipped in controller v0.189.0.
An app's state is three different kinds of information, and conflating them is what produced R-157 mechanism B and F-CRIT-1. They are stored differently on purpose.
| Kind | Question it answers | Where it lives | Persisted? |
|---|---|---|---|
| Desired | What did the customer ask for? | app.yaml → desired_state |
Yes, beside the app's other settings |
| In-flight | Is an operation part-way through, and did it finish? | its own marker file under <data_dir> |
Yes, written before the operation and cleared after |
| Observed | Is it running, unhealthy, restarting, is its drive gone? | nowhere | No — rebuilt by looking |
The rule that ties them together: never derive one from another. The defect this replaced did exactly that — it derived desired from observed (zero containers ⇒ "the customer stopped it"), and zero containers is equally what a power cut mid-compose, an interrupted deploy and an interrupted backup leave behind. Two real faults were therefore read as deliberate stops and stranded silently.
Desired — app.yaml, desired_state
Tri-state: "" (unknown) · "running" · "stopped".
- ONE OWNER: the customer's own action. Writers are the
/api/stacks/{name}/{action}switch,DeployStack,UpdateOptionalConfig's redeploy branch, and the.fabimport.StartStackandStopStackare NOT writers — a census found 14 callers of which only 2 are the customer; the rest are quiesce, the backup volume dump, offbox reconstitution, app export/restore, the storage gate, migration and the boot reconciler. Intent recorded in the primitive would make a nightly backup indistinguishable from the customer pressing Stop. - Written BEFORE the act; a failed write REFUSES the act.
- Absent means UNKNOWN — never "running". Every
app.yamlpredating v0.189.0 lacks the field, so consumers must fall back to the pre-v0.189.0 behaviour rather than assume. A running-only backfill converges the unambiguous cases;stoppedis never inferred, from any signal.
In-flight — a marker file, one per owner
Two exist and they are deliberately separate files: quiesce-state.json (the whole-guest backup
window, internal/quiesce) and appstop-state.json (app-data operations that stop an app —
backup.AppStopGuard, covering the volume dump, offbox reconstitution and .fab export). One
file, one writer: sharing would give one record two lifetimes, and one owner clearing the other's
note is a stranded app by a different route.
- Written before the stop; cleared only after a restart that succeeded; a failed restart keeps the marker so the next startup retries.
- A
deferis not the mechanism. A SIGKILL runs no deferred function — established on live hardware by Campaign 8 fault 10, where what brought the stacks back was the marker read at startup. - Recovery runs at startup and completes before the boot reconciler is launched, so an app the marker explains is not also reported as an unexplained boot orphan.
Observed — not persisted, by design
aggregateState walks every container of a stack and any unhealthy or mixed result wins, so a
partly-dead app cannot read as healthy (F-CRIT-1's shape). This requirement is met here and must not
be re-implemented downstream. Nothing about observed state is written to disk: a controller restart
re-observes it within one refresh, whereas persisting it risks carrying a stale verdict across the
very restart that fixed it (the same argument as RestartingSince).
The binding safety rule (verbatim, from decision D-b)
Losing the state store must never cause an app to be deleted, restarted wrongly, or reported healthy when it is not — the worst acceptable outcome is re-running a backup that already ran.
Applied: a lost or corrupt marker means the app is not auto-restarted by that mechanism, which is
the pre-v0.189.0 position, not a new hazard. A lost app.yaml already means the app is not deployed.
Nothing here may make an absent file more dangerous than a present one.
Boot recovery reads desired, and asks before it acts (v0.190.0)
Both boot gates now read desired state. bootrecon.isBootOrphan (the R-52 sweep) and
shouldRecreateOnBoot (the drive-backed recreate gate) answer the same question — did the customer
want this running? — with the same three-way table, absent falling back to the pre-v0.190.0
container count in both. Their agreement is pinned from both sides against one fixture table, because
an import cycle prevents testing them together. R-170 closed the last gate that still guessed.
The sweep observes a SETTLED fleet, not a single early sample. It samples (name, state, container count) every 5 s, calls the fleet settled after 3 identical samples, and sweeps once, at the end. The window ends on settled or a 50 s budget, and the log says which. Two constraints bound it:
bootReconcileSettle + budget + one DefaultRetryDelaymust stay insidedeadAppBootGrace, or a successful recovery stops being silent. This is arithmetic, pinned by a test.- each sample must REFRESH first.
GetStacks()is the Manager's in-memory map, refreshed by the scheduler on its own 10 s cadence; sampling it faster without refreshing lets "settled" mean "the cache did not update". Found by live validation, not review.
A recovery completing after the grace emits a LATE RECOVERY warning naming the apps. The grace is
never widened to make a late recovery look silent.
Nothing is started without asking whether it may be. bootrecon.StartGate is the one question
the sweep asks per candidate, and it is fail-safe: cannot determine ⇒ do not start. Three holders
answer it, and the last two only became reachable once the window widened past T+5 s:
| Holder | Why starting would be wrong |
|---|---|
| the drive-absent gate | compose creates bind sources wherever the mountpoint points — the guest rootfs |
| a quiesce | a running app inside a snapshot meant to be clean-shutdown-consistent |
| an in-flight app-data operation | restarting an app under its own tar |
The rule is not new — the API's startGatedByMissingDrive already refused a customer's start on
an absent drive. The sweep bypassed it by calling Manager.StartStack directly, which is what R-171
closed. Held apps are reported separately from "still down": they are not a fault the sweep failed to
fix, and reporting them as one is a false alarm.
Read this before adding a fourth caller of Manager.StartStack. That method has no gate of its
own; every caller that is not the customer must decide for itself whether the app may run.
1. v0.33 module inventory (package → purpose, key deps)
| Package | Purpose | Key internal deps |
|---|---|---|
cmd/controller/main.go |
Entry point; wires all subsystems; 6 adapters break import cycles; branches into setup mode | imports every package |
api/ |
REST API (router.go) + geo endpoints (geo.go) |
stacks, backup, metrics, notify, selfupdate, sync, system, assets, integrations, cloudflare, config, settings |
appexport/ |
.fab app export/import (config+DB+volumes, AES-256-CTR+scrypt) |
backup (DB dump), (provider iface → stacks) |
assets/ |
Download/cache app assets from Hub API | — (HTTP only) |
backup/ |
DB dumps, Docker-volume archive, restic, cross-drive rsync, per-app restore, drive mount, disk-layout, infra-backup metadata | config, monitor, settings, system, util |
cloudflare/ |
Geo-restriction via Cloudflare WAF (zone/waf/geosync/countries) — enforcement → hub (S4) | settings |
config/ |
controller.yaml schema + load |
— |
crypto/ |
AES-256-GCM for app.yaml secrets | — |
integrations/ |
App-to-app (OnlyOffice→FileBrowser/Nextcloud) via docker exec / config patch | stacks, crypto, settings |
metrics/ |
SQLite time-series: system + container metrics, log scan | system |
monitor/ |
App health (healthcheck,pinger) + storage/USB watchdog |
config, notify, settings, system |
notify/ |
Hub event push (direct, own API key) | settings |
recovery/ |
Generate recovery-info.txt (DR guide) |
— |
report/ |
Build+push hub report; infra-backup payload; recovery pull | backup, config, metrics, monitor, scheduler, settings, stacks, system |
scheduler/ |
Cron/interval jobs, Budapest TZ | — |
selftest/ |
Startup checks (docker/dirs/catalog/hub/restic repos/mountpoint) | backup, config, settings, system |
selfupdate/ |
Self-update: pull image, edit compose, up -d |
config |
settings/ |
settings.json persistent state: storage paths/disconnect/decommission, cross-drive cfg, notif prefs, geo, integration state, DB-validation cache |
— |
setup/ |
First-run wizard (scan drives, hub-restore, manual config) | backup, config, report, settings, web |
stacks/ |
Docker Compose lifecycle, deploy + memory validation, metadata (.felhom.yml), HDD-data delete |
config, crypto, system |
storage/ |
Physical disk scan/format/attach/mount/migrate/fstab/safety | backup, settings, util |
sync/ |
Catalog git-sync (pull templates) | config |
system/ |
Resource info: mem/cpu/load (guest) + temp/disk-model/USB/mount topology (host) | — |
util/ |
String helper | — |
web/ |
Hungarian dashboard: pages, auth, deploy, backup UI, storage/disk UI, DR restore UI, export UI, debug | appexport, backup, config, crypto, integrations, monitor, notify, scheduler, selfupdate, settings, stacks, storage, system |
2. Classification table (per package/file)
cmd/
| File | Class | Reason | Risk |
|---|---|---|---|
cmd/controller/main.go |
MODIFY | Wiring stays, but drop the setup-mode branch, the storage/watchdog/drive-migrator/restic/cross-drive/infra-backup wiring, and add the agent local-API client. 6 adapters shrink. | hazard |
api/
| File | Class | Reason | Risk |
|---|---|---|---|
api/router.go |
PORT/MODIFY | Keep stacks/deploy/integrations/metrics/sync/assets/selfupdate routes; remove /api/storage/* (disk); backup routes become agent-coordinated guest-backup requests; config/apply (hub-pushes-yaml) changes since the agent now injects config at provision. |
needs-rework |
api/geo.go |
PORT/MODIFY | Keep the customer-facing geo preference endpoints (set/get global + per-app); drop the Cloudflare-sync trigger — enforcement → hub (S4). The controller reports geo desired-state up instead of calling the CF API. | needs-rework |
appexport/ — KEEP/PORT (Docker-volume + DB level, no disk ops)
| File | Class | Reason | Risk |
|---|---|---|---|
crypto.go |
KEEP | Self-contained AES-256-CTR+HMAC+scrypt for .fab. |
clean |
manifest.go, provider.go |
KEEP | Bundle metadata; provider interface (impl in main). | clean |
export.go |
PORT | Docker-volume tar, DB dump via backup.DumpOne, config copy. Depends on the retained app-data-backup subset of backup/; HDD-mount enumeration reworked to per-volume placement. |
needs-rework |
restore.go |
PORT | docker volume create/tar xf, DB import, compose up. Same per-volume rework. |
needs-rework |
estimate.go |
PORT | du/df on mounts → per-volume sizing. |
clean |
assets/
| File | Class | Reason | Risk |
|---|---|---|---|
syncer.go |
KEEP | Hub API download + checksum cache; already a direct hub channel. | clean |
backup/ — THE SPLIT (delete side interleaved with keep side; see §3)
| File | Class | Reason | Risk |
|---|---|---|---|
dbdump.go |
KEEP | Pure docker exec pg_dump/mariadb-dump — app/DB data layer; the retained per-app backup. |
clean |
appdata.go |
PORT | App-data discovery (stacks/volumes/DB containers, du). "HDD mount" concept → per-volume. |
needs-rework |
backup.go (1478 L) |
MODIFY (split) | Mixes keep (RunDBDumps, DumpAppVolumes(Safe), app restore) with delete→agent (RunBackup/backupDrive/restic snapshot/prune/check on per-drive repos). Must be torn in two. |
hazard |
restore.go (442 L) |
MODIFY (split) | RestoreApp restic path → agent; Docker-volume + Tier-2 rsync restore (app layer) → keep. |
hazard |
restore_app_linux.go/_other.go |
PORT | Per-app restore: compose pull/up, rsync app data, DB-dump restore. App layer; depends on backup location that changes. | needs-rework |
paths.go |
MODIFY (split) | AppDBDumpPath/AppVolumeDumpPath keep; Primary/SecondaryResticRepoPath, InfraBackupDir → agent. |
needs-rework |
restic.go |
DELETE (→agent) | restic repos on drives = infra backup tier; agent does vzdump/PBS. | hazard |
crossdrive.go |
DELETE (→agent) | Tier-2 cross-drive rsync to secondary storage = storage-tier (agent + storage manifest). | hazard |
restore_drives_linux.go/_other.go |
DELETE (→agent) | lsblk/blkid/mount/fstab — pure host disk. |
hazard |
disk_layout.go |
DELETE (→agent) | Disk topology for DR → agent. | clean |
local_infra.go |
DELETE (→agent) | Per-drive infra-backup metadata → agent. | clean |
restore_scan.go |
DELETE (→agent) | Scans drives to build a DR restore plan = agent-tier DR. | needs-rework |
cloudflare/ — DELETE (→hub): CF-API enforcement moves to the hub (S4)
| File | Class | Reason | Risk |
|---|---|---|---|
client.go,zone.go,waf.go,geosync.go,countries.go |
DELETE (→hub) | The hub holds the CF API token and reconciles geo desired-state → WAF (doc 01 §5, doc 03 §2). The controller no longer calls the Cloudflare API — it reports geo desired-state up. The customer-facing geo preference UI/data stays (see api/geo.go). |
needs-rework |
config/, crypto/, util/
| File | Class | Reason | Risk |
|---|---|---|---|
config/config.go |
MODIFY | Drop BackupConfig (restic/retention), storage-drive keys, and InfrastructureConfig.cf_api_token (→hub, S4); keep customer/paths/web/git/stacks/monitoring/hub/assets/system; add agent local-API endpoint+token. |
needs-rework |
crypto/crypto.go |
KEEP | App.yaml secret encryption. | clean |
util/strings.go |
KEEP | Trivial helper. | clean |
integrations/ — all KEEP (pure app-domain)
| File | Class | Reason | Risk |
|---|---|---|---|
integrations.go,lifecycle.go,manager.go,onlyoffice_filebrowser.go,onlyoffice_nextcloud.go |
KEEP | App-to-app via docker exec / compose-config patch; no host ops. |
clean |
metrics/
| File | Class | Reason | Risk |
|---|---|---|---|
store.go,logscanner.go,telemetry.go,types.go |
KEEP | SQLite store, docker logs scan, container telemetry — app-domain. |
clean |
collector.go |
PORT | Container metrics (docker stats) keep; host metrics via system.GetInfo (temp, physical disk) become agent-provided or dropped. |
needs-rework |
sysinfo.go/sysinfo_other.go |
MODIFY | Reads /host/etc, /proc/cpuinfo, uptime — host static info; in-guest some is meaningful, hardware identity via agent. |
needs-rework |
monitor/
| File | Class | Reason | Risk |
|---|---|---|---|
healthcheck.go |
PORT (split) | Keep guest health (mem/cpu/docker/protected-containers); host health (temp, physical disk, storage-path mount status) becomes agent-fed. | needs-rework |
pinger.go |
DELETE (obsolete) | Legacy Healthchecks.io; main.go itself marks it "deprecated… now handled by Hub". (Corrects the task's KEEP/PORT guess.) |
clean |
watchdog.go (902 L) |
DELETE (→agent) | Storage/USB disconnect monitoring: umount -l, mount -T /host-fstab, UUID probing, restic-lock cleanup — pure host storage. |
hazard |
notify/, recovery/, scheduler/, selftest/
| File | Class | Reason | Risk |
|---|---|---|---|
notify/notifier.go |
KEEP/MODIFY | Direct hub event channel (own API key) — keep; prune infra event types that move to the agent (storage_disconnected, crossdrive_*, disaster_recovery_*). |
clean |
recovery/info.go |
DELETE (obsolete) | Generates a DR text guide (OS install, docker-setup.sh, hub restore UI); DR is now agent+hub provisioning. | clean |
scheduler/scheduler.go |
KEEP | Generic cron/interval, Budapest TZ. | clean |
selftest/selftest.go |
PORT | Keep docker/dirs/catalog/hub checks; drop restic-repo + system-data mountpoint checks (→agent). | needs-rework |
report/
| File | Class | Reason | Risk |
|---|---|---|---|
pusher.go |
KEEP | Direct hub push (/api/v1/report, Bearer). |
clean |
telemetry.go |
KEEP | Per-app telemetry section. | clean |
builder.go (326 L) |
MODIFY | Keep containers/telemetry/stacks/geo/app-health; drop/relocate host system info, physical storage, restic backup status incl. restic password. | hazard |
types.go |
MODIFY | Schema: drop infra fields (restic password, physical storage), keep app-domain. |
needs-rework |
infra_backup.go/_linux.go/_other.go |
DELETE (→agent) | Builds infra-backup payload (disk layout, restic/enc passwords) for hub. | hazard |
infra_pull.go |
DELETE (→agent) | Pulls recovery config + infra backup from hub (setup-wizard DR). | needs-rework |
selfupdate/ — controller is agent-managed (doc 03 §11)
| File | Class | Reason | Risk |
|---|---|---|---|
version.go |
KEEP | Semver parse / version string (still used for reporting). | clean |
state.go |
DELETE (obsolete) | Self-update audit state — the agent owns controller updates now (doc 03 §11). | clean |
updater.go |
DELETE (→agent) | Resolved (doc 03 §11): the controller is agent-managed — the agent snapshots → redeploys → health-gates → rolls back the controller. The controller's old self-update path (image pull + compose edit) is removed. | clean |
settings/
| File | Class | Reason | Risk |
|---|---|---|---|
settings/settings.go (1101 L) |
MODIFY (split) | Keep notif prefs, integration state, geo, DB-validation cache, cross-drive intent. The storage-path registry (StoragePath with Disconnected/DisconnectedAt/StoppedStacks/decommission) is disk-management state → reshape to per-volume placement fed by the agent's storage manifest; disconnect/decommission/migrate state leaves. (UUID is not a persisted field — runtime-derived from fstab.) |
hazard |
setup/ — all DELETE (obsolete); the agent provisions the controller
| File | Class | Reason | Risk |
|---|---|---|---|
handlers.go,setup.go,csrf.go,network.go |
DELETE (obsolete) | First-run wizard (hub-restore, manual config, LAN-IP detection). | needs-rework |
scanner.go |
DELETE (→agent) | Drive scan (lsblk+temp mounts) for backup discovery — host op; its capability informs the agent. |
clean |
stacks/ — core app domain (KEEP/PORT)
| File | Class | Reason | Risk |
|---|---|---|---|
manager.go (1074 L) |
KEEP/PORT | Docker Compose orchestration, scan/state/start/stop/logs — the heart. Minor port. | clean |
deploy.go |
PORT | Memory validation (system.GetMemoryMB — guest mem, fine in LXC), secret gen, encrypted app.yaml. Add snapshot-before-deploy → agent hook. |
needs-rework |
healthprobe.go |
KEEP | TCP/HTTP app probes. | clean |
metadata.go |
PORT | .felhom.yml parse. Add per-volume hot/bulk classification (doc 01 §8). |
needs-rework |
delete.go |
PORT | Stack delete + HDD-data os.RemoveAll on bind mounts → per-volume cleanup. |
needs-rework |
storage/ — entire package DELETE (→agent)
| File | Class | Reason | Risk |
|---|---|---|---|
scan*,format*,attach*,migrate*,migrate_drive*,safety* |
DELETE (→agent) | Physical disk: lsblk/sfdisk/wipefs/mkfs.ext4/partprobe/mount/umount/fstab/blkid/drive-rsync. The agent owns all of this (doc 01 §3, §8). |
hazard |
sync/
| File | Class | Reason | Risk |
|---|---|---|---|
sync/sync.go |
KEEP | Catalog git-sync (clone/fetch/reset). Since controller v0.235.0 it RENDERS docker-compose.yml rather than copying it — verbatim while the catalog still offers the app's pinned version, from the app's stored applied-compose.yml once the catalog moves past it. .felhom.yml is still copied verbatim always, and app.yaml is still never touched. Reasoning: 09-update-architecture.md §5. |
clean |
system/ — split per-function (not per-file)
| File | Class | Reason | Risk |
|---|---|---|---|
cpu_linux.go/cpu_other.go |
KEEP | /proc/stat works inside an LXC. |
clean |
info.go/info_other.go |
KEEP | Structs/stubs. | clean |
info_linux.go |
MODIFY (split) | Keep mem (/proc/meminfo)/load/statfs (guest); temp via /host/sys, hwmon → agent. |
needs-rework |
mounts_linux.go/mounts_other.go |
DELETE (→agent) mostly | Mount-point detection, USB, disk model, fstab, probe — host/disk. Guest-meaningful statfs disk-usage is the only keep-candidate → fold into the kept info. |
hazard |
web/ — split by UI surface
| File | Class | Reason | Risk |
|---|---|---|---|
auth.go,csrf.go,logbuffer.go,embed.go,templates.go |
KEEP | Session/CSRF, log ring buffer, embeds/logo. | clean |
funcmap.go |
KEEP/PORT | Template helpers; a few backup/state labels track the backup rework. | clean |
server.go (559 L) |
MODIFY | Routing/wiring; remove storage/DR-restore/watchdog wiring; keep app/deploy/backup/settings/export/debug. | needs-rework |
handlers.go (1883 L) |
PORT/MODIFY | Core pages keep; the embedded storage-path management (add/remove/label/schedulable, storage bars, FileBrowser mount sync) → per-volume / agent-fed. | hazard |
handler_export.go |
KEEP/PORT | .fab UI. |
clean |
handler_debug.go (823 L) |
PORT | Drop storage-simulate/infra-push/DR debug; keep the rest. | needs-rework |
alerts.go |
PORT/MODIFY | Storage-disconnect alert now sourced from agent status; backup/update alerts keep. | needs-rework |
handler_restore.go |
DELETE (→agent) / MODIFY | DR restore-mode UI; DR is agent-tier — replace with an agent-status view or remove. | needs-rework |
storage_handlers.go (1600 L) |
DELETE (→agent) | Format/attach/mount/disconnect/migrate-drive/decommission disk UI. Any survivor is a thin client calling the agent API (e.g. per-volume placement requests). | hazard |
templates/ (HTML, non-Go) |
PORT | Remove disk-wizard + DR pages; keep app/deploy/backup/settings pages. | needs-rework |
scripts/
| File | Class | Reason | Risk |
|---|---|---|---|
scripts/hashpass.go |
KEEP | Standalone bcrypt helper. | clean |
3. Coupling hazards (delete-targets depended on by keep/port)
-
backup/is half-deleted but split inside files, not across them.backup.gocontains bothRunDBDumps/DumpAppVolumesSafe/app-restore (keep) andRunBackup/backupDrive+ restic (delete→agent);restore.goandpaths.goare likewise mixed. Keep/port consumers reach into this same package:appexport/export.go:295→backup.DiscoverDatabases/DumpOne(DB dump is app-layer — must survive)report/builder.go:buildBackupReport→ backup status (MODIFY)web/handlers.go(backups page,buildAppBackupRows),web/funcmap.go,web/alerts.go,web/handler_restore.go,web/handler_debug.goselftest/selftest.go:217→checkResticRepos(restic path — delete)main.goscheduler chainRunFullBackup(DB→volume→restic→infra-push) interleaves both sides. Action: extract the app-data-backup subset (DB dump, volume archive, per-app restore) into a clean retained package before deleting the restic/cross-drive code, or every keep consumer breaks.
-
backup/crossdrive.go(delete→agent) is wired ascrossDriveRunnerintomain.go,api/router.go,web/server.go, and surfaced byreport/builder.goand the backups page. Removing it requires reworking the backup UI/report to the agent's guest-backup status. -
storage/(delete→agent) depended on by keep/port UI:web/storage_handlers.go(delete) andweb/server.go/web/handlers.go(port) — the latter renders storage labels/bars and runs FileBrowser mount sync off the storage-path registry.storage/migrate*.goalso importsbackup(also being split). Untangle the per-volume placement UI from the disk-management UI. -
monitor/watchdog.go(delete→agent) depended on byweb/alerts.go(port),web/server.go,web/handler_debug.go,main.go. The disconnect alert must instead consume agent-reported storage status. -
system/mixed-per-function, consumed by both sides. Keep consumers —stacks/deploy.go(GetMemoryMB, guest),metrics/collector.go(container) — must not drag in the host-disk/temp/USB code that goes to the agent (mounts_linux.go,info_linux.gotemp). Also consumed byreport/builder.go(MODIFY),monitor/healthcheck.go(PORT),selftest,crossdrive(delete). Splitsystem/cleanly into guest-info vs host-info first. -
settings/StoragePathcarries disk state into an app-domain store. Disk fields (Disconnected,DisconnectedAt,StoppedStacks, decommission — UUID is not persisted, it's runtime-derived from fstab viasystem.ParseFstabUUID/watchdog.go) are written bywatchdog.go/storage_handlers.go/crossdrive.go(all delete) but the same struct is read bystacks/webfor labels and placement (keep). ReshapeStoragePathto a placement record fed by the agent manifest. -
report/builder.goimports almost everything (backup, monitor, scheduler, stacks, system, metrics, settings, config). Its MODIFY must land after the backup and system splits, or it pulls deleted code along. -
backup/paths.goshared both ways —appexport+selftest+ the kept DB-dump flow use the app-dump path helpers; the same file holds the restic/secondary helpers that leave. -
DR/provisioning chain is cross-cut:
setup/(obsolete) →report/infra_pull+recovery/info+backup.MountDrivesFromLayout+backup.ReadLocalInfraBackup. All obsolete/→agent, butmain.go's setup branch andweb/handler_restore.goreference them; remove together.
4. Moves to the host agent (consolidated — feeds the future agent design)
Reporting only; not designing the agent here.
- All physical-disk management —
storage/in full: scan/classify, format (wipefs/sfdisk/mkfs.ext4/partprobe), attach (raw mount + bind + fstab), per-app and full-drive migration (rsync), safety checks (system-disk detection). - Storage/USB watchdog —
monitor/watchdog.go: disconnect/reconnect detection,umount -l,mount -T /host-fstab, UUID-by-id probing, safe-disconnect, restic-lock cleanup. - Infra/disk backup tier —
backup/restic.go,crossdrive.go,restore_drives_*,disk_layout.go,local_infra.go,restore_scan.go, plus the restic-snapshot half ofbackup.go, the restic-restore half ofrestore.go, and the restic/secondary path helpers inpaths.go. (Maps to the agent'svzdump→tiers→PBS in doc 01 §8.) - Infra-backup payload + recovery pull —
report/infra_backup*,report/infra_pull. - Host-physical telemetry —
system/mounts_linux.go(mount topology, USB, disk model), the temp/hwmon parts ofsystem/info_linux.go, and the host-hardware parts ofmetrics/sysinfo.go. - Drive scanning for provisioning/DR —
setup/scanner.go. - Self-restore-test execution — the agent performs the restore-to-scratch-guest; the controller only orchestrates/validates (see §5).
5. New components to build (no v0.33 equivalent)
- Agent local-API client — the controller's only path to guest-level Proxmox
operations (doc 01 §3, §5):
snapshot-before-deploy+ rollback, "grow my RAM", request guest backup/restore, read the storage manifest / mount placement, query per-target storage status. Replaces the deleted direct host/disk code with constrained RPC. The controller holds no Proxmox creds — only a local-API token. - Per-volume storage placement (doc 01 §8) —
.felhom.ymlhot/bulkvolume classification (extendstacks/metadata.go), enforcement at deploy (extendstacks/deploy.go), and a placement record insettings. Replaces the per-app HDD-path + cross-drive model. Abulkvolume must be realized as abackup=0mount point, never a rootfs Docker named volume (validated recipe:phase3-findings.mdB2 / doc 03 §7). - Self-restore-test status display (read-only) — the agent owns orchestration (it
holds the PBS key and creates the scratch guest — operator-tier, doc 03 §8); the controller
only surfaces
GET /restore-test/statusin its UI. (Round-trip validated: Phase 2, ../proxmox-platform.md §4.) - Snapshot-before-deploy/rollback flow in the deploy path — wraps the existing
compose deploy with agent snapshot → health check → agent rollback-on-failure
(doc 01 §9). New behaviour on top of
stacks/deploy.go+stacks/healthprobe.go. - Agent-provisioning bootstrap receiver — the controller accepts its injected hub API
key + local-API token from the agent at provision time (doc 01 §6), replacing the
deleted
setup/wizard.
6. Open / blocked items
- Geo — resolved (S4): CF-API enforcement moves to the hub (it holds the CF token and
reconciles geo → WAF); the controller keeps the geo preference UI/data and reports
desired-state up. Tunnel placement is settled (host, agent-managed, doc 03 §3/§5). The
cloudflare/package +api/geo.go's CF-sync are DELETE-from-controller → hub. - Self-update — resolved (doc 03 §11): the controller is agent-managed; its self-update path is removed.
settings/stacksper-volume reshape — depends on the storage-manifest contract between hub ↔ agent ↔ controller (doc 01 §8), not yet specified.- Backup UI/report surface — depends on the agent's guest-backup status API shape (what the controller can see about vzdump/PBS state) — undefined.
- Notification event taxonomy — which infra events (
storage_disconnected,crossdrive_*,disaster_recovery_*) the agent emits vs the controller, once those responsibilities move.
Changelog — design-review + Phase-3 fold-in (2026-06-08)
- M1: removed
UUIDfrom thesettings.StoragePathfield lists (§ settings, hazard #6) — it is runtime-derived from fstab, not persisted. - S4 (geo):
cloudflare/reclassified PORT(blocked) → DELETE(→hub) (CF-API enforcement moves to the hub);api/geo.go→ PORT/MODIFY (keep geo preference endpoints, drop the CF-sync trigger);config/config.goalso dropscf_api_token. §6 + §1 updated. - S5: cloudflare/geo no longer "blocked on tunnel placement" (resolved).
- S6: §5(3) self-restore-test → status-display only; the agent owns orchestration.
- Self-update resolved (03 §11):
updater.go→ DELETE(→agent),state.go→ DELETE(obsolete),version.goKEEP; §6 + §5(2) updated (bulk =backup=0mountpoint recipe).
The app-definition seam — what happens to a deployed app when its compose file is rewritten
STATUS: MEASURED BEHAVIOUR, NOT A RECORDED DESIGN. This section describes what the system was observed to do on 2026-09-01 (
audits/SPIKE-app-update-2026-09-01.md), with controls. It is deliberately not marked[DESIGN], because the operator has not ruled on whether the behaviour was intended. Do not read this as an endorsement, and do not spec against it as though it were settled. The open questions are R-438 and R-441.
Until this was measured, no architecture document said what happens here, and the gap itself is R-438. The three facts below are the ones a reader needs before touching any of it.
⚠ SECTIONS 1 AND 2 DESCRIBE THE BEHAVIOUR UP TO CONTROLLER v0.234.0. Controller v0.235.0 (2026-09-06) CHANGED IT, on an operator ruling. They are kept because they are the measured account of why it was changed, and because every box below v0.235.0 still behaves this way. What ships now:
09-update-architecture.md§5. In one sentence — an app's VERSION is frozen to what the customer has and only a deliberate Update moves it, while template CORRECTIONS and the self-healing below still arrive on the 15-minute cycle. Nothing was added to the thirteen call sites in §2; they were made safe by removing the reason.
1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle
Syncer.copyTemplates (sync/sync.go:319) walks every directory in the catalog cache and copies
docker-compose.yml and .felhom.yml into the matching stack folder. There is no test of whether
the app is deployed. The only guard is a sha256 content compare (copyIfChanged) and the only
exclusion is app.yaml. Interval is git.sync_interval, default 15m (config/config.go:351), plus
one immediate sync at controller start (sync.go:98).
SINCE v0.235.0 this walk still happens and .felhom.yml is still copied unconditionally, but the
compose file goes through Syncer.renderSource, which consults a per-app plan supplied by the stack
manager (Manager.RenderPlanFor) through a nil-safe seam. A nil seam is byte-for-byte the behaviour
described above.
It restarts nothing. The post-sync hook is stackMgr.InjectMissingFields(updated) and nothing
else. So from the moment it runs, a deployed app's definition and its running containers disagree,
and they stay that way until something else acts.
2. Every lifecycle path resolves that disagreement, silently, by upgrading
StartStack, RestartStack and UpdateStack (stacks/manager.go:1029 / 1133 / 1170) all end in
docker compose up -d. up -d makes the container match the file, and pulls the image itself if it
is absent — measured at 18.3 s with a pull versus 0.5 s without, against a negative control
(unchanged file) that did not even recreate the container.
RestartStack carries an explicit in-source comment saying this is deliberate — "so that … any
template changes (new images, healthchecks) are picked up". That is a fact about the source, not an
operator ruling, and it speaks only for the customer-pressed restart.
Thirteen non-API call sites across nine files reach up -d without anyone pressing anything. The
full table is §8 of the spike doc; the three that matter most are:
bootrecon.Reconciler.Run→StartStack—bootrecon/bootrecon.go:269backup.AppStopGuard.Recover→StartStack—backup/appstop_marker.go:283- the drive-return gate,
web.Server.restartStacks→StartStack—web/intermediary.go:222
A plain power cut does NOT trigger this. Docker's restart: unless-stopped restores the existing
containers on the old image, the reconciler finds no orphan, and it logs so
(no boot-orphaned apps (nothing to start)). The unattended upgrade needs the narrower precondition
"and the app did not come back" — which was measured, and does upgrade.
3. Nothing takes a copy first, and the tag cannot be put back afterwards
UpdateStack runs compose pull then compose up -d --remove-orphans and does nothing else — no
dump, no copy, no hold. The R-361 safety dump (backup/offbox_reconstitute.go:207) is database-only
and is not on this path at all; an app with no database gets nothing from it even on the paths where it
does run.
And restoring the old tag is not a rollback. Once an app has migrated its data, the old image refuses to start — measured on Nextcloud: "the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported". The data itself survives; only the downgrade is blocked. So the only route back is a data restore from a copy taken before the update.
And that route fights this seam (R-441): stackAdapter.RecreateStackDefinitionFromUnit
(cmd/controller/main.go:2570) writes the recovery unit's captured compose — with the OLD pin — into
the live stack dir, and copyIfChanged overwrites it again on the next tick. The overwrite is
measured; that the restore writes to that path is read, not measured.
What a change here must not break
- The syncer's overwrite is also the repair path: a locally-corrupted compose is replaced by the catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that.
up -d-on-restart is what injectsapp.yamlenv into a running stack. Reverting todocker compose restartwould silently stop doing that.