SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s

THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
This commit is contained in:
2026-09-01 21:35:32 +02:00
parent ac079b8c43
commit 56c7e373a3
43 changed files with 2026 additions and 287 deletions
@@ -335,7 +335,7 @@ own; every caller that is not the customer must decide for itself whether the ap
### `sync/`
| File | Class | Reason | Risk |
|---|---|---|---|
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). | clean |
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). **It copies into EVERY stack folder, deployed or not — see "the app-definition seam" below.** | clean |
### `system/` — split per-function (not per-file)
| File | Class | Reason | Risk |
@@ -503,3 +503,77 @@ own; every caller that is not the customer must decide for itself whether the ap
- **S6:** §5(3) self-restore-test → **status-display only**; the agent owns orchestration.
- **Self-update resolved (03 §11):** `updater.go` → **DELETE(→agent)**, `state.go` →
DELETE(obsolete), `version.go` KEEP; §6 + §5(2) updated (bulk = `backup=0` mountpoint recipe).
---
## The app-definition seam — what happens to a deployed app when its compose file is rewritten
> **STATUS: MEASURED BEHAVIOUR, NOT A RECORDED DESIGN.** This section describes what the system was
> observed to do on 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`), with controls. It is
> deliberately **not** marked `[DESIGN]`, because the operator has not ruled on whether the behaviour
> was intended. **Do not read this as an endorsement, and do not spec against it as though it were
> settled.** The open questions are R-438 and R-441.
Until this was measured, no architecture document said what happens here, and the gap itself is
R-438. The three facts below are the ones a reader needs before touching any of it.
### 1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle
`Syncer.copyTemplates` (`sync/sync.go:319`) walks every directory in the catalog cache and copies
`docker-compose.yml` and `.felhom.yml` into the matching stack folder. **There is no test of whether
the app is deployed.** The only guard is a sha256 content compare (`copyIfChanged`) and the only
exclusion is `app.yaml`. Interval is `git.sync_interval`, default `15m` (`config/config.go:351`), plus
one immediate sync at controller start (`sync.go:98`).
**It restarts nothing.** The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing
else. So from the moment it runs, a deployed app's *definition* and its *running containers* disagree,
and they stay that way until something else acts.
### 2. Every lifecycle path resolves that disagreement, silently, by upgrading
`StartStack`, `RestartStack` and `UpdateStack` (`stacks/manager.go:1029 / 1133 / 1170`) all end in
`docker compose up -d`. `up -d` makes the container match the file, **and pulls the image itself if it
is absent** — measured at 18.3 s with a pull versus 0.5 s without, against a negative control
(unchanged file) that did not even recreate the container.
`RestartStack` carries an explicit in-source comment saying this is deliberate — *"so that … any
template changes (new images, healthchecks) are picked up"*. **That is a fact about the source, not an
operator ruling**, and it speaks only for the customer-pressed restart.
**Thirteen non-API call sites across nine files** reach `up -d` without anyone pressing anything. The
full table is §8 of the spike doc; the three that matter most are:
- `bootrecon.Reconciler.Run` → `StartStack` — `bootrecon/bootrecon.go:269`
- `backup.AppStopGuard.Recover` → `StartStack` — `backup/appstop_marker.go:283`
- the drive-return gate, `web.Server.restartStacks` → `StartStack` — `web/intermediary.go:222`
**A plain power cut does NOT trigger this.** Docker's `restart: unless-stopped` restores the existing
containers on the old image, the reconciler finds no orphan, and it logs so
(`no boot-orphaned apps (nothing to start)`). The unattended upgrade needs the narrower precondition
*"and the app did not come back"* — which was measured, and does upgrade.
### 3. Nothing takes a copy first, and the tag cannot be put back afterwards
`UpdateStack` runs `compose pull` then `compose up -d --remove-orphans` and does nothing else — no
dump, no copy, no hold. The R-361 safety dump (`backup/offbox_reconstitute.go:207`) is **database-only**
and is not on this path at all; an app with no database gets nothing from it even on the paths where it
does run.
**And restoring the old tag is not a rollback.** Once an app has migrated its data, the old image
refuses to start — measured on Nextcloud: *"the version of the data (32.0.9.2) is higher than the
docker image version (31.0.14.1) and downgrading is not supported"*. The data itself survives; only the
downgrade is blocked. So the only route back is a **data** restore from a copy taken before the
update.
**And that route fights this seam** (R-441): `stackAdapter.RecreateStackDefinitionFromUnit`
(`cmd/controller/main.go:2570`) writes the recovery unit's captured compose — with the OLD pin — into
the live stack dir, and `copyIfChanged` overwrites it again on the next tick. The overwrite is
measured; that the restore writes to that path is read, not measured.
### What a change here must not break
- The syncer's overwrite is also the **repair** path: a locally-corrupted compose is replaced by the
catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that.
- `up -d`-on-restart is what injects `app.yaml` env into a running stack. Reverting to
`docker compose restart` would silently stop doing that.