SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has already moved, and the Restart button does it — 18.3s with a network pull when the target image is absent, 0.5s when present, against a negative control that did not even recreate the container. The boot reconciler does the same thing unattended when an app fails to come back (bootrecon.go:269 -> StartStack). AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades nothing. Docker restores the old containers and the reconciler logs 'no boot-orphaned apps (nothing to start)'. AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run, the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported'. 'Rollback' is the wrong word for this arc and is struck. Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator confirmation. Peti's box was never contacted. No production code written in any repo: felhom-controller is at 960d29b0612c before and after, tree clean, and build/vet/test are green — run at the end to prove exactly that. Register: R-438/439/440 updated with live evidence; R-441..R-445 opened (restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB left behind; update reports success over a broken app; no fleet fstrim; hub telemetry outlives the app). Capability map gains three measured rows. The mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR, not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision. Five claims in the brief are named as wrong, including two of my own method.
This commit is contained in:
@@ -335,7 +335,7 @@ own; every caller that is not the customer must decide for itself whether the ap
|
||||
### `sync/`
|
||||
| File | Class | Reason | Risk |
|
||||
|---|---|---|---|
|
||||
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). | clean |
|
||||
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). **It copies into EVERY stack folder, deployed or not — see "the app-definition seam" below.** | clean |
|
||||
|
||||
### `system/` — split per-function (not per-file)
|
||||
| File | Class | Reason | Risk |
|
||||
@@ -503,3 +503,77 @@ own; every caller that is not the customer must decide for itself whether the ap
|
||||
- **S6:** §5(3) self-restore-test → **status-display only**; the agent owns orchestration.
|
||||
- **Self-update resolved (03 §11):** `updater.go` → **DELETE(→agent)**, `state.go` →
|
||||
DELETE(obsolete), `version.go` KEEP; §6 + §5(2) updated (bulk = `backup=0` mountpoint recipe).
|
||||
|
||||
|
||||
---
|
||||
|
||||
## The app-definition seam — what happens to a deployed app when its compose file is rewritten
|
||||
|
||||
> **STATUS: MEASURED BEHAVIOUR, NOT A RECORDED DESIGN.** This section describes what the system was
|
||||
> observed to do on 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`), with controls. It is
|
||||
> deliberately **not** marked `[DESIGN]`, because the operator has not ruled on whether the behaviour
|
||||
> was intended. **Do not read this as an endorsement, and do not spec against it as though it were
|
||||
> settled.** The open questions are R-438 and R-441.
|
||||
|
||||
Until this was measured, no architecture document said what happens here, and the gap itself is
|
||||
R-438. The three facts below are the ones a reader needs before touching any of it.
|
||||
|
||||
### 1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle
|
||||
|
||||
`Syncer.copyTemplates` (`sync/sync.go:319`) walks every directory in the catalog cache and copies
|
||||
`docker-compose.yml` and `.felhom.yml` into the matching stack folder. **There is no test of whether
|
||||
the app is deployed.** The only guard is a sha256 content compare (`copyIfChanged`) and the only
|
||||
exclusion is `app.yaml`. Interval is `git.sync_interval`, default `15m` (`config/config.go:351`), plus
|
||||
one immediate sync at controller start (`sync.go:98`).
|
||||
|
||||
**It restarts nothing.** The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing
|
||||
else. So from the moment it runs, a deployed app's *definition* and its *running containers* disagree,
|
||||
and they stay that way until something else acts.
|
||||
|
||||
### 2. Every lifecycle path resolves that disagreement, silently, by upgrading
|
||||
|
||||
`StartStack`, `RestartStack` and `UpdateStack` (`stacks/manager.go:1029 / 1133 / 1170`) all end in
|
||||
`docker compose up -d`. `up -d` makes the container match the file, **and pulls the image itself if it
|
||||
is absent** — measured at 18.3 s with a pull versus 0.5 s without, against a negative control
|
||||
(unchanged file) that did not even recreate the container.
|
||||
|
||||
`RestartStack` carries an explicit in-source comment saying this is deliberate — *"so that … any
|
||||
template changes (new images, healthchecks) are picked up"*. **That is a fact about the source, not an
|
||||
operator ruling**, and it speaks only for the customer-pressed restart.
|
||||
|
||||
**Thirteen non-API call sites across nine files** reach `up -d` without anyone pressing anything. The
|
||||
full table is §8 of the spike doc; the three that matter most are:
|
||||
|
||||
- `bootrecon.Reconciler.Run` → `StartStack` — `bootrecon/bootrecon.go:269`
|
||||
- `backup.AppStopGuard.Recover` → `StartStack` — `backup/appstop_marker.go:283`
|
||||
- the drive-return gate, `web.Server.restartStacks` → `StartStack` — `web/intermediary.go:222`
|
||||
|
||||
**A plain power cut does NOT trigger this.** Docker's `restart: unless-stopped` restores the existing
|
||||
containers on the old image, the reconciler finds no orphan, and it logs so
|
||||
(`no boot-orphaned apps (nothing to start)`). The unattended upgrade needs the narrower precondition
|
||||
*"and the app did not come back"* — which was measured, and does upgrade.
|
||||
|
||||
### 3. Nothing takes a copy first, and the tag cannot be put back afterwards
|
||||
|
||||
`UpdateStack` runs `compose pull` then `compose up -d --remove-orphans` and does nothing else — no
|
||||
dump, no copy, no hold. The R-361 safety dump (`backup/offbox_reconstitute.go:207`) is **database-only**
|
||||
and is not on this path at all; an app with no database gets nothing from it even on the paths where it
|
||||
does run.
|
||||
|
||||
**And restoring the old tag is not a rollback.** Once an app has migrated its data, the old image
|
||||
refuses to start — measured on Nextcloud: *"the version of the data (32.0.9.2) is higher than the
|
||||
docker image version (31.0.14.1) and downgrading is not supported"*. The data itself survives; only the
|
||||
downgrade is blocked. So the only route back is a **data** restore from a copy taken before the
|
||||
update.
|
||||
|
||||
**And that route fights this seam** (R-441): `stackAdapter.RecreateStackDefinitionFromUnit`
|
||||
(`cmd/controller/main.go:2570`) writes the recovery unit's captured compose — with the OLD pin — into
|
||||
the live stack dir, and `copyIfChanged` overwrites it again on the next tick. The overwrite is
|
||||
measured; that the restore writes to that path is read, not measured.
|
||||
|
||||
### What a change here must not break
|
||||
|
||||
- The syncer's overwrite is also the **repair** path: a locally-corrupted compose is replaced by the
|
||||
catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that.
|
||||
- `up -d`-on-restart is what injects `app.yaml` env into a running stack. Reverting to
|
||||
`docker compose restart` would silently stop doing that.
|
||||
|
||||
Reference in New Issue
Block a user