# 09 — How an app update works, and what it is becoming > **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.** > Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen > deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who > did not know it was one. **This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The source is the truth.** Nothing here is invented: every mechanism claim is cited either to `audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`. --- ## 1. How an update works today, as measured ### 1.1 The button `Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose commands and nothing else: ``` compose pull → compose up -d --remove-orphans ``` **No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike §4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443. ### 1.2 The catalog syncer moves the file underneath a deployed app `Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies `docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is `app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the sync does not restart anything.** That is why a deployed app's compose file and its running containers can disagree **indefinitely**. Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3). ### 1.3 Thirteen other paths end in `compose up -d` Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`, and every one ends in `compose up -d` against the live compose file (spike §8 — the task that commissioned the spike said five; the count is thirteen). They include the **boot reconciler** (`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every restore path. **So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where a boot reconciliation started an app on a newer image at 17:55:44Z. ### 1.4 One fear is measured SMALLER than it was stated A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs `up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant 1c, a positive observable and not an absent log line). **The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying so is more useful than leaving the scarier version standing. --- ## 2. What was chosen, and by whom `Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the whole arc: > *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template > changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends > SIGTERM+start to existing containers without re-reading the compose file or env."* **So the restart behaviour was chosen, deliberately, and written down. A design decision is not a defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism; it does not change it. --- ## 3. The three operator decisions (2026-09-02) These are rulings, not proposals. Anything specced against a different assumption is wrong. 1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the update path. The guest-snapshot alternative is to be **spiked before anything is designed around it**. Context: the existing safety machinery (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike §6 — the file half was never priced, and demo-hp is too young a box to price it. 2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.** A customer on the newest version is supported however old that version is. This is why `catalog_since` exists and why no version string is shown. 3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human, because §4 says it cannot be undone. --- ## 4. The vocabulary ruling — "rollback" is struck **App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually run, putting the old image tag back produces a container that refuses to start — > *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and > downgrading is not supported"* with a positive control proving the data is intact, only unreachable by the old version (§7 6d). **So "rollback" must not appear in any spec for this arc.** The two shapes actually available are: | shape | when it applies | what it does | |---|---|---| | **ABORT** | before anything migrated | stop, put the old image back, the app runs again | | **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy | There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute half-life because the syncer overwrites it (**R-441**). --- ## 5. The target shape **The live `docker-compose.yml` becomes DERIVED from a pin recorded in `app.yaml`** — the one file the syncer never touches (`sync.go:319`'s exclusion). The catalog then proposes; `app.yaml` decides; the rendered compose file is an output rather than an input, and the thirteen unattended `up -d` paths stop being able to change a version by accident. **Nothing in slices 1 or 2 implements this.** They make the current state *visible*, which is the prerequisite for judging how urgent it is. --- ## 6. The seven slices | # | slice | status | |---|---|---| | **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** | | **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** | | **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 | | **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 | | **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 | | **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 | | **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 | **The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge, never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it breaks. --- ## 7. What slices 1 and 2 actually built ### 7.1 The record (slice 1) `Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`, and writes `app.yaml`: ```yaml installed_images: web: ref: lscr.io/linuxserver/bookstack:26.05.2 digest: sha256:… # "" if the image was never pulled from a registry at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service ``` Three rules, each with its reason: - **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has already moved. A record built from it would answer "what will happen next time something runs `up -d`", which is a different question. - **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`. Intent refused, observation logged. Refusing to start a customer's app because we could not write down which version it is trades a real outage for a bookkeeping gap. - **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a complete record with a partial one. **Nothing reads it to take a decision.** Slice 2 reads it to render a label. ### 7.2 The label (slice 2) `web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against what the current template pins, and renders through the existing `meta_badge` partial — no new markup, no new CSS. **Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by a test with a companion red-proof. **No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`). Version strings stay in the logs, the API and the hub. --- ## 8. Known limitations, stated plainly 1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and queries no registry — a customer's box must not depend on reaching eight upstream registries to render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be identical while the image behind it has moved. **Measured, not theorised:** spike §5 found `mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls holding. Digest-level comparison needs a registry query and is deferred — **R-446**. 2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent commit to diff against, so the gate needs a deeper fetch — **R-452**. 3. **The record only appears after the next lifecycle action.** An app that is running and untouched keeps a legacy `app.yaml` and therefore no badge, until someone restarts, updates or redeploys it. That is correct — the alternative is inventing a record from the file §1.2 says has already moved — but it means the fleet view fills in gradually rather than at upgrade. 4. **The hub does not record image tags at all.** Its report's container payload carries name, state, CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side change; it is not derivable from what is already reported. --- ## 9. Where the rest lives | what | where | |---|---| | the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` | | the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 | | the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` | | what the lifecycle actions are proven to do | `architecture/00-capability-map.md` | | the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |