# 09 — How an app update works, and what it is becoming > **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.** > Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen > deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who > did not know it was one. **This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The source is the truth.** Nothing here is invented: every mechanism claim is cited either to `audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`. --- ## 1. How an update works today, as measured ### 1.1 The button `Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose commands and nothing else: ``` compose pull → compose up -d --remove-orphans ``` **No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike §4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443. ### 1.2 The catalog syncer moves the file underneath a deployed app `Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies `docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is `app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the sync does not restart anything.** That is why a deployed app's compose file and its running containers can disagree **indefinitely**. Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3). ### 1.3 Thirteen other paths end in `compose up -d` Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`, and every one ends in `compose up -d` against the live compose file (spike §8 — the task that commissioned the spike said five; the count is thirteen). They include the **boot reconciler** (`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every restore path. **So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where a boot reconciliation started an app on a newer image at 17:55:44Z. ### 1.4 One fear is measured SMALLER than it was stated A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs `up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant 1c, a positive observable and not an absent log line). **The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying so is more useful than leaving the scarier version standing. --- ## 2. What was chosen, and by whom `Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the whole arc: > *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template > changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends > SIGTERM+start to existing containers without re-reading the compose file or env."* **So the restart behaviour was chosen, deliberately, and written down. A design decision is not a defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism; it does not change it. --- ## 3. The operator decisions These are rulings, not proposals. Anything specced against a different assumption is wrong. ### 2026-09-02 1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the update path. The guest-snapshot alternative is to be **spiked before anything is designed around it**. Context: the existing safety machinery (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike §6 — the file half was never priced, and demo-hp is too young a box to price it. 2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.** A customer on the newest version is supported however old that version is. This is why `catalog_since` exists and why no version string is shown. 3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human, because §4 says it cannot be undone. ### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0 4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it — while corrections to its definition keep arriving on the 15-minute cycle exactly as they do today.** **This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in its own comment. Reversing a chosen behaviour is a decision, not a bug fix. **What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".** The old behaviour had two halves and only one of them was unwanted: | half | verdict | |---|---| | a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone | | template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 | So the ruling keeps the second and removes the first. In the operator's own words: *while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update.* **It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup precondition goes. Slice 3 only stops the other twelve paths from doing the update's job. --- ## 4. The vocabulary ruling — "rollback" is struck **App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually run, putting the old image tag back produces a container that refuses to start — > *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and > downgrading is not supported"* with a positive control proving the data is intact, only unreachable by the old version (§7 6d). **So "rollback" must not appear in any spec for this arc.** The two shapes actually available are: | shape | when it applies | what it does | |---|---|---| | **ABORT** | before anything migrated | stop, put the old image back, the app runs again | | **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy | There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute half-life because the syncer overwrites it (**R-441**). --- ## 5. The shape, as SHIPPED in v0.235.0 The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather than an input, and the thirteen unattended `up -d` paths stop being able to change a version by accident. ### 5.1 Nothing was added to the thirteen call sites, and that is deliberate **They are made safe by removing the reason, not by gating them.** The most important of them are REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`), the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a customer's app down, which is worse than the problem this slice solves.** Since the file they act on no longer changes version, every one of them became safe without being touched. ### 5.2 The pin, and what it is not `AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref. **It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a DECISION — what should run. Letting an observation feed a decision would make a bad reading become a bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over. They will normally agree; when they disagree that is a signal, not a bug to paper over. **Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.** Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is safe from the catalog, and keeping it beside the app means it travels with every path that already moves a stack dir. ### 5.3 The four writers — the only acts entitled to move a version | writer | pin source | |---|---| | the deploy path (`runComposeDeploy`) | the template just deployed from | | **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull | | the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** | | `Manager.AdoptPins` | the observation, once, and only when complete AND matching | **`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards would pull the frozen version and change nothing — while reporting success, and a button that lies is worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of `recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent. ### 5.4 The render table, complete | app state | result | |---|---| | not deployed / protected / seam not wired | the catalog template — today's behaviour | | deployed, **unpinned** | the catalog template + one DEBUG | | deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** | | deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE | | pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it | | mid-deploy | the compose file is left alone this cycle | **`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries `catalog_since`, which the badge needs. See §8.5. **The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration the older template cannot supply, so an old image under a new template is broken in a way neither version is. **And this is not "skip deployed apps".** That option was considered and rejected: it also stops health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in the spike §3 — both halves the ruling explicitly kept. ### 5.5 Adoption, and why the startup order matters `AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app to what it is already running. **It reads and writes files only** — no container is started, stopped or touched. It skips, loudly, when the observation is incomplete or when the app runs something the current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a guessed pin. **`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and handed the next restart a version change — the exact behaviour this slice removes, once per boot. ### 5.6 The trap this slice set for the previous one `Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one. On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test still green, because the new field has the same type and shape. The badge now reads `Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an earlier feature is the failure mode to look for whenever a file changes meaning.** --- ## 6. The seven slices | # | slice | status | |---|---|---| | **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** | | **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** | | **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** | | **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 | | **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 | | **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 | | **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 | | **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 | **The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge, never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it breaks. --- ## 7. What slices 1 and 2 actually built ### 7.1 The record (slice 1) `Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`, and writes `app.yaml`: ```yaml installed_images: web: ref: lscr.io/linuxserver/bookstack:26.05.2 digest: sha256:… # "" if the image was never pulled from a registry at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service ``` Three rules, each with its reason: - **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has already moved. A record built from it would answer "what will happen next time something runs `up -d`", which is a different question. - **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`. Intent refused, observation logged. Refusing to start a customer's app because we could not write down which version it is trades a real outage for a bookkeeping gap. - **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a complete record with a partial one. **Seeded at startup (v0.234.0).** `Manager.BackfillInstalledImages` runs once at boot, beside the desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY on. It only READS containers. **This was not a refinement — without it the feature did not reach a quiet box at all:** see §8.3, which was written as a known limitation on 2026-09-02 and was a defect by the next morning. Two admission rules, and the second is the design: - **It never overwrites an existing record.** The bring-up paths own updates; this fills gaps only. - **It refuses to seed a PARTIAL observation.** §7.2's comparison reads a service-count mismatch as BEHIND, so a degraded or crash-looping app seeded from its visible containers would render „Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial because they follow a SUCCESSFUL `up -d`, where a gap is real news and is logged; a backfill meets a box in whatever state it is in. **Same field, two writers, two different admission rules — that is deliberate and must not be "made consistent".** **Nothing reads it to take a decision.** Slice 2 reads it to render a label. ### 7.2 The label (slice 2) `web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against what the current template pins, and renders through the existing `meta_badge` partial — no new markup, no new CSS. **Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by a test with a companion red-proof. **No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`). Version strings stay in the logs, the API and the hub. --- ## 8. Known limitations, stated plainly 1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and queries no registry — a customer's box must not depend on reaching eight upstream registries to render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be identical while the image behind it has moved. **Measured, not theorised:** spike §5 found `mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls holding. Digest-level comparison needs a registry query and is deferred — **R-446**. 2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent commit to diff against, so the gate needs a deeper fetch — **R-452**. 3. ~~**The record only appears after the next lifecycle action.**~~ **CLOSED in v0.234.0, and the way it closed is worth keeping.** This was written on 2026-09-02 as an accepted limitation — *"the fleet view fills in gradually"*. The operator looked at demo-felhom the next morning and found OpenGist, up 15 hours, running exactly the catalog pin, showing **nothing at all**. **On a quiet box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped.** `BackfillInstalledImages` now seeds the absences at startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller START, so a box between upgrade and its next restart still shows nothing — bounded by one restart rather than unbounded. 4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the `wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken state. Recorded so it is a choice, not a surprise. 5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in §5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded. **The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break the update badge by withholding `catalog_since`. **R-458.** 6. **The Update button is still unguarded.** It takes no backup, has no rollback, and can still attempt a multi-major jump the app will refuse (R-40). **Slice 3 did not change that and must not be read as having done so** — the precondition is slice 4 (R-448). 7. **The hub does not record image tags at all.** Its report's container payload carries name, state, CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side change; it is not derivable from what is already reported. --- ## 9. Where the rest lives | what | where | |---|---| | the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` | | the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 | | the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` | | what the lifecycle actions are proven to do | `architecture/00-capability-map.md` | | the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |