Files
felhom.eu/documentation/architecture/09-update-architecture.md
T
admin 417df06f35
gates / gates (push) Successful in 17s
slice 3 docs: the ruling, the shipped mechanism, and four rows closed
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06,
Option 1) and its section 5 is rewritten from a proposed shape into the shipped
one: the pin, the stored definition, the render table, the four writers, the
startup ordering, and the trap this slice set for slice 2 - the live compose file
is now the frozen one, so a badge comparing against it would answer Naprakesz on
exactly the apps that are behind.

02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being
true today, so it is corrected, and the two sections describing the old seam now
carry a banner saying they describe v0.234.0 and below - kept because every box
under v0.235.0 still behaves that way and because they are the measured account
of why it changed.

R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458
opened for the .felhom.yml asymmetry, with what would settle it by measurement.

Live evidence: two real catalog pushes travelling the real 15-minute cycle, both
reverted, the tree byte-identical afterwards. The restart that used to take 18.3
seconds and pull a new image now takes 0.1 seconds and pulls nothing.
2026-09-06 10:37:33 +02:00

371 lines
22 KiB
Markdown

# 09 — How an app update works, and what it is becoming
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
> did not know it was one.
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
---
## 1. How an update works today, as measured
### 1.1 The button
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
commands and nothing else:
```
compose pull → compose up -d --remove-orphans
```
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
### 1.2 The catalog syncer moves the file underneath a deployed app
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
sync does not restart anything.**
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
### 1.3 Thirteen other paths end in `compose up -d`
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
restore path.
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
a boot reconciliation started an app on a newer image at 17:55:44Z.
### 1.4 One fear is measured SMALLER than it was stated
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
1c, a positive observable and not an absent log line).
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
so is more useful than leaving the scarier version standing.
---
## 2. What was chosen, and by whom
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
whole arc:
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
> SIGTERM+start to existing containers without re-reading the compose file or env."*
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
it does not change it.
---
## 3. The operator decisions
These are rulings, not proposals. Anything specced against a different assumption is wrong.
### 2026-09-02
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
§6 — the file half was never priced, and demo-hp is too young a box to price it.
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
A customer on the newest version is supported however old that version is. This is why
`catalog_since` exists and why no version string is shown.
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
because §4 says it cannot be undone.
### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0
4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it —
while corrections to its definition keep arriving on the 15-minute cycle exactly as they do
today.**
**This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes
that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in
its own comment. Reversing a chosen behaviour is a decision, not a bug fix.
**What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".**
The old behaviour had two halves and only one of them was unwanted:
| half | verdict |
|---|---|
| a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone |
| template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 |
So the ruling keeps the second and removes the first. In the operator's own words: *while the
catalog is offering the same version you are running, its fixes flow to you; the moment it moves to
a newer version, you are frozen at what you have until you choose to update.*
**It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup
precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.
---
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
run, putting the old image tag back produces a container that refuses to start —
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
> downgrading is not supported"*
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
| shape | when it applies | what it does |
|---|---|---|
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute
half-life because the syncer overwrites it (**R-441**).
---
## 5. The shape, as SHIPPED in v0.235.0
The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the
syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather
than an input, and the thirteen unattended `up -d` paths stop being able to change a version by
accident.
### 5.1 Nothing was added to the thirteen call sites, and that is deliberate
**They are made safe by removing the reason, not by gating them.** The most important of them are
REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`),
the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a
customer's app down, which is worse than the problem this slice solves.** Since the file they act on
no longer changes version, every one of them became safe without being touched.
### 5.2 The pin, and what it is not
`AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref.
**It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a
DECISION — what should run. Letting an observation feed a decision would make a bad reading become a
bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over.
They will normally agree; when they disagree that is a signal, not a bug to paper over.
**Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.**
Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came
from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is
safe from the catalog, and keeping it beside the app means it travels with every path that already
moves a stack dir.
### 5.3 The four writers — the only acts entitled to move a version
| writer | pin source |
|---|---|
| the deploy path (`runComposeDeploy`) | the template just deployed from |
| **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull |
| the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** |
| `Manager.AdoptPins` | the observation, once, and only when complete AND matching |
**`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file
on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards
would pull the frozen version and change nothing — while reporting success, and a button that lies is
worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of
`recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent.
### 5.4 The render table, complete
| app state | result |
|---|---|
| not deployed / protected / seam not wired | the catalog template — today's behaviour |
| deployed, **unpinned** | the catalog template + one DEBUG |
| deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** |
| deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE |
| pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it |
| mid-deploy | the compose file is left alone this cycle |
**`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries
`catalog_since`, which the badge needs. See §8.5.
**The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and
putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration
the older template cannot supply, so an old image under a new template is broken in a way neither
version is.
**And this is not "skip deployed apps".** That option was considered and rejected: it also stops
health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in
the spike §3 — both halves the ruling explicitly kept.
### 5.5 Adoption, and why the startup order matters
`AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app
to what it is already running. **It reads and writes files only** — no container is started, stopped or
touched. It skips, loudly, when the observation is incomplete or when the app runs something the
current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a
guessed pin.
**`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position
that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and
handed the next restart a version change — the exact behaviour this slice removes, once per boot.
### 5.6 The trap this slice set for the previous one
`Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one.
On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find
installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test
still green, because the new field has the same type and shape. The badge now reads
`Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an
earlier feature is the failure mode to look for whenever a file changes meaning.**
---
## 6. The seven slices
| # | slice | status |
|---|---|---|
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
| **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
breaks.
---
## 7. What slices 1 and 2 actually built
### 7.1 The record (slice 1)
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
and writes `app.yaml`:
```yaml
installed_images:
web:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:… # "" if the image was never pulled from a registry
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
```
Three rules, each with its reason:
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
already moved. A record built from it would answer "what will happen next time something runs
`up -d`", which is a different question.
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
Intent refused, observation logged. Refusing to start a customer's app because we could not write
down which version it is trades a real outage for a bookkeeping gap.
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
complete record with a partial one.
**Seeded at startup (v0.234.0).** `Manager.BackfillInstalledImages` runs once at boot, beside the
desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY
on. It only READS containers. **This was not a refinement — without it the feature did not reach a
quiet box at all:** see §8.3, which was written as a known limitation on 2026-09-02 and was a defect
by the next morning.
Two admission rules, and the second is the design:
- **It never overwrites an existing record.** The bring-up paths own updates; this fills gaps only.
- **It refuses to seed a PARTIAL observation.** §7.2's comparison reads a service-count mismatch as
BEHIND, so a degraded or crash-looping app seeded from its visible containers would render
„Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial
because they follow a SUCCESSFUL `up -d`, where a gap is real news and is logged; a backfill meets a
box in whatever state it is in. **Same field, two writers, two different admission rules — that is
deliberate and must not be "made consistent".**
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
### 7.2 The label (slice 2)
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
no new CSS.
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
a test with a companion red-proof.
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
Version strings stay in the logs, the API and the hub.
---
## 8. Known limitations, stated plainly
1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and
queries no registry — a customer's box must not depend on reaching eight upstream registries to
render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be
identical while the image behind it has moved. **Measured, not theorised:** spike §5 found
`mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls
holding. Digest-level comparison needs a registry query and is deferred — **R-446**.
2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
commit to diff against, so the gate needs a deeper fetch — **R-452**.
3. ~~**The record only appears after the next lifecycle action.**~~ **CLOSED in v0.234.0, and the way
it closed is worth keeping.** This was written on 2026-09-02 as an accepted limitation — *"the
fleet view fills in gradually"*. The operator looked at demo-felhom the next morning and found
OpenGist, up 15 hours, running exactly the catalog pin, showing **nothing at all**. **On a quiet
box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is,
on the quiet installations, not shipped.** `BackfillInstalledImages` now seeds the absences at
startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller
START, so a box between upgrade and its next restart still shows nothing — bounded by one restart
rather than unbounded.
4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches
that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the
`wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken
state. Recorded so it is a choice, not a surprise.
5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
the update badge by withholding `catalog_since`. **R-458.**
6. **The Update button is still unguarded.** It takes no backup, has no rollback, and can still
attempt a multi-major jump the app will refuse (R-40). **Slice 3 did not change that and must not
be read as having done so** — the precondition is slice 4 (R-448).
7. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.
---
## 9. Where the rest lives
| what | where |
|---|---|
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |