Files
felhom.eu/documentation/architecture/09-update-architecture.md
T
admin 6035dfcc3a
gates / gates (push) Successful in 17s
09-update-architecture.md: the update path finally has a document, and it is a living one
R-438's document half. It records how an update works AS MEASURED, quotes the
RestartStack comment that proves the restart half was CHOSEN (a design decision
is not a defect), carries the three operator rulings of 2026-09-02, strikes the
word 'rollback' (once a migration has run the old image will not start), states
the target shape, and lists the seven slices with a status each.

R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not
changed. Nothing closed, so CLOSED-ITEMS.md is untouched.

Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23
floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner),
R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and
R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what
stopped the badge render from being validated live).

Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN
LIVE through the boot reconciler on demo-hp - one entry per compose service,
digests matching ground truth read independently. The badge RENDER is not, and
the five attempts are listed rather than summarised.
2026-09-02 20:32:30 +02:00

231 lines
12 KiB
Markdown

# 09 — How an app update works, and what it is becoming
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
> did not know it was one.
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
---
## 1. How an update works today, as measured
### 1.1 The button
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
commands and nothing else:
```
compose pull → compose up -d --remove-orphans
```
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
### 1.2 The catalog syncer moves the file underneath a deployed app
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
sync does not restart anything.**
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
### 1.3 Thirteen other paths end in `compose up -d`
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
restore path.
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
a boot reconciliation started an app on a newer image at 17:55:44Z.
### 1.4 One fear is measured SMALLER than it was stated
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
1c, a positive observable and not an absent log line).
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
so is more useful than leaving the scarier version standing.
---
## 2. What was chosen, and by whom
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
whole arc:
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
> SIGTERM+start to existing containers without re-reading the compose file or env."*
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
it does not change it.
---
## 3. The three operator decisions (2026-09-02)
These are rulings, not proposals. Anything specced against a different assumption is wrong.
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
§6 — the file half was never priced, and demo-hp is too young a box to price it.
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
A customer on the newest version is supported however old that version is. This is why
`catalog_since` exists and why no version string is shown.
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
because §4 says it cannot be undone.
---
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
run, putting the old image tag back produces a container that refuses to start —
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
> downgrading is not supported"*
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
| shape | when it applies | what it does |
|---|---|---|
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute
half-life because the syncer overwrites it (**R-441**).
---
## 5. The target shape
**The live `docker-compose.yml` becomes DERIVED from a pin recorded in `app.yaml`** — the one file the
syncer never touches (`sync.go:319`'s exclusion). The catalog then proposes; `app.yaml` decides; the
rendered compose file is an output rather than an input, and the thirteen unattended `up -d` paths
stop being able to change a version by accident.
**Nothing in slices 1 or 2 implements this.** They make the current state *visible*, which is the
prerequisite for judging how urgent it is.
---
## 6. The seven slices
| # | slice | status |
|---|---|---|
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
| **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
breaks.
---
## 7. What slices 1 and 2 actually built
### 7.1 The record (slice 1)
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
and writes `app.yaml`:
```yaml
installed_images:
web:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:… # "" if the image was never pulled from a registry
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
```
Three rules, each with its reason:
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
already moved. A record built from it would answer "what will happen next time something runs
`up -d`", which is a different question.
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
Intent refused, observation logged. Refusing to start a customer's app because we could not write
down which version it is trades a real outage for a bookkeeping gap.
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
complete record with a partial one.
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
### 7.2 The label (slice 2)
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
no new CSS.
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
a test with a companion red-proof.
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
Version strings stay in the logs, the API and the hub.
---
## 8. Known limitations, stated plainly
1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and
queries no registry — a customer's box must not depend on reaching eight upstream registries to
render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be
identical while the image behind it has moved. **Measured, not theorised:** spike §5 found
`mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls
holding. Digest-level comparison needs a registry query and is deferred — **R-446**.
2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
commit to diff against, so the gate needs a deeper fetch — **R-452**.
3. **The record only appears after the next lifecycle action.** An app that is running and untouched
keeps a legacy `app.yaml` and therefore no badge, until someone restarts, updates or redeploys it.
That is correct — the alternative is inventing a record from the file §1.2 says has already moved —
but it means the fleet view fills in gradually rather than at upgrade.
4. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.
---
## 9. Where the rest lives
| what | where |
|---|---|
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |