Files
felhom.eu/documentation/architecture/09-update-architecture.md
T

604 lines
40 KiB
Markdown

# 09 — How an app update works, and what it is becoming
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
> did not know it was one.
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
---
## 1. How an update works today, as measured
### 1.1 The button
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
commands and nothing else:
```
compose pull → compose up -d --remove-orphans
```
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
### 1.2 The catalog syncer moves the file underneath a deployed app
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
sync does not restart anything.**
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
### 1.3 Thirteen other paths end in `compose up -d`
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
restore path.
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
a boot reconciliation started an app on a newer image at 17:55:44Z.
### 1.4 One fear is measured SMALLER than it was stated
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
1c, a positive observable and not an absent log line).
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
so is more useful than leaving the scarier version standing.
---
## 2. What was chosen, and by whom
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
whole arc:
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
> SIGTERM+start to existing containers without re-reading the compose file or env."*
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
it does not change it.
---
## 3. The operator decisions
These are rulings, not proposals. Anything specced against a different assumption is wrong.
### 2026-09-02
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
§6 — the file half was never priced, and demo-hp is too young a box to price it.
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
A customer on the newest version is supported however old that version is. This is why
`catalog_since` exists and why no version string is shown.
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
because §4 says it cannot be undone.
### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0
4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it —
while corrections to its definition keep arriving on the 15-minute cycle exactly as they do
today.**
**This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes
that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in
its own comment. Reversing a chosen behaviour is a decision, not a bug fix.
**What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".**
The old behaviour had two halves and only one of them was unwanted:
| half | verdict |
|---|---|
| a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone |
| template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 |
So the ruling keeps the second and removes the first. In the operator's own words: *while the
catalog is offering the same version you are running, its fixes flow to you; the moment it moves to
a newer version, you are frozen at what you have until you choose to update.*
**It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup
precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.
### 2026-09-13 — the database engine finishes its own conversion, and the upgrade test goes wide
5. **DBs should be updated when the app moves, with proper precautions, tests and backoff plans**
— the operator's own words, ruling on `SPIKE-r459-mariadb-upgrade-2026-09-06.md`. SHIPPED in the
catalog the same day: every `mariadb:` sidecar (`bookstack-db`, `kimai-db`, `nextcloud-db`,
`romm-db`) carries `MARIADB_AUTO_UPGRADE=1`; `MARIADB_DISABLE_UPGRADE_BACKUP` stays unset. **Not an
image change, so `catalog_since` does not move.** The three precautions, because they are the real
content of the ruling:
1. **Proven before it ships** — `upgrade-test.py` re-ran E3 and E3b on the changed template and the
engine-state field shows the conversion RAN (`mariadb_upgrade_info` reads the new version, the
engine's own check says nothing further is needed, the entrypoint no longer prints
`skipped due to $MARIADB_AUTO_UPGRADE`), with the seeded data reading back after. C3 still
returns `failed`. Evidence: `audits/r459-close-2026-09-13/`.
2. **Watched as it lands** — the change travelled the real 15-minute cycle to demo-hp: the live
compose gained the setting, the sync recreated nothing, and one deliberate restart logged
`MariaDB upgrade not required` with the app serving (same evidence directory).
3. **A rule until Slice 4 is built** — the Update button still takes no backup, so **no template
may move a database-engine image across a major version until R-448 ships.** Catalog
`CLAUDE.md` states it; `scripts/check-engine-major.py` enforces it in the pre-push hook (the
CI half cannot, R-452); its removal is tracked as **R-469** so it is a deliberate act.
The setting is inert until an engine major moves, and precaution 3 keeps it that way.
6. **The upgrade test goes as wide as possible, through the nightly unattended sessions** — the
ruling on `STATUS.md` item 11. Not "the ~25 database apps first": all of them, as the nightly
rotation reaches them, one fixture per app through the app's own interface. **Browser-only apps
become reachable when CC runs on the operator's Windows workstation with Chrome** — the
`claude-in-chrome` route that DooPlex does not have — so an app recorded `inconclusive` for want
of a headless seed route (bookstack's file half, R-460) is deferred to that venue, not faked.
---
### 2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts
7. **A floor carries a release past the vouched golden when the release's MinAgent is declared with
it** (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it,
the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which
`minagent_header_gate.py` now guarantees — goes into the same per-box agent comparison. An
undeclared floor above the golden is still HELD, and both floor forms refuse to save one. **Why:** a
controller image is pulled by tag and needs no golden to be delivered; what the floor was missing
was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release
delivery coexist. **Proven live:** both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s
from the save, the hub logging `SERVED … from declared`
(`audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt`).
8. **Any backup tier lets an app update** (R-475; controller v0.239.0). The precondition takes the first
FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site
(Tier 3, bounded; unreachable = absent). `update.backup_max_age` applies to whichever tier is chosen.
An app with nothing anywhere is backed up first; it is refused only when no backup can be taken
either. The hold names the tier and the date. **Tier 2 is required nowhere in the update path.**
This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a
verified recent backup as a precondition, not a new copy) stands. **Proven live:**
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy
holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit
that holds the definition and the database dumps and NOT the files (measured: gokapi restored from
„helyi" came back with settings and no data). For such an app the update walks second drive →
off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive →
own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a
customer is never sent to a copy that cannot bring the data back without being told so.
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
run, putting the old image tag back produces a container that refuses to start —
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
> downgrading is not supported"*
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
| shape | when it applies | what it does |
|---|---|---|
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
There is no third.
### 4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades
`SPIKE-upgrade-test-2026-09-06.md` upgraded three real apps with real data in them and then attempted
the abort on each. **All five real catalog upgrades kept the customer's data.** The abort did not
behave the same way twice:
| app | abort | why |
|---|---|---|
| **docmost** `0.25.3`→`0.95.0` | **REFUSES** | the old code finds migration ledger entries it does not know: *"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*, then *"Failed to run database migration. Exiting program."* |
| **privatebin** `1.7.5`→`2.0.5` | **works** | file-backed, no database, no schema — a major version moves no data |
| **bookstack** app+engine | **works, misleadingly** | only because the MariaDB datadir upgrade was skipped and never happened — see R-459 |
**This puts TWO independent measurements behind the ruling above, by two unrelated mechanisms:**
Nextcloud refused on an explicit version comparison; docmost refuses on its migration ledger. The word
"rollback" was already struck; it is now struck on evidence rather than on one case.
**And it adds a distinction this document did not have: there is no single answer to "can this update
be undone". There are apps where it can and apps where it cannot, and the only way to know which is to
MEASURE THAT APP.** Any design that assumes one answer for all 53 is designing against a fact that was
checked and is false.
Per §5 below, a restore's image-level undo used to have a ≤15-minute half-life because the syncer
overwrote it (**R-441**) — closed in v0.235.0.
---
## 5. The shape, as SHIPPED in v0.235.0
The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the
syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather
than an input, and the thirteen unattended `up -d` paths stop being able to change a version by
accident.
### 5.1 Nothing was added to the thirteen call sites, and that is deliberate
**They are made safe by removing the reason, not by gating them.** The most important of them are
REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`),
the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a
customer's app down, which is worse than the problem this slice solves.** Since the file they act on
no longer changes version, every one of them became safe without being touched.
### 5.2 The pin, and what it is not
`AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref.
**It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a
DECISION — what should run. Letting an observation feed a decision would make a bad reading become a
bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over.
They will normally agree; when they disagree that is a signal, not a bug to paper over.
**Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.**
Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came
from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is
safe from the catalog, and keeping it beside the app means it travels with every path that already
moves a stack dir.
### 5.3 The four writers — the only acts entitled to move a version
| writer | pin source |
|---|---|
| the deploy path (`runComposeDeploy`) | the template just deployed from |
| **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull |
| the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** |
| `Manager.AdoptPins` | the observation, once, and only when complete AND matching |
**`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file
on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards
would pull the frozen version and change nothing — while reporting success, and a button that lies is
worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of
`recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent.
### 5.4 The render table, complete
| app state | result |
|---|---|
| not deployed / protected / seam not wired | the catalog template — today's behaviour |
| deployed, **unpinned** | the catalog template + one DEBUG |
| deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** |
| deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE |
| pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it |
| mid-deploy | the compose file is left alone this cycle |
**`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries
`catalog_since`, which the badge needs. See §8.5.
**The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and
putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration
the older template cannot supply, so an old image under a new template is broken in a way neither
version is.
**And this is not "skip deployed apps".** That option was considered and rejected: it also stops
health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in
the spike §3 — both halves the ruling explicitly kept.
### 5.5 Adoption, and why the startup order matters
`AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app
to what it is already running. **It reads and writes files only** — no container is started, stopped or
touched. It skips, loudly, when the observation is incomplete or when the app runs something the
current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a
guessed pin.
**`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position
that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and
handed the next restart a version change — the exact behaviour this slice removes, once per boot.
### 5.6 The trap this slice set for the previous one
`Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one.
On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find
installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test
still green, because the new field has the same type and shape. The badge now reads
`Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an
earlier feature is the failure mode to look for whenever a file changes meaning.**
---
## 6. The seven slices
| # | slice | status |
|---|---|---|
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 |
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)
`POST /api/stacks/{name}/update` is a guarded job. It answers **202** at once; the outcome exists only
on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`, `update_error`,
`hold_reason`), and `update_phase=done` is written only after the app's health is known. **R-443 is
closed by construction: nothing reports an update complete on the compose exit code.**
**The sequence, and the order is the design:**
| # | phase | what happens | on failure |
|---|---|---|---|
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no copy on ANY tier and no backup can be taken now** (since v0.239.0; before it, no restorable Tier-2 copy) | nothing moves, nothing is recorded |
| 1 | `checking` | walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than `update.backup_max_age` (v0.239.0) | nothing moves |
| 2 | `backing-up` — only when no tier holds a fresh copy | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 | refused with the backup's own error; nothing moves |
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
| 8 | `done` | installed images recorded, journal cleared | — |
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
`update.health_timeout` (default `5m`).
**The precondition is the existing verified backup, not a new copy** (§3 decision 1). It is
`backup.Tier2UnitRestorePoint` — the SAME predicate that permits the destructive „Teljes
visszaállítás", extracted from the backups page rather than copied. **The copy is aged by the last
SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was measured before it was
designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp
bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the
manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).** **SUPERSEDED in
v0.239.0 by §3 decision 8:** `backup.Manager.UpdateRestorePoints` walks all three tiers and the update
leans on the first fresh copy; `Tier2UnitRestorePoint` is still the Tier-2 half and still the page's
predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum
skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit
proven current. The hold stores `copy_tier` and names „második meghajtó" / „saját meghajtó" /
„távoli mentés"; a successful off-site restore now lifts an update hold too.
**The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and
gate as R-379, so every start path that already honoured a restore hold honours this one. A successful
unit restore lifts an update hold (only that kind). **Three unattended paths honoured no hold before
v0.237.0 and now do:** the drive-return gate's restart and boot recreate, and the nightly volume dump
(which ends in `StartStack`). The nightly capture and Tier-2 run skip a held app, so the restore point
the hold text names is never overwritten.
**Crash safety is a journal**, `<data>/update-journal.json`, written before every phase. `RecoverUpdates`
runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back;
after `up` → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by
`ResumeInterruptedUpdates` once the backup side is wired, ending healthy or held.
**THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT.** Whether an old image
starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot
be predicted. So the box never puts the old version back by itself. **The route back is the restore**,
and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where
the harness has proven it, is slice 6's.
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
both demo boxes by the floor alone, with its MinAgent declared.
**Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not
only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute
health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit.
The Tier-2 mirror the hold names survived only because Tier 2 is daily. `backup.Manager.isHeld` now also
answers true for an app a guarded update is moving (`SetUpdatingCheck`).
**v0.240.0 (2026-09-13, evening) — what the afternoon's proof and the first nightly rotation found, fixed.**
Seven rows: removal with backups kept now keeps the Tier-2 RECORD, so the second-drive restore is not
refused over an intact mirror (R-486, P1 — the disaster the second copy exists for); PostGIS/pgvector/
TimescaleDB images are Postgres, so such apps get their logical dump (R-484); "delete backups" deletes
the unit, the mirror(s) and the prefs (R-474/R-466); the backup card sizes them (R-485); a held
update's sentence leaves the card with the hold (R-480); the Tier-3 lookup is one `snapshots` call
(R-477); a unit older than the app's `deployed_at` does not count (R-478). Delivered by the floor in
16 s / 18 s; every row proven live with a throwaway adventurelog. `audits/v0240-2026-09-13/`.
**Proven live on demo-hp, 2026-09-13**, with a throwaway uptime-kuma and real catalog tag changes (each
reverted in the same phase): A (2.3.2→2.4.0, done after health), B (`backup_max_age: 2m` → backup first),
E (non-existent tag → pin back, container untouched), F (`alpine:3.20` → held), H (three buttons and the
boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold
cleared). Live evidence: `audits/slice4-2026-09-13/`.
### The verdict record — the contract Slice 6 carries
Decided here rather than invented twice. The harness writes one of these per edge, beside its
evidence; Slice 6 puts the same shape in the catalog.
```json
{"harness_version": 1, "app": "bookstack",
"from": {"bookstack": "…:25.02.2", "bookstack-db": "mariadb:11.6"},
"to": {"bookstack": "…:26.05.2", "bookstack-db": "mariadb:12.3"},
"verdict": "proven | failed | inconclusive",
"seed_read_before": true, "seed_read_after": true, "healthy_after": true,
"migration_observed": "verbatim log line, or null",
"abort": "starts-and-serves | refuses | starts-data-gone | not-attempted",
"abort_detail": "the refusal quoted verbatim, or null",
"duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"}
```
**`inconclusive` is a first-class verdict and must never be collapsed into `failed`.** "We could not
measure it" and "it does not work" are different facts, and only one of them is about the app.
**`migration_observed` is a quoted line, never an inference from timing** — the value of both the
Nextcloud and the docmost findings was the exact sentence the app printed.
### Database engines under an upgrade — MEASURED 2026-09-06
**The arc's standing rule that an engine change gets its OWN edge now has measured evidence behind
it**, and the evidence is stronger than the rule's original argument. The rule was justified by
*"two migrations behind one edge is an unreadable failure when it breaks"* — a readability argument.
What was measured is that **an engine change can be applied and silently NOT happen**, which the
app-half edge cannot produce and which no amount of readability would have surfaced:
- `SPIKE-upgrade-test-2026-09-06.md` §4 — MariaDB 12.3 starts on an 11.6 datadir, logs that the
conversion it requires was **skipped**, and serves. **Assigned to the engine half by decomposition:**
the app half alone produces no such line.
- `SPIKE-r459-mariadb-upgrade-2026-09-06.md` — it is **stable but never self-resolving** (5 of 5
restarts, no degradation, and the engine says `Check required!` every time, forever). Converting
properly **succeeds**, costs **7 s**, takes its own system-database backup, and **does not** cost the
ability to abort. **The trade that was expected here does not exist.**
- **2026-09-13 — the setting is in the catalog.** All four `mariadb:` sidecars carry
`MARIADB_AUTO_UPGRADE=1` (operator ruling, §3 decision 5), and `upgrade-test.py`'s engine-state field
now shows the conversion RUNNING on the bookstack edges. **And a gate holds the engines inside their
major until Slice 4:** `app-catalog-felhom.eu/scripts/check-engine-major.py` (R-469).
**Two rules for anything this arc builds around a database engine:**
1. **Ask the engine, not the log.** MariaDB's entrypoint prints `MariaDB upgrade not required` on an
unsupported **downgrade**; `mariadb-upgrade --check-if-upgrade-is-needed` names it exactly
(**R-464**). A cheap instrument built on the log line would report "fine" for the broken case.
2. **The two engines fail in opposite directions, so one check will not do.** MariaDB starts anyway
and skips quietly; **PostgreSQL refuses to start** on a datadir from an older major, and the image
performs no `pg_upgrade`. Eleven templates carry PostgreSQL and **eight sit on `postgres:16-alpine`**
(**R-463**).
**And an engine-state field belongs BESIDE a verdict, never inside it.** `upgrade-test.py` reports
`engine_state_after` next to `verdict`, because an unconverted datadir is not *known* to be a failure
and a verdict that said so would encode an unproven judgement.
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
breaks.
---
## 7. What slices 1 and 2 actually built
### 7.1 The record (slice 1)
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
and writes `app.yaml`:
```yaml
installed_images:
web:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:… # "" if the image was never pulled from a registry
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
```
Three rules, each with its reason:
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
already moved. A record built from it would answer "what will happen next time something runs
`up -d`", which is a different question.
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
Intent refused, observation logged. Refusing to start a customer's app because we could not write
down which version it is trades a real outage for a bookkeeping gap.
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
complete record with a partial one.
**Seeded at startup (v0.234.0).** `Manager.BackfillInstalledImages` runs once at boot, beside the
desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY
on. It only READS containers. **This was not a refinement — without it the feature did not reach a
quiet box at all:** see §8.3, which was written as a known limitation on 2026-09-02 and was a defect
by the next morning.
Two admission rules, and the second is the design:
- **It never overwrites an existing record.** The bring-up paths own updates; this fills gaps only.
- **It refuses to seed a PARTIAL observation.** §7.2's comparison reads a service-count mismatch as
BEHIND, so a degraded or crash-looping app seeded from its visible containers would render
„Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial
because they follow a SUCCESSFUL `up -d`, where a gap is real news and is logged; a backfill meets a
box in whatever state it is in. **Same field, two writers, two different admission rules — that is
deliberate and must not be "made consistent".**
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
### 7.2 The label (slice 2)
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
no new CSS.
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
a test with a companion red-proof.
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
Version strings stay in the logs, the API and the hub.
---
## 8. Known limitations, stated plainly
1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and
queries no registry — a customer's box must not depend on reaching eight upstream registries to
render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be
identical while the image behind it has moved. **Measured, not theorised:** spike §5 found
`mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls
holding. Digest-level comparison needs a registry query and is deferred — **R-446**.
2. **~~Nothing enforces `catalog_since`.~~ Enforced by the pre-push hook since 2026-09-13 (R-452, `app-catalog-felhom.eu/scripts/check-catalog-since.py`); CI's shallow clone still skips it out loud.** A commit that moves an `image:` line and forgets the date
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
commit to diff against, so the gate needs a deeper fetch — **R-452**.
3. ~~**The record only appears after the next lifecycle action.**~~ **CLOSED in v0.234.0, and the way
it closed is worth keeping.** This was written on 2026-09-02 as an accepted limitation — *"the
fleet view fills in gradually"*. The operator looked at demo-felhom the next morning and found
OpenGist, up 15 hours, running exactly the catalog pin, showing **nothing at all**. **On a quiet
box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is,
on the quiet installations, not shipped.** `BackfillInstalledImages` now seeds the absences at
startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller
START, so a box between upgrade and its next restart still shows nothing — bounded by one restart
rather than unbounded.
4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches
that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the
`wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken
state. Recorded so it is a choice, not a surprise.
5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
the update badge by withholding `catalog_since`. **R-458.**
6. ~~**The Update button is still unguarded.**~~ **CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1).** It
refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a
safety dump, and holds an app that does not come up. **What stays true:** it still has no automatic
rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) —
that now ends HELD rather than crash-looping behind a green button.
7. ~~**An engine major can be applied without its datadir upgrade, and nothing notices.**~~ **CLOSED
2026-09-13 for MariaDB (R-459):** every `mariadb:` sidecar carries `MARIADB_AUTO_UPGRADE=1`, and the
harness shows the conversion running on the E3/E3b edges (§3 decision 5). **What stays true:** the
PostgreSQL half (R-463) has no equivalent — the image performs no `pg_upgrade` — and the
engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until
Slice 4 gives the Update button a backup.
8. **Only three of 53 apps have ever had an upgrade measured**, and one of them (bookstack) can only
be half-proven headlessly (**R-460**). The widening is **R-462**, costed with real numbers.
9. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.
---
## 9. Where the rest lives
| what | where |
|---|---|
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |