The operator looked at demo-felhom and found OpenGist - up 15 hours, running exactly the catalog pin, showing no badge at all. 09-update-architecture.md had recorded that as an accepted limitation the day before: 'the fleet view fills in gradually'. On a quiet box gradually means never, and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped. That limitation row is now struck with the reason kept. The living document gains slice 1b, the two admission rules of the backfill (it never overwrites, and it refuses to seed a partial observation because the badge reads a service-count mismatch as BEHIND), and the note that the same field having two writers with two different admission rules is deliberate. Live evidence added: all nine apps already had records by the time 0.234.0 was ready, so the natural fleet state could no longer exercise the new code - said plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp by stripping two records; the backfill re-seeded exactly those two with digests matching independently-read ground truth and left the other seven alone. The refusal half was deliberately NOT staged live: it needs a degraded app, and manufacturing one risks the false-customer-email class that already cost 61 mails (R-330). Unit-tested with a red-proof, and recorded as unproven-live. R-457: a test that hardcodes a date and asserts an age derived from it is green only on the day it is written. Mine was, and it went red overnight. Six other files carry both a date literal and time.Now() - named as candidates, not accused.
14 KiB
09 — How an app update works, and what it is becoming
LIVING DOCUMENT. Every slice of the update arc updates this file in the same session. Opened 2026-09-02 with slices 1 and 2. Its absence was R-438: the update mechanism was chosen deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who did not know it was one.
This file carries the REASONING. The register (backlog/OPEN-ITEMS.md) carries the work. The
source is the truth. Nothing here is invented: every mechanism claim is cited either to
audits/SPIKE-app-update-2026-09-01.md, which measured it live, or to live source at file:symbol.
1. How an update works today, as measured
1.1 The button
Manager.UpdateStack (felhom-controller/controller/internal/stacks/manager.go:1199) is two compose
commands and nothing else:
compose pull → compose up -d --remove-orphans
No safety copy. No rollback. No hold. No verification. Confirmed by reading and across six live
updates (spike §10 item 6). A pull FAILURE is handled correctly — UpdateStack returns after the
failed pull and never reaches up -d, so the running app survives untouched (measured twice, spike
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
1.2 The catalog syncer moves the file underneath a deployed app
Syncer.copyTemplates (felhom-controller/controller/internal/sync/sync.go:319) copies
docker-compose.yml and .felhom.yml into every stack folder on a 15-minute cycle
(internal/config/config.go:351, default 15m). It has no deployed check of any kind. Its only
guard is a sha256 content compare in copyIfChanged (sync.go:403) and its only exclusion is
app.yaml. The post-sync hook is stackMgr.InjectMissingFields(updated) and nothing else — the
sync does not restart anything.
That is why a deployed app's compose file and its running containers can disagree indefinitely. Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
1.3 Thirteen other paths end in compose up -d
Excluding the three API actions, 13 call sites across 9 files call StartStack or RestartStack,
and every one ends in compose up -d against the live compose file (spike §8 — the task that
commissioned the spike said five; the count is thirteen). They include the boot reconciler
(bootrecon.go:269), the app-stop guard's recovery (appstop_marker.go:283), the drive-return
gate (intermediary.go:222), the quiesce restart-after-backup, off-site reconstitution, and every
restore path.
So an upgrade can happen with nobody pressing anything — measured, spike §2 variant 1c-ii, where a boot reconciliation started an app on a newer image at 17:55:44Z.
1.4 One fear is measured SMALLER than it was stated
A plain power cut does not upgrade anything. Docker's own restart: unless-stopped puts the
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
up -d — it says so in its own words: no boot-orphaned apps (nothing to start) (spike §2, variant
1c, a positive observable and not an absent log line).
The unattended upgrade needs the narrower precondition: "and the app did not come back." Saying so is more useful than leaving the scarier version standing.
2. What was chosen, and by whom
Manager.RestartStack (internal/stacks/manager.go:1161) carries this comment, and it predates the
whole arc:
"Use
up -dinstead of barerestartso that env vars from app.yaml are injected and any template changes (new images, healthchecks) are picked up. Plaindocker compose restartonly sends SIGTERM+start to existing containers without re-reading the compose file or env."
So the restart behaviour was chosen, deliberately, and written down. A design decision is not a defect. What was never decided — and is recorded nowhere — is what happens once the catalog syncer moves the file underneath a deployed app, and whether the choice was meant to extend to the thirteen unattended call sites. That gap is R-438, and it stays open: this document records the mechanism; it does not change it.
3. The three operator decisions (2026-09-02)
These are rulings, not proposals. Anything specced against a different assumption is wrong.
-
The safety copy is a verified recent backup as a PRECONDITION — not a new copy invented for the update path. The guest-snapshot alternative is to be spiked before anything is designed around it. Context: the existing safety machinery (
Manager.writeSafetyDump,internal/backup/offbox_reconstitute.go:207) is database-only, which is the headline of spike §6 — the file half was never priced, and demo-hp is too young a box to price it. -
The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is. A customer on the newest version is supported however old that version is. This is why
catalog_sinceexists and why no version string is shown. -
Updates are automatic WITHIN a major, never ACROSS one. The cross-major case needs a human, because §4 says it cannot be undone.
4. The vocabulary ruling — "rollback" is struck
App data CANNOT be rolled back. Measured on Nextcloud (spike §7): once a migration has actually run, putting the old image tag back produces a container that refuses to start —
"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
So "rollback" must not appear in any spec for this arc. The two shapes actually available are:
| shape | when it applies | what it does |
|---|---|---|
| ABORT | before anything migrated | stop, put the old image back, the app runs again |
| RESTORE FROM A COPY | after a migration ran | the data restore is the whole remedy |
There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute half-life because the syncer overwrites it (R-441).
5. The target shape
The live docker-compose.yml becomes DERIVED from a pin recorded in app.yaml — the one file the
syncer never touches (sync.go:319's exclusion). The catalog then proposes; app.yaml decides; the
rendered compose file is an output rather than an input, and the thirteen unattended up -d paths
stop being able to change a version by accident.
Nothing in slices 1 or 2 implements this. They make the current state visible, which is the prerequisite for judging how urgent it is.
6. The seven slices
| # | slice | status |
|---|---|---|
| 1 | The box records what it actually installed — app.yaml.installed_images, per compose service, ref + digest + first-seen. |
SHIPPED, controller v0.233.0 (2026-09-02) |
| 1b | Seed the record for apps nobody touches — a startup backfill, so the label is not restricted to apps that happen to get restarted. | SHIPPED, controller v0.234.0 (2026-09-03) |
| 2 | One badge says whether the app is current — „Naprakész" / „Frissítés elérhető — N napja", from catalog_since. No version number. |
SHIPPED, controller v0.233.0 + catalog 69761cf (2026-09-02) |
| 3 | The compose file becomes DERIVED — stop the syncer overwriting a deployed app's file; the pin in app.yaml wins. Needs the operator's ruling on R-438 first. |
OPEN — R-447 |
| 4 | A guarded update — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
| 5 | An upgrade test — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 |
| 6 | A version sequence — updates automatic within a major, a human across one; an engine change gets its own edge. | OPEN — R-450 |
| 7 | A fleet sweep pipeline — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
The rule slice 6 inherits, recorded now while it is cheap: an engine change gets its own edge,
never bundled with an app version bump. bookstack moved the application and MariaDB 11.6 → 12.3 in
one commit (0b73e5e); that is two migrations behind one edge, and an unreadable failure when it
breaks.
7. What slices 1 and 2 actually built
7.1 The record (slice 1)
Manager.recordInstalledImages (felhom-controller/controller/internal/stacks/installed.go) runs
after a successful compose up from StartStack, RestartStack, UpdateStack and runComposeDeploy,
and writes app.yaml:
installed_images:
web:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:… # "" if the image was never pulled from a registry
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
Three rules, each with its reason:
- It reads the CONTAINER, never
docker-compose.yml. §1.2 is why: that file is the value that has already moved. A record built from it would answer "what will happen next time something runsup -d", which is a different question. - A failed write NEVER refuses the action — deliberately the opposite of
SetDesiredState. Intent refused, observation logged. Refusing to start a customer's app because we could not write down which version it is trades a real outage for a bookkeeping gap. - It is NOT called from
StartStackServices— the R-47 DB-only restore window would overwrite a complete record with a partial one.
Seeded at startup (v0.234.0). Manager.BackfillInstalledImages runs once at boot, beside the
desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY
on. It only READS containers. This was not a refinement — without it the feature did not reach a
quiet box at all: see §8.3, which was written as a known limitation on 2026-09-02 and was a defect
by the next morning.
Two admission rules, and the second is the design:
- It never overwrites an existing record. The bring-up paths own updates; this fills gaps only.
- It refuses to seed a PARTIAL observation. §7.2's comparison reads a service-count mismatch as
BEHIND, so a degraded or crash-looping app seeded from its visible containers would render
„Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial
because they follow a SUCCESSFUL
up -d, where a gap is real news and is logged; a backfill meets a box in whatever state it is in. Same field, two writers, two different admission rules — that is deliberate and must not be "made consistent".
Nothing reads it to take a decision. Slice 2 reads it to render a label.
7.2 The label (slice 2)
web.updateBadge (internal/web/updatebadge.go) compares the recorded reference per service against
what the current template pins, and renders through the existing meta_badge partial — no new markup,
no new CSS.
Absent means UNKNOWN and never means current. Every app.yaml written before v0.233.0 has no
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
a test with a companion red-proof.
No version number reaches the customer (operator ruling: a household cannot act on 26.05.2).
Version strings stay in the logs, the API and the hub.
8. Known limitations, stated plainly
- „Naprakész" can be FALSE for the 23 floating pins. The comparison is reference-to-reference and
queries no registry — a customer's box must not depend on reaching eight upstream registries to
render a page. For
postgres:16-alpine,mariadb:11.6and 21 others the reference can be identical while the image behind it has moved. Measured, not theorised: spike §5 foundmariadb:11.4andmariadb:12.3had both already moved upstream, with two fully-pinned controls holding. Digest-level comparison needs a registry query and is deferred — R-446. - Nothing enforces
catalog_since. A commit that moves animage:line and forgets the date under-reports how far behind a box is. The gates runner fetches at--depth 1and has no parent commit to diff against, so the gate needs a deeper fetch — R-452. The record only appears after the next lifecycle action.CLOSED in v0.234.0, and the way it closed is worth keeping. This was written on 2026-09-02 as an accepted limitation — "the fleet view fills in gradually". The operator looked at demo-felhom the next morning and found OpenGist, up 15 hours, running exactly the catalog pin, showing nothing at all. On a quiet box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped.BackfillInstalledImagesnow seeds the absences at startup by reading containers (§7.1). The residue that stays: the seed happens at controller START, so a box between upgrade and its next restart still shows nothing — bounded by one restart rather than unbounded.- The hub does not record image tags at all. Its report's container payload carries name, state, CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side change; it is not derivable from what is already reported.
9. Where the rest lives
| what | where |
|---|---|
| the measurements this document rests on | audits/SPIKE-app-update-2026-09-01.md |
| the work | backlog/OPEN-ITEMS.md — R-438..R-445, R-446..R-452 |
| the syncer, described accurately but without the consequence | architecture/02-controller-module-map.md |
| what the lifecycle actions are proven to do | architecture/00-capability-map.md |
| the implementation | felhom-controller/controller/README.md §"What is installed, and is it current?" |