Files
felhom.eu/documentation/architecture/09-update-architecture.md
T
admin c85262111c
gates / gates (push) Successful in 23s
The update arc's two missing measurements, the lock, and the floor to 0.260.0
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all
down or blocked; both demo boxes SERVED.

Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and
two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live
compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s
after the cut decision and the seeded data read back intact — so the branch that is
one step from old-binary-on-migrated-database is now evidence, not argument.
Instrument limit stated: `starting` lasts under a second; all three landed in
`verifying`, which RecoverUpdates handles in the same branch.

Part 3 (R-611) — the night the previous session skipped without saying so. An app
updated with nobody pressing anything; a terminally-refused app was pressed exactly
once and never again over three passes. The unattended HOLD was NOT produced: the
within-a-major rule correctly refused the broken edge before it was attempted, so
Q4 still rests on the attended hold from slice 4. Said plainly rather than implied.

Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh
install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard),
R-614 (stale update phase survives a redeploy). R-520's pointer corrected.

Catalog: two drill pairs, both reverted; every image line byte-identical to
ff9717d3. The alpine:3.20 negative control a security review flagged is cleared.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 15:00:09 +02:00

62 KiB
Raw Blame History

09 — How an app update works, and what it is becoming

LIVING DOCUMENT. Every slice of the update arc updates this file in the same session. Opened 2026-09-02 with slices 1 and 2. Its absence was R-438: the update mechanism was chosen deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who did not know it was one.

This file carries the REASONING. The register (backlog/OPEN-ITEMS.md) carries the work. The source is the truth. Nothing here is invented: every mechanism claim is cited either to audits/SPIKE-app-update-2026-09-01.md, which measured it live, or to live source at file:symbol.


1. How an update works today, as measured

1.1 The button

Manager.UpdateStack (felhom-controller/controller/internal/stacks/manager.go:1199) is two compose commands and nothing else:

compose pull                       →  compose up -d --remove-orphans

No safety copy. No rollback. No hold. No verification. Confirmed by reading and across six live updates (spike §10 item 6). A pull FAILURE is handled correctly — UpdateStack returns after the failed pull and never reaches up -d, so the running app survives untouched (measured twice, spike §4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.

1.2 The catalog syncer moves the file underneath a deployed app

Syncer.copyTemplates (felhom-controller/controller/internal/sync/sync.go:319) copies docker-compose.yml and .felhom.yml into every stack folder on a 15-minute cycle (internal/config/config.go:351, default 15m). It has no deployed check of any kind. Its only guard is a sha256 content compare in copyIfChanged (sync.go:403) and its only exclusion is app.yaml. The post-sync hook is stackMgr.InjectMissingFields(updated) and nothing else — the sync does not restart anything.

That is why a deployed app's compose file and its running containers can disagree indefinitely. Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).

1.3 Thirteen other paths end in compose up -d

Excluding the three API actions, 13 call sites across 9 files call StartStack or RestartStack, and every one ends in compose up -d against the live compose file (spike §8 — the task that commissioned the spike said five; the count is thirteen). They include the boot reconciler (bootrecon.go:269), the app-stop guard's recovery (appstop_marker.go:283), the drive-return gate (intermediary.go:222), the quiesce restart-after-backup, off-site reconstitution, and every restore path.

So an upgrade can happen with nobody pressing anything — measured, spike §2 variant 1c-ii, where a boot reconciliation started an app on a newer image at 17:55:44Z.

1.4 One fear is measured SMALLER than it was stated

A plain power cut does not upgrade anything. Docker's own restart: unless-stopped puts the existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs up -d — it says so in its own words: no boot-orphaned apps (nothing to start) (spike §2, variant 1c, a positive observable and not an absent log line).

The unattended upgrade needs the narrower precondition: "and the app did not come back." Saying so is more useful than leaving the scarier version standing.


2. What was chosen, and by whom

Manager.RestartStack (internal/stacks/manager.go:1161) carries this comment, and it predates the whole arc:

"Use up -d instead of bare restart so that env vars from app.yaml are injected and any template changes (new images, healthchecks) are picked up. Plain docker compose restart only sends SIGTERM+start to existing containers without re-reading the compose file or env."

So the restart behaviour was chosen, deliberately, and written down. A design decision is not a defect. What was never decided — and is recorded nowhere — is what happens once the catalog syncer moves the file underneath a deployed app, and whether the choice was meant to extend to the thirteen unattended call sites. That gap is R-438, and it stays open: this document records the mechanism; it does not change it.


3. The operator decisions

These are rulings, not proposals. Anything specced against a different assumption is wrong.

2026-09-02

  1. The safety copy is a verified recent backup as a PRECONDITION — not a new copy invented for the update path. The guest-snapshot alternative is to be spiked before anything is designed around it. Context: the existing safety machinery (Manager.writeSafetyDump, internal/backup/offbox_reconstitute.go:207) is database-only, which is the headline of spike §6 — the file half was never priced, and demo-hp is too young a box to price it.

  2. The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is. A customer on the newest version is supported however old that version is. This is why catalog_since exists and why no version string is shown.

  3. Updates are automatic WITHIN a major, never ACROSS one. The cross-major case needs a human, because §4 says it cannot be undone.

2026-09-06 — Option 1: freeze the version, keep the fixes flowing. SHIPPED, v0.235.0

  1. An app's version is frozen to what the customer has, and only a deliberate Update moves it — while corrections to its definition keep arriving on the 15-minute cycle exactly as they do today.

    This is the ruling R-447 was blocked on, and it was blocked for a good reason: §2 establishes that RestartStack's use of up -d to pick up template changes was chosen and written down in its own comment. Reversing a chosen behaviour is a decision, not a bug fix.

    What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps". The old behaviour had two halves and only one of them was unwanted:

    half verdict
    a restart/repair silently changes the app's VERSION unwanted — nobody chose it, nobody is told, and §4 says it cannot be undone
    template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes worth keeping — both measured in the spike §3

    So the ruling keeps the second and removes the first. In the operator's own words: while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update.

    It does NOT make the Update button safer. That is slice 4 (R-448), and it is where the backup precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.

2026-09-13 — the database engine finishes its own conversion, and the upgrade test goes wide

  1. DBs should be updated when the app moves, with proper precautions, tests and backoff plans — the operator's own words, ruling on SPIKE-r459-mariadb-upgrade-2026-09-06.md. SHIPPED in the catalog the same day: every mariadb: sidecar (bookstack-db, kimai-db, nextcloud-db, romm-db) carries MARIADB_AUTO_UPGRADE=1; MARIADB_DISABLE_UPGRADE_BACKUP stays unset. Not an image change, so catalog_since does not move. The three precautions, because they are the real content of the ruling:

    1. Proven before it ships — upgrade-test.py re-ran E3 and E3b on the changed template and the engine-state field shows the conversion RAN (mariadb_upgrade_info reads the new version, the engine's own check says nothing further is needed, the entrypoint no longer prints skipped due to $MARIADB_AUTO_UPGRADE), with the seeded data reading back after. C3 still returns failed. Evidence: audits/r459-close-2026-09-13/.
    2. Watched as it lands — the change travelled the real 15-minute cycle to demo-hp: the live compose gained the setting, the sync recreated nothing, and one deliberate restart logged MariaDB upgrade not required with the app serving (same evidence directory).
    3. A rule until Slice 4 is built — the Update button still takes no backup, so no template may move a database-engine image across a major version until R-448 ships. Catalog CLAUDE.md states it; scripts/check-engine-major.py enforces it in the pre-push hook (the CI half cannot, R-452); its removal is tracked as R-469 so it is a deliberate act. The setting is inert until an engine major moves, and precaution 3 keeps it that way.
  2. The upgrade test goes as wide as possible, through the nightly unattended sessions — the ruling on STATUS.md item 11. Not "the ~25 database apps first": all of them, as the nightly rotation reaches them, one fixture per app through the app's own interface. Browser-only apps become reachable when CC runs on the operator's Windows workstation with Chrome — the claude-in-chrome route that DooPlex does not have — so an app recorded inconclusive for want of a headless seed route (bookstack's file half, R-460) is deferred to that venue, not faked.


2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts

  1. A floor carries a release past the vouched golden when the release's MinAgent is declared with it (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it, the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which minagent_header_gate.py now guarantees — goes into the same per-box agent comparison. An undeclared floor above the golden is still HELD, and both floor forms refuse to save one. Why: a controller image is pulled by tag and needs no golden to be delivered; what the floor was missing was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release delivery coexist. Proven live: both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s from the save, the hub logging SERVED … from declared (audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt).

  2. Any backup tier lets an app update (R-475; controller v0.239.0). The precondition takes the first FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site (Tier 3, bounded; unreachable = absent). update.backup_max_age applies to whichever tier is chosen. An app with nothing anywhere is backed up first; it is refused only when no backup can be taken either. The hold names the tier and the date. Tier 2 is required nowhere in the update path. This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a verified recent backup as a precondition, not a new copy) stands. Proven live: audits/rulings-r472-r475-2026-09-13/ 04 (nothing anywhere → backup first), 05 (Tier 1 alone), 07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).

  3. A bind-data app's route back is off-site before its own unit, and the hold says what the copy holds (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit that holds the definition and the database dumps and NOT the files (measured: gokapi restored from „helyi" came back with settings and no data). For such an app the update walks second drive → off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive → own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a customer is never sent to a copy that cannot bring the data back without being told so.

2026-09-21 — decided by CC unattended, operator may reverse

  1. A box AHEAD of the catalog reads „Naprakész", and the guarded Update refuses to move a pin backwards (R-524, controller v0.260.0). One sentence: when the catalog is reverted under a box that already updated, is that a "Frissítés elérhető"? Options: (a) leave it — the label compares for difference, as §5.4 says; (b) show „Naprakész" and let the button still run; (c) show „Naprakész" and refuse the button. Costs: (a) is free and offers a household a downgrade onto a datadir the newer version may have migrated, which §4 says cannot be undone; (b) removes the invitation but leaves the loaded gun; (c) costs one comparison and can, wrongly applied, block a legitimate update. Why (c): the direction was already settled — §3 decision 3 says a version change that cannot be undone needs a human, and this is one. The risk in (c) is bounded by making the Ahead verdict NARROW: every differing service must be orderable AND newer, or the answer falls back to today's behaviour. Reversible, no customer-data risk, and it only ever withholds an act. Implementation: stacks.CatalogOrder, one verdict read by both the badge and UpdatePreflight.

3b. OPEN — the seven questions Slices 6 and 7 need answered

These are questions, not rulings. CC does not decide them. Each is one answerable sentence, the options, what each costs, the recommendation, and what happens if nothing is decided. The measurement behind them is audits/UPDATE-ARC-STATE-2026-09-21.md; the short version is that 46 of the catalog's 58 exact pins are behind upstream today and 39 of those are within a major — the population §3 decision 3 already says may move without a human, and nobody presses 39 buttons.

Q1 — When may a box update itself?

May the box run the guarded Update by itself between 02:30 and 05:00, nightly?

option cost
02:30–05:00 nightly, after the backup legs the update leans on a copy made hours earlier the same night, which is the freshest the box ever has. The app is down for the health wait in the middle of the night.
a weekly window fewer interruptions; a box sits up to 7 days on a version the catalog already moved past, which widens the support window §3 decision 2 runs on
the household picks the window one more setting on a page that already has several, for a choice almost nobody will change

Recommendation: 02:30–05:00 nightly. The DB dump runs 02:30 and restic 03:00 on a demo box, so a window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy of the night without a new mechanism. If nothing is decided: Slice 6 cannot be built at all — every other question below is downstream of this one.

⚠ THE WINDOW CONTAINS 04:30, AND 04:30 IS WHEN THE BOX UPDATES ITSELF. Found 2026-09-21 (R-608) by reading the clock rather than by a failure: the controller self-updates daily at self_update.auto_update_time, default 04:30, and again from MaybeAutoUpdate after ANY hub report once a floor sits above the box — so at any hour, not only at 04:30. Either path restarts the controller container, which is the supervisor of a running app update.

What v0.261.0 now guarantees, so this question can be answered without also solving that one: the two cannot overlap in either direction. The controller defers its own swap while a guarded app update is in flight (retrying on the next report, exactly as it already did for a running backup), and UpdatePreflight refuses self_updating while a swap is in progress. The lock does not latch — a held app does not block the controller's updates for ever. So the window may contain 04:30; the two jobs will queue behind one another rather than meet. What it does NOT do is reorder them: if the operator prefers the box to take its own update first, that is a scheduling choice still open here.

Q2 — May an automatic update run on a bind-data app when no copy holds its FILES?

The button's rule and the automatic rule can differ. Should they?

The mechanism, verified at source this session, because an earlier draft had it backwards: the guard does not refuse these apps. Since v0.239.0/v0.241.0 (§3 decisions 8–9) the Update is refused only when no copy exists on ANY tier and none can be taken. For an app whose data is bind-mounted files, Manager.UpdateTierOrderFor (controller/internal/backup/update_guard.go:136-141) walks second drive → off-site → own unit last, and when the own unit is the copy chosen, UpdateCopyHolds (:145-165) ends the hold sentence with „csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem" — it holds the settings and the database and not the files. So the update proceeds, and the household is told what the copy holds.

With a human pressing, that is an informed choice. With nobody pressing, nobody was informed.

option cost
automatic requires a fresh copy that HOLDS THE FILES; the button keeps today's rule the nine file-leg apps (and any other bind-data app) update automatically only on a box with a second drive or off-site; on a one-drive box they wait for a person. Two rules to hold in one's head.
one rule for both — automatic follows the button simpler; a file-leg app can be updated unattended against a copy that cannot bring its files back, and the sentence saying so is read by nobody
automatic skips bind-data apps entirely simplest; the seven file-leg apps that are behind never move by themselves even when a good copy exists

Recommendation: the first. It is the smallest rule that keeps the promise the hold sentence makes. If nothing is decided: Slice 6 must be built for the safe subset only, and the file-leg apps stay manual — which is the third option by default, without anyone choosing it.

Q3 — What counts as "within a major" when the tag is not a version number?

§3 decision 3 says automatic within a major, never across. What about postgres:16-alpine, kimai/kimai2:apache-2.57.0, a date stamp, a digest?

And the test is per compose SERVICE, with ALL of them having to pass. An app bump that is minor while its mariadb: sidecar moves a major is ACROSS — that sidecar now converts the customer's datadir by itself (R-459), so the edge carries a migration whatever the app's own number says.

option cost
an unorderable tag on ANY service makes the whole edge ACROSS → human the 8 floating pins and every suffix-versioned image stay manual. Conservative, and it is the same rule v0.260.0's CompareImageRefs already implements and tests.
teach the comparator each shape every new shape is a new rule, and a wrong rule silently automates a major
compare digests instead needs Q6 first, and a digest carries no order at all — it can say "different", never "newer"

Recommendation: the first, reusing stacks.CompareImageRefs rather than writing a second rule. One small extension is needed and is named here so it is not discovered late: v0.260.0's CompareImageRefs answers orderable? and newer?, which is all R-524 needed. Slice 6 also needs same major?, so the parsed major has to be exposed from the same normaliser — an addition to the one comparator, never a second one. If nothing is decided: Slice 6 would have to invent a rule under time pressure, which is how a major gets automated by accident.

Q4 — A held app: who is told, when, and does the box try again?

An automatic update that ends HELD happened while everyone was asleep.

option cost
the household on the app page and by mail ONCE; the operator by event; NO retry until the catalog moves again or a person presses one mail per held app. The app stays down until someone acts — which is already true of a held update today.
retry the next night a broken edge takes the app down every night and mails every morning; the hold exists precisely because the box cannot fix it
tell only the operator the household finds their app down and has no sentence explaining it

Recommendation: the first. It is what the manual hold already does (settings.RestoreHold with reason: update_failed), plus one mail. If nothing is decided: the safe default is no automatic update at all, because a hold nobody is told about is worse than a version nobody moved.

MEASURED 2026-09-21, and the honest answer is that HALF of this is still unmeasured. The unattended night ran (audits/update-arc-gaps-2026-09-21/09-unattended-night.md). What it proved: an app updates itself end to end with nobody pressing anything; one app takes 51 s – 1 m 26 s including the health wait; the caller needs no new controller code, only the existing guarded Update plus UpdateRefusal.Reason on the wire (v0.261.0, R-609); and a terminally-refused app is pressed exactly ONCE and never again — four apps, three passes, proven.

What it did NOT produce is a HOLD, and the reason is instructive rather than a failure of the run. The only failing edge available was vikunja → alpine:3.20, and the caller correctly refused to attempt it: different repositories cannot be ordered, so the edge is "across" and belongs to a human by decision 3. The rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. Measuring the unattended hold needs an edge that PASSES the within-a-major test and still fails its health check — same repository, same major, a tag that starts and does not serve — which probably means a purpose-built image rather than a catalog move. So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).

Q5 — PostgreSQL: what has to exist before the catalog may move postgres:16 to 17?

Eleven templates, and the image performs no conversion — it refuses to start on an older major's datadir (R-463).

option cost
a scripted pg_upgrade edge in the harness, proven on all eleven, before the catalog may move real work: eleven fixtures, and pg_upgrade needs both major's binaries present. The engine-major gate keeps the rule until it exists.
move the pin and let the update HOLD honestly every one of the eleven apps goes down on the same night and comes back only by a restore
never move PostgreSQL majors the fleet sits on an engine that eventually loses upstream support

Recommendation: the first, and the gate stays until it lands. As of 2026-09-21 the engine-major rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of it and MARIADB_AUTO_UPGRADE=1); this half is exactly what stays. If nothing is decided: nothing breaks — the gate refuses the move — but the eleven apps drift further from upstream every month.

Q6 — Should the catalog record each pin's DIGEST at push time?

So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.

This is no longer theoretical. Measured 2026-09-21: six of the seven measurable floating pins have been repushed upstream since the catalog set them — postgres:16-alpine (8 apps), postgres:15-alpine, redis:7-alpine (6 apps), mariadb:11.4, mariadb:12.3, postgis:16-3.5-alpine. On demo-hp today, four apps read „Naprakész" over a database engine image that has demonstrably moved.

option cost
the catalog records the digest at push time; the box compares digests one field per pin. check-image-resolvable.py already resolves the digest, so the producer exists. §8.1's rule — the box never queries a registry — is untouched.
the box queries registries breaks §8.1 outright: a page that cannot render without eight upstream registries
leave it the badge stays right about the question it asks and wrong about the one a household hears

Recommendation: yes. It is the cheapest real improvement on this list and it closes R-446. If nothing is decided: „Naprakész" keeps meaning "the reference matches", which is measurably not what it sounds like.

Q7 — What does the hub's report need to carry for a fleet view?

Slice 7 lets the operator SEE and MOVE how far behind every box is.

Verified both sides this session: the controller's report payload carries name, state, CPU and memory and no image (controller/internal/report/types.go L98–103), and the hub's Store.SaveReport (hub/internal/store/store.go:965) denormalises only container counts. But the hub stores the raw report JSON whole, so a new controller field lands there the day it is sent — what is missing is the denormalisation and the page, not the transport.

option cost
per app: installed reference + catalog reference + badge state; the hub lists boxes behind, with a "move" that is the same guarded Update, operator-triggered additive on both sides; the report grows by a few fields per app
badge state only smaller payload; the operator cannot see WHAT is behind, only that something is
leave it to per-box pages free today at two boxes; unusable at twenty

Recommendation: the first, and it stays P3-LOW until the fleet grows. If nothing is decided: the only way to answer "is the fleet current?" is what this session did — read both boxes' files by hand.


Not a question — already ruled

R-462's scope was decided on 2026-09-13. §3 decision 6: the upgrade test goes to all apps through the nightly rotation, explicitly not "database apps first". The register row R-462 still says "VIKTOR rules on scope" — that row is stale and is corrected to cite decision 6. The update night below proposes an ORDER inside that ruling; it does not reopen it.

4. The vocabulary ruling — "rollback" is struck

App data CANNOT be rolled back. Measured on Nextcloud (spike §7): once a migration has actually run, putting the old image tag back produces a container that refuses to start —

"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"

with a positive control proving the data is intact, only unreachable by the old version (§7 6d).

So "rollback" must not appear in any spec for this arc. The two shapes actually available are:

shape when it applies what it does
ABORT before anything migrated stop, put the old image back, the app runs again
RESTORE FROM A COPY after a migration ran the data restore is the whole remedy

There is no third.

4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades

SPIKE-upgrade-test-2026-09-06.md upgraded three real apps with real data in them and then attempted the abort on each. All five real catalog upgrades kept the customer's data. The abort did not behave the same way twice:

app abort why
docmost 0.25.3→0.95.0 REFUSES the old code finds migration ledger entries it does not know: "corrupted migrations: previously executed migration 20260213T085259-notifications is missing", then "Failed to run database migration. Exiting program."
privatebin 1.7.5→2.0.5 works file-backed, no database, no schema — a major version moves no data
bookstack app+engine works, misleadingly only because the MariaDB datadir upgrade was skipped and never happened — see R-459

This puts TWO independent measurements behind the ruling above, by two unrelated mechanisms: Nextcloud refused on an explicit version comparison; docmost refuses on its migration ledger. The word "rollback" was already struck; it is now struck on evidence rather than on one case.

And it adds a distinction this document did not have: there is no single answer to "can this update be undone". There are apps where it can and apps where it cannot, and the only way to know which is to MEASURE THAT APP. Any design that assumes one answer for all 53 is designing against a fact that was checked and is false.

Per §5 below, a restore's image-level undo used to have a ≤15-minute half-life because the syncer overwrote it (R-441) — closed in v0.235.0.


5. The shape, as SHIPPED in v0.235.0

The live docker-compose.yml is DERIVED from a pin recorded in app.yaml — the one file the syncer never touches. The catalog proposes; app.yaml decides; the compose file is an output rather than an input, and the thirteen unattended up -d paths stop being able to change a version by accident.

5.1 Nothing was added to the thirteen call sites, and that is deliberate

They are made safe by removing the reason, not by gating them. The most important of them are REPAIRS — the boot reconciler (bootrecon.go:269), the drive-return gate (intermediary.go:222), the app-stop guard (appstop_marker.go:283). A repair path that refuses to repair leaves a customer's app down, which is worse than the problem this slice solves. Since the file they act on no longer changes version, every one of them became safe without being touched.

5.2 The pin, and what it is not

AppConfig.PinnedImages (app.yaml, pinned_images:), service → image ref.

It is NOT InstalledImages. That field is an OBSERVATION — what containers report. This one is a DECISION — what should run. Letting an observation feed a decision would make a bad reading become a bad deployment, which is the category error desired_state exists to avoid (R-166), one field over. They will normally agree; when they disagree that is a signal, not a bug to paper over.

Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.

Beside it, applied-compose.yml in the stack directory stores the exact definition that pin came from. Syncer.copyTemplates copies exactly docker-compose.yml and .felhom.yml, so that name is safe from the catalog, and keeping it beside the app means it travels with every path that already moves a stack dir.

5.3 The four writers — the only acts entitled to move a version

writer pin source
the deploy path (runComposeDeploy) the template just deployed from
UpdateStack the catalog's current template, written BEFORE the pull
the restore (stackAdapter.RecreateStackDefinitionFromUnit) the recovery unit's captured compose — this closes R-441
Manager.AdoptPins the observation, once, and only when complete AND matching

UpdateStack's ordering is load-bearing, not stylistic. compose pull and up -d act on the file on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards would pull the frozen version and change nothing — while reporting success, and a button that lies is worse than a button that refuses. A failed pin write REFUSES the update, which is the opposite of recordInstalledImages and for the same reason desired_state refuses: this field is intent.

5.4 The render table, complete

app state result
not deployed / protected / seam not wired the catalog template — today's behaviour
deployed, unpinned the catalog template + one DEBUG
deployed, pinned, catalog images equal the catalog template — fixes flow, self-healing works
deployed, pinned, catalog images differ the stored applied definition — frozen WHOLE
pinned, differ, nothing stored the catalog template + one WARN. We cannot freeze what we do not have and must not invent it
mid-deploy the compose file is left alone this cycle

.felhom.yml is copied verbatim in every case — it carries no image, and it carries catalog_since, which the badge needs. See §8.5.

The frozen branch writes a WHOLE file and never a substitution. Taking the new template and putting the old refs back creates a third state nobody chose: wger 2.6 needs a full DB configuration the older template cannot supply, so an old image under a new template is broken in a way neither version is.

And this is not "skip deployed apps". That option was considered and rejected: it also stops health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in the spike §3 — both halves the ruling explicitly kept.

5.5 Adoption, and why the startup order matters

AdoptPins runs once at boot, immediately after BackfillInstalledImages, and pins every deployed app to what it is already running. It reads and writes files only — no container is started, stopped or touched. It skips, loudly, when the observation is incomplete or when the app runs something the current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a guessed pin.

syncer.Start() was moved to after adoption. It fires an immediate sync; at its previous position that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and handed the next restart a version change — the exact behaviour this slice removes, once per boot.

5.6 The trap this slice set for the previous one

Stack.TemplateImages is read from the app's live compose file — which is now the RENDERED one. On a frozen app that file names the OLD version, so web.compareInstalledToTemplate would find installed == template and answer „Naprakész" on exactly the apps that are behind — with every test still green, because the new field has the same type and shape. The badge now reads Stack.CatalogImages, taken from the syncer's own git clone. A feature that silently inverts an earlier feature is the failure mode to look for whenever a file changes meaning.


6. The seven slices

# slice status
1 The box records what it actually installed — app.yaml.installed_images, per compose service, ref + digest + first-seen. SHIPPED, controller v0.233.0 (2026-09-02)
1b Seed the record for apps nobody touches — a startup backfill, so the label is not restricted to apps that happen to get restarted. SHIPPED, controller v0.234.0 (2026-09-03)
2 One badge says whether the app is current — „Naprakész" / „Frissítés elérhető — N napja", from catalog_since. No version number. SHIPPED, controller v0.233.0 + catalog 69761cf (2026-09-02); English since v0.258.0 (R-589); a FOURTH verdict — AHEAD — and the downgrade refusal in v0.260.0 (R-524, §3 decision 10)
3 The compose file becomes DERIVED — the pin in app.yaml wins; the syncer renders instead of copying. SHIPPED, controller v0.235.0 (2026-09-06) — operator ruling §3.4
4 A guarded update — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8) — §6.1
5 An upgrade test that runs again — a harness that upgrades a real app with real data in it and asks the app for the data back. SHIPPED, app-catalog/scripts/upgrade-test.py (2026-09-06) — 7 edges, 3 apps; see §4.1 and §10
6 A version sequence — updates automatic within a major, a human across one; an engine change gets its own edge. OPEN — R-450
7 A fleet sweep pipeline — the operator can see, and move, how far behind every box is. OPEN — R-451

6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)

POST /api/stacks/{name}/update is a guarded job. It answers 202 at once; the outcome exists only on GET /api/stacks/{name} (updating, update_phase, update_phase_label, update_error, hold_reason), and update_phase=done is written only after the app's health is known. R-443 is closed by construction: nothing reports an update complete on the compose exit code.

The sequence, and the order is the design:

# phase what happens on failure
0 refusals (409, before the intent is recorded) held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's memoryVerdict, releasing the app's own request), disk (fixed 2 GB floor — image size unknown without a registry), no copy on ANY tier and no backup can be taken now (since v0.239.0; before it, no restorable Tier-2 copy) nothing moves, nothing is recorded
1 checking walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than update.backup_max_age (v0.239.0) nothing moves
2 backing-up — only when no tier holds a fresh copy RunAppBackupNow: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 refused with the backup's own error; nothing moves
3 safety-dump WriteUpdateSafetyDump (R-361's undo copy) — before the pin moves refused; nothing moves
4 pinning the previous definition is copied aside and journaled, then the pin advances pin put back
5 pulling compose pull pin and definition PUT BACK — nothing ran (Scenario E)
6 starting compose up -d --remove-orphans stop + HOLD
7 verifying the .felhom.yml health check through the existing probe, or 60 s of every container running and none restarting; bounded by update.health_timeout stop + HOLD; the pin STAYS — the migration may have run (Scenario F)
8 done installed images recorded, journal cleared —

The two knobs (controller.yaml, operator-owned): update.backup_max_age (default 24h) and update.health_timeout (default 5m).

The precondition is the existing verified backup, not a new copy (§3 decision 1). It is backup.Tier2UnitRestorePoint — the SAME predicate that permits the destructive „Teljes visszaállítás", extracted from the backups page rather than copied. The copy is aged by the last SUCCESSFUL Tier-2 copy, not by the unit manifest's created_at, and that was measured before it was designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the manifest, a quiet app would be "stale" forever and a backup-first would not fix it. The predicate is Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475). SUPERSEDED in v0.239.0 by §3 decision 8: backup.Manager.UpdateRestorePoints walks all three tiers and the update leans on the first fresh copy; Tier2UnitRestorePoint is still the Tier-2 half and still the page's predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit proven current. The hold stores copy_tier and names „második meghajtó" / „saját meghajtó" / „távoli mentés"; a successful off-site restore now lifts an update hold too.

The hold is settings.RestoreHold with reason: update_failed and copy_date — the SAME store and gate as R-379, so every start path that already honoured a restore hold honours this one. A successful unit restore lifts an update hold (only that kind). Three unattended paths honoured no hold before v0.237.0 and now do: the drive-return gate's restart and boot recreate, and the nightly volume dump (which ends in StartStack). The nightly capture and Tier-2 run skip a held app, so the restore point the hold text names is never overwritten.

Crash safety is a journal, <data>/update-journal.json, written before every phase. RecoverUpdates runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back; after up → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by ResumeInterruptedUpdates once the backup side is wired, ending healthy or held.

THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT. Whether an old image starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot be predicted. So the box never puts the old version back by itself. The route back is the restore, and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where the harness has proven it, is slice 6's.

Not gated here: a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.

The release could not reach the fleet by floor — R-472. The hub holds a controller floor above the vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0 were hand-deployed to the demo guests. RESOLVED by §3 decision 7 (hub v0.112.0): v0.239.0 reached both demo boxes by the floor alone, with its MinAgent declared.

Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not only once it is held. In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit. The Tier-2 mirror the hold names survived only because Tier 2 is daily. backup.Manager.isHeld now also answers true for an app a guarded update is moving (SetUpdatingCheck).

v0.240.0 (2026-09-13, evening) — what the afternoon's proof and the first nightly rotation found, fixed. Seven rows: removal with backups kept now keeps the Tier-2 RECORD, so the second-drive restore is not refused over an intact mirror (R-486, P1 — the disaster the second copy exists for); PostGIS/pgvector/ TimescaleDB images are Postgres, so such apps get their logical dump (R-484); "delete backups" deletes the unit, the mirror(s) and the prefs (R-474/R-466); the backup card sizes them (R-485); a held update's sentence leaves the card with the hold (R-480); the Tier-3 lookup is one snapshots call (R-477); a unit older than the app's deployed_at does not count (R-478). Delivered by the floor in 16 s / 18 s; every row proven live with a throwaway adventurelog. audits/v0240-2026-09-13/.

Proven live on demo-hp, 2026-09-13, with a throwaway uptime-kuma and real catalog tag changes (each reverted in the same phase): A (2.3.2→2.4.0, done after health), B (backup_max_age: 2m → backup first), E (non-existent tag → pin back, container untouched), F (alpine:3.20 → held), H (three buttons and the boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold cleared). Live evidence: audits/slice4-2026-09-13/.

The verdict record — the contract Slice 6 carries

Decided here rather than invented twice. The harness writes one of these per edge, beside its evidence; Slice 6 puts the same shape in the catalog.

{"harness_version": 1, "app": "bookstack",
 "from": {"bookstack": "…:25.02.2", "bookstack-db": "mariadb:11.6"},
 "to":   {"bookstack": "…:26.05.2", "bookstack-db": "mariadb:12.3"},
 "verdict": "proven | failed | inconclusive",
 "seed_read_before": true, "seed_read_after": true, "healthy_after": true,
 "migration_observed": "verbatim log line, or null",
 "abort": "starts-and-serves | refuses | starts-data-gone | not-attempted",
 "abort_detail": "the refusal quoted verbatim, or null",
 "duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"}

inconclusive is a first-class verdict and must never be collapsed into failed. "We could not measure it" and "it does not work" are different facts, and only one of them is about the app. migration_observed is a quoted line, never an inference from timing — the value of both the Nextcloud and the docmost findings was the exact sentence the app printed.

Database engines under an upgrade — MEASURED 2026-09-06

The arc's standing rule that an engine change gets its OWN edge now has measured evidence behind it, and the evidence is stronger than the rule's original argument. The rule was justified by "two migrations behind one edge is an unreadable failure when it breaks" — a readability argument. What was measured is that an engine change can be applied and silently NOT happen, which the app-half edge cannot produce and which no amount of readability would have surfaced:

  • SPIKE-upgrade-test-2026-09-06.md §4 — MariaDB 12.3 starts on an 11.6 datadir, logs that the conversion it requires was skipped, and serves. Assigned to the engine half by decomposition: the app half alone produces no such line.
  • SPIKE-r459-mariadb-upgrade-2026-09-06.md — it is stable but never self-resolving (5 of 5 restarts, no degradation, and the engine says Check required! every time, forever). Converting properly succeeds, costs 7 s, takes its own system-database backup, and does not cost the ability to abort. The trade that was expected here does not exist.
  • 2026-09-13 — the setting is in the catalog. All four mariadb: sidecars carry MARIADB_AUTO_UPGRADE=1 (operator ruling, §3 decision 5), and upgrade-test.py's engine-state field now shows the conversion RUNNING on the bookstack edges. And a gate holds the engines inside their major until Slice 4: app-catalog-felhom.eu/scripts/check-engine-major.py (R-469).

Two rules for anything this arc builds around a database engine:

  1. Ask the engine, not the log. MariaDB's entrypoint prints MariaDB upgrade not required on an unsupported downgrade; mariadb-upgrade --check-if-upgrade-is-needed names it exactly (R-464). A cheap instrument built on the log line would report "fine" for the broken case.
  2. The two engines fail in opposite directions, so one check will not do. MariaDB starts anyway and skips quietly; PostgreSQL refuses to start on a datadir from an older major, and the image performs no pg_upgrade. Eleven templates carry PostgreSQL and eight sit on postgres:16-alpine (R-463).

And an engine-state field belongs BESIDE a verdict, never inside it. upgrade-test.py reports engine_state_after next to verdict, because an unconverted datadir is not known to be a failure and a verdict that said so would encode an unproven judgement.

The rule slice 6 inherits, recorded now while it is cheap: an engine change gets its own edge, never bundled with an app version bump. bookstack moved the application and MariaDB 11.6 → 12.3 in one commit (0b73e5e); that is two migrations behind one edge, and an unreadable failure when it breaks.


6.2 Slice 6, as it would be built (OPEN — R-450; needs Q1–Q4)

Not a design yet; the shape the seven questions bound. Written down so the answers have somewhere to land.

Nothing new happens to the app. Inside the window, for an app that qualifies, the box runs exactly the guarded Update of §6.1 — same precondition, same safety dump, same pin journal, same health wait, same hold. Slice 6 adds a caller, not a path. That is the whole reason it is affordable: every failure mode was measured in slice 4 and every one of them already ends in a hold the household can read.

An app qualifies when ALL of these hold (each clause is a question above, not a decision taken):

  1. the window is open (Q1);
  2. stacks.CatalogOrder says Behind — never Unknown, never Ahead (v0.260.0 gives all four);
  3. the edge is within a major for EVERY compose service, the engine sidecar included (Q3), judged by stacks.CompareImageRefs — one unorderable service makes the whole edge across;
  4. a fresh copy exists on a tier that holds what this app's data actually is (Q2);
  5. the app's own switch is on (Q2's default: on).

What the household sees. An event and a line on the app page's timeline, before and after, in both languages: „Automatikus frissítés 03:12-kor — sikeres" / „— megállítva, a másolat 2026-09-20-i". A held app is not retried until the catalog moves again or a person presses (Q4).

Where it would live. A scheduler beside the existing nightly legs, reading settings for the window and the per-app switch, and calling Manager.StartGuardedUpdate. It must respect the same isHeld/SetUpdatingCheck interlocks v0.238.1 added — the nightly capture running inside an update's health wait is the defect that release fixed, and a second unattended caller is exactly the shape that finds it again.

The in-process caller reads UpdateRefusal.Reason, and the split is measured, not assumed (v0.261.0, R-609): busy, updating, deploying, migrating and self_updating are transient — try again on the next pass; held and downgrade are terminal — never press that app again until a person acts; memory, disk and no_backup need a person and should be surfaced, not retried. Before the reason reached the wire the only safe readings were "give up on everything" or "press for ever", which is why this is listed as a dependency of the slice rather than a detail of it. A working caller in this exact shape exists as evidence, not product: audits/update-arc-gaps-2026-09-21/unattended-caller.py.

Ships behind app_update.unattended: false with no UI until the operator answers Q1.

⚠ NOT auto_update — that name is TAKEN, and by the very thing this must not collide with. self_update.auto_update / self_update.auto_update_time (config/config.go L280-281, default 04:30 at L422) are the CONTROLLER's own update. An earlier draft of this section said Slice 6 "ships behind auto_update: off"; two settings with that name, one meaning the controller and one meaning apps, is the kind of collision that is only discovered by an operator who turned off the wrong one. The app-scoped key is app_update.*.

6.3 Slice 7, as it would be built (OPEN — R-451; needs Q7)

Three additive pieces, and the transport already exists (§3b Q7):

  1. controller — the report's per-app object gains installed reference, catalog reference and badge state. Additive; an older hub ignores it.
  2. hub — denormalise those out of the raw report it already stores whole, and list boxes by how far behind they are.
  3. hub → box — a "move" button that is the same guarded Update, operator-triggered, through the existing command path. Not a second update mechanism, and not automatic.

Rank stays P3-LOW at two enrolled boxes. It rises with the fleet, and §2 of the state audit is what that looks like today: the only way to answer "is the fleet current?" was to read both boxes' files by hand.

6.4 The update night — a drill brief outline, costed from R-462's real numbers

The ruling is decision 6: all 53 apps, through the nightly rotation. This is an ORDER inside that ruling, not a scope change. The database apps go first because they are the ones where a wrong answer costs data rather than uptime.

The real numbers this rests on (R-462, measured 2026-09-06): a successful edge takes 6.4 s – 305.1 s, median 71.8 s; a FAILING edge takes 556 s, roughly 8×, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost 5.07 GB. Machine time is not the cost — fixtures are. Two of the three apps needed a bespoke non-browser seed route, one needed two attempts and a discarded approach, and one (bookstack) can only ever be half-proven headlessly (R-460).

leg what cost
A the 15 database services — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus adventurelog's postgis — one edge each, fixture per app 15–25 CC-hours, dominated by seed routes; ~30 min machine time at the median; ~25 GB
B one power cut mid-update DONE 2026-09-21 (R-610) — measured THREE times, two cut mechanisms, three apps: pulling (R-520) and the dangerous post-start case three times over. All ended honest; vikunja's 2.6.0 migration had already run when the power went and the data read back intact. What remains: a cut landing inside starting itself (it lasts well under a second; needs an in-process fault injector, not a faster shell) 0 — spent
C one PostgreSQL pg_upgrade rehearsal, the Q5 edge, on one app before any of the eleven 3–4 CC-hours
D one downgrade refusal already done — v0.260.0, proven live 2026-09-21
E the automatic night MOSTLY DONE 2026-09-21 (R-611) — the success night and the no-retry proof both measured. What remains: the unattended HOLD, which needs an edge that passes the within-a-major test and still fails health (see Q4) ~1 CC-hour + a purpose-built image
F the remaining 38 apps, through the nightly rotation as decision 6 directs ~1 app/night; fixtures amortised

Total for legs A–E: roughly 21–34 CC-hours, plus ~25–30 GB of images on a scratch host. Legs C and E are the ones that unblock a decision; leg A is the one that takes the time.

Venue: a scratch host, never a customer box — demo-hp's guest 9202 for the box-side legs, the harness on DooPlex for the image-side ones.

7. What slices 1 and 2 actually built

7.1 The record (slice 1)

Manager.recordInstalledImages (felhom-controller/controller/internal/stacks/installed.go) runs after a successful compose up from StartStack, RestartStack, UpdateStack and runComposeDeploy, and writes app.yaml:

installed_images:
  web:
    ref: lscr.io/linuxserver/bookstack:26.05.2
    digest: sha256:…              # "" if the image was never pulled from a registry
    at: "2026-09-02T18:41:03Z"    # when this ref+digest was FIRST seen for this service

Three rules, each with its reason:

  • It reads the CONTAINER, never docker-compose.yml. §1.2 is why: that file is the value that has already moved. A record built from it would answer "what will happen next time something runs up -d", which is a different question.
  • A failed write NEVER refuses the action — deliberately the opposite of SetDesiredState. Intent refused, observation logged. Refusing to start a customer's app because we could not write down which version it is trades a real outage for a bookkeeping gap.
  • It is NOT called from StartStackServices — the R-47 DB-only restore window would overwrite a complete record with a partial one.

Seeded at startup (v0.234.0). Manager.BackfillInstalledImages runs once at boot, beside the desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY on. It only READS containers. This was not a refinement — without it the feature did not reach a quiet box at all: see §8.3, which was written as a known limitation on 2026-09-02 and was a defect by the next morning.

Two admission rules, and the second is the design:

  • It never overwrites an existing record. The bring-up paths own updates; this fills gaps only.
  • It refuses to seed a PARTIAL observation. §7.2's comparison reads a service-count mismatch as BEHIND, so a degraded or crash-looping app seeded from its visible containers would render „Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial because they follow a SUCCESSFUL up -d, where a gap is real news and is logged; a backfill meets a box in whatever state it is in. Same field, two writers, two different admission rules — that is deliberate and must not be "made consistent".

Nothing reads it to take a decision. Slice 2 reads it to render a label.

7.2 The label (slice 2)

web.updateBadge (internal/web/updatebadge.go) compares the recorded reference per service against what the current template pins, and renders through the existing meta_badge partial — no new markup, no new CSS.

Absent means UNKNOWN and never means current. Every app.yaml written before v0.233.0 has no record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by a test with a companion red-proof.

No version number reaches the customer (operator ruling: a household cannot act on 26.05.2). Version strings stay in the logs, the API and the hub.


8. Known limitations, stated plainly

  1. „Naprakész" can be FALSE for the floating pins, and 2026-09-21 measured HOW false. The comparison is reference-to-reference and queries no registry — a customer's box must not depend on reaching eight upstream registries to render a page. For postgres:16-alpine, mariadb:11.6 and the others the reference can be identical while the image behind it has moved. NUMBERS, 2026-09-21 (audits/UPDATE-ARC-STATE-2026-09-21.md §3.3). The count this document carried — "23 of 66" — is STALE and matched no definition the catalog supports today. Recounted at catalog 18a6d2d8, with the definition stated so it can be rechecked: a pin FLOATS when its tag names a version LINE rather than an exact release. Of 66 unique pins, 48 are full X.Y.Z, 6 are two-part lines (mariadb:11.4/11.6/12.3, claper:2.5, opengist:1.13, wger/server:2.6) and 4 are major lines (postgres:15-alpine, postgres:16-alpine, redis:7-alpine, postgis:16-3.5-alpine) — 10 float. The remaining 8 are exact versions wearing a variant suffix (ghost:6.53.0-alpine, nextcloud:34.0.1-apache, …), which do not float by this definition. Of the 8 database and cache engine pins the sweep measured, 7 were measurable and 6 have been repushed upstream since the catalog set them — postgres:16-alpine (8 apps), postgres:15-alpine, redis:7-alpine (6 apps), mariadb:11.4, mariadb:12.3, postgis:16-3.5-alpine. Only mariadb:11.6 has not. The 8th, immich's own ghcr build, is UNMEASURED — ghcr exposes no anonymous last-modified. So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved. Digest-level comparison needs the catalog to record the digest at push time — R-446, put to the operator as §3b Q6, recommended YES.
  2. Nothing enforces catalog_since. Enforced by the pre-push hook since 2026-09-13 (R-452, app-catalog-felhom.eu/scripts/check-catalog-since.py); CI's shallow clone still skips it out loud. A commit that moves an image: line and forgets the date under-reports how far behind a box is. The gates runner fetches at --depth 1 and has no parent commit to diff against, so the gate needs a deeper fetch — R-452.
  3. The record only appears after the next lifecycle action. CLOSED in v0.234.0, and the way it closed is worth keeping. This was written on 2026-09-02 as an accepted limitation — "the fleet view fills in gradually". The operator looked at demo-felhom the next morning and found OpenGist, up 15 hours, running exactly the catalog pin, showing nothing at all. On a quiet box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped. BackfillInstalledImages now seeds the absences at startup by reading containers (§7.1). The residue that stays: the seed happens at controller START, so a box between upgrade and its next restart still shows nothing — bounded by one restart rather than unbounded.
  4. A frozen app is frozen WHOLE. While the catalog is ahead, no template correction reaches that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the wger 2.6 hazard, and it is the right trade: a new template around an old image is a third broken state. Recorded so it is a choice, not a surprise.
  5. .felhom.yml keeps flowing while the compose file is frozen — the deliberate asymmetry in §5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded. The failure direction is a false alarm, never data loss, and freezing .felhom.yml would break the update badge by withholding catalog_since. R-458.
  6. The Update button is still unguarded. CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1). It refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a safety dump, and holds an app that does not come up. What stays true: it still has no automatic rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) — that now ends HELD rather than crash-looping behind a green button.
  7. An engine major can be applied without its datadir upgrade, and nothing notices. CLOSED 2026-09-13 for MariaDB (R-459): every mariadb: sidecar carries MARIADB_AUTO_UPGRADE=1, and the harness shows the conversion running on the E3/E3b edges (§3 decision 5). What stays true: the PostgreSQL half (R-463) has no equivalent — the image performs no pg_upgrade — and the engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until Slice 4 gives the Update button a backup.
  8. Only three of 53 apps have ever had an upgrade measured, and one of them (bookstack) can only be half-proven headlessly (R-460). The widening is R-462, costed with real numbers.
  9. The hub does not record image tags at all. Its report's container payload carries name, state, CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side change; it is not derivable from what is already reported.

9. Where the rest lives

what where
the measurements this document rests on audits/SPIKE-app-update-2026-09-01.md
the work backlog/OPEN-ITEMS.md — R-438..R-445, R-446..R-452
the syncer, described accurately but without the consequence architecture/02-controller-module-map.md
what the lifecycle actions are proven to do architecture/00-capability-map.md
the implementation felhom-controller/controller/README.md §"What is installed, and is it current?"