slice 3 docs: the ruling, the shipped mechanism, and four rows closed
gates / gates (push) Successful in 17s

09-update-architecture.md gains the fourth dated operator ruling (2026-09-06,
Option 1) and its section 5 is rewritten from a proposed shape into the shipped
one: the pin, the stored definition, the render table, the four writers, the
startup ordering, and the trap this slice set for slice 2 - the live compose file
is now the frozen one, so a badge comparing against it would answer Naprakesz on
exactly the apps that are behind.

02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being
true today, so it is corrected, and the two sections describing the old seam now
carry a banner saying they describe v0.234.0 and below - kept because every box
under v0.235.0 still behaves that way and because they are the measured account
of why it changed.

R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458
opened for the .felhom.yml asymmetry, with what would settle it by measurement.

Live evidence: two real catalog pushes travelling the real 15-minute cycle, both
reverted, the tree byte-identical afterwards. The restart that used to take 18.3
seconds and pull a new image now takes 0.1 seconds and pulls nothing.
This commit is contained in:
2026-09-06 10:37:33 +02:00
parent bc47dd4ef9
commit 417df06f35
9 changed files with 411 additions and 76 deletions
@@ -100,11 +100,12 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|---|---|---|---|---|
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. What they do to app DATA is now measured too, and it is a separate row-worth of facts (below).** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3`; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller | **PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
| **What VERSION a box is running, and whether it is behind the catalog** | controller **v0.234.0** + catalog `69761cf` | **PROVEN-LIVE — both the record and the rendered badge.** | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Badge evidence: same file §4. **v0.234.0 startup backfill, PROVEN LIVE 2026-09-03 on demo-hp:** the record was stripped from `privatebin` (1 service) and `romm` (3 services) to recreate the pre-0.233.0 shape, the controller restarted, and the backfill re-seeded **exactly** those two — every digest matching the ground truth read from the containers beforehand — while logging `2 app(s) recorded, 7 already had a record, 0 left unrecorded`. **All 9 deployed apps then carried „Naprakész" on `/stacks`** (ASCII fragments with a negative control at 0). On demo-felhom the operator's own case, OpenGist, now renders the badge. **Why the backfill exists at all: without it the label never reached an app that simply runs**, which the operator found the morning after v0.233.0. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: one badge STATE of four.** „Frissítés elérhető" WITHOUT an age needs an app whose `catalog_since` is absent, malformed or future-dated, and all 53 now carry a valid one — unit-tested only (`TestGroupF`). The other three are live: „Naprakész" ×2 on `/stacks` and on `/apps/bookstack`; **NO badge at all on `/apps/docmost`**, a deployed app with no record — absent is UNKNOWN and is not rendered as current; and „Frissítés elérhető — **52 napja**" on both surfaces, the age being real arithmetic on bentopdf's `catalog_since` 2026-07-12. **The behind state was staged by editing bentopdf's compose tag ONLY** — no container restarted, no `up -d` — then reverted byte-identically (`sha256` equal, `diff` empty); so the RENDER is measured and the syncer's own half stays measured separately in the spike. Searched with `grep -oF` ASCII fragments plus positive AND negative controls. **The Frissítés/Újraindítás/Leállítás buttons are unchanged on the live page in the behind state.** **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 |
| **An app's VERSION is frozen to what the customer has; only a deliberate Update moves it, while template CORRECTIONS still arrive** | controller v0.235.0 | **PROVEN-LIVE (2026-09-06)** — by two REAL catalog pushes travelling the REAL 15-minute cycle, not a hand-edited file | **`tests/VALIDATION-update-slice3-2026-09-06.md`** — a non-image catalog change REACHED the pinned app (08:01:51Z) with the container untouched; an image change did NOT (08:20:29Z); and the restart afterwards took **0.1 s, did not recreate the container, and never pulled the new image**, against the spike's **18.3 s with a pull** for the identical sequence before. The Update button still moved the version (pin advanced 17 s BEFORE the pull completed) and the teardown update returned the container to the baseline digest byte for byte. The freeze holds in BOTH directions. Unit side: `internal/sync/render_test.go` (the whole render table, incl. self-healing in BOTH branches), `internal/stacks/pin_test.go` (adoption never guesses; the update advances the pin BEFORE the pull), `TestGroupG` (the badge reads the catalog, not the frozen file), `TestGroupH` (an AST walk of `cmd/controller/main.go` asserting the seam, adoption, and their ORDER against `syncer.Start()`). **Three companion red-proofs**, each run, failing, and reverted. | Operator ruling 2026-09-06, Option 1 (`architecture/09-update-architecture.md` §3.4). **Nothing was added to the thirteen `compose up -d` call sites** — most are repairs, and a repair that refuses to repair leaves an app down; they were made safe by removing the reason. `pinned_images` is INTENT, `installed_images` is an OBSERVATION — never fed from each other (R-166, one field over). **Known limitations, all recorded rather than fixed:** a frozen app is frozen WHOLE (§8.4); `.felhom.yml` keeps flowing, so a frozen app can get a probe for a newer version — false alarm, never data loss (**R-458**); and **the Update button is still unguarded** (R-448 is slice 4). Closes R-447, R-441, R-438 |
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
| App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven |
@@ -335,7 +335,7 @@ own; every caller that is not the customer must decide for itself whether the ap
### `sync/`
| File | Class | Reason | Risk |
|---|---|---|---|
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). **It copies into EVERY stack folder, deployed or not — see "the app-definition seam" below.** | clean |
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset). **Since controller v0.235.0 it RENDERS `docker-compose.yml` rather than copying it** — verbatim while the catalog still offers the app's pinned version, from the app's stored `applied-compose.yml` once the catalog moves past it. `.felhom.yml` is still copied verbatim always, and `app.yaml` is still never touched. Reasoning: `09-update-architecture.md` §5. | clean |
### `system/` — split per-function (not per-file)
| File | Class | Reason | Risk |
@@ -518,6 +518,14 @@ own; every caller that is not the customer must decide for itself whether the ap
Until this was measured, no architecture document said what happens here, and the gap itself is
R-438. The three facts below are the ones a reader needs before touching any of it.
> **⚠ SECTIONS 1 AND 2 DESCRIBE THE BEHAVIOUR UP TO CONTROLLER v0.234.0. Controller v0.235.0
> (2026-09-06) CHANGED IT, on an operator ruling.** They are kept because they are the measured
> account of why it was changed, and because every box below v0.235.0 still behaves this way. **What
> ships now: `09-update-architecture.md` §5.** In one sentence — an app's VERSION is frozen to what the
> customer has and only a deliberate Update moves it, while template CORRECTIONS and the self-healing
> below still arrive on the 15-minute cycle. Nothing was added to the thirteen call sites in §2; they
> were made safe by removing the reason.
### 1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle
`Syncer.copyTemplates` (`sync/sync.go:319`) walks every directory in the catalog cache and copies
@@ -526,6 +534,11 @@ the app is deployed.** The only guard is a sha256 content compare (`copyIfChange
exclusion is `app.yaml`. Interval is `git.sync_interval`, default `15m` (`config/config.go:351`), plus
one immediate sync at controller start (`sync.go:98`).
**SINCE v0.235.0** this walk still happens and `.felhom.yml` is still copied unconditionally, but the
compose file goes through `Syncer.renderSource`, which consults a per-app plan supplied by the stack
manager (`Manager.RenderPlanFor`) through a nil-safe seam. A nil seam is byte-for-byte the behaviour
described above.
**It restarts nothing.** The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing
else. So from the moment it runs, a deployed app's *definition* and its *running containers* disagree,
and they stay that way until something else acts.
@@ -81,10 +81,12 @@ it does not change it.
---
## 3. The three operator decisions (2026-09-02)
## 3. The operator decisions
These are rulings, not proposals. Anything specced against a different assumption is wrong.
### 2026-09-02
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
@@ -98,6 +100,31 @@ These are rulings, not proposals. Anything specced against a different assumptio
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
because §4 says it cannot be undone.
### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0
4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it —
while corrections to its definition keep arriving on the 15-minute cycle exactly as they do
today.**
**This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes
that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in
its own comment. Reversing a chosen behaviour is a decision, not a bug fix.
**What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".**
The old behaviour had two halves and only one of them was unwanted:
| half | verdict |
|---|---|
| a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone |
| template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 |
So the ruling keeps the second and removes the first. In the operator's own words: *while the
catalog is offering the same version you are running, its fixes flow to you; the moment it moves to
a newer version, you are frozen at what you have until you choose to update.*
**It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup
precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.
---
## 4. The vocabulary ruling — "rollback" is struck
@@ -122,15 +149,95 @@ half-life because the syncer overwrites it (**R-441**).
---
## 5. The target shape
## 5. The shape, as SHIPPED in v0.235.0
**The live `docker-compose.yml` becomes DERIVED from a pin recorded in `app.yaml`** — the one file the
syncer never touches (`sync.go:319`'s exclusion). The catalog then proposes; `app.yaml` decides; the
rendered compose file is an output rather than an input, and the thirteen unattended `up -d` paths
stop being able to change a version by accident.
The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the
syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather
than an input, and the thirteen unattended `up -d` paths stop being able to change a version by
accident.
**Nothing in slices 1 or 2 implements this.** They make the current state *visible*, which is the
prerequisite for judging how urgent it is.
### 5.1 Nothing was added to the thirteen call sites, and that is deliberate
**They are made safe by removing the reason, not by gating them.** The most important of them are
REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`),
the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a
customer's app down, which is worse than the problem this slice solves.** Since the file they act on
no longer changes version, every one of them became safe without being touched.
### 5.2 The pin, and what it is not
`AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref.
**It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a
DECISION — what should run. Letting an observation feed a decision would make a bad reading become a
bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over.
They will normally agree; when they disagree that is a signal, not a bug to paper over.
**Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.**
Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came
from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is
safe from the catalog, and keeping it beside the app means it travels with every path that already
moves a stack dir.
### 5.3 The four writers — the only acts entitled to move a version
| writer | pin source |
|---|---|
| the deploy path (`runComposeDeploy`) | the template just deployed from |
| **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull |
| the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** |
| `Manager.AdoptPins` | the observation, once, and only when complete AND matching |
**`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file
on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards
would pull the frozen version and change nothing — while reporting success, and a button that lies is
worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of
`recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent.
### 5.4 The render table, complete
| app state | result |
|---|---|
| not deployed / protected / seam not wired | the catalog template — today's behaviour |
| deployed, **unpinned** | the catalog template + one DEBUG |
| deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** |
| deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE |
| pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it |
| mid-deploy | the compose file is left alone this cycle |
**`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries
`catalog_since`, which the badge needs. See §8.5.
**The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and
putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration
the older template cannot supply, so an old image under a new template is broken in a way neither
version is.
**And this is not "skip deployed apps".** That option was considered and rejected: it also stops
health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in
the spike §3 — both halves the ruling explicitly kept.
### 5.5 Adoption, and why the startup order matters
`AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app
to what it is already running. **It reads and writes files only** — no container is started, stopped or
touched. It skips, loudly, when the observation is incomplete or when the app runs something the
current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a
guessed pin.
**`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position
that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and
handed the next restart a version change — the exact behaviour this slice removes, once per boot.
### 5.6 The trap this slice set for the previous one
`Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one.
On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find
installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test
still green, because the new field has the same type and shape. The badge now reads
`Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an
earlier feature is the failure mode to look for whenever a file changes meaning.**
---
@@ -141,7 +248,7 @@ prerequisite for judging how urgent it is.
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 |
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
| **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
@@ -235,7 +342,18 @@ Version strings stay in the logs, the API and the hub.
startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller
START, so a box between upgrade and its next restart still shows nothing — bounded by one restart
rather than unbounded.
4. **The hub does not record image tags at all.** Its report's container payload carries name, state,
4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches
that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the
`wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken
state. Recorded so it is a choice, not a surprise.
5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
the update badge by withholding `catalog_since`. **R-458.**
6. **The Update button is still unguarded.** It takes no backup, has no rollback, and can still
attempt a multi-major jump the app will refuse (R-40). **Slice 3 did not change that and must not
be read as having done so** — the precondition is slice 4 (R-448).
7. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.