604 lines
40 KiB
Markdown
604 lines
40 KiB
Markdown
# 09 — How an app update works, and what it is becoming
|
|
|
|
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
|
|
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
|
|
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
|
|
> did not know it was one.
|
|
|
|
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
|
|
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
|
|
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
|
|
|
|
---
|
|
|
|
## 1. How an update works today, as measured
|
|
|
|
### 1.1 The button
|
|
|
|
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
|
|
commands and nothing else:
|
|
|
|
```
|
|
compose pull → compose up -d --remove-orphans
|
|
```
|
|
|
|
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
|
|
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
|
|
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
|
|
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
|
|
|
|
### 1.2 The catalog syncer moves the file underneath a deployed app
|
|
|
|
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
|
|
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
|
|
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
|
|
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
|
|
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
|
|
sync does not restart anything.**
|
|
|
|
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
|
|
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
|
|
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
|
|
|
|
### 1.3 Thirteen other paths end in `compose up -d`
|
|
|
|
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
|
|
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
|
|
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
|
|
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
|
|
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
|
|
restore path.
|
|
|
|
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
|
|
a boot reconciliation started an app on a newer image at 17:55:44Z.
|
|
|
|
### 1.4 One fear is measured SMALLER than it was stated
|
|
|
|
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
|
|
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
|
|
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
|
|
1c, a positive observable and not an absent log line).
|
|
|
|
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
|
|
so is more useful than leaving the scarier version standing.
|
|
|
|
---
|
|
|
|
## 2. What was chosen, and by whom
|
|
|
|
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
|
|
whole arc:
|
|
|
|
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
|
|
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
|
|
> SIGTERM+start to existing containers without re-reading the compose file or env."*
|
|
|
|
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
|
|
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
|
|
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
|
|
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
|
|
it does not change it.
|
|
|
|
---
|
|
|
|
## 3. The operator decisions
|
|
|
|
These are rulings, not proposals. Anything specced against a different assumption is wrong.
|
|
|
|
### 2026-09-02
|
|
|
|
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
|
|
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
|
|
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
|
|
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
|
|
§6 — the file half was never priced, and demo-hp is too young a box to price it.
|
|
|
|
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
|
|
A customer on the newest version is supported however old that version is. This is why
|
|
`catalog_since` exists and why no version string is shown.
|
|
|
|
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
|
|
because §4 says it cannot be undone.
|
|
|
|
### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0
|
|
|
|
4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it —
|
|
while corrections to its definition keep arriving on the 15-minute cycle exactly as they do
|
|
today.**
|
|
|
|
**This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes
|
|
that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in
|
|
its own comment. Reversing a chosen behaviour is a decision, not a bug fix.
|
|
|
|
**What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".**
|
|
The old behaviour had two halves and only one of them was unwanted:
|
|
|
|
| half | verdict |
|
|
|---|---|
|
|
| a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone |
|
|
| template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 |
|
|
|
|
So the ruling keeps the second and removes the first. In the operator's own words: *while the
|
|
catalog is offering the same version you are running, its fixes flow to you; the moment it moves to
|
|
a newer version, you are frozen at what you have until you choose to update.*
|
|
|
|
**It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup
|
|
precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.
|
|
|
|
### 2026-09-13 — the database engine finishes its own conversion, and the upgrade test goes wide
|
|
|
|
5. **DBs should be updated when the app moves, with proper precautions, tests and backoff plans**
|
|
— the operator's own words, ruling on `SPIKE-r459-mariadb-upgrade-2026-09-06.md`. SHIPPED in the
|
|
catalog the same day: every `mariadb:` sidecar (`bookstack-db`, `kimai-db`, `nextcloud-db`,
|
|
`romm-db`) carries `MARIADB_AUTO_UPGRADE=1`; `MARIADB_DISABLE_UPGRADE_BACKUP` stays unset. **Not an
|
|
image change, so `catalog_since` does not move.** The three precautions, because they are the real
|
|
content of the ruling:
|
|
1. **Proven before it ships** — `upgrade-test.py` re-ran E3 and E3b on the changed template and the
|
|
engine-state field shows the conversion RAN (`mariadb_upgrade_info` reads the new version, the
|
|
engine's own check says nothing further is needed, the entrypoint no longer prints
|
|
`skipped due to $MARIADB_AUTO_UPGRADE`), with the seeded data reading back after. C3 still
|
|
returns `failed`. Evidence: `audits/r459-close-2026-09-13/`.
|
|
2. **Watched as it lands** — the change travelled the real 15-minute cycle to demo-hp: the live
|
|
compose gained the setting, the sync recreated nothing, and one deliberate restart logged
|
|
`MariaDB upgrade not required` with the app serving (same evidence directory).
|
|
3. **A rule until Slice 4 is built** — the Update button still takes no backup, so **no template
|
|
may move a database-engine image across a major version until R-448 ships.** Catalog
|
|
`CLAUDE.md` states it; `scripts/check-engine-major.py` enforces it in the pre-push hook (the
|
|
CI half cannot, R-452); its removal is tracked as **R-469** so it is a deliberate act.
|
|
The setting is inert until an engine major moves, and precaution 3 keeps it that way.
|
|
|
|
6. **The upgrade test goes as wide as possible, through the nightly unattended sessions** — the
|
|
ruling on `STATUS.md` item 11. Not "the ~25 database apps first": all of them, as the nightly
|
|
rotation reaches them, one fixture per app through the app's own interface. **Browser-only apps
|
|
become reachable when CC runs on the operator's Windows workstation with Chrome** — the
|
|
`claude-in-chrome` route that DooPlex does not have — so an app recorded `inconclusive` for want
|
|
of a headless seed route (bookstack's file half, R-460) is deferred to that venue, not faked.
|
|
|
|
---
|
|
|
|
### 2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts
|
|
|
|
7. **A floor carries a release past the vouched golden when the release's MinAgent is declared with
|
|
it** (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it,
|
|
the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which
|
|
`minagent_header_gate.py` now guarantees — goes into the same per-box agent comparison. An
|
|
undeclared floor above the golden is still HELD, and both floor forms refuse to save one. **Why:** a
|
|
controller image is pulled by tag and needs no golden to be delivered; what the floor was missing
|
|
was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release
|
|
delivery coexist. **Proven live:** both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s
|
|
from the save, the hub logging `SERVED … from declared`
|
|
(`audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt`).
|
|
8. **Any backup tier lets an app update** (R-475; controller v0.239.0). The precondition takes the first
|
|
FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site
|
|
(Tier 3, bounded; unreachable = absent). `update.backup_max_age` applies to whichever tier is chosen.
|
|
An app with nothing anywhere is backed up first; it is refused only when no backup can be taken
|
|
either. The hold names the tier and the date. **Tier 2 is required nowhere in the update path.**
|
|
This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a
|
|
verified recent backup as a precondition, not a new copy) stands. **Proven live:**
|
|
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
|
|
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
|
|
|
|
9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy
|
|
holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit
|
|
that holds the definition and the database dumps and NOT the files (measured: gokapi restored from
|
|
„helyi" came back with settings and no data). For such an app the update walks second drive →
|
|
off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive →
|
|
own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a
|
|
customer is never sent to a copy that cannot bring the data back without being told so.
|
|
|
|
## 4. The vocabulary ruling — "rollback" is struck
|
|
|
|
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
|
|
run, putting the old image tag back produces a container that refuses to start —
|
|
|
|
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
|
|
> downgrading is not supported"*
|
|
|
|
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
|
|
|
|
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
|
|
|
|
| shape | when it applies | what it does |
|
|
|---|---|---|
|
|
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
|
|
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
|
|
|
|
There is no third.
|
|
|
|
### 4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades
|
|
|
|
`SPIKE-upgrade-test-2026-09-06.md` upgraded three real apps with real data in them and then attempted
|
|
the abort on each. **All five real catalog upgrades kept the customer's data.** The abort did not
|
|
behave the same way twice:
|
|
|
|
| app | abort | why |
|
|
|---|---|---|
|
|
| **docmost** `0.25.3`→`0.95.0` | **REFUSES** | the old code finds migration ledger entries it does not know: *"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*, then *"Failed to run database migration. Exiting program."* |
|
|
| **privatebin** `1.7.5`→`2.0.5` | **works** | file-backed, no database, no schema — a major version moves no data |
|
|
| **bookstack** app+engine | **works, misleadingly** | only because the MariaDB datadir upgrade was skipped and never happened — see R-459 |
|
|
|
|
**This puts TWO independent measurements behind the ruling above, by two unrelated mechanisms:**
|
|
Nextcloud refused on an explicit version comparison; docmost refuses on its migration ledger. The word
|
|
"rollback" was already struck; it is now struck on evidence rather than on one case.
|
|
|
|
**And it adds a distinction this document did not have: there is no single answer to "can this update
|
|
be undone". There are apps where it can and apps where it cannot, and the only way to know which is to
|
|
MEASURE THAT APP.** Any design that assumes one answer for all 53 is designing against a fact that was
|
|
checked and is false.
|
|
|
|
Per §5 below, a restore's image-level undo used to have a ≤15-minute half-life because the syncer
|
|
overwrote it (**R-441**) — closed in v0.235.0.
|
|
|
|
---
|
|
|
|
## 5. The shape, as SHIPPED in v0.235.0
|
|
|
|
The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the
|
|
syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather
|
|
than an input, and the thirteen unattended `up -d` paths stop being able to change a version by
|
|
accident.
|
|
|
|
### 5.1 Nothing was added to the thirteen call sites, and that is deliberate
|
|
|
|
**They are made safe by removing the reason, not by gating them.** The most important of them are
|
|
REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`),
|
|
the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a
|
|
customer's app down, which is worse than the problem this slice solves.** Since the file they act on
|
|
no longer changes version, every one of them became safe without being touched.
|
|
|
|
### 5.2 The pin, and what it is not
|
|
|
|
`AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref.
|
|
|
|
**It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a
|
|
DECISION — what should run. Letting an observation feed a decision would make a bad reading become a
|
|
bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over.
|
|
They will normally agree; when they disagree that is a signal, not a bug to paper over.
|
|
|
|
**Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.**
|
|
|
|
Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came
|
|
from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is
|
|
safe from the catalog, and keeping it beside the app means it travels with every path that already
|
|
moves a stack dir.
|
|
|
|
### 5.3 The four writers — the only acts entitled to move a version
|
|
|
|
| writer | pin source |
|
|
|---|---|
|
|
| the deploy path (`runComposeDeploy`) | the template just deployed from |
|
|
| **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull |
|
|
| the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** |
|
|
| `Manager.AdoptPins` | the observation, once, and only when complete AND matching |
|
|
|
|
**`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file
|
|
on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards
|
|
would pull the frozen version and change nothing — while reporting success, and a button that lies is
|
|
worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of
|
|
`recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent.
|
|
|
|
### 5.4 The render table, complete
|
|
|
|
| app state | result |
|
|
|---|---|
|
|
| not deployed / protected / seam not wired | the catalog template — today's behaviour |
|
|
| deployed, **unpinned** | the catalog template + one DEBUG |
|
|
| deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** |
|
|
| deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE |
|
|
| pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it |
|
|
| mid-deploy | the compose file is left alone this cycle |
|
|
|
|
**`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries
|
|
`catalog_since`, which the badge needs. See §8.5.
|
|
|
|
**The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and
|
|
putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration
|
|
the older template cannot supply, so an old image under a new template is broken in a way neither
|
|
version is.
|
|
|
|
**And this is not "skip deployed apps".** That option was considered and rejected: it also stops
|
|
health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in
|
|
the spike §3 — both halves the ruling explicitly kept.
|
|
|
|
### 5.5 Adoption, and why the startup order matters
|
|
|
|
`AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app
|
|
to what it is already running. **It reads and writes files only** — no container is started, stopped or
|
|
touched. It skips, loudly, when the observation is incomplete or when the app runs something the
|
|
current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a
|
|
guessed pin.
|
|
|
|
**`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position
|
|
that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and
|
|
handed the next restart a version change — the exact behaviour this slice removes, once per boot.
|
|
|
|
### 5.6 The trap this slice set for the previous one
|
|
|
|
`Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one.
|
|
On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find
|
|
installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test
|
|
still green, because the new field has the same type and shape. The badge now reads
|
|
`Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an
|
|
earlier feature is the failure mode to look for whenever a file changes meaning.**
|
|
|
|
---
|
|
|
|
## 6. The seven slices
|
|
|
|
| # | slice | status |
|
|
|---|---|---|
|
|
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
|
|
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
|
|
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
|
|
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
|
|
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 |
|
|
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
|
|
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
|
|
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
|
|
|
|
### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)
|
|
|
|
`POST /api/stacks/{name}/update` is a guarded job. It answers **202** at once; the outcome exists only
|
|
on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`, `update_error`,
|
|
`hold_reason`), and `update_phase=done` is written only after the app's health is known. **R-443 is
|
|
closed by construction: nothing reports an update complete on the compose exit code.**
|
|
|
|
**The sequence, and the order is the design:**
|
|
|
|
| # | phase | what happens | on failure |
|
|
|---|---|---|---|
|
|
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no copy on ANY tier and no backup can be taken now** (since v0.239.0; before it, no restorable Tier-2 copy) | nothing moves, nothing is recorded |
|
|
| 1 | `checking` | walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than `update.backup_max_age` (v0.239.0) | nothing moves |
|
|
| 2 | `backing-up` — only when no tier holds a fresh copy | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 | refused with the backup's own error; nothing moves |
|
|
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
|
|
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
|
|
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
|
|
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
|
|
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
|
|
| 8 | `done` | installed images recorded, journal cleared | — |
|
|
|
|
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
|
|
`update.health_timeout` (default `5m`).
|
|
|
|
**The precondition is the existing verified backup, not a new copy** (§3 decision 1). It is
|
|
`backup.Tier2UnitRestorePoint` — the SAME predicate that permits the destructive „Teljes
|
|
visszaállítás", extracted from the backups page rather than copied. **The copy is aged by the last
|
|
SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was measured before it was
|
|
designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp
|
|
bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the
|
|
manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is
|
|
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).** **SUPERSEDED in
|
|
v0.239.0 by §3 decision 8:** `backup.Manager.UpdateRestorePoints` walks all three tiers and the update
|
|
leans on the first fresh copy; `Tier2UnitRestorePoint` is still the Tier-2 half and still the page's
|
|
predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum
|
|
skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit
|
|
proven current. The hold stores `copy_tier` and names „második meghajtó" / „saját meghajtó" /
|
|
„távoli mentés"; a successful off-site restore now lifts an update hold too.
|
|
|
|
**The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and
|
|
gate as R-379, so every start path that already honoured a restore hold honours this one. A successful
|
|
unit restore lifts an update hold (only that kind). **Three unattended paths honoured no hold before
|
|
v0.237.0 and now do:** the drive-return gate's restart and boot recreate, and the nightly volume dump
|
|
(which ends in `StartStack`). The nightly capture and Tier-2 run skip a held app, so the restore point
|
|
the hold text names is never overwritten.
|
|
|
|
**Crash safety is a journal**, `<data>/update-journal.json`, written before every phase. `RecoverUpdates`
|
|
runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back;
|
|
after `up` → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by
|
|
`ResumeInterruptedUpdates` once the backup side is wired, ending healthy or held.
|
|
|
|
**THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT.** Whether an old image
|
|
starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot
|
|
be predicted. So the box never puts the old version back by itself. **The route back is the restore**,
|
|
and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where
|
|
the harness has proven it, is slice 6's.
|
|
|
|
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
|
|
|
|
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
|
|
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
|
|
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
|
|
both demo boxes by the floor alone, with its MinAgent declared.
|
|
|
|
**Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not
|
|
only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute
|
|
health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit.
|
|
The Tier-2 mirror the hold names survived only because Tier 2 is daily. `backup.Manager.isHeld` now also
|
|
answers true for an app a guarded update is moving (`SetUpdatingCheck`).
|
|
|
|
**v0.240.0 (2026-09-13, evening) — what the afternoon's proof and the first nightly rotation found, fixed.**
|
|
Seven rows: removal with backups kept now keeps the Tier-2 RECORD, so the second-drive restore is not
|
|
refused over an intact mirror (R-486, P1 — the disaster the second copy exists for); PostGIS/pgvector/
|
|
TimescaleDB images are Postgres, so such apps get their logical dump (R-484); "delete backups" deletes
|
|
the unit, the mirror(s) and the prefs (R-474/R-466); the backup card sizes them (R-485); a held
|
|
update's sentence leaves the card with the hold (R-480); the Tier-3 lookup is one `snapshots` call
|
|
(R-477); a unit older than the app's `deployed_at` does not count (R-478). Delivered by the floor in
|
|
16 s / 18 s; every row proven live with a throwaway adventurelog. `audits/v0240-2026-09-13/`.
|
|
|
|
**Proven live on demo-hp, 2026-09-13**, with a throwaway uptime-kuma and real catalog tag changes (each
|
|
reverted in the same phase): A (2.3.2→2.4.0, done after health), B (`backup_max_age: 2m` → backup first),
|
|
E (non-existent tag → pin back, container untouched), F (`alpine:3.20` → held), H (three buttons and the
|
|
boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold
|
|
cleared). Live evidence: `audits/slice4-2026-09-13/`.
|
|
|
|
### The verdict record — the contract Slice 6 carries
|
|
|
|
Decided here rather than invented twice. The harness writes one of these per edge, beside its
|
|
evidence; Slice 6 puts the same shape in the catalog.
|
|
|
|
```json
|
|
{"harness_version": 1, "app": "bookstack",
|
|
"from": {"bookstack": "…:25.02.2", "bookstack-db": "mariadb:11.6"},
|
|
"to": {"bookstack": "…:26.05.2", "bookstack-db": "mariadb:12.3"},
|
|
"verdict": "proven | failed | inconclusive",
|
|
"seed_read_before": true, "seed_read_after": true, "healthy_after": true,
|
|
"migration_observed": "verbatim log line, or null",
|
|
"abort": "starts-and-serves | refuses | starts-data-gone | not-attempted",
|
|
"abort_detail": "the refusal quoted verbatim, or null",
|
|
"duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"}
|
|
```
|
|
|
|
**`inconclusive` is a first-class verdict and must never be collapsed into `failed`.** "We could not
|
|
measure it" and "it does not work" are different facts, and only one of them is about the app.
|
|
**`migration_observed` is a quoted line, never an inference from timing** — the value of both the
|
|
Nextcloud and the docmost findings was the exact sentence the app printed.
|
|
|
|
### Database engines under an upgrade — MEASURED 2026-09-06
|
|
|
|
**The arc's standing rule that an engine change gets its OWN edge now has measured evidence behind
|
|
it**, and the evidence is stronger than the rule's original argument. The rule was justified by
|
|
*"two migrations behind one edge is an unreadable failure when it breaks"* — a readability argument.
|
|
What was measured is that **an engine change can be applied and silently NOT happen**, which the
|
|
app-half edge cannot produce and which no amount of readability would have surfaced:
|
|
|
|
- `SPIKE-upgrade-test-2026-09-06.md` §4 — MariaDB 12.3 starts on an 11.6 datadir, logs that the
|
|
conversion it requires was **skipped**, and serves. **Assigned to the engine half by decomposition:**
|
|
the app half alone produces no such line.
|
|
- `SPIKE-r459-mariadb-upgrade-2026-09-06.md` — it is **stable but never self-resolving** (5 of 5
|
|
restarts, no degradation, and the engine says `Check required!` every time, forever). Converting
|
|
properly **succeeds**, costs **7 s**, takes its own system-database backup, and **does not** cost the
|
|
ability to abort. **The trade that was expected here does not exist.**
|
|
- **2026-09-13 — the setting is in the catalog.** All four `mariadb:` sidecars carry
|
|
`MARIADB_AUTO_UPGRADE=1` (operator ruling, §3 decision 5), and `upgrade-test.py`'s engine-state field
|
|
now shows the conversion RUNNING on the bookstack edges. **And a gate holds the engines inside their
|
|
major until Slice 4:** `app-catalog-felhom.eu/scripts/check-engine-major.py` (R-469).
|
|
|
|
**Two rules for anything this arc builds around a database engine:**
|
|
|
|
1. **Ask the engine, not the log.** MariaDB's entrypoint prints `MariaDB upgrade not required` on an
|
|
unsupported **downgrade**; `mariadb-upgrade --check-if-upgrade-is-needed` names it exactly
|
|
(**R-464**). A cheap instrument built on the log line would report "fine" for the broken case.
|
|
2. **The two engines fail in opposite directions, so one check will not do.** MariaDB starts anyway
|
|
and skips quietly; **PostgreSQL refuses to start** on a datadir from an older major, and the image
|
|
performs no `pg_upgrade`. Eleven templates carry PostgreSQL and **eight sit on `postgres:16-alpine`**
|
|
(**R-463**).
|
|
|
|
**And an engine-state field belongs BESIDE a verdict, never inside it.** `upgrade-test.py` reports
|
|
`engine_state_after` next to `verdict`, because an unconverted datadir is not *known* to be a failure
|
|
and a verdict that said so would encode an unproven judgement.
|
|
|
|
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
|
|
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
|
|
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
|
|
breaks.
|
|
|
|
---
|
|
|
|
## 7. What slices 1 and 2 actually built
|
|
|
|
### 7.1 The record (slice 1)
|
|
|
|
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
|
|
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
|
|
and writes `app.yaml`:
|
|
|
|
```yaml
|
|
installed_images:
|
|
web:
|
|
ref: lscr.io/linuxserver/bookstack:26.05.2
|
|
digest: sha256:… # "" if the image was never pulled from a registry
|
|
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
|
|
```
|
|
|
|
Three rules, each with its reason:
|
|
|
|
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
|
|
already moved. A record built from it would answer "what will happen next time something runs
|
|
`up -d`", which is a different question.
|
|
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
|
|
Intent refused, observation logged. Refusing to start a customer's app because we could not write
|
|
down which version it is trades a real outage for a bookkeeping gap.
|
|
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
|
|
complete record with a partial one.
|
|
|
|
**Seeded at startup (v0.234.0).** `Manager.BackfillInstalledImages` runs once at boot, beside the
|
|
desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY
|
|
on. It only READS containers. **This was not a refinement — without it the feature did not reach a
|
|
quiet box at all:** see §8.3, which was written as a known limitation on 2026-09-02 and was a defect
|
|
by the next morning.
|
|
|
|
Two admission rules, and the second is the design:
|
|
|
|
- **It never overwrites an existing record.** The bring-up paths own updates; this fills gaps only.
|
|
- **It refuses to seed a PARTIAL observation.** §7.2's comparison reads a service-count mismatch as
|
|
BEHIND, so a degraded or crash-looping app seeded from its visible containers would render
|
|
„Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial
|
|
because they follow a SUCCESSFUL `up -d`, where a gap is real news and is logged; a backfill meets a
|
|
box in whatever state it is in. **Same field, two writers, two different admission rules — that is
|
|
deliberate and must not be "made consistent".**
|
|
|
|
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
|
|
|
|
### 7.2 The label (slice 2)
|
|
|
|
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
|
|
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
|
|
no new CSS.
|
|
|
|
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
|
|
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
|
|
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
|
|
a test with a companion red-proof.
|
|
|
|
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
|
|
Version strings stay in the logs, the API and the hub.
|
|
|
|
---
|
|
|
|
## 8. Known limitations, stated plainly
|
|
|
|
1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and
|
|
queries no registry — a customer's box must not depend on reaching eight upstream registries to
|
|
render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be
|
|
identical while the image behind it has moved. **Measured, not theorised:** spike §5 found
|
|
`mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls
|
|
holding. Digest-level comparison needs a registry query and is deferred — **R-446**.
|
|
2. **~~Nothing enforces `catalog_since`.~~ Enforced by the pre-push hook since 2026-09-13 (R-452, `app-catalog-felhom.eu/scripts/check-catalog-since.py`); CI's shallow clone still skips it out loud.** A commit that moves an `image:` line and forgets the date
|
|
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
|
|
commit to diff against, so the gate needs a deeper fetch — **R-452**.
|
|
3. ~~**The record only appears after the next lifecycle action.**~~ **CLOSED in v0.234.0, and the way
|
|
it closed is worth keeping.** This was written on 2026-09-02 as an accepted limitation — *"the
|
|
fleet view fills in gradually"*. The operator looked at demo-felhom the next morning and found
|
|
OpenGist, up 15 hours, running exactly the catalog pin, showing **nothing at all**. **On a quiet
|
|
box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is,
|
|
on the quiet installations, not shipped.** `BackfillInstalledImages` now seeds the absences at
|
|
startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller
|
|
START, so a box between upgrade and its next restart still shows nothing — bounded by one restart
|
|
rather than unbounded.
|
|
4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches
|
|
that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the
|
|
`wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken
|
|
state. Recorded so it is a choice, not a surprise.
|
|
5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in
|
|
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
|
|
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
|
|
the update badge by withholding `catalog_since`. **R-458.**
|
|
6. ~~**The Update button is still unguarded.**~~ **CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1).** It
|
|
refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a
|
|
safety dump, and holds an app that does not come up. **What stays true:** it still has no automatic
|
|
rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) —
|
|
that now ends HELD rather than crash-looping behind a green button.
|
|
7. ~~**An engine major can be applied without its datadir upgrade, and nothing notices.**~~ **CLOSED
|
|
2026-09-13 for MariaDB (R-459):** every `mariadb:` sidecar carries `MARIADB_AUTO_UPGRADE=1`, and the
|
|
harness shows the conversion running on the E3/E3b edges (§3 decision 5). **What stays true:** the
|
|
PostgreSQL half (R-463) has no equivalent — the image performs no `pg_upgrade` — and the
|
|
engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until
|
|
Slice 4 gives the Update button a backup.
|
|
8. **Only three of 53 apps have ever had an upgrade measured**, and one of them (bookstack) can only
|
|
be half-proven headlessly (**R-460**). The widening is **R-462**, costed with real numbers.
|
|
9. **The hub does not record image tags at all.** Its report's container payload carries name, state,
|
|
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
|
|
change; it is not derivable from what is already reported.
|
|
|
|
---
|
|
|
|
## 9. Where the rest lives
|
|
|
|
| what | where |
|
|
|---|---|
|
|
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
|
|
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
|
|
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
|
|
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
|
|
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |
|