Files
felhom.eu/documentation/architecture/09-update-architecture.md
T
admin 4c92beab8f
gates / gates (push) Successful in 26s
Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
  (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
  back with data written before AND after the backup. The product's loader
  cannot do it: over a migrated PG database it fails on the new tables'
  foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
  rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
  Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.

No product code. Live catalog untouched; 9202 back on it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 08:54:41 +02:00

1349 lines
100 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 09 — How an app update works, and what it is becoming
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
> did not know it was one.
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
---
## 1. How an update works today, as measured
### 1.1 The button
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
commands and nothing else:
```
compose pull → compose up -d --remove-orphans
```
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
### 1.2 The catalog syncer moves the file underneath a deployed app
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
sync does not restart anything.**
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
### 1.3 Thirteen other paths end in `compose up -d`
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
restore path.
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
a boot reconciliation started an app on a newer image at 17:55:44Z.
### 1.4 One fear is measured SMALLER than it was stated
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
1c, a positive observable and not an absent log line).
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
so is more useful than leaving the scarier version standing.
---
## 2. What was chosen, and by whom
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
whole arc:
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
> SIGTERM+start to existing containers without re-reading the compose file or env."*
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
it does not change it.
---
## 3. The operator decisions
These are rulings, not proposals. Anything specced against a different assumption is wrong.
### 2026-09-02
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
§6 — the file half was never priced, and demo-hp is too young a box to price it.
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
A customer on the newest version is supported however old that version is. This is why
`catalog_since` exists and why no version string is shown.
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
because §4 says it cannot be undone. **Second half REPLACED 2026-09-23 by decision 13** (the test
decides, not the tag); the first half is confirmed by decision 12.
### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0
4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it —
while corrections to its definition keep arriving on the 15-minute cycle exactly as they do
today.**
**This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes
that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in
its own comment. Reversing a chosen behaviour is a decision, not a bug fix.
**What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".**
The old behaviour had two halves and only one of them was unwanted:
| half | verdict |
|---|---|
| a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone |
| template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 |
So the ruling keeps the second and removes the first. In the operator's own words: *while the
catalog is offering the same version you are running, its fixes flow to you; the moment it moves to
a newer version, you are frozen at what you have until you choose to update.*
**It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup
precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.
### 2026-09-13 — the database engine finishes its own conversion, and the upgrade test goes wide
5. **DBs should be updated when the app moves, with proper precautions, tests and backoff plans**
— the operator's own words, ruling on `SPIKE-r459-mariadb-upgrade-2026-09-06.md`. SHIPPED in the
catalog the same day: every `mariadb:` sidecar (`bookstack-db`, `kimai-db`, `nextcloud-db`,
`romm-db`) carries `MARIADB_AUTO_UPGRADE=1`; `MARIADB_DISABLE_UPGRADE_BACKUP` stays unset. **Not an
image change, so `catalog_since` does not move.** The three precautions, because they are the real
content of the ruling:
1. **Proven before it ships** — `upgrade-test.py` re-ran E3 and E3b on the changed template and the
engine-state field shows the conversion RAN (`mariadb_upgrade_info` reads the new version, the
engine's own check says nothing further is needed, the entrypoint no longer prints
`skipped due to $MARIADB_AUTO_UPGRADE`), with the seeded data reading back after. C3 still
returns `failed`. Evidence: `audits/r459-close-2026-09-13/`.
2. **Watched as it lands** — the change travelled the real 15-minute cycle to demo-hp: the live
compose gained the setting, the sync recreated nothing, and one deliberate restart logged
`MariaDB upgrade not required` with the app serving (same evidence directory).
3. **A rule until Slice 4 is built** — the Update button still takes no backup, so **no template
may move a database-engine image across a major version until R-448 ships.** Catalog
`CLAUDE.md` states it; `scripts/check-engine-major.py` enforces it in the pre-push hook (the
CI half cannot, R-452); its removal is tracked as **R-469** so it is a deliberate act.
The setting is inert until an engine major moves, and precaution 3 keeps it that way.
6. **The upgrade test goes as wide as possible, through the nightly unattended sessions** — the
ruling on `STATUS.md` item 11. Not "the ~25 database apps first": all of them, as the nightly
rotation reaches them, one fixture per app through the app's own interface. **Browser-only apps
become reachable when CC runs on the operator's Windows workstation with Chrome** — the
`claude-in-chrome` route that DooPlex does not have — so an app recorded `inconclusive` for want
of a headless seed route (bookstack's file half, R-460) is deferred to that venue, not faked.
---
### 2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts
7. **A floor carries a release past the vouched golden when the release's MinAgent is declared with
it** (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it,
the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which
`minagent_header_gate.py` now guarantees — goes into the same per-box agent comparison. An
undeclared floor above the golden is still HELD, and both floor forms refuse to save one. **Why:** a
controller image is pulled by tag and needs no golden to be delivered; what the floor was missing
was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release
delivery coexist. **Proven live:** both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s
from the save, the hub logging `SERVED … from declared`
(`audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt`).
8. **Any backup tier lets an app update** (R-475; controller v0.239.0). The precondition takes the first
FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site
(Tier 3, bounded; unreachable = absent). `update.backup_max_age` applies to whichever tier is chosen.
An app with nothing anywhere is backed up first; it is refused only when no backup can be taken
either. The hold names the tier and the date. **Tier 2 is required nowhere in the update path.**
This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a
verified recent backup as a precondition, not a new copy) stands. **Proven live:**
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy
holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit
that holds the definition and the database dumps and NOT the files (measured: gokapi restored from
„helyi" came back with settings and no data). For such an app the update walks second drive →
off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive →
own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a
customer is never sent to a copy that cannot bring the data back without being told so.
### 2026-09-21 — decided by CC unattended, operator may reverse
10. **A box AHEAD of the catalog reads „Naprakész", and the guarded Update refuses to move a pin
backwards** (R-524, controller v0.260.0). *One sentence:* when the catalog is reverted under a box
that already updated, is that a "Frissítés elérhető"? **Options:** (a) leave it — the label
compares for difference, as §5.4 says; (b) show „Naprakész" and let the button still run;
(c) show „Naprakész" and refuse the button. **Costs:** (a) is free and offers a household a
downgrade onto a datadir the newer version may have migrated, which §4 says cannot be undone;
(b) removes the invitation but leaves the loaded gun; (c) costs one comparison and can, wrongly
applied, block a legitimate update. **Why (c):** the direction was already settled — §3 decision 3
says a version change that cannot be undone needs a human, and this is one. The risk in (c) is
bounded by making the Ahead verdict NARROW: every differing service must be orderable AND newer,
or the answer falls back to today's behaviour. Reversible, no customer-data risk, and it only ever
withholds an act. **Implementation:** `stacks.CatalogOrder`, one verdict read by both the badge
and `UpdatePreflight`.
### 2026-09-23 — the operator answers §3b's seven questions
Operator rulings, not CC decisions. Each answers one question in §3b, which is kept, marked ANSWERED,
because the costs written there are the reasoning behind these. **Nothing below is built yet** — the
build order is §6.4, and the two new mechanisms (15 and 14) were spiked the same day before anyone
builds them (`audits/update-rulings-2026-09-23/`).
11. **One maintenance window, and it is the backup window the household already sets** (Q1). App
updates are one more leg of the nightly chain, **after the off-site copy and before the
full-system backup**. An update not finished when the full-system backup is due waits for the
next night. **Why:** the update then leans on the freshest copy the box ever has, and there is no
second window for anyone to set or misread. **Replaces:** §3b Q1's recommendation of a fixed
02:30–05:00. *For the builder:* `stacks.update_window` (`config.go` L195, default `03:00-05:00`)
exists and **no Go code reads it** — it appears only in `config.go` (field + default),
`setup/handlers.go` L482 (written into a new box's `controller.yaml`) and
`configs/controller.yaml.example`. Remove it or fold it into the chain; do not build a second
window.
12. **Automatic updates stay, behind a per-box switch that is ON by default** (Q2). Confirms the first
half of decision 3. **Today every update is still manual; nothing automatic is built.**
13. **The test decides, not the tag** (Q3). **REPLACES decision 3's second half, "never across a
major".** The box applies by itself every step the catalog holds, because **the catalog holds only
tested steps**; the size of the version jump does not matter, the test does. The box does not
parse tags to decide. The tag rule (`stacks.CompareImageRefs`) moves to the **catalog gate** as a
push-time safety net: an image move with no test record is refused there. Two exceptions, both
**marks the test sets on the step**: *files may change* → automatic only when a fresh copy holds
the files, else a person (this is Q2's refinement); *needs a person* → the tester writes why.
**Why:** decision 3 was a proxy. "Within a major" was standing in for "known to work", and the
update night measured the proxy wrong in both directions — minor moves that held
(adventurelog, outline) and a tag shape it cannot read at all (§3b Q3). The test record is the
thing the proxy was guessing at.
14. **The update ladder** (Q3). A box more than one step behind **climbs one tested step at a time, in
order, and never jumps**. Each step is the full guarded update of §6.1. **Why:** a step was tested
from the version before it, not from three versions before it; R-40's multi-major jump is what
this prevents. This is the "stepping" §6.1 already assigned to slice 6.
15. **The box undoes a failed update itself** (Q4). **REPLACES §6.1's "the box never puts the old
version back by itself".** On a failed health check: the old definition and pin back + the
database copy taken seconds before the pin moved loaded back + the health check again. The app
stays stopped (HOLD) **only if the undo itself fails**. **Why the old ruling does not bind this:**
the old version refused to start on data the new one had migrated (Nextcloud, docmost — §4). The
undo puts the **pre-migration data** back too, so the old version meets the data it knows. The
word "rollback" stays struck; **the name is undo**. The household is told on the app page and by
one mail; the operator by event; **no retry** until the catalog moves or a person presses.
Spiked 2026-09-23 before any build — §6.1a.
16. **PostgreSQL majors are converted by the box** (Q5). A guarded-update step: save everything from
the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on
the test bench before the catalog may move it. The engine-major gate stays until then. **Why:** the
update night costed it at ~9 s of engine work for 49 MB (§3b Q5); moving the pin and letting it
hold would take eleven apps down on one night.
17. **The catalog records the image digest of every pin at push time** (Q6). The box compares against
it; where the catalog carries one, the box pulls **that exact image** — which also makes a floating
tag reproducible, not only the badge honest. *Whether compose can pull by a recorded digest while
the definition names a tag is a claim for the build plan to verify, not a ruling on mechanism.*
18. **Fleet view** (Q7). The report carries, per compose service (so the database too), the installed
reference, the catalog reference and the badge state. **Built later**, when the fleet grows.
**RomM follow-ups, operator-agreed the same day:** the test bench watches memory after an update
(`upgrade-test.py`, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4);
R-636's louder repeated alarm.
---
## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed
> **ALL SEVEN ANSWERED by the operator on 2026-09-23 — §3 decisions 11–18.** Q1 → 11; Q2 → 12 and
> 13's *files may change* mark; Q3 → 13 and 14; Q4 → 15; Q5 → 16; Q6 → 17; Q7 → 18. Kept below
> unedited, because the costs and measurements here are the reasoning behind those rulings. Where a
> ruling differs from the recommendation below, the ruling wins: Q1 (the household's backup window,
> not a fixed 02:30–05:00), Q3 (the test decides, not `CompareImageRefs`), Q4 (the box UNDOES before
> it holds).
**These were questions, not rulings. CC did not decide them.** Each is one answerable sentence, the
options, what each costs, the recommendation, and what happens if nothing is decided. The measurement
behind them is `audits/UPDATE-ARC-STATE-2026-09-21.md`; the short version is that **46 of the
catalog's 58 exact pins are behind upstream today and 39 of those are within a major** — the
population §3 decision 3 already says may move without a human, and nobody presses 39 buttons.
### Q1 — When may a box update itself? — ANSWERED 2026-09-23: §3 decision 11
*May the box run the guarded Update by itself between 02:30 and 05:00, nightly?*
| option | cost |
|---|---|
| **02:30–05:00 nightly, after the backup legs** | the update leans on a copy made hours earlier the same night, which is the freshest the box ever has. The app is down for the health wait in the middle of the night. |
| a weekly window | fewer interruptions; a box sits up to 7 days on a version the catalog already moved past, which widens the support window §3 decision 2 runs on |
| the household picks the window | one more setting on a page that already has several, for a choice almost nobody will change |
**Recommendation: 02:30–05:00 nightly.** The DB dump runs 02:30 and restic 03:00 on a demo box, so a
window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy of the night without a
new mechanism. **If nothing is decided:** Slice 6 cannot be built at all — every other question below
is downstream of this one.
**⚠ THE WINDOW CONTAINS 04:30, AND 04:30 IS WHEN THE BOX UPDATES ITSELF.** Found 2026-09-21 (R-608) by
reading the clock rather than by a failure: the controller self-updates daily at
`self_update.auto_update_time`, **default 04:30**, and again from `MaybeAutoUpdate` after ANY hub
report once a floor sits above the box — so **at any hour, not only at 04:30**. Either path restarts
the controller container, which is the supervisor of a running app update.
**What v0.261.0 now guarantees, so this question can be answered without also solving that one:** the
two cannot overlap in either direction. The controller defers its own swap while a guarded app update
is in flight (retrying on the next report, exactly as it already did for a running backup), and
`UpdatePreflight` refuses `self_updating` while a swap is in progress. **The lock does not latch** — a
held app does not block the controller's updates for ever. So the window may contain 04:30; the two
jobs will queue behind one another rather than meet. **What it does NOT do is reorder them**: if the
operator prefers the box to take its own update first, that is a scheduling choice still open here.
### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES? — ANSWERED 2026-09-23: §3 decisions 12 and 13 (the *files may change* mark)
*The button's rule and the automatic rule can differ. Should they?*
**The mechanism, verified at source this session, because an earlier draft had it backwards:** the
guard does **not** refuse these apps. Since v0.239.0/v0.241.0 (§3 decisions 8–9) the Update is refused
only when no copy exists on ANY tier and none can be taken. For an app whose data is bind-mounted
files, `Manager.UpdateTierOrderFor` (`controller/internal/backup/update_guard.go:136-141`) walks
second drive → off-site → **own unit last**, and when the own unit is the copy chosen,
`UpdateCopyHolds` (`:145-165`) ends the hold sentence with *„csak a beállításokat és az adatbázist
tartalmazza, a fájlokat nem"* — it holds the settings and the database and **not the files**. So the
update **proceeds**, and the household is told what the copy holds.
**With a human pressing, that is an informed choice. With nobody pressing, nobody was informed.**
| option | cost |
|---|---|
| **automatic requires a fresh copy that HOLDS THE FILES; the button keeps today's rule** | the nine file-leg apps (and any other bind-data app) update automatically only on a box with a second drive or off-site; on a one-drive box they wait for a person. Two rules to hold in one's head. |
| one rule for both — automatic follows the button | simpler; a file-leg app can be updated unattended against a copy that cannot bring its files back, and the sentence saying so is read by nobody |
| automatic skips bind-data apps entirely | simplest; the seven file-leg apps that are behind never move by themselves even when a good copy exists |
**Recommendation: the first.** It is the smallest rule that keeps the promise the hold sentence makes.
**If nothing is decided:** Slice 6 must be built for the safe subset only, and the file-leg apps stay
manual — which is the third option by default, without anyone choosing it.
**MEASURED 2026-09-21 (update night).** The hold sentence this question turns on was read verbatim
off a REAL failure rather than from source. `adventurelog v0.12.1 -> v0.13.0` applied nine database
migrations successfully, never bound its port, and held:
> „A(z) adventurelog frissitese 2026-09-21 20:53-kor nem sikerult, es az alkalmazas nem indult el az
> uj verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek.
> Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: **sajat meghajto, 2026-09-21 20:47
> — ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza.**"
(ASCII fragments here; the live page carries its accents.) So the machinery this question's first
option would key on **exists and works**: the sentence names the tier, the date and **what the copy
holds**, unprompted, on a real edge. Whether the AUTOMATIC rule should differ from the button's is
untouched by that and remains the operator's.
**And one thing Q2 did not ask, which tonight makes urgent: after the hold, nobody can find out WHY.**
`failAndHold` removes the containers, so the failing version's own output is gone within seconds
(**R-621**). With a person pressing, they at least watched it happen.
### Q3 — What counts as "within a major" when the tag is not a version number? — ANSWERED 2026-09-23: §3 decisions 13 and 14
*§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`,
`kimai/kimai2:apache-2.57.0`, a date stamp, a digest?*
**And the test is per compose SERVICE, with ALL of them having to pass.** An app bump that is minor
while its `mariadb:` sidecar moves a major is **ACROSS** — that sidecar now converts the customer's
datadir by itself (R-459), so the edge carries a migration whatever the app's own number says.
| option | cost |
|---|---|
| **an unorderable tag on ANY service makes the whole edge ACROSS → human** | the 8 floating pins and every suffix-versioned image stay manual. Conservative, and it is the same rule v0.260.0's `CompareImageRefs` already implements and tests. |
| teach the comparator each shape | every new shape is a new rule, and a wrong rule silently automates a major |
| compare digests instead | needs Q6 first, and a digest carries no order at all — it can say "different", never "newer" |
**Recommendation: the first**, reusing `stacks.CompareImageRefs` rather than writing a second rule.
*One small extension is needed and is named here so it is not discovered late:* v0.260.0's
`CompareImageRefs` answers *orderable?* and *newer?*, which is all R-524 needed. Slice 6 also needs
*same major?*, so the parsed major has to be exposed from the same normaliser — **an addition to the
one comparator, never a second one.**
**If nothing is decided:** Slice 6 would have to invent a rule under time pressure, which is how a
major gets automated by accident.
**MEASURED 2026-09-21 by RUNNING the comparator rather than reading it.** `CompareImageRefs` orders a
reference carrying a `host:port/` prefix correctly — `splitImageRef` takes the LAST colon and rejects
it only when a `/` follows, so a registry port is never mistaken for a tag. Four positive cases and
one negative control (different repositories are not orderable). **This is what made the unattended
hold measurable at all**: the drill edge `localhost:5000/drill/glance:1.0.0 -> :1.0.1` PASSES the
within-a-major test and still fails, which no real catalog move does. The recommendation is
unchanged; the *same major?* extension it already names is still owed.
### Q4 — A held app: who is told, when, and does the box try again? — ANSWERED 2026-09-23: §3 decision 15
*An automatic update that ends HELD happened while everyone was asleep.*
| option | cost |
|---|---|
| **the household on the app page and by mail ONCE; the operator by event; NO retry until the catalog moves again or a person presses** | one mail per held app. The app stays down until someone acts — which is already true of a held update today. |
| retry the next night | a broken edge takes the app down every night and mails every morning; the hold exists precisely because the box cannot fix it |
| tell only the operator | the household finds their app down and has no sentence explaining it |
**Recommendation: the first.** It is what the manual hold already does (`settings.RestoreHold` with
`reason: update_failed`), plus one mail. **If nothing is decided:** the safe default is no automatic
update at all, because a hold nobody is told about is worse than a version nobody moved.
**MEASURED 2026-09-21, and the honest answer is that HALF of this is still unmeasured.** The
unattended night ran (`audits/update-arc-gaps-2026-09-21/09-unattended-night.md`). What it proved:
an app updates itself end to end with nobody pressing anything; one app takes **51 s – 1 m 26 s**
including the health wait; the caller needs **no new controller code**, only the existing guarded
Update plus `UpdateRefusal.Reason` on the wire (v0.261.0, R-609); and **a terminally-refused app is
pressed exactly ONCE and never again** — four apps, three passes, proven.
**What it did NOT produce is a HOLD, and the reason is instructive rather than a failure of the
run.** The only failing edge available was `vikunja → alpine:3.20`, and the caller **correctly
refused to attempt it**: different repositories cannot be ordered, so the edge is "across" and
belongs to a human by decision 3. **The rule that makes automatic updates safe is the same rule that
refuses the obvious way to break one.** Measuring the unattended hold needs an edge that PASSES the
within-a-major test and still fails its health check — same repository, same major, a tag that
starts and does not serve — which probably means a purpose-built image rather than a catalog move.
**So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).**
**MEASURED 2026-09-21 (update night) — and this is the half that was missing.** The caller pressed
ONCE with nobody watching; the app held after **312.9 s**; passes 2 and 3 pressed nothing at all
(`outcomes={'glance': ('held', 312.9)} never_again=['glance']`).
| the question | the answer, measured |
|---|---|
| does an unattended update ever produce a HOLD? | **yes** — 312.9 s, the full health wait plus the phases |
| does the box try again? | **no** — two further passes pressed nothing |
| is the household told? | **on the screen, yes** — the app page, and a banner on EVERY authenticated page carrying every held app at once |
| told what? | what happened, when, **which copy** and **what that copy holds** — all four scored True |
| by MAIL? | **still unmeasured** — the scratch guest runs `hub.enabled: false` and the notifier returns before it logs (**R-620**) |
| in ENGLISH? | **no** — the sentence is Hungarian on the English page (**R-606**, confirmed on the hold sentence itself) |
**So the mechanism this question's recommended option rests on is already there and already behaves
that way.** What remains in Q4 is the MAIL and the ENGLISH, not the hold.
**Two further facts this measurement produced, neither of which the question anticipated.**
**(1) There is NO single-flight** — five Updates pressed within 0.45 s all ran at once and all ended
honest, so a caller pressing N apps runs N updates simultaneously. **(2) A held app keeps inviting
the household to update it and the button then refuses** (`409 reason='held'`), even after the
catalog publishes a FIXED newer version — the household's only route out is the restore. Correct per
§6.1, and the page says otherwise (**R-625**).
**Also proven across a genuine power cut:** the boot sweep met a held app after an unclean shutdown
and deliberately left it alone — *„whatever is holding it owns its recovery"*.
### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`? — ANSWERED 2026-09-23: §3 decision 16
*Eleven templates, and the image performs no conversion — it refuses to start on an older major's
datadir (R-463).*
| option | cost |
|---|---|
| **a scripted `pg_upgrade` edge in the harness, proven on all eleven, before the catalog may move** | real work: eleven fixtures, and `pg_upgrade` needs both major's binaries present. The engine-major gate keeps the rule until it exists. |
| move the pin and let the update HOLD honestly | every one of the eleven apps goes down on the same night and comes back only by a restore |
| never move PostgreSQL majors | the fleet sits on an engine that eventually loses upstream support |
**Recommendation: the first, and the gate stays until it lands.** As of 2026-09-21 the engine-major
rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of it and
`MARIADB_AUTO_UPGRADE=1`); this half is exactly what stays. **If nothing is decided:** nothing breaks
— the gate refuses the move — but the eleven apps drift further from upstream every month.
**MEASURED 2026-09-21, both halves, on a real seeded datadir.**
**(a) What a household would see today — as predicted, and now observed.** The guarded Update of
`postgres:16-alpine -> 17-alpine` ended **`failed` in 5.1 s**; the app was stopped and held; **the pin
named 17 while `installed_images` still said 16 and nothing was running**; the data was intact; and
the restore the hold sentence names brought it back in **29.1 s**. The engine's refusal had to be
REPRODUCED independently, because `failAndHold` destroyed it before any probe could read it
(**R-621**) — *FATAL: database files are incompatible with server / DETAIL: The data directory was
initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir
was still `16` afterwards; the positive control (the same copy under 16) started and held 48 tables.
**(b) The conversion rehearsal, COSTED.** Logical dump and restore, 49 MB / 48 tables:
`pg_dumpall` **2.6 s / 132 201 B**; fresh 17 datadir plus replay **6.5 s / 48 tables restored**; the
app up on 17 saying *Database connection successful*; **the seeded account read back**; **total
155.9 s, of which ~9 s is engine work.** For eleven apps that is a maintenance window, not a project.
`pg_upgrade` was NOT run — it needs both majors' binaries in one image and no such image exists in
this project; the logical route may make it unnecessary at this size. Full paragraph:
`audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`.
**(c) A fact about the INSTRUMENT, not the engine.** `upgrade-test.py`'s PostgreSQL probe is
`cat /var/lib/postgresql/data/PG_VERSION` **inside the container**. Against the converted datadir it
answered `17`, exit 0 — it works. **But it is blind in exactly the case that matters**: when
PostgreSQL refuses, the container is not running, so `docker exec` cannot ask it anything. Tonight it
recorded `No such container`, which its own honesty rule covers — but it must never be read as *the
engine is content*.
**The recommendation is unchanged.** Tonight gives it a price rather than a new opinion.
### Q6 — Should the catalog record each pin's DIGEST at push time? — ANSWERED 2026-09-23: §3 decision 17
*So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.*
**This is no longer theoretical. Measured 2026-09-21: six of the seven measurable floating pins have
been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps),
`postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`,
`postgis:16-3.5-alpine`. On demo-hp today, four apps read „Naprakész" over a database engine image
that has demonstrably moved.
| option | cost |
|---|---|
| **the catalog records the digest at push time; the box compares digests** | one field per pin. `check-image-resolvable.py` already resolves the digest, so the producer exists. §8.1's rule — the box never queries a registry — is untouched. |
| the box queries registries | breaks §8.1 outright: a page that cannot render without eight upstream registries |
| leave it | the badge stays right about the question it asks and wrong about the one a household hears |
**Recommendation: yes.** It is the cheapest real improvement on this list and it closes R-446.
**If nothing is decided:** „Naprakész" keeps meaning "the reference matches", which is measurably not
what it sounds like.
**MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it refines the picture in two ways.**
§8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a
customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins
were read as `installed_images` records them and compared with the upstream digests measured the same
night: `postgres:16-alpine` -> `sha256:721873c34ceb9…` **on both sides**; `redis:7-alpine` ->
`sha256:858f009f9709c…` **on both sides**. **Identical — so „Naprakesz" is TRUE for this box.**
**(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for
a box that pulled BEFORE the tag moved. R-446's six repushed pins measure the tag against the date
the CATALOG set it, which is the right measure for the catalog and not for a box.
**(2) The producer this question needs ALREADY EXISTS on the box.** `installed_images` records a real
`digest` per service — the box knows exactly what it is running. What it cannot do is COMPARE,
because the catalog carries no digest. That is precisely this question's proposal, and only the
catalog half is missing. **The recommendation is unchanged.**
### Q7 — What does the hub's report need to carry for a fleet view? — ANSWERED 2026-09-23: §3 decision 18
*Slice 7 lets the operator SEE and MOVE how far behind every box is.*
**Verified both sides this session:** the controller's report payload carries name, state, CPU and
memory and no image (`controller/internal/report/types.go` L98–103), and the hub's
`Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts**. **But
the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent —
what is missing is the denormalisation and the page, not the transport.
| option | cost |
|---|---|
| **per app: installed reference + catalog reference + badge state; the hub lists boxes behind, with a "move" that is the same guarded Update, operator-triggered** | additive on both sides; the report grows by a few fields per app |
| badge state only | smaller payload; the operator cannot see WHAT is behind, only that something is |
| leave it to per-box pages | free today at two boxes; unusable at twenty |
**Recommendation: the first, and it stays P3-LOW until the fleet grows.** **If nothing is decided:**
the only way to answer "is the fleet current?" is what this session did — read both boxes' files by
hand.
---
### Not a question — already ruled
**R-462's scope was decided on 2026-09-13.** §3 decision 6: the upgrade test goes to **all** apps
through the nightly rotation, explicitly *not* "database apps first". The register row R-462 still
says *"VIKTOR rules on scope"* — **that row is stale and is corrected to cite decision 6.** The update
night below proposes an ORDER *inside* that ruling; it does not reopen it.
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
run, putting the old image tag back produces a container that refuses to start —
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
> downgrading is not supported"*
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
| shape | when it applies | what it does |
|---|---|---|
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
There is no third. **2026-09-23 (§3 decision 15): the UNDO is the two combined** — the abort's old
image plus a restore from the one copy that is seconds old, the pre-pin safety dump. It is not a
third shape and it is not a rollback: nothing is migrated backwards, the pre-migration state is put
back.
### 4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades
`SPIKE-upgrade-test-2026-09-06.md` upgraded three real apps with real data in them and then attempted
the abort on each. **All five real catalog upgrades kept the customer's data.** The abort did not
behave the same way twice:
| app | abort | why |
|---|---|---|
| **docmost** `0.25.3`→`0.95.0` | **REFUSES** | the old code finds migration ledger entries it does not know: *"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*, then *"Failed to run database migration. Exiting program."* |
| **privatebin** `1.7.5`→`2.0.5` | **works** | file-backed, no database, no schema — a major version moves no data |
| **bookstack** app+engine | **works, misleadingly** | only because the MariaDB datadir upgrade was skipped and never happened — see R-459 |
**This puts TWO independent measurements behind the ruling above, by two unrelated mechanisms:**
Nextcloud refused on an explicit version comparison; docmost refuses on its migration ledger. The word
"rollback" was already struck; it is now struck on evidence rather than on one case.
**And it adds a distinction this document did not have: there is no single answer to "can this update
be undone". There are apps where it can and apps where it cannot, and the only way to know which is to
MEASURE THAT APP.** Any design that assumes one answer for all 53 is designing against a fact that was
checked and is false.
Per §5 below, a restore's image-level undo used to have a ≤15-minute half-life because the syncer
overwrote it (**R-441**) — closed in v0.235.0.
---
## 5. The shape, as SHIPPED in v0.235.0
The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the
syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather
than an input, and the thirteen unattended `up -d` paths stop being able to change a version by
accident.
### 5.1 Nothing was added to the thirteen call sites, and that is deliberate
**They are made safe by removing the reason, not by gating them.** The most important of them are
REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`),
the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a
customer's app down, which is worse than the problem this slice solves.** Since the file they act on
no longer changes version, every one of them became safe without being touched.
### 5.2 The pin, and what it is not
`AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref.
**It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a
DECISION — what should run. Letting an observation feed a decision would make a bad reading become a
bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over.
They will normally agree; when they disagree that is a signal, not a bug to paper over.
**Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.**
Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came
from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is
safe from the catalog, and keeping it beside the app means it travels with every path that already
moves a stack dir.
### 5.3 The four writers — the only acts entitled to move a version
| writer | pin source |
|---|---|
| the deploy path (`runComposeDeploy`) | the template just deployed from |
| **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull |
| the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** |
| `Manager.AdoptPins` | the observation, once, and only when complete AND matching |
**`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file
on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards
would pull the frozen version and change nothing — while reporting success, and a button that lies is
worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of
`recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent.
### 5.4 The render table, complete
| app state | result |
|---|---|
| not deployed / protected / seam not wired | the catalog template — today's behaviour |
| deployed, **unpinned** | the catalog template + one DEBUG |
| deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** |
| deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE |
| pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it |
| mid-deploy | the compose file is left alone this cycle |
**`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries
`catalog_since`, which the badge needs. See §8.5.
**The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and
putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration
the older template cannot supply, so an old image under a new template is broken in a way neither
version is.
**And this is not "skip deployed apps".** That option was considered and rejected: it also stops
health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in
the spike §3 — both halves the ruling explicitly kept.
### 5.5 Adoption, and why the startup order matters
`AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app
to what it is already running. **It reads and writes files only** — no container is started, stopped or
touched. It skips, loudly, when the observation is incomplete or when the app runs something the
current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a
guessed pin.
**`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position
that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and
handed the next restart a version change — the exact behaviour this slice removes, once per boot.
### 5.6 The trap this slice set for the previous one
`Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one.
On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find
installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test
still green, because the new field has the same type and shape. The badge now reads
`Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an
earlier feature is the failure mode to look for whenever a file changes meaning.**
---
## 6. The seven slices
| # | slice | status |
|---|---|---|
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02); English since v0.258.0 (R-589); a FOURTH verdict — AHEAD — and the downgrade refusal in v0.260.0 (R-524, §3 decision 10)** |
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 |
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
| **6** | **Automatic updates** — ~~within a major, a human across one~~ **the catalog's tested steps, climbed one at a time, as a leg of the backup chain, undone by the box on failure** (§3 decisions 11–17); **an engine change gets its own edge.** | OPEN — R-450; rulings 2026-09-23; build order §6.4 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451; ruled 2026-09-23 (decision 18), built later |
### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)
`POST /api/stacks/{name}/update` is a guarded job. It answers **202** at once; the outcome exists only
on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`, `update_error`,
`hold_reason`), and `update_phase=done` is written only after the app's health is known. **R-443 is
closed by construction: nothing reports an update complete on the compose exit code.**
**The sequence, and the order is the design:**
| # | phase | what happens | on failure |
|---|---|---|---|
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no copy on ANY tier and no backup can be taken now** (since v0.239.0; before it, no restorable Tier-2 copy) | nothing moves, nothing is recorded |
| 1 | `checking` | walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than `update.backup_max_age` (v0.239.0) | nothing moves |
| 2 | `backing-up` — only when no tier holds a fresh copy | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 | refused with the backup's own error; nothing moves |
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
| 8 | `done` | installed images recorded, journal cleared | — |
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
`update.health_timeout` (default `5m`).
**The precondition is the existing verified backup, not a new copy** (§3 decision 1). It is
`backup.Tier2UnitRestorePoint` — the SAME predicate that permits the destructive „Teljes
visszaállítás", extracted from the backups page rather than copied. **The copy is aged by the last
SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was measured before it was
designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp
bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the
manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).** **SUPERSEDED in
v0.239.0 by §3 decision 8:** `backup.Manager.UpdateRestorePoints` walks all three tiers and the update
leans on the first fresh copy; `Tier2UnitRestorePoint` is still the Tier-2 half and still the page's
predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum
skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit
proven current. The hold stores `copy_tier` and names „második meghajtó" / „saját meghajtó" /
„távoli mentés"; a successful off-site restore now lifts an update hold too.
**The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and
gate as R-379, so every start path that already honoured a restore hold honours this one. A successful
unit restore lifts an update hold (only that kind). **Three unattended paths honoured no hold before
v0.237.0 and now do:** the drive-return gate's restart and boot recreate, and the nightly volume dump
(which ends in `StartStack`). The nightly capture and Tier-2 run skip a held app, so the restore point
the hold text names is never overwritten.
**Crash safety is a journal**, `<data>/update-journal.json`, written before every phase. `RecoverUpdates`
runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back;
after `up` → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by
`ResumeInterruptedUpdates` once the backup side is wired, ending healthy or held.
**THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT.** Whether an old image
starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot
be predicted. So the box never puts the old version back by itself. **The route back is the restore**,
and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where
the harness has proven it, is slice 6's.
> **REPLACED 2026-09-23 by §3 decision 15 — the box UNDOES a failed update itself.** The measurement
> above still stands and is the reason the replacement has a different shape: putting the old IMAGE
> back alone is what refuses (§4.1). The undo puts the old image back **together with the database
> copy taken before the pin moved**, so the old version meets pre-migration data. Spiked by hand on
> 9202 the same day — §6.1a.
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
### 6.1a The undo (decision 15) — SPIKED BY HAND 2026-09-23, not built
Evidence: `audits/update-rulings-2026-09-23/README.md`. Three real migrating edges on 9202, each made
to fail a deliberately wrong probe, each held by today's product, each then undone by hand.
| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (SQLite in a volume) |
|---|---|---|---|
| old version on the migrated data, nothing loaded | **refuses** (migration ledger) | **refuses** (alembic revision) | starts and serves |
| the safety dump | DB only, holds the post-backup write | DB only, holds it | **none — no-op** |
| undo, load + start → healthy | **≈ 16 s** | **≈ 38 s** | ≈ 1 s |
| data written before AND after the backup read back | yes / yes | yes / yes | yes / yes |
**The undo works — and not with the loader the product has.** `ImportDump` over a migrated
PostgreSQL database FAILS (the new version's foreign keys block the dump's own drops); over MariaDB it
succeeds and leaves the new version's tables behind. The load that worked empties the schema and loads
the copy in one transaction. **And a truncated PostgreSQL copy loads with exit 0 into an EMPTY
database** — the copy's completion marker must be checked first, which `ValidateDump` does not do. A
product path that loads a safety dump back exists (`rollbackSafetyDump`), but only the off-site
restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices
them.
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
both demo boxes by the floor alone, with its MinAgent declared.
**Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not
only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute
health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit.
The Tier-2 mirror the hold names survived only because Tier 2 is daily. `backup.Manager.isHeld` now also
answers true for an app a guarded update is moving (`SetUpdatingCheck`).
**v0.240.0 (2026-09-13, evening) — what the afternoon's proof and the first nightly rotation found, fixed.**
Seven rows: removal with backups kept now keeps the Tier-2 RECORD, so the second-drive restore is not
refused over an intact mirror (R-486, P1 — the disaster the second copy exists for); PostGIS/pgvector/
TimescaleDB images are Postgres, so such apps get their logical dump (R-484); "delete backups" deletes
the unit, the mirror(s) and the prefs (R-474/R-466); the backup card sizes them (R-485); a held
update's sentence leaves the card with the hold (R-480); the Tier-3 lookup is one `snapshots` call
(R-477); a unit older than the app's `deployed_at` does not count (R-478). Delivered by the floor in
16 s / 18 s; every row proven live with a throwaway adventurelog. `audits/v0240-2026-09-13/`.
**Proven live on demo-hp, 2026-09-13**, with a throwaway uptime-kuma and real catalog tag changes (each
reverted in the same phase): A (2.3.2→2.4.0, done after health), B (`backup_max_age: 2m` → backup first),
E (non-existent tag → pin back, container untouched), F (`alpine:3.20` → held), H (three buttons and the
boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold
cleared). Live evidence: `audits/slice4-2026-09-13/`.
### The verdict record — the contract Slice 6 carries
Decided here rather than invented twice. The harness writes one of these per edge, beside its
evidence; Slice 6 puts the same shape in the catalog.
```json
{"harness_version": 1, "app": "bookstack",
"from": {"bookstack": "…:25.02.2", "bookstack-db": "mariadb:11.6"},
"to": {"bookstack": "…:26.05.2", "bookstack-db": "mariadb:12.3"},
"verdict": "proven | failed | inconclusive",
"seed_read_before": true, "seed_read_after": true, "healthy_after": true,
"migration_observed": "verbatim log line, or null",
"abort": "starts-and-serves | refuses | starts-data-gone | not-attempted",
"memory": {"soak_s": 600, "containers": {"<name>": {"limit": 0, "peak": 0, "peak_pct": 0.0,
"oom_kills": 0, "restarts": 0}}, "first_kill": null},
"marks": ["memory_tight"],
"abort_detail": "the refusal quoted verbatim, or null",
"duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"}
```
**Harness version 2 (2026-09-23, R-635) adds `memory` and `marks`.** After a successful readback the
harness runs the new version for `--soak` seconds (default 600) under light load and reads the
kernel's own `oom_kill` counter host-side. A kill or a restart turns `proven` into `failed`; a peak
above 80 % of the compose limit adds `memory_tight`. The two fields are the test record's memory half
(decision 13).
**`inconclusive` is a first-class verdict and must never be collapsed into `failed`.** "We could not
measure it" and "it does not work" are different facts, and only one of them is about the app.
**`migration_observed` is a quoted line, never an inference from timing** — the value of both the
Nextcloud and the docmost findings was the exact sentence the app printed.
### Database engines under an upgrade — MEASURED 2026-09-06
**The arc's standing rule that an engine change gets its OWN edge now has measured evidence behind
it**, and the evidence is stronger than the rule's original argument. The rule was justified by
*"two migrations behind one edge is an unreadable failure when it breaks"* — a readability argument.
What was measured is that **an engine change can be applied and silently NOT happen**, which the
app-half edge cannot produce and which no amount of readability would have surfaced:
- `SPIKE-upgrade-test-2026-09-06.md` §4 — MariaDB 12.3 starts on an 11.6 datadir, logs that the
conversion it requires was **skipped**, and serves. **Assigned to the engine half by decomposition:**
the app half alone produces no such line.
- `SPIKE-r459-mariadb-upgrade-2026-09-06.md` — it is **stable but never self-resolving** (5 of 5
restarts, no degradation, and the engine says `Check required!` every time, forever). Converting
properly **succeeds**, costs **7 s**, takes its own system-database backup, and **does not** cost the
ability to abort. **The trade that was expected here does not exist.**
- **2026-09-13 — the setting is in the catalog.** All four `mariadb:` sidecars carry
`MARIADB_AUTO_UPGRADE=1` (operator ruling, §3 decision 5), and `upgrade-test.py`'s engine-state field
now shows the conversion RUNNING on the bookstack edges. **And a gate holds the engines inside their
major until Slice 4:** `app-catalog-felhom.eu/scripts/check-engine-major.py` (R-469).
**Two rules for anything this arc builds around a database engine:**
1. **Ask the engine, not the log.** MariaDB's entrypoint prints `MariaDB upgrade not required` on an
unsupported **downgrade**; `mariadb-upgrade --check-if-upgrade-is-needed` names it exactly
(**R-464**). A cheap instrument built on the log line would report "fine" for the broken case.
2. **The two engines fail in opposite directions, so one check will not do.** MariaDB starts anyway
and skips quietly; **PostgreSQL refuses to start** on a datadir from an older major, and the image
performs no `pg_upgrade`. Eleven templates carry PostgreSQL and **eight sit on `postgres:16-alpine`**
(**R-463**).
**And an engine-state field belongs BESIDE a verdict, never inside it.** `upgrade-test.py` reports
`engine_state_after` next to `verdict`, because an unconverted datadir is not *known* to be a failure
and a verdict that said so would encode an unproven judgement.
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
breaks.
---
### 6.2 Slice 6, as it will be built (OPEN — R-450; ruled 2026-09-23, decisions 11–17)
**The shape the rulings fix. The build order and costs are §6.4.** Rewritten 2026-09-23; the
earlier "as it would be built" draft keyed on the window of Q1's recommendation and on
`CompareImageRefs`, and both were ruled differently.
**Nothing new happens to a single step.** Each step is **exactly the guarded Update of §6.1** — same
precondition, same safety dump, same pin journal, same health wait — with **one change to its end**:
a failed health check runs the **undo** (decision 15, §6.1a) before it holds. Slice 6 adds a
*caller*, a *ladder* and the *undo*; it adds no second update path.
**When — a leg of the chain, not a window of its own (decision 11).** The household's backup window
start W already drives every nightly leg at fixed offsets (`07` §6.1): DB dump at W, Tier 2 at
W+60m, off-site at W+105m, and the full-system backup's gate opens at **W+2h** (`quiesce.go`
`gateOpenOffsetMin = 120`, span to W+6h). The update leg starts **when the off-site leg has
finished** and stops starting new steps **when the full-system backup starts**; whatever is left
waits for the next night. **Measured consequence the builder must face, not discover:** the gap
between "off-site finished" and W+2h is **at most 15 minutes** and is zero on a night the off-site
leg runs long, while one step takes 51 s – 1 m 26 s when it succeeds and ~5 m when it fails (§3b
Q4). So either the full-system gate learns to wait for the update leg (it has a four-hour span to
spend) or the leg gets almost no time. That is a build choice inside decision 11, named in §6.4.
**Which apps — the catalog decides, not the tag (decision 13).** An app qualifies when ALL hold:
1. the per-box switch is on (decision 12; **on by default**);
2. `stacks.CatalogOrder` says **Behind** — never Unknown, never Ahead;
3. the **next step** from the app's installed state carries a **test record** in the catalog —
a step with none is never applied by a box (the catalog gate refuses to publish it, decision 13);
4. the step's **marks** allow it: *needs a person* → never automatic; *files may change* → automatic
only when a fresh copy on some tier holds the app's FILES (`UpdateCopyHolds`), else a person;
5. the app is not held (`held` is terminal until a person acts or the catalog moves, decision 15).
**How far — one step at a time (decision 14).** A box two steps behind applies step A→B, then B→C,
each the full guarded update, each with its own health check and undo. A failed step stops the
ladder for that app. **Measured 2026-09-23:** today one press jumps A → C and B never runs, and the
box cannot see B at all — its catalog clone is `--depth 1` (`sync.go:283`, `:300`; one commit
visible on both demo guests).
**The ladder's format — recommended, not ruled** (`audits/update-rulings-2026-09-23/README.md` Part 2):
an `update_ladder:` list in `.felhom.yml`, one entry per step — `from`/`to` refs per service, the
digest per ref, the test record, the marks — and, for every step but the last, the step's OWN
complete definition in `templates/<app>/steps/<to>.yml`. **Not the git history:** romm's image-moving
commit `15f9ebf` is the definition that OOM-looped on demo-hp; the step that works is its images with
the later `f4eb94f` template, and no commit holds that pair.
**When it fails — undo, then hold only if the undo fails (decision 15).** The household is told on
the app page and by **one** mail, in the box's language (R-606 is a precondition — an automatic
update's mail must be in the household's language); the operator by event; **no retry** until the
catalog moves or a person presses.
**What the household sees.** An event and a line on the app page's timeline, before and after, in both
languages: *„Automatikus frissítés 03:12-kor — sikeres"* / *„— visszaállítva az előző verzióra"* /
*„— megállítva, a másolat 2026-09-20-i"*.
**Where it lives.** A leg in the nightly chain beside the existing ones, calling
`Manager.StartGuardedUpdate`. **It must respect the `isHeld`/`SetUpdatingCheck` interlocks v0.238.1
added** — the nightly capture running *inside* an update's health wait is the defect that release
fixed, and a second unattended caller is exactly the shape that finds it again. And it must run
**one app at a time**: there is no single-flight today (§3b Q4: five Updates pressed within 0.45 s
all ran at once).
**The in-process caller reads `UpdateRefusal.Reason`, and the split is measured, not assumed**
(v0.261.0, R-609): `busy`, `updating`, `deploying`, `migrating` and `self_updating` are **transient**
— try again on the next pass; `held` and `downgrade` are **terminal** — never press that app again
until a person acts; `memory`, `disk` and `no_backup` need a person and should be surfaced, not
retried. A working caller in this exact shape exists as evidence, not product:
`audits/update-arc-gaps-2026-09-21/unattended-caller.py`.
**The per-box switch key is `app_update.unattended`**, default **true** (decision 12).
⚠ **NOT `auto_update` — that name is TAKEN, and by the very thing this must not collide with.**
`self_update.auto_update` / `self_update.auto_update_time` (`config/config.go` L280-281, default
**04:30** at L422) are the CONTROLLER's own update. Two settings with that name, one meaning the
controller and one meaning apps, is the kind of collision that is only discovered by an operator who
turned off the wrong one. The app-scoped key is `app_update.*`. **And `stacks.update_window` is
removed or folded in, never a second window** (decision 11).
### 6.3 Slice 7, as it would be built (OPEN — R-451; ruled 2026-09-23, decision 18 — built later)
Three additive pieces, and the transport already exists (§3b Q7):
1. **controller** — the report's per-app object gains installed reference, catalog reference and badge
state. Additive; an older hub ignores it.
2. **hub** — denormalise those out of the raw report it already stores whole, and list boxes by how
far behind they are.
3. **hub → box** — a "move" button that is the same guarded Update, operator-triggered, through the
existing command path. **Not a second update mechanism**, and not automatic.
Rank stays P3-LOW at two enrolled boxes. It rises with the fleet, and §2 of the state audit is what
that looks like today: the only way to answer *"is the fleet current?"* was to read both boxes' files
by hand.
### 6.4 The build order for the 2026-09-23 rulings (PLAN — each part returns to the operator for go/no-go)
Costed in **CC-evenings** (one evening ≈ one unattended session: build, red-proofs, live proof on 9202,
release). Written from the two spikes and the memory watch of 2026-09-23
(`audits/update-rulings-2026-09-23/`), not from source reading alone. **Risk to customer data** is
what the part can do to a household's data if it is wrong, not how likely that is.
| # | part | rulings / rows | cost | depends on | risk to customer data |
|---|---|---|---|---|---|
| **1** | **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
| **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none |
| **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none |
| **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only |
| **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/<to>.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape |
| **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low |
| **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
| **8** | **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows |
**Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10**, part 11 when the fleet grows. **Total
for 1–10: ≈ 22 evenings.** The automatic caller (7) is deliberately late: it is the only part that
acts with nobody watching, and every part before it is what makes that safe. Parts 8 and 9 are small
and independent and can fill any short evening.
**The one open point the build cannot settle alone — part 7, inside decision 11.** The ruled chain is
*off-site copy → updates → full-system backup*. Today the off-site leg starts at **W+105m** and the
full-system backup's gate opens at **W+2h** (`quiesce.go` `gateOpenOffsetMin = 120`, span to W+6h). So
the update leg has **at most 15 minutes**, and none on a night the off-site copy runs long — while one
step takes ~1 min when it works and ~2–6 min when it fails and is undone.
| option | cost |
|---|---|
| **the full-system backup waits for the update leg, inside its own window; the leg stops starting new steps at W+5h** | the full-system backup starts later on update nights, still inside its four-hour window, with an hour kept; one more interlock between two nightly jobs |
| the leg stops at W+2h as the chain stands | ≤ 15 min a night — about ten steps on a good night, none on a slow one; a box far behind takes weeks to climb |
**Recommendation: the first.** It keeps the ruling's order and its promise that the full-system backup
is never skipped for an update; only the start time inside its existing window moves.
### 6.4.1 (record) The update night — the drill brief that preceded the rulings, costed and re-costed
**The ruling is decision 6: all 53 apps, through the nightly rotation.** This is an ORDER inside that
ruling, not a scope change. The database apps go first because they are the ones where a wrong answer
costs data rather than uptime.
**The real numbers this rests on** (R-462, measured 2026-09-06): a successful edge takes
**6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly 8×, because a negative is
only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**. **Machine
time is not the cost — fixtures are.** Two of the three apps needed a bespoke non-browser seed route,
one needed two attempts and a discarded approach, and one (bookstack) can only ever be half-proven
headlessly (R-460).
| leg | what | cost |
|---|---|---|
| A | the **15 database services** — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus `adventurelog`'s postgis — one edge each, fixture per app | **15–25 CC-hours**, dominated by seed routes; ~30 min machine time at the median; ~25 GB |
| B | ~~one power cut mid-update~~ **DONE 2026-09-21 (R-610)** — measured THREE times, two cut mechanisms, three apps: `pulling` (R-520) and the dangerous post-start case three times over. All ended honest; vikunja's 2.6.0 migration had already run when the power went and the data read back intact. **What remains: a cut landing inside `starting` itself** (it lasts well under a second; needs an in-process fault injector, not a faster shell) | 0 — spent |
| C | one **PostgreSQL `pg_upgrade` rehearsal**, the Q5 edge, on one app before any of the eleven | 3–4 CC-hours |
| D | one **downgrade refusal** | **already done** — v0.260.0, proven live 2026-09-21 |
| E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image |
| F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised |
**RE-COSTED 2026-09-21 FROM THE NIGHT'S REAL NUMBERS** (`audits/DRILL-update-night-2026-09-21.md`):
| leg | status after the update night |
|---|---|
| A — the 15 database services | **LARGELY DONE.** 21 edges across 19 apps walked box-side in one night, including both engines and 8 database-carrying apps. **Machine time was never the cost and is now known: a proven edge took 11-218 s, median ~45 s.** The cost was fixtures, exactly as costed — and the real surprise is that two apps can NEVER be seeded headlessly while the catalog rightly closes their sign-up (R-624) |
| B — the power cut | **COMPLETE.** The two EARLY phases nobody had cut in were cut tonight: `backing-up` (the box recovered and said so) and `safety-dump` (nothing moved, nothing to say). Only a cut inside `starting` itself remains, and it still needs an in-process fault injector |
| C — the PostgreSQL rehearsal | **DONE and COSTED**: ~9 s of engine work, 155.9 s end to end for 49 MB / 48 tables. `pg_upgrade` still owed and may prove unnecessary |
| D — the downgrade refusal | already done, v0.260.0 |
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
| F — the remaining apps | **DONE 2026-09-22 (`audits/DRILL-the-28-2026-09-22.md`): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now **53 of 53 attempted**, not 25.** 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (`outline 1.9.1 -> 1.10.1`, HELD with the right sentence); 2 could not be deployed, one of them (`plant-it`) **by design** — it is `lifecycle: abandoned` and the product's lifecycle gate refused it, proven live for the first time. **The cost is now known and it is not machine time:** three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that `POST /api/backup/run` is BOX-WIDE, so concurrent walks serialise on it. **This leg also added the night's biggest finding**, which no count would have produced: R-630 |
**The words "when it resolves to a container" are v0.262.0's, and they are the whole of R-630.**
Before it, the probe branch had no exit: `findProbeContainer` returning `""` set a message and
looped, while the settle path that judges an app declaring NO check sat in the outer `else`,
unreachable. So a stack whose probe resolved to nothing could only ever time out — and
`failAndHold` then stopped an app whose containers were all healthy. **A stack with no probe is not
healthy and not failing; it is settled on container state (§3), and never a reason to stop a running
app.** Proven live on paperless-ngx 2026-09-22: the identical Update that ended **`failed` at
+313.0 s with the app stopped** now ends **`done` at +53.4 s**
(`audits/v0262-live-2026-09-22/live262.json`). Which container is probed is now decidable too —
exact stack name, then `healthcheck.container`, then a UNIQUE prefix, else nothing with the
candidates logged; the old rule took the FIRST prefix match.
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is
now the first thing Slice 6 has to be safe against, ahead of everything in this table.
**Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C
and E are the ones that unblock a decision; leg A is the one that takes the time.
**Venue:** a scratch host, never a customer box — `demo-hp`'s guest 9202 for the box-side legs, the
harness on DooPlex for the image-side ones.
## 6.5 The drill catalog and the image store — the standing method for update drills
**Why this section exists.** On 2026-09-21 an afternoon session put a deliberately broken image into
the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached
a customer, but the method was wrong and the brief that asked for it said so. This is the method that
replaces it, proven the same night.
**The rule, and it has no exception:** *nothing broken, dummy, cross-repo or engine-major ever enters
the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that,
the leg is skipped and named.*
### The two mechanisms
| | what it is | what it makes possible |
|---|---|---|
| **the drill catalog** | `admin/app-catalog-drill` on Gitea — private, a copy of the live catalog's `main` | a scratch box can be pointed at a catalog where a failing edge is *committable*, because it carries none of the live repo's gates |
| **the image store** | a `registry:2` container on the scratch guest at `127.0.0.1:5000` | an edge that **passes the within-a-major test and still fails** — the one shape a real catalog move cannot produce |
**The image store is not a convenience.** `09` §3b Q4 could not be measured for a year of drills
because the only failing edges available were across-a-major, and the within-a-major rule — correctly
— refuses those before the guarded update is ever reached. *The rule that makes automatic updates
safe is the same rule that refuses the obvious way to break one.* Measuring an unattended HOLD needs
`drill/<app>:X.Y.Z` (the real image, retagged) against `drill/<app>:X.Y.(Z+1)` (a built image that
starts, stays up and never serves) — same repository, same major, plain version tags. A third
flavour, a tag simply **absent** from the store, gives the pull-failure leg.
`stacks.CompareImageRefs` orders a `host:port/` reference correctly: `splitImageRef` takes the last
colon and rejects it only when a `/` follows, so a registry port is never read as a tag. **Proven by
running it**, four positive cases and a negative control, 2026-09-21.
### Creating the drill repo — and the step that mails the operator if you skip it
`POST /api/v1/repos/migrate` is the route that works (the project's Gitea tokens carry
`write:repository` but not `write:user`, so `POST /user/repos` answers 403). **A migrated repo
inherits `has_actions: true`, and the catalog's CI workflow comes with it.** Every drill push then
runs that workflow, it fails — the drill repo is deliberately gate-less — and each failure **mails
`admin@felhom.eu`**. The 2026-09-21 night sent **47 such alarms overnight**, into the same mailbox
that was carrying real off-site alarms at the time (**R-629**).
So creating the drill repo is two acts, not one:
```
POST /api/v1/repos/migrate {"repo_name":"app-catalog-drill","private":true, …}
PATCH /api/v1/repos/admin/app-catalog-drill {"has_actions": false}
```
then read the repo back and quote `has_actions: False` — an alarm channel trained to be ignored is
worse than no alarm channel, and R-168 made CI mail the thing that notices a bypassed gate.
### Pointing a box at the drill catalog — the step that is NOT obvious
**`git.repo_url` alone is inert.** `Syncer.gitCloneOrPull` clones only when the cache has no `.git`;
otherwise it fetches from the remote the clone already stores. The cache directory must be removed as
well, or the box goes on following the live catalog and reports success. Filed as **R-615**; until it
is fixed, the drill procedure is:
1. save `controller.yaml` as `controller.yaml.pre-update-night`;
2. set `git.repo_url` (and `username`/`token` — the drill repo is private);
3. **remove `<data>/catalog-cache`**;
4. restart the controller, sync, **rescan** (R-607: a sync can answer „nincs változás" while the
catalog has moved, and the badge answers from the stale value until the rescan);
5. **three controls, all quoted in the report** — the drill bump appears on the scratch box; the
other boxes' caches are unchanged; the live catalog's `main` hash is unchanged.
### What the drill must leave behind
- `controller.yaml` restored from the saved copy, the controller restarted, and `git.repo_url` **read
back and quoted** as the live catalog.
- The registry container and its volume removed; drill images removed **by name**. Never `prune`.
- The drill repo **kept**, private, reset to the live catalog's `main`, so the next drill starts clean.
- A diff of every `image:` line against the live catalog's `main` — expected: identical.
### The fence
Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no
customer box has credentials for it. The store listens on the guest's loopback only.
## 7. What slices 1 and 2 actually built
### 7.1 The record (slice 1)
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
and writes `app.yaml`:
```yaml
installed_images:
web:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:… # "" if the image was never pulled from a registry
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
```
Three rules, each with its reason:
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
already moved. A record built from it would answer "what will happen next time something runs
`up -d`", which is a different question.
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
Intent refused, observation logged. Refusing to start a customer's app because we could not write
down which version it is trades a real outage for a bookkeeping gap.
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
complete record with a partial one.
**Seeded at startup (v0.234.0).** `Manager.BackfillInstalledImages` runs once at boot, beside the
desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY
on. It only READS containers. **This was not a refinement — without it the feature did not reach a
quiet box at all:** see §8.3, which was written as a known limitation on 2026-09-02 and was a defect
by the next morning.
Two admission rules, and the second is the design:
- **It never overwrites an existing record.** The bring-up paths own updates; this fills gaps only.
- **It refuses to seed a PARTIAL observation.** §7.2's comparison reads a service-count mismatch as
BEHIND, so a degraded or crash-looping app seeded from its visible containers would render
„Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial
because they follow a SUCCESSFUL `up -d`, where a gap is real news and is logged; a backfill meets a
box in whatever state it is in. **Same field, two writers, two different admission rules — that is
deliberate and must not be "made consistent".**
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
### 7.2 The label (slice 2)
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
no new CSS.
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
a test with a companion red-proof.
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
Version strings stay in the logs, the API and the hub.
---
## 8. Known limitations, stated plainly
1. **„Naprakész" can be FALSE for the floating pins, and 2026-09-21 measured HOW false.** The
comparison is reference-to-reference and queries no registry — a customer's box must not depend on
reaching eight upstream registries to render a page. For `postgres:16-alpine`, `mariadb:11.6` and
the others the reference can be identical while the image behind it has moved.
**NUMBERS, 2026-09-21** (`audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3). **The count this document
carried — "23 of 66" — is STALE and matched no definition the catalog supports today.** Recounted
at catalog `18a6d2d8`, with the definition stated so it can be rechecked: a pin FLOATS when its
tag names a version LINE rather than an exact release. Of 66 unique pins, 48 are full `X.Y.Z`,
**6 are two-part lines** (`mariadb:11.4`/`11.6`/`12.3`, `claper:2.5`, `opengist:1.13`,
`wger/server:2.6`) and **4 are major lines** (`postgres:15-alpine`, `postgres:16-alpine`,
`redis:7-alpine`, `postgis:16-3.5-alpine`) — **10 float**. The remaining 8 are exact versions
wearing a variant suffix (`ghost:6.53.0-alpine`, `nextcloud:34.0.1-apache`, …), which do not float
by this definition. Of the 8 database and cache engine pins the sweep measured, 7 were measurable
and **6 have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine`
(6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`. Only `mariadb:11.6` has not.
The 8th, immich's own ghcr build, is UNMEASURED — ghcr exposes no anonymous last-modified. So on
demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.
Digest-level comparison needs the catalog to record the digest at push time — **R-446**, put to the
operator as §3b **Q6**, recommended YES.
2. **~~Nothing enforces `catalog_since`.~~ Enforced by the pre-push hook since 2026-09-13 (R-452, `app-catalog-felhom.eu/scripts/check-catalog-since.py`); CI's shallow clone still skips it out loud.** A commit that moves an `image:` line and forgets the date
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
commit to diff against, so the gate needs a deeper fetch — **R-452**.
3. ~~**The record only appears after the next lifecycle action.**~~ **CLOSED in v0.234.0, and the way
it closed is worth keeping.** This was written on 2026-09-02 as an accepted limitation — *"the
fleet view fills in gradually"*. The operator looked at demo-felhom the next morning and found
OpenGist, up 15 hours, running exactly the catalog pin, showing **nothing at all**. **On a quiet
box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is,
on the quiet installations, not shipped.** `BackfillInstalledImages` now seeds the absences at
startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller
START, so a box between upgrade and its next restart still shows nothing — bounded by one restart
rather than unbounded.
4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches
that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the
`wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken
state. Recorded so it is a choice, not a surprise.
5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
the update badge by withholding `catalog_since`. **R-458.**
6. ~~**The Update button is still unguarded.**~~ **CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1).** It
refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a
safety dump, and holds an app that does not come up. **What stays true:** it still has no automatic
rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) —
that now ends HELD rather than crash-looping behind a green button.
7. ~~**An engine major can be applied without its datadir upgrade, and nothing notices.**~~ **CLOSED
2026-09-13 for MariaDB (R-459):** every `mariadb:` sidecar carries `MARIADB_AUTO_UPGRADE=1`, and the
harness shows the conversion running on the E3/E3b edges (§3 decision 5). **What stays true:** the
PostgreSQL half (R-463) has no equivalent — the image performs no `pg_upgrade` — and the
engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until
Slice 4 gives the Update button a backup.
8. ~~**Only three of 53 apps have ever had an upgrade measured.**~~ **WIDENED 2026-09-21 to 21
EDGES ACROSS 19 APPS** (`audits/DRILL-update-night-2026-09-21.md`), on scratch guest 9202
through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog
carried no test reference at any point: **14 proven, 3 failed, 4 inconclusive**, each app
seeded and read back through its OWN front door with a negative control on every readback.
Ten of the fourteen printed a verbatim migration line. **What stays true:** bookstack is still
only half-provable headlessly (**R-460**), and **two apps cannot be seeded AT ALL** while the
catalog rightly closes their sign-up — vaultwarden (`SIGNUPS_ALLOWED=false`, R-512) and
zipline — which is a permanent ceiling on R-462's scope rather than a fixture nobody has
written (**R-624**). **And one thing this widening FOUND that no count would have:** the
`verifying` phase trusts the `.felhom.yml` probe absolutely, and **three of the 53 templates
name a probe the app does not answer**, so a SUCCESSFUL update of those apps ends by STOPPING
a working app (**R-618**, P1 — tandoor measured serving HTTP 200 on the new version at four
samples across five minutes, then stopped).
**CLOSED 2026-09-22, and the closing changed the numbers above:** the three probes were
corrected (`app-catalog-felhom.eu@793c4fb`), red-proofed live on 9202 in both directions, and a
static gate now refuses a probe that does not match the same service's own compose healthcheck.
**tandoor's edge was re-walked with nothing else changed and ended `done` at +41.1 s** where it
had ended `failed` at +361.9 s — so the tally is **15 proven, 2 failed, 4 inconclusive**, and all
fifteen are on the live catalog since 2026-09-22.
**WHAT THE CLOSING FOUND, and it is the part worth carrying forward:** a wrong probe was the
LOUD failure. Two quiet ones sit beside it. `paperless-ngx` has no container whose name matches
its stack name, so `findProbeContainer` returns nothing and **its probe has never run on any box**
— an absence, with no badge to contradict it (**R-630**). And five more templates cannot be judged
statically at all, one of which (`home-assistant`) is correct only because its check type cannot
fail (**R-631**). **So "the probe is right" is now enforced for 47 of 53 templates and still
unknown for six.**
**AND THE SWEEP'S REAL CEILING, counted rather than felt: 28 of the 53 templates have never been
deployed by any drill** (**R-632**) — the widening above went from 3 apps to 21, and 21 is not 53.
**CLOSED THE NEXT NIGHT, 2026-09-22: all 28 were walked** (`audits/DRILL-the-28-2026-09-22.md`),
so every template in the catalog has now been attempted at least once. **And the walk that closed
it found something the probe work had left open.** `paperless-ngx` has no container whose name
matches its stack name, so `findProbeContainer` returns nothing and its probe has **never run on
any box**. Asked what `verifying` does with no probe to wait on, the answer is the worst of the
three: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three
containers all read `healthy`. The controller names it itself — *`not healthy within 5m0s (last:
no probe container) — stopping and HOLDING the app`* — at **+313.0 s**, front door 404 afterwards.
**So limitation 8 now has two shapes, not one:** a probe that names the wrong target (R-618,
fixed) and **no probe at all** (**R-630, raised to P1**), and the static gate can see the first
but not the second, because there is nothing to compare.
**FIXED 2026-09-22 in controller v0.262.0**, and the fix is in the phase table above: the
no-probe case now settles on container state instead of looping, and the probe TARGET is
decidable (exact name → `healthcheck.container` → a UNIQUE prefix → nothing, candidates logged).
`paperless-ngx` and `immich` carry the explicit field, and the catalog gate REFUSES a probe that
resolves to nothing rather than warning about it. **So limitation 8's two shapes are both closed
in the product**; what remains is that six templates still cannot be judged STATICALLY (R-631,
all five read live and correct) and that a probe can still be right about the port and wrong
about what a 200 means — `romm` answered 200 from nginx while its workers were being OOM-killed
for six hours (**R-635**), which is a third shape again and the reason "the update is guarded"
must never be read as "the new version runs".
**Two more things the same night measured, both about state rather than health:** a `remove` sent
while a restore is still running reports success and leaves a container restarting with a live
public route (**R-633**) — and the product already has exactly that guard for `update` and for
`restore`, which name the blocking operation, but not for `remove`; and an app can be **running,
healthy and serving while recorded as `deployed: false`**, in which state the product refuses to
remove it at all (**R-634**, reproducible alone on `sparkyfitness`).
9. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.
---
## 9. Where the rest lives
| what | where |
|---|---|
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |