Files
felhom.eu/documentation/architecture/09-update-architecture.md
T

157 KiB
Raw Blame History

09 — How an app update works, and what it is becoming

LIVING DOCUMENT. Every slice of the update arc updates this file in the same session. Opened 2026-09-02 with slices 1 and 2. Its absence was R-438: the update mechanism was chosen deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who did not know it was one.

This file carries the REASONING. The register (backlog/OPEN-ITEMS.md) carries the work. The source is the truth. Nothing here is invented: every mechanism claim is cited either to audits/SPIKE-app-update-2026-09-01.md, which measured it live, or to live source at file:symbol.


1. How an update works today, as measured

1.1 The button

Manager.UpdateStack (felhom-controller/controller/internal/stacks/manager.go:1199) is two compose commands and nothing else:

compose pull                       →  compose up -d --remove-orphans

No safety copy. No rollback. No hold. No verification. Confirmed by reading and across six live updates (spike §10 item 6). A pull FAILURE is handled correctly — UpdateStack returns after the failed pull and never reaches up -d, so the running app survives untouched (measured twice, spike §4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.

1.2 The catalog syncer moves the file underneath a deployed app

Syncer.copyTemplates (felhom-controller/controller/internal/sync/sync.go:319) copies docker-compose.yml and .felhom.yml into every stack folder on a 15-minute cycle (internal/config/config.go:351, default 15m). It has no deployed check of any kind. Its only guard is a sha256 content compare in copyIfChanged (sync.go:403) and its only exclusion is app.yaml. The post-sync hook is stackMgr.InjectMissingFields(updated) and nothing else — the sync does not restart anything.

That is why a deployed app's compose file and its running containers can disagree indefinitely. Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).

1.3 Thirteen other paths end in compose up -d

Excluding the three API actions, 13 call sites across 9 files call StartStack or RestartStack, and every one ends in compose up -d against the live compose file (spike §8 — the task that commissioned the spike said five; the count is thirteen). They include the boot reconciler (bootrecon.go:269), the app-stop guard's recovery (appstop_marker.go:283), the drive-return gate (intermediary.go:222), the quiesce restart-after-backup, off-site reconstitution, and every restore path.

So an upgrade can happen with nobody pressing anything — measured, spike §2 variant 1c-ii, where a boot reconciliation started an app on a newer image at 17:55:44Z.

1.4 One fear is measured SMALLER than it was stated

A plain power cut does not upgrade anything. Docker's own restart: unless-stopped puts the existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs up -d — it says so in its own words: no boot-orphaned apps (nothing to start) (spike §2, variant 1c, a positive observable and not an absent log line).

The unattended upgrade needs the narrower precondition: "and the app did not come back." Saying so is more useful than leaving the scarier version standing.


2. What was chosen, and by whom

Manager.RestartStack (internal/stacks/manager.go:1161) carries this comment, and it predates the whole arc:

"Use up -d instead of bare restart so that env vars from app.yaml are injected and any template changes (new images, healthchecks) are picked up. Plain docker compose restart only sends SIGTERM+start to existing containers without re-reading the compose file or env."

So the restart behaviour was chosen, deliberately, and written down. A design decision is not a defect. What was never decided — and is recorded nowhere — is what happens once the catalog syncer moves the file underneath a deployed app, and whether the choice was meant to extend to the thirteen unattended call sites. That gap is R-438, and it stays open: this document records the mechanism; it does not change it.


3. The operator decisions

These are rulings, not proposals. Anything specced against a different assumption is wrong.

2026-09-02

  1. The safety copy is a verified recent backup as a PRECONDITION — not a new copy invented for the update path. The guest-snapshot alternative is to be spiked before anything is designed around it. Context: the existing safety machinery (Manager.writeSafetyDump, internal/backup/offbox_reconstitute.go:207) is database-only, which is the headline of spike §6 — the file half was never priced, and demo-hp is too young a box to price it.

  2. The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is. A customer on the newest version is supported however old that version is. This is why catalog_since exists and why no version string is shown.

  3. Updates are automatic WITHIN a major, never ACROSS one. The cross-major case needs a human, because §4 says it cannot be undone. Second half REPLACED 2026-09-23 by decision 13 (the test decides, not the tag); the first half is confirmed by decision 12.

2026-09-06 — Option 1: freeze the version, keep the fixes flowing. SHIPPED, v0.235.0

  1. An app's version is frozen to what the customer has, and only a deliberate Update moves it — while corrections to its definition keep arriving on the 15-minute cycle exactly as they do today.

    This is the ruling R-447 was blocked on, and it was blocked for a good reason: §2 establishes that RestartStack's use of up -d to pick up template changes was chosen and written down in its own comment. Reversing a chosen behaviour is a decision, not a bug fix.

    What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps". The old behaviour had two halves and only one of them was unwanted:

    half verdict
    a restart/repair silently changes the app's VERSION unwanted — nobody chose it, nobody is told, and §4 says it cannot be undone
    template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes worth keeping — both measured in the spike §3

    So the ruling keeps the second and removes the first. In the operator's own words: while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update.

    It does NOT make the Update button safer. That is slice 4 (R-448), and it is where the backup precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.

2026-09-13 — the database engine finishes its own conversion, and the upgrade test goes wide

  1. DBs should be updated when the app moves, with proper precautions, tests and backoff plans — the operator's own words, ruling on SPIKE-r459-mariadb-upgrade-2026-09-06.md. SHIPPED in the catalog the same day: every mariadb: sidecar (bookstack-db, kimai-db, nextcloud-db, romm-db) carries MARIADB_AUTO_UPGRADE=1; MARIADB_DISABLE_UPGRADE_BACKUP stays unset. Not an image change, so catalog_since does not move. The three precautions, because they are the real content of the ruling:

    1. Proven before it ships — upgrade-test.py re-ran E3 and E3b on the changed template and the engine-state field shows the conversion RAN (mariadb_upgrade_info reads the new version, the engine's own check says nothing further is needed, the entrypoint no longer prints skipped due to $MARIADB_AUTO_UPGRADE), with the seeded data reading back after. C3 still returns failed. Evidence: audits/r459-close-2026-09-13/.
    2. Watched as it lands — the change travelled the real 15-minute cycle to demo-hp: the live compose gained the setting, the sync recreated nothing, and one deliberate restart logged MariaDB upgrade not required with the app serving (same evidence directory).
    3. A rule until Slice 4 is built — the Update button still takes no backup, so no template may move a database-engine image across a major version until R-448 ships. Catalog CLAUDE.md states it; scripts/check-engine-major.py enforces it in the pre-push hook (the CI half cannot, R-452); its removal is tracked as R-469 so it is a deliberate act. The setting is inert until an engine major moves, and precaution 3 keeps it that way.
  2. The upgrade test goes as wide as possible, through the nightly unattended sessions — the ruling on STATUS.md item 11. Not "the ~25 database apps first": all of them, as the nightly rotation reaches them, one fixture per app through the app's own interface. Browser-only apps become reachable when CC runs on the operator's Windows workstation with Chrome — the claude-in-chrome route that DooPlex does not have — so an app recorded inconclusive for want of a headless seed route (bookstack's file half, R-460) is deferred to that venue, not faked.


2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts

  1. A floor carries a release past the vouched golden when the release's MinAgent is declared with it (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it, the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which minagent_header_gate.py now guarantees — goes into the same per-box agent comparison. An undeclared floor above the golden is still HELD, and both floor forms refuse to save one. Why: a controller image is pulled by tag and needs no golden to be delivered; what the floor was missing was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release delivery coexist. Proven live: both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s from the save, the hub logging SERVED … from declared (audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt).

  2. Any backup tier lets an app update (R-475; controller v0.239.0). The precondition takes the first FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site (Tier 3, bounded; unreachable = absent). update.backup_max_age applies to whichever tier is chosen. An app with nothing anywhere is backed up first; it is refused only when no backup can be taken either. The hold names the tier and the date. Tier 2 is required nowhere in the update path. This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a verified recent backup as a precondition, not a new copy) stands. Proven live: audits/rulings-r472-r475-2026-09-13/ 04 (nothing anywhere → backup first), 05 (Tier 1 alone), 07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).

  3. A bind-data app's route back is off-site before its own unit, and the hold says what the copy holds (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit that holds the definition and the database dumps and NOT the files (measured: gokapi restored from „helyi" came back with settings and no data). For such an app the update walks second drive → off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive → own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a customer is never sent to a copy that cannot bring the data back without being told so.

2026-09-21 — decided by CC unattended, operator may reverse

  1. A box AHEAD of the catalog reads „Naprakész", and the guarded Update refuses to move a pin backwards (R-524, controller v0.260.0). One sentence: when the catalog is reverted under a box that already updated, is that a "Frissítés elérhető"? Options: (a) leave it — the label compares for difference, as §5.4 says; (b) show „Naprakész" and let the button still run; (c) show „Naprakész" and refuse the button. Costs: (a) is free and offers a household a downgrade onto a datadir the newer version may have migrated, which §4 says cannot be undone; (b) removes the invitation but leaves the loaded gun; (c) costs one comparison and can, wrongly applied, block a legitimate update. Why (c): the direction was already settled — §3 decision 3 says a version change that cannot be undone needs a human, and this is one. The risk in (c) is bounded by making the Ahead verdict NARROW: every differing service must be orderable AND newer, or the answer falls back to today's behaviour. Reversible, no customer-data risk, and it only ever withholds an act. Implementation: stacks.CatalogOrder, one verdict read by both the badge and UpdatePreflight.

2026-09-23 — the operator answers §3b's seven questions

Operator rulings, not CC decisions. Each answers one question in §3b, which is kept, marked ANSWERED, because the costs written there are the reasoning behind these. Nothing below is built yet — the build order is §6.4, and the two new mechanisms (15 and 14) were spiked the same day before anyone builds them (audits/update-rulings-2026-09-23/).

  1. One maintenance window, and it is the backup window the household already sets (Q1). App updates are one more leg of the nightly chain, after the off-site copy and before the full-system backup. An update not finished when the full-system backup is due waits for the next night. Why: the update then leans on the freshest copy the box ever has, and there is no second window for anyone to set or misread. Replaces: §3b Q1's recommendation of a fixed 02:30–05:00. For the builder: stacks.update_window (config.go L195, default 03:00-05:00) exists and no Go code reads it — it appears only in config.go (field + default), setup/handlers.go L482 (written into a new box's controller.yaml) and configs/controller.yaml.example. Remove it or fold it into the chain; do not build a second window.

  2. Automatic updates stay, behind a per-box switch that is ON by default (Q2). Confirms the first half of decision 3. Today every update is still manual; nothing automatic is built.

  3. The test decides, not the tag (Q3). REPLACES decision 3's second half, "never across a major". The box applies by itself every step the catalog holds, because the catalog holds only tested steps; the size of the version jump does not matter, the test does. The box does not parse tags to decide. The tag rule (stacks.CompareImageRefs) moves to the catalog gate as a push-time safety net: an image move with no test record is refused there. Two exceptions, both marks the test sets on the step: files may change → automatic only when a fresh copy holds the files, else a person (this is Q2's refinement); needs a person → the tester writes why. Why: decision 3 was a proxy. "Within a major" was standing in for "known to work", and the update night measured the proxy wrong in both directions — minor moves that held (adventurelog, outline) and a tag shape it cannot read at all (§3b Q3). The test record is the thing the proxy was guessing at.

  4. The update ladder (Q3). A box more than one step behind climbs one tested step at a time, in order, and never jumps. Each step is the full guarded update of §6.1. Why: a step was tested from the version before it, not from three versions before it; R-40's multi-major jump is what this prevents. This is the "stepping" §6.1 already assigned to slice 6.

  5. The box undoes a failed update itself (Q4). REPLACES §6.1's "the box never puts the old version back by itself". On a failed health check: the old definition and pin back + the database copy taken seconds before the pin moved loaded back + the health check again. The app stays stopped (HOLD) only if the undo itself fails. Why the old ruling does not bind this: the old version refused to start on data the new one had migrated (Nextcloud, docmost — §4). The undo puts the pre-migration data back too, so the old version meets the data it knows. The word "rollback" stays struck; the name is undo. The household is told on the app page and by one mail; the operator by event; no retry until the catalog moves or a person presses. Spiked 2026-09-23 before any build — §6.1a.

  6. PostgreSQL majors are converted by the box (Q5). A guarded-update step: save everything from the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on the test bench before the catalog may move it. The engine-major gate stays until then. Why: the update night costed it at ~9 s of engine work for 49 MB (§3b Q5); moving the pin and letting it hold would take eleven apps down on one night.

  7. The catalog records the image digest of every pin at push time (Q6). The box compares against it; where the catalog carries one, the box pulls that exact image — which also makes a floating tag reproducible, not only the badge honest. Whether compose can pull by a recorded digest while the definition names a tag is a claim for the build plan to verify, not a ruling on mechanism.

  8. Fleet view (Q7). The report carries, per compose service (so the database too), the installed reference, the catalog reference and the badge state. Built later, when the fleet grows.

2026-09-23 (afternoon) — two more operator rulings

  1. The undo's copy method is chosen by a bake-off, not by assumption. Two methods are measured on the same apps — dump and load (the database as a text file, loaded back) and copy the folder (the app's named volumes copied with its containers stopped, put back on failure). The simpler one that passes every case is built; if both pass, the folder copy wins on simplicity unless its downtime or disk cost fails the bake-off's limits. Why: the morning spike found two traps in the dump route (a cut-off file loads as success; a failed MariaDB load is half old, half new) and the folder route had not been measured at all. Result: §6.1a.

  2. On update nights, the full-system backup waits for the update leg, inside its own window (answers R-643 / §6.4 part 7's open point). The leg stops starting new steps at W+5h, so the full-system backup keeps at least one hour of its [W+2h, W+6h) window. Built with §6.4 part 7, not before.

2026-09-23 (night) — one operator word, and two decisions taken by CC unattended

  1. Tonight the catalog may move every app whose within-a-major upstream edge is proven on both venues, each with its test record; nothing else moves (operator word, 2026-09-23 evening). The bar: proven on the test bench (seed through the front door, update, read back, ten-minute memory watch) AND on a scratch box through the real guarded Update; no across-a-major edge, no PostgreSQL engine major, no inconclusive, no app whose seed has no front-door route. This is decision 13 applied, not a new rule. Result: audits/DRILL-night-2026-09-23.md.

  2. The memory watch's memory_tight mark reads the app's OWN memory (anon), not the cgroup peak — decided by CC unattended 2026-09-23 night — operator may reverse. One sentence: when the watch says "tight", should the file cache count? Options: (a) the cgroup memory.peak as built (R-635 follow-up); (b) the anon figure of memory.stat, sampled every 15 s, with the cgroup peak kept beside it. Costs: (a) marks every app that reads files — measured tonight: nextcloud and immich's PostgreSQL at 100 % with 0 kernel kills, because the kernel fills the limit with cache it drops before killing anything — and the gate then demands a raised limit, which inflates mem_limit (a customer-box capacity figure) for no reason; (b) can miss a cache-heavy app whose own memory is tight only if a kill never happens — and a kill or a restart still fails the edge under both options. Why (b): decision 13's mark exists for "does not fit the memory" (RomM, R-635), and only the app's own memory decides that. The cgroup peak stays in every record (memory_cgroup_peak_pct in the ladder entry). R-652.

  3. The push-time test-record gate asks the registry — for moved refs only — decided by CC unattended 2026-09-23 night — operator may reverse. One sentence: may a --fast (hook) gate use the network? Options: (a) no — digests compared only in the slow periodic run; (b) yes, but only for the refs a push MOVES, and an unreachable registry is INCONCLUSIVE (the push is refused until it can ask). Costs: (a) a move whose digest changed between the test and the push is published; (b) a push that moves an image needs the registry — zero requests for every other push. Why (b): decision 17 says the box pulls the recorded digest; recording one the registry no longer serves makes every box's pull of that step fail (Scenario E). app-catalog-felhom.eu/scripts/check-test-record-move.py.

2026-09-24 — two operator rulings

  1. The fleet takes controller v0.267.0 although the chaos hour's stop rule fired — both faults the rule caught (R-658, R-659) predate that release; v0.267.0 causes neither and adds R-640's protection to the same restore path. Floor saved with MinAgent 0.131.0 at 05:12:14Z; both demo boxes on v0.267.0, healthy, 10 s later (audits/ladder-2026-09-24/part0/).

  2. R-659, option A: a held app's page names only a copy that can bring the app back WHOLE; when this box has none, the page says so and that support is informed, and the operator gets an urgent event. The database-only restore under existing files (option B) is NOT built. As built (v0.268.0 + hub v0.122.0): "whole" is read from the restores' OWN refusals, not from what a tier stores — measured from source, an app with declared drive files (DeclaredDriveFileLegs) is refused by both the own-unit and the second-drive unit restore (R-538's guard sits in the shared function), so for those apps only the off-site copy counts (the second-drive gap is R-661). The hold names the newest whole copy with what it holds; with none, hold.update.no_whole_copy, no Mentések button, app_hold_no_whole_copy (critical, operator-only). The operator's English was used with one word dropped („please" — the house rule, i18n_missing_gate.py); the Hungarian verbatim, its formal register recorded against R-516.

2026-09-24 (afternoon) — three more operator rulings

  1. R-661, option A: one action brings a file app back WHOLE from the second drive — the unit (settings + database) and the drive files together. The second drive then counts as a whole copy in decision 25's truth table. R-538's guard stays: this is a new action beside it, not its removal.
  2. R-666, option B: while a held app's page says support is informed, Remove offers only „remove the app, keep my data". The household keeps control; the data stays. The no-whole-copy sentence moves to the informal voice, like every other screen.
  3. An app in a crash loop, or in an out-of-memory storm, is STOPPED by the box, and the household and the operator are told. A press on Start gives it one more try. Operator's words: "if it is in a crashloop, or consuming resources, then yes, definitely stopped at least."

RomM follow-ups, operator-agreed the same day: the test bench watches memory after an update (upgrade-test.py, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4); R-636's louder repeated alarm.

2026-09-24 (night) — two decisions taken by CC unattended

  1. Decision 28's crash loop is ≥ 6 restarts within 10 minutes, counted from Docker's RestartCount — decided by CC unattended 2026-09-24 — operator may reverse. One sentence: how many restarts in what window make a crash loop the box stops? Options: (a) the brief's 10 in 10 min; (b) 6 in 10 min. Costs: (a) misses a STEADY loop — Docker's restart back-off caps it at about one restart a minute (gokapi measured 539 → 546 in 7 min), so a loop that has run for an hour can sit at 9–10 per 10 min forever; (b) could stop a slow first start that restarts 6 times — none in the drill evidence (1,831 harness samples, 40 live containers; the one ≥ 6 case, immich's first-start import, was broken, R-676 watches it), and Start always gives one more try. Why (b): (a) fails the ruling's own case. 08 §6.2; audits/night-2026-09-24/A3/.

  2. An INSTALLED app keeps the image digest it runs until a guarded Update moves it; the sync never moves it — decided by CC unattended 2026-09-24 — operator may reverse. One sentence: when the catalog re-tests a floating tag at a new digest, does the sync write it into an installed app's compose? Options: (a) yes, as v0.269.0 did; (b) no — the sync carries the running digest over, and only the guarded Update (backup, undo copy, health check) renders the new one; a fresh install takes the tested digest. Costs: (a) the next restart (a backup's stop/start, a power cut) pulls a new image with no backup and no undo — measured live on 9202; (b) an installed app runs the older tested image until someone (or the automatic leg) presses Update — the badge says so. Why (b): decision 17 wants the TESTED image pulled, and slice 3's ruling freezes the version until a deliberate Update; (a) broke the second. Shipped as controller v0.269.1 — a second release in the session, instead of shipping (a) to the fleet with the floor. Note 2026-09-30 (R-740) — the cost line above is not what the code does. „(or the automatic leg)": the automatic leg never presses a digest-only change (unattended.go:433–435, skip no_test_record). And the catalog never writes a re-test of a floating tag at a new digest (the writer refuses a step that moves no image), so a same-name upstream fix reaches no box by either route today. Measured, with the options, in R-740; the decision is the operator's (STATUS 2026-09-30). This decision's ruling — the sync never moves an installed app's digest — is unchanged. Note 2026-09-30 evening: decided — decision 52 (option A): the automatic leg takes a same-tag fix once the catalog has re-tested it and written it as a step.

2026-09-25 (night) — three decisions taken by CC unattended, building §6.4 part 7

  1. After W+5h the full-system backup waits only for an automatic step ALREADY in flight, and never past W+5h30m — decided by CC unattended 2026-09-25 — operator may reverse. One sentence: when the leg is still running at W+5h, does the full-system backup start at once or wait for the step? Options: (a) start at once — the backup's quiesce stops an app in the middle of its update's verify, which fails the step into an undo during a backup; (b) wait for the step in flight, capped at W+5h30m. Costs: (a) a spurious undo and a backup of an app mid-change; (b) the full-system backup keeps at least 30 minutes of its window instead of 60 on the worst night; a stuck flag cannot hold it past the cap. Why (b): §6.4.2 point 3 already says "a step already running finishes"; this makes the gate agree with it. quiesce.updateLegDefers, TestD20_UpdateLegDefers.
  2. The controller's own self-update waits for the whole automatic leg, not only for a step in flight — decided by CC unattended 2026-09-25 — operator may reverse. One sentence: may the controller swap itself between two automatic steps? Options: (a) yes, as for manual updates (the lock covers only a step in flight); (b) no, the lock covers the whole leg. Costs: (a) the swap restarts the controller and the rest of the night's leg is lost (it is not resumed); (b) a floor-served controller release waits up to the leg's length (at most until W+5h) on an update night. Why (b): the leg is the night's only update window (decision 11); a release can wait an hour, a household's app cannot wait a day for nothing.
  3. The automatic leg takes ONE step per app per night — follows the operator's own brief of 2026-09-24 ("one app at a time, one tested step per app"), recorded here because §6.4.2 point 6 (a) said "rescan between steps", which reads as climbing several. Cost: an app three steps behind takes three nights. Why: each step then gets a night of the household using it before the next one, and a failure is one step wide. Operator may widen it. TestLeg_OneStepPerAppPerNight.

2026-09-25 — an operator ruling

  1. A controller restart during the night's update leg does NOT resume the leg (operator ruling 2026-09-25, option B of R-686). The apps the leg had not reached wait for the next night. A step already pressed is finished or put back by the guarded update's own journal, exactly as before. Nothing is built: the behaviour shipped in v0.271.0 is now the rule. Why: one night's delay for an app is cheap, and a resume would be one more mechanism acting with nobody watching.

2026-09-25 (evening) — two operator rulings

  1. Part 10 is built now, one app first (operator ruling 2026-09-25 evening, D1 option A). The box converts a PostgreSQL major as a step of the guarded update. docmost is the first app. Each other app needs its own proof on both venues (the test bench and a scratch box through the real guarded Update) before the catalog may move it. The engine-major gate stays for every app without that proof. Why: decision 16 said "each of the eleven apps is proven before the catalog may move it"; one app first makes the first proof small enough to read. As built: §6.4 part 10.
  2. A reinstall over kept data offers a choice (operator ruling 2026-09-25 evening, D2 option A): "use my kept data" (the database from a backup, with the kept files) or "start fresh" (the kept files move to a dated folder; nothing is deleted). And, from the operator's follow-up question: kept data must never be a dead end. The household can see it (read only), load it into a later install, and delete it. Support is for the rare case, not the normal one. Not ruled (D3, in STATUS.md): whether the box ever deletes kept data by itself. Until it is ruled, nothing deletes kept data automatically — only a household's explicit Delete. Where it lives and who deletes it: 07-backup-architecture.md §"Kept data".

2026-09-25 (evening) — two decisions taken by CC unattended, building part 10

  1. docmost converts PostgreSQL 16 → 18, not 16 → 17 — decided by CC unattended 2026-09-25 — operator may reverse. One sentence: which major does the first conversion target? Options: (a) 17 — no mount change; (b) 18 — the data volume's mount moves to /var/lib/postgresql in the same step. Costs: (a) every box converts again later, and docmost's own upstream compose already ships postgres:18 at /var/lib/postgresql; (b) the step changes the mount point — measured: postgres:18 REFUSES (exit 1) even an EMPTY volume at /var/lib/postgresql/data — so the step's definition carries the mount, and the undo must put the old mount back with the old bytes, which it does (it restores the old definition and every volume). Why (b): one conversion is better than two when the app's upstream runs the newer major. Reversible: the catalog could pin 17 instead before any customer box moves. audits/night-2026-09-26/A/README.md A1.
  2. The box loads from a pg_dumpall of the OLD engine, with no error tolerated — decided by CC unattended 2026-09-25 — operator may reverse. One sentence: what does the conversion load into the new engine? Options: (a) the existing safety dump (pg_dump per database, --no-owner --no-privileges); (b) a pg_dumpall taken from the old engine at conversion time, loaded with ON_ERROR_STOP. Costs: (a) is free, and silently drops any role, grant or database setting beyond the bootstrap ones (for docmost both routes gave identical databases — it has one role and one database; the other ten are not measured); (b) one more dump (0.6 s for docmost) and a load that must step around exactly two objects the new engine's entrypoint makes (the bootstrap role and database — measured: loaded raw, exactly two already exists errors). Why (b): decision 16 says "save everything"; the two collisions are removed precisely (the entrypoint's databases dropped only when they hold no table; CREATE ROLE <x>; skipped only for a role that exists — its ALTER ROLE … PASSWORD still runs), so ANY other error stops the load and the undo runs. Proven live: an extension 18 lacks (adminpack) stopped the load and the box undid it. audits/night-2026-09-26/A/README.md A2.
  3. docmost's own memory limit rises 384M → 512M with its PostgreSQL 18 step — decided by CC unattended 2026-09-25 — operator may reverse. One sentence: the bench marked the step memory_tight (the docmost app, not the engine) — raise the limit, or leave docmost unmoved? Options: (a) keep 384M — the move gate refuses a tight step whose limit does not move, so docmost stays on 16; (b) raise to 512M in the same commit (the gate's own remedy, the RomM precedent). Costs: (a) the conversion that was built and proven tonight never reaches a box; (b) +128 MB of reservation per docmost box (mem_limit 768M → 896M), and the mark stays (80.4 % at 512M): measured twice, Node sizes its heap from the limit — 349 MB at 384M, 431 MB at 512M, 0 kills and 0 restarts in both 10-minute watches (~12 000 requests each). Why (b): 91 % of the old limit is too close for a household box, the gate's rule is followed rather than bypassed, and the mark's weakness for such apps is filed (R-693). Reversible.

2026-09-26 — two operator rulings

  1. The box never deletes kept data by itself (operator ruling 2026-09-26, D3 option A). Only the household deletes it, by Delete on the „Megőrzött adatok" page with the app's name typed. No age limit, no automatic clean-up with warnings. Kept data can fill a drive; the drive-full warning names the kept folders and their sizes as space the household can free (as built in v0.274.0). Why: the household decided to keep it; a box that deletes it later reverses their decision without them. Closes D3 in STATUS.md.
  2. adventurelog keeps its world-data download (operator ruling 2026-09-26, R-655 option B). The template does NOT set SKIP_WORLD_DATA=1, so a fresh install still gets the countries/regions data. The move to v0.13.0 is proven with the file already in the adventurelog_media volume (an update re-uses what v0.12.1 left), plus a check that a truncated file is replaced (download-countries --force is the command's own repair). The healthcheck override fix (/usr/bin/node) rides the image move in the same commit. Why: a fresh install without world data is a smaller product; the internet dependence is at first start, which the update's own health wait and undo already cover.

2026-09-26/27 — decisions taken by CC unattended (version-travel brief)

  1. The next PostgreSQL apps: paperless-ngx → 18, tandoor → 17 — decided by CC unattended 2026-09-27 — operator may reverse. One sentence: which major does each app's conversion target? Options: (a) 17 for both — no mount change; (b) 18 for both; (c) per app, from what the app's own upstream runs (decision 37's reasoning). Costs: (a) paperless converts twice (its upstream compose ships postgres:18 at /var/lib/postgresql); (b) tandoor would run a major its own upstream compose (postgres:16-alpine) and its Django (5.2.16) have not documented; (c) two different targets to remember. Why (c): decision 37 said one conversion is better than two WHEN the upstream runs the newer major — paperless's does (18, Django ~5.2.5, psycopg 3); tandoor's does not, and 17 is the newest major its Django documents, with no mount change. Read 2026-09-27 from each upstream's compose and requirements (audits/version-travel-2026-09-26/B/). Reversible: the catalog pins the major.
  2. claper → PostgreSQL 17, calcom → 18 — decided by CC unattended 2026-09-28 — operator may reverse. One sentence: which major does claper's conversion target? Options: (a) 17 — no mount change; (b) 18 — the mount moves to /var/lib/postgresql. Costs: (a) a later second conversion if claper's upstream moves to 18; (b) claper would run a major its own upstream compose (postgres:15, v2.5.0) has not run. Why (a): decision 42's rule — the newer major only when the app's own upstream runs it; claper's does not. Proven on both venues (bench: 32 tables equal, peaks 75.8 %/78.3 %; box 9202: converted through the guarded Update in 7.2 s), catalog 4a249b9, audits/pg-calcom-claper-2026-09-28/. calcom → 18 by the same rule: its upstream compose runs untagged postgres at /var/lib/postgresql (Prisma 6.16.1 in v6.2.0). First its memory limit had to be fixed (768M → 1536M, R-703); then both venues proved it (bench: 122 tables equal, anon 61.5 %; box: converted through the guarded Update in 68 s), catalog 037f956. Reversible: the catalog pins the major.
  3. demo-hp's scheduled restore test restores onto nvme-scratch — operator ruling 2026-09-28 (R-701 option (b)). local-lvm (53.9 GiB, the guest's own pool) cannot hold a 31 GiB restore; the NVMe can (~880 GiB free). Needed, and added with the operator's word: the agent's storage role on /storage/nvme-scratch (user + token) — without it Proxmox refused the restore with 403 Datastore.AllocateSpace. Proven: one restore test passed there in 8m46s, local-lvm untouched. 03-host-agent.md, audits/logins-nvme-2026-09-28/C/.
  4. No app is published with a login a stranger knows — operator ruling 2026-09-28 (R-702 widened). Where the box can set the first admin password, it generates one at install and shows it on the app page (after_install:, controller v0.279.0). Where it cannot, the app stays in the catalog; the install dialog and the app page say what the default login is and to change it at once. Operator: "I don't think we should exclude apps if we can't change the first PW." The per-app audit and status: app-catalog-felhom.eu/FIRST-ADMIN.md.
  5. A setup gate for open-first-run apps is SPIKED before it is built — operator ruling 2026-09-29 ("I was leaning towards B, but let's test A"). The 34 class-4 apps let the first visitor create the admin. Option A: while such an app is not yet set up, the box lets only a person logged in to the household's dashboard reach it; the gate opens when the app's own status says an admin exists, or when the household presses "Done, I set it up". Option B: one fix per app (route (b), R-707). A is built only if the spike passes its exit test in writing; otherwise B continues and the spike's result is the recorded reason. Outcome (2026-09-29): the spike PASSED (audits/login-gate-2026-09-29/ B/B-VERDICT.md, written before any build): on 9202 a stranger never reached a first-setup screen (~530 polls during two installs, 0 app answers), the household passed with its dashboard session in 0.2 s, both probes flipped on the setup, immich's phone-app API worked unchanged once the gate's router was removed, and the dashboard cookie was never widened (a redirect handshake mints a 60-second, one-use token per app host instead). Built in controller v0.280.0 and proven live on immich, n8n, audiobookshelf (probe) and uptime-kuma (button). Costs, stated: ~2 ms per gated request; a gated app answers 500 while the controller is down (closed, not open); a phone app cannot reach a gated app; an app with no probe waits for the household's press, and the press trusts the household. Not covered by the gate: open sign-up after the setup (R-711). Design record: 01-topology-and-trust.md §5.
  6. Open sign-up is closed once an app's first admin exists — operator ruling 2026-09-29 (R-711, option A). After the household's first admin exists, a stranger can no longer make an account; only the admin adds people, from the app's own user page, and the app page says how. Where an app cannot close sign-up, it stays in the catalog and its page says plainly that anyone who finds the address can make an account; STATUS asks the operator about it. Mechanism — decided by CC unattended 2026-09-29, operator may reverse. How does the box close sign-up in apps that keep the switch only in their own admin settings? Options: (a) per-app switches — an env read at start plus a restart, or the app's admin API; (b) a box-side block of the app's own sign-up address once the gate opens, with a household window to let a family member in. Costs: (a) measured impossible for most — opengist and wishlist keep it only in their admin settings (no env, no CLI), vikunja has an env but then no way to add a user except its CLI; (b) one more traefik router per app, and a family member joins through a 15-minute window instead of an in-app invite. Chose (b): it works the same for every app, needs no credentials and no restart, and leaves installed apps untouched (only an app whose gate this box opened gets a block). Built in controller v0.281.0 (internal/stacks/signup_block.go): the block goes up before the gate comes down; a failed write keeps the gate closed. Proven on 9202: 11 apps let a stranger sign up after the setup and refused it with the block; 11 more refuse by themselves; the window let a family member in and closed again. wanderer cannot be gated yet (R-714). Evidence audits/gate-rollout-2026-09-29/.
  7. wanderer stays in the catalog, with its page's warning, until its gate ships — operator ruling 2026-09-29 evening (R-714, option A). Outcome (2026-09-29 evening): no gate — measured: wanderer has ONE database address for its web server and the browser (no internal URL), so a gate would refuse its own server's calls; and it needs none for the first admin (PocketBase's superuser installer needs the one-time link from the server log). Its door is sign-up, in the web AND straight through PocketBase's API. Closed by the household's "Close sign-up now" (decision 49's press, offered because wanderer is never gated): a case-insensitive block on /register and PocketBase's POST /api/collections/(users|_pb_users_auth_)/records (the collection id and a letter-case change both got in before the fix), plus PUBLIC_DISABLE_SIGNUP. Its first step now says: make your account, then close sign-up.
  8. "Close sign-up now" for apps installed before decision 47 — operator ruling 2026-09-29 evening (R-716, option A). The app page of an installed app whose template has a sign-up lock and whose install has none offers one press that applies exactly what a fresh install gets after its setup. The box never applies it by itself: a catalog change never touches an installed app (Part 0 of 2026-09-29 stays the rule). Outcome (2026-09-29 evening, controller v0.282.0): POST /apps/<slug>/close-signup writes a lock record (opened_by: close-signup, never a gate), the block, then the app's own switch (after_setup), which restarts the app once. Pressed on the demo boxes (demo-hp adventurelog and opengist, demo-felhom opengist): before, each served its sign-up; after, every sign-up answered "closed"; adventurelog's own switch went on (its backend restarted ~30 s, same images). Also 2026-09-29 evening (decision 47, second lock): 9 of the 11 apps' OWN sign-up switch is set by after_setup when the gate opens (6 of them also refuse the household's own first account, so never at install), and every block is case-insensitive (termix's router ignores case). Final trick run: 113 tries on 11 apps, 0 got in. Same day, operator: CC changes the admin passwords of demo-hp's installed bookstack and calibre-web and stores them in the operator's credentials file (not in any repo).
  9. Every newly installed app goes off-site by itself when the customer has off-site — operator ruling 2026-09-30 (R-720, option A). It restores the intent of the 2026-09-16 ruling (07 §6: Tier 3 ON for every new customer), which a per-app switch starting OFF had undone. If the apps will not fit the customer's quota, the page says so, names the largest, and the household chooses which stay off-site. Apps already installed are not switched by a release (the 2026-09-29 Part 0 rule); the backups page offers one press to switch them all on. The box never deletes off-site history to make room without the household's choice. Outcome (2026-09-30, controller v0.283.0): measured first — over the quota the box already refused NEW pushes and ran only the ruled retention, and an app whose files would cross the quota went up settings + database only; no history was ever deleted to make room (now pinned). A fresh install switches the app ON (DefaultOffboxOnForNewApp; an earlier choice is kept); older apps get one press on both backup pages; the size card names the three largest. Live on 9202: the press, and a fresh app joining by itself.
  10. ep0: the three whole-guest archives of deleted drill boxes in tester-1's namespace are removed, and only those — operator ruling 2026-09-30 (R-727). With before/after controls on every other namespace. The restore test is fixed to take only the current box's own archives, so an archive of an earlier box in the same namespace can never be tested again. Outcome (2026-09-30): the three archives (2026-09-16T17:27:32Z, 2026-09-16T21:59:54Z, 2026-09-29T19:37:07Z), mapped to their drill boxes by key fingerprint and host record, were forgotten on ep0 one by one; every other namespace byte-identical before and after; chunks go at ep0's own weekly GC. Agent v0.138.0: an archive carries its key FINGERPRINT (not a host id), and the restore test skips one written with another key. Same day, operator: the first real tester gets a NEW customer record (Tester-2, domain sajatfelhom.hu), not tester-1; and CC rewrites the volunteer guide as measured (R-722).

2026-09-30 (evening) — two operator rulings

  1. A box takes, at night, a same-tag security fix of an image — only after the catalog has re-tested that tag at the new digest — operator ruling 2026-09-30 evening (R-740, option A). The re-test is the full method on both venues (bench with the memory watch, box 9202 through the guarded Update, read-back, negative control) and is written by the only writer as a ladder step like any other; the box then takes it like any other step. Start with the database and redis lines. The re-test is ONE command, run monthly. Why: a same-name upstream fix otherwise reached no box at all (R-740: the catalog never recorded a same-tag re-test, and the leg skips a digest-only change); measured that the box's leg presses such an entry with no controller change. Decision 30's cost line ("or the automatic leg") is corrected by the dated note there; decision 30's ruling is unchanged. Outcome (2026-09-30 late): catalog 6a3ead9 — the re-test entry (from == to, digest_from, box_evidence), its gates and decoys, upgrade-test.py --retest, and the ONE command scripts/retest-floating.py. Proven end to end on 9202: docmost at an older redis:7-alpine digest → the re-test → "run tonight's chain now" → the leg pressed the step, the new digest ran, the data read back, the badge went back to current. The box needed no change. No database/redis line differs today; two exact tags do (R-743). The monthly run is a runbook, not a cron job (runbooks/monthly-floating-retest.md).
  2. A box keeps, per app service, the image it runs now and the image before it (the undo's); it deletes older images of that app by itself — operator ruling 2026-09-30 evening (R-736, option A). It never deletes an image that any container (running or stopped) or any installed app's compose still names. Removing an app deletes that app's images under the same rule. Kept data (decision 40) is unaffected — it is data, not images. Why: a remove ran compose down --rmi local, which keeps every registry-pulled image, and no other code deleted any; on 9202 that filled the Docker disk until the box refused an install (R-736). Cost, stated: a restore to a version older than the previous one re-pulls it — as every restore already does (R-698: a backup stores the image's name, not the image). Outcome (2026-09-30 late): controller v0.284.2 (floor 0.284.2; 0.284.0 and 0.284.1 never floored — two wiring faults found live on 9202). The one-time sweep: 9202 26.6 → 5.7 GB, demo-hp 24.3 → 13.5 GB, the N100 6.15 → 6.07 GB, every app healthy. Old CONTROLLER images are outside this decision and stay (R-745).

2026-10-01 — three operator rulings (recorded before the work)

  1. The monthly re-test (decision 52) is run by a CC session the operator starts with the standing brief claude/MONTHLY-security-retest.md (in the planning project) — not by an unattended job — operator ruling 2026-10-01 (R-743, option 1A). Why: the runbook's reasons hold (a fresh bench, 9202 on the drill catalog, a drill reset, pushes to the live catalog — nothing of that should run with nobody watching). STATUS carries a standing line: "Monthly security re-test: last run , next due ".
  2. The monthly re-test covers every app with a proven ladder, not only the database and redis lines — operator ruling 2026-10-01 (R-743, option 2A). --engines-only stays as a switch, not the default. Why: the database containers are not reachable from the internet; the web apps are — a same-name fix to an exposed app matters at least as much. Cost, stated: linuxserver images are rebuilt weekly under the same tag, so their apps come up most months.
  3. A box keeps the controller image it runs and the one before it (the self-update's roll-back target); older controller images are deleted by the same in-use rule as decision 53 — operator ruling 2026-10-01 (R-745, option 3A). Why: about 50 old controller versions sat on each demo box (R-745); a release is ~400 MB unpacked and several ship a day. Registry tags are never deleted by this.

2026-10-01 — decided by CC unattended, operator may reverse

  1. mealie's account lock lasts 1 hour (in practice 1–2 h), not 24 — decided by CC unattended 2026-10-01 (R-747) — operator may reverse. Question: how long may a stranger's five wrong passwords lock the household out of mealie? Options: (a) keep mealie's default 24 h — the strongest guard against guessing, and a stranger who knows the public login (admin / changeme@example.com) locks the household out for a day; (b) SECURITY_USER_LOCKOUT_TIME=1 — the unit's minimum; mealie's hourly job lifts it, so 1–2 h; still 5 tries per lock against guessing; (c) no lock (a huge attempt limit) — strangers cannot lock anyone out, and nothing slows guessing but the hash; (d) rename the login to something unguessable after install — the guard stays and strangers cannot aim it, but it changes what the household types and what the app page shows. Why (b): the brief prefers a short lockout over none; a household would not notice the weaker guard (5 tries per 1–2 h against a generated 30-character first password); (d) changes a promise on the page and stays open as an option. Measured on 9202: 5 wrong → the right password 423 for 120 min, then 200; a wrong one 401 after. Limit, stated: the lock is per account and the name is public, so a stranger can renew it every hour. Catalog a4597cd. Operator, 2026-10-01 (afternoon): kept — mealie stays on the 1-hour lock; no secret login name (option d not taken).

2026-10-01 (late afternoon) — two operator rulings (recorded before the work)

  1. calibre-web: the box generates the admin LOGIN NAME at install and shows it on the app page next to the generated password, so a stranger cannot aim the per-name lock at it — operator ruling 2026-10-01 (R-752, option A). The lock itself stays (the guessing guard is kept). Why: calibre-web locks a NAME for a day after 40 wrong tries (cps/web.py:2218, measured), has no knob for that length, and its default name admin is public. Outcome (same day): catalog e9f50b5 — ADMIN_USER (secret, hex:5, no controller change); proven on 9202; demo-hp renamed by hand. An installed app is given a made-up name by the box's missing-field injection (R-757).
  2. Registry retention: the prune keeps the newest 20 versions of each image, plus every version named by the vouched golden, the controller floor and the vouched agent (golden_version, the floor, agent_version, min_agent) and the controller image the golden names — operator ruling 2026-10-01 (R-750, option A). It never runs on a schedule unless the operator says so; --apply is a person's act. Why: a manual --keep 7 run (HM-024, August) kept about a day of controller releases and protected nothing in use. Outcome (same day): admin/misc-scripts c9d5ed5; live dry-run protected controller 0.285.0, golden 0.285.0, agent 0.138.0 + 0.131.0, hub 0.126.0 (the running hub, added: a version in use) and felhom-samba 1.1.0; nothing deleted.

2026-10-01 (afternoon) — decided by CC unattended, operator may reverse (R-752)

The question for all three: how long may a stranger's wrong passwords, aimed at a PUBLIC login name, keep the household out? Behind the tunnel every visitor has one address (R-753), so no fix may lean on the visitor's address or on any header a client can send. Each measured on 9202 (audits/lockouts-2026-10-01/B/).

  1. wger: lock only the targeted name, for 5 minutes, counted in the database — AXES_LOCKOUT_PARAMETERS=username, AXES_COOLOFF_TIME=5, AXES_HANDLER=axes.handlers.database.AxesDatabaseHandler (catalog 82fff32). Options: (a) keep ip_address — 10 wrong tries lock EVERY member for 30 min (measured: the second member locked too); (b) username, 30 min; (c) username, 5 min; (d) axes off — no guard. Why (c): wger 2.7 hard-codes AXES_RESET_COOL_OFF_ON_FAILURE_DURING_LOCKOUT = True, so every try during a lock restarts it — measured: retrying every 8 minutes kept a 15-minute lock closed for 40+ minutes; at 5 minutes one retry still let the household in at 7.5 min. The guard stays: 10 tries per 5 minutes per name. The database handler answers axes' own warning W001 (the default cache is per process) and keeps the count over a restart.
  2. BookStack: no change — its throttle is 5 tries then 60 s, hard-coded (ThrottlesLogins.php:82,90), keyed email|ip where ip is traefik's (APP_PROXIES empty), so in effect per name. Measured: locked 1.0 min, then the right password works. APP_PROXIES would only move the key to the tunnel's one address (R-753) — no gain.
  3. Grafana: no change — per name, 5 failures in a sliding 5 minutes (loginattemptimpl/login_attempt.go:14, defaults.ini:498-507); a blocked try is not counted and a successful login resets the count. Measured: locked 5.0 min; under a one-try-a-minute trickle the household got in after the burst aged out. The lock shows as "wrong password" to the household (Grafana hides it). Raising the attempt limit would not stop a script and weakens the guard.

calibre-web-automated is NOT decided here: its only long lock (40 tries a day per name, cps/web.py:2218) has no knob for its length, and both fixes cost something the household would notice — operator decision in STATUS.

2026-10-01 (evening) — an operator ruling (recorded before the work)

  1. Login problems are solved by extending the box's OWN gate (option A), not by an identity app on every box (option B: Authentik/Authelia) — operator ruling 2026-10-01 evening. B stays a possible later step: the gate's traefik forwardAuth hook could later point to one. Order: first the box tells visitors apart (R-753) — built; then a SPIKE of a permanent gate with family accounts (Grimmory, MeTube) — measured, not built; a build only after the operator's go. Why: one gate the box already runs and the controller already answers, against a new service with its own database, upgrades and failure modes on every box. Design record: 01-topology-and-trust.md §5.

2026-10-02 (morning) — two operator rulings (recorded before the work)

  1. The permanent family gate: GO — build per audits/permanent-gate-2026-10-01/VERDICT.md §"Build plan", both sessions (controller, then catalog: Grimmory and MeTube published behind it) — operator ruling 2026-10-02 (R-780). The family login never opens the dashboard; every exception is anchored (finding F1) and still needs the app's own login. Why: the spike passed every exit item; one gate the box already runs answers Grimmory's lock (R-775) and MeTube's "no login at all" (R-767) at once. Built 2026-10-02: controller v0.287.0 (floored, golden 0.287.0), catalog fields family_gate / family_gate_except / min_controller and gate check-family-gate.py; Grimmory and MeTube published — audits/family-gate-2026-10-02/.
  2. SparkyFitness stays offered (R-784, option B); the operator asks its author for written permission now; if there is no written permission before the first paying customer, it is hidden (lifecycle: hidden) — operator ruling 2026-10-02. Why, recorded: the operator's model is that Felhom sells installation and care, not software; that holds for OSI licences, but this licence names indirect use "in any … service … resulting in … monetary compensation", so it needs the permission. STATUS carries it as a standing item with that trigger.

2026-09-30 (day) — operator notes, recorded before the work

  • The day brief runs by day. Every backup and automatic-update test is started by hand — the night chain's debug action (POST /api/debug/backup/night-chain), or demo-hp's backup window moved for one test and put back. Nothing waits for a real night.
  • Tester-2's pre-checks (the tunnel route in Cloudflare, the two connect mails) were done by the operator (handoff 2026-09-30 §4).
  • A decision 52 was offered by the reviewer and NOT taken — may a PostgreSQL conversion proof seed through the database when an app has no headless front-door seed? It was to be recorded only if used. It was not needed: all three apps it was written for have a front-door seed after all (outline's self-hosted installation.create, rallly's own sign-up with the e-mail code read from its own table in place of a mailbox, sparkyfitness's better-auth). The number 52 stays free.

Outcome (same day): decision 42's rule applied to the last six PostgreSQL apps from each upstream's compose — rallly, outline and sparkyfitness (from 15) move to 18 (each upstream runs 18), each proven on both venues with one undo case on 9202 (catalog 25ffd89, aeb0cd6, 1666572); zipline, adventurelog and immich stay on 16 (their upstreams run 16, 16 and 14). All eleven PostgreSQL apps are now decided (R-463 closed). audits/pg-last-six-2026-09-30/.


3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed

ALL SEVEN ANSWERED by the operator on 2026-09-23 — §3 decisions 11–18. Q1 → 11; Q2 → 12 and 13's files may change mark; Q3 → 13 and 14; Q4 → 15; Q5 → 16; Q6 → 17; Q7 → 18. Kept below unedited, because the costs and measurements here are the reasoning behind those rulings. Where a ruling differs from the recommendation below, the ruling wins: Q1 (the household's backup window, not a fixed 02:30–05:00), Q3 (the test decides, not CompareImageRefs), Q4 (the box UNDOES before it holds).

These were questions, not rulings. CC did not decide them. Each is one answerable sentence, the options, what each costs, the recommendation, and what happens if nothing is decided. The measurement behind them is audits/UPDATE-ARC-STATE-2026-09-21.md; the short version is that 46 of the catalog's 58 exact pins are behind upstream today and 39 of those are within a major — the population §3 decision 3 already says may move without a human, and nobody presses 39 buttons.

Q1 — When may a box update itself? — ANSWERED 2026-09-23: §3 decision 11

May the box run the guarded Update by itself between 02:30 and 05:00, nightly?

option cost
02:30–05:00 nightly, after the backup legs the update leans on a copy made hours earlier the same night, which is the freshest the box ever has. The app is down for the health wait in the middle of the night.
a weekly window fewer interruptions; a box sits up to 7 days on a version the catalog already moved past, which widens the support window §3 decision 2 runs on
the household picks the window one more setting on a page that already has several, for a choice almost nobody will change

Recommendation: 02:30–05:00 nightly. The DB dump runs 02:30 and restic 03:00 on a demo box, so a window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy of the night without a new mechanism. If nothing is decided: Slice 6 cannot be built at all — every other question below is downstream of this one.

⚠ THE WINDOW CONTAINS 04:30, AND 04:30 IS WHEN THE BOX UPDATES ITSELF. Found 2026-09-21 (R-608) by reading the clock rather than by a failure: the controller self-updates daily at self_update.auto_update_time, default 04:30, and again from MaybeAutoUpdate after ANY hub report once a floor sits above the box — so at any hour, not only at 04:30. Either path restarts the controller container, which is the supervisor of a running app update.

What v0.261.0 now guarantees, so this question can be answered without also solving that one: the two cannot overlap in either direction. The controller defers its own swap while a guarded app update is in flight (retrying on the next report, exactly as it already did for a running backup), and UpdatePreflight refuses self_updating while a swap is in progress. The lock does not latch — a held app does not block the controller's updates for ever. So the window may contain 04:30; the two jobs will queue behind one another rather than meet. What it does NOT do is reorder them: if the operator prefers the box to take its own update first, that is a scheduling choice still open here.

Q2 — May an automatic update run on a bind-data app when no copy holds its FILES? — ANSWERED 2026-09-23: §3 decisions 12 and 13 (the files may change mark)

The button's rule and the automatic rule can differ. Should they?

The mechanism, verified at source this session, because an earlier draft had it backwards: the guard does not refuse these apps. Since v0.239.0/v0.241.0 (§3 decisions 8–9) the Update is refused only when no copy exists on ANY tier and none can be taken. For an app whose data is bind-mounted files, Manager.UpdateTierOrderFor (controller/internal/backup/update_guard.go:136-141) walks second drive → off-site → own unit last, and when the own unit is the copy chosen, UpdateCopyHolds (:145-165) ends the hold sentence with „csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem" — it holds the settings and the database and not the files. So the update proceeds, and the household is told what the copy holds.

With a human pressing, that is an informed choice. With nobody pressing, nobody was informed.

option cost
automatic requires a fresh copy that HOLDS THE FILES; the button keeps today's rule the nine file-leg apps (and any other bind-data app) update automatically only on a box with a second drive or off-site; on a one-drive box they wait for a person. Two rules to hold in one's head.
one rule for both — automatic follows the button simpler; a file-leg app can be updated unattended against a copy that cannot bring its files back, and the sentence saying so is read by nobody
automatic skips bind-data apps entirely simplest; the seven file-leg apps that are behind never move by themselves even when a good copy exists

Recommendation: the first. It is the smallest rule that keeps the promise the hold sentence makes. If nothing is decided: Slice 6 must be built for the safe subset only, and the file-leg apps stay manual — which is the third option by default, without anyone choosing it.

MEASURED 2026-09-21 (update night). The hold sentence this question turns on was read verbatim off a REAL failure rather than from source. adventurelog v0.12.1 -> v0.13.0 applied nine database migrations successfully, never bound its port, and held:

„A(z) adventurelog frissitese 2026-09-21 20:53-kor nem sikerult, es az alkalmazas nem indult el az uj verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek. Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: sajat meghajto, 2026-09-21 20:47 — ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza."

(ASCII fragments here; the live page carries its accents.) So the machinery this question's first option would key on exists and works: the sentence names the tier, the date and what the copy holds, unprompted, on a real edge. Whether the AUTOMATIC rule should differ from the button's is untouched by that and remains the operator's.

And one thing Q2 did not ask, which tonight makes urgent: after the hold, nobody can find out WHY. failAndHold removes the containers, so the failing version's own output is gone within seconds (R-621). With a person pressing, they at least watched it happen.

Q3 — What counts as "within a major" when the tag is not a version number? — ANSWERED 2026-09-23: §3 decisions 13 and 14

§3 decision 3 says automatic within a major, never across. What about postgres:16-alpine, kimai/kimai2:apache-2.57.0, a date stamp, a digest?

And the test is per compose SERVICE, with ALL of them having to pass. An app bump that is minor while its mariadb: sidecar moves a major is ACROSS — that sidecar now converts the customer's datadir by itself (R-459), so the edge carries a migration whatever the app's own number says.

option cost
an unorderable tag on ANY service makes the whole edge ACROSS → human the 8 floating pins and every suffix-versioned image stay manual. Conservative, and it is the same rule v0.260.0's CompareImageRefs already implements and tests.
teach the comparator each shape every new shape is a new rule, and a wrong rule silently automates a major
compare digests instead needs Q6 first, and a digest carries no order at all — it can say "different", never "newer"

Recommendation: the first, reusing stacks.CompareImageRefs rather than writing a second rule. One small extension is needed and is named here so it is not discovered late: v0.260.0's CompareImageRefs answers orderable? and newer?, which is all R-524 needed. Slice 6 also needs same major?, so the parsed major has to be exposed from the same normaliser — an addition to the one comparator, never a second one. If nothing is decided: Slice 6 would have to invent a rule under time pressure, which is how a major gets automated by accident.

MEASURED 2026-09-21 by RUNNING the comparator rather than reading it. CompareImageRefs orders a reference carrying a host:port/ prefix correctly — splitImageRef takes the LAST colon and rejects it only when a / follows, so a registry port is never mistaken for a tag. Four positive cases and one negative control (different repositories are not orderable). This is what made the unattended hold measurable at all: the drill edge localhost:5000/drill/glance:1.0.0 -> :1.0.1 PASSES the within-a-major test and still fails, which no real catalog move does. The recommendation is unchanged; the same major? extension it already names is still owed.

Q4 — A held app: who is told, when, and does the box try again? — ANSWERED 2026-09-23: §3 decision 15

An automatic update that ends HELD happened while everyone was asleep.

option cost
the household on the app page and by mail ONCE; the operator by event; NO retry until the catalog moves again or a person presses one mail per held app. The app stays down until someone acts — which is already true of a held update today.
retry the next night a broken edge takes the app down every night and mails every morning; the hold exists precisely because the box cannot fix it
tell only the operator the household finds their app down and has no sentence explaining it

Recommendation: the first. It is what the manual hold already does (settings.RestoreHold with reason: update_failed), plus one mail. If nothing is decided: the safe default is no automatic update at all, because a hold nobody is told about is worse than a version nobody moved.

MEASURED 2026-09-21, and the honest answer is that HALF of this is still unmeasured. The unattended night ran (audits/update-arc-gaps-2026-09-21/09-unattended-night.md). What it proved: an app updates itself end to end with nobody pressing anything; one app takes 51 s – 1 m 26 s including the health wait; the caller needs no new controller code, only the existing guarded Update plus UpdateRefusal.Reason on the wire (v0.261.0, R-609); and a terminally-refused app is pressed exactly ONCE and never again — four apps, three passes, proven.

What it did NOT produce is a HOLD, and the reason is instructive rather than a failure of the run. The only failing edge available was vikunja → alpine:3.20, and the caller correctly refused to attempt it: different repositories cannot be ordered, so the edge is "across" and belongs to a human by decision 3. The rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. Measuring the unattended hold needs an edge that PASSES the within-a-major test and still fails its health check — same repository, same major, a tag that starts and does not serve — which probably means a purpose-built image rather than a catalog move. So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).

MEASURED 2026-09-21 (update night) — and this is the half that was missing. The caller pressed ONCE with nobody watching; the app held after 312.9 s; passes 2 and 3 pressed nothing at all (outcomes={'glance': ('held', 312.9)} never_again=['glance']).

the question the answer, measured
does an unattended update ever produce a HOLD? yes — 312.9 s, the full health wait plus the phases
does the box try again? no — two further passes pressed nothing
is the household told? on the screen, yes — the app page, and a banner on EVERY authenticated page carrying every held app at once
told what? what happened, when, which copy and what that copy holds — all four scored True
by MAIL? still unmeasured — the scratch guest runs hub.enabled: false and the notifier returns before it logs (R-620)
in ENGLISH? no — the sentence is Hungarian on the English page (R-606, confirmed on the hold sentence itself)

So the mechanism this question's recommended option rests on is already there and already behaves that way. What remains in Q4 is the MAIL and the ENGLISH, not the hold.

Two further facts this measurement produced, neither of which the question anticipated. (1) There is NO single-flight — five Updates pressed within 0.45 s all ran at once and all ended honest, so a caller pressing N apps runs N updates simultaneously. (2) A held app keeps inviting the household to update it and the button then refuses (409 reason='held'), even after the catalog publishes a FIXED newer version — the household's only route out is the restore. Correct per §6.1, and the page says otherwise (R-625).

Also proven across a genuine power cut: the boot sweep met a held app after an unclean shutdown and deliberately left it alone — „whatever is holding it owns its recovery".

Q5 — PostgreSQL: what has to exist before the catalog may move postgres:16 to 17? — ANSWERED 2026-09-23: §3 decision 16

Eleven templates, and the image performs no conversion — it refuses to start on an older major's datadir (R-463).

option cost
a scripted pg_upgrade edge in the harness, proven on all eleven, before the catalog may move real work: eleven fixtures, and pg_upgrade needs both major's binaries present. The engine-major gate keeps the rule until it exists.
move the pin and let the update HOLD honestly every one of the eleven apps goes down on the same night and comes back only by a restore
never move PostgreSQL majors the fleet sits on an engine that eventually loses upstream support

Recommendation: the first, and the gate stays until it lands. As of 2026-09-21 the engine-major rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of it and MARIADB_AUTO_UPGRADE=1); this half is exactly what stays. If nothing is decided: nothing breaks — the gate refuses the move — but the eleven apps drift further from upstream every month.

MEASURED 2026-09-21, both halves, on a real seeded datadir.

(a) What a household would see today — as predicted, and now observed. The guarded Update of postgres:16-alpine -> 17-alpine ended failed in 5.1 s; the app was stopped and held; the pin named 17 while installed_images still said 16 and nothing was running; the data was intact; and the restore the hold sentence names brought it back in 29.1 s. The engine's refusal had to be REPRODUCED independently, because failAndHold destroyed it before any probe could read it (R-621) — FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11. The datadir was still 16 afterwards; the positive control (the same copy under 16) started and held 48 tables.

(b) The conversion rehearsal, COSTED. Logical dump and restore, 49 MB / 48 tables: pg_dumpall 2.6 s / 132 201 B; fresh 17 datadir plus replay 6.5 s / 48 tables restored; the app up on 17 saying Database connection successful; the seeded account read back; total 155.9 s, of which ~9 s is engine work. For eleven apps that is a maintenance window, not a project. pg_upgrade was NOT run — it needs both majors' binaries in one image and no such image exists in this project; the logical route may make it unnecessary at this size. Full paragraph: audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md.

(c) A fact about the INSTRUMENT, not the engine. upgrade-test.py's PostgreSQL probe is cat /var/lib/postgresql/data/PG_VERSION inside the container. Against the converted datadir it answered 17, exit 0 — it works. But it is blind in exactly the case that matters: when PostgreSQL refuses, the container is not running, so docker exec cannot ask it anything. Tonight it recorded No such container, which its own honesty rule covers — but it must never be read as the engine is content.

The recommendation is unchanged. Tonight gives it a price rather than a new opinion.

Q6 — Should the catalog record each pin's DIGEST at push time? — ANSWERED 2026-09-23: §3 decision 17

So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.

This is no longer theoretical. Measured 2026-09-21: six of the seven measurable floating pins have been repushed upstream since the catalog set them — postgres:16-alpine (8 apps), postgres:15-alpine, redis:7-alpine (6 apps), mariadb:11.4, mariadb:12.3, postgis:16-3.5-alpine. On demo-hp today, four apps read „Naprakész" over a database engine image that has demonstrably moved.

option cost
the catalog records the digest at push time; the box compares digests one field per pin. check-image-resolvable.py already resolves the digest, so the producer exists. §8.1's rule — the box never queries a registry — is untouched.
the box queries registries breaks §8.1 outright: a page that cannot render without eight upstream registries
leave it the badge stays right about the question it asks and wrong about the one a household hears

Recommendation: yes. It is the cheapest real improvement on this list and it closes R-446. If nothing is decided: „Naprakész" keeps meaning "the reference matches", which is measurably not what it sounds like.

MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it refines the picture in two ways. §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, docmost's two floating pins were read as installed_images records them and compared with the upstream digests measured the same night: postgres:16-alpine -> sha256:721873c34ceb9… on both sides; redis:7-alpine -> sha256:858f009f9709c… on both sides. Identical — so „Naprakesz" is TRUE for this box.

(1) The defect's size is set by INSTALL AGE, not by the catalog. A floating pin is wrong only for a box that pulled BEFORE the tag moved. R-446's six repushed pins measure the tag against the date the CATALOG set it, which is the right measure for the catalog and not for a box.

(2) The producer this question needs ALREADY EXISTS on the box. installed_images records a real digest per service — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest. That is precisely this question's proposal, and only the catalog half is missing. The recommendation is unchanged.

Q7 — What does the hub's report need to carry for a fleet view? — ANSWERED 2026-09-23: §3 decision 18

Slice 7 lets the operator SEE and MOVE how far behind every box is.

Verified both sides this session: the controller's report payload carries name, state, CPU and memory and no image (controller/internal/report/types.go L98–103), and the hub's Store.SaveReport (hub/internal/store/store.go:965) denormalises only container counts. But the hub stores the raw report JSON whole, so a new controller field lands there the day it is sent — what is missing is the denormalisation and the page, not the transport.

option cost
per app: installed reference + catalog reference + badge state; the hub lists boxes behind, with a "move" that is the same guarded Update, operator-triggered additive on both sides; the report grows by a few fields per app
badge state only smaller payload; the operator cannot see WHAT is behind, only that something is
leave it to per-box pages free today at two boxes; unusable at twenty

Recommendation: the first, and it stays P3-LOW until the fleet grows. If nothing is decided: the only way to answer "is the fleet current?" is what this session did — read both boxes' files by hand.


Not a question — already ruled

R-462's scope was decided on 2026-09-13. §3 decision 6: the upgrade test goes to all apps through the nightly rotation, explicitly not "database apps first". The register row R-462 still says "VIKTOR rules on scope" — that row is stale and is corrected to cite decision 6. The update night below proposes an ORDER inside that ruling; it does not reopen it.

4. The vocabulary ruling — "rollback" is struck

App data CANNOT be rolled back. Measured on Nextcloud (spike §7): once a migration has actually run, putting the old image tag back produces a container that refuses to start —

"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"

with a positive control proving the data is intact, only unreachable by the old version (§7 6d).

So "rollback" must not appear in any spec for this arc. The two shapes actually available are:

shape when it applies what it does
ABORT before anything migrated stop, put the old image back, the app runs again
RESTORE FROM A COPY after a migration ran the data restore is the whole remedy

There is no third. 2026-09-23 (§3 decision 15): the UNDO is the two combined — the abort's old image plus a restore from the one copy that is seconds old, the pre-pin safety dump. It is not a third shape and it is not a rollback: nothing is migrated backwards, the pre-migration state is put back.

4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades

SPIKE-upgrade-test-2026-09-06.md upgraded three real apps with real data in them and then attempted the abort on each. All five real catalog upgrades kept the customer's data. The abort did not behave the same way twice:

app abort why
docmost 0.25.3→0.95.0 REFUSES the old code finds migration ledger entries it does not know: "corrupted migrations: previously executed migration 20260213T085259-notifications is missing", then "Failed to run database migration. Exiting program."
privatebin 1.7.5→2.0.5 works file-backed, no database, no schema — a major version moves no data
bookstack app+engine works, misleadingly only because the MariaDB datadir upgrade was skipped and never happened — see R-459

This puts TWO independent measurements behind the ruling above, by two unrelated mechanisms: Nextcloud refused on an explicit version comparison; docmost refuses on its migration ledger. The word "rollback" was already struck; it is now struck on evidence rather than on one case.

And it adds a distinction this document did not have: there is no single answer to "can this update be undone". There are apps where it can and apps where it cannot, and the only way to know which is to MEASURE THAT APP. Any design that assumes one answer for all 53 is designing against a fact that was checked and is false.

Per §5 below, a restore's image-level undo used to have a ≤15-minute half-life because the syncer overwrote it (R-441) — closed in v0.235.0.


5. The shape, as SHIPPED in v0.235.0

The live docker-compose.yml is DERIVED from a pin recorded in app.yaml — the one file the syncer never touches. The catalog proposes; app.yaml decides; the compose file is an output rather than an input, and the thirteen unattended up -d paths stop being able to change a version by accident.

5.1 Nothing was added to the thirteen call sites, and that is deliberate

They are made safe by removing the reason, not by gating them. The most important of them are REPAIRS — the boot reconciler (bootrecon.go:269), the drive-return gate (intermediary.go:222), the app-stop guard (appstop_marker.go:283). A repair path that refuses to repair leaves a customer's app down, which is worse than the problem this slice solves. Since the file they act on no longer changes version, every one of them became safe without being touched.

5.2 The pin, and what it is not

AppConfig.PinnedImages (app.yaml, pinned_images:), service → image ref.

It is NOT InstalledImages. That field is an OBSERVATION — what containers report. This one is a DECISION — what should run. Letting an observation feed a decision would make a bad reading become a bad deployment, which is the category error desired_state exists to avoid (R-166), one field over. They will normally agree; when they disagree that is a signal, not a bug to paper over.

Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.

Beside it, applied-compose.yml in the stack directory stores the exact definition that pin came from. Syncer.copyTemplates copies exactly docker-compose.yml and .felhom.yml, so that name is safe from the catalog, and keeping it beside the app means it travels with every path that already moves a stack dir.

5.3 The four writers — the only acts entitled to move a version

writer pin source
the deploy path (runComposeDeploy) the template just deployed from
UpdateStack the catalog's current template, written BEFORE the pull
the restore (stackAdapter.RecreateStackDefinitionFromUnit) the recovery unit's captured compose — this closes R-441
Manager.AdoptPins the observation, once, and only when complete AND matching

UpdateStack's ordering is load-bearing, not stylistic. compose pull and up -d act on the file on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards would pull the frozen version and change nothing — while reporting success, and a button that lies is worse than a button that refuses. A failed pin write REFUSES the update, which is the opposite of recordInstalledImages and for the same reason desired_state refuses: this field is intent.

5.4 The render table, complete

app state result
not deployed / protected / seam not wired the catalog template — today's behaviour
deployed, unpinned the catalog template + one DEBUG
deployed, pinned, catalog images equal the catalog template — fixes flow, self-healing works
deployed, pinned, catalog images differ the stored applied definition — frozen WHOLE
pinned, differ, nothing stored the catalog template + one WARN. We cannot freeze what we do not have and must not invent it
mid-deploy the compose file is left alone this cycle

.felhom.yml is copied verbatim in every case — it carries no image, and it carries catalog_since, which the badge needs. See §8.5.

The frozen branch writes a WHOLE file and never a substitution. Taking the new template and putting the old refs back creates a third state nobody chose: wger 2.6 needs a full DB configuration the older template cannot supply, so an old image under a new template is broken in a way neither version is.

And this is not "skip deployed apps". That option was considered and rejected: it also stops health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in the spike §3 — both halves the ruling explicitly kept.

5.5 Adoption, and why the startup order matters

AdoptPins runs once at boot, immediately after BackfillInstalledImages, and pins every deployed app to what it is already running. It reads and writes files only — no container is started, stopped or touched. It skips, loudly, when the observation is incomplete or when the app runs something the current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a guessed pin.

syncer.Start() was moved to after adoption. It fires an immediate sync; at its previous position that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and handed the next restart a version change — the exact behaviour this slice removes, once per boot.

5.6 The trap this slice set for the previous one

Stack.TemplateImages is read from the app's live compose file — which is now the RENDERED one. On a frozen app that file names the OLD version, so web.compareInstalledToTemplate would find installed == template and answer „Naprakész" on exactly the apps that are behind — with every test still green, because the new field has the same type and shape. The badge now reads Stack.CatalogImages, taken from the syncer's own git clone. A feature that silently inverts an earlier feature is the failure mode to look for whenever a file changes meaning.


6. The seven slices

# slice status
1 The box records what it actually installed — app.yaml.installed_images, per compose service, ref + digest + first-seen. SHIPPED, controller v0.233.0 (2026-09-02)
1b Seed the record for apps nobody touches — a startup backfill, so the label is not restricted to apps that happen to get restarted. SHIPPED, controller v0.234.0 (2026-09-03)
2 One badge says whether the app is current — „Naprakész" / „Frissítés elérhető — N napja", from catalog_since. No version number. SHIPPED, controller v0.233.0 + catalog 69761cf (2026-09-02); English since v0.258.0 (R-589); a FOURTH verdict — AHEAD — and the downgrade refusal in v0.260.0 (R-524, §3 decision 10)
3 The compose file becomes DERIVED — the pin in app.yaml wins; the syncer renders instead of copying. SHIPPED, controller v0.235.0 (2026-09-06) — operator ruling §3.4
4 A guarded update — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8) — §6.1
5 An upgrade test that runs again — a harness that upgrades a real app with real data in it and asks the app for the data back. SHIPPED, app-catalog/scripts/upgrade-test.py (2026-09-06) — 7 edges, 3 apps; see §4.1 and §10
6 Automatic updates — within a major, a human across one the catalog's tested steps, climbed one at a time, as a leg of the backup chain, undone by the box on failure (§3 decisions 11–17); an engine change gets its own edge. OPEN — R-450; rulings 2026-09-23; build order §6.4
7 A fleet sweep pipeline — the operator can see, and move, how far behind every box is. OPEN — R-451; ruled 2026-09-23 (decision 18), built later

6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)

POST /api/stacks/{name}/update is a guarded job. It answers 202 at once; the outcome exists only on GET /api/stacks/{name} (updating, update_phase, update_phase_label, update_error, hold_reason), and update_phase=done is written only after the app's health is known. R-443 is closed by construction: nothing reports an update complete on the compose exit code.

The sequence, and the order is the design:

# phase what happens on failure
0 refusals (409, before the intent is recorded) held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's memoryVerdict, releasing the app's own request), disk (fixed 2 GB floor — image size unknown without a registry), no copy on ANY tier and no backup can be taken now (since v0.239.0; before it, no restorable Tier-2 copy) nothing moves, nothing is recorded
1 checking walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than update.backup_max_age (v0.239.0) nothing moves
2 backing-up — only when no tier holds a fresh copy RunAppBackupNow: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 refused with the backup's own error; nothing moves
3 safety-dump WriteUpdateSafetyDump (R-361's undo copy) — before the pin moves refused; nothing moves
4 pinning the previous definition is copied aside and journaled, then the pin advances pin put back
5 pulling compose pull pin and definition PUT BACK — nothing ran (Scenario E)
5a copying (v0.263.0) compose stop, then each NAMED volume cp -a into <volume>.pre-update-<stamp> by a helper that writes a finished-marker last (bind folders never) copies removed, pin put back, the old version started again — nothing new ran
6 starting compose up -d --remove-orphans UNDO (below), HOLD only if the undo fails
7 verifying the .felhom.yml health check through the existing probe when it resolves to a container, else 60 s of every container running and none restarting; bounded by update.health_timeout UNDO (below), HOLD only if the undo fails. Before v0.263.0: stop + HOLD, the pin stays (Scenario F)
7a undoing (v0.263.0) every copy validated (marker) BEFORE anything is poured back; volumes emptied and refilled; definition, pin and the pinned version's .felhom.yml record put back from the job's own copies; up; health with the OLD version's probe HOLD, the sentence prefixed „A frissítés nem sikerült, és az automatikus visszaállítás sem." + the data state (untouched / half / not_started); copies kept
7b undone (v0.263.0) installed images recorded, copies removed, app.yaml last_update_undone, journal cleared —
8 done installed images recorded, copies removed, last_update_undone cleared, journal cleared —

The two knobs (controller.yaml, operator-owned): update.backup_max_age (default 24h) and update.health_timeout (default 5m).

The precondition is the existing verified backup, not a new copy (§3 decision 1). It is backup.Tier2UnitRestorePoint — the SAME predicate that permits the destructive „Teljes visszaállítás", extracted from the backups page rather than copied. The copy is aged by the last SUCCESSFUL Tier-2 copy, not by the unit manifest's created_at, and that was measured before it was designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the manifest, a quiet app would be "stale" forever and a backup-first would not fix it. The predicate is Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475). SUPERSEDED in v0.239.0 by §3 decision 8: backup.Manager.UpdateRestorePoints walks all three tiers and the update leans on the first fresh copy; Tier2UnitRestorePoint is still the Tier-2 half and still the page's predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit proven current. The hold stores copy_tier and names „második meghajtó" / „saját meghajtó" / „távoli mentés"; a successful off-site restore now lifts an update hold too.

The hold is settings.RestoreHold with reason: update_failed and copy_date — the SAME store and gate as R-379, so every start path that already honoured a restore hold honours this one. A successful unit restore lifts an update hold (only that kind). Three unattended paths honoured no hold before v0.237.0 and now do: the drive-return gate's restart and boot recreate, and the nightly volume dump (which ends in StartStack). The nightly capture and Tier-2 run skip a held app, so the restore point the hold text names is never overwritten.

Crash safety is a journal, <data>/update-journal.json, written before every phase. RecoverUpdates runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back; after up → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by ResumeInterruptedUpdates once the backup side is wired, ending healthy or held.

THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT. Whether an old image starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot be predicted. So the box never puts the old version back by itself. The route back is the restore, and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where the harness has proven it, is slice 6's.

REPLACED 2026-09-23 by §3 decision 15 — the box UNDOES a failed update itself. The measurement above still stands and is the reason the replacement has a different shape: putting the old IMAGE back alone is what refuses (§4.1). The undo puts the old image back together with the database copy taken before the pin moved, so the old version meets pre-migration data. Spiked by hand on 9202 the same day — §6.1a.

Not gated here: a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.

6.1a The undo (decision 15) — SPIKED BY HAND, then BUILT: controller v0.263.2 (2026-09-23)

SHIPPED AND PROVEN LIVE on 9202 (audits/undo-live-2026-09-23/README.md): docmost, romm and vikunja each made a real migrating update fail its (deliberately wrong) probe; the product undid all three in 30–52 s, with data written before the backup, after it, and seconds before the press all read back through each app's front door and the ledgers equal; the page carries one line in the request's language. A cut-off copy → HOLD saying the data is as the new version left it; a power cut during the undo → resumed after boot and completed; a person's press after an undo → done. Two defects only the live box could show, fixed the same day: the undo's probe was never asked while the current probe held the app unhealthy (v0.263.1), and the "old" .felhom.yml taken at update time was already the new one, because .felhom.yml flows in on every catalog sync (v0.263.2: the pinned version's file is now recorded in applied-meta/ whenever a version is pinned). Residual (R-646): an app pinned before v0.263.2 has no such record until its next pin. Two more, found by the chaos hour of 2026-09-23 night (audits/DRILL-night-2026-09-23.md Part D): the undo finds an app's volumes by their compose label, and a restore recreates them WITHOUT it — so after any restore the undo copies nothing (R-658, P1); and a held file-leg app is pointed at a restore that refuses a database-only copy, with no route left on a box without an off-site tier (R-659, P1). A controller kill during verifying resumed and undid correctly (round 9). Both fixed in v0.268.0 (2026-09-24): the undo selects volumes from the app's compose definition and the restore labels what it creates (R-658); the hold names only a copy that brings the app back whole, else says support is informed (R-659, decision 25). Proven live on 9202, audits/ladder-2026-09-24/.

The spike, as it was run by hand before any build:

Evidence: audits/update-rulings-2026-09-23/README.md. Three real migrating edges on 9202, each made to fail a deliberately wrong probe, each held by today's product, each then undone by hand.

docmost (PostgreSQL) romm (MariaDB) vikunja (SQLite in a volume)
old version on the migrated data, nothing loaded refuses (migration ledger) refuses (alembic revision) starts and serves
the safety dump DB only, holds the post-backup write DB only, holds it none — no-op
undo, load + start → healthy ≈ 16 s ≈ 38 s ≈ 1 s
data written before AND after the backup read back yes / yes yes / yes yes / yes

The undo works — and not with the loader the product has. ImportDump over a migrated PostgreSQL database FAILS (the new version's foreign keys block the dump's own drops); over MariaDB it succeeds and leaves the new version's tables behind. The load that worked empties the schema and loads the copy in one transaction. And a truncated PostgreSQL copy loads with exit 0 into an EMPTY database — the copy's completion marker must be checked first, which ValidateDump does not do. A product path that loads a safety dump back exists (rollbackSafetyDump), but only the off-site restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices them.

THE COPY METHOD, chosen by the bake-off of 2026-09-23 afternoon (decision 19): COPY THE FOLDER. audits/undo-bakeoff-2026-09-23/README.md. Both methods passed every case on all three apps (seeds before and after the backup, ledger equal, a cut-off copy caught before anything is swapped or loaded). The folder copy wins because an app with no database server gets no dump at all, so the dump route would have needed the folder copy anyway. Shape: after the pull and just before up — where the app is stopped anyway to be recreated — each NAMED volume the app owns is copied with cp -a into a sibling volume <volume>.pre-update-<stamp> by a helper container, which writes a finished-marker last; a bind-mounted user folder is never copied and never touched. On a failed health check the copy is validated (marker present) and put back, the old definition and pin are restored from the journal's own copies, and the old version is checked with the OLD .felhom.yml probe. Measured cost: ≈ 1–5 s extra downtime for the three apps, ≈ 420 MB/s, disk = the volumes' size (the update refuses before moving anything when the copy would breach the 2 GB floor).

The release could not reach the fleet by floor — R-472. The hub holds a controller floor above the vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0 were hand-deployed to the demo guests. RESOLVED by §3 decision 7 (hub v0.112.0): v0.239.0 reached both demo boxes by the floor alone, with its MinAgent declared.

Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not only once it is held. In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit. The Tier-2 mirror the hold names survived only because Tier 2 is daily. backup.Manager.isHeld now also answers true for an app a guarded update is moving (SetUpdatingCheck).

v0.240.0 (2026-09-13, evening) — what the afternoon's proof and the first nightly rotation found, fixed. Seven rows: removal with backups kept now keeps the Tier-2 RECORD, so the second-drive restore is not refused over an intact mirror (R-486, P1 — the disaster the second copy exists for); PostGIS/pgvector/ TimescaleDB images are Postgres, so such apps get their logical dump (R-484); "delete backups" deletes the unit, the mirror(s) and the prefs (R-474/R-466); the backup card sizes them (R-485); a held update's sentence leaves the card with the hold (R-480); the Tier-3 lookup is one snapshots call (R-477); a unit older than the app's deployed_at does not count (R-478). Delivered by the floor in 16 s / 18 s; every row proven live with a throwaway adventurelog. audits/v0240-2026-09-13/.

Proven live on demo-hp, 2026-09-13, with a throwaway uptime-kuma and real catalog tag changes (each reverted in the same phase): A (2.3.2→2.4.0, done after health), B (backup_max_age: 2m → backup first), E (non-existent tag → pin back, container untouched), F (alpine:3.20 → held), H (three buttons and the boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold cleared). Live evidence: audits/slice4-2026-09-13/.

The verdict record — the contract Slice 6 carries

Decided here rather than invented twice. The harness writes one of these per edge, beside its evidence; Slice 6 puts the same shape in the catalog.

{"harness_version": 1, "app": "bookstack",
 "from": {"bookstack": "…:25.02.2", "bookstack-db": "mariadb:11.6"},
 "to":   {"bookstack": "…:26.05.2", "bookstack-db": "mariadb:12.3"},
 "verdict": "proven | failed | inconclusive",
 "seed_read_before": true, "seed_read_after": true, "healthy_after": true,
 "migration_observed": "verbatim log line, or null",
 "abort": "starts-and-serves | refuses | starts-data-gone | not-attempted",
 "memory": {"soak_s": 600, "containers": {"<name>": {"limit": 0, "peak": 0, "peak_pct": 0.0,
            "oom_kills": 0, "restarts": 0}}, "first_kill": null},
 "marks": ["memory_tight"],
 "abort_detail": "the refusal quoted verbatim, or null",
 "duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"}

Harness version 2 (2026-09-23, R-635) adds memory and marks. After a successful readback the harness runs the new version for --soak seconds (default 600) under light load and reads the kernel's own oom_kill counter host-side. A kill or a restart turns proven into failed; a peak above 80 % of the compose limit adds memory_tight. The two fields are the test record's memory half (decision 13).

inconclusive is a first-class verdict and must never be collapsed into failed. "We could not measure it" and "it does not work" are different facts, and only one of them is about the app. migration_observed is a quoted line, never an inference from timing — the value of both the Nextcloud and the docmost findings was the exact sentence the app printed.

Database engines under an upgrade — MEASURED 2026-09-06

The arc's standing rule that an engine change gets its OWN edge now has measured evidence behind it, and the evidence is stronger than the rule's original argument. The rule was justified by "two migrations behind one edge is an unreadable failure when it breaks" — a readability argument. What was measured is that an engine change can be applied and silently NOT happen, which the app-half edge cannot produce and which no amount of readability would have surfaced:

  • SPIKE-upgrade-test-2026-09-06.md §4 — MariaDB 12.3 starts on an 11.6 datadir, logs that the conversion it requires was skipped, and serves. Assigned to the engine half by decomposition: the app half alone produces no such line.
  • SPIKE-r459-mariadb-upgrade-2026-09-06.md — it is stable but never self-resolving (5 of 5 restarts, no degradation, and the engine says Check required! every time, forever). Converting properly succeeds, costs 7 s, takes its own system-database backup, and does not cost the ability to abort. The trade that was expected here does not exist.
  • 2026-09-13 — the setting is in the catalog. All four mariadb: sidecars carry MARIADB_AUTO_UPGRADE=1 (operator ruling, §3 decision 5), and upgrade-test.py's engine-state field now shows the conversion RUNNING on the bookstack edges. And a gate holds the engines inside their major until Slice 4: app-catalog-felhom.eu/scripts/check-engine-major.py (R-469).

Two rules for anything this arc builds around a database engine:

  1. Ask the engine, not the log. MariaDB's entrypoint prints MariaDB upgrade not required on an unsupported downgrade; mariadb-upgrade --check-if-upgrade-is-needed names it exactly (R-464). A cheap instrument built on the log line would report "fine" for the broken case.
  2. The two engines fail in opposite directions, so one check will not do. MariaDB starts anyway and skips quietly; PostgreSQL refuses to start on a datadir from an older major, and the image performs no pg_upgrade. Eleven templates carry PostgreSQL and eight sit on postgres:16-alpine (R-463).

And an engine-state field belongs BESIDE a verdict, never inside it. upgrade-test.py reports engine_state_after next to verdict, because an unconverted datadir is not known to be a failure and a verdict that said so would encode an unproven judgement.

The rule slice 6 inherits, recorded now while it is cheap: an engine change gets its own edge, never bundled with an app version bump. bookstack moved the application and MariaDB 11.6 → 12.3 in one commit (0b73e5e); that is two migrations behind one edge, and an unreadable failure when it breaks.


6.2 Slice 6, as it will be built (OPEN — R-450; ruled 2026-09-23, decisions 11–17)

The shape the rulings fix. The build order and costs are §6.4. Rewritten 2026-09-23; the earlier "as it would be built" draft keyed on the window of Q1's recommendation and on CompareImageRefs, and both were ruled differently.

Nothing new happens to a single step. Each step is exactly the guarded Update of §6.1 — same precondition, same safety dump, same pin journal, same health wait — with one change to its end: a failed health check runs the undo (decision 15, §6.1a) before it holds. Slice 6 adds a caller, a ladder and the undo; it adds no second update path.

When — a leg of the chain, not a window of its own (decision 11). The household's backup window start W already drives every nightly leg at fixed offsets (07 §6.1): DB dump at W, Tier 2 at W+60m, off-site at W+105m, and the full-system backup's gate opens at W+2h (quiesce.go gateOpenOffsetMin = 120, span to W+6h). The update leg starts when the off-site leg has finished and stops starting new steps when the full-system backup starts; whatever is left waits for the next night. Measured consequence the builder must face, not discover: the gap between "off-site finished" and W+2h is at most 15 minutes and is zero on a night the off-site leg runs long, while one step takes 51 s – 1 m 26 s when it succeeds and ~5 m when it fails (§3b Q4). So either the full-system gate learns to wait for the update leg (it has a four-hour span to spend) or the leg gets almost no time. That is a build choice inside decision 11, named in §6.4.

Which apps — the catalog decides, not the tag (decision 13). An app qualifies when ALL hold:

  1. the per-box switch is on (decision 12; on by default);
  2. stacks.CatalogOrder says Behind — never Unknown, never Ahead;
  3. the next step from the app's installed state carries a test record in the catalog — a step with none is never applied by a box (the catalog gate refuses to publish it, decision 13);
  4. the step's marks allow it: needs a person → never automatic; files may change → automatic only when a fresh copy on some tier holds the app's FILES (UpdateCopyHolds), else a person;
  5. the app is not held (held is terminal until a person acts or the catalog moves, decision 15).

How far — one step at a time (decision 14). A box two steps behind applies step A→B, then B→C, each the full guarded update, each with its own health check and undo. A failed step stops the ladder for that app. Measured 2026-09-23: today one press jumps A → C and B never runs, and the box cannot see B at all — its catalog clone is --depth 1 (sync.go:283, :300; one commit visible on both demo guests).

The ladder's format — recommended, not ruled (audits/update-rulings-2026-09-23/README.md Part 2): an update_ladder: list in .felhom.yml, one entry per step — from/to refs per service, the digest per ref, the test record, the marks — and, for every step but the last, the step's OWN complete definition in templates/<app>/steps/<to>.yml. Not the git history: romm's image-moving commit 15f9ebf is the definition that OOM-looped on demo-hp; the step that works is its images with the later f4eb94f template, and no commit holds that pair.

When it fails — undo, then hold only if the undo fails (decision 15). The household is told on the app page and by one mail, in the box's language (R-606 is a precondition — an automatic update's mail must be in the household's language); the operator by event; no retry until the catalog moves or a person presses.

What the household sees. An event and a line on the app page's timeline, before and after, in both languages: „Automatikus frissítés 03:12-kor — sikeres" / „— visszaállítva az előző verzióra" / „— megállítva, a másolat 2026-09-20-i".

Where it lives. A leg in the nightly chain beside the existing ones, calling Manager.StartGuardedUpdate. It must respect the isHeld/SetUpdatingCheck interlocks v0.238.1 added — the nightly capture running inside an update's health wait is the defect that release fixed, and a second unattended caller is exactly the shape that finds it again. And it must run one app at a time: there is no single-flight today (§3b Q4: five Updates pressed within 0.45 s all ran at once).

The in-process caller reads UpdateRefusal.Reason, and the split is measured, not assumed (v0.261.0, R-609): busy, updating, deploying, migrating and self_updating are transient — try again on the next pass; held and downgrade are terminal — never press that app again until a person acts; memory, disk and no_backup need a person and should be surfaced, not retried. A working caller in this exact shape exists as evidence, not product: audits/update-arc-gaps-2026-09-21/unattended-caller.py.

The per-box switch key is app_update.unattended, default true (decision 12).

⚠ NOT auto_update — that name is TAKEN, and by the very thing this must not collide with. self_update.auto_update / self_update.auto_update_time (config/config.go L280-281, default 04:30 at L422) are the CONTROLLER's own update. Two settings with that name, one meaning the controller and one meaning apps, is the kind of collision that is only discovered by an operator who turned off the wrong one. The app-scoped key is app_update.*. And stacks.update_window is removed or folded in, never a second window (decision 11).

6.3 Slice 7, as it would be built (OPEN — R-451; ruled 2026-09-23, decision 18 — built later)

Three additive pieces, and the transport already exists (§3b Q7):

  1. controller — the report's per-app object gains installed reference, catalog reference and badge state. Additive; an older hub ignores it.
  2. hub — denormalise those out of the raw report it already stores whole, and list boxes by how far behind they are.
  3. hub → box — a "move" button that is the same guarded Update, operator-triggered, through the existing command path. Not a second update mechanism, and not automatic.

Rank stays P3-LOW at two enrolled boxes. It rises with the fleet, and §2 of the state audit is what that looks like today: the only way to answer "is the fleet current?" was to read both boxes' files by hand.

6.4 The build order for the 2026-09-23 rulings (PLAN — each part returns to the operator for go/no-go)

Costed in CC-evenings (one evening ≈ one unattended session: build, red-proofs, live proof on 9202, release). Written from the two spikes and the memory watch of 2026-09-23 (audits/update-rulings-2026-09-23/), not from source reading alone. Risk to customer data is what the part can do to a household's data if it is wrong, not how likely that is.

# part rulings / rows cost depends on risk to customer data
1 SHIPPED — controller v0.263.2, proven live on 9202 2026-09-23 (audits/undo-live-2026-09-23/). The undo. Keep the pre-update copies (compose, applied, pin, old .felhom.yml) until the undo is over; in failAndHold: pin back → DB up alone → validate the copy's completion marker → empty-then-load in one transaction (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → health with the OLD probe → undone, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. 15; the audit's 8-point list 4 — HIGH by nature — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. It also makes the manual button safer on its own, which is why it goes first.
2 SHIPPED — controller v0.264.0 + hub v0.120.0, proven live on 9202 and 9201 2026-09-23 (audits/undo-fleet-2026-09-23/). The update sentences in the household's language (R-606) and a mail when an automatic update is undone or held: events app_update_undone (warning) and app_update_held (error), one per app per outcome, on by default, mailed in the household's language with the app named in the subject; per-app cooldown on both legs. Leftovers: R-647. R-606, 15 1 — none
3 SHIPPED — controller v0.264.0, proven live on 9202 2026-09-23. A disabled notifier says so (R-620), so the mail of part 2 can be measured on a scratch box at all. R-620 0.5 — none
4 SHIPPED — catalog 6db08a5, night 2026-09-23 (audits/DRILL-night-2026-09-23.md): update_ladder: in .felhom.yml, one JSON entry per line (spiked on controller v0.266.0 and v0.267.0 first — the controller ignores the key); check-test-record.py (static, CI) + check-test-record-move.py (history + the registry for moved refs only, decision 23); the only writer upgrade-test.py --write-ladder; the 21 moves of 2026-09-22 backfilled from their records (21 proven). Not built: steps/<to>.yml (part 5 needs it once an app has two steps); CompareImageRefs did NOT move to the gate — the gate asks for a proven test instead, which is decision 13's own test. The test record + the catalog gate + the memory check. The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a failed verdict, or one with no memory watch; CompareImageRefs' rule moves here as the push-time safety net. Backfill: one entry per current pin — the 21 proven moves from their records, every other pin needs_person: "never tested", which is honest and keeps them manual. A version move re-checks mem_limit against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). 13, R-635 follow-up 2.5 the memory watch (shipped 2026-09-23) none on a box — catalog-side only
5 SHIPPED — controller v0.268.0 (206b035) + catalog 5ed599c, proven live on 9202 2026-09-24 (audits/ladder-2026-09-24/partD/): romm 5.3.0/11.4 → 5.3.1/11.4 (its app step, from steps/90dd9d68258286ef.yml) → 5.3.1/11.8 (the engine step; mariadb-upgrade ran) in two presses, the seeded account read back after each, the page's „Hátralévő frissítési lépések" 2 → 1 → none. Step files: templates/<app>/steps/<StepKey(to)>.yml (sha256 of to as canonical JSON, 16 hex; stacks.StepKey = ladder.step_key), refused absent or wrong by check-test-record.py rule 4, written by --write-ladder when a step is superseded, 8 backfilled from the NEWEST commit naming each step's images (romm's from f4eb94f, not the OOM-looping 15f9ebf). The box reads them from its --depth 1 clone (the whole tree is there — measured). One press = one step; a missing step file refuses before anything moves; an installed version matching no entry jumps, logged by name (measured live: vikunja 2.5.0). Limitation: a step has no .felhom.yml of its own (R-664). The ladder on the box. Read update_ladder: from the clone, find the installed step, apply ONE step with its OWN definition (steps/<to>.yml, the last step the current template), repeat next night; a failed step stops the ladder for that app. CatalogOrder compares refs with the digest stripped (see part 7). 14 2.5 1, 4 medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape
6 BOTH HALVES SHIPPED. BOX HALF — controller v0.269.0 + v0.269.1, proven live on 9202 2026-09-24 (audits/night-2026-09-24/B/): the compose that runs pins tag@sha256 from the ladder entry that tested those refs; pins and records stay digest-free; the badge reads behind for a newer TESTED digest; an installed app keeps its digest until a guarded Update moves it (v0.269.0 let the sync move it — found live, fixed in v0.269.1, CarryDigests); Compose accepts the form for all 25 ladder apps. CATALOG HALF SHIPPED — night 2026-09-23: every ladder entry carries the digest per to ref (scripts/image_digest.py, stdlib; equals Docker's RepoDigests on a box), and the move gate refuses a digest the registry no longer serves. The box half (compare, render name:tag@sha256) is not built. Digests. The catalog records sha256 per pin at push time (check-image-resolvable.py already resolves it); the box compares it for the badge and renders name:tag@sha256:… when present. Measured 2026-09-23 on 9202: Docker and Compose both pull and run redis:7-alpine@sha256:858f…, and refuse a digest that does not exist (audits/update-rulings-2026-09-23/70-…). Build trap, read from source: splitImageRef returns "unorderable" for ANY ref containing @ (updateorder.go:134), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. 17, R-446 2 4 (the entry carries the digest) low
7 SHIPPED — controller v0.271.0 (9cf13a3), proven on 9202 over four simulated nights 2026-09-24/25 and watched through the demo boxes' first real automatic night (audits/DRILL-night-2026-09-25.md). Decisions 31–33 taken unattended. Gaps a scratch box cannot close: R-687; no resume after a restart: R-686. (Before:) SPIKED 2026-09-24 (night, Part C) — the measured build brief is §6.4.2 below. The update leg in the chain + the automatic caller + the switch. A leg that starts when the off-site leg has ENDED (on every exit path, skips included), one app at a time, one step per press, app_update.unattended default ON (decision 12), stacks.update_window removed, reads UpdateRefusal.Reason, remembers a failed step so it never re-presses it, and the full-system backup's gate waits for it until W+5h (decision 20). 11, 12, 14, 15, 20 3 1, 2, 5 medium — the only part that acts with nobody watching; everything above is what makes it safe
8 SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23 (audits/cleanup-2026-09-23/; a kernel oom_kill counter, not the sticky flag — 08 §6.2). R-636 — the same OOM key re-firing escalates instead of staying one warning for six hours. R-636 1 — none
9 SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23 (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 held unchanged). R-625 — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. R-625 0.5 1 none
10 SHIPPED FOR DOCMOST — controller v0.273.0 + catalog 6a4a5f0 (harness v4, the gate's proof clause), proven live on 9202 2026-09-25 (audits/night-2026-09-26/D/): one press converted docmost 16 → 18 in 9.9 s of engine work (42 s end to end), the check equal (2 databases, 48 tables, 71 rows), PG_VERSION 18, the seed read back; the load failing (an extension 18 lacks) → undone in 43.6 s; the app unhealthy on 18 → undone in 148 s (90 s drill health timeout); the controller SIGKILLed one second after the volume was emptied → the restart undid it in 40 s; each ending on 16 with the seed read back. The old datadir's copy is kept until a backup is proven after the conversion. Decisions 35, 37, 38. Every other PostgreSQL app stays refused by the gate until its own two-venue proof. (Before:) PostgreSQL majors converted by the box. A guarded-update step: pg_dumpall from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. 16, R-463 2 + 3 1 (the same load discipline), 4 HIGH — it rebuilds the datadir; bounded by keeping the old datadir aside
11 Fleet view — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. 18, R-451 2 — none — deferred by the ruling until the fleet grows

Dated note 2026-09-30 — how current the catalog is (audits/catalog-currency-2026-09-30.md, measured from the registries, the 2026-09-21 method). 25 of 53 apps were behind upstream inside a major and 19 across one (37 behind in some way) that morning. The night rule, read from the leg's code (legCandidate): an app can update itself at night when the catalog's ladder holds a proven step from the box's pin, not marked needs_person (and, for a files_may_change step, only with a fresh whole copy). 28 apps + nextcloud (conditional) in the morning; 31 + nextcloud after the day's moves (rallly, outline, sparkyfitness got their first steps; bookstack and kimai new heads); 21 have no ladder at all and are never updated at night. Part 10 is now decided for all eleven PostgreSQL apps (8 moved, 3 stay); part 7's gate wait (decision 20) was seen live on demo-hp the same day (R-687 item 4).

Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10, part 11 when the fleet grows. Total for 1–10: ≈ 22 evenings. The automatic caller (7) is deliberately late: it is the only part that acts with nobody watching, and every part before it is what makes that safe. Parts 8 and 9 are small and independent and can fill any short evening.

The one open point the build cannot settle alone — part 7, inside decision 11. The ruled chain is off-site copy → updates → full-system backup. Today the off-site leg starts at W+105m and the full-system backup's gate opens at W+2h (quiesce.go gateOpenOffsetMin = 120, span to W+6h). So the update leg has at most 15 minutes, and none on a night the off-site copy runs long — while one step takes ~1 min when it works and ~2–6 min when it fails and is undone.

option cost
the full-system backup waits for the update leg, inside its own window; the leg stops starting new steps at W+5h the full-system backup starts later on update nights, still inside its four-hour window, with an hour kept; one more interlock between two nightly jobs
the leg stops at W+2h as the chain stands ≤ 15 min a night — about ten steps on a good night, none on a slow one; a box far behind takes weeks to climb

Recommendation: the first. It keeps the ruling's order and its promise that the full-system backup is never skipped for an update; only the start time inside its existing window moves.

6.4.2 Part 7 — the build brief, from measurement (night 2026-09-24, Part C)

Evidence: audits/night-2026-09-24/C/ and C-00…C-07. Spike only — no product code for the caller was written. Read with decisions 11, 12, 14, 15 and 20.

1. The chain, as it runs today (source + a real night). The legs are clock-scheduled from one window start W and NOTHING waits on the leg before it: db-dump at W, Tier 2 at W+60m, off-site at W+105m (backupwindow.go:16-21, cmd/controller/main.go:1048/1134/1267); the full-system backup's gate is a poll loop that opens at W+2h for 4 h (quiesce.go:656-657, window check quiesce.go:238-245). Measured on demo-hp 9201's real night (W = 02:30): the off-site leg ran 04:15:05 → 04:18:33 CEST (3 min 28 s, ok); the gate opened at 04:30. So a leg placed after the off-site leg has ~11 minutes before the gate — decision 20's wait is required, not optional.

2. The signal that the off-site leg has ended. There is no event. The nearest thing is the persisted offbox status in settings.json — LastStatus (running at start, ok/incomplete/ error at the end) with LastRun beside it, written by UpdateOffboxStatus (offbox.go:1003-1125). The status travels with the timestamp (presence-is-not-success holds), but it is not a reliable "finished" signal: five early returns (offbox.go:863-897 — not configured, escrow pending, orphaned, migration, another backup running) and the scheduler's own skip (main.go:1268-1271) write nothing, and a box with no off-site target (9202 tonight) never runs the leg at all. Build: CHAIN, do not poll. The offbox-backup job function calls the update leg after RunOffboxBackup returns, on EVERY path including every skip — one call site, no new signal to go stale. A box with no off-site target runs the update leg at W+105m.

3. The gate's interlock (decision 20). Insert after the window check in quiesce.runOnce, as an Options func beside WindowStartFn (quiesce.go:75-95): update leg active and now < W+5h → defer (log and return, exactly like the window deferral; the loop re-polls every 5 min). The leg itself stops STARTING steps at W+5h; a step already running finishes (≤ health timeout + undo). Manual whole-box runs (TriggerNow) bypass it, as today.

4. One simulated night (9202, v0.269.1; the caller = tools/unattended_caller_ladder.py, which presses only the public Update). Four apps: wishlist 1 step behind, navidrome 2, romm 3, vikunja 1 with a failing step. Measured per step: navidrome 10.6 s and 8.5 s (no database); wishlist 55.8 s (its first update: the backup ran); romm 77.4 s, 71.2 s, 93.9 s (database + volume dump each step); vikunja's failing step undone in 104 s with the drill's 90 s health timeout — with the default 5 min it is ≈ 6–11 min (verify up to 5 min, the undo's own verify up to 5 min). The whole leg: 7 min 18 s for four steps and one undo. The household's pages: badges „Naprakész" for the three that climbed, and for vikunja the undone sentence in both languages („…A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el." / "…put back the previous version and its data automatically — nothing was lost."); event app_update_undone (dropped on 9202 — no hub, as designed). During the db-dump leg every press was refused busy (transient, retried next pass).

5. What the spike found that the build must handle (rows filed):

  • ladder_steps_left and the badge are STALE after a step ends done until the next scan — the first caller run re-pressed a current app four times (R-678). The leg rescans after every step, and the update's own finish should refresh them.
  • An Update pressed on an app already at the head runs the whole guarded update — backup, pull, restart, done — with no refusal (R-679: navidrome restarted four times for nothing). The leg never presses an app whose pin equals the head; the preflight should refuse with reason current.
  • The product does not remember a failed step. After vikunja's undo the badge still offers the same step; only the caller's memory kept it from a re-press. Decision 15 needs it on the box: the failed to is recorded per app and the leg skips it until the catalog's ladder changes.

6. The build, in order (≈ 3 evenings, unchanged): (a) the leg as a function called from the off-site job on every path; one app at a time, one step per press, rescan between steps, stop at W+5h; (b) the failed-step record + current refusal (R-679); (c) the gate's UpdateLegActiveFn; (d) the per-box switch app_update.unattended, default ON (decision 12), and stacks.update_window removed; (e) the live proof: this night's four apps, with the window moved forward through the product's own form (POST /backups/window — measured: the three legs reschedule at once).

6.4.2 as BUILT (v0.271.0) — where the build differs from the brief above, named: (a) the leg is stacks.RunUpdateLeg, chained by chainUpdateLeg inside the offbox-backup job (every path, a panic included); one step per app per night (decision 33 — not "rescan between steps" in one night); (b) the failed-step record is app.yaml failed_update_step tied to the ladder's print (R-680); R-679's current refusal shipped in v0.270.0; (c) the gate interlock is quiesce.Options.UpdateLegFn with decision 31's in-flight grace; (d) the switch is settings.json app_update.unattended, absent = ON, a card on the settings page; stacks.update_window removed; (e) the live proof: audits/DRILL-night-2026-09-25.md Part C (four nights on 9202) and Part D (the demo boxes). The demo boxes' first real automatic night (2026-09-25): demo-felhom's leg ran at 04:15:46 after the off-site copy and took opengist 1.13 → 1.15 in 20 s; demo-hp's found nothing to take. Both summaries reached the hub. Found live, not in the brief: the leg presses the next app in the same second the previous step ends, so "a controller killed BETWEEN two apps" lands inside the next step's early phases — which the guarded update puts back (Scenario G); the rest of the night is lost (R-686).

6.4.1 (record) The update night — the drill brief that preceded the rulings, costed and re-costed

The ruling is decision 6: all 53 apps, through the nightly rotation. This is an ORDER inside that ruling, not a scope change. The database apps go first because they are the ones where a wrong answer costs data rather than uptime.

The real numbers this rests on (R-462, measured 2026-09-06): a successful edge takes 6.4 s – 305.1 s, median 71.8 s; a FAILING edge takes 556 s, roughly 8×, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost 5.07 GB. Machine time is not the cost — fixtures are. Two of the three apps needed a bespoke non-browser seed route, one needed two attempts and a discarded approach, and one (bookstack) can only ever be half-proven headlessly (R-460).

leg what cost
A the 15 database services — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus adventurelog's postgis — one edge each, fixture per app 15–25 CC-hours, dominated by seed routes; ~30 min machine time at the median; ~25 GB
B one power cut mid-update DONE 2026-09-21 (R-610) — measured THREE times, two cut mechanisms, three apps: pulling (R-520) and the dangerous post-start case three times over. All ended honest; vikunja's 2.6.0 migration had already run when the power went and the data read back intact. What remains: a cut landing inside starting itself (it lasts well under a second; needs an in-process fault injector, not a faster shell) 0 — spent
C one PostgreSQL pg_upgrade rehearsal, the Q5 edge, on one app before any of the eleven 3–4 CC-hours
D one downgrade refusal already done — v0.260.0, proven live 2026-09-21
E the automatic night MOSTLY DONE 2026-09-21 (R-611) — the success night and the no-retry proof both measured. What remains: the unattended HOLD, which needs an edge that passes the within-a-major test and still fails health (see Q4) ~1 CC-hour + a purpose-built image
F the remaining 38 apps, through the nightly rotation as decision 6 directs ~1 app/night; fixtures amortised

RE-COSTED 2026-09-21 FROM THE NIGHT'S REAL NUMBERS (audits/DRILL-update-night-2026-09-21.md):

leg status after the update night
A — the 15 database services LARGELY DONE. 21 edges across 19 apps walked box-side in one night, including both engines and 8 database-carrying apps. Machine time was never the cost and is now known: a proven edge took 11-218 s, median ~45 s. The cost was fixtures, exactly as costed — and the real surprise is that two apps can NEVER be seeded headlessly while the catalog rightly closes their sign-up (R-624)
B — the power cut COMPLETE. The two EARLY phases nobody had cut in were cut tonight: backing-up (the box recovered and said so) and safety-dump (nothing moved, nothing to say). Only a cut inside starting itself remains, and it still needs an in-process fault injector
C — the PostgreSQL rehearsal DONE and COSTED: ~9 s of engine work, 155.9 s end to end for 49 MB / 48 tables. pg_upgrade still owed and may prove unnecessary
D — the downgrade refusal already done, v0.260.0
E — the automatic night COMPLETE. The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5
F — the remaining apps DONE 2026-09-22 (audits/DRILL-the-28-2026-09-22.md): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now 53 of 53 attempted, not 25. 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (outline 1.9.1 -> 1.10.1, HELD with the right sentence); 2 could not be deployed, one of them (plant-it) by design — it is lifecycle: abandoned and the product's lifecycle gate refused it, proven live for the first time. The cost is now known and it is not machine time: three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that POST /api/backup/run is BOX-WIDE, so concurrent walks serialise on it. This leg also added the night's biggest finding, which no count would have produced: R-630

The words "when it resolves to a container" are v0.262.0's, and they are the whole of R-630. Before it, the probe branch had no exit: findProbeContainer returning "" set a message and looped, while the settle path that judges an app declaring NO check sat in the outer else, unreachable. So a stack whose probe resolved to nothing could only ever time out — and failAndHold then stopped an app whose containers were all healthy. A stack with no probe is not healthy and not failing; it is settled on container state (§3), and never a reason to stop a running app. Proven live on paperless-ngx 2026-09-22: the identical Update that ended failed at +313.0 s with the app stopped now ends done at +53.4 s (audits/v0262-live-2026-09-22/live262.json). Which container is probed is now decidable too — exact stack name, then healthcheck.container, then a UNIQUE prefix, else nothing with the candidates logged; the old rule took the FIRST prefix match.

What the night ADDED to this table, which none of the legs anticipated: the verifying phase trusts the .felhom.yml probe absolutely, and three of 53 templates name a probe the app does not answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (R-618, P1). That is now the first thing Slice 6 has to be safe against, ahead of everything in this table.

Total for legs A–E: roughly 21–34 CC-hours, plus ~25–30 GB of images on a scratch host. Legs C and E are the ones that unblock a decision; leg A is the one that takes the time.

Venue: a scratch host, never a customer box — demo-hp's guest 9202 for the box-side legs, the harness on DooPlex for the image-side ones.

6.5 The drill catalog and the image store — the standing method for update drills

Why this section exists. On 2026-09-21 an afternoon session put a deliberately broken image into the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached a customer, but the method was wrong and the brief that asked for it said so. This is the method that replaces it, proven the same night.

The rule, and it has no exception: nothing broken, dummy, cross-repo or engine-major ever enters the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that, the leg is skipped and named.

The two mechanisms

what it is what it makes possible
the drill catalog admin/app-catalog-drill on Gitea — private, a copy of the live catalog's main a scratch box can be pointed at a catalog where a failing edge is committable, because it carries none of the live repo's gates
the image store a registry:2 container on the scratch guest at 127.0.0.1:5000 an edge that passes the within-a-major test and still fails — the one shape a real catalog move cannot produce

The image store is not a convenience. 09 §3b Q4 could not be measured for a year of drills because the only failing edges available were across-a-major, and the within-a-major rule — correctly — refuses those before the guarded update is ever reached. The rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. Measuring an unattended HOLD needs drill/<app>:X.Y.Z (the real image, retagged) against drill/<app>:X.Y.(Z+1) (a built image that starts, stays up and never serves) — same repository, same major, plain version tags. A third flavour, a tag simply absent from the store, gives the pull-failure leg.

stacks.CompareImageRefs orders a host:port/ reference correctly: splitImageRef takes the last colon and rejects it only when a / follows, so a registry port is never read as a tag. Proven by running it, four positive cases and a negative control, 2026-09-21.

Creating the drill repo — and the step that mails the operator if you skip it

POST /api/v1/repos/migrate is the route that works (the project's Gitea tokens carry write:repository but not write:user, so POST /user/repos answers 403). A migrated repo inherits has_actions: true, and the catalog's CI workflow comes with it. Every drill push then runs that workflow, it fails — the drill repo is deliberately gate-less — and each failure mails admin@felhom.eu. The 2026-09-21 night sent 47 such alarms overnight, into the same mailbox that was carrying real off-site alarms at the time (R-629).

So creating the drill repo is two acts, not one:

POST  /api/v1/repos/migrate      {"repo_name":"app-catalog-drill","private":true, …}
PATCH /api/v1/repos/admin/app-catalog-drill   {"has_actions": false}

then read the repo back and quote has_actions: False — an alarm channel trained to be ignored is worse than no alarm channel, and R-168 made CI mail the thing that notices a bypassed gate.

Pointing a box at the drill catalog — the step that is NOT obvious

git.repo_url alone is inert. Syncer.gitCloneOrPull clones only when the cache has no .git; otherwise it fetches from the remote the clone already stores. The cache directory must be removed as well, or the box goes on following the live catalog and reports success. Filed as R-615; until it is fixed, the drill procedure is:

  1. save controller.yaml as controller.yaml.pre-update-night;
  2. set git.repo_url (and username/token — the drill repo is private);
  3. remove <data>/catalog-cache;
  4. restart the controller, sync, rescan (R-607: a sync can answer „nincs változás" while the catalog has moved, and the badge answers from the stale value until the rescan);
  5. three controls, all quoted in the report — the drill bump appears on the scratch box; the other boxes' caches are unchanged; the live catalog's main hash is unchanged.

What the drill must leave behind

  • controller.yaml restored from the saved copy, the controller restarted, and git.repo_url read back and quoted as the live catalog.
  • The registry container and its volume removed; drill images removed by name. Never prune.
  • The drill repo kept, private, reset to the live catalog's main, so the next drill starts clean.
  • A diff of every image: line against the live catalog's main — expected: identical.

The fence

Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no customer box has credentials for it. The store listens on the guest's loopback only.

7. What slices 1 and 2 actually built

7.1 The record (slice 1)

Manager.recordInstalledImages (felhom-controller/controller/internal/stacks/installed.go) runs after a successful compose up from StartStack, RestartStack, UpdateStack and runComposeDeploy, and writes app.yaml:

installed_images:
  web:
    ref: lscr.io/linuxserver/bookstack:26.05.2
    digest: sha256:…              # "" if the image was never pulled from a registry
    at: "2026-09-02T18:41:03Z"    # when this ref+digest was FIRST seen for this service

Three rules, each with its reason:

  • It reads the CONTAINER, never docker-compose.yml. §1.2 is why: that file is the value that has already moved. A record built from it would answer "what will happen next time something runs up -d", which is a different question.
  • A failed write NEVER refuses the action — deliberately the opposite of SetDesiredState. Intent refused, observation logged. Refusing to start a customer's app because we could not write down which version it is trades a real outage for a bookkeeping gap.
  • It is NOT called from StartStackServices — the R-47 DB-only restore window would overwrite a complete record with a partial one.

Seeded at startup (v0.234.0). Manager.BackfillInstalledImages runs once at boot, beside the desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY on. It only READS containers. This was not a refinement — without it the feature did not reach a quiet box at all: see §8.3, which was written as a known limitation on 2026-09-02 and was a defect by the next morning.

Two admission rules, and the second is the design:

  • It never overwrites an existing record. The bring-up paths own updates; this fills gaps only.
  • It refuses to seed a PARTIAL observation. §7.2's comparison reads a service-count mismatch as BEHIND, so a degraded or crash-looping app seeded from its visible containers would render „Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial because they follow a SUCCESSFUL up -d, where a gap is real news and is logged; a backfill meets a box in whatever state it is in. Same field, two writers, two different admission rules — that is deliberate and must not be "made consistent".

Nothing reads it to take a decision. Slice 2 reads it to render a label.

7.2 The label (slice 2)

web.updateBadge (internal/web/updatebadge.go) compares the recorded reference per service against what the current template pins, and renders through the existing meta_badge partial — no new markup, no new CSS.

Absent means UNKNOWN and never means current. Every app.yaml written before v0.233.0 has no record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by a test with a companion red-proof.

No version number reaches the customer (operator ruling: a household cannot act on 26.05.2). Version strings stay in the logs, the API and the hub.


8. Known limitations, stated plainly

  1. „Naprakész" can be FALSE for the floating pins, and 2026-09-21 measured HOW false. The comparison is reference-to-reference and queries no registry — a customer's box must not depend on reaching eight upstream registries to render a page. For postgres:16-alpine, mariadb:11.6 and the others the reference can be identical while the image behind it has moved. NUMBERS, 2026-09-21 (audits/UPDATE-ARC-STATE-2026-09-21.md §3.3). The count this document carried — "23 of 66" — is STALE and matched no definition the catalog supports today. Recounted at catalog 18a6d2d8, with the definition stated so it can be rechecked: a pin FLOATS when its tag names a version LINE rather than an exact release. Of 66 unique pins, 48 are full X.Y.Z, 6 are two-part lines (mariadb:11.4/11.6/12.3, claper:2.5, opengist:1.13, wger/server:2.6) and 4 are major lines (postgres:15-alpine, postgres:16-alpine, redis:7-alpine, postgis:16-3.5-alpine) — 10 float. The remaining 8 are exact versions wearing a variant suffix (ghost:6.53.0-alpine, nextcloud:34.0.1-apache, …), which do not float by this definition. Of the 8 database and cache engine pins the sweep measured, 7 were measurable and 6 have been repushed upstream since the catalog set them — postgres:16-alpine (8 apps), postgres:15-alpine, redis:7-alpine (6 apps), mariadb:11.4, mariadb:12.3, postgis:16-3.5-alpine. Only mariadb:11.6 has not. The 8th, immich's own ghcr build, is UNMEASURED — ghcr exposes no anonymous last-modified. So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved. Digest-level comparison needs the catalog to record the digest at push time — R-446, put to the operator as §3b Q6, recommended YES.
  2. Nothing enforces catalog_since. Enforced by the pre-push hook since 2026-09-13 (R-452, app-catalog-felhom.eu/scripts/check-catalog-since.py); CI's shallow clone still skips it out loud. A commit that moves an image: line and forgets the date under-reports how far behind a box is. The gates runner fetches at --depth 1 and has no parent commit to diff against, so the gate needs a deeper fetch — R-452.
  3. The record only appears after the next lifecycle action. CLOSED in v0.234.0, and the way it closed is worth keeping. This was written on 2026-09-02 as an accepted limitation — "the fleet view fills in gradually". The operator looked at demo-felhom the next morning and found OpenGist, up 15 hours, running exactly the catalog pin, showing nothing at all. On a quiet box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped. BackfillInstalledImages now seeds the absences at startup by reading containers (§7.1). The residue that stays: the seed happens at controller START, so a box between upgrade and its next restart still shows nothing — bounded by one restart rather than unbounded.
  4. A frozen app is frozen WHOLE. While the catalog is ahead, no template correction reaches that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the wger 2.6 hazard, and it is the right trade: a new template around an old image is a third broken state. Recorded so it is a choice, not a surprise.
  5. .felhom.yml keeps flowing while the compose file is frozen — the deliberate asymmetry in §5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded. The failure direction is a false alarm, never data loss, and freezing .felhom.yml would break the update badge by withholding catalog_since. R-458.
  6. The Update button is still unguarded. CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1). It refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a safety dump, and holds an app that does not come up. What stays true: it still has no automatic rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) — that now ends HELD rather than crash-looping behind a green button.
  7. An engine major can be applied without its datadir upgrade, and nothing notices. CLOSED 2026-09-13 for MariaDB (R-459): every mariadb: sidecar carries MARIADB_AUTO_UPGRADE=1, and the harness shows the conversion running on the E3/E3b edges (§3 decision 5). What stays true: the PostgreSQL half (R-463) has no equivalent — the image performs no pg_upgrade — and the engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until Slice 4 gives the Update button a backup.
  8. Only three of 53 apps have ever had an upgrade measured. WIDENED 2026-09-21 to 21 EDGES ACROSS 19 APPS (audits/DRILL-update-night-2026-09-21.md), on scratch guest 9202 through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at any point: 14 proven, 3 failed, 4 inconclusive, each app seeded and read back through its OWN front door with a negative control on every readback. Ten of the fourteen printed a verbatim migration line. What stays true: bookstack is still only half-provable headlessly (R-460), and two apps cannot be seeded AT ALL while the catalog rightly closes their sign-up — vaultwarden (SIGNUPS_ALLOWED=false, R-512) and zipline — which is a permanent ceiling on R-462's scope rather than a fixture nobody has written (R-624). And one thing this widening FOUND that no count would have: the verifying phase trusts the .felhom.yml probe absolutely, and three of the 53 templates name a probe the app does not answer, so a SUCCESSFUL update of those apps ends by STOPPING a working app (R-618, P1 — tandoor measured serving HTTP 200 on the new version at four samples across five minutes, then stopped). CLOSED 2026-09-22, and the closing changed the numbers above: the three probes were corrected (app-catalog-felhom.eu@793c4fb), red-proofed live on 9202 in both directions, and a static gate now refuses a probe that does not match the same service's own compose healthcheck. tandoor's edge was re-walked with nothing else changed and ended done at +41.1 s where it had ended failed at +361.9 s — so the tally is 15 proven, 2 failed, 4 inconclusive, and all fifteen are on the live catalog since 2026-09-22. WHAT THE CLOSING FOUND, and it is the part worth carrying forward: a wrong probe was the LOUD failure. Two quiet ones sit beside it. paperless-ngx has no container whose name matches its stack name, so findProbeContainer returns nothing and its probe has never run on any box — an absence, with no badge to contradict it (R-630). And five more templates cannot be judged statically at all, one of which (home-assistant) is correct only because its check type cannot fail (R-631). So "the probe is right" is now enforced for 47 of 53 templates and still unknown for six. AND THE SWEEP'S REAL CEILING, counted rather than felt: 28 of the 53 templates have never been deployed by any drill (R-632) — the widening above went from 3 apps to 21, and 21 is not 53. CLOSED THE NEXT NIGHT, 2026-09-22: all 28 were walked (audits/DRILL-the-28-2026-09-22.md), so every template in the catalog has now been attempted at least once. And the walk that closed it found something the probe work had left open. paperless-ngx has no container whose name matches its stack name, so findProbeContainer returns nothing and its probe has never run on any box. Asked what verifying does with no probe to wait on, the answer is the worst of the three: it waits out the full update.health_timeout and HOLDS, stopping an app whose three containers all read healthy. The controller names it itself — not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app — at +313.0 s, front door 404 afterwards. So limitation 8 now has two shapes, not one: a probe that names the wrong target (R-618, fixed) and no probe at all (R-630, raised to P1), and the static gate can see the first but not the second, because there is nothing to compare. FIXED 2026-09-22 in controller v0.262.0, and the fix is in the phase table above: the no-probe case now settles on container state instead of looping, and the probe TARGET is decidable (exact name → healthcheck.container → a UNIQUE prefix → nothing, candidates logged). paperless-ngx and immich carry the explicit field, and the catalog gate REFUSES a probe that resolves to nothing rather than warning about it. So limitation 8's two shapes are both closed in the product; what remains is that six templates still cannot be judged STATICALLY (R-631, all five read live and correct) and that a probe can still be right about the port and wrong about what a 200 means — romm answered 200 from nginx while its workers were being OOM-killed for six hours (R-635), which is a third shape again and the reason "the update is guarded" must never be read as "the new version runs". Two more things the same night measured, both about state rather than health: a remove sent while a restore is still running reports success and leaves a container restarting with a live public route (R-633) — and the product already has exactly that guard for update and for restore, which name the blocking operation, but not for remove; and an app can be running, healthy and serving while recorded as deployed: false, in which state the product refuses to remove it at all (R-634, reproducible alone on sparkyfitness).
  9. The hub does not record image tags at all. Its report's container payload carries name, state, CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side change; it is not derivable from what is already reported.

9. Where the rest lives

what where
the measurements this document rests on audits/SPIKE-app-update-2026-09-01.md
the work backlog/OPEN-ITEMS.md — R-438..R-445, R-446..R-452
the syncer, described accurately but without the consequence architecture/02-controller-module-map.md
what the lifecycle actions are proven to do architecture/00-capability-map.md
the implementation felhom-controller/controller/README.md §"What is installed, and is it current?"