1749e73e50
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
1719 lines
138 KiB
Markdown
1719 lines
138 KiB
Markdown
# 09 — How an app update works, and what it is becoming
|
||
|
||
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
|
||
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
|
||
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
|
||
> did not know it was one.
|
||
|
||
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
|
||
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
|
||
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
|
||
|
||
---
|
||
|
||
## 1. How an update works today, as measured
|
||
|
||
### 1.1 The button
|
||
|
||
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
|
||
commands and nothing else:
|
||
|
||
```
|
||
compose pull → compose up -d --remove-orphans
|
||
```
|
||
|
||
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
|
||
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
|
||
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
|
||
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
|
||
|
||
### 1.2 The catalog syncer moves the file underneath a deployed app
|
||
|
||
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
|
||
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
|
||
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
|
||
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
|
||
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
|
||
sync does not restart anything.**
|
||
|
||
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
|
||
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
|
||
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
|
||
|
||
### 1.3 Thirteen other paths end in `compose up -d`
|
||
|
||
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
|
||
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
|
||
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
|
||
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
|
||
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
|
||
restore path.
|
||
|
||
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
|
||
a boot reconciliation started an app on a newer image at 17:55:44Z.
|
||
|
||
### 1.4 One fear is measured SMALLER than it was stated
|
||
|
||
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
|
||
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
|
||
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
|
||
1c, a positive observable and not an absent log line).
|
||
|
||
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
|
||
so is more useful than leaving the scarier version standing.
|
||
|
||
---
|
||
|
||
## 2. What was chosen, and by whom
|
||
|
||
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
|
||
whole arc:
|
||
|
||
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
|
||
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
|
||
> SIGTERM+start to existing containers without re-reading the compose file or env."*
|
||
|
||
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
|
||
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
|
||
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
|
||
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
|
||
it does not change it.
|
||
|
||
---
|
||
|
||
## 3. The operator decisions
|
||
|
||
These are rulings, not proposals. Anything specced against a different assumption is wrong.
|
||
|
||
### 2026-09-02
|
||
|
||
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
|
||
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
|
||
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
|
||
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
|
||
§6 — the file half was never priced, and demo-hp is too young a box to price it.
|
||
|
||
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
|
||
A customer on the newest version is supported however old that version is. This is why
|
||
`catalog_since` exists and why no version string is shown.
|
||
|
||
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
|
||
because §4 says it cannot be undone. **Second half REPLACED 2026-09-23 by decision 13** (the test
|
||
decides, not the tag); the first half is confirmed by decision 12.
|
||
|
||
### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0
|
||
|
||
4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it —
|
||
while corrections to its definition keep arriving on the 15-minute cycle exactly as they do
|
||
today.**
|
||
|
||
**This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes
|
||
that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in
|
||
its own comment. Reversing a chosen behaviour is a decision, not a bug fix.
|
||
|
||
**What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".**
|
||
The old behaviour had two halves and only one of them was unwanted:
|
||
|
||
| half | verdict |
|
||
|---|---|
|
||
| a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone |
|
||
| template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 |
|
||
|
||
So the ruling keeps the second and removes the first. In the operator's own words: *while the
|
||
catalog is offering the same version you are running, its fixes flow to you; the moment it moves to
|
||
a newer version, you are frozen at what you have until you choose to update.*
|
||
|
||
**It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup
|
||
precondition goes. Slice 3 only stops the other twelve paths from doing the update's job.
|
||
|
||
### 2026-09-13 — the database engine finishes its own conversion, and the upgrade test goes wide
|
||
|
||
5. **DBs should be updated when the app moves, with proper precautions, tests and backoff plans**
|
||
— the operator's own words, ruling on `SPIKE-r459-mariadb-upgrade-2026-09-06.md`. SHIPPED in the
|
||
catalog the same day: every `mariadb:` sidecar (`bookstack-db`, `kimai-db`, `nextcloud-db`,
|
||
`romm-db`) carries `MARIADB_AUTO_UPGRADE=1`; `MARIADB_DISABLE_UPGRADE_BACKUP` stays unset. **Not an
|
||
image change, so `catalog_since` does not move.** The three precautions, because they are the real
|
||
content of the ruling:
|
||
1. **Proven before it ships** — `upgrade-test.py` re-ran E3 and E3b on the changed template and the
|
||
engine-state field shows the conversion RAN (`mariadb_upgrade_info` reads the new version, the
|
||
engine's own check says nothing further is needed, the entrypoint no longer prints
|
||
`skipped due to $MARIADB_AUTO_UPGRADE`), with the seeded data reading back after. C3 still
|
||
returns `failed`. Evidence: `audits/r459-close-2026-09-13/`.
|
||
2. **Watched as it lands** — the change travelled the real 15-minute cycle to demo-hp: the live
|
||
compose gained the setting, the sync recreated nothing, and one deliberate restart logged
|
||
`MariaDB upgrade not required` with the app serving (same evidence directory).
|
||
3. **A rule until Slice 4 is built** — the Update button still takes no backup, so **no template
|
||
may move a database-engine image across a major version until R-448 ships.** Catalog
|
||
`CLAUDE.md` states it; `scripts/check-engine-major.py` enforces it in the pre-push hook (the
|
||
CI half cannot, R-452); its removal is tracked as **R-469** so it is a deliberate act.
|
||
The setting is inert until an engine major moves, and precaution 3 keeps it that way.
|
||
|
||
6. **The upgrade test goes as wide as possible, through the nightly unattended sessions** — the
|
||
ruling on `STATUS.md` item 11. Not "the ~25 database apps first": all of them, as the nightly
|
||
rotation reaches them, one fixture per app through the app's own interface. **Browser-only apps
|
||
become reachable when CC runs on the operator's Windows workstation with Chrome** — the
|
||
`claude-in-chrome` route that DooPlex does not have — so an app recorded `inconclusive` for want
|
||
of a headless seed route (bookstack's file half, R-460) is deferred to that venue, not faked.
|
||
|
||
---
|
||
|
||
### 2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts
|
||
|
||
7. **A floor carries a release past the vouched golden when the release's MinAgent is declared with
|
||
it** (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it,
|
||
the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which
|
||
`minagent_header_gate.py` now guarantees — goes into the same per-box agent comparison. An
|
||
undeclared floor above the golden is still HELD, and both floor forms refuse to save one. **Why:** a
|
||
controller image is pulled by tag and needs no golden to be delivered; what the floor was missing
|
||
was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release
|
||
delivery coexist. **Proven live:** both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s
|
||
from the save, the hub logging `SERVED … from declared`
|
||
(`audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt`).
|
||
8. **Any backup tier lets an app update** (R-475; controller v0.239.0). The precondition takes the first
|
||
FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site
|
||
(Tier 3, bounded; unreachable = absent). `update.backup_max_age` applies to whichever tier is chosen.
|
||
An app with nothing anywhere is backed up first; it is refused only when no backup can be taken
|
||
either. The hold names the tier and the date. **Tier 2 is required nowhere in the update path.**
|
||
This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a
|
||
verified recent backup as a precondition, not a new copy) stands. **Proven live:**
|
||
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
|
||
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
|
||
|
||
9. **A bind-data app's route back is off-site before its own unit, and the hold says what the copy
|
||
holds** (R-479, controller v0.241.0). An app whose data is bind-mounted files has a recovery unit
|
||
that holds the definition and the database dumps and NOT the files (measured: gokapi restored from
|
||
„helyi" came back with settings and no data). For such an app the update walks second drive →
|
||
off-site → own unit; for an app whose data is in named volumes the v0.239.0 order (second drive →
|
||
own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a
|
||
customer is never sent to a copy that cannot bring the data back without being told so.
|
||
|
||
### 2026-09-21 — decided by CC unattended, operator may reverse
|
||
|
||
10. **A box AHEAD of the catalog reads „Naprakész", and the guarded Update refuses to move a pin
|
||
backwards** (R-524, controller v0.260.0). *One sentence:* when the catalog is reverted under a box
|
||
that already updated, is that a "Frissítés elérhető"? **Options:** (a) leave it — the label
|
||
compares for difference, as §5.4 says; (b) show „Naprakész" and let the button still run;
|
||
(c) show „Naprakész" and refuse the button. **Costs:** (a) is free and offers a household a
|
||
downgrade onto a datadir the newer version may have migrated, which §4 says cannot be undone;
|
||
(b) removes the invitation but leaves the loaded gun; (c) costs one comparison and can, wrongly
|
||
applied, block a legitimate update. **Why (c):** the direction was already settled — §3 decision 3
|
||
says a version change that cannot be undone needs a human, and this is one. The risk in (c) is
|
||
bounded by making the Ahead verdict NARROW: every differing service must be orderable AND newer,
|
||
or the answer falls back to today's behaviour. Reversible, no customer-data risk, and it only ever
|
||
withholds an act. **Implementation:** `stacks.CatalogOrder`, one verdict read by both the badge
|
||
and `UpdatePreflight`.
|
||
|
||
### 2026-09-23 — the operator answers §3b's seven questions
|
||
|
||
Operator rulings, not CC decisions. Each answers one question in §3b, which is kept, marked ANSWERED,
|
||
because the costs written there are the reasoning behind these. **Nothing below is built yet** — the
|
||
build order is §6.4, and the two new mechanisms (15 and 14) were spiked the same day before anyone
|
||
builds them (`audits/update-rulings-2026-09-23/`).
|
||
|
||
11. **One maintenance window, and it is the backup window the household already sets** (Q1). App
|
||
updates are one more leg of the nightly chain, **after the off-site copy and before the
|
||
full-system backup**. An update not finished when the full-system backup is due waits for the
|
||
next night. **Why:** the update then leans on the freshest copy the box ever has, and there is no
|
||
second window for anyone to set or misread. **Replaces:** §3b Q1's recommendation of a fixed
|
||
02:30–05:00. *For the builder:* `stacks.update_window` (`config.go` L195, default `03:00-05:00`)
|
||
exists and **no Go code reads it** — it appears only in `config.go` (field + default),
|
||
`setup/handlers.go` L482 (written into a new box's `controller.yaml`) and
|
||
`configs/controller.yaml.example`. Remove it or fold it into the chain; do not build a second
|
||
window.
|
||
|
||
12. **Automatic updates stay, behind a per-box switch that is ON by default** (Q2). Confirms the first
|
||
half of decision 3. **Today every update is still manual; nothing automatic is built.**
|
||
|
||
13. **The test decides, not the tag** (Q3). **REPLACES decision 3's second half, "never across a
|
||
major".** The box applies by itself every step the catalog holds, because **the catalog holds only
|
||
tested steps**; the size of the version jump does not matter, the test does. The box does not
|
||
parse tags to decide. The tag rule (`stacks.CompareImageRefs`) moves to the **catalog gate** as a
|
||
push-time safety net: an image move with no test record is refused there. Two exceptions, both
|
||
**marks the test sets on the step**: *files may change* → automatic only when a fresh copy holds
|
||
the files, else a person (this is Q2's refinement); *needs a person* → the tester writes why.
|
||
**Why:** decision 3 was a proxy. "Within a major" was standing in for "known to work", and the
|
||
update night measured the proxy wrong in both directions — minor moves that held
|
||
(adventurelog, outline) and a tag shape it cannot read at all (§3b Q3). The test record is the
|
||
thing the proxy was guessing at.
|
||
|
||
14. **The update ladder** (Q3). A box more than one step behind **climbs one tested step at a time, in
|
||
order, and never jumps**. Each step is the full guarded update of §6.1. **Why:** a step was tested
|
||
from the version before it, not from three versions before it; R-40's multi-major jump is what
|
||
this prevents. This is the "stepping" §6.1 already assigned to slice 6.
|
||
|
||
15. **The box undoes a failed update itself** (Q4). **REPLACES §6.1's "the box never puts the old
|
||
version back by itself".** On a failed health check: the old definition and pin back + the
|
||
database copy taken seconds before the pin moved loaded back + the health check again. The app
|
||
stays stopped (HOLD) **only if the undo itself fails**. **Why the old ruling does not bind this:**
|
||
the old version refused to start on data the new one had migrated (Nextcloud, docmost — §4). The
|
||
undo puts the **pre-migration data** back too, so the old version meets the data it knows. The
|
||
word "rollback" stays struck; **the name is undo**. The household is told on the app page and by
|
||
one mail; the operator by event; **no retry** until the catalog moves or a person presses.
|
||
Spiked 2026-09-23 before any build — §6.1a.
|
||
|
||
16. **PostgreSQL majors are converted by the box** (Q5). A guarded-update step: save everything from
|
||
the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on
|
||
the test bench before the catalog may move it. The engine-major gate stays until then. **Why:** the
|
||
update night costed it at ~9 s of engine work for 49 MB (§3b Q5); moving the pin and letting it
|
||
hold would take eleven apps down on one night.
|
||
|
||
17. **The catalog records the image digest of every pin at push time** (Q6). The box compares against
|
||
it; where the catalog carries one, the box pulls **that exact image** — which also makes a floating
|
||
tag reproducible, not only the badge honest. *Whether compose can pull by a recorded digest while
|
||
the definition names a tag is a claim for the build plan to verify, not a ruling on mechanism.*
|
||
|
||
18. **Fleet view** (Q7). The report carries, per compose service (so the database too), the installed
|
||
reference, the catalog reference and the badge state. **Built later**, when the fleet grows.
|
||
|
||
### 2026-09-23 (afternoon) — two more operator rulings
|
||
|
||
19. **The undo's copy method is chosen by a bake-off, not by assumption.** Two methods are measured on
|
||
the same apps — *dump and load* (the database as a text file, loaded back) and *copy the folder*
|
||
(the app's named volumes copied with its containers stopped, put back on failure). The simpler one
|
||
that passes every case is built; if both pass, the folder copy wins on simplicity unless its
|
||
downtime or disk cost fails the bake-off's limits. **Why:** the morning spike found two traps in the
|
||
dump route (a cut-off file loads as success; a failed MariaDB load is half old, half new) and the
|
||
folder route had not been measured at all. Result: §6.1a.
|
||
|
||
20. **On update nights, the full-system backup waits for the update leg, inside its own window**
|
||
(answers R-643 / §6.4 part 7's open point). The leg stops starting new steps at **W+5h**, so the
|
||
full-system backup keeps at least one hour of its [W+2h, W+6h) window. **Built with §6.4 part 7,
|
||
not before.**
|
||
|
||
### 2026-09-23 (night) — one operator word, and two decisions taken by CC unattended
|
||
|
||
21. **Tonight the catalog may move every app whose within-a-major upstream edge is proven on both
|
||
venues, each with its test record; nothing else moves** (operator word, 2026-09-23 evening). The
|
||
bar: `proven` on the test bench (seed through the front door, update, read back, ten-minute memory
|
||
watch) AND on a scratch box through the real guarded Update; no across-a-major edge, no PostgreSQL
|
||
engine major, no `inconclusive`, no app whose seed has no front-door route. **This is decision 13
|
||
applied, not a new rule.** Result: `audits/DRILL-night-2026-09-23.md`.
|
||
|
||
22. **The memory watch's `memory_tight` mark reads the app's OWN memory (anon), not the cgroup peak**
|
||
— *decided by CC unattended 2026-09-23 night — operator may reverse.* *One sentence:* when the watch
|
||
says "tight", should the file cache count? **Options:** (a) the cgroup `memory.peak` as built
|
||
(R-635 follow-up); (b) the `anon` figure of `memory.stat`, sampled every 15 s, with the cgroup peak
|
||
kept beside it. **Costs:** (a) marks every app that reads files — measured tonight: nextcloud and
|
||
immich's PostgreSQL at **100 %** with **0** kernel kills, because the kernel fills the limit with
|
||
cache it drops before killing anything — and the gate then demands a raised limit, which inflates
|
||
`mem_limit` (a customer-box capacity figure) for no reason; (b) can miss a cache-heavy app whose own
|
||
memory is tight only if a kill never happens — and a kill or a restart still fails the edge under
|
||
both options. **Why (b):** decision 13's mark exists for "does not fit the memory" (RomM, R-635), and
|
||
only the app's own memory decides that. The cgroup peak stays in every record (`memory_cgroup_peak_pct`
|
||
in the ladder entry). R-652.
|
||
|
||
23. **The push-time test-record gate asks the registry — for moved refs only** — *decided by CC
|
||
unattended 2026-09-23 night — operator may reverse.* *One sentence:* may a `--fast` (hook) gate use
|
||
the network? **Options:** (a) no — digests compared only in the slow periodic run; (b) yes, but only
|
||
for the refs a push MOVES, and an unreachable registry is INCONCLUSIVE (the push is refused until it
|
||
can ask). **Costs:** (a) a move whose digest changed between the test and the push is published;
|
||
(b) a push that moves an image needs the registry — zero requests for every other push. **Why (b):**
|
||
decision 17 says the box pulls the recorded digest; recording one the registry no longer serves makes
|
||
every box's pull of that step fail (Scenario E). `app-catalog-felhom.eu/scripts/check-test-record-move.py`.
|
||
|
||
### 2026-09-24 — two operator rulings
|
||
|
||
24. **The fleet takes controller v0.267.0 although the chaos hour's stop rule fired** — both faults the
|
||
rule caught (R-658, R-659) predate that release; v0.267.0 causes neither and adds R-640's protection to
|
||
the same restore path. Floor saved with MinAgent 0.131.0 at 05:12:14Z; both demo boxes on v0.267.0,
|
||
healthy, 10 s later (`audits/ladder-2026-09-24/part0/`).
|
||
|
||
25. **R-659, option A: a held app's page names only a copy that can bring the app back WHOLE; when this box
|
||
has none, the page says so and that support is informed, and the operator gets an urgent event.** The
|
||
database-only restore under existing files (option B) is NOT built. *As built (v0.268.0 + hub
|
||
v0.122.0):* "whole" is read from the restores' OWN refusals, not from what a tier stores — measured from
|
||
source, an app with declared drive files (`DeclaredDriveFileLegs`) is refused by both the own-unit and the
|
||
second-drive unit restore (R-538's guard sits in the shared function), so for those apps only the
|
||
off-site copy counts (the second-drive gap is R-661). The hold names the newest whole copy with what it
|
||
holds; with none, `hold.update.no_whole_copy`, no Mentések button, `app_hold_no_whole_copy` (critical,
|
||
operator-only). The operator's English was used with one word dropped („please" — the house rule,
|
||
`i18n_missing_gate.py`); the Hungarian verbatim, its formal register recorded against R-516.
|
||
|
||
### 2026-09-24 (afternoon) — three more operator rulings
|
||
|
||
26. **R-661, option A: one action brings a file app back WHOLE from the second drive** — the unit (settings +
|
||
database) and the drive files together. The second drive then counts as a whole copy in decision 25's
|
||
truth table. R-538's guard stays: this is a new action beside it, not its removal.
|
||
27. **R-666, option B: while a held app's page says support is informed, Remove offers only „remove the app,
|
||
keep my data".** The household keeps control; the data stays. The no-whole-copy sentence moves to the
|
||
informal voice, like every other screen.
|
||
28. **An app in a crash loop, or in an out-of-memory storm, is STOPPED by the box, and the household and the
|
||
operator are told. A press on Start gives it one more try.** Operator's words: *"if it is in a crashloop,
|
||
or consuming resources, then yes, definitely stopped at least."*
|
||
|
||
**RomM follow-ups, operator-agreed the same day:** the test bench watches memory after an update
|
||
(`upgrade-test.py`, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4);
|
||
R-636's louder repeated alarm.
|
||
|
||
|
||
### 2026-09-24 (night) — two decisions taken by CC unattended
|
||
|
||
29. **Decision 28's crash loop is ≥ 6 restarts within 10 minutes, counted from Docker's `RestartCount`** —
|
||
*decided by CC unattended 2026-09-24 — operator may reverse.* *One sentence:* how many restarts in
|
||
what window make a crash loop the box stops? **Options:** (a) the brief's 10 in 10 min; (b) 6 in
|
||
10 min. **Costs:** (a) misses a STEADY loop — Docker's restart back-off caps it at about one restart a
|
||
minute (gokapi measured 539 → 546 in 7 min), so a loop that has run for an hour can sit at 9–10 per 10
|
||
min forever; (b) could stop a slow first start that restarts 6 times — none in the drill evidence
|
||
(1,831 harness samples, 40 live containers; the one ≥ 6 case, immich's first-start import, was broken,
|
||
R-676 watches it), and Start always gives one more try. **Why (b):** (a) fails the ruling's own case.
|
||
`08` §6.2; `audits/night-2026-09-24/A3/`.
|
||
|
||
30. **An INSTALLED app keeps the image digest it runs until a guarded Update moves it; the sync never moves
|
||
it** — *decided by CC unattended 2026-09-24 — operator may reverse.* *One sentence:* when the catalog
|
||
re-tests a floating tag at a new digest, does the sync write it into an installed app's compose?
|
||
**Options:** (a) yes, as v0.269.0 did; (b) no — the sync carries the running digest over, and only
|
||
the guarded Update (backup, undo copy, health check) renders the new one; a fresh install takes the
|
||
tested digest. **Costs:** (a) the next restart (a backup's stop/start, a power cut) pulls a new image
|
||
with no backup and no undo — measured live on 9202; (b) an installed app runs the older tested image
|
||
until someone (or the automatic leg) presses Update — the badge says so. **Why (b):** decision 17
|
||
wants the TESTED image pulled, and slice 3's ruling freezes the version until a deliberate Update;
|
||
(a) broke the second. Shipped as controller **v0.269.1** — a second release in the session, instead of
|
||
shipping (a) to the fleet with the floor.
|
||
|
||
### 2026-09-25 (night) — three decisions taken by CC unattended, building §6.4 part 7
|
||
|
||
31. **After W+5h the full-system backup waits only for an automatic step ALREADY in flight, and never past
|
||
W+5h30m** — *decided by CC unattended 2026-09-25 — operator may reverse.* *One sentence:* when the leg is
|
||
still running at W+5h, does the full-system backup start at once or wait for the step? **Options:** (a) start
|
||
at once — the backup's quiesce stops an app in the middle of its update's verify, which fails the step into
|
||
an undo during a backup; (b) wait for the step in flight, capped at W+5h30m. **Costs:** (a) a spurious undo
|
||
and a backup of an app mid-change; (b) the full-system backup keeps at least 30 minutes of its window instead
|
||
of 60 on the worst night; a stuck flag cannot hold it past the cap. **Why (b):** §6.4.2 point 3 already says
|
||
"a step already running finishes"; this makes the gate agree with it. `quiesce.updateLegDefers`,
|
||
`TestD20_UpdateLegDefers`.
|
||
32. **The controller's own self-update waits for the whole automatic leg, not only for a step in flight** —
|
||
*decided by CC unattended 2026-09-25 — operator may reverse.* *One sentence:* may the controller swap itself
|
||
between two automatic steps? **Options:** (a) yes, as for manual updates (the lock covers only a step in
|
||
flight); (b) no, the lock covers the whole leg. **Costs:** (a) the swap restarts the controller and the rest
|
||
of the night's leg is lost (it is not resumed); (b) a floor-served controller release waits up to the leg's
|
||
length (at most until W+5h) on an update night. **Why (b):** the leg is the night's only update window
|
||
(decision 11); a release can wait an hour, a household's app cannot wait a day for nothing.
|
||
33. **The automatic leg takes ONE step per app per night** — *follows the operator's own brief of 2026-09-24
|
||
("one app at a time, one tested step per app"), recorded here because §6.4.2 point 6 (a) said "rescan between
|
||
steps", which reads as climbing several.* **Cost:** an app three steps behind takes three nights. **Why:**
|
||
each step then gets a night of the household using it before the next one, and a failure is one step wide.
|
||
Operator may widen it. `TestLeg_OneStepPerAppPerNight`.
|
||
|
||
### 2026-09-25 — an operator ruling
|
||
|
||
34. **A controller restart during the night's update leg does NOT resume the leg** (operator ruling 2026-09-25,
|
||
option B of R-686). The apps the leg had not reached wait for the next night. A step already pressed is
|
||
finished or put back by the guarded update's own journal, exactly as before. **Nothing is built:** the
|
||
behaviour shipped in v0.271.0 is now the rule. **Why:** one night's delay for an app is cheap, and a resume
|
||
would be one more mechanism acting with nobody watching.
|
||
|
||
### 2026-09-25 (evening) — two operator rulings
|
||
|
||
35. **Part 10 is built now, one app first** (operator ruling 2026-09-25 evening, D1 option A). The box
|
||
converts a PostgreSQL major as a step of the guarded update. **docmost is the first app.** Each other app
|
||
needs its own proof on both venues (the test bench and a scratch box through the real guarded Update)
|
||
before the catalog may move it. The engine-major gate stays for every app without that proof. **Why:**
|
||
decision 16 said "each of the eleven apps is proven before the catalog may move it"; one app first makes
|
||
the first proof small enough to read. As built: §6.4 part 10.
|
||
36. **A reinstall over kept data offers a choice** (operator ruling 2026-09-25 evening, D2 option A): "use
|
||
my kept data" (the database from a backup, with the kept files) or "start fresh" (the kept files move to
|
||
a dated folder; nothing is deleted). **And, from the operator's follow-up question: kept data must never
|
||
be a dead end.** The household can see it (read only), load it into a later install, and delete it.
|
||
Support is for the rare case, not the normal one. **Not ruled (D3, in `STATUS.md`):** whether the box
|
||
ever deletes kept data by itself. Until it is ruled, **nothing deletes kept data automatically** — only a
|
||
household's explicit Delete. Where it lives and who deletes it: `07-backup-architecture.md` §"Kept data".
|
||
|
||
### 2026-09-25 (evening) — two decisions taken by CC unattended, building part 10
|
||
|
||
37. **docmost converts PostgreSQL 16 → 18, not 16 → 17** — *decided by CC unattended 2026-09-25 — operator may
|
||
reverse.* *One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount
|
||
change; (b) 18 — the data volume's mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) every
|
||
box converts again later, and docmost's own upstream compose already ships `postgres:18` at
|
||
`/var/lib/postgresql`; (b) the step changes the mount point — measured: `postgres:18` REFUSES (exit 1) even an
|
||
EMPTY volume at `/var/lib/postgresql/data` — so the step's definition carries the mount, and the undo must put
|
||
the old mount back with the old bytes, which it does (it restores the old definition and every volume).
|
||
**Why (b):** one conversion is better than two when the app's upstream runs the newer major. Reversible: the
|
||
catalog could pin 17 instead before any customer box moves. `audits/night-2026-09-26/A/README.md` A1.
|
||
38. **The box loads from a `pg_dumpall` of the OLD engine, with no error tolerated** — *decided by CC unattended
|
||
2026-09-25 — operator may reverse.* *One sentence:* what does the conversion load into the new engine?
|
||
**Options:** (a) the existing safety dump (`pg_dump` per database, `--no-owner --no-privileges`); (b) a
|
||
`pg_dumpall` taken from the old engine at conversion time, loaded with `ON_ERROR_STOP`. **Costs:** (a) is free,
|
||
and silently drops any role, grant or database setting beyond the bootstrap ones (for docmost both routes gave
|
||
identical databases — it has one role and one database; the other ten are not measured); (b) one more dump
|
||
(0.6 s for docmost) and a load that must step around exactly two objects the new engine's entrypoint makes (the
|
||
bootstrap role and database — measured: loaded raw, exactly two `already exists` errors). **Why (b):** decision
|
||
16 says "save everything"; the two collisions are removed precisely (the entrypoint's databases dropped only
|
||
when they hold no table; `CREATE ROLE <x>;` skipped only for a role that exists — its `ALTER ROLE … PASSWORD`
|
||
still runs), so ANY other error stops the load and the undo runs. Proven live: an extension 18 lacks
|
||
(`adminpack`) stopped the load and the box undid it. `audits/night-2026-09-26/A/README.md` A2.
|
||
39. **docmost's own memory limit rises 384M → 512M with its PostgreSQL 18 step** — *decided by CC unattended
|
||
2026-09-25 — operator may reverse.* *One sentence:* the bench marked the step `memory_tight` (the docmost app, not
|
||
the engine) — raise the limit, or leave docmost unmoved? **Options:** (a) keep 384M — the move gate refuses a tight
|
||
step whose limit does not move, so docmost stays on 16; (b) raise to 512M in the same commit (the gate's own
|
||
remedy, the RomM precedent). **Costs:** (a) the conversion that was built and proven tonight never reaches a box;
|
||
(b) +128 MB of reservation per docmost box (`mem_limit` 768M → 896M), and the mark stays (80.4 % at 512M): measured
|
||
twice, Node sizes its heap from the limit — 349 MB at 384M, 431 MB at 512M, 0 kills and 0 restarts in both 10-minute
|
||
watches (~12 000 requests each). **Why (b):** 91 % of the old limit is too close for a household box, the gate's
|
||
rule is followed rather than bypassed, and the mark's weakness for such apps is filed (R-693). Reversible.
|
||
|
||
### 2026-09-26 — two operator rulings
|
||
|
||
40. **The box never deletes kept data by itself** (operator ruling 2026-09-26, D3 option A). Only the household
|
||
deletes it, by Delete on the „Megőrzött adatok" page with the app's name typed. No age limit, no automatic
|
||
clean-up with warnings. Kept data can fill a drive; the drive-full warning names the kept folders and their
|
||
sizes as space the household can free (as built in v0.274.0). **Why:** the household decided to keep it; a box
|
||
that deletes it later reverses their decision without them. Closes D3 in `STATUS.md`.
|
||
41. **adventurelog keeps its world-data download** (operator ruling 2026-09-26, R-655 option B). The template does
|
||
NOT set `SKIP_WORLD_DATA=1`, so a fresh install still gets the countries/regions data. The move to v0.13.0 is
|
||
proven with the file already in the `adventurelog_media` volume (an update re-uses what v0.12.1 left), plus a
|
||
check that a truncated file is replaced (`download-countries --force` is the command's own repair). The
|
||
healthcheck override fix (`/usr/bin/node`) rides the image move in the same commit. **Why:** a fresh install
|
||
without world data is a smaller product; the internet dependence is at first start, which the update's own
|
||
health wait and undo already cover.
|
||
|
||
### 2026-09-26/27 — decisions taken by CC unattended (version-travel brief)
|
||
|
||
42. **The next PostgreSQL apps: paperless-ngx → 18, tandoor → 17** — *decided by CC unattended 2026-09-27 — operator
|
||
may reverse.* *One sentence:* which major does each app's conversion target? **Options:** (a) 17 for both — no
|
||
mount change; (b) 18 for both; (c) per app, from what the app's own upstream runs (decision 37's reasoning).
|
||
**Costs:** (a) paperless converts twice (its upstream compose ships `postgres:18` at `/var/lib/postgresql`);
|
||
(b) tandoor would run a major its own upstream compose (`postgres:16-alpine`) and its Django (5.2.16) have not
|
||
documented; (c) two different targets to remember. **Why (c):** decision 37 said one conversion is better than
|
||
two WHEN the upstream runs the newer major — paperless's does (18, Django ~5.2.5, psycopg 3); tandoor's does not,
|
||
and 17 is the newest major its Django documents, with no mount change. Read 2026-09-27 from each upstream's
|
||
compose and requirements (`audits/version-travel-2026-09-26/B/`). Reversible: the catalog pins the major.
|
||
43. **claper → PostgreSQL 17, calcom → 18** — *decided by CC unattended 2026-09-28 — operator may reverse.*
|
||
*One sentence:* which major does claper's conversion target? **Options:** (a) 17 — no mount change; (b) 18 —
|
||
the mount moves to `/var/lib/postgresql`. **Costs:** (a) a later second conversion if claper's upstream moves to
|
||
18; (b) claper would run a major its own upstream compose (`postgres:15`, v2.5.0) has not run. **Why (a):**
|
||
decision 42's rule — the newer major only when the app's own upstream runs it; claper's does not. Proven on both
|
||
venues (bench: 32 tables equal, peaks 75.8 %/78.3 %; box 9202: converted through the guarded Update in 7.2 s),
|
||
catalog `4a249b9`, `audits/pg-calcom-claper-2026-09-28/`. **calcom → 18** by the same rule: its upstream compose
|
||
runs untagged `postgres` at `/var/lib/postgresql` (Prisma 6.16.1 in v6.2.0). First its memory limit had to be fixed
|
||
(768M → 1536M, R-703); then both venues proved it (bench: 122 tables equal, `anon` 61.5 %; box: converted through the
|
||
guarded Update in 68 s), catalog `037f956`. Reversible: the catalog pins the major.
|
||
44. **demo-hp's scheduled restore test restores onto `nvme-scratch`** — *operator ruling 2026-09-28 (R-701 option (b)).*
|
||
`local-lvm` (53.9 GiB, the guest's own pool) cannot hold a 31 GiB restore; the NVMe can (~880 GiB free). Needed, and
|
||
added with the operator's word: the agent's storage role on `/storage/nvme-scratch` (user + token) — without it
|
||
Proxmox refused the restore with 403 `Datastore.AllocateSpace`. Proven: one restore test passed there in 8m46s,
|
||
`local-lvm` untouched. `03-host-agent.md`, `audits/logins-nvme-2026-09-28/C/`.
|
||
45. **No app is published with a login a stranger knows** — *operator ruling 2026-09-28 (R-702 widened).* Where the box
|
||
can set the first admin password, it generates one at install and shows it on the app page (`after_install:`,
|
||
controller v0.279.0). Where it cannot, the app stays in the catalog; the install dialog and the app page say what
|
||
the default login is and to change it at once. Operator: *"I don't think we should exclude apps if we can't change
|
||
the first PW."* The per-app audit and status: `app-catalog-felhom.eu/FIRST-ADMIN.md`.
|
||
46. **A setup gate for open-first-run apps is SPIKED before it is built** — *operator ruling 2026-09-29 ("I was leaning
|
||
towards B, but let's test A").* The 34 class-4 apps let the first visitor create the admin. Option A: while such an
|
||
app is not yet set up, the box lets only a person logged in to the household's dashboard reach it; the gate opens
|
||
when the app's own status says an admin exists, or when the household presses "Done, I set it up". Option B: one
|
||
fix per app (route (b), R-707). A is built only if the spike passes its exit test in writing; otherwise B continues
|
||
and the spike's result is the recorded reason. **Outcome (2026-09-29): the spike PASSED** (`audits/login-gate-2026-09-29/
|
||
B/B-VERDICT.md`, written before any build): on 9202 a stranger never reached a first-setup screen (~530 polls during
|
||
two installs, 0 app answers), the household passed with its dashboard session in 0.2 s, both probes flipped on the
|
||
setup, immich's phone-app API worked unchanged once the gate's router was removed, and the dashboard cookie was never
|
||
widened (a redirect handshake mints a 60-second, one-use token per app host instead). **Built in controller v0.280.0**
|
||
and proven live on immich, n8n, audiobookshelf (probe) and uptime-kuma (button). Costs, stated: ~2 ms per gated
|
||
request; a gated app answers 500 while the controller is down (closed, not open); a phone app cannot reach a gated
|
||
app; an app with no probe waits for the household's press, and the press trusts the household. Not covered by the
|
||
gate: open sign-up after the setup (R-711). Design record: `01-topology-and-trust.md` §5.
|
||
47. **Open sign-up is closed once an app's first admin exists** — *operator ruling 2026-09-29 (R-711, option A).* After
|
||
the household's first admin exists, a stranger can no longer make an account; only the admin adds people, from the
|
||
app's own user page, and the app page says how. Where an app cannot close sign-up, it stays in the catalog and its
|
||
page says plainly that anyone who finds the address can make an account; STATUS asks the operator about it.
|
||
**Mechanism — decided by CC unattended 2026-09-29, operator may reverse.** *How does the box close sign-up in apps
|
||
that keep the switch only in their own admin settings?* Options: (a) per-app switches — an env read at start plus a
|
||
restart, or the app's admin API; (b) a box-side block of the app's own sign-up address once the gate opens, with a
|
||
household window to let a family member in. Costs: (a) measured impossible for most — opengist and wishlist keep
|
||
it only in their admin settings (no env, no CLI), vikunja has an env but then no way to add a user except its CLI;
|
||
(b) one more traefik router per app, and a family member joins through a 15-minute window instead of an in-app
|
||
invite. **Chose (b)**: it works the same for every app, needs no credentials and no restart, and leaves installed
|
||
apps untouched (only an app whose gate this box opened gets a block). Built in controller v0.281.0
|
||
(`internal/stacks/signup_block.go`): the block goes up before the gate comes down; a failed write keeps the gate
|
||
closed. **Proven on 9202:** 11 apps let a stranger sign up after the setup and refused it with the block; 11 more
|
||
refuse by themselves; the window let a family member in and closed again. wanderer cannot be gated yet (R-714).
|
||
Evidence `audits/gate-rollout-2026-09-29/`.
|
||
Same day, operator: CC changes the admin passwords of demo-hp's installed bookstack and calibre-web and stores them
|
||
in the operator's credentials file (not in any repo).
|
||
|
||
---
|
||
|
||
## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed
|
||
|
||
> **ALL SEVEN ANSWERED by the operator on 2026-09-23 — §3 decisions 11–18.** Q1 → 11; Q2 → 12 and
|
||
> 13's *files may change* mark; Q3 → 13 and 14; Q4 → 15; Q5 → 16; Q6 → 17; Q7 → 18. Kept below
|
||
> unedited, because the costs and measurements here are the reasoning behind those rulings. Where a
|
||
> ruling differs from the recommendation below, the ruling wins: Q1 (the household's backup window,
|
||
> not a fixed 02:30–05:00), Q3 (the test decides, not `CompareImageRefs`), Q4 (the box UNDOES before
|
||
> it holds).
|
||
|
||
**These were questions, not rulings. CC did not decide them.** Each is one answerable sentence, the
|
||
options, what each costs, the recommendation, and what happens if nothing is decided. The measurement
|
||
behind them is `audits/UPDATE-ARC-STATE-2026-09-21.md`; the short version is that **46 of the
|
||
catalog's 58 exact pins are behind upstream today and 39 of those are within a major** — the
|
||
population §3 decision 3 already says may move without a human, and nobody presses 39 buttons.
|
||
|
||
### Q1 — When may a box update itself? — ANSWERED 2026-09-23: §3 decision 11
|
||
|
||
*May the box run the guarded Update by itself between 02:30 and 05:00, nightly?*
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **02:30–05:00 nightly, after the backup legs** | the update leans on a copy made hours earlier the same night, which is the freshest the box ever has. The app is down for the health wait in the middle of the night. |
|
||
| a weekly window | fewer interruptions; a box sits up to 7 days on a version the catalog already moved past, which widens the support window §3 decision 2 runs on |
|
||
| the household picks the window | one more setting on a page that already has several, for a choice almost nobody will change |
|
||
|
||
**Recommendation: 02:30–05:00 nightly.** The DB dump runs 02:30 and restic 03:00 on a demo box, so a
|
||
window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy of the night without a
|
||
new mechanism. **If nothing is decided:** Slice 6 cannot be built at all — every other question below
|
||
is downstream of this one.
|
||
|
||
**⚠ THE WINDOW CONTAINS 04:30, AND 04:30 IS WHEN THE BOX UPDATES ITSELF.** Found 2026-09-21 (R-608) by
|
||
reading the clock rather than by a failure: the controller self-updates daily at
|
||
`self_update.auto_update_time`, **default 04:30**, and again from `MaybeAutoUpdate` after ANY hub
|
||
report once a floor sits above the box — so **at any hour, not only at 04:30**. Either path restarts
|
||
the controller container, which is the supervisor of a running app update.
|
||
|
||
**What v0.261.0 now guarantees, so this question can be answered without also solving that one:** the
|
||
two cannot overlap in either direction. The controller defers its own swap while a guarded app update
|
||
is in flight (retrying on the next report, exactly as it already did for a running backup), and
|
||
`UpdatePreflight` refuses `self_updating` while a swap is in progress. **The lock does not latch** — a
|
||
held app does not block the controller's updates for ever. So the window may contain 04:30; the two
|
||
jobs will queue behind one another rather than meet. **What it does NOT do is reorder them**: if the
|
||
operator prefers the box to take its own update first, that is a scheduling choice still open here.
|
||
|
||
### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES? — ANSWERED 2026-09-23: §3 decisions 12 and 13 (the *files may change* mark)
|
||
|
||
*The button's rule and the automatic rule can differ. Should they?*
|
||
|
||
**The mechanism, verified at source this session, because an earlier draft had it backwards:** the
|
||
guard does **not** refuse these apps. Since v0.239.0/v0.241.0 (§3 decisions 8–9) the Update is refused
|
||
only when no copy exists on ANY tier and none can be taken. For an app whose data is bind-mounted
|
||
files, `Manager.UpdateTierOrderFor` (`controller/internal/backup/update_guard.go:136-141`) walks
|
||
second drive → off-site → **own unit last**, and when the own unit is the copy chosen,
|
||
`UpdateCopyHolds` (`:145-165`) ends the hold sentence with *„csak a beállításokat és az adatbázist
|
||
tartalmazza, a fájlokat nem"* — it holds the settings and the database and **not the files**. So the
|
||
update **proceeds**, and the household is told what the copy holds.
|
||
|
||
**With a human pressing, that is an informed choice. With nobody pressing, nobody was informed.**
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **automatic requires a fresh copy that HOLDS THE FILES; the button keeps today's rule** | the nine file-leg apps (and any other bind-data app) update automatically only on a box with a second drive or off-site; on a one-drive box they wait for a person. Two rules to hold in one's head. |
|
||
| one rule for both — automatic follows the button | simpler; a file-leg app can be updated unattended against a copy that cannot bring its files back, and the sentence saying so is read by nobody |
|
||
| automatic skips bind-data apps entirely | simplest; the seven file-leg apps that are behind never move by themselves even when a good copy exists |
|
||
|
||
**Recommendation: the first.** It is the smallest rule that keeps the promise the hold sentence makes.
|
||
**If nothing is decided:** Slice 6 must be built for the safe subset only, and the file-leg apps stay
|
||
manual — which is the third option by default, without anyone choosing it.
|
||
|
||
|
||
**MEASURED 2026-09-21 (update night).** The hold sentence this question turns on was read verbatim
|
||
off a REAL failure rather than from source. `adventurelog v0.12.1 -> v0.13.0` applied nine database
|
||
migrations successfully, never bound its port, and held:
|
||
|
||
> „A(z) adventurelog frissitese 2026-09-21 20:53-kor nem sikerult, es az alkalmazas nem indult el az
|
||
> uj verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek.
|
||
> Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: **sajat meghajto, 2026-09-21 20:47
|
||
> — ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza.**"
|
||
|
||
(ASCII fragments here; the live page carries its accents.) So the machinery this question's first
|
||
option would key on **exists and works**: the sentence names the tier, the date and **what the copy
|
||
holds**, unprompted, on a real edge. Whether the AUTOMATIC rule should differ from the button's is
|
||
untouched by that and remains the operator's.
|
||
|
||
**And one thing Q2 did not ask, which tonight makes urgent: after the hold, nobody can find out WHY.**
|
||
`failAndHold` removes the containers, so the failing version's own output is gone within seconds
|
||
(**R-621**). With a person pressing, they at least watched it happen.
|
||
|
||
### Q3 — What counts as "within a major" when the tag is not a version number? — ANSWERED 2026-09-23: §3 decisions 13 and 14
|
||
|
||
*§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`,
|
||
`kimai/kimai2:apache-2.57.0`, a date stamp, a digest?*
|
||
|
||
**And the test is per compose SERVICE, with ALL of them having to pass.** An app bump that is minor
|
||
while its `mariadb:` sidecar moves a major is **ACROSS** — that sidecar now converts the customer's
|
||
datadir by itself (R-459), so the edge carries a migration whatever the app's own number says.
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **an unorderable tag on ANY service makes the whole edge ACROSS → human** | the 8 floating pins and every suffix-versioned image stay manual. Conservative, and it is the same rule v0.260.0's `CompareImageRefs` already implements and tests. |
|
||
| teach the comparator each shape | every new shape is a new rule, and a wrong rule silently automates a major |
|
||
| compare digests instead | needs Q6 first, and a digest carries no order at all — it can say "different", never "newer" |
|
||
|
||
**Recommendation: the first**, reusing `stacks.CompareImageRefs` rather than writing a second rule.
|
||
*One small extension is needed and is named here so it is not discovered late:* v0.260.0's
|
||
`CompareImageRefs` answers *orderable?* and *newer?*, which is all R-524 needed. Slice 6 also needs
|
||
*same major?*, so the parsed major has to be exposed from the same normaliser — **an addition to the
|
||
one comparator, never a second one.**
|
||
**If nothing is decided:** Slice 6 would have to invent a rule under time pressure, which is how a
|
||
major gets automated by accident.
|
||
|
||
|
||
**MEASURED 2026-09-21 by RUNNING the comparator rather than reading it.** `CompareImageRefs` orders a
|
||
reference carrying a `host:port/` prefix correctly — `splitImageRef` takes the LAST colon and rejects
|
||
it only when a `/` follows, so a registry port is never mistaken for a tag. Four positive cases and
|
||
one negative control (different repositories are not orderable). **This is what made the unattended
|
||
hold measurable at all**: the drill edge `localhost:5000/drill/glance:1.0.0 -> :1.0.1` PASSES the
|
||
within-a-major test and still fails, which no real catalog move does. The recommendation is
|
||
unchanged; the *same major?* extension it already names is still owed.
|
||
|
||
### Q4 — A held app: who is told, when, and does the box try again? — ANSWERED 2026-09-23: §3 decision 15
|
||
|
||
*An automatic update that ends HELD happened while everyone was asleep.*
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **the household on the app page and by mail ONCE; the operator by event; NO retry until the catalog moves again or a person presses** | one mail per held app. The app stays down until someone acts — which is already true of a held update today. |
|
||
| retry the next night | a broken edge takes the app down every night and mails every morning; the hold exists precisely because the box cannot fix it |
|
||
| tell only the operator | the household finds their app down and has no sentence explaining it |
|
||
|
||
**Recommendation: the first.** It is what the manual hold already does (`settings.RestoreHold` with
|
||
`reason: update_failed`), plus one mail. **If nothing is decided:** the safe default is no automatic
|
||
update at all, because a hold nobody is told about is worse than a version nobody moved.
|
||
|
||
**MEASURED 2026-09-21, and the honest answer is that HALF of this is still unmeasured.** The
|
||
unattended night ran (`audits/update-arc-gaps-2026-09-21/09-unattended-night.md`). What it proved:
|
||
an app updates itself end to end with nobody pressing anything; one app takes **51 s – 1 m 26 s**
|
||
including the health wait; the caller needs **no new controller code**, only the existing guarded
|
||
Update plus `UpdateRefusal.Reason` on the wire (v0.261.0, R-609); and **a terminally-refused app is
|
||
pressed exactly ONCE and never again** — four apps, three passes, proven.
|
||
|
||
**What it did NOT produce is a HOLD, and the reason is instructive rather than a failure of the
|
||
run.** The only failing edge available was `vikunja → alpine:3.20`, and the caller **correctly
|
||
refused to attempt it**: different repositories cannot be ordered, so the edge is "across" and
|
||
belongs to a human by decision 3. **The rule that makes automatic updates safe is the same rule that
|
||
refuses the obvious way to break one.** Measuring the unattended hold needs an edge that PASSES the
|
||
within-a-major test and still fails its health check — same repository, same major, a tag that
|
||
starts and does not serve — which probably means a purpose-built image rather than a catalog move.
|
||
**So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).**
|
||
|
||
|
||
**MEASURED 2026-09-21 (update night) — and this is the half that was missing.** The caller pressed
|
||
ONCE with nobody watching; the app held after **312.9 s**; passes 2 and 3 pressed nothing at all
|
||
(`outcomes={'glance': ('held', 312.9)} never_again=['glance']`).
|
||
|
||
| the question | the answer, measured |
|
||
|---|---|
|
||
| does an unattended update ever produce a HOLD? | **yes** — 312.9 s, the full health wait plus the phases |
|
||
| does the box try again? | **no** — two further passes pressed nothing |
|
||
| is the household told? | **on the screen, yes** — the app page, and a banner on EVERY authenticated page carrying every held app at once |
|
||
| told what? | what happened, when, **which copy** and **what that copy holds** — all four scored True |
|
||
| by MAIL? | **still unmeasured** — the scratch guest runs `hub.enabled: false` and the notifier returns before it logs (**R-620**) |
|
||
| in ENGLISH? | **no** — the sentence is Hungarian on the English page (**R-606**, confirmed on the hold sentence itself) |
|
||
|
||
**So the mechanism this question's recommended option rests on is already there and already behaves
|
||
that way.** What remains in Q4 is the MAIL and the ENGLISH, not the hold.
|
||
|
||
**Two further facts this measurement produced, neither of which the question anticipated.**
|
||
**(1) There is NO single-flight** — five Updates pressed within 0.45 s all ran at once and all ended
|
||
honest, so a caller pressing N apps runs N updates simultaneously. **(2) A held app keeps inviting
|
||
the household to update it and the button then refuses** (`409 reason='held'`), even after the
|
||
catalog publishes a FIXED newer version — the household's only route out is the restore. Correct per
|
||
§6.1, and the page says otherwise (**R-625**).
|
||
|
||
**Also proven across a genuine power cut:** the boot sweep met a held app after an unclean shutdown
|
||
and deliberately left it alone — *„whatever is holding it owns its recovery"*.
|
||
|
||
### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`? — ANSWERED 2026-09-23: §3 decision 16
|
||
|
||
*Eleven templates, and the image performs no conversion — it refuses to start on an older major's
|
||
datadir (R-463).*
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **a scripted `pg_upgrade` edge in the harness, proven on all eleven, before the catalog may move** | real work: eleven fixtures, and `pg_upgrade` needs both major's binaries present. The engine-major gate keeps the rule until it exists. |
|
||
| move the pin and let the update HOLD honestly | every one of the eleven apps goes down on the same night and comes back only by a restore |
|
||
| never move PostgreSQL majors | the fleet sits on an engine that eventually loses upstream support |
|
||
|
||
**Recommendation: the first, and the gate stays until it lands.** As of 2026-09-21 the engine-major
|
||
rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of it and
|
||
`MARIADB_AUTO_UPGRADE=1`); this half is exactly what stays. **If nothing is decided:** nothing breaks
|
||
— the gate refuses the move — but the eleven apps drift further from upstream every month.
|
||
|
||
|
||
**MEASURED 2026-09-21, both halves, on a real seeded datadir.**
|
||
|
||
**(a) What a household would see today — as predicted, and now observed.** The guarded Update of
|
||
`postgres:16-alpine -> 17-alpine` ended **`failed` in 5.1 s**; the app was stopped and held; **the pin
|
||
named 17 while `installed_images` still said 16 and nothing was running**; the data was intact; and
|
||
the restore the hold sentence names brought it back in **29.1 s**. The engine's refusal had to be
|
||
REPRODUCED independently, because `failAndHold` destroyed it before any probe could read it
|
||
(**R-621**) — *FATAL: database files are incompatible with server / DETAIL: The data directory was
|
||
initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir
|
||
was still `16` afterwards; the positive control (the same copy under 16) started and held 48 tables.
|
||
|
||
**(b) The conversion rehearsal, COSTED.** Logical dump and restore, 49 MB / 48 tables:
|
||
`pg_dumpall` **2.6 s / 132 201 B**; fresh 17 datadir plus replay **6.5 s / 48 tables restored**; the
|
||
app up on 17 saying *Database connection successful*; **the seeded account read back**; **total
|
||
155.9 s, of which ~9 s is engine work.** For eleven apps that is a maintenance window, not a project.
|
||
`pg_upgrade` was NOT run — it needs both majors' binaries in one image and no such image exists in
|
||
this project; the logical route may make it unnecessary at this size. Full paragraph:
|
||
`audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`.
|
||
|
||
**(c) A fact about the INSTRUMENT, not the engine.** `upgrade-test.py`'s PostgreSQL probe is
|
||
`cat /var/lib/postgresql/data/PG_VERSION` **inside the container**. Against the converted datadir it
|
||
answered `17`, exit 0 — it works. **But it is blind in exactly the case that matters**: when
|
||
PostgreSQL refuses, the container is not running, so `docker exec` cannot ask it anything. Tonight it
|
||
recorded `No such container`, which its own honesty rule covers — but it must never be read as *the
|
||
engine is content*.
|
||
|
||
**The recommendation is unchanged.** Tonight gives it a price rather than a new opinion.
|
||
|
||
### Q6 — Should the catalog record each pin's DIGEST at push time? — ANSWERED 2026-09-23: §3 decision 17
|
||
|
||
*So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.*
|
||
|
||
**This is no longer theoretical. Measured 2026-09-21: six of the seven measurable floating pins have
|
||
been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps),
|
||
`postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`,
|
||
`postgis:16-3.5-alpine`. On demo-hp today, four apps read „Naprakész" over a database engine image
|
||
that has demonstrably moved.
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **the catalog records the digest at push time; the box compares digests** | one field per pin. `check-image-resolvable.py` already resolves the digest, so the producer exists. §8.1's rule — the box never queries a registry — is untouched. |
|
||
| the box queries registries | breaks §8.1 outright: a page that cannot render without eight upstream registries |
|
||
| leave it | the badge stays right about the question it asks and wrong about the one a household hears |
|
||
|
||
**Recommendation: yes.** It is the cheapest real improvement on this list and it closes R-446.
|
||
**If nothing is decided:** „Naprakész" keeps meaning "the reference matches", which is measurably not
|
||
what it sounds like.
|
||
|
||
|
||
**MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it refines the picture in two ways.**
|
||
§8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a
|
||
customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins
|
||
were read as `installed_images` records them and compared with the upstream digests measured the same
|
||
night: `postgres:16-alpine` -> `sha256:721873c34ceb9…` **on both sides**; `redis:7-alpine` ->
|
||
`sha256:858f009f9709c…` **on both sides**. **Identical — so „Naprakesz" is TRUE for this box.**
|
||
|
||
**(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for
|
||
a box that pulled BEFORE the tag moved. R-446's six repushed pins measure the tag against the date
|
||
the CATALOG set it, which is the right measure for the catalog and not for a box.
|
||
|
||
**(2) The producer this question needs ALREADY EXISTS on the box.** `installed_images` records a real
|
||
`digest` per service — the box knows exactly what it is running. What it cannot do is COMPARE,
|
||
because the catalog carries no digest. That is precisely this question's proposal, and only the
|
||
catalog half is missing. **The recommendation is unchanged.**
|
||
|
||
### Q7 — What does the hub's report need to carry for a fleet view? — ANSWERED 2026-09-23: §3 decision 18
|
||
|
||
*Slice 7 lets the operator SEE and MOVE how far behind every box is.*
|
||
|
||
**Verified both sides this session:** the controller's report payload carries name, state, CPU and
|
||
memory and no image (`controller/internal/report/types.go` L98–103), and the hub's
|
||
`Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts**. **But
|
||
the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent —
|
||
what is missing is the denormalisation and the page, not the transport.
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **per app: installed reference + catalog reference + badge state; the hub lists boxes behind, with a "move" that is the same guarded Update, operator-triggered** | additive on both sides; the report grows by a few fields per app |
|
||
| badge state only | smaller payload; the operator cannot see WHAT is behind, only that something is |
|
||
| leave it to per-box pages | free today at two boxes; unusable at twenty |
|
||
|
||
**Recommendation: the first, and it stays P3-LOW until the fleet grows.** **If nothing is decided:**
|
||
the only way to answer "is the fleet current?" is what this session did — read both boxes' files by
|
||
hand.
|
||
|
||
---
|
||
|
||
### Not a question — already ruled
|
||
|
||
**R-462's scope was decided on 2026-09-13.** §3 decision 6: the upgrade test goes to **all** apps
|
||
through the nightly rotation, explicitly *not* "database apps first". The register row R-462 still
|
||
says *"VIKTOR rules on scope"* — **that row is stale and is corrected to cite decision 6.** The update
|
||
night below proposes an ORDER *inside* that ruling; it does not reopen it.
|
||
|
||
## 4. The vocabulary ruling — "rollback" is struck
|
||
|
||
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
|
||
run, putting the old image tag back produces a container that refuses to start —
|
||
|
||
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
|
||
> downgrading is not supported"*
|
||
|
||
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
|
||
|
||
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
|
||
|
||
| shape | when it applies | what it does |
|
||
|---|---|---|
|
||
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
|
||
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
|
||
|
||
There is no third. **2026-09-23 (§3 decision 15): the UNDO is the two combined** — the abort's old
|
||
image plus a restore from the one copy that is seconds old, the pre-pin safety dump. It is not a
|
||
third shape and it is not a rollback: nothing is migrated backwards, the pre-migration state is put
|
||
back.
|
||
|
||
### 4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades
|
||
|
||
`SPIKE-upgrade-test-2026-09-06.md` upgraded three real apps with real data in them and then attempted
|
||
the abort on each. **All five real catalog upgrades kept the customer's data.** The abort did not
|
||
behave the same way twice:
|
||
|
||
| app | abort | why |
|
||
|---|---|---|
|
||
| **docmost** `0.25.3`→`0.95.0` | **REFUSES** | the old code finds migration ledger entries it does not know: *"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*, then *"Failed to run database migration. Exiting program."* |
|
||
| **privatebin** `1.7.5`→`2.0.5` | **works** | file-backed, no database, no schema — a major version moves no data |
|
||
| **bookstack** app+engine | **works, misleadingly** | only because the MariaDB datadir upgrade was skipped and never happened — see R-459 |
|
||
|
||
**This puts TWO independent measurements behind the ruling above, by two unrelated mechanisms:**
|
||
Nextcloud refused on an explicit version comparison; docmost refuses on its migration ledger. The word
|
||
"rollback" was already struck; it is now struck on evidence rather than on one case.
|
||
|
||
**And it adds a distinction this document did not have: there is no single answer to "can this update
|
||
be undone". There are apps where it can and apps where it cannot, and the only way to know which is to
|
||
MEASURE THAT APP.** Any design that assumes one answer for all 53 is designing against a fact that was
|
||
checked and is false.
|
||
|
||
Per §5 below, a restore's image-level undo used to have a ≤15-minute half-life because the syncer
|
||
overwrote it (**R-441**) — closed in v0.235.0.
|
||
|
||
---
|
||
|
||
## 5. The shape, as SHIPPED in v0.235.0
|
||
|
||
The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the
|
||
syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather
|
||
than an input, and the thirteen unattended `up -d` paths stop being able to change a version by
|
||
accident.
|
||
|
||
### 5.1 Nothing was added to the thirteen call sites, and that is deliberate
|
||
|
||
**They are made safe by removing the reason, not by gating them.** The most important of them are
|
||
REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`),
|
||
the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a
|
||
customer's app down, which is worse than the problem this slice solves.** Since the file they act on
|
||
no longer changes version, every one of them became safe without being touched.
|
||
|
||
### 5.2 The pin, and what it is not
|
||
|
||
`AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref.
|
||
|
||
**It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a
|
||
DECISION — what should run. Letting an observation feed a decision would make a bad reading become a
|
||
bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over.
|
||
They will normally agree; when they disagree that is a signal, not a bug to paper over.
|
||
|
||
**Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.**
|
||
|
||
Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came
|
||
from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is
|
||
safe from the catalog, and keeping it beside the app means it travels with every path that already
|
||
moves a stack dir.
|
||
|
||
### 5.3 The four writers — the only acts entitled to move a version
|
||
|
||
| writer | pin source |
|
||
|---|---|
|
||
| the deploy path (`runComposeDeploy`) | the template just deployed from |
|
||
| **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull |
|
||
| the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** |
|
||
| `Manager.AdoptPins` | the observation, once, and only when complete AND matching |
|
||
|
||
**`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file
|
||
on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards
|
||
would pull the frozen version and change nothing — while reporting success, and a button that lies is
|
||
worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of
|
||
`recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent.
|
||
|
||
### 5.4 The render table, complete
|
||
|
||
| app state | result |
|
||
|---|---|
|
||
| not deployed / protected / seam not wired | the catalog template — today's behaviour |
|
||
| deployed, **unpinned** | the catalog template + one DEBUG |
|
||
| deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** |
|
||
| deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE |
|
||
| pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it |
|
||
| mid-deploy | the compose file is left alone this cycle |
|
||
|
||
**`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries
|
||
`catalog_since`, which the badge needs. See §8.5.
|
||
|
||
**The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and
|
||
putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration
|
||
the older template cannot supply, so an old image under a new template is broken in a way neither
|
||
version is.
|
||
|
||
**And this is not "skip deployed apps".** That option was considered and rejected: it also stops
|
||
health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in
|
||
the spike §3 — both halves the ruling explicitly kept.
|
||
|
||
### 5.5 Adoption, and why the startup order matters
|
||
|
||
`AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app
|
||
to what it is already running. **It reads and writes files only** — no container is started, stopped or
|
||
touched. It skips, loudly, when the observation is incomplete or when the app runs something the
|
||
current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a
|
||
guessed pin.
|
||
|
||
**`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position
|
||
that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and
|
||
handed the next restart a version change — the exact behaviour this slice removes, once per boot.
|
||
|
||
### 5.6 The trap this slice set for the previous one
|
||
|
||
`Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one.
|
||
On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find
|
||
installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test
|
||
still green, because the new field has the same type and shape. The badge now reads
|
||
`Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an
|
||
earlier feature is the failure mode to look for whenever a file changes meaning.**
|
||
|
||
---
|
||
|
||
## 6. The seven slices
|
||
|
||
| # | slice | status |
|
||
|---|---|---|
|
||
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
|
||
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
|
||
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02); English since v0.258.0 (R-589); a FOURTH verdict — AHEAD — and the downgrade refusal in v0.260.0 (R-524, §3 decision 10)** |
|
||
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
|
||
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 |
|
||
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
|
||
| **6** | **Automatic updates** — ~~within a major, a human across one~~ **the catalog's tested steps, climbed one at a time, as a leg of the backup chain, undone by the box on failure** (§3 decisions 11–17); **an engine change gets its own edge.** | OPEN — R-450; rulings 2026-09-23; build order §6.4 |
|
||
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451; ruled 2026-09-23 (decision 18), built later |
|
||
|
||
### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)
|
||
|
||
`POST /api/stacks/{name}/update` is a guarded job. It answers **202** at once; the outcome exists only
|
||
on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`, `update_error`,
|
||
`hold_reason`), and `update_phase=done` is written only after the app's health is known. **R-443 is
|
||
closed by construction: nothing reports an update complete on the compose exit code.**
|
||
|
||
**The sequence, and the order is the design:**
|
||
|
||
| # | phase | what happens | on failure |
|
||
|---|---|---|---|
|
||
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no copy on ANY tier and no backup can be taken now** (since v0.239.0; before it, no restorable Tier-2 copy) | nothing moves, nothing is recorded |
|
||
| 1 | `checking` | walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than `update.backup_max_age` (v0.239.0) | nothing moves |
|
||
| 2 | `backing-up` — only when no tier holds a fresh copy | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 | refused with the backup's own error; nothing moves |
|
||
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
|
||
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
|
||
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
|
||
| 5a | `copying` **(v0.263.0)** | `compose stop`, then each NAMED volume `cp -a` into `<volume>.pre-update-<stamp>` by a helper that writes a finished-marker last (bind folders never) | copies removed, **pin put back, the old version started again** — nothing new ran |
|
||
| 6 | `starting` | `compose up -d --remove-orphans` | **UNDO** (below), HOLD only if the undo fails |
|
||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **UNDO** (below), HOLD only if the undo fails. Before v0.263.0: stop + HOLD, the pin stays (Scenario F) |
|
||
| 7a | `undoing` **(v0.263.0)** | every copy validated (marker) BEFORE anything is poured back; volumes emptied and refilled; definition, pin and the pinned version's `.felhom.yml` record put back from the job's own copies; `up`; health with the OLD version's probe | HOLD, the sentence prefixed *„A frissítés nem sikerült, és az automatikus visszaállítás sem."* + the data state (`untouched` / `half` / `not_started`); copies kept |
|
||
| 7b | `undone` **(v0.263.0)** | installed images recorded, copies removed, `app.yaml` `last_update_undone`, journal cleared | — |
|
||
| 8 | `done` | installed images recorded, copies removed, `last_update_undone` cleared, journal cleared | — |
|
||
|
||
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
|
||
`update.health_timeout` (default `5m`).
|
||
|
||
**The precondition is the existing verified backup, not a new copy** (§3 decision 1). It is
|
||
`backup.Tier2UnitRestorePoint` — the SAME predicate that permits the destructive „Teljes
|
||
visszaállítás", extracted from the backups page rather than copied. **The copy is aged by the last
|
||
SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was measured before it was
|
||
designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp
|
||
bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the
|
||
manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is
|
||
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).** **SUPERSEDED in
|
||
v0.239.0 by §3 decision 8:** `backup.Manager.UpdateRestorePoints` walks all three tiers and the update
|
||
leans on the first fresh copy; `Tier2UnitRestorePoint` is still the Tier-2 half and still the page's
|
||
predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum
|
||
skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit
|
||
proven current. The hold stores `copy_tier` and names „második meghajtó" / „saját meghajtó" /
|
||
„távoli mentés"; a successful off-site restore now lifts an update hold too.
|
||
|
||
**The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and
|
||
gate as R-379, so every start path that already honoured a restore hold honours this one. A successful
|
||
unit restore lifts an update hold (only that kind). **Three unattended paths honoured no hold before
|
||
v0.237.0 and now do:** the drive-return gate's restart and boot recreate, and the nightly volume dump
|
||
(which ends in `StartStack`). The nightly capture and Tier-2 run skip a held app, so the restore point
|
||
the hold text names is never overwritten.
|
||
|
||
**Crash safety is a journal**, `<data>/update-journal.json`, written before every phase. `RecoverUpdates`
|
||
runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back;
|
||
after `up` → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by
|
||
`ResumeInterruptedUpdates` once the backup side is wired, ending healthy or held.
|
||
|
||
**THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT.** Whether an old image
|
||
starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot
|
||
be predicted. So the box never puts the old version back by itself. **The route back is the restore**,
|
||
and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where
|
||
the harness has proven it, is slice 6's.
|
||
|
||
> **REPLACED 2026-09-23 by §3 decision 15 — the box UNDOES a failed update itself.** The measurement
|
||
> above still stands and is the reason the replacement has a different shape: putting the old IMAGE
|
||
> back alone is what refuses (§4.1). The undo puts the old image back **together with the database
|
||
> copy taken before the pin moved**, so the old version meets pre-migration data. Spiked by hand on
|
||
> 9202 the same day — §6.1a.
|
||
|
||
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
|
||
|
||
### 6.1a The undo (decision 15) — SPIKED BY HAND, then BUILT: controller v0.263.2 (2026-09-23)
|
||
|
||
**SHIPPED AND PROVEN LIVE on 9202** (`audits/undo-live-2026-09-23/README.md`): docmost, romm and
|
||
vikunja each made a real migrating update fail its (deliberately wrong) probe; the product undid all
|
||
three in 30–52 s, with data written before the backup, after it, and seconds before the press all read
|
||
back through each app's front door and the ledgers equal; the page carries one line in the request's
|
||
language. A cut-off copy → HOLD saying *the data is as the new version left it*; a power cut during
|
||
the undo → resumed after boot and completed; a person's press after an undo → `done`. **Two defects
|
||
only the live box could show, fixed the same day:** the undo's probe was never asked while the
|
||
current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` taken at update time
|
||
was already the new one, because `.felhom.yml` flows in on every catalog sync (v0.263.2: the pinned
|
||
version's file is now recorded in `applied-meta/` whenever a version is pinned). **Residual (R-646):**
|
||
an app pinned before v0.263.2 has no such record until its next pin. **Two more, found by the chaos hour of
|
||
2026-09-23 night (`audits/DRILL-night-2026-09-23.md` Part D):** the undo finds an app's volumes by their compose
|
||
label, and a restore recreates them WITHOUT it — so after any restore the undo copies nothing (R-658, P1); and a
|
||
held file-leg app is pointed at a restore that refuses a database-only copy, with no route left on a box without
|
||
an off-site tier (R-659, P1). A controller kill during `verifying` resumed and undid correctly (round 9). **Both fixed in v0.268.0 (2026-09-24):** the undo selects volumes from the app's compose definition and the restore
|
||
labels what it creates (R-658); the hold names only a copy that brings the app back whole, else says support is
|
||
informed (R-659, decision 25). Proven live on 9202, `audits/ladder-2026-09-24/`.
|
||
|
||
The spike, as it was run by hand before any build:
|
||
|
||
Evidence: `audits/update-rulings-2026-09-23/README.md`. Three real migrating edges on 9202, each made
|
||
to fail a deliberately wrong probe, each held by today's product, each then undone by hand.
|
||
|
||
| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (SQLite in a volume) |
|
||
|---|---|---|---|
|
||
| old version on the migrated data, nothing loaded | **refuses** (migration ledger) | **refuses** (alembic revision) | starts and serves |
|
||
| the safety dump | DB only, holds the post-backup write | DB only, holds it | **none — no-op** |
|
||
| undo, load + start → healthy | **≈ 16 s** | **≈ 38 s** | ≈ 1 s |
|
||
| data written before AND after the backup read back | yes / yes | yes / yes | yes / yes |
|
||
|
||
**The undo works — and not with the loader the product has.** `ImportDump` over a migrated
|
||
PostgreSQL database FAILS (the new version's foreign keys block the dump's own drops); over MariaDB it
|
||
succeeds and leaves the new version's tables behind. The load that worked empties the schema and loads
|
||
the copy in one transaction. **And a truncated PostgreSQL copy loads with exit 0 into an EMPTY
|
||
database** — the copy's completion marker must be checked first, which `ValidateDump` does not do. A
|
||
product path that loads a safety dump back exists (`rollbackSafetyDump`), but only the off-site
|
||
restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices
|
||
them.
|
||
|
||
**THE COPY METHOD, chosen by the bake-off of 2026-09-23 afternoon (decision 19):
|
||
COPY THE FOLDER.** `audits/undo-bakeoff-2026-09-23/README.md`. Both methods passed every case on all
|
||
three apps (seeds before and after the backup, ledger equal, a cut-off copy caught before anything is
|
||
swapped or loaded). The folder copy wins because an app with no database server gets no dump at all,
|
||
so the dump route would have needed the folder copy anyway. **Shape:** after the pull and just before
|
||
`up` — where the app is stopped anyway to be recreated — each NAMED volume the app owns is copied with
|
||
`cp -a` into a sibling volume `<volume>.pre-update-<stamp>` by a helper container, which writes a
|
||
finished-marker last; a bind-mounted user folder is never copied and never touched. On a failed health
|
||
check the copy is validated (marker present) and put back, the old definition and pin are restored
|
||
from the journal's own copies, and the old version is checked with the OLD `.felhom.yml` probe.
|
||
Measured cost: ≈ 1–5 s extra downtime for the three apps, ≈ 420 MB/s, disk = the volumes' size (the
|
||
update refuses before moving anything when the copy would breach the 2 GB floor).
|
||
|
||
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
|
||
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
|
||
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
|
||
both demo boxes by the floor alone, with its MinAgent declared.
|
||
|
||
**Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not
|
||
only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute
|
||
health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit.
|
||
The Tier-2 mirror the hold names survived only because Tier 2 is daily. `backup.Manager.isHeld` now also
|
||
answers true for an app a guarded update is moving (`SetUpdatingCheck`).
|
||
|
||
**v0.240.0 (2026-09-13, evening) — what the afternoon's proof and the first nightly rotation found, fixed.**
|
||
Seven rows: removal with backups kept now keeps the Tier-2 RECORD, so the second-drive restore is not
|
||
refused over an intact mirror (R-486, P1 — the disaster the second copy exists for); PostGIS/pgvector/
|
||
TimescaleDB images are Postgres, so such apps get their logical dump (R-484); "delete backups" deletes
|
||
the unit, the mirror(s) and the prefs (R-474/R-466); the backup card sizes them (R-485); a held
|
||
update's sentence leaves the card with the hold (R-480); the Tier-3 lookup is one `snapshots` call
|
||
(R-477); a unit older than the app's `deployed_at` does not count (R-478). Delivered by the floor in
|
||
16 s / 18 s; every row proven live with a throwaway adventurelog. `audits/v0240-2026-09-13/`.
|
||
|
||
**Proven live on demo-hp, 2026-09-13**, with a throwaway uptime-kuma and real catalog tag changes (each
|
||
reverted in the same phase): A (2.3.2→2.4.0, done after health), B (`backup_max_age: 2m` → backup first),
|
||
E (non-existent tag → pin back, container untouched), F (`alpine:3.20` → held), H (three buttons and the
|
||
boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold
|
||
cleared). Live evidence: `audits/slice4-2026-09-13/`.
|
||
|
||
### The verdict record — the contract Slice 6 carries
|
||
|
||
Decided here rather than invented twice. The harness writes one of these per edge, beside its
|
||
evidence; Slice 6 puts the same shape in the catalog.
|
||
|
||
```json
|
||
{"harness_version": 1, "app": "bookstack",
|
||
"from": {"bookstack": "…:25.02.2", "bookstack-db": "mariadb:11.6"},
|
||
"to": {"bookstack": "…:26.05.2", "bookstack-db": "mariadb:12.3"},
|
||
"verdict": "proven | failed | inconclusive",
|
||
"seed_read_before": true, "seed_read_after": true, "healthy_after": true,
|
||
"migration_observed": "verbatim log line, or null",
|
||
"abort": "starts-and-serves | refuses | starts-data-gone | not-attempted",
|
||
"memory": {"soak_s": 600, "containers": {"<name>": {"limit": 0, "peak": 0, "peak_pct": 0.0,
|
||
"oom_kills": 0, "restarts": 0}}, "first_kill": null},
|
||
"marks": ["memory_tight"],
|
||
"abort_detail": "the refusal quoted verbatim, or null",
|
||
"duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"}
|
||
```
|
||
|
||
**Harness version 2 (2026-09-23, R-635) adds `memory` and `marks`.** After a successful readback the
|
||
harness runs the new version for `--soak` seconds (default 600) under light load and reads the
|
||
kernel's own `oom_kill` counter host-side. A kill or a restart turns `proven` into `failed`; a peak
|
||
above 80 % of the compose limit adds `memory_tight`. The two fields are the test record's memory half
|
||
(decision 13).
|
||
|
||
**`inconclusive` is a first-class verdict and must never be collapsed into `failed`.** "We could not
|
||
measure it" and "it does not work" are different facts, and only one of them is about the app.
|
||
**`migration_observed` is a quoted line, never an inference from timing** — the value of both the
|
||
Nextcloud and the docmost findings was the exact sentence the app printed.
|
||
|
||
### Database engines under an upgrade — MEASURED 2026-09-06
|
||
|
||
**The arc's standing rule that an engine change gets its OWN edge now has measured evidence behind
|
||
it**, and the evidence is stronger than the rule's original argument. The rule was justified by
|
||
*"two migrations behind one edge is an unreadable failure when it breaks"* — a readability argument.
|
||
What was measured is that **an engine change can be applied and silently NOT happen**, which the
|
||
app-half edge cannot produce and which no amount of readability would have surfaced:
|
||
|
||
- `SPIKE-upgrade-test-2026-09-06.md` §4 — MariaDB 12.3 starts on an 11.6 datadir, logs that the
|
||
conversion it requires was **skipped**, and serves. **Assigned to the engine half by decomposition:**
|
||
the app half alone produces no such line.
|
||
- `SPIKE-r459-mariadb-upgrade-2026-09-06.md` — it is **stable but never self-resolving** (5 of 5
|
||
restarts, no degradation, and the engine says `Check required!` every time, forever). Converting
|
||
properly **succeeds**, costs **7 s**, takes its own system-database backup, and **does not** cost the
|
||
ability to abort. **The trade that was expected here does not exist.**
|
||
- **2026-09-13 — the setting is in the catalog.** All four `mariadb:` sidecars carry
|
||
`MARIADB_AUTO_UPGRADE=1` (operator ruling, §3 decision 5), and `upgrade-test.py`'s engine-state field
|
||
now shows the conversion RUNNING on the bookstack edges. **And a gate holds the engines inside their
|
||
major until Slice 4:** `app-catalog-felhom.eu/scripts/check-engine-major.py` (R-469).
|
||
|
||
**Two rules for anything this arc builds around a database engine:**
|
||
|
||
1. **Ask the engine, not the log.** MariaDB's entrypoint prints `MariaDB upgrade not required` on an
|
||
unsupported **downgrade**; `mariadb-upgrade --check-if-upgrade-is-needed` names it exactly
|
||
(**R-464**). A cheap instrument built on the log line would report "fine" for the broken case.
|
||
2. **The two engines fail in opposite directions, so one check will not do.** MariaDB starts anyway
|
||
and skips quietly; **PostgreSQL refuses to start** on a datadir from an older major, and the image
|
||
performs no `pg_upgrade`. Eleven templates carry PostgreSQL and **eight sit on `postgres:16-alpine`**
|
||
(**R-463**).
|
||
|
||
**And an engine-state field belongs BESIDE a verdict, never inside it.** `upgrade-test.py` reports
|
||
`engine_state_after` next to `verdict`, because an unconverted datadir is not *known* to be a failure
|
||
and a verdict that said so would encode an unproven judgement.
|
||
|
||
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
|
||
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
|
||
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
|
||
breaks.
|
||
|
||
---
|
||
|
||
### 6.2 Slice 6, as it will be built (OPEN — R-450; ruled 2026-09-23, decisions 11–17)
|
||
|
||
**The shape the rulings fix. The build order and costs are §6.4.** Rewritten 2026-09-23; the
|
||
earlier "as it would be built" draft keyed on the window of Q1's recommendation and on
|
||
`CompareImageRefs`, and both were ruled differently.
|
||
|
||
**Nothing new happens to a single step.** Each step is **exactly the guarded Update of §6.1** — same
|
||
precondition, same safety dump, same pin journal, same health wait — with **one change to its end**:
|
||
a failed health check runs the **undo** (decision 15, §6.1a) before it holds. Slice 6 adds a
|
||
*caller*, a *ladder* and the *undo*; it adds no second update path.
|
||
|
||
**When — a leg of the chain, not a window of its own (decision 11).** The household's backup window
|
||
start W already drives every nightly leg at fixed offsets (`07` §6.1): DB dump at W, Tier 2 at
|
||
W+60m, off-site at W+105m, and the full-system backup's gate opens at **W+2h** (`quiesce.go`
|
||
`gateOpenOffsetMin = 120`, span to W+6h). The update leg starts **when the off-site leg has
|
||
finished** and stops starting new steps **when the full-system backup starts**; whatever is left
|
||
waits for the next night. **Measured consequence the builder must face, not discover:** the gap
|
||
between "off-site finished" and W+2h is **at most 15 minutes** and is zero on a night the off-site
|
||
leg runs long, while one step takes 51 s – 1 m 26 s when it succeeds and ~5 m when it fails (§3b
|
||
Q4). So either the full-system gate learns to wait for the update leg (it has a four-hour span to
|
||
spend) or the leg gets almost no time. That is a build choice inside decision 11, named in §6.4.
|
||
|
||
**Which apps — the catalog decides, not the tag (decision 13).** An app qualifies when ALL hold:
|
||
|
||
1. the per-box switch is on (decision 12; **on by default**);
|
||
2. `stacks.CatalogOrder` says **Behind** — never Unknown, never Ahead;
|
||
3. the **next step** from the app's installed state carries a **test record** in the catalog —
|
||
a step with none is never applied by a box (the catalog gate refuses to publish it, decision 13);
|
||
4. the step's **marks** allow it: *needs a person* → never automatic; *files may change* → automatic
|
||
only when a fresh copy on some tier holds the app's FILES (`UpdateCopyHolds`), else a person;
|
||
5. the app is not held (`held` is terminal until a person acts or the catalog moves, decision 15).
|
||
|
||
**How far — one step at a time (decision 14).** A box two steps behind applies step A→B, then B→C,
|
||
each the full guarded update, each with its own health check and undo. A failed step stops the
|
||
ladder for that app. **Measured 2026-09-23:** today one press jumps A → C and B never runs, and the
|
||
box cannot see B at all — its catalog clone is `--depth 1` (`sync.go:283`, `:300`; one commit
|
||
visible on both demo guests).
|
||
|
||
**The ladder's format — recommended, not ruled** (`audits/update-rulings-2026-09-23/README.md` Part 2):
|
||
an `update_ladder:` list in `.felhom.yml`, one entry per step — `from`/`to` refs per service, the
|
||
digest per ref, the test record, the marks — and, for every step but the last, the step's OWN
|
||
complete definition in `templates/<app>/steps/<to>.yml`. **Not the git history:** romm's image-moving
|
||
commit `15f9ebf` is the definition that OOM-looped on demo-hp; the step that works is its images with
|
||
the later `f4eb94f` template, and no commit holds that pair.
|
||
|
||
**When it fails — undo, then hold only if the undo fails (decision 15).** The household is told on
|
||
the app page and by **one** mail, in the box's language (R-606 is a precondition — an automatic
|
||
update's mail must be in the household's language); the operator by event; **no retry** until the
|
||
catalog moves or a person presses.
|
||
|
||
**What the household sees.** An event and a line on the app page's timeline, before and after, in both
|
||
languages: *„Automatikus frissítés 03:12-kor — sikeres"* / *„— visszaállítva az előző verzióra"* /
|
||
*„— megállítva, a másolat 2026-09-20-i"*.
|
||
|
||
**Where it lives.** A leg in the nightly chain beside the existing ones, calling
|
||
`Manager.StartGuardedUpdate`. **It must respect the `isHeld`/`SetUpdatingCheck` interlocks v0.238.1
|
||
added** — the nightly capture running *inside* an update's health wait is the defect that release
|
||
fixed, and a second unattended caller is exactly the shape that finds it again. And it must run
|
||
**one app at a time**: there is no single-flight today (§3b Q4: five Updates pressed within 0.45 s
|
||
all ran at once).
|
||
|
||
**The in-process caller reads `UpdateRefusal.Reason`, and the split is measured, not assumed**
|
||
(v0.261.0, R-609): `busy`, `updating`, `deploying`, `migrating` and `self_updating` are **transient**
|
||
— try again on the next pass; `held` and `downgrade` are **terminal** — never press that app again
|
||
until a person acts; `memory`, `disk` and `no_backup` need a person and should be surfaced, not
|
||
retried. A working caller in this exact shape exists as evidence, not product:
|
||
`audits/update-arc-gaps-2026-09-21/unattended-caller.py`.
|
||
|
||
**The per-box switch key is `app_update.unattended`**, default **true** (decision 12).
|
||
|
||
⚠ **NOT `auto_update` — that name is TAKEN, and by the very thing this must not collide with.**
|
||
`self_update.auto_update` / `self_update.auto_update_time` (`config/config.go` L280-281, default
|
||
**04:30** at L422) are the CONTROLLER's own update. Two settings with that name, one meaning the
|
||
controller and one meaning apps, is the kind of collision that is only discovered by an operator who
|
||
turned off the wrong one. The app-scoped key is `app_update.*`. **And `stacks.update_window` is
|
||
removed or folded in, never a second window** (decision 11).
|
||
|
||
### 6.3 Slice 7, as it would be built (OPEN — R-451; ruled 2026-09-23, decision 18 — built later)
|
||
|
||
Three additive pieces, and the transport already exists (§3b Q7):
|
||
|
||
1. **controller** — the report's per-app object gains installed reference, catalog reference and badge
|
||
state. Additive; an older hub ignores it.
|
||
2. **hub** — denormalise those out of the raw report it already stores whole, and list boxes by how
|
||
far behind they are.
|
||
3. **hub → box** — a "move" button that is the same guarded Update, operator-triggered, through the
|
||
existing command path. **Not a second update mechanism**, and not automatic.
|
||
|
||
Rank stays P3-LOW at two enrolled boxes. It rises with the fleet, and §2 of the state audit is what
|
||
that looks like today: the only way to answer *"is the fleet current?"* was to read both boxes' files
|
||
by hand.
|
||
|
||
### 6.4 The build order for the 2026-09-23 rulings (PLAN — each part returns to the operator for go/no-go)
|
||
|
||
Costed in **CC-evenings** (one evening ≈ one unattended session: build, red-proofs, live proof on 9202,
|
||
release). Written from the two spikes and the memory watch of 2026-09-23
|
||
(`audits/update-rulings-2026-09-23/`), not from source reading alone. **Risk to customer data** is
|
||
what the part can do to a household's data if it is wrong, not how likely that is.
|
||
|
||
| # | part | rulings / rows | cost | depends on | risk to customer data |
|
||
|---|---|---|---|---|---|
|
||
| **1** | **SHIPPED — controller v0.263.2, proven live on 9202 2026-09-23** (`audits/undo-live-2026-09-23/`). **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
|
||
| **2** | **SHIPPED — controller v0.264.0 + hub v0.120.0, proven live on 9202 and 9201 2026-09-23** (`audits/undo-fleet-2026-09-23/`). **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held: events `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, on by default, mailed in the household's language with the app named in the subject; per-app cooldown on both legs. Leftovers: R-647. | R-606, 15 | **1** | — | none |
|
||
| **3** | **SHIPPED — controller v0.264.0, proven live on 9202 2026-09-23.** **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none |
|
||
| **4** | **SHIPPED — catalog `6db08a5`, night 2026-09-23** (`audits/DRILL-night-2026-09-23.md`): `update_ladder:` in `.felhom.yml`, one JSON entry per line (spiked on controller v0.266.0 and v0.267.0 first — the controller ignores the key); `check-test-record.py` (static, CI) + `check-test-record-move.py` (history + the registry for moved refs only, decision 23); the only writer `upgrade-test.py --write-ladder`; the 21 moves of 2026-09-22 backfilled from their records (21 proven). Not built: `steps/<to>.yml` (part 5 needs it once an app has two steps); `CompareImageRefs` did NOT move to the gate — the gate asks for a proven test instead, which is decision 13's own test. **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only |
|
||
| **5** | **SHIPPED — controller v0.268.0 (`206b035`) + catalog `5ed599c`, proven live on 9202 2026-09-24** (`audits/ladder-2026-09-24/partD/`): romm 5.3.0/11.4 → 5.3.1/11.4 (its app step, from `steps/90dd9d68258286ef.yml`) → 5.3.1/11.8 (the engine step; `mariadb-upgrade` ran) in two presses, the seeded account read back after each, the page's „Hátralévő frissítési lépések" 2 → 1 → none. Step files: `templates/<app>/steps/<StepKey(to)>.yml` (sha256 of `to` as canonical JSON, 16 hex; `stacks.StepKey` = `ladder.step_key`), refused absent or wrong by `check-test-record.py` rule 4, written by `--write-ladder` when a step is superseded, 8 backfilled from the NEWEST commit naming each step's images (romm's from `f4eb94f`, not the OOM-looping `15f9ebf`). The box reads them from its `--depth 1` clone (the whole tree is there — measured). One press = one step; a missing step file refuses before anything moves; an installed version matching no entry jumps, logged by name (measured live: vikunja 2.5.0). Limitation: a step has no `.felhom.yml` of its own (R-664). **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/<to>.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape |
|
||
| **6** | **BOTH HALVES SHIPPED. BOX HALF — controller v0.269.0 + v0.269.1, proven live on 9202 2026-09-24** (`audits/night-2026-09-24/B/`): the compose that runs pins `tag@sha256` from the ladder entry that tested those refs; pins and records stay digest-free; the badge reads behind for a newer TESTED digest; **an installed app keeps its digest until a guarded Update moves it** (v0.269.0 let the sync move it — found live, fixed in v0.269.1, `CarryDigests`); Compose accepts the form for all 25 ladder apps. **CATALOG HALF SHIPPED — night 2026-09-23:** every ladder entry carries the digest per `to` ref (`scripts/image_digest.py`, stdlib; equals Docker's `RepoDigests` on a box), and the move gate refuses a digest the registry no longer serves. **The box half (compare, render `name:tag@sha256`) is not built.** **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low |
|
||
| **7** | **SHIPPED — controller v0.271.0 (`9cf13a3`), proven on 9202 over four simulated nights 2026-09-24/25 and watched through the demo boxes' first real automatic night (`audits/DRILL-night-2026-09-25.md`).** Decisions 31–33 taken unattended. Gaps a scratch box cannot close: R-687; no resume after a restart: R-686. *(Before:)* **SPIKED 2026-09-24 (night, Part C) — the measured build brief is §6.4.2 below.** **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has ENDED (on every exit path, skips included), one app at a time, one step per press, `app_update.unattended` default ON (decision 12), `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it, and the full-system backup's gate waits for it until W+5h (decision 20). | 11, 12, 14, 15, 20 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
|
||
| **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
|
||
| **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
|
||
| **10** | **SHIPPED FOR DOCMOST — controller v0.273.0 + catalog `6a4a5f0` (harness v4, the gate's proof clause), proven live on 9202 2026-09-25** (`audits/night-2026-09-26/D/`): one press converted docmost 16 → 18 in **9.9 s** of engine work (**42 s** end to end), the check equal (2 databases, 48 tables, 71 rows), `PG_VERSION` 18, the seed read back; the load failing (an extension 18 lacks) → **undone in 43.6 s**; the app unhealthy on 18 → **undone in 148 s** (90 s drill health timeout); the controller SIGKILLed one second after the volume was emptied → the restart **undid it in 40 s**; each ending on 16 with the seed read back. The old datadir's copy is kept until a backup is proven after the conversion. Decisions 35, 37, 38. Every other PostgreSQL app stays refused by the gate until its own two-venue proof. *(Before:)* **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
|
||
| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows |
|
||
|
||
**Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10**, part 11 when the fleet grows. **Total
|
||
for 1–10: ≈ 22 evenings.** The automatic caller (7) is deliberately late: it is the only part that
|
||
acts with nobody watching, and every part before it is what makes that safe. Parts 8 and 9 are small
|
||
and independent and can fill any short evening.
|
||
|
||
**The one open point the build cannot settle alone — part 7, inside decision 11.** The ruled chain is
|
||
*off-site copy → updates → full-system backup*. Today the off-site leg starts at **W+105m** and the
|
||
full-system backup's gate opens at **W+2h** (`quiesce.go` `gateOpenOffsetMin = 120`, span to W+6h). So
|
||
the update leg has **at most 15 minutes**, and none on a night the off-site copy runs long — while one
|
||
step takes ~1 min when it works and ~2–6 min when it fails and is undone.
|
||
|
||
| option | cost |
|
||
|---|---|
|
||
| **the full-system backup waits for the update leg, inside its own window; the leg stops starting new steps at W+5h** | the full-system backup starts later on update nights, still inside its four-hour window, with an hour kept; one more interlock between two nightly jobs |
|
||
| the leg stops at W+2h as the chain stands | ≤ 15 min a night — about ten steps on a good night, none on a slow one; a box far behind takes weeks to climb |
|
||
|
||
**Recommendation: the first.** It keeps the ruling's order and its promise that the full-system backup
|
||
is never skipped for an update; only the start time inside its existing window moves.
|
||
|
||
### 6.4.2 Part 7 — the build brief, from measurement (night 2026-09-24, Part C)
|
||
|
||
Evidence: `audits/night-2026-09-24/C/` and `C-00…C-07`. Spike only — no product code for the caller was
|
||
written. Read with decisions 11, 12, 14, 15 and 20.
|
||
|
||
**1. The chain, as it runs today (source + a real night).** The legs are clock-scheduled from one window
|
||
start W and NOTHING waits on the leg before it: db-dump at W, Tier 2 at W+60m, off-site at W+105m
|
||
(`backupwindow.go:16-21`, `cmd/controller/main.go:1048/1134/1267`); the full-system backup's gate is a
|
||
poll loop that opens at W+2h for 4 h (`quiesce.go:656-657`, window check `quiesce.go:238-245`).
|
||
**Measured on demo-hp 9201's real night (W = 02:30):** the off-site leg ran 04:15:05 → 04:18:33 CEST
|
||
(3 min 28 s, `ok`); the gate opened at 04:30. **So a leg placed after the off-site leg has ~11 minutes
|
||
before the gate** — decision 20's wait is required, not optional.
|
||
|
||
**2. The signal that the off-site leg has ended.** There is no event. The nearest thing is the
|
||
persisted `offbox` status in `settings.json` — `LastStatus` (`running` at start, `ok`/`incomplete`/
|
||
`error` at the end) with `LastRun` beside it, written by `UpdateOffboxStatus` (`offbox.go:1003-1125`).
|
||
The status travels with the timestamp (presence-is-not-success holds), **but it is not a reliable
|
||
"finished" signal:** five early returns (`offbox.go:863-897` — not configured, escrow pending, orphaned,
|
||
migration, another backup running) and the scheduler's own skip (`main.go:1268-1271`) write nothing, and
|
||
a box with no off-site target (9202 tonight) never runs the leg at all. **Build: CHAIN, do not poll.**
|
||
The `offbox-backup` job function calls the update leg after `RunOffboxBackup` returns, on EVERY path
|
||
including every skip — one call site, no new signal to go stale. A box with no off-site target runs
|
||
the update leg at W+105m.
|
||
|
||
**3. The gate's interlock (decision 20).** Insert after the window check in `quiesce.runOnce`, as an
|
||
`Options` func beside `WindowStartFn` (`quiesce.go:75-95`): *update leg active and now < W+5h → defer*
|
||
(log and return, exactly like the window deferral; the loop re-polls every 5 min). The leg itself stops
|
||
STARTING steps at W+5h; a step already running finishes (≤ health timeout + undo). Manual whole-box runs
|
||
(`TriggerNow`) bypass it, as today.
|
||
|
||
**4. One simulated night (9202, v0.269.1; the caller = `tools/unattended_caller_ladder.py`, which presses
|
||
only the public Update).** Four apps: wishlist 1 step behind, navidrome 2, romm 3, vikunja 1 with a
|
||
failing step. Measured per step: navidrome 10.6 s and 8.5 s (no database); wishlist 55.8 s (its first
|
||
update: the backup ran); romm 77.4 s, 71.2 s, 93.9 s (database + volume dump each step); vikunja's
|
||
failing step **undone in 104 s** with the drill's 90 s health timeout — **with the default 5 min it is
|
||
≈ 6–11 min** (verify up to 5 min, the undo's own verify up to 5 min). The whole leg: 7 min 18 s for
|
||
four steps and one undo. The household's pages: badges „Naprakész" for the three that climbed, and for
|
||
vikunja the undone sentence in both languages („…A doboz automatikusan visszaállította az előző
|
||
változatot és az adatokat — semmi nem veszett el." / "…put back the previous version and its data
|
||
automatically — nothing was lost."); event `app_update_undone` (dropped on 9202 — no hub, as designed).
|
||
During the db-dump leg every press was refused `busy` (transient, retried next pass).
|
||
|
||
**5. What the spike found that the build must handle** (rows filed):
|
||
- **`ladder_steps_left` and the badge are STALE after a step ends `done`** until the next scan — the first
|
||
caller run re-pressed a current app four times (R-678). The leg rescans after every step, and the
|
||
update's own finish should refresh them.
|
||
- **An Update pressed on an app already at the head runs the whole guarded update** — backup, pull,
|
||
restart, `done` — with no refusal (R-679: navidrome restarted four times for nothing). The leg never
|
||
presses an app whose pin equals the head; the preflight should refuse with reason `current`.
|
||
- **The product does not remember a failed step.** After vikunja's undo the badge still offers the
|
||
same step; only the caller's memory kept it from a re-press. Decision 15 needs it on the box: the
|
||
failed `to` is recorded per app and the leg skips it until the catalog's ladder changes.
|
||
|
||
**6. The build, in order (≈ 3 evenings, unchanged):** (a) the leg as a function called from the
|
||
off-site job on every path; one app at a time, one step per press, rescan between steps, stop at W+5h;
|
||
(b) the failed-step record + `current` refusal (R-679); (c) the gate's `UpdateLegActiveFn`; (d) the
|
||
per-box switch `app_update.unattended`, **default ON (decision 12)**, and `stacks.update_window`
|
||
removed; (e) the live proof: this night's four apps, with the window moved forward through the
|
||
product's own form (`POST /backups/window` — measured: the three legs reschedule at once).
|
||
|
||
**6.4.2 as BUILT (v0.271.0) — where the build differs from the brief above, named:** (a) the leg is
|
||
`stacks.RunUpdateLeg`, chained by `chainUpdateLeg` inside the `offbox-backup` job (every path, a panic included);
|
||
**one step per app per night** (decision 33 — not "rescan between steps" in one night); (b) the failed-step record is
|
||
`app.yaml` `failed_update_step` tied to the ladder's print (R-680); R-679's `current` refusal shipped in v0.270.0;
|
||
(c) the gate interlock is `quiesce.Options.UpdateLegFn` with decision 31's in-flight grace; (d) the switch is
|
||
`settings.json` `app_update.unattended`, absent = ON, a card on the settings page; `stacks.update_window` removed;
|
||
(e) the live proof: `audits/DRILL-night-2026-09-25.md` Part C (four nights on 9202) and Part D (the demo boxes).
|
||
**The demo boxes' first real automatic night (2026-09-25):** demo-felhom's leg ran at 04:15:46 after the off-site
|
||
copy and took opengist 1.13 → 1.15 in 20 s; demo-hp's found nothing to take. Both summaries reached the hub.
|
||
**Found live, not in the brief:** the leg presses the next app in the same second the previous step ends, so "a
|
||
controller killed BETWEEN two apps" lands inside the next step's early phases — which the guarded update puts back
|
||
(Scenario G); the rest of the night is lost (R-686).
|
||
|
||
### 6.4.1 (record) The update night — the drill brief that preceded the rulings, costed and re-costed
|
||
|
||
**The ruling is decision 6: all 53 apps, through the nightly rotation.** This is an ORDER inside that
|
||
ruling, not a scope change. The database apps go first because they are the ones where a wrong answer
|
||
costs data rather than uptime.
|
||
|
||
**The real numbers this rests on** (R-462, measured 2026-09-06): a successful edge takes
|
||
**6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly 8×, because a negative is
|
||
only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**. **Machine
|
||
time is not the cost — fixtures are.** Two of the three apps needed a bespoke non-browser seed route,
|
||
one needed two attempts and a discarded approach, and one (bookstack) can only ever be half-proven
|
||
headlessly (R-460).
|
||
|
||
| leg | what | cost |
|
||
|---|---|---|
|
||
| A | the **15 database services** — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus `adventurelog`'s postgis — one edge each, fixture per app | **15–25 CC-hours**, dominated by seed routes; ~30 min machine time at the median; ~25 GB |
|
||
| B | ~~one power cut mid-update~~ **DONE 2026-09-21 (R-610)** — measured THREE times, two cut mechanisms, three apps: `pulling` (R-520) and the dangerous post-start case three times over. All ended honest; vikunja's 2.6.0 migration had already run when the power went and the data read back intact. **What remains: a cut landing inside `starting` itself** (it lasts well under a second; needs an in-process fault injector, not a faster shell) | 0 — spent |
|
||
| C | one **PostgreSQL `pg_upgrade` rehearsal**, the Q5 edge, on one app before any of the eleven | 3–4 CC-hours |
|
||
| D | one **downgrade refusal** | **already done** — v0.260.0, proven live 2026-09-21 |
|
||
| E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image |
|
||
| F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised |
|
||
|
||
**RE-COSTED 2026-09-21 FROM THE NIGHT'S REAL NUMBERS** (`audits/DRILL-update-night-2026-09-21.md`):
|
||
|
||
| leg | status after the update night |
|
||
|---|---|
|
||
| A — the 15 database services | **LARGELY DONE.** 21 edges across 19 apps walked box-side in one night, including both engines and 8 database-carrying apps. **Machine time was never the cost and is now known: a proven edge took 11-218 s, median ~45 s.** The cost was fixtures, exactly as costed — and the real surprise is that two apps can NEVER be seeded headlessly while the catalog rightly closes their sign-up (R-624) |
|
||
| B — the power cut | **COMPLETE.** The two EARLY phases nobody had cut in were cut tonight: `backing-up` (the box recovered and said so) and `safety-dump` (nothing moved, nothing to say). Only a cut inside `starting` itself remains, and it still needs an in-process fault injector |
|
||
| C — the PostgreSQL rehearsal | **DONE and COSTED**: ~9 s of engine work, 155.9 s end to end for 49 MB / 48 tables. `pg_upgrade` still owed and may prove unnecessary |
|
||
| D — the downgrade refusal | already done, v0.260.0 |
|
||
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
|
||
| F — the remaining apps | **DONE 2026-09-22 (`audits/DRILL-the-28-2026-09-22.md`): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now **53 of 53 attempted**, not 25.** 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (`outline 1.9.1 -> 1.10.1`, HELD with the right sentence); 2 could not be deployed, one of them (`plant-it`) **by design** — it is `lifecycle: abandoned` and the product's lifecycle gate refused it, proven live for the first time. **The cost is now known and it is not machine time:** three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that `POST /api/backup/run` is BOX-WIDE, so concurrent walks serialise on it. **This leg also added the night's biggest finding**, which no count would have produced: R-630 |
|
||
|
||
**The words "when it resolves to a container" are v0.262.0's, and they are the whole of R-630.**
|
||
Before it, the probe branch had no exit: `findProbeContainer` returning `""` set a message and
|
||
looped, while the settle path that judges an app declaring NO check sat in the outer `else`,
|
||
unreachable. So a stack whose probe resolved to nothing could only ever time out — and
|
||
`failAndHold` then stopped an app whose containers were all healthy. **A stack with no probe is not
|
||
healthy and not failing; it is settled on container state (§3), and never a reason to stop a running
|
||
app.** Proven live on paperless-ngx 2026-09-22: the identical Update that ended **`failed` at
|
||
+313.0 s with the app stopped** now ends **`done` at +53.4 s**
|
||
(`audits/v0262-live-2026-09-22/live262.json`). Which container is probed is now decidable too —
|
||
exact stack name, then `healthcheck.container`, then a UNIQUE prefix, else nothing with the
|
||
candidates logged; the old rule took the FIRST prefix match.
|
||
|
||
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
|
||
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
|
||
answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is
|
||
now the first thing Slice 6 has to be safe against, ahead of everything in this table.
|
||
|
||
**Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C
|
||
and E are the ones that unblock a decision; leg A is the one that takes the time.
|
||
|
||
**Venue:** a scratch host, never a customer box — `demo-hp`'s guest 9202 for the box-side legs, the
|
||
harness on DooPlex for the image-side ones.
|
||
|
||
|
||
## 6.5 The drill catalog and the image store — the standing method for update drills
|
||
|
||
**Why this section exists.** On 2026-09-21 an afternoon session put a deliberately broken image into
|
||
the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached
|
||
a customer, but the method was wrong and the brief that asked for it said so. This is the method that
|
||
replaces it, proven the same night.
|
||
|
||
**The rule, and it has no exception:** *nothing broken, dummy, cross-repo or engine-major ever enters
|
||
the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that,
|
||
the leg is skipped and named.*
|
||
|
||
### The two mechanisms
|
||
|
||
| | what it is | what it makes possible |
|
||
|---|---|---|
|
||
| **the drill catalog** | `admin/app-catalog-drill` on Gitea — private, a copy of the live catalog's `main` | a scratch box can be pointed at a catalog where a failing edge is *committable*, because it carries none of the live repo's gates |
|
||
| **the image store** | a `registry:2` container on the scratch guest at `127.0.0.1:5000` | an edge that **passes the within-a-major test and still fails** — the one shape a real catalog move cannot produce |
|
||
|
||
**The image store is not a convenience.** `09` §3b Q4 could not be measured for a year of drills
|
||
because the only failing edges available were across-a-major, and the within-a-major rule — correctly
|
||
— refuses those before the guarded update is ever reached. *The rule that makes automatic updates
|
||
safe is the same rule that refuses the obvious way to break one.* Measuring an unattended HOLD needs
|
||
`drill/<app>:X.Y.Z` (the real image, retagged) against `drill/<app>:X.Y.(Z+1)` (a built image that
|
||
starts, stays up and never serves) — same repository, same major, plain version tags. A third
|
||
flavour, a tag simply **absent** from the store, gives the pull-failure leg.
|
||
|
||
`stacks.CompareImageRefs` orders a `host:port/` reference correctly: `splitImageRef` takes the last
|
||
colon and rejects it only when a `/` follows, so a registry port is never read as a tag. **Proven by
|
||
running it**, four positive cases and a negative control, 2026-09-21.
|
||
|
||
### Creating the drill repo — and the step that mails the operator if you skip it
|
||
|
||
`POST /api/v1/repos/migrate` is the route that works (the project's Gitea tokens carry
|
||
`write:repository` but not `write:user`, so `POST /user/repos` answers 403). **A migrated repo
|
||
inherits `has_actions: true`, and the catalog's CI workflow comes with it.** Every drill push then
|
||
runs that workflow, it fails — the drill repo is deliberately gate-less — and each failure **mails
|
||
`admin@felhom.eu`**. The 2026-09-21 night sent **47 such alarms overnight**, into the same mailbox
|
||
that was carrying real off-site alarms at the time (**R-629**).
|
||
|
||
So creating the drill repo is two acts, not one:
|
||
|
||
```
|
||
POST /api/v1/repos/migrate {"repo_name":"app-catalog-drill","private":true, …}
|
||
PATCH /api/v1/repos/admin/app-catalog-drill {"has_actions": false}
|
||
```
|
||
|
||
then read the repo back and quote `has_actions: False` — an alarm channel trained to be ignored is
|
||
worse than no alarm channel, and R-168 made CI mail the thing that notices a bypassed gate.
|
||
|
||
### Pointing a box at the drill catalog — the step that is NOT obvious
|
||
|
||
**`git.repo_url` alone is inert.** `Syncer.gitCloneOrPull` clones only when the cache has no `.git`;
|
||
otherwise it fetches from the remote the clone already stores. The cache directory must be removed as
|
||
well, or the box goes on following the live catalog and reports success. Filed as **R-615**; until it
|
||
is fixed, the drill procedure is:
|
||
|
||
1. save `controller.yaml` as `controller.yaml.pre-update-night`;
|
||
2. set `git.repo_url` (and `username`/`token` — the drill repo is private);
|
||
3. **remove `<data>/catalog-cache`**;
|
||
4. restart the controller, sync, **rescan** (R-607: a sync can answer „nincs változás" while the
|
||
catalog has moved, and the badge answers from the stale value until the rescan);
|
||
5. **three controls, all quoted in the report** — the drill bump appears on the scratch box; the
|
||
other boxes' caches are unchanged; the live catalog's `main` hash is unchanged.
|
||
|
||
### What the drill must leave behind
|
||
|
||
- `controller.yaml` restored from the saved copy, the controller restarted, and `git.repo_url` **read
|
||
back and quoted** as the live catalog.
|
||
- The registry container and its volume removed; drill images removed **by name**. Never `prune`.
|
||
- The drill repo **kept**, private, reset to the live catalog's `main`, so the next drill starts clean.
|
||
- A diff of every `image:` line against the live catalog's `main` — expected: identical.
|
||
|
||
### The fence
|
||
|
||
Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no
|
||
customer box has credentials for it. The store listens on the guest's loopback only.
|
||
|
||
## 7. What slices 1 and 2 actually built
|
||
|
||
### 7.1 The record (slice 1)
|
||
|
||
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
|
||
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
|
||
and writes `app.yaml`:
|
||
|
||
```yaml
|
||
installed_images:
|
||
web:
|
||
ref: lscr.io/linuxserver/bookstack:26.05.2
|
||
digest: sha256:… # "" if the image was never pulled from a registry
|
||
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
|
||
```
|
||
|
||
Three rules, each with its reason:
|
||
|
||
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
|
||
already moved. A record built from it would answer "what will happen next time something runs
|
||
`up -d`", which is a different question.
|
||
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
|
||
Intent refused, observation logged. Refusing to start a customer's app because we could not write
|
||
down which version it is trades a real outage for a bookkeeping gap.
|
||
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
|
||
complete record with a partial one.
|
||
|
||
**Seeded at startup (v0.234.0).** `Manager.BackfillInstalledImages` runs once at boot, beside the
|
||
desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY
|
||
on. It only READS containers. **This was not a refinement — without it the feature did not reach a
|
||
quiet box at all:** see §8.3, which was written as a known limitation on 2026-09-02 and was a defect
|
||
by the next morning.
|
||
|
||
Two admission rules, and the second is the design:
|
||
|
||
- **It never overwrites an existing record.** The bring-up paths own updates; this fills gaps only.
|
||
- **It refuses to seed a PARTIAL observation.** §7.2's comparison reads a service-count mismatch as
|
||
BEHIND, so a degraded or crash-looping app seeded from its visible containers would render
|
||
„Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial
|
||
because they follow a SUCCESSFUL `up -d`, where a gap is real news and is logged; a backfill meets a
|
||
box in whatever state it is in. **Same field, two writers, two different admission rules — that is
|
||
deliberate and must not be "made consistent".**
|
||
|
||
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
|
||
|
||
### 7.2 The label (slice 2)
|
||
|
||
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
|
||
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
|
||
no new CSS.
|
||
|
||
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
|
||
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
|
||
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
|
||
a test with a companion red-proof.
|
||
|
||
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
|
||
Version strings stay in the logs, the API and the hub.
|
||
|
||
---
|
||
|
||
## 8. Known limitations, stated plainly
|
||
|
||
1. **„Naprakész" can be FALSE for the floating pins, and 2026-09-21 measured HOW false.** The
|
||
comparison is reference-to-reference and queries no registry — a customer's box must not depend on
|
||
reaching eight upstream registries to render a page. For `postgres:16-alpine`, `mariadb:11.6` and
|
||
the others the reference can be identical while the image behind it has moved.
|
||
**NUMBERS, 2026-09-21** (`audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3). **The count this document
|
||
carried — "23 of 66" — is STALE and matched no definition the catalog supports today.** Recounted
|
||
at catalog `18a6d2d8`, with the definition stated so it can be rechecked: a pin FLOATS when its
|
||
tag names a version LINE rather than an exact release. Of 66 unique pins, 48 are full `X.Y.Z`,
|
||
**6 are two-part lines** (`mariadb:11.4`/`11.6`/`12.3`, `claper:2.5`, `opengist:1.13`,
|
||
`wger/server:2.6`) and **4 are major lines** (`postgres:15-alpine`, `postgres:16-alpine`,
|
||
`redis:7-alpine`, `postgis:16-3.5-alpine`) — **10 float**. The remaining 8 are exact versions
|
||
wearing a variant suffix (`ghost:6.53.0-alpine`, `nextcloud:34.0.1-apache`, …), which do not float
|
||
by this definition. Of the 8 database and cache engine pins the sweep measured, 7 were measurable
|
||
and **6 have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine`
|
||
(6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`. Only `mariadb:11.6` has not.
|
||
The 8th, immich's own ghcr build, is UNMEASURED — ghcr exposes no anonymous last-modified. So on
|
||
demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.
|
||
Digest-level comparison needs the catalog to record the digest at push time — **R-446**, put to the
|
||
operator as §3b **Q6**, recommended YES.
|
||
2. **~~Nothing enforces `catalog_since`.~~ Enforced by the pre-push hook since 2026-09-13 (R-452, `app-catalog-felhom.eu/scripts/check-catalog-since.py`); CI's shallow clone still skips it out loud.** A commit that moves an `image:` line and forgets the date
|
||
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
|
||
commit to diff against, so the gate needs a deeper fetch — **R-452**.
|
||
3. ~~**The record only appears after the next lifecycle action.**~~ **CLOSED in v0.234.0, and the way
|
||
it closed is worth keeping.** This was written on 2026-09-02 as an accepted limitation — *"the
|
||
fleet view fills in gradually"*. The operator looked at demo-felhom the next morning and found
|
||
OpenGist, up 15 hours, running exactly the catalog pin, showing **nothing at all**. **On a quiet
|
||
box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is,
|
||
on the quiet installations, not shipped.** `BackfillInstalledImages` now seeds the absences at
|
||
startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller
|
||
START, so a box between upgrade and its next restart still shows nothing — bounded by one restart
|
||
rather than unbounded.
|
||
4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches
|
||
that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the
|
||
`wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken
|
||
state. Recorded so it is a choice, not a surprise.
|
||
5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in
|
||
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
|
||
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
|
||
the update badge by withholding `catalog_since`. **R-458.**
|
||
6. ~~**The Update button is still unguarded.**~~ **CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1).** It
|
||
refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a
|
||
safety dump, and holds an app that does not come up. **What stays true:** it still has no automatic
|
||
rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) —
|
||
that now ends HELD rather than crash-looping behind a green button.
|
||
7. ~~**An engine major can be applied without its datadir upgrade, and nothing notices.**~~ **CLOSED
|
||
2026-09-13 for MariaDB (R-459):** every `mariadb:` sidecar carries `MARIADB_AUTO_UPGRADE=1`, and the
|
||
harness shows the conversion running on the E3/E3b edges (§3 decision 5). **What stays true:** the
|
||
PostgreSQL half (R-463) has no equivalent — the image performs no `pg_upgrade` — and the
|
||
engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until
|
||
Slice 4 gives the Update button a backup.
|
||
8. ~~**Only three of 53 apps have ever had an upgrade measured.**~~ **WIDENED 2026-09-21 to 21
|
||
EDGES ACROSS 19 APPS** (`audits/DRILL-update-night-2026-09-21.md`), on scratch guest 9202
|
||
through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog
|
||
carried no test reference at any point: **14 proven, 3 failed, 4 inconclusive**, each app
|
||
seeded and read back through its OWN front door with a negative control on every readback.
|
||
Ten of the fourteen printed a verbatim migration line. **What stays true:** bookstack is still
|
||
only half-provable headlessly (**R-460**), and **two apps cannot be seeded AT ALL** while the
|
||
catalog rightly closes their sign-up — vaultwarden (`SIGNUPS_ALLOWED=false`, R-512) and
|
||
zipline — which is a permanent ceiling on R-462's scope rather than a fixture nobody has
|
||
written (**R-624**). **And one thing this widening FOUND that no count would have:** the
|
||
`verifying` phase trusts the `.felhom.yml` probe absolutely, and **three of the 53 templates
|
||
name a probe the app does not answer**, so a SUCCESSFUL update of those apps ends by STOPPING
|
||
a working app (**R-618**, P1 — tandoor measured serving HTTP 200 on the new version at four
|
||
samples across five minutes, then stopped).
|
||
**CLOSED 2026-09-22, and the closing changed the numbers above:** the three probes were
|
||
corrected (`app-catalog-felhom.eu@793c4fb`), red-proofed live on 9202 in both directions, and a
|
||
static gate now refuses a probe that does not match the same service's own compose healthcheck.
|
||
**tandoor's edge was re-walked with nothing else changed and ended `done` at +41.1 s** where it
|
||
had ended `failed` at +361.9 s — so the tally is **15 proven, 2 failed, 4 inconclusive**, and all
|
||
fifteen are on the live catalog since 2026-09-22.
|
||
**WHAT THE CLOSING FOUND, and it is the part worth carrying forward:** a wrong probe was the
|
||
LOUD failure. Two quiet ones sit beside it. `paperless-ngx` has no container whose name matches
|
||
its stack name, so `findProbeContainer` returns nothing and **its probe has never run on any box**
|
||
— an absence, with no badge to contradict it (**R-630**). And five more templates cannot be judged
|
||
statically at all, one of which (`home-assistant`) is correct only because its check type cannot
|
||
fail (**R-631**). **So "the probe is right" is now enforced for 47 of 53 templates and still
|
||
unknown for six.**
|
||
**AND THE SWEEP'S REAL CEILING, counted rather than felt: 28 of the 53 templates have never been
|
||
deployed by any drill** (**R-632**) — the widening above went from 3 apps to 21, and 21 is not 53.
|
||
**CLOSED THE NEXT NIGHT, 2026-09-22: all 28 were walked** (`audits/DRILL-the-28-2026-09-22.md`),
|
||
so every template in the catalog has now been attempted at least once. **And the walk that closed
|
||
it found something the probe work had left open.** `paperless-ngx` has no container whose name
|
||
matches its stack name, so `findProbeContainer` returns nothing and its probe has **never run on
|
||
any box**. Asked what `verifying` does with no probe to wait on, the answer is the worst of the
|
||
three: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three
|
||
containers all read `healthy`. The controller names it itself — *`not healthy within 5m0s (last:
|
||
no probe container) — stopping and HOLDING the app`* — at **+313.0 s**, front door 404 afterwards.
|
||
**So limitation 8 now has two shapes, not one:** a probe that names the wrong target (R-618,
|
||
fixed) and **no probe at all** (**R-630, raised to P1**), and the static gate can see the first
|
||
but not the second, because there is nothing to compare.
|
||
**FIXED 2026-09-22 in controller v0.262.0**, and the fix is in the phase table above: the
|
||
no-probe case now settles on container state instead of looping, and the probe TARGET is
|
||
decidable (exact name → `healthcheck.container` → a UNIQUE prefix → nothing, candidates logged).
|
||
`paperless-ngx` and `immich` carry the explicit field, and the catalog gate REFUSES a probe that
|
||
resolves to nothing rather than warning about it. **So limitation 8's two shapes are both closed
|
||
in the product**; what remains is that six templates still cannot be judged STATICALLY (R-631,
|
||
all five read live and correct) and that a probe can still be right about the port and wrong
|
||
about what a 200 means — `romm` answered 200 from nginx while its workers were being OOM-killed
|
||
for six hours (**R-635**), which is a third shape again and the reason "the update is guarded"
|
||
must never be read as "the new version runs".
|
||
**Two more things the same night measured, both about state rather than health:** a `remove` sent
|
||
while a restore is still running reports success and leaves a container restarting with a live
|
||
public route (**R-633**) — and the product already has exactly that guard for `update` and for
|
||
`restore`, which name the blocking operation, but not for `remove`; and an app can be **running,
|
||
healthy and serving while recorded as `deployed: false`**, in which state the product refuses to
|
||
remove it at all (**R-634**, reproducible alone on `sparkyfitness`).
|
||
9. **The hub does not record image tags at all.** Its report's container payload carries name, state,
|
||
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
|
||
change; it is not derivable from what is already reported.
|
||
|
||
---
|
||
|
||
## 9. Where the rest lives
|
||
|
||
| what | where |
|
||
|---|---|
|
||
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
|
||
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
|
||
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
|
||
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
|
||
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |
|