docs(update arc): record the operator's rulings of 2026-09-23 (09 §3 decisions 11-18)
gates / gates (push) Successful in 29s

- 09 §3: decisions 11 (one window = a leg of the backup chain), 12 (automatic,
  per-box switch on by default), 13 (the test decides, not the tag - replaces
  decision 3's "never across a major"), 14 (the ladder), 15 (the box undoes a
  failed update - replaces §6.1's no-auto-undo), 16 (Postgres majors converted
  by the box), 17 (digests), 18 (fleet view, later).
- 09 §3b marked ANSWERED with a pointer per question; kept as the reasoning.
- 09 §6.2 rewritten to the ruled shape; §6.1 abort paragraph and §4 point at 15.
- Register: R-450, R-451, R-446, R-463 cite the decisions.
- STATUS: the seven questions no longer wait on the operator.

Documents only. No product code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 07:55:04 +02:00
parent 267dcad01b
commit 805ad1e962
3 changed files with 155 additions and 48 deletions
+4
View File
@@ -1,5 +1,9 @@
# STATUS — what works, what's broken, what's next # STATUS — what works, what's broken, what's next
**Updated 2026-09-23 (morning) — your seven answers on automatic updates are written into the design notes as rulings.** No later session will ask them again. The two new pieces your answers need — the automatic undo and the step-by-step climb — are being tested by hand on the scratch machine today, before anyone builds them. The build plan comes back to you part by part.
---
**Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.** **Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.**
**Decisions I took on my own: none.** **Decisions I took on my own: none.**
@@ -98,7 +98,8 @@ These are rulings, not proposals. Anything specced against a different assumptio
`catalog_since` exists and why no version string is shown. `catalog_since` exists and why no version string is shown.
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human, 3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
because §4 says it cannot be undone. because §4 says it cannot be undone. **Second half REPLACED 2026-09-23 by decision 13** (the test
decides, not the tag); the first half is confirmed by decision 12.
### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0 ### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0
@@ -202,17 +203,90 @@ These are rulings, not proposals. Anything specced against a different assumptio
withholds an act. **Implementation:** `stacks.CatalogOrder`, one verdict read by both the badge withholds an act. **Implementation:** `stacks.CatalogOrder`, one verdict read by both the badge
and `UpdatePreflight`. and `UpdatePreflight`.
### 2026-09-23 — the operator answers §3b's seven questions
Operator rulings, not CC decisions. Each answers one question in §3b, which is kept, marked ANSWERED,
because the costs written there are the reasoning behind these. **Nothing below is built yet** — the
build order is §6.4, and the two new mechanisms (15 and 14) were spiked the same day before anyone
builds them (`audits/update-rulings-2026-09-23/`).
11. **One maintenance window, and it is the backup window the household already sets** (Q1). App
updates are one more leg of the nightly chain, **after the off-site copy and before the
full-system backup**. An update not finished when the full-system backup is due waits for the
next night. **Why:** the update then leans on the freshest copy the box ever has, and there is no
second window for anyone to set or misread. **Replaces:** §3b Q1's recommendation of a fixed
02:30–05:00. *For the builder:* `stacks.update_window` (`config.go` L195, default `03:00-05:00`)
exists and **no Go code reads it** — it appears only in `config.go` (field + default),
`setup/handlers.go` L482 (written into a new box's `controller.yaml`) and
`configs/controller.yaml.example`. Remove it or fold it into the chain; do not build a second
window.
12. **Automatic updates stay, behind a per-box switch that is ON by default** (Q2). Confirms the first
half of decision 3. **Today every update is still manual; nothing automatic is built.**
13. **The test decides, not the tag** (Q3). **REPLACES decision 3's second half, "never across a
major".** The box applies by itself every step the catalog holds, because **the catalog holds only
tested steps**; the size of the version jump does not matter, the test does. The box does not
parse tags to decide. The tag rule (`stacks.CompareImageRefs`) moves to the **catalog gate** as a
push-time safety net: an image move with no test record is refused there. Two exceptions, both
**marks the test sets on the step**: *files may change* → automatic only when a fresh copy holds
the files, else a person (this is Q2's refinement); *needs a person* → the tester writes why.
**Why:** decision 3 was a proxy. "Within a major" was standing in for "known to work", and the
update night measured the proxy wrong in both directions — minor moves that held
(adventurelog, outline) and a tag shape it cannot read at all (§3b Q3). The test record is the
thing the proxy was guessing at.
14. **The update ladder** (Q3). A box more than one step behind **climbs one tested step at a time, in
order, and never jumps**. Each step is the full guarded update of §6.1. **Why:** a step was tested
from the version before it, not from three versions before it; R-40's multi-major jump is what
this prevents. This is the "stepping" §6.1 already assigned to slice 6.
15. **The box undoes a failed update itself** (Q4). **REPLACES §6.1's "the box never puts the old
version back by itself".** On a failed health check: the old definition and pin back + the
database copy taken seconds before the pin moved loaded back + the health check again. The app
stays stopped (HOLD) **only if the undo itself fails**. **Why the old ruling does not bind this:**
the old version refused to start on data the new one had migrated (Nextcloud, docmost — §4). The
undo puts the **pre-migration data** back too, so the old version meets the data it knows. The
word "rollback" stays struck; **the name is undo**. The household is told on the app page and by
one mail; the operator by event; **no retry** until the catalog moves or a person presses.
Spiked 2026-09-23 before any build — §6.1a.
16. **PostgreSQL majors are converted by the box** (Q5). A guarded-update step: save everything from
the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on
the test bench before the catalog may move it. The engine-major gate stays until then. **Why:** the
update night costed it at ~9 s of engine work for 49 MB (§3b Q5); moving the pin and letting it
hold would take eleven apps down on one night.
17. **The catalog records the image digest of every pin at push time** (Q6). The box compares against
it; where the catalog carries one, the box pulls **that exact image** — which also makes a floating
tag reproducible, not only the badge honest. *Whether compose can pull by a recorded digest while
the definition names a tag is a claim for the build plan to verify, not a ruling on mechanism.*
18. **Fleet view** (Q7). The report carries, per compose service (so the database too), the installed
reference, the catalog reference and the badge state. **Built later**, when the fleet grows.
**RomM follow-ups, operator-agreed the same day:** the test bench watches memory after an update
(`upgrade-test.py`, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4);
R-636's louder repeated alarm.
--- ---
## 3b. OPEN — the seven questions Slices 6 and 7 need answered ## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed
**These are questions, not rulings. CC does not decide them.** Each is one answerable sentence, the > **ALL SEVEN ANSWERED by the operator on 2026-09-23 — §3 decisions 11–18.** Q1 → 11; Q2 → 12 and
> 13's *files may change* mark; Q3 → 13 and 14; Q4 → 15; Q5 → 16; Q6 → 17; Q7 → 18. Kept below
> unedited, because the costs and measurements here are the reasoning behind those rulings. Where a
> ruling differs from the recommendation below, the ruling wins: Q1 (the household's backup window,
> not a fixed 02:30–05:00), Q3 (the test decides, not `CompareImageRefs`), Q4 (the box UNDOES before
> it holds).
**These were questions, not rulings. CC did not decide them.** Each is one answerable sentence, the
options, what each costs, the recommendation, and what happens if nothing is decided. The measurement options, what each costs, the recommendation, and what happens if nothing is decided. The measurement
behind them is `audits/UPDATE-ARC-STATE-2026-09-21.md`; the short version is that **46 of the behind them is `audits/UPDATE-ARC-STATE-2026-09-21.md`; the short version is that **46 of the
catalog's 58 exact pins are behind upstream today and 39 of those are within a major** — the catalog's 58 exact pins are behind upstream today and 39 of those are within a major** — the
population §3 decision 3 already says may move without a human, and nobody presses 39 buttons. population §3 decision 3 already says may move without a human, and nobody presses 39 buttons.
### Q1 — When may a box update itself? ### Q1 — When may a box update itself? — ANSWERED 2026-09-23: §3 decision 11
*May the box run the guarded Update by itself between 02:30 and 05:00, nightly?* *May the box run the guarded Update by itself between 02:30 and 05:00, nightly?*
@@ -241,7 +315,7 @@ held app does not block the controller's updates for ever. So the window may con
jobs will queue behind one another rather than meet. **What it does NOT do is reorder them**: if the jobs will queue behind one another rather than meet. **What it does NOT do is reorder them**: if the
operator prefers the box to take its own update first, that is a scheduling choice still open here. operator prefers the box to take its own update first, that is a scheduling choice still open here.
### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES? ### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES? — ANSWERED 2026-09-23: §3 decisions 12 and 13 (the *files may change* mark)
*The button's rule and the automatic rule can differ. Should they?* *The button's rule and the automatic rule can differ. Should they?*
@@ -285,7 +359,7 @@ untouched by that and remains the operator's.
`failAndHold` removes the containers, so the failing version's own output is gone within seconds `failAndHold` removes the containers, so the failing version's own output is gone within seconds
(**R-621**). With a person pressing, they at least watched it happen. (**R-621**). With a person pressing, they at least watched it happen.
### Q3 — What counts as "within a major" when the tag is not a version number? ### Q3 — What counts as "within a major" when the tag is not a version number? — ANSWERED 2026-09-23: §3 decisions 13 and 14
*§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`, *§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`,
`kimai/kimai2:apache-2.57.0`, a date stamp, a digest?* `kimai/kimai2:apache-2.57.0`, a date stamp, a digest?*
@@ -317,7 +391,7 @@ hold measurable at all**: the drill edge `localhost:5000/drill/glance:1.0.0 -> :
within-a-major test and still fails, which no real catalog move does. The recommendation is within-a-major test and still fails, which no real catalog move does. The recommendation is
unchanged; the *same major?* extension it already names is still owed. unchanged; the *same major?* extension it already names is still owed.
### Q4 — A held app: who is told, when, and does the box try again? ### Q4 — A held app: who is told, when, and does the box try again? — ANSWERED 2026-09-23: §3 decision 15
*An automatic update that ends HELD happened while everyone was asleep.* *An automatic update that ends HELD happened while everyone was asleep.*
@@ -374,7 +448,7 @@ catalog publishes a FIXED newer version — the household's only route out is th
**Also proven across a genuine power cut:** the boot sweep met a held app after an unclean shutdown **Also proven across a genuine power cut:** the boot sweep met a held app after an unclean shutdown
and deliberately left it alone — *„whatever is holding it owns its recovery"*. and deliberately left it alone — *„whatever is holding it owns its recovery"*.
### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`? ### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`? — ANSWERED 2026-09-23: §3 decision 16
*Eleven templates, and the image performs no conversion — it refuses to start on an older major's *Eleven templates, and the image performs no conversion — it refuses to start on an older major's
datadir (R-463).* datadir (R-463).*
@@ -419,7 +493,7 @@ engine is content*.
**The recommendation is unchanged.** Tonight gives it a price rather than a new opinion. **The recommendation is unchanged.** Tonight gives it a price rather than a new opinion.
### Q6 — Should the catalog record each pin's DIGEST at push time? ### Q6 — Should the catalog record each pin's DIGEST at push time? — ANSWERED 2026-09-23: §3 decision 17
*So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.* *So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.*
@@ -456,7 +530,7 @@ the CATALOG set it, which is the right measure for the catalog and not for a box
because the catalog carries no digest. That is precisely this question's proposal, and only the because the catalog carries no digest. That is precisely this question's proposal, and only the
catalog half is missing. **The recommendation is unchanged.** catalog half is missing. **The recommendation is unchanged.**
### Q7 — What does the hub's report need to carry for a fleet view? ### Q7 — What does the hub's report need to carry for a fleet view? — ANSWERED 2026-09-23: §3 decision 18
*Slice 7 lets the operator SEE and MOVE how far behind every box is.* *Slice 7 lets the operator SEE and MOVE how far behind every box is.*
@@ -502,7 +576,10 @@ with a positive control proving the data is intact, only unreachable by the old
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again | | **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy | | **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
There is no third. There is no third. **2026-09-23 (§3 decision 15): the UNDO is the two combined** — the abort's old
image plus a restore from the one copy that is seconds old, the pre-pin safety dump. It is not a
third shape and it is not a rollback: nothing is migrated backwards, the pre-migration state is put
back.
### 4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades ### 4.1 MEASURED 2026-09-06 — and the abort turns out to be a property of the APP, not of upgrades
@@ -632,8 +709,8 @@ earlier feature is the failure mode to look for whenever a file changes meaning.
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 | | **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 | | **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 |
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 | | **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 | | **6** | **Automatic updates** — ~~within a major, a human across one~~ **the catalog's tested steps, climbed one at a time, as a leg of the backup chain, undone by the box on failure** (§3 decisions 11–17); **an engine change gets its own edge.** | OPEN — R-450; rulings 2026-09-23; build order §6.4 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 | | **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451; ruled 2026-09-23 (decision 18), built later |
### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13) ### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)
@@ -692,6 +769,12 @@ be predicted. So the box never puts the old version back by itself. **The route
and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where
the harness has proven it, is slice 6's. the harness has proven it, is slice 6's.
> **REPLACED 2026-09-23 by §3 decision 15 — the box UNDOES a failed update itself.** The measurement
> above still stands and is the reason the replacement has a different shape: putting the old IMAGE
> back alone is what refuses (§4.1). The undo puts the old image back **together with the database
> copy taken before the pin moved**, so the old version meets pre-migration data. Spiked by hand on
> 9202 the same day — §6.1a.
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6. **Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the **The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
@@ -783,55 +866,75 @@ breaks.
--- ---
### 6.2 Slice 6, as it would be built (OPEN — R-450; needs Q1–Q4) ### 6.2 Slice 6, as it will be built (OPEN — R-450; ruled 2026-09-23, decisions 11–17)
**Not a design yet; the shape the seven questions bound.** Written down so the answers have somewhere **The shape the rulings fix. The build order and costs are §6.4.** Rewritten 2026-09-23; the
to land. earlier "as it would be built" draft keyed on the window of Q1's recommendation and on
`CompareImageRefs`, and both were ruled differently.
**Nothing new happens to the app.** Inside the window, for an app that qualifies, the box runs **Nothing new happens to a single step.** Each step is **exactly the guarded Update of §6.1** — same
**exactly the guarded Update of §6.1** — same precondition, same safety dump, same pin journal, same precondition, same safety dump, same pin journal, same health wait — with **one change to its end**:
health wait, same hold. Slice 6 adds a *caller*, not a *path*. That is the whole reason it is a failed health check runs the **undo** (decision 15, §6.1a) before it holds. Slice 6 adds a
affordable: every failure mode was measured in slice 4 and every one of them already ends in a hold *caller*, a *ladder* and the *undo*; it adds no second update path.
the household can read.
**An app qualifies when ALL of these hold** (each clause is a question above, not a decision taken): **When — a leg of the chain, not a window of its own (decision 11).** The household's backup window
start W already drives every nightly leg at fixed offsets (`07` §6.1): DB dump at W, Tier 2 at
W+60m, off-site at W+105m, and the full-system backup's gate opens at **W+2h** (`quiesce.go`
`gateOpenOffsetMin = 120`, span to W+6h). The update leg starts **when the off-site leg has
finished** and stops starting new steps **when the full-system backup starts**; whatever is left
waits for the next night. **Measured consequence the builder must face, not discover:** the gap
between "off-site finished" and W+2h is **at most 15 minutes** and is zero on a night the off-site
leg runs long, while one step takes 51 s – 1 m 26 s when it succeeds and ~5 m when it fails (§3b
Q4). So either the full-system gate learns to wait for the update leg (it has a four-hour span to
spend) or the leg gets almost no time. That is a build choice inside decision 11, named in §6.4.
1. the window is open (Q1); **Which apps — the catalog decides, not the tag (decision 13).** An app qualifies when ALL hold:
2. `stacks.CatalogOrder` says **Behind** — never Unknown, never Ahead (v0.260.0 gives all four);
3. the edge is **within a major for EVERY compose service**, the engine sidecar included (Q3), 1. the per-box switch is on (decision 12; **on by default**);
judged by `stacks.CompareImageRefs` — one unorderable service makes the whole edge *across*; 2. `stacks.CatalogOrder` says **Behind** — never Unknown, never Ahead;
4. a fresh copy exists on a tier that holds what this app's data actually is (Q2); 3. the **next step** from the app's installed state carries a **test record** in the catalog —
5. the app's own switch is on (Q2's default: on). a step with none is never applied by a box (the catalog gate refuses to publish it, decision 13);
4. the step's **marks** allow it: *needs a person* → never automatic; *files may change* → automatic
only when a fresh copy on some tier holds the app's FILES (`UpdateCopyHolds`), else a person;
5. the app is not held (`held` is terminal until a person acts or the catalog moves, decision 15).
**How far — one step at a time (decision 14).** A box two steps behind applies step A→B, then B→C,
each the full guarded update, each with its own health check and undo. A failed step stops the
ladder for that app. The ladder's format is §6.4 / `audits/update-rulings-2026-09-23/`.
**When it fails — undo, then hold only if the undo fails (decision 15).** The household is told on
the app page and by **one** mail, in the box's language (R-606 is a precondition — an automatic
update's mail must be in the household's language); the operator by event; **no retry** until the
catalog moves or a person presses.
**What the household sees.** An event and a line on the app page's timeline, before and after, in both **What the household sees.** An event and a line on the app page's timeline, before and after, in both
languages: *„Automatikus frissítés 03:12-kor — sikeres"* / *„— megállítva, a másolat 2026-09-20-i"*. languages: *„Automatikus frissítés 03:12-kor — sikeres"* / *„— visszaállítva az előző verzióra"* /
A held app is not retried until the catalog moves again or a person presses (Q4). *„— megállítva, a másolat 2026-09-20-i"*.
**Where it would live.** A scheduler beside the existing nightly legs, reading `settings` for the **Where it lives.** A leg in the nightly chain beside the existing ones, calling
window and the per-app switch, and calling `Manager.StartGuardedUpdate`. **It must respect the same `Manager.StartGuardedUpdate`. **It must respect the `isHeld`/`SetUpdatingCheck` interlocks v0.238.1
`isHeld`/`SetUpdatingCheck` interlocks v0.238.1 added** — the nightly capture running *inside* an added** — the nightly capture running *inside* an update's health wait is the defect that release
update's health wait is the defect that release fixed, and a second unattended caller is exactly the fixed, and a second unattended caller is exactly the shape that finds it again. And it must run
shape that finds it again. **one app at a time**: there is no single-flight today (§3b Q4: five Updates pressed within 0.45 s
all ran at once).
**The in-process caller reads `UpdateRefusal.Reason`, and the split is measured, not assumed** **The in-process caller reads `UpdateRefusal.Reason`, and the split is measured, not assumed**
(v0.261.0, R-609): `busy`, `updating`, `deploying`, `migrating` and `self_updating` are **transient** (v0.261.0, R-609): `busy`, `updating`, `deploying`, `migrating` and `self_updating` are **transient**
— try again on the next pass; `held` and `downgrade` are **terminal** — never press that app again — try again on the next pass; `held` and `downgrade` are **terminal** — never press that app again
until a person acts; `memory`, `disk` and `no_backup` need a person and should be surfaced, not until a person acts; `memory`, `disk` and `no_backup` need a person and should be surfaced, not
retried. **Before the reason reached the wire the only safe readings were "give up on everything" or retried. A working caller in this exact shape exists as evidence, not product:
"press for ever"**, which is why this is listed as a dependency of the slice rather than a detail of
it. A working caller in this exact shape exists as evidence, not product:
`audits/update-arc-gaps-2026-09-21/unattended-caller.py`. `audits/update-arc-gaps-2026-09-21/unattended-caller.py`.
**Ships behind `app_update.unattended: false` with no UI until the operator answers Q1.** **The per-box switch key is `app_update.unattended`**, default **true** (decision 12).
⚠ **NOT `auto_update` — that name is TAKEN, and by the very thing this must not collide with.** ⚠ **NOT `auto_update` — that name is TAKEN, and by the very thing this must not collide with.**
`self_update.auto_update` / `self_update.auto_update_time` (`config/config.go` L280-281, default `self_update.auto_update` / `self_update.auto_update_time` (`config/config.go` L280-281, default
**04:30** at L422) are the CONTROLLER's own update. An earlier draft of this section said Slice 6 **04:30** at L422) are the CONTROLLER's own update. Two settings with that name, one meaning the
"ships behind `auto_update: off`"; two settings with that name, one meaning the controller and one controller and one meaning apps, is the kind of collision that is only discovered by an operator who
meaning apps, is the kind of collision that is only discovered by an operator who turned off the turned off the wrong one. The app-scoped key is `app_update.*`. **And `stacks.update_window` is
wrong one. The app-scoped key is `app_update.*`. removed or folded in, never a second window** (decision 11).
### 6.3 Slice 7, as it would be built (OPEN — R-451; needs Q7) ### 6.3 Slice 7, as it would be built (OPEN — R-451; ruled 2026-09-23, decision 18 — built later)
Three additive pieces, and the transport already exists (§3b Q7): Three additive pieces, and the transport already exists (§3b Q7):
+4 -4
View File
@@ -666,15 +666,15 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | | **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. | **WAITING-ON-OPERATOR — `09` §3b Q6; owner: CC once answered** | | **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. **-- RULED 2026-09-23 (`09` §3 decision 17):** YES — the catalog records the image digest of every pin at push time; the box compares against it and, where the catalog carries one, pulls **that exact image**, which makes a floating tag reproducible, not only the badge honest. *Pull-by-digest while the definition names a tag is a claim to verify in the build, not a ruling on mechanism.* | **READY TO BUILD — owner: CC; `09` §6.4** |
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. | **WAITING-ON-OPERATOR — `09` §3b Q1–Q4 (own-edge half SHIPPED); owner: CC once answered** | | **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. **-- RULED 2026-09-23 (operator, `09` §3 decisions 11–15):** Q1–Q4 answered. The update is a leg of the backup chain after off-site and before the full-system backup (11); automatic with a per-box switch ON by default (12); **the TEST decides, not the tag** — the box applies every step the catalog holds because the catalog holds only tested steps, and `CompareImageRefs` moves to the catalog gate (13, REPLACES "never across a major"); a box behind climbs **one tested step at a time** (14); **the box UNDOES a failed update itself** — old definition + the pre-pin safety dump + health check again, HOLD only if the undo fails (15, REPLACES §6.1's no-auto-undo). The undo and the ladder were SPIKED the same day before any build (`audits/update-rulings-2026-09-23/`); build order and costs in `09` §6.4. | **READY TO BUILD — owner: CC; `09` §6.4 part by part, each part returns to the operator for go/no-go** |
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. | **WAITING-ON-OPERATOR — `09` §3b Q7; owner: CC once answered** | | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. **-- RULED 2026-09-23 (`09` §3 decision 18):** the report carries, per compose service (database included), the installed reference, the catalog reference and the badge state. **Built later, when the fleet grows** — Q7's recommendation, confirmed. | **RULED — build deferred until the fleet grows; owner: CC** |
| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** |
| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 **— UPDATE NIGHT 2026-09-21:** **MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states.** A `.felhom.yml`-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN `bentopdf` — installed `v2.8.6`, catalog ahead. §5.4's asymmetry is confirmed live: the new `.felhom.yml` reached the box while the compose `image:` line stayed `v2.8.6`. **But no false alarm was produced**: ten samples over two minutes all read `state=running` with the front door at `200`. The reason is the probe's own semantics, not luck — `healthprobe.go:258-261` treats **any response** as healthy for `type: http`, and the bogus path answers 404, which is a response. **So this row's false-alarm risk exists only for `type: api` probes carrying an `expect` block**, where the status is compared; for every `type: http` template and every `type: api` without `expect`, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: `audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md`. | **READY — rank P3-LOW; owner: CC** | | **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 **— UPDATE NIGHT 2026-09-21:** **MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states.** A `.felhom.yml`-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN `bentopdf` — installed `v2.8.6`, catalog ahead. §5.4's asymmetry is confirmed live: the new `.felhom.yml` reached the box while the compose `image:` line stayed `v2.8.6`. **But no false alarm was produced**: ten samples over two minutes all read `state=running` with the front door at `200`. The reason is the probe's own semantics, not luck — `healthprobe.go:258-261` treats **any response** as healthy for `type: http`, and the bogus path answers 404, which is a response. **So this row's false-alarm risk exists only for `type: api` probes carrying an `expect` block**, where the status is compared; for every `type: http` template and every `type: api` without `expect`, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: `audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md`. | **READY — rank P3-LOW; owner: CC** |
| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 **-- UPDATE NIGHT 2026-09-21:** bookstack's edge was walked again on 2026-09-21 and is again **half-proven**: the database half read back through `php artisan` with its own negative control, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness. Two more apps joined the same class tonight for a different reason (R-624). | **READY — rank P3-LOW; owner: CC** | | **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 **-- UPDATE NIGHT 2026-09-21:** bookstack's edge was walked again on 2026-09-21 and is again **half-proven**: the database half read back through `php artisan` with its own negative control, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness. Two more apps joined the same class tonight for a different reason (R-624). | **READY — rank P3-LOW; owner: CC** |
| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** | | **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** |
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. | **READY — rank P2-MEDIUM; owner: CC** | | **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. **-- RULED 2026-09-23 (`09` §3 decision 16):** PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on the test bench before the catalog may move it; the engine-major gate stays until then. | **READY TO BUILD — owner: CC; `09` §6.4; the gate stays until all eleven are proven** |
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** | | **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | | **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** | | **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** |