SPIKE: an upgrade test that runs again — and a real defect in our own bookstack template
gates / gates (push) Successful in 19s

R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and
the whole update arc was designed against that single data point.

C3 first: the negative control, whose TO image exits immediately, came back
failed. That is what makes the greens mean anything, and it cost 556s because a
negative is only honest if it waits out the full settle window.

Seven edges, three apps. All five real catalog upgrades kept the customer's data.

The finding that changes an assumption the arc was carrying: whether an upgrade
can be UNDONE is a property of the individual APP, not of upgrades. Docmost
refuses - 'corrupted migrations: previously executed migration
20260213T085259-notifications is missing' - and privatebin does not. That
reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the
struck word 'rollback' now rests on two measurements instead of one.

The finding nobody was looking for, R-459: our own bookstack template moves
MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that
the datadir upgrade it requires is being skipped, and serves anyway. The cause is
assigned rather than guessed - the app half alone produces no upgrade line, both
edges that move the engine produce it - which is exactly what decomposing E3 into
E3a and E3b was for. It also explains why E3's abort looked like it worked: the
datadir was never converted. Whether that ever breaks is NOT established, and the
row says so.

Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461
(target-selection.md names a venue that does not exist and fences a VM that is
gone), R-462 (the widening, costed with this run's real numbers - and the cost is
dominated by fixtures, which do not amortise).

Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50
percent before and after. The capability map was deliberately NOT edited: this
measured apps, not the product.
This commit is contained in:
2026-09-06 11:48:57 +02:00
parent 417df06f35
commit a1a6c73fe1
55 changed files with 2234 additions and 54 deletions
+4 -1
View File
@@ -683,7 +683,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** |
| **R-449** | **[P2-MEDIUM] UPDATE ARC SLICE 5 — an upgrade test that proves a real one-major upgrade end to end, INCLUDING the abort path.** Spike §7 performed the pieces by hand on Nextcloud (31.0.14 → 32.0.9 ran the migration; 31.0.14 → 34.0.1 was refused by the app and reported as SUCCESS by the product; the downgrade attempt was refused with the data intact). **None of it is a test that runs again.** A one-off measurement that nothing repeats decays into a claim — this project's most-repeated defect class. The test must assert the CONSEQUENCE (does the app serve after the upgrade? does the abort put it back?), not the mechanism. Venue: a Tier-0 box; `runbooks/target-selection.md` names which. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: CC** |
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** |
@@ -692,6 +691,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** |
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** |
| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** |
| **R-459** | **[P2-MEDIUM] Our own bookstack template moves MariaDB across a MAJOR while the image tells us, in its own words, that the datadir upgrade it needs is being SKIPPED.** MEASURED 2026-09-06 in a throwaway guest, on the catalog's own transition `0b73e5e` (mariadb `11.6` → `12.3`): the moment 12.3 starts on the 11.6 datadir it logs **`[Note] [Entrypoint]: MariaDB upgrade (mariadb-upgrade or creating healthcheck users) required, but skipped due to $MARIADB_AUTO_UPGRADE setting`** — and then serves normally. `templates/bookstack/docker-compose.yml` sets **no `MARIADB_*` environment at all**, so `MARIADB_AUTO_UPGRADE` is unset and the entrypoint declines to run `mariadb-upgrade`. **THE CAUSE IS ASSIGNED, NOT GUESSED:** the edge was decomposed, and the app half alone (`E3a`, engine held at 11.6) produces **no** upgrade line while both edges that move the engine (`E3`, `E3b`) produce it. A bundled edge could never have said which half. **AND IT EXPLAINS A RESULT THAT LOOKED LIKE GOOD NEWS:** the abort of E3 "worked" — 11.6 came back and served the data, logging `MariaDB upgrade not required` — **because the datadir was never converted.** The reversibility is a side-effect of an upgrade that did not fully happen. **WHAT IS NOT ESTABLISHED, and this row must not be read past: whether running 12.3 on an unconverted 11.6 datadir ever actually breaks.** It did not break here. MariaDB calls the upgrade required; we measured that it is skipped and did **not** measure a consequence. **What would settle it, cheaply:** run the E3 edge, then restart the stack several times and exercise the app, watching for the entrypoint's own complaint to become an error — one guest, no new mechanism. **The fix, if the consequence is real, is one line of template env**, but setting `MARIADB_AUTO_UPGRADE` fleet-wide is a change to how every MariaDB app upgrades and is the operator's call, not a quiet edit. Owner: **CC measures, VIKTOR rules on the fleet-wide env.** `audits/SPIKE-upgrade-test-2026-09-06.md` §4 | **READY — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** |
| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 | **READY — rank P3-LOW; owner: CC** |
| **R-461** | **[P3-LOW] `runbooks/target-selection.md` names a venue that does not exist and fences a fixture that is gone.** MEASURED 2026-09-06 on demo-hp while siting the R-449 guest. (a) The runbook says to put VM disks on a dir storage at **`/mnt/nvme-1tb`, at its root**. **There is no `/mnt/nvme-1tb`** — the 1 TB NVMe is mounted at **`/mnt/hdd_1`**, which is the enrolled user-data drive and the `felhom-backup` target, i.e. the same disk under a different path. The instruction's REASON is still exactly right (`local-lvm` is an over-subscribed thin pool backing the live guest 9201, and this run kept off it — `local-lvm` read **30.50 % before and after**), so the fence held; only its address is stale. (b) The runbook forbids destroying **`drill-r50` (VM 300)**, "the only drift fixture (R-93)". **`qm list` returns nothing on demo-hp** — there are no VMs at all. **The fence currently protects nothing, and R-93's premise that a drift fixture exists is false.** **Why it is a row and not a quiet edit:** a runbook that names a path nobody can find is one a session works around, and working around a safety instruction is how the instruction stops being followed. Both halves need checking against the box before the text is changed — (b) in particular may mean R-93 should be closed or reopened as "the drift fixture is gone", which is a different fact from "do not destroy it". Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §7 | **READY — rank P3-LOW; owner: CC** |
| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 | **READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.