R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and the whole update arc was designed against that single data point. C3 first: the negative control, whose TO image exits immediately, came back failed. That is what makes the greens mean anything, and it cost 556s because a negative is only honest if it waits out the full settle window. Seven edges, three apps. All five real catalog upgrades kept the customer's data. The finding that changes an assumption the arc was carrying: whether an upgrade can be UNDONE is a property of the individual APP, not of upgrades. Docmost refuses - 'corrupted migrations: previously executed migration 20260213T085259-notifications is missing' - and privatebin does not. That reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the struck word 'rollback' now rests on two measurements instead of one. The finding nobody was looking for, R-459: our own bookstack template moves MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that the datadir upgrade it requires is being skipped, and serves anyway. The cause is assigned rather than guessed - the app half alone produces no upgrade line, both edges that move the engine produce it - which is exactly what decomposing E3 into E3a and E3b was for. It also explains why E3's abort looked like it worked: the datadir was never converted. Whether that ever breaks is NOT established, and the row says so. Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461 (target-selection.md names a venue that does not exist and fences a VM that is gone), R-462 (the widening, costed with this run's real numbers - and the cost is dominated by fixtures, which do not amortise). Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50 percent before and after. The capability map was deliberately NOT edited: this measured apps, not the product.
10 KiB
REPORT — SPIKE: an upgrade test that runs again (2026-09-06)
Overwritten each session. Nothing durable lives only here.
C3 FIRST, because everything else is conditional on it. The negative control — an edge whose TO image is
alpine:3.20, which pulls cleanly and exits immediately — came backfailed(healthy_after: false,seed_read_after: false,seed_read_before: true). The harness can say no, so its greens mean something. It also cost the most wall clock of any edge, 556 s, because a negative is only honest if it waits out the full settle window.
1. Confirmed baselines — none had moved
| repo | task's baseline | found |
|---|---|---|
| app-catalog-felhom.eu | 7b9b9b34a5ee |
7b9b9b34a5ee |
| felhom.eu | 417df06f3529 |
417df06f3529 |
| felhom-controller | bab82c4 (v0.235.0) |
not touched |
Highest R- id: 458, confirmed. Minted R-459 … R-462. No version bump, no release, no golden.
2. The verdict record — all seven edges
Full JSON per edge in documentation/audits/upgrade-spike-2026-09-06/evidence/<EDGE>/verdict.json.
| edge | app | from → to | verdict | data after | abort | TO settle | total |
|---|---|---|---|---|---|---|---|
| C3 | privatebin | 2.0.5 → alpine:3.20 |
failed | no | starts-and-serves | 421.1 s | 556.0 s |
| C2 | privatebin | 2.0.5 → 2.0.5 |
proven | yes | starts-and-serves | 0.1 s | 6.4 s |
| E1 | privatebin | 1.7.5 → 2.0.5 |
proven | yes | starts-and-serves | 5.2 s | 21.5 s |
| E2 | docmost | 0.25.3 → 0.95.0 |
proven | yes | REFUSES | 10.7 s | 305.1 s |
| E3 | bookstack | app 25.02.2→26.05.2 + mariadb 11.6→12.3 |
proven | yes | starts-and-serves¹ | 15.7 s | 77.5 s |
| E3a | bookstack | app only, engine held | proven | yes | starts-and-serves | 15.7 s | 71.8 s |
| E3b | bookstack | engine only, app held | proven | yes | starts-and-serves¹ | 0.2 s | 48.8 s |
¹ and §4 is why that is not the good news it looks like.
C1 passed on every edge, including C3. A fixture that cannot prove itself first proves nothing after.
3. Quoted verbatim
E2's refusal — the abort of docmost:
{"level":"error","context":"DatabaseMigrationService",
"msg":"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"}
{"level":"error","context":"DatabaseMigrationService","msg":"Failed to run database migration. Exiting program."}
E2's migration, at the TO step (six such lines, one shown):
{"level":"info","context":"DatabaseMigrationService","msg":"Migration \"20260213T085259-notifications\" executed successfully"}
E3/E3b, at the moment MariaDB 12.3 first started on the 11.6 datadir:
[Note] [Entrypoint]: MariaDB upgrade (mariadb-upgrade or creating healthcheck users) required,
but skipped due to $MARIADB_AUTO_UPGRADE setting
E3a, the app half alone: no such line at all.
4. The two findings
(a) Whether an upgrade can be UNDONE is a property of the app, not of upgrades. docmost refuses;
privatebin does not. This independently reproduces the Nextcloud finding on a second app by a
different mechanism — Nextcloud refused on a version comparison, docmost on its migration ledger. So
09-update-architecture.md §4's ruling now rests on two measurements, not one. And the arc was
carrying an assumption that there is one answer for all 53 apps. There is not.
(b) R-459 — a real defect in our own catalog, found by accident. The bookstack template moves
MariaDB across a major and sets no MARIADB_* env at all, so the image skips the datadir upgrade
it says it requires. The cause is assigned, not guessed: E3a (app half, engine held) produces no
upgrade line; E3 and E3b (both move the engine) produce it. A bundled edge could never have said
which half — which is exactly what the decomposition existed for. It also explains why E3's abort
"worked": the datadir was never converted, so 11.6 could still read it. Whether that ever breaks is
not established and the row says so.
5. What it cost — measured, for costing the widening
| successful edge | 6.4 – 305.1 s, median 71.8 s |
| failing edge | 556 s — ~8× a positive |
| 7 edges total | ~18 min harness time + ~35 min build-out and two fixture iterations |
| disk, 3 apps / 11 images | 5.07 GB images, 6.0 GB guest |
| naive extrapolation to 53 | ~90 GB, ~1 h harness time |
The extrapolation understates it by an order of magnitude, and that is the finding. Two of three apps needed a bespoke seed route, one needed two attempts and a discarded approach, and one can only ever be half-proven. Fixture time scales with apps and does not amortise. → R-462, which asks the operator for scope rather than proposing one.
6. Apps with no non-browser seed route
None was fully blocked; one is half-blocked. privatebin and docmost have clean HTTP APIs.
BookStack has neither an API token nor a usable HTTP login headlessly — its template's https
APP_URL makes the session cookies secure, so curl over plain http gets 419 Page Expired on
every login, which looks exactly like a wrong password. Its database half is provable through
php artisan (seed and readback are different commands, and the readback runs its own negative
control on every call). Its FILE half cannot be seeded headlessly at all → R-460. Nothing was
planted by hand anywhere.
7. Evidence, copied off after EACH edge
felhom.eu/documentation/audits/upgrade-spike-2026-09-06/evidence/ — 48 files, 340 KB, pulled to
DooPlex after each edge and again at the end, before any teardown. Scanned for secrets before commit:
no token, no password, no generated key. The only spike-* strings are seeded account names from a
guest that no longer exists.
8. Teardown — all three layers
| layer | result |
|---|---|
| 1 — machine | guest 9401 destroyed; pct list shows only 9201. scratch-upg storage removed. Downloaded LXC template deleted. |
| 2 — host | local-lvm 30.50 % before and after — never touched, which was the whole point of siting the guest off it. local 19 595 164 → 19 605 084 KiB (+9.7 MB). pct fstrim 9401 before destroy: 52.8 GiB trimmed. |
| 3 — hub | Checked, not asserted: /hosts lists exactly demo-felhom-8363b5 and demo-hp-bb76ea; 0 customers created. This run created no customer, no appliance and no host record. |
No felhom-controller was in the path at any point. Raw docker compose throughout.
9. Register — 203 open before, 206 after; closed 167 → 168
| row | disposition |
|---|---|
| R-449 | CLOSED — harness built, run, and proven by a red negative control |
| R-459 | OPENED, P2-MEDIUM — the skipped MariaDB datadir upgrade in our own bookstack template. CC measures the consequence, VIKTOR rules on a fleet-wide env change |
| R-460 | OPENED, P3-LOW, CC — bookstack's file half is unprovable headlessly |
| R-461 | OPENED, P3-LOW, CC — target-selection.md names a venue that does not exist and fences a fixture that is gone |
| R-462 | OPENED, P2-MEDIUM — the widening, costed with real numbers. VIKTOR rules on scope |
10. The capability map was NOT edited, and that is deliberate
This run measured apps, not the product. Nothing the platform can do changed: no controller code, no version, no behaviour. Editing the map reflexively would record a capability the product did not gain. Said here rather than left silent.
11. Claims in the task that turned out to be wrong, named
- §11's venue is stale, in two ways.
/mnt/nvme-1tbdoes not exist — the 1 TB NVMe is at/mnt/hdd_1, the enrolled user-data drive, i.e. the same disk under a different path; the scratch storage went at its root, honouring the rule's reason. Anddrill-r50(VM 300) is gone —qm listreturns nothing on demo-hp, so that fence protects nothing today. → R-461. - §6's edges were all real and all resolvable. Every one of the eleven images was verified against its registry before use; none had to be substituted. The task asked to say so if any had.
- §8 expected bookstack to be "the one most likely to be hard" — correct, and for a reason the task
did not name. The blocker was not the missing API token; it was that the template's
httpsAPP_URLmakes the session cookiessecure, so no http login can ever work. Two independent blockers, and only one was anticipated. - §10.2's "gate on each command's own exit code" had to be broken once, deliberately and in the
open.
bookstack:reset-mfaexits 1 for a user it found and 1 for one it did not, so the exit code carries no information; the discriminator is the output, required positive with the not-found sentence required absent. Stated in the fixture's docstring rather than done quietly. - §9's verdict shape needed no change and is now recorded in
09-update-architecture.md§6 as the contract Slice 6 carries.
12. Observations — noticed, documented, NOT acted on
docker compose logsonly shows containers that currently exist, so the abort erases the TO step's output from any later capture — the single most important line of this run survived only because it had already been extracted. Fixed mid-run (to-full.logis now written at the TO step) and E3 re-run to get clean evidence. FILED: R-449's closure records it; the harness change is committed. NOT-A-FINDING as a separate row: it is a harness bug that was found and fixed inside the same session, with the fix committed and the affected edge re-measured — there is no residue for a row to track.- A negative edge costs ~8× a positive. NOT-A-FINDING: it is a measured cost recorded in R-462, which is where the widening will be scheduled from; a second row would duplicate it.
bookstack:reset-mfa's help says[options]but a positional argument is rejected with "No arguments expected" in 25.02.2 — correct behaviour that reads as a missing feature. NOT-A-FINDING: it is upstream's interface, not ours, and it costs us nothing now that the fixture documents it.