Files
felhom.eu/documentation/audits/SPIKE-upgrade-test-2026-09-06.md
T
admin a1a6c73fe1
gates / gates (push) Successful in 19s
SPIKE: an upgrade test that runs again — and a real defect in our own bookstack template
R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and
the whole update arc was designed against that single data point.

C3 first: the negative control, whose TO image exits immediately, came back
failed. That is what makes the greens mean anything, and it cost 556s because a
negative is only honest if it waits out the full settle window.

Seven edges, three apps. All five real catalog upgrades kept the customer's data.

The finding that changes an assumption the arc was carrying: whether an upgrade
can be UNDONE is a property of the individual APP, not of upgrades. Docmost
refuses - 'corrupted migrations: previously executed migration
20260213T085259-notifications is missing' - and privatebin does not. That
reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the
struck word 'rollback' now rests on two measurements instead of one.

The finding nobody was looking for, R-459: our own bookstack template moves
MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that
the datadir upgrade it requires is being skipped, and serves anyway. The cause is
assigned rather than guessed - the app half alone produces no upgrade line, both
edges that move the engine produce it - which is exactly what decomposing E3 into
E3a and E3b was for. It also explains why E3's abort looked like it worked: the
datadir was never converted. Whether that ever breaks is NOT established, and the
row says so.

Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461
(target-selection.md names a venue that does not exist and fences a VM that is
gone), R-462 (the widening, costed with this run's real numbers - and the cost is
dominated by fixtures, which do not amortise).

Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50
percent before and after. The capability map was deliberately NOT edited: this
measured apps, not the product.
2026-09-06 11:48:57 +02:00

14 KiB
Raw Blame History

SPIKE — does a real app upgrade keep the customer's data, and can it be undone? (2026-09-06)

THE ANSWER, IN ONE SENTENCE: on all five real catalog upgrades measured, the customer's data survived — and whether the upgrade can be UNDONE is not a property of upgrades at all, it is a property of the individual app, which is why it has to be measured per app rather than reasoned about.

AND THE FINDING NOBODY WAS LOOKING FOR, which is a defect in our own catalog: this repo's bookstack template moves MariaDB 11.6 → 12.3, and MariaDB 12.3 comes up, says in its own words that a datadir upgrade is required, skips it, and serves anyway — because the template sets no MARIADB_AUTO_UPGRADE. The app works. The engine is running on a datadir it itself calls un-upgraded. R-459.

Class: spike. No controller code, no product change, no customer's box. Everything ran inside a throwaway LXC destroyed at the end (§7). R-449 asked for "an upgrade test that runs again"; the harness is app-catalog-felhom.eu/scripts/upgrade-test.py and it is now a maintained script, not an audit artifact.


1. C3 first — the harness can say no

Everything below is conditional on this, so it is reported first. C3 is an edge whose TO image is alpine:3.20 — a real image that pulls cleanly and exits immediately.

verdict: failed    seed_read_before: true    seed_read_after: false    healthy_after: false
TO settled=False in 421.1s :: {"privatebin": {"status": "restarting", "health": "unhealthy"}}

C3 came back RED. A harness that cannot fail a known-broken upgrade proves nothing with its greens. It also cost the most wall-clock of any edge (556 s), because a negative is only honest if it waits out the full settle window.

C1 (the seed reads back BEFORE the upgrade) passed on every edge, including C3 — a fixture that cannot prove itself first proves nothing after.

2. The verdict table

edge app from → to verdict data after abort TO settle total
C3 privatebin 2.0.5 → alpine:3.20 failed no starts-and-serves 421.1 s 556.0 s
C2 privatebin 2.0.5 → 2.0.5 (no-op) proven yes starts-and-serves 0.1 s 6.4 s
E1 privatebin 1.7.5 → 2.0.5 (major) proven yes starts-and-serves 5.2 s 21.5 s
E2 docmost 0.25.3 → 0.95.0 proven yes REFUSES 10.7 s 305.1 s
E3 bookstack app 25.02.2→26.05.2 + mariadb 11.6→12.3 proven yes starts-and-serves¹ 15.7 s 77.5 s
E3a bookstack app only, engine held at 11.6 proven yes starts-and-serves 15.7 s 71.8 s
E3b bookstack engine only 11.6→12.3, app held proven yes starts-and-serves¹ 0.2 s 48.8 s

¹ and §4 is why that is not the good news it looks like.

3. The abort is a property of the APP, not of upgrades

E2, docmost — the abort REFUSES. The old image will not start on the migrated database. Verbatim, from the app's own log:

{"level":"error","context":"DatabaseMigrationService",
 "msg":"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"}
{"level":"error","context":"DatabaseMigrationService","msg":"Failed to run database migration. Exiting program."}

The container then crash-loops: {"docmost": {"status": "restarting", "health": "unhealthy", "exit": 1}} while docmost-postgres and docmost-redis stay healthy beside it.

This independently reproduces the Nextcloud finding on a second app — and by a DIFFERENT mechanism. Nextcloud refused on an explicit version comparison; docmost refuses because its migration ledger contains entries the older code does not know about. Two apps, two unrelated causes, same outcome. 09-update-architecture.md §4's ruling that "rollback" is the wrong word is now supported by two independent measurements rather than one.

E1, privatebin — the abort WORKS, and the reason is structural: privatebin is file-backed with no database and no schema, so a major version change moves no data. Its migration-line capture is empty, which is a true negative rather than a missed observation.

So the vocabulary needs one more distinction than the arc currently has. There is no single answer to "can an update be undone". There are apps where it can and apps where it cannot, and the only way to know which is to measure that app. This is the finding R-449 existed to produce.

4. The finding nobody was looking for — R-459

E3 is the catalog's own bookstack transition, and it moves the database engine across a major at the same time as the app. It came back proven, and its abort came back starts-and-serves. Both are true and both are misleading, and the decomposition is what showed it.

Verbatim, at the moment MariaDB 12.3 first started on the 11.6 datadir:

[Note] [Entrypoint]: MariaDB upgrade (mariadb-upgrade or creating healthcheck users) required,
                     but skipped due to $MARIADB_AUTO_UPGRADE setting

Confirmed by reading which engine actually served in each edge — not inferred:

edge engine that served the TO step upgrade line
E3 (app and engine) 12.3.3-MariaDB-ubu2404 required, but skipped
E3a (app only) 11.6.2-MariaDB-ubu2404 none
E3b (engine only) 12.3.3-MariaDB-ubu2404 required, but skipped

E3a is the control that assigns the cause. The app half alone produces no engine upgrade line at all; both edges that move the engine produce it. This is exactly what the decomposition was for, and a bundled edge could never have said it.

templates/bookstack/docker-compose.yml sets no MARIADB_* environment at all, so MARIADB_AUTO_UPGRADE is unset and the image's entrypoint declines to run mariadb-upgrade.

And this is why the abort "worked": the datadir was never converted, so MariaDB 11.6 could still read it — on the way back it says MariaDB upgrade not required. The reversibility of E3 is a side-effect of an upgrade that did not fully happen.

What is NOT established, and the row says so: whether running 12.3 on an unconverted 11.6 datadir ever actually breaks. It did not break here. MariaDB itself calls the upgrade required; we measured that it is skipped, and we did not measure a consequence. Naming the gap is the finding; the consequence is a separate measurement.

5. What each edge cost — the numbers the widening must be costed with

Measured, not estimated, which is the whole point of §15.5 of the task.

wall-clock, successful edges 6.4 s – 305.1 s, median 71.8 s
wall-clock, the failing edge 556 s — a negative costs ~8× a positive, because it must wait out the full settle window
total for 7 edges ~18 minutes of harness time, plus ~35 minutes of build-out and two fixture iterations
disk, 3 apps / 11 images 5.07 GB of images, 6.0 GB guest total
naive extrapolation to 53 apps ~90 GB of images and, at the median, ~1 h of harness time for one edge each — but see the caveat below, which is the real cost

The real cost is not the machine, it is the fixture. Two of the three apps needed a bespoke non-browser seed route, and one of those (bookstack) needed two attempts and a discarded approach (§6). Harness time scales with apps; fixture time scales with apps too, and it does not amortise. Costing the widening from the 71.8 s median alone would understate it by an order of magnitude.

6. Fixtures — what worked, what did not, and what cannot be done at all

Every seed went in through the app's own interface. Nothing was written into a volume or a database by hand (R-156).

app seed route readback route notes
privatebin its JSON paste API (HTTP POST) HTTP GET of the same paste file-backed, no database — this single seed IS the file half; there is no database half to seed
docmost POST /api/auth/setup POST /api/auth/login as the seeded user the readback deliberately uses the app's own front door, and is version-stable across the 0.25→0.95 API churn
bookstack php artisan bookstack:create-admin php artisan bookstack:reset-mfa --email=… database half only — see below

Two dead ends, recorded because they cost time and will cost the next person the same:

  1. PrivateBin refuses a malformed paste envelope with {"status":1,"message":"Invalid data."}. ct, the IV and the salt must all be real base64. The first fixture attempt used placeholder strings and was rejected — which read like an app failure and was a harness bug.
  2. BookStack cannot be verified over HTTP without TLS. Its APP_URL comes from the template as https://${SUBDOMAIN}.${DOMAIN}, so it marks its session and XSRF cookies secure; curl over plain http stores neither, sends neither, and every login POST returns 419 Page Expired — which looks exactly like a wrong password. The container serves no TLS. There is no http route to a logged-in session without changing the app's own configuration, which would be measuring a different app.

So bookstack's readback is artisan on both sides, and two things were done to keep that honest: the readback command is a different command from the seed and must FIND the record the other one created; and the fixture runs its own negative control on every call — it also looks up an email that cannot exist and requires the answer A user where email=… could not be found. A readback that had broken into always saying "found" therefore fails instead of passing everything.

One thing genuinely cannot be done headlessly, and it is a gap, not a defect: the FILE half of bookstack. Seeding an uploaded image or attachment needs an API token that BookStack only mints through a browser. So for bookstack this run proves the database survived and says nothing about uploaded files. R-460.

A note on the standing "gate on the exit code" rule, because this run had to break it once. bookstack:reset-mfa asks for interactive confirmation, finds no TTY, and exits 1 in both cases — for a user it found and one it did not. The exit code carries no information there, so the discriminator is the output, required positive AND with the not-found sentence required absent. Stated rather than done quietly.

7. Venue and teardown — all three layers

Throwaway LXC 9401 upgrade-spike on demo-hp, Debian 13, 4 cores, 6 GB, created for this run.

Two things the runbook says about this venue are now out of date, and were worked around rather than followed blindly:

  • /mnt/nvme-1tb does not exist. The 1 TB NVMe is mounted at /mnt/hdd_1 and is the enrolled user-data drive. A dir storage scratch-upg was created at its root (the runbook's own rule — a subdirectory fails the agent's exactMount check) and removed at teardown. The guest's disk was deliberately kept off local-lvm, whose over-subscription is the runbook's real warning.
  • drill-r50 (VM 300) no longer exists. qm list returns nothing on this host. The runbook's "do not destroy it" fence currently protects nothing. R-461.
layer before after
1 — the machine guest 9401 created destroyed; pct list shows only 9201. scratch-upg storage removed. Downloaded LXC template deleted.
2 — the host local-lvm 30.50 %, local 19 595 164 KiB local-lvm 30.50 % — unchanged, never touched; local 19 605 084 KiB (+9.7 MB). pct fstrim 9401 before destroy: 52.8 GiB trimmed
3 — the hub 2 enrolled hosts 2 enrolled hosts, 0 customers created. Checked, not assumed: /hosts lists exactly demo-felhom-8363b5 and demo-hp-bb76ea. This run created no customer, no appliance and no host record.

No felhom-controller was in the path at any point — raw docker compose throughout. The property under test belongs to the app and its images; putting the controller in the path would have confounded the two, which is the choice the persistence sweep made and stated.

8. What was NOT measured

  • 50 of 53 apps. Three is the sample that decides whether the idea is sound, not a fleet survey.
  • Whether an unconverted MariaDB datadir ever breaks (§4). Measured that the upgrade is skipped; did not measure a consequence.
  • The FILE half of bookstack and docmost. privatebin's seed is a file; the other two prove the database only.
  • Any edge with realistic data VOLUME. Every seed here is one record. An upgrade that migrates 10 GB may behave differently, and the durations above are therefore floors, not estimates.
  • Anything through felhom-controller, deliberately (§7).

9. Observations — noticed, documented, not acted on

  1. docker compose logs only shows the containers that currently exist, so the abort — which replaces them — erases the TO step's output from any capture taken afterwards. The single most important line of this run survived only because it had already been extracted. The harness now writes to-full.log at the TO step. Same class as R-320, one layer down.
  2. A negative edge costs ~8× a positive because it must wait out the settle window. If the widening runs many edges, that asymmetry dominates the schedule, not the average.
  3. bookstack:reset-mfa's help output disagrees with its argument parsing in 25.02.2 — the usage line says [options] and a positional argument is rejected with "No arguments expected", which is correct but reads as a missing feature. Upstream's problem, recorded because it cost a minute.