Files
felhom.eu/REPORT.md
T
admin a1a6c73fe1
gates / gates (push) Successful in 19s
SPIKE: an upgrade test that runs again — and a real defect in our own bookstack template
R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and
the whole update arc was designed against that single data point.

C3 first: the negative control, whose TO image exits immediately, came back
failed. That is what makes the greens mean anything, and it cost 556s because a
negative is only honest if it waits out the full settle window.

Seven edges, three apps. All five real catalog upgrades kept the customer's data.

The finding that changes an assumption the arc was carrying: whether an upgrade
can be UNDONE is a property of the individual APP, not of upgrades. Docmost
refuses - 'corrupted migrations: previously executed migration
20260213T085259-notifications is missing' - and privatebin does not. That
reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the
struck word 'rollback' now rests on two measurements instead of one.

The finding nobody was looking for, R-459: our own bookstack template moves
MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that
the datadir upgrade it requires is being skipped, and serves anyway. The cause is
assigned rather than guessed - the app half alone produces no upgrade line, both
edges that move the engine produce it - which is exactly what decomposing E3 into
E3a and E3b was for. It also explains why E3's abort looked like it worked: the
datadir was never converted. Whether that ever breaks is NOT established, and the
row says so.

Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461
(target-selection.md names a venue that does not exist and fences a VM that is
gone), R-462 (the widening, costed with this run's real numbers - and the cost is
dominated by fixtures, which do not amortise).

Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50
percent before and after. The capability map was deliberately NOT edited: this
measured apps, not the product.
2026-09-06 11:48:57 +02:00

10 KiB
Raw Blame History

REPORT — SPIKE: an upgrade test that runs again (2026-09-06)

Overwritten each session. Nothing durable lives only here.

C3 FIRST, because everything else is conditional on it. The negative control — an edge whose TO image is alpine:3.20, which pulls cleanly and exits immediately — came back failed (healthy_after: false, seed_read_after: false, seed_read_before: true). The harness can say no, so its greens mean something. It also cost the most wall clock of any edge, 556 s, because a negative is only honest if it waits out the full settle window.

1. Confirmed baselines — none had moved

repo task's baseline found
app-catalog-felhom.eu 7b9b9b34a5ee 7b9b9b34a5ee
felhom.eu 417df06f3529 417df06f3529
felhom-controller bab82c4 (v0.235.0) not touched

Highest R- id: 458, confirmed. Minted R-459 … R-462. No version bump, no release, no golden.

2. The verdict record — all seven edges

Full JSON per edge in documentation/audits/upgrade-spike-2026-09-06/evidence/<EDGE>/verdict.json.

edge app from → to verdict data after abort TO settle total
C3 privatebin 2.0.5 → alpine:3.20 failed no starts-and-serves 421.1 s 556.0 s
C2 privatebin 2.0.5 → 2.0.5 proven yes starts-and-serves 0.1 s 6.4 s
E1 privatebin 1.7.5 → 2.0.5 proven yes starts-and-serves 5.2 s 21.5 s
E2 docmost 0.25.3 → 0.95.0 proven yes REFUSES 10.7 s 305.1 s
E3 bookstack app 25.02.2→26.05.2 + mariadb 11.6→12.3 proven yes starts-and-serves¹ 15.7 s 77.5 s
E3a bookstack app only, engine held proven yes starts-and-serves 15.7 s 71.8 s
E3b bookstack engine only, app held proven yes starts-and-serves¹ 0.2 s 48.8 s

¹ and §4 is why that is not the good news it looks like.

C1 passed on every edge, including C3. A fixture that cannot prove itself first proves nothing after.

3. Quoted verbatim

E2's refusal — the abort of docmost:

{"level":"error","context":"DatabaseMigrationService",
 "msg":"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"}
{"level":"error","context":"DatabaseMigrationService","msg":"Failed to run database migration. Exiting program."}

E2's migration, at the TO step (six such lines, one shown):

{"level":"info","context":"DatabaseMigrationService","msg":"Migration \"20260213T085259-notifications\" executed successfully"}

E3/E3b, at the moment MariaDB 12.3 first started on the 11.6 datadir:

[Note] [Entrypoint]: MariaDB upgrade (mariadb-upgrade or creating healthcheck users) required,
                     but skipped due to $MARIADB_AUTO_UPGRADE setting

E3a, the app half alone: no such line at all.

4. The two findings

(a) Whether an upgrade can be UNDONE is a property of the app, not of upgrades. docmost refuses; privatebin does not. This independently reproduces the Nextcloud finding on a second app by a different mechanism — Nextcloud refused on a version comparison, docmost on its migration ledger. So 09-update-architecture.md §4's ruling now rests on two measurements, not one. And the arc was carrying an assumption that there is one answer for all 53 apps. There is not.

(b) R-459 — a real defect in our own catalog, found by accident. The bookstack template moves MariaDB across a major and sets no MARIADB_* env at all, so the image skips the datadir upgrade it says it requires. The cause is assigned, not guessed: E3a (app half, engine held) produces no upgrade line; E3 and E3b (both move the engine) produce it. A bundled edge could never have said which half — which is exactly what the decomposition existed for. It also explains why E3's abort "worked": the datadir was never converted, so 11.6 could still read it. Whether that ever breaks is not established and the row says so.

5. What it cost — measured, for costing the widening

successful edge 6.4 – 305.1 s, median 71.8 s
failing edge 556 s — ~8× a positive
7 edges total ~18 min harness time + ~35 min build-out and two fixture iterations
disk, 3 apps / 11 images 5.07 GB images, 6.0 GB guest
naive extrapolation to 53 ~90 GB, ~1 h harness time

The extrapolation understates it by an order of magnitude, and that is the finding. Two of three apps needed a bespoke seed route, one needed two attempts and a discarded approach, and one can only ever be half-proven. Fixture time scales with apps and does not amortise. → R-462, which asks the operator for scope rather than proposing one.

6. Apps with no non-browser seed route

None was fully blocked; one is half-blocked. privatebin and docmost have clean HTTP APIs. BookStack has neither an API token nor a usable HTTP login headlessly — its template's https APP_URL makes the session cookies secure, so curl over plain http gets 419 Page Expired on every login, which looks exactly like a wrong password. Its database half is provable through php artisan (seed and readback are different commands, and the readback runs its own negative control on every call). Its FILE half cannot be seeded headlessly at all → R-460. Nothing was planted by hand anywhere.

7. Evidence, copied off after EACH edge

felhom.eu/documentation/audits/upgrade-spike-2026-09-06/evidence/ — 48 files, 340 KB, pulled to DooPlex after each edge and again at the end, before any teardown. Scanned for secrets before commit: no token, no password, no generated key. The only spike-* strings are seeded account names from a guest that no longer exists.

8. Teardown — all three layers

layer result
1 — machine guest 9401 destroyed; pct list shows only 9201. scratch-upg storage removed. Downloaded LXC template deleted.
2 — host local-lvm 30.50 % before and after — never touched, which was the whole point of siting the guest off it. local 19 595 164 → 19 605 084 KiB (+9.7 MB). pct fstrim 9401 before destroy: 52.8 GiB trimmed.
3 — hub Checked, not asserted: /hosts lists exactly demo-felhom-8363b5 and demo-hp-bb76ea; 0 customers created. This run created no customer, no appliance and no host record.

No felhom-controller was in the path at any point. Raw docker compose throughout.

9. Register — 203 open before, 206 after; closed 167 → 168

row disposition
R-449 CLOSED — harness built, run, and proven by a red negative control
R-459 OPENED, P2-MEDIUM — the skipped MariaDB datadir upgrade in our own bookstack template. CC measures the consequence, VIKTOR rules on a fleet-wide env change
R-460 OPENED, P3-LOW, CC — bookstack's file half is unprovable headlessly
R-461 OPENED, P3-LOW, CC — target-selection.md names a venue that does not exist and fences a fixture that is gone
R-462 OPENED, P2-MEDIUM — the widening, costed with real numbers. VIKTOR rules on scope

10. The capability map was NOT edited, and that is deliberate

This run measured apps, not the product. Nothing the platform can do changed: no controller code, no version, no behaviour. Editing the map reflexively would record a capability the product did not gain. Said here rather than left silent.

11. Claims in the task that turned out to be wrong, named

  1. §11's venue is stale, in two ways. /mnt/nvme-1tb does not exist — the 1 TB NVMe is at /mnt/hdd_1, the enrolled user-data drive, i.e. the same disk under a different path; the scratch storage went at its root, honouring the rule's reason. And drill-r50 (VM 300) is gone — qm list returns nothing on demo-hp, so that fence protects nothing today. → R-461.
  2. §6's edges were all real and all resolvable. Every one of the eleven images was verified against its registry before use; none had to be substituted. The task asked to say so if any had.
  3. §8 expected bookstack to be "the one most likely to be hard" — correct, and for a reason the task did not name. The blocker was not the missing API token; it was that the template's https APP_URL makes the session cookies secure, so no http login can ever work. Two independent blockers, and only one was anticipated.
  4. §10.2's "gate on each command's own exit code" had to be broken once, deliberately and in the open. bookstack:reset-mfa exits 1 for a user it found and 1 for one it did not, so the exit code carries no information; the discriminator is the output, required positive with the not-found sentence required absent. Stated in the fixture's docstring rather than done quietly.
  5. §9's verdict shape needed no change and is now recorded in 09-update-architecture.md §6 as the contract Slice 6 carries.

12. Observations — noticed, documented, NOT acted on

  1. docker compose logs only shows containers that currently exist, so the abort erases the TO step's output from any later capture — the single most important line of this run survived only because it had already been extracted. Fixed mid-run (to-full.log is now written at the TO step) and E3 re-run to get clean evidence. FILED: R-449's closure records it; the harness change is committed. NOT-A-FINDING as a separate row: it is a harness bug that was found and fixed inside the same session, with the fix committed and the affected edge re-measured — there is no residue for a row to track.
  2. A negative edge costs ~8× a positive. NOT-A-FINDING: it is a measured cost recorded in R-462, which is where the widening will be scheduled from; a second row would duplicate it.
  3. bookstack:reset-mfa's help says [options] but a positional argument is rejected with "No arguments expected" in 25.02.2 — correct behaviour that reads as a missing feature. NOT-A-FINDING: it is upstream's interface, not ours, and it costs us nothing now that the fixture documents it.