SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s

THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
This commit is contained in:
2026-09-01 21:35:32 +02:00
parent ac079b8c43
commit 56c7e373a3
43 changed files with 2026 additions and 287 deletions
+43 -2
View File
@@ -18,7 +18,7 @@ not an evening's work.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **One thing is waiting on you — item 4 (send two e-mails).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision, new today).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
the real machines:
- the background job that could delete a live restore's lock now waits its turn — and the check
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
@@ -70,7 +70,48 @@ nothing.*
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
7. **Where should the safety go before an app updates? This is the one decision from today's
measurement, and it is a design choice, not a bug report.**
**What I measured.** The box downloads new app versions by itself every 15 minutes and writes them
into the customer's files, whether the app is running or not. Nothing tells the customer. Then the
**Restart** button — not just Update — installs that new version. I watched it download a version
that was not on the machine and swap the app onto it in 18 seconds. **And the box does it on its
own** when an app fails to come back after a crash: nobody pressed anything.
**One fear is smaller than we thought, and you should have that too.** A plain power cut does
**not** upgrade anything. The apps come back on their old version. It only happens when an app
fails to return.
**One fear is bigger.** I tested whether we can undo an app update. **We cannot.** Once an app has
moved its data to the new version, putting the old version back gives an app that will not start at
all. So **"rollback" is the wrong word** and I have struck it. The only way back is to restore the
customer's data from a copy taken *before* the update — and today no update takes one.
**The decision, in one sentence: should the safety copy sit under the Update button only, or under
everything that can install a new version?**
- **Under the button only.** Cheap and quick. Covers the case a customer causes. **Leaves the
unattended path uncovered** — the one where an app that failed to come back is upgraded with
nobody watching.
- **Under everything.** Covers all of it. Costs more, and it has a hard limit I measured: for a
big app a copy is roughly 30 minutes and about twice the app's size, against a standard box that
ships with 20 GB. **A copy of a large app does not fit.** So this option cannot be built without
also answering where the copy lives.
**My pick: under everything — but decide the "where does it live" question first**, because the
answer decides whether the rest is even buildable.
**If you do nothing:** nothing breaks today, and no customer is at risk this week — the fleet is
young and its running versions match the catalog. But the exposure is real and dated: **Peti's box
has a one-major upgrade of `rallly` queued behind its next boot**, from a catalog change made three
days after it went quiet. **Who else is blocked:** nobody can spec the safe-update work until this
is answered, because the two options produce different products.
Full measurement, with the controls and the quoted output:
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.