SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has already moved, and the Restart button does it — 18.3s with a network pull when the target image is absent, 0.5s when present, against a negative control that did not even recreate the container. The boot reconciler does the same thing unattended when an app fails to come back (bootrecon.go:269 -> StartStack). AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades nothing. Docker restores the old containers and the reconciler logs 'no boot-orphaned apps (nothing to start)'. AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run, the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported'. 'Rollback' is the wrong word for this arc and is struck. Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator confirmation. Peti's box was never contacted. No production code written in any repo: felhom-controller is at 960d29b0612c before and after, tree clean, and build/vet/test are green — run at the end to prove exactly that. Register: R-438/439/440 updated with live evidence; R-441..R-445 opened (restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB left behind; update reports success over a broken app; no fleet fstrim; hub telemetry outlives the app). Capability map gains three measured rows. The mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR, not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision. Five claims in the brief are named as wrong, including two of my own method.
This commit is contained in:
@@ -18,7 +18,7 @@ not an evening's work.**
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **One thing is waiting on you — item 4 (send two e-mails).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision, new today).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
the real machines:
|
||||
- the background job that could delete a live restore's lock now waits its turn — and the check
|
||||
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
|
||||
@@ -70,7 +70,48 @@ nothing.*
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
|
||||
7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
7. **Where should the safety go before an app updates? This is the one decision from today's
|
||||
measurement, and it is a design choice, not a bug report.**
|
||||
|
||||
**What I measured.** The box downloads new app versions by itself every 15 minutes and writes them
|
||||
into the customer's files, whether the app is running or not. Nothing tells the customer. Then the
|
||||
**Restart** button — not just Update — installs that new version. I watched it download a version
|
||||
that was not on the machine and swap the app onto it in 18 seconds. **And the box does it on its
|
||||
own** when an app fails to come back after a crash: nobody pressed anything.
|
||||
|
||||
**One fear is smaller than we thought, and you should have that too.** A plain power cut does
|
||||
**not** upgrade anything. The apps come back on their old version. It only happens when an app
|
||||
fails to return.
|
||||
|
||||
**One fear is bigger.** I tested whether we can undo an app update. **We cannot.** Once an app has
|
||||
moved its data to the new version, putting the old version back gives an app that will not start at
|
||||
all. So **"rollback" is the wrong word** and I have struck it. The only way back is to restore the
|
||||
customer's data from a copy taken *before* the update — and today no update takes one.
|
||||
|
||||
**The decision, in one sentence: should the safety copy sit under the Update button only, or under
|
||||
everything that can install a new version?**
|
||||
|
||||
- **Under the button only.** Cheap and quick. Covers the case a customer causes. **Leaves the
|
||||
unattended path uncovered** — the one where an app that failed to come back is upgraded with
|
||||
nobody watching.
|
||||
- **Under everything.** Covers all of it. Costs more, and it has a hard limit I measured: for a
|
||||
big app a copy is roughly 30 minutes and about twice the app's size, against a standard box that
|
||||
ships with 20 GB. **A copy of a large app does not fit.** So this option cannot be built without
|
||||
also answering where the copy lives.
|
||||
|
||||
**My pick: under everything — but decide the "where does it live" question first**, because the
|
||||
answer decides whether the rest is even buildable.
|
||||
|
||||
**If you do nothing:** nothing breaks today, and no customer is at risk this week — the fleet is
|
||||
young and its running versions match the catalog. But the exposure is real and dated: **Peti's box
|
||||
has a one-major upgrade of `rallly` queued behind its next boot**, from a catalog change made three
|
||||
days after it went quiet. **Who else is blocked:** nobody can spec the safe-update work until this
|
||||
is answered, because the two options produce different products.
|
||||
|
||||
Full measurement, with the controls and the quoted output:
|
||||
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
|
||||
|
||||
8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user