Files
felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md
T
admin 56c7e373a3
gates / gates (push) Successful in 18s
SPIKE: what an app update actually does, and which other paths do it too
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
2026-09-01 21:35:32 +02:00

36 KiB
Raw Blame History

SPIKE — what an app update actually does, and which other paths do it too (2026-09-01)

THE ANSWER TO PHASE 1, IN ONE SENTENCE: YES — the Restart button upgrades the app, and so does the box's own repair when an app fails to come back, because every one of these paths ends in docker compose up -d, and up -d makes the container match whatever the file now says, pulling the image itself if it is missing.

AND THE SECOND ANSWER, WHICH THE OPERATOR PAGE GOT WRONG IN THE HELPFUL DIRECTION: a plain power cut does NOT do this. Docker's own restart: unless-stopped puts the existing containers back on the OLD image, so the boot reconciler finds nothing to repair and never runs up -d. The unattended upgrade happens only in the narrower case where an app does not come back by itself — and there it happens with nobody pressing anything.

AND THE THIRD, WHICH DECIDES THE VOCABULARY OF THIS WHOLE ARC: app data CANNOT be rolled back. Measured on Nextcloud: once a migration has actually run, putting the old image tag back produces a container that refuses to start — "the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported". So "rollback" is the wrong word and should be struck before anyone specs against it. The shape that is actually available is pre-update copy plus refuse-and-explain.

Scope. A spike. No production code was written, in any repo. The controller was read only: felhom-controller is at 960d29b0612c before and after, working tree clean, and go build ./... && go vet ./... && go test ./... is green (28 packages, 0 FAIL) — run at the end precisely to prove the tree was left untouched.

Method. Endpoint-level throughout, which is the standard method here because there is no browser on DooPlex: every action was the exact endpoint the UI invokes (POST /api/stacks/{name}/restart, /update, /start, /stop, /deploy, /remove), driven over HTTPS with the customer's own session cookie and CSRF token. No raw agent attach, no hand-set state — the F9 bypass was not used. The two places where state was staged by hand are named as such at the point of use (Phase 1's compose edits, and Phase 1c-ii's docker rm -f), and each says exactly what was staged and why.

Target. All destructive work ran on demo-hp (192.168.0.104, ssh hp; guest 9201 = 192.168.0.138), which runbooks/target-selection.md classes Tier 0 — disposable. DooPlex, demo-felhom, ep0 and Peti's box were not mutated. Peti's box was not touched at all.

Evidence. audits/evidence-spike-app-update-2026-09-01/ — written straight onto DooPlex as each phase produced it, before every revert, so R-96 rule 5 is satisfied continuously rather than at the end. Nothing was lost and nothing had to be reproduced.


1. Confirmed baselines — all four matched §1 of the task, none had moved

Repo main @ commit at start at end note
felhom-controller 960d29b0612c 960d29b0612c read only — unchanged
felhom.eu 1d59353df437 moved (this spike's documents) Phase 0 + Phase 7
app-catalog-felhom.eu 29edad9c5bf4 moved, then reverted to the same tree 214d448 → 30bd892 → 5d8f25f
felhom-agent 4586f0f7f6d1 4586f0f7f6d1 not touched

Live controller on demo-hp: 0.232.0, matching the newest release. Sync interval 15m (internal/config/config.go:351, the default, no override on the box).


2. Phase 1 — THE GATE. Does compose up -d upgrade an app whose compose file already moved?

App: bentopdf — one container, no database, no volume, no data of any kind, and deployed on demo-hp only. Chosen so the gate could be measured with zero data risk anywhere.

Baseline: container 9473b6e4a18f, ghcr.io/alam00000/bentopdf:v2.8.6, digest sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3, started 04:05:50Z.

# Variant Result Discriminator
1d negative control — file NOT edited, then restart no change container id, image id, digest and StartedAt all IDENTICAL. up -d saw no change and did not even recreate.
1a restart, target image ABSENT from the local store UPGRADED, and it pulled new container 93e9db74e97f, image v2.8.5, digest sha256:2d867aac…, and v2.8.5 appeared in docker images where it had not been. 18.3 s.
1b restart, target image ALREADY PRESENT UPGRADED, no pull new container 222f8178dbd0, image id unchanged from baseline 4baadf01bd68. 0.5 s.
1c boot recovery — hard guest reset, app left running NO upgrade SAME container 222f8178dbd0 restored by Docker at 17:47:45Z on v2.8.6 while the file said v2.8.5.
1c-ii boot recovery where the app did not come back UPGRADED, unattended new container aa443836a4cd on v2.8.5 at 17:55:44Z. Nobody pressed anything.

The 18.3 s vs 0.5 s split is the quantitative discriminator and it is worth more than the tags: the restart path contains no pull step in the source, yet variant 1a spent 18 seconds and ended with a new image in the local store. up -d pulled it.

Quoted live, variant 1a (phase1-1a-controller-log.txt):

17:35:45 router.go:566: [INFO] [api] restart requested for stack: bentopdf
17:35:45 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
17:36:03 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 18.3s)
17:36:07 manager.go:1403: [INFO] [stacks]   bentopdf   ghcr.io/alam00000/bentopdf:v2.8.5   running

Quoted live, variant 1c-ii — the unattended upgrade, attributed to the exact symbol (phase1-1cii-orphan.txt):

17:55:44 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 10s — sweeping
17:55:44 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]
17:55:44 manager.go:1039: [INFO] [stacks] Starting stack: bentopdf
17:55:44 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
17:55:48 manager.go:1403: [INFO] [stacks]   bentopdf   ghcr.io/alam00000/bentopdf:v2.8.5   running

1c's negative result is a POSITIVE observable, not an absent log line — the reconciler says so in its own words, which is exactly what it was built to do:

17:48:34 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)

What 1c-ii staged, stated plainly

docker rm -f bentopdf removed the container so it could not be restored by Docker's restart policy, which is the shape a power cut leaves when an app does not come back. desired_state: running was left untouched in app.yaml. The controller was then restarted, because that is what runs the boot reconciler. Nothing else was staged, and the reconciler selected the app on its own.

Why 1c could not use a hand-edited file, and how that was handled

At controller start the initial catalog sync runs immediately (sync.go:98, 17:47:49Z) while the boot reconciler waits out bootReconcileSettle = 5 s plus a settle window (17:48:34Z). The sync wins by ~45 s, so a hand-edited compose would have been overwritten before the reconciler ever saw it, and 1c would have measured nothing. 1c and 1c-ii were therefore run inside Phase 2's window, where the catalog itself carried the new tag — which is also the faithful production shape.

The design intent is already stated in the source, and it is not hidden

Manager.RestartStack (internal/stacks/manager.go:1133) carries this comment:

"Use up -d instead of bare restart so that env vars from app.yaml are injected and any template changes (new images, healthchecks) are picked up."

So the restart behaviour was chosen, deliberately, and written down. What is NOT written down anywhere is the consequence once the catalog syncer moves the file underneath a deployed app. That gap is R-438, and whether the choice should extend to the unattended paths is the operator's ruling.


3. Phase 2 — the sync overwrite, confirmed live

A real tag change (bentopdf v2.8.6 → v2.8.5) was pushed to the catalog main at 17:38:22Z (214d448) and left to travel the real 15-minute cycle. No hand-edit; the catalog was the only input.

time (UTC) file on demo-hp running container
17:38:35 (pre-sync) v2.8.6 v2.8.6
17:45:15 v2.8.6 v2.8.6
17:45:37 (post-sync) v2.8.5 v2.8.6

The sync's own line, and the file's mtime, agree to the second:

17:45:17 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
mtime: 2026-09-01 17:45:17.646017366 +0000  /opt/docker/stacks/bentopdf/docker-compose.yml
container started: 2026-09-01T17:36:35Z   (unchanged)

Syncer.copyTemplates (internal/sync/sync.go:319) has no deployed check of any kind. Its only guard is a sha256 content compare in copyIfChanged, and its only exclusion is app.yaml. The post-sync hook is stackMgr.InjectMissingFields(updated) and nothing else — the sync does not restart anything, which is why the file and the container can disagree indefinitely.

Was the customer told? No — and the page does not even show a version.

Searched with ASCII fragments and both controls, per the standing rule.

page BentoPDF (positive control) zzz-never-present (negative control) v2.8.5 v2.8.6
/apps/bentopdf (40 360 B) 4 0 0 0
/stacks (140 883 B) 1 0 0 0

A correction to my own method, made here rather than buried: the first pass used grep -o "2.8.6", where . is a regex wildcard, and it reported 2 hits. Re-run with grep -F the count is 0. The controls are what exposed it. The customer's pages carry no version string at all — not the running one, not the pending one.

No event, no notification and no email were emitted by the sync. The only Friss… strings on the pages are the catalog-sync toast and the Frissítés button label.

The revert, verified on the box

30bd892 reverted the pin; the sync landed it at 18:10:29Z ([INFO] [sync] Updated bentopdf/docker-compose.yml, Sablonok frissítve — frissítve: bentopdf) and the file read v2.8.6 again. The catalog tree is byte-identical to 29edad9c5bf4.

An unlooked-for second result: that same sync also overwrote a hand-broken compose file (Phase 3's alpine:3.20 edit) with the catalog's version. So the syncer is also the repair path for a broken app definition — within 15 minutes, automatically. The container stays broken until something runs up -d, but the definition heals itself.


4. Phase 3 — what a failed update looks like to the customer

3a — the pull fails. Loud, honest, and harmless.

Compose pointed at ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist, then POST .../update:

HTTP 500
{"ok":false,"error":"pulling images for bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ... Error manifest unknown\nError response from daemon: manifest unknown"}

The app was untouched: same container aa443836a4cd, v2.8.5, status=running, RestartCount=0. UpdateStack returns after the failed pull and never reaches up -d.

3a-ii, the landmine that was not one. The failed update leaves the file pointing at an image that does not exist, while the page shows a green healthy app with a Restart button beside the Update button. Pressing Restart in that state was measured: HTTP 500, and the app still ran. compose up -d resolves every image before it touches a container, so a bad reference fails before anything stops. This is the reassuring half of Phase 3 and it should be said as plainly as the bad half.

What the customer is shown, though, is raw English Docker output — exit code 1, stderr:, manifest unknown, Error response from daemon — on a product whose every other error string is Hungarian.

3b — the pull succeeds and the app does not. This is the one that matters.

Compose pointed at alpine:3.20 — a real image that pulls cleanly and then exits at once, so restart: unless-stopped puts it in a crash loop. This is the shape of a real upstream image whose configuration the template can no longer supply, and the catalog's own history contains exactly that case (7350cd9, "wger: revert 2.6 -> 2.3 (2.6 needs a full DB config the template cannot supply)").

POST /api/stacks/bentopdf/update
HTTP 200
{"ok":true,"message":"Stack bentopdf update completed"}

18:00:47 manager.go:1195: [INFO] [stacks] Stack bentopdf updated successfully (took 3.5s)

Reality 37 seconds later: status=restarting, RestartCount=9, alpine:3.20.

The controller's own post-start line told the truth — and it ran after the API had already answered ok:true:

18:00:50 manager.go:1403: [INFO] [stacks]   bentopdf   alpine:3.20   restarting   Restarting (0) Less than a second ago

What the customer's page said (/stacks, live HTML):

BentoPDF · pdf.enkisfelhom.hu · „URL nem elérhető – útvonal nincs publikálva" · badge „Újraindítás…" · Restarting (0) 15 seconds ago · buttons: Frissítés Újraindítás Leállítás Naplók Részletek

„Újraindítás…" reads as transient, not as failure. Nothing on the page says the update broke the app. isOperationalState (internal/web/funcmap.go:90) counts StateRestarting and StateDegraded as operational, so the full button row — including the green Frissítés — is rendered over a crash-looping app.

Is there a route back? Not from the page. Every button offered re-runs the same broken definition or stops the app. The compose file is not customer-editable. The routes that exist are (a) the catalog being corrected, which then heals the file within 15 minutes, or (b) a restore, which has its own problem — see §7.

The alarm DOES fire — 5 minutes 16 seconds later, and by a different road

This was measured to a positive observable rather than inferred from silence:

18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down

Update was at 18:00:43; the alarm at 18:05:59. The delay is crashLoopAfter = 5 * time.Minute (internal/stacks/manager.go), and it is deliberate and well-argued in its own comment — a shorter threshold would alarm on every routine deploy. The severity word is warning, which is inside the hub's exact vocabulary, so this one really does reach the customer (R-328/R-329 class avoided).

So the honest summary of 3b is not "silently broken". It is: the button lied at the moment it was pressed, the page then described a failure as a restart, and the truth arrived five minutes later through the dead-app alarm rather than through the update the customer actually performed.


5. Phase 4 — how far behind is the real fleet

Read from the running container, never from the file — the file is the thing that has already moved. Both tag and digest recorded.

demo-hp — zero drift by tag. Nine deployed apps, every running tag equal to the catalog pin.

This is not luck: the box was reinstalled 2026-08-21 and its apps were deployed ~13 h before this run. It is a young box, and it is the reason Phase 5 could not be priced on it (§6).

demo-felhom — one deployed app, opengist, running ghcr.io/thomiceli/opengist:1.13 = catalog pin.

But "up to date by tag" is not up to date. Two floating pins have already moved upstream.

Running digests compared against what the registry serves for the same tag today, with fully-pinned tags as the control:

image kind verdict
postgres:16-alpine floating SAME
redis:7-alpine floating SAME
mariadb:11.6 floating SAME
ghcr.io/thomiceli/opengist:1.13 floating SAME
mariadb:11.4 floating MOVED — running sha256:4f1d8d20…, upstream now sha256:611a2fcc…
mariadb:12.3 floating MOVED — running sha256:a02fe89c…, upstream now sha256:dd9b303a…
rommapp/romm:5.0.0 pinned (CONTROL) SAME
privatebin/pdo:2.0.5 pinned (CONTROL) SAME

Both controls held and both positives are floating tags. So on a box with no visible drift at all, pressing Frissítés today would silently swap the database engine build under romm (MariaDB 11.4) and bookstack (MariaDB 12.3) — with no catalog change, no version change on any screen, and no record anywhere of what it was before. That is R-440, no longer as an argument but as a measurement.

Peti's box — NOT measurable, and the task's premise here was wrong

The task asks Phase 4 to inspect three boxes. runbooks/target-selection.md:161 states plainly: "Currently DOWN, no enrolled host. No access route from DooPlex, and nothing here needs one." The hub agrees: the Hosts page lists exactly two enrolled hosts (demo-felhom-8363b5, demo-hp-bb76ea), and the peti-felhom customer shows status DOWN, last report 48 d ago, controller 0.115.0.

So there is no running container to read, and the hub does not record image tags at all — the report's container payload carries name, state, CPU and memory, and no image field. The row is therefore recorded as UNKNOWN, not guessed.

What IS knowable read-only, and it matters: the box last reported ~2026-07-15 running rallly + rallly-postgres. On 2026-07-18 — three days after it went quiet — the catalog moved rallly 3.11.2 → 4.11.1 [MAJOR] (e3f3a81). So a one-major upgrade is queued behind that box's next boot. On this spike's own measurements that upgrade will NOT fire on the power-on itself (1c: Docker restores the containers on the old image), but it will fire the moment any app fails to come back, or anyone presses Restart or Update. Nothing was touched on that box to establish this; it is the catalog's git history plus the hub's own record.


6. Phase 5 — what a pre-update copy would cost

The existing safety machinery is DATABASE-ONLY, and that is the headline

Manager.writeSafetyDump (internal/backup/offbox_reconstitute.go:207) discovers the app's databases and dumps each one. An app with no database gets nothing at all — len(mine) == 0 returns an empty set. There is no file-level safety copy on the R-361 path.

The database half is nearly free. Measured on demo-hp:

Every .sql dump on the box, including the real pre-restore-* undo copies from the August restore work: 48 KB – 395 KB. Largest is paperless-ngx-postgres.sql at 395 065 bytes.

The file half is the cost, and demo-hp cannot price it. Stated, not papered over.

The box holds a few MB of app content; the bulk is engine data directories:

app biggest components total
kimai kimai_db_data 166 M + kimai_var 56 M ≈ 222 MB
romm romm_db_data 166 M + romm_redis_data 25 M + roms 1.1 M ≈ 192 MB
bookstack bookstack_db_data 166 M + bookstack_config 6.8 M ≈ 173 MB

These are not customer-scale numbers — 166 MB is a fresh MariaDB's preallocated files, not anyone's data. So rather than over-claim from a young box, the arc's already-measured figures are cited:

CAMPAIGN-10-two-storage-soak-2026-07-31.md §6, 66 restores plus an M-band point 327× larger: backup ≈ 29 s + 17.4 s/GB, and a DB-backed app's recovery unit is 1.90× its data (21.1 GB of data produced a 40.2 GB unit).

Applying that to demo-hp's three heaviest: ≈ 32–33 s each, unit ≈ 330–420 MB. Trivial.

And that is exactly what makes the ceiling the real finding. For a large catalog app — Immich, Nextcloud, Plex — the same arithmetic gives, for 100 GB of data, ≈ 29 minutes and a ~190 GB copy, against a default appliance whose /mnt/sys_drive ships at 20 GB. A pre-update copy of a large app does not fit on a default box, and no amount of tuning the copy changes that. Whatever is designed here has to answer that before it answers anything else. Where the copy lives, and whether it is a copy at all or a snapshot, is a design decision and is not proposed here.


7. Phase 6 — can app data be rolled back at all? NO.

Run on operator confirmation, on a throwaway Nextcloud on demo-hp. Nothing else on any box was involved; the app was created, used and destroyed inside this phase.

Seeded on 31.0.14.1 (occ status), two independent markers — a database-backed system config value and a real file on disk:

System config value spike_marker set to string SPIKE-SEEDED-ON-31.0.14-2026-09-01
/var/www/html/data/admin/files/spike-marker.txt = SPIKE-FILE-CONTENT-31.0.14-2026-09-01
occ files:scan admin → 5 Folders, 53 Files, 2 Updated, 0 Errors

6a — the 3-major jump (31.0.14 → 34.0.1). Refused by the app, reported as success by the product.

POST /api/stacks/nextcloud/update  →  HTTP 200  {"ok":true,"message":"Stack nextcloud update completed"}

The container then crash-looped, saying:

Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported. It is only possible to upgrade one major version at a time.

This is R-40, measured. The catalog really does carry this jump today: 5e2c1ae, 2026-07-18, "nextcloud: 31.0.14-apache -> 34.0.1-apache [MAJOR]".

It was fully recoverable — putting 31.0.14 back gave a healthy app in 9 seconds with both markers intact. But that is only because nothing migrated. The refusal is Nextcloud's own safety net doing its job, and it is why 6a is the easy case.

6b — a migration that actually runs (31.0.14 → 32.0.9, one major). It ran.

Initializing nextcloud 32.0.9.2 ...
Upgrading nextcloud from 31.0.14.1 ...
Updated database
Updated <dav> to 1.34.2 · <files> to 2.4.0 · <files_sharing> to 1.24.1 · … (20+ apps)

Healthy in 31 s on 32.0.9.2, both markers still readable.

6c — THE ROLLBACK ATTEMPT. Refused. The old image will not start on migrated data.

Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker
image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the
newest image version?

Crash loop, indefinitely.

6d — positive control: the data is not destroyed, only the downgrade is blocked

Putting 32.0.9 back gave a healthy app in 46 s, occ status 32.0.9.2, and both markers read back byte-identical to what was seeded on 31. So the failure in 6c is a refusal, not corruption — which is the distinction that decides the remedy.

What this settles

Putting the old image tag back is not a rollback and must not be described as one. The only route back from a migration that has run is restoring the DATA from a copy taken before the update — which is precisely the thing the update path does not take (§2 of the task, confirmed by reading Manager.UpdateStack, and confirmed live: no dump, no copy, no hold, in any of the six updates run here).


8. The exact symbols — found by reading, as required

Every path that ends in compose up -d on a customer's stack folder, with the enclosing function.

what brings the app back symbol line
the four compose callers Manager.StartStack / StopStack / RestartStack / UpdateStack internal/stacks/manager.go:1029 / 1106 / 1133 / 1170
partial start (R-47 DB-only window) Manager.StartStackServices internal/stacks/manager.go:1082
the boot reconciler Reconciler.Run → r.stacks.StartStack(name) internal/bootrecon/bootrecon.go:223 → :269
its scheduler runBootReconcile → bootReconcileFn cmd/controller/main.go:2127 → :2172, called at :450
the app-stop guard AppStopGuard.Recover → g.starter.StartStack(name) internal/backup/appstop_marker.go:264 → :283
its hold-aware wrapper gatedAppStopStarter.StartStack cmd/controller/main.go:2008
the drive-return gate Server.restartStacks → s.stackMgr.StartStack(name) internal/web/intermediary.go:220 → :222
guest-boot change handler Server.processGuestBootChange internal/web/intermediary.go:395 → :458
quiesce restart-after-backup Loop.restartAll internal/quiesce/quiesce.go:730 → :733
off-site reconstitution Manager.ReconstituteFromOffsite internal/backup/offbox_reconstitute.go:480 → :692
restore from unit / local / tier-2 restore_unit.go:381, restore.go:76, tier2_restore.go:418, backup.go:895 —
.fab export / import appexport/export.go:286, appexport/restore.go:461 —
integrations onlyoffice_filebrowser.go:61, :96 (RestartStack) —

CORRECTION TO THE TASK: not five other paths — thirteen. Excluding the three API actions (start/restart/update) and excluding interface declarations and adapters, there are 13 call sites across 9 files that call StartStack or RestartStack, and every one of them ends in docker compose up -d against the live compose file.


9. R-439, confirmed by reading and refined

Router.actionStack (internal/api/router.go:565) checks the hold under if action == "start" || action == "restart" — update is absent and falls through to UpdateStack. The design intent is stated in RestoreHoldFor's own comment (internal/backup/offbox_reconstitute.go:323):

"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."

The severity argument in the task is correct, but for a narrower reason than it states. The task says the UI hides Frissítés unless the app is operational, and a held app is stopped. That is right — but isOperationalState (internal/web/funcmap.go:90) counts StateRestarting and StateDegraded as operational too, which was observed live in Phase 3b: the green Frissítés button was rendered over a crash-looping app. So the button is hidden specifically because a held app is StateStopped, not because broken apps hide it. The conclusion (LOW, not customer-reachable) survives; the reason needs stating precisely, and the fix needs a test pinning it or the comment stays a wish.


10. Every claim in the task that turned out to be wrong, named

  1. "a power cut on a sleeping customer's box is an unattended three-major-version upgrade" — NOT AS STATED. Measured: a hard guest reset upgraded nothing, because Docker restored the containers itself and the reconciler had no orphan. The unattended upgrade is real but needs the narrower precondition "and the app did not come back". The exposure is smaller than the operator page claims, and saying so is more useful than leaving the scarier version standing.
  2. "Five other code paths end in compose up -d" — thirteen non-API call sites across nine files (§8).
  3. "Phase 4 — demo-hp, demo-felhom and Peti's box" — Peti's box is DOWN, not enrolled, and has no access route from DooPlex by the project's own runbook. Two boxes were measured live; Peti's row is UNKNOWN, with what is knowable recorded from the hub and the catalog history (§5).
  4. "R-440 — 23 catalog image pins float" — CONFIRMED EXACTLY (79 image: lines, 53 apps, 66 distinct; 23 with no patch component). One arguable 24th is recorded rather than rounded away: ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 pins both extensions exactly and leaves the PostgreSQL patch floating.
  5. The catalog history numbers — spot-checked and all correct: 153 commits touching templates/, 53 apps, 0 .felhom.yml carrying any upgrade metadata (the only upgrade matches are prose about STARTTLS), and all four multi-major bumps confirmed on 2026-07-18 — nextcloud 5e2c1ae, grafana b789acc, calcom 147cee7, vikunja 3fa63cd.
  6. "The Update button … takes no safety copy, cannot undo itself, does not stop the app if it goes wrong" — CONFIRMED, by reading and across six live updates.
  7. A methodological correction of my own, not the task's: the first customer-page search used grep -o "2.8.6" and the unescaped . produced two false hits. Re-run with grep -F: zero. The controls caught it (§3).

11. What was NOT measured

  • Whether the drive-return gate and AppStopGuard.Recover upgrade in practice. Both were located by reading (§8) and both call StartStack, which is the same function variant 1c-ii measured upgrading an app. The mechanism is measured; these two specific entry points were not exercised live. Naming them as read-not-measured is deliberate.
  • Whether RecreateStackDefinitionFromUnit's rollback is really undone by the sync in a live restore — see §12, which states which half is measured and which is read.
  • Any behaviour on Peti's box (§5).

12. Observations — noticed, documented, not acted on

O1 — the restore path and the catalog sync disagree about the image, and the sync wins. stackAdapter.RecreateStackDefinitionFromUnit (cmd/controller/main.go:2570) writes the recovery unit's captured docker-compose.yml — carrying the OLD image pin — straight into the live stack dir (os.WriteFile(filepath.Join(stackDir, fname), data, 0644)), and restore_unit.go:317 says so: "Resolved from the UNIT's compose, because that file is about to BECOME the live one." But copyIfChanged overwrites any file whose content differs from the catalog, on the next 15-minute tick. MEASURED: a locally-modified compose (Phase 3's alpine:3.20) was overwritten by the sync at 18:10:29Z. READ, not measured: that the restore writes to that same path. So a restore's image-level rollback has a ≤15-minute half-life, and then the next up -d from any source re-applies the catalog pin. Filed as R-441.

O2 — remove_hdd_data: true is inert on this box, and the response says neither removed nor preserved. Removing the Phase 6 Nextcloud with {"remove_hdd_data":true,"remove_backups":true} returned HTTP 200 with "hdd_paths_removed":null,"hdd_paths_preserved":null, and left 128 MB at /mnt/felhom-drives/hdd_1/appdata/nextcloud. Root cause, with controls: Paths.HDDPath (internal/config/config.go:117) has no default — only an env override at :403 — and demo-hp's controller.yaml paths: block contains only data_dir, stacks_dir, system_data_path. The container has no FELHOM_PATHS_* variable at all (measured, count 0). So cfg.Paths.HDDPath == "" and ParseComposeHDDMounts (internal/stacks/delete.go:600-603) returns nil on its first line — [INFO] found 0 HDD mounts — for a compose that plainly contains - ${HDD_PATH}/appdata/nextcloud:/var/www/html/data. The second half of the same removal also no-op'd: [WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps. Filed as R-442.

O3 — an update can report success over a broken app, and the truth arrives by another road 5 minutes later. §4. Filed as R-443.

O4 — demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed. This run added ~1.05 GiB that local-lvm did not reclaim; fstrim inside the unprivileged container is refused (FITRIM ioctl failed: Operation not permitted), and pct fstrim 9201 from the host then trimmed 30.2 GiB + 57 GiB and took local-lvm from 70.91% → 26.78% — i.e. 23.8 GB below this run's own starting point. Nothing runs pct fstrim on the fleet. Filed as R-444.

O5 — the hub keeps app telemetry for an app that no longer exists anywhere. The throwaway Nextcloud now sets a fleet-wide Suggested Limit (P95×1.2) = 352 MB for Nextcloud, from ~15 minutes of a crash-looping instance, plus three MariaDB io_uring "Known Issues" attributed to demo-hp. Retained deliberately, not cleared — see §13. Filed as R-445.

O6 — NOT-A-FINDING: the sync's debug hash line cannot show what changed. logFileHashes (internal/sync/sync.go:386) reads the destination after the write, so it prints src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed) — the same hash twice, with the word "changed". Harmless (DEBUG only, and the Updated <app>/<file> INFO line above it carries the fact), but it cannot serve the purpose its name implies. Not filed; recorded here so the next person does not trust it.


13. Teardown — all three layers

Layer 1 — the machine.

  • bentopdf restored to its catalog tag v2.8.6, digest sha256:eaeea1e447205a79… — byte-identical to the run's baseline. Container 02c80375fcba, running, healthy.
  • The throwaway nextcloud stack removed via POST /api/stacks/nextcloud/remove (after the required stop): containers gone, all three named volumes gone, app.yaml gone. The stack dir holds only the catalog template (docker-compose.yml, .felhom.yml), i.e. not deployed.
  • The 128 MB the product did not remove (O2) was deleted by hand, together with the nextcloud backup dirs on both drives. find /mnt -iname "*nextcloud*" returns nothing.
  • Four images this run pulled were removed by targeted docker rmi — nextcloud:31.0.14-apache, 32.0.9-apache, 34.0.1-apache, bentopdf:v2.8.5, plus alpine:3.20. No prune of any kind was run, anywhere.
  • No guest was created. Guest 9201 was hard-reset once, by design (variant 1c), and came back with all nine apps.

Layer 2 — the host. pvesm status, before → after:

pool before after trim note
local 44.42% 44.44% unchanged in substance
local-lvm 68.97% (38 959 729 KiB) 26.78% (15 127 469 KiB) the run's ~1.05 GiB was returned, and pct fstrim 9201 reclaimed 23.8 GB more that predated this run (O4)

Guest: / 957 M used (baseline 957 M), /mnt/sys_drive 12 G / 18%, /mnt/felhom-drives/hdd_1 5.5 G — all back to their pre-run values.

Layer 3 — the hub. This run provisioned NOTHING.

  • No customer record and no appliance record was created. The customers list is unchanged at five rows: demo-felhom, demo-hp, drill-r50, peti-felhom, tester-1 — identical to the list read at the start of the run. The existing demo-hp customer was used throughout.
  • What the run DID create is events — app_start_failed (warning) for BentoPDF at 18:05:59Z, plus deploy/remove events for the throwaway Nextcloud. These are retained deliberately: the event log is an append-only record and deleting from it to tidy up a test would damage the very surface this project relies on for history.
  • One piece of residue is retained rather than cleared, and the reason is stated: the Nextcloud app telemetry row (O5). The hub offers POST /apps/nextcloud/reset-telemetry, whose own confirm text is "Delete all telemetry data for nextcloud? This cannot be undone." That is an irreversible write on the operator's own surface, and the operator authorised Phase 6, not this. The exact command is recorded here so it is a one-line decision: curl -u ":$HUB_PW" -X POST http://<hub-clusterIP>:8080/apps/nextcloud/reset-telemetry

Nothing on demo-felhom, ep0, DooPlex or Peti's box was modified. Peti's box was never contacted.


14. What goes to the operator

One decision, and it is not a bug report. §2 shows the restart behaviour was chosen and is stated in the source. §7 shows the word "rollback" does not describe anything this product can do. The two together mean the safety question is not "fix the Update button" — it is where the safety belongs, and that is in STATUS.md, phrased as one answerable question with what happens if nothing is done.

No design is proposed here, deliberately. Four production designs in this project were specced against unvalidated mechanisms and all four were wrong; this spike exists so the fifth is not.