THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has already moved, and the Restart button does it — 18.3s with a network pull when the target image is absent, 0.5s when present, against a negative control that did not even recreate the container. The boot reconciler does the same thing unattended when an app fails to come back (bootrecon.go:269 -> StartStack). AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades nothing. Docker restores the old containers and the reconciler logs 'no boot-orphaned apps (nothing to start)'. AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run, the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported'. 'Rollback' is the wrong word for this arc and is struck. Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator confirmation. Peti's box was never contacted. No production code written in any repo: felhom-controller is at 960d29b0612c before and after, tree clean, and build/vet/test are green — run at the end to prove exactly that. Register: R-438/439/440 updated with live evidence; R-441..R-445 opened (restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB left behind; update reports success over a broken app; no fleet fstrim; hub telemetry outlives the app). Capability map gains three measured rows. The mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR, not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision. Five claims in the brief are named as wrong, including two of my own method.
36 KiB
SPIKE — what an app update actually does, and which other paths do it too (2026-09-01)
THE ANSWER TO PHASE 1, IN ONE SENTENCE: YES — the Restart button upgrades the app, and so does the box's own repair when an app fails to come back, because every one of these paths ends in
docker compose up -d, andup -dmakes the container match whatever the file now says, pulling the image itself if it is missing.AND THE SECOND ANSWER, WHICH THE OPERATOR PAGE GOT WRONG IN THE HELPFUL DIRECTION: a plain power cut does NOT do this. Docker's own
restart: unless-stoppedputs the existing containers back on the OLD image, so the boot reconciler finds nothing to repair and never runsup -d. The unattended upgrade happens only in the narrower case where an app does not come back by itself — and there it happens with nobody pressing anything.AND THE THIRD, WHICH DECIDES THE VOCABULARY OF THIS WHOLE ARC: app data CANNOT be rolled back. Measured on Nextcloud: once a migration has actually run, putting the old image tag back produces a container that refuses to start — "the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported". So "rollback" is the wrong word and should be struck before anyone specs against it. The shape that is actually available is pre-update copy plus refuse-and-explain.
Scope. A spike. No production code was written, in any repo. The controller was read only:
felhom-controller is at 960d29b0612c before and after, working tree clean, and
go build ./... && go vet ./... && go test ./... is green (28 packages, 0 FAIL) — run at the end
precisely to prove the tree was left untouched.
Method. Endpoint-level throughout, which is the standard method here because there is no browser on
DooPlex: every action was the exact endpoint the UI invokes (POST /api/stacks/{name}/restart,
/update, /start, /stop, /deploy, /remove), driven over HTTPS with the customer's own session
cookie and CSRF token. No raw agent attach, no hand-set state — the F9 bypass was not used. The two
places where state was staged by hand are named as such at the point of use (Phase 1's compose edits,
and Phase 1c-ii's docker rm -f), and each says exactly what was staged and why.
Target. All destructive work ran on demo-hp (192.168.0.104, ssh hp; guest 9201 =
192.168.0.138), which runbooks/target-selection.md classes Tier 0 — disposable. DooPlex,
demo-felhom, ep0 and Peti's box were not mutated. Peti's box was not touched at all.
Evidence. audits/evidence-spike-app-update-2026-09-01/ — written straight onto DooPlex as each
phase produced it, before every revert, so R-96 rule 5 is satisfied continuously rather than at the
end. Nothing was lost and nothing had to be reproduced.
1. Confirmed baselines — all four matched §1 of the task, none had moved
| Repo | main @ commit at start |
at end | note |
|---|---|---|---|
| felhom-controller | 960d29b0612c |
960d29b0612c |
read only — unchanged |
| felhom.eu | 1d59353df437 |
moved (this spike's documents) | Phase 0 + Phase 7 |
| app-catalog-felhom.eu | 29edad9c5bf4 |
moved, then reverted to the same tree | 214d448 → 30bd892 → 5d8f25f |
| felhom-agent | 4586f0f7f6d1 |
4586f0f7f6d1 |
not touched |
Live controller on demo-hp: 0.232.0, matching the newest release. Sync interval 15m
(internal/config/config.go:351, the default, no override on the box).
2. Phase 1 — THE GATE. Does compose up -d upgrade an app whose compose file already moved?
App: bentopdf — one container, no database, no volume, no data of any kind, and deployed on
demo-hp only. Chosen so the gate could be measured with zero data risk anywhere.
Baseline: container 9473b6e4a18f, ghcr.io/alam00000/bentopdf:v2.8.6, digest
sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3, started 04:05:50Z.
| # | Variant | Result | Discriminator |
|---|---|---|---|
| 1d | negative control — file NOT edited, then restart |
no change | container id, image id, digest and StartedAt all IDENTICAL. up -d saw no change and did not even recreate. |
| 1a | restart, target image ABSENT from the local store |
UPGRADED, and it pulled | new container 93e9db74e97f, image v2.8.5, digest sha256:2d867aac…, and v2.8.5 appeared in docker images where it had not been. 18.3 s. |
| 1b | restart, target image ALREADY PRESENT |
UPGRADED, no pull | new container 222f8178dbd0, image id unchanged from baseline 4baadf01bd68. 0.5 s. |
| 1c | boot recovery — hard guest reset, app left running | NO upgrade | SAME container 222f8178dbd0 restored by Docker at 17:47:45Z on v2.8.6 while the file said v2.8.5. |
| 1c-ii | boot recovery where the app did not come back | UPGRADED, unattended | new container aa443836a4cd on v2.8.5 at 17:55:44Z. Nobody pressed anything. |
The 18.3 s vs 0.5 s split is the quantitative discriminator and it is worth more than the tags:
the restart path contains no pull step in the source, yet variant 1a spent 18 seconds and ended with a
new image in the local store. up -d pulled it.
Quoted live, variant 1a (phase1-1a-controller-log.txt):
17:35:45 router.go:566: [INFO] [api] restart requested for stack: bentopdf
17:35:45 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
17:36:03 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 18.3s)
17:36:07 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running
Quoted live, variant 1c-ii — the unattended upgrade, attributed to the exact symbol
(phase1-1cii-orphan.txt):
17:55:44 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 10s — sweeping
17:55:44 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]
17:55:44 manager.go:1039: [INFO] [stacks] Starting stack: bentopdf
17:55:44 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
17:55:48 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running
1c's negative result is a POSITIVE observable, not an absent log line — the reconciler says so in its own words, which is exactly what it was built to do:
17:48:34 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
What 1c-ii staged, stated plainly
docker rm -f bentopdf removed the container so it could not be restored by Docker's restart policy,
which is the shape a power cut leaves when an app does not come back. desired_state: running was left
untouched in app.yaml. The controller was then restarted, because that is what runs the boot
reconciler. Nothing else was staged, and the reconciler selected the app on its own.
Why 1c could not use a hand-edited file, and how that was handled
At controller start the initial catalog sync runs immediately (sync.go:98, 17:47:49Z) while the
boot reconciler waits out bootReconcileSettle = 5 s plus a settle window (17:48:34Z). The sync
wins by ~45 s, so a hand-edited compose would have been overwritten before the reconciler ever saw
it, and 1c would have measured nothing. 1c and 1c-ii were therefore run inside Phase 2's window, where
the catalog itself carried the new tag — which is also the faithful production shape.
The design intent is already stated in the source, and it is not hidden
Manager.RestartStack (internal/stacks/manager.go:1133) carries this comment:
"Use
up -dinstead of barerestartso that env vars from app.yaml are injected and any template changes (new images, healthchecks) are picked up."
So the restart behaviour was chosen, deliberately, and written down. What is NOT written down anywhere is the consequence once the catalog syncer moves the file underneath a deployed app. That gap is R-438, and whether the choice should extend to the unattended paths is the operator's ruling.
3. Phase 2 — the sync overwrite, confirmed live
A real tag change (bentopdf v2.8.6 → v2.8.5) was pushed to the catalog main at 17:38:22Z
(214d448) and left to travel the real 15-minute cycle. No hand-edit; the catalog was the only input.
| time (UTC) | file on demo-hp | running container |
|---|---|---|
| 17:38:35 (pre-sync) | v2.8.6 |
v2.8.6 |
| 17:45:15 | v2.8.6 |
v2.8.6 |
| 17:45:37 (post-sync) | v2.8.5 |
v2.8.6 |
The sync's own line, and the file's mtime, agree to the second:
17:45:17 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
mtime: 2026-09-01 17:45:17.646017366 +0000 /opt/docker/stacks/bentopdf/docker-compose.yml
container started: 2026-09-01T17:36:35Z (unchanged)
Syncer.copyTemplates (internal/sync/sync.go:319) has no deployed check of any kind. Its only
guard is a sha256 content compare in copyIfChanged, and its only exclusion is app.yaml. The
post-sync hook is stackMgr.InjectMissingFields(updated) and nothing else — the sync does not
restart anything, which is why the file and the container can disagree indefinitely.
Was the customer told? No — and the page does not even show a version.
Searched with ASCII fragments and both controls, per the standing rule.
| page | BentoPDF (positive control) |
zzz-never-present (negative control) |
v2.8.5 |
v2.8.6 |
|---|---|---|---|---|
/apps/bentopdf (40 360 B) |
4 | 0 | 0 | 0 |
/stacks (140 883 B) |
1 | 0 | 0 | 0 |
A correction to my own method, made here rather than buried: the first pass used
grep -o "2.8.6", where . is a regex wildcard, and it reported 2 hits. Re-run with grep -F the
count is 0. The controls are what exposed it. The customer's pages carry no version string at
all — not the running one, not the pending one.
No event, no notification and no email were emitted by the sync. The only Friss… strings on the
pages are the catalog-sync toast and the Frissítés button label.
The revert, verified on the box
30bd892 reverted the pin; the sync landed it at 18:10:29Z
([INFO] [sync] Updated bentopdf/docker-compose.yml, Sablonok frissítve — frissítve: bentopdf) and
the file read v2.8.6 again. The catalog tree is byte-identical to 29edad9c5bf4.
An unlooked-for second result: that same sync also overwrote a hand-broken compose file (Phase 3's
alpine:3.20 edit) with the catalog's version. So the syncer is also the repair path for a broken
app definition — within 15 minutes, automatically. The container stays broken until something runs
up -d, but the definition heals itself.
4. Phase 3 — what a failed update looks like to the customer
3a — the pull fails. Loud, honest, and harmless.
Compose pointed at ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist, then POST .../update:
HTTP 500
{"ok":false,"error":"pulling images for bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ... Error manifest unknown\nError response from daemon: manifest unknown"}
The app was untouched: same container aa443836a4cd, v2.8.5, status=running, RestartCount=0.
UpdateStack returns after the failed pull and never reaches up -d.
3a-ii, the landmine that was not one. The failed update leaves the file pointing at an image that
does not exist, while the page shows a green healthy app with a Restart button beside the Update
button. Pressing Restart in that state was measured: HTTP 500, and the app still ran. compose up -d resolves every image before it touches a container, so a bad reference fails before anything stops.
This is the reassuring half of Phase 3 and it should be said as plainly as the bad half.
What the customer is shown, though, is raw English Docker output — exit code 1, stderr:,
manifest unknown, Error response from daemon — on a product whose every other error string is
Hungarian.
3b — the pull succeeds and the app does not. This is the one that matters.
Compose pointed at alpine:3.20 — a real image that pulls cleanly and then exits at once, so
restart: unless-stopped puts it in a crash loop. This is the shape of a real upstream image whose
configuration the template can no longer supply, and the catalog's own history contains exactly that
case (7350cd9, "wger: revert 2.6 -> 2.3 (2.6 needs a full DB config the template cannot supply)").
POST /api/stacks/bentopdf/update
HTTP 200
{"ok":true,"message":"Stack bentopdf update completed"}
18:00:47 manager.go:1195: [INFO] [stacks] Stack bentopdf updated successfully (took 3.5s)
Reality 37 seconds later: status=restarting, RestartCount=9, alpine:3.20.
The controller's own post-start line told the truth — and it ran after the API had already
answered ok:true:
18:00:50 manager.go:1403: [INFO] [stacks] bentopdf alpine:3.20 restarting Restarting (0) Less than a second ago
What the customer's page said (/stacks, live HTML):
BentoPDF·pdf.enkisfelhom.hu· „URL nem elérhető – útvonal nincs publikálva" · badge „Újraindítás…" ·Restarting (0) 15 seconds ago· buttons:FrissítésÚjraindításLeállításNaplókRészletek
„Újraindítás…" reads as transient, not as failure. Nothing on the page says the update broke the
app. isOperationalState (internal/web/funcmap.go:90) counts StateRestarting and StateDegraded
as operational, so the full button row — including the green Frissítés — is rendered over a
crash-looping app.
Is there a route back? Not from the page. Every button offered re-runs the same broken definition or stops the app. The compose file is not customer-editable. The routes that exist are (a) the catalog being corrected, which then heals the file within 15 minutes, or (b) a restore, which has its own problem — see §7.
The alarm DOES fire — 5 minutes 16 seconds later, and by a different road
This was measured to a positive observable rather than inferred from silence:
18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
Update was at 18:00:43; the alarm at 18:05:59. The delay is crashLoopAfter = 5 * time.Minute
(internal/stacks/manager.go), and it is deliberate and well-argued in its own comment — a shorter
threshold would alarm on every routine deploy. The severity word is warning, which is inside the
hub's exact vocabulary, so this one really does reach the customer (R-328/R-329 class avoided).
So the honest summary of 3b is not "silently broken". It is: the button lied at the moment it was pressed, the page then described a failure as a restart, and the truth arrived five minutes later through the dead-app alarm rather than through the update the customer actually performed.
5. Phase 4 — how far behind is the real fleet
Read from the running container, never from the file — the file is the thing that has already moved. Both tag and digest recorded.
demo-hp — zero drift by tag. Nine deployed apps, every running tag equal to the catalog pin.
This is not luck: the box was reinstalled 2026-08-21 and its apps were deployed ~13 h before this run. It is a young box, and it is the reason Phase 5 could not be priced on it (§6).
demo-felhom — one deployed app, opengist, running ghcr.io/thomiceli/opengist:1.13 = catalog pin.
But "up to date by tag" is not up to date. Two floating pins have already moved upstream.
Running digests compared against what the registry serves for the same tag today, with fully-pinned tags as the control:
| image | kind | verdict |
|---|---|---|
postgres:16-alpine |
floating | SAME |
redis:7-alpine |
floating | SAME |
mariadb:11.6 |
floating | SAME |
ghcr.io/thomiceli/opengist:1.13 |
floating | SAME |
mariadb:11.4 |
floating | MOVED — running sha256:4f1d8d20…, upstream now sha256:611a2fcc… |
mariadb:12.3 |
floating | MOVED — running sha256:a02fe89c…, upstream now sha256:dd9b303a… |
rommapp/romm:5.0.0 |
pinned (CONTROL) | SAME |
privatebin/pdo:2.0.5 |
pinned (CONTROL) | SAME |
Both controls held and both positives are floating tags. So on a box with no visible drift at all,
pressing Frissítés today would silently swap the database engine build under romm (MariaDB
11.4) and bookstack (MariaDB 12.3) — with no catalog change, no version change on any screen, and no
record anywhere of what it was before. That is R-440, no longer as an argument but as a measurement.
Peti's box — NOT measurable, and the task's premise here was wrong
The task asks Phase 4 to inspect three boxes. runbooks/target-selection.md:161 states plainly:
"Currently DOWN, no enrolled host. No access route from DooPlex, and nothing here needs one." The
hub agrees: the Hosts page lists exactly two enrolled hosts (demo-felhom-8363b5,
demo-hp-bb76ea), and the peti-felhom customer shows status DOWN, last report 48 d ago, controller
0.115.0.
So there is no running container to read, and the hub does not record image tags at all — the report's container payload carries name, state, CPU and memory, and no image field. The row is therefore recorded as UNKNOWN, not guessed.
What IS knowable read-only, and it matters: the box last reported ~2026-07-15 running rallly +
rallly-postgres. On 2026-07-18 — three days after it went quiet — the catalog moved
rallly 3.11.2 → 4.11.1 [MAJOR] (e3f3a81). So a one-major upgrade is queued behind that box's
next boot. On this spike's own measurements that upgrade will NOT fire on the power-on itself
(1c: Docker restores the containers on the old image), but it will fire the moment any app fails to
come back, or anyone presses Restart or Update. Nothing was touched on that box to establish this;
it is the catalog's git history plus the hub's own record.
6. Phase 5 — what a pre-update copy would cost
The existing safety machinery is DATABASE-ONLY, and that is the headline
Manager.writeSafetyDump (internal/backup/offbox_reconstitute.go:207) discovers the app's databases
and dumps each one. An app with no database gets nothing at all — len(mine) == 0 returns an empty
set. There is no file-level safety copy on the R-361 path.
The database half is nearly free. Measured on demo-hp:
Every .sql dump on the box, including the real pre-restore-* undo copies from the August restore
work: 48 KB – 395 KB. Largest is paperless-ngx-postgres.sql at 395 065 bytes.
The file half is the cost, and demo-hp cannot price it. Stated, not papered over.
The box holds a few MB of app content; the bulk is engine data directories:
| app | biggest components | total |
|---|---|---|
| kimai | kimai_db_data 166 M + kimai_var 56 M |
≈ 222 MB |
| romm | romm_db_data 166 M + romm_redis_data 25 M + roms 1.1 M |
≈ 192 MB |
| bookstack | bookstack_db_data 166 M + bookstack_config 6.8 M |
≈ 173 MB |
These are not customer-scale numbers — 166 MB is a fresh MariaDB's preallocated files, not anyone's data. So rather than over-claim from a young box, the arc's already-measured figures are cited:
CAMPAIGN-10-two-storage-soak-2026-07-31.md§6, 66 restores plus an M-band point 327× larger: backup ≈ 29 s + 17.4 s/GB, and a DB-backed app's recovery unit is 1.90× its data (21.1 GB of data produced a 40.2 GB unit).
Applying that to demo-hp's three heaviest: ≈ 32–33 s each, unit ≈ 330–420 MB. Trivial.
And that is exactly what makes the ceiling the real finding. For a large catalog app — Immich,
Nextcloud, Plex — the same arithmetic gives, for 100 GB of data, ≈ 29 minutes and a ~190 GB copy,
against a default appliance whose /mnt/sys_drive ships at 20 GB. A pre-update copy of a large
app does not fit on a default box, and no amount of tuning the copy changes that. Whatever is
designed here has to answer that before it answers anything else. Where the copy lives, and whether it
is a copy at all or a snapshot, is a design decision and is not proposed here.
7. Phase 6 — can app data be rolled back at all? NO.
Run on operator confirmation, on a throwaway Nextcloud on demo-hp. Nothing else on any box was
involved; the app was created, used and destroyed inside this phase.
Seeded on 31.0.14.1 (occ status), two independent markers — a database-backed system config value
and a real file on disk:
System config value spike_marker set to string SPIKE-SEEDED-ON-31.0.14-2026-09-01
/var/www/html/data/admin/files/spike-marker.txt = SPIKE-FILE-CONTENT-31.0.14-2026-09-01
occ files:scan admin → 5 Folders, 53 Files, 2 Updated, 0 Errors
6a — the 3-major jump (31.0.14 → 34.0.1). Refused by the app, reported as success by the product.
POST /api/stacks/nextcloud/update → HTTP 200 {"ok":true,"message":"Stack nextcloud update completed"}
The container then crash-looped, saying:
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.It is only possible to upgrade one major version at a time.
This is R-40, measured. The catalog really does carry this jump today: 5e2c1ae, 2026-07-18,
"nextcloud: 31.0.14-apache -> 34.0.1-apache [MAJOR]".
It was fully recoverable — putting 31.0.14 back gave a healthy app in 9 seconds with both markers
intact. But that is only because nothing migrated. The refusal is Nextcloud's own safety net doing
its job, and it is why 6a is the easy case.
6b — a migration that actually runs (31.0.14 → 32.0.9, one major). It ran.
Initializing nextcloud 32.0.9.2 ...
Upgrading nextcloud from 31.0.14.1 ...
Updated database
Updated <dav> to 1.34.2 · <files> to 2.4.0 · <files_sharing> to 1.24.1 · … (20+ apps)
Healthy in 31 s on 32.0.9.2, both markers still readable.
6c — THE ROLLBACK ATTEMPT. Refused. The old image will not start on migrated data.
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker
image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the
newest image version?
Crash loop, indefinitely.
6d — positive control: the data is not destroyed, only the downgrade is blocked
Putting 32.0.9 back gave a healthy app in 46 s, occ status 32.0.9.2, and both markers read back
byte-identical to what was seeded on 31. So the failure in 6c is a refusal, not corruption — which is
the distinction that decides the remedy.
What this settles
Putting the old image tag back is not a rollback and must not be described as one. The only route
back from a migration that has run is restoring the DATA from a copy taken before the update —
which is precisely the thing the update path does not take (§2 of the task, confirmed by reading
Manager.UpdateStack, and confirmed live: no dump, no copy, no hold, in any of the six updates run
here).
8. The exact symbols — found by reading, as required
Every path that ends in compose up -d on a customer's stack folder, with the enclosing function.
| what brings the app back | symbol | line |
|---|---|---|
| the four compose callers | Manager.StartStack / StopStack / RestartStack / UpdateStack |
internal/stacks/manager.go:1029 / 1106 / 1133 / 1170 |
| partial start (R-47 DB-only window) | Manager.StartStackServices |
internal/stacks/manager.go:1082 |
| the boot reconciler | Reconciler.Run → r.stacks.StartStack(name) |
internal/bootrecon/bootrecon.go:223 → :269 |
| its scheduler | runBootReconcile → bootReconcileFn |
cmd/controller/main.go:2127 → :2172, called at :450 |
| the app-stop guard | AppStopGuard.Recover → g.starter.StartStack(name) |
internal/backup/appstop_marker.go:264 → :283 |
| its hold-aware wrapper | gatedAppStopStarter.StartStack |
cmd/controller/main.go:2008 |
| the drive-return gate | Server.restartStacks → s.stackMgr.StartStack(name) |
internal/web/intermediary.go:220 → :222 |
| guest-boot change handler | Server.processGuestBootChange |
internal/web/intermediary.go:395 → :458 |
| quiesce restart-after-backup | Loop.restartAll |
internal/quiesce/quiesce.go:730 → :733 |
| off-site reconstitution | Manager.ReconstituteFromOffsite |
internal/backup/offbox_reconstitute.go:480 → :692 |
| restore from unit / local / tier-2 | restore_unit.go:381, restore.go:76, tier2_restore.go:418, backup.go:895 |
— |
.fab export / import |
appexport/export.go:286, appexport/restore.go:461 |
— |
| integrations | onlyoffice_filebrowser.go:61, :96 (RestartStack) |
— |
CORRECTION TO THE TASK: not five other paths — thirteen. Excluding the three API actions
(start/restart/update) and excluding interface declarations and adapters, there are 13 call
sites across 9 files that call StartStack or RestartStack, and every one of them ends in
docker compose up -d against the live compose file.
9. R-439, confirmed by reading and refined
Router.actionStack (internal/api/router.go:565) checks the hold under
if action == "start" || action == "restart" — update is absent and falls through to
UpdateStack. The design intent is stated in RestoreHoldFor's own comment
(internal/backup/offbox_reconstitute.go:323):
"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."
The severity argument in the task is correct, but for a narrower reason than it states. The task
says the UI hides Frissítés unless the app is operational, and a held app is stopped. That is right —
but isOperationalState (internal/web/funcmap.go:90) counts StateRestarting and StateDegraded
as operational too, which was observed live in Phase 3b: the green Frissítés button was rendered
over a crash-looping app. So the button is hidden specifically because a held app is StateStopped,
not because broken apps hide it. The conclusion (LOW, not customer-reachable) survives; the reason
needs stating precisely, and the fix needs a test pinning it or the comment stays a wish.
10. Every claim in the task that turned out to be wrong, named
- "a power cut on a sleeping customer's box is an unattended three-major-version upgrade" — NOT AS STATED. Measured: a hard guest reset upgraded nothing, because Docker restored the containers itself and the reconciler had no orphan. The unattended upgrade is real but needs the narrower precondition "and the app did not come back". The exposure is smaller than the operator page claims, and saying so is more useful than leaving the scarier version standing.
- "Five other code paths end in
compose up -d" — thirteen non-API call sites across nine files (§8). - "Phase 4 — demo-hp, demo-felhom and Peti's box" — Peti's box is DOWN, not enrolled, and has no access route from DooPlex by the project's own runbook. Two boxes were measured live; Peti's row is UNKNOWN, with what is knowable recorded from the hub and the catalog history (§5).
- "R-440 — 23 catalog image pins float" — CONFIRMED EXACTLY (79
image:lines, 53 apps, 66 distinct; 23 with no patch component). One arguable 24th is recorded rather than rounded away:ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0pins both extensions exactly and leaves the PostgreSQL patch floating. - The catalog history numbers — spot-checked and all correct: 153 commits touching
templates/, 53 apps, 0.felhom.ymlcarrying any upgrade metadata (the onlyupgradematches are prose about STARTTLS), and all four multi-major bumps confirmed on 2026-07-18 — nextcloud5e2c1ae, grafanab789acc, calcom147cee7, vikunja3fa63cd. - "The Update button … takes no safety copy, cannot undo itself, does not stop the app if it goes wrong" — CONFIRMED, by reading and across six live updates.
- A methodological correction of my own, not the task's: the first customer-page search used
grep -o "2.8.6"and the unescaped.produced two false hits. Re-run withgrep -F: zero. The controls caught it (§3).
11. What was NOT measured
- Whether the drive-return gate and
AppStopGuard.Recoverupgrade in practice. Both were located by reading (§8) and both callStartStack, which is the same function variant 1c-ii measured upgrading an app. The mechanism is measured; these two specific entry points were not exercised live. Naming them as read-not-measured is deliberate. - Whether
RecreateStackDefinitionFromUnit's rollback is really undone by the sync in a live restore — see §12, which states which half is measured and which is read. - Any behaviour on Peti's box (§5).
12. Observations — noticed, documented, not acted on
O1 — the restore path and the catalog sync disagree about the image, and the sync wins.
stackAdapter.RecreateStackDefinitionFromUnit (cmd/controller/main.go:2570) writes the recovery
unit's captured docker-compose.yml — carrying the OLD image pin — straight into the live stack
dir (os.WriteFile(filepath.Join(stackDir, fname), data, 0644)), and restore_unit.go:317 says so:
"Resolved from the UNIT's compose, because that file is about to BECOME the live one." But
copyIfChanged overwrites any file whose content differs from the catalog, on the next 15-minute tick.
MEASURED: a locally-modified compose (Phase 3's alpine:3.20) was overwritten by the sync at
18:10:29Z. READ, not measured: that the restore writes to that same path. So a restore's
image-level rollback has a ≤15-minute half-life, and then the next up -d from any source
re-applies the catalog pin. Filed as R-441.
O2 — remove_hdd_data: true is inert on this box, and the response says neither removed nor
preserved. Removing the Phase 6 Nextcloud with {"remove_hdd_data":true,"remove_backups":true}
returned HTTP 200 with "hdd_paths_removed":null,"hdd_paths_preserved":null, and left 128 MB at
/mnt/felhom-drives/hdd_1/appdata/nextcloud. Root cause, with controls: Paths.HDDPath
(internal/config/config.go:117) has no default — only an env override at :403 — and demo-hp's
controller.yaml paths: block contains only data_dir, stacks_dir, system_data_path. The
container has no FELHOM_PATHS_* variable at all (measured, count 0). So cfg.Paths.HDDPath == "" and
ParseComposeHDDMounts (internal/stacks/delete.go:600-603) returns nil on its first line —
[INFO] found 0 HDD mounts — for a compose that plainly contains
- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data. The second half of the same removal also
no-op'd: [WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps. Filed as R-442.
O3 — an update can report success over a broken app, and the truth arrives by another road 5 minutes later. §4. Filed as R-443.
O4 — demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed. This run
added ~1.05 GiB that local-lvm did not reclaim; fstrim inside the unprivileged container is refused
(FITRIM ioctl failed: Operation not permitted), and pct fstrim 9201 from the host then trimmed
30.2 GiB + 57 GiB and took local-lvm from 70.91% → 26.78% — i.e. 23.8 GB below this run's own
starting point. Nothing runs pct fstrim on the fleet. Filed as R-444.
O5 — the hub keeps app telemetry for an app that no longer exists anywhere. The throwaway Nextcloud
now sets a fleet-wide Suggested Limit (P95×1.2) = 352 MB for Nextcloud, from ~15 minutes of a
crash-looping instance, plus three MariaDB io_uring "Known Issues" attributed to demo-hp. Retained
deliberately, not cleared — see §13. Filed as R-445.
O6 — NOT-A-FINDING: the sync's debug hash line cannot show what changed. logFileHashes
(internal/sync/sync.go:386) reads the destination after the write, so it prints
src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed) — the same hash twice, with the word "changed".
Harmless (DEBUG only, and the Updated <app>/<file> INFO line above it carries the fact), but it
cannot serve the purpose its name implies. Not filed; recorded here so the next person does not trust
it.
13. Teardown — all three layers
Layer 1 — the machine.
bentopdfrestored to its catalog tagv2.8.6, digestsha256:eaeea1e447205a79…— byte-identical to the run's baseline. Container02c80375fcba, running, healthy.- The throwaway
nextcloudstack removed viaPOST /api/stacks/nextcloud/remove(after the required stop): containers gone, all three named volumes gone,app.yamlgone. The stack dir holds only the catalog template (docker-compose.yml,.felhom.yml), i.e. not deployed. - The 128 MB the product did not remove (O2) was deleted by hand, together with the nextcloud
backup dirs on both drives.
find /mnt -iname "*nextcloud*"returns nothing. - Four images this run pulled were removed by targeted
docker rmi—nextcloud:31.0.14-apache,32.0.9-apache,34.0.1-apache,bentopdf:v2.8.5, plusalpine:3.20. Nopruneof any kind was run, anywhere. - No guest was created. Guest 9201 was hard-reset once, by design (variant 1c), and came back with all nine apps.
Layer 2 — the host. pvesm status, before → after:
| pool | before | after trim | note |
|---|---|---|---|
local |
44.42% | 44.44% | unchanged in substance |
local-lvm |
68.97% (38 959 729 KiB) | 26.78% (15 127 469 KiB) | the run's ~1.05 GiB was returned, and pct fstrim 9201 reclaimed 23.8 GB more that predated this run (O4) |
Guest: / 957 M used (baseline 957 M), /mnt/sys_drive 12 G / 18%, /mnt/felhom-drives/hdd_1 5.5 G —
all back to their pre-run values.
Layer 3 — the hub. This run provisioned NOTHING.
- No customer record and no appliance record was created. The customers list is unchanged at five
rows:
demo-felhom,demo-hp,drill-r50,peti-felhom,tester-1— identical to the list read at the start of the run. The existingdemo-hpcustomer was used throughout. - What the run DID create is events —
app_start_failed(warning) for BentoPDF at18:05:59Z, plus deploy/remove events for the throwaway Nextcloud. These are retained deliberately: the event log is an append-only record and deleting from it to tidy up a test would damage the very surface this project relies on for history. - One piece of residue is retained rather than cleared, and the reason is stated: the Nextcloud app
telemetry row (O5). The hub offers
POST /apps/nextcloud/reset-telemetry, whose own confirm text is "Delete all telemetry data for nextcloud? This cannot be undone." That is an irreversible write on the operator's own surface, and the operator authorised Phase 6, not this. The exact command is recorded here so it is a one-line decision:curl -u ":$HUB_PW" -X POST http://<hub-clusterIP>:8080/apps/nextcloud/reset-telemetry
Nothing on demo-felhom, ep0, DooPlex or Peti's box was modified. Peti's box was never contacted.
14. What goes to the operator
One decision, and it is not a bug report. §2 shows the restart behaviour was chosen and is stated
in the source. §7 shows the word "rollback" does not describe anything this product can do. The two
together mean the safety question is not "fix the Update button" — it is where the safety belongs,
and that is in STATUS.md, phrased as one answerable question with what happens if nothing is done.
No design is proposed here, deliberately. Four production designs in this project were specced against unvalidated mechanisms and all four were wrong; this spike exists so the fifth is not.