THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has already moved, and the Restart button does it — 18.3s with a network pull when the target image is absent, 0.5s when present, against a negative control that did not even recreate the container. The boot reconciler does the same thing unattended when an app fails to come back (bootrecon.go:269 -> StartStack). AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades nothing. Docker restores the old containers and the reconciler logs 'no boot-orphaned apps (nothing to start)'. AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run, the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported'. 'Rollback' is the wrong word for this arc and is struck. Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator confirmation. Peti's box was never contacted. No production code written in any repo: felhom-controller is at 960d29b0612c before and after, tree clean, and build/vet/test are green — run at the end to prove exactly that. Register: R-438/439/440 updated with live evidence; R-441..R-445 opened (restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB left behind; update reports success over a broken app; no fleet fstrim; hub telemetry outlives the app). Capability map gains three measured rows. The mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR, not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision. Five claims in the brief are named as wrong, including two of my own method.
13 KiB
REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01)
Task class: Spike. No production code was written in any repo. Output is a findings doc, register rows, two architecture updates and one operator decision.
Findings doc: documentation/audits/SPIKE-app-update-2026-09-01.md
Evidence: documentation/audits/evidence-spike-app-update-2026-09-01/ (19 files)
1. The Phase 1 answer, stated unambiguously
YES — docker compose up -d upgrades an app whose compose file has already moved, and the Restart
button does it.
Produced by variants 1a and 1b, and independently by 1c-ii:
- 1a (target image ABSENT locally): the restart took 18.3 s, ended with the container on the
new tag, and left the new image in the local store where it had not been — so
up -dperformed a network pull, in a code path that contains no pull step. - 1b (target image PRESENT): 0.5 s, container recreated onto the new tag, image id unchanged.
- 1d, the negative control (file NOT edited): container id, image id, digest and
StartedAtall identical —up -ddid not even recreate. The method can show "no change" when nothing changed. - 1c-ii (unattended): the boot reconciler brought a failed app back on the new version with nobody pressing anything.
And the answer to 1c is NO, which narrows the exposure: a hard guest reset upgraded nothing.
Docker's restart: unless-stopped restored the existing containers on the old image and the
reconciler logged its own verdict — Boot reconciliation: no boot-orphaned apps (nothing to start).
Per the task's own rule, no fix is proposed. The behaviour may have been chosen — RestartStack
says so in a comment — so it goes to the operator as a decision.
2. Confirmed baselines actually used
| Repo | at §1 of the task | measured at start | end | moved? |
|---|---|---|---|---|
| felhom-controller | 960d29b0612c |
960d29b0612c |
960d29b0612c |
no — read only |
| felhom.eu | 1d59353df437 |
1d59353df437 |
this session's documents | as planned |
| app-catalog-felhom.eu | 29edad9c5bf4 |
29edad9c5bf4 |
5d8f25f, tree identical to 29edad9c5bf4 |
reverted |
| felhom-agent | 4586f0f7f6d1 |
4586f0f7f6d1 |
4586f0f7f6d1 |
not touched |
None had moved since the task was written. Live controller on demo-hp: 0.232.0.
3. Per-phase results, each with its control
| Phase | Result | Control |
|---|---|---|
| 0 | R-438, R-439, R-440 filed, each with rank and owner | ids confirmed free by grepping all three register files (R-437 was highest) |
| 1 | THE GATE: YES. 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended | 1d — unchanged file, no change at all |
| 2 | Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer | positive (BentoPDF, 4 hits) + negative (zzz-never-present, 0) on both customer pages |
| 3a | Pull failure: HTTP 500, app untouched and still running | re-run as a Restart: also 500, app still ran |
| 3b | HTTP 200 "update completed" over a crash-looping app; alarm fired 5m16s later | the alarm is a positive observable (Event pushed: app_start_failed), plus the detector heartbeat 1 currently down |
| 4 | demo-hp and demo-felhom: zero drift by tag. But 2 of 6 floating pins have already MOVED upstream | both fully-pinned images (romm:5.0.0, pdo:2.0.5) were SAME |
| 5 | Existing safety dump is database-only; DB dumps are 48 KB–395 KB; the file half is the cost, and it does not fit for a large app | demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead |
| 6 | App data CANNOT be rolled back. Old image refuses to start on migrated data | returning to 32.0.9 restored both seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked |
4. The exact symbols that bring an app back — found by reading
| path | file:symbol |
|---|---|
| boot reconciler | felhom-controller/controller/internal/bootrecon/bootrecon.go:269 — Reconciler.Run → StartStack |
| its scheduler | controller/cmd/controller/main.go:2127 — runBootReconcile, called at :450 |
| app-stop guard | controller/internal/backup/appstop_marker.go:283 — AppStopGuard.Recover → StartStack |
| its hold-aware wrapper | controller/cmd/controller/main.go:2008 — gatedAppStopStarter.StartStack |
| drive-return gate | controller/internal/web/intermediary.go:222 — Server.restartStacks → StartStack |
| guest-boot change | controller/internal/web/intermediary.go:458 — Server.processGuestBootChange |
| quiesce restart | controller/internal/quiesce/quiesce.go:733 — Loop.restartAll |
| off-site reconstitution | controller/internal/backup/offbox_reconstitute.go:692 — Manager.ReconstituteFromOffsite |
Full table of all 13 non-API call sites: findings doc §8.
5. Register rows opened, updated or re-ranked
| row | what | owner |
|---|---|---|
| R-438 | opened P1-HIGH, then updated with the live evidence (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) | VIKTOR rules, CC measures |
| R-439 | opened P3-LOW, then updated — the severity survives but its stated reason was imprecise (isOperationalState counts restarting/degraded as operational) |
CC |
| R-440 | opened P2-MEDIUM, then updated: two floating pins have ALREADY moved, with a passing control | CC |
| R-441 | NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes | CC measures, VIKTOR rules |
| R-442 | NEW — P1-HIGH: remove_hdd_data:true is inert with no paths.hdd_path; 128 MB left, API said neither removed nor preserved |
CC |
| R-443 | NEW — the Update button reports success over an app it has broken | CC proposes, VIKTOR rules |
| R-444 | NEW — nothing runs pct fstrim; demo-hp's thin pool held ~23.8 GB of freed blocks |
CC |
| R-445 | NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation | VIKTOR rules, CC implements |
Nothing was closed and nothing was re-ranked. The ranking of R-438 relative to existing rows is Viktor's, and I have not moved anything.
6. Claims in the task that turned out to be wrong, named
- "a power cut … is an unattended three-major-version upgrade" — not as stated. A plain power cut upgraded nothing (measured). The unattended upgrade needs "and the app did not come back". The exposure is smaller than the operator page claims.
- "Five other code paths end in
compose up -d" — thirteen non-API call sites, nine files. - "Phase 4 — demo-hp, demo-felhom and Peti's box" — Peti's box is DOWN, not enrolled, and
runbooks/target-selection.md:161says "No access route from DooPlex". Two boxes measured live; Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history. - R-439's severity reason — right conclusion, imprecise reason (see §5).
- R-440's "23 pins" — exactly right (79 image lines / 53 apps / 66 distinct; 23 with no patch component). One arguable 24th named rather than rounded away.
- The catalog history figures — spot-checked and all correct: 153 commits, 53 apps, 0 files with upgrade metadata, and all four multi-major bumps confirmed by commit hash.
- My own method, corrected in-flight: the first customer-page search used
grep -o "2.8.6", whose unescaped.produced two false hits;grep -Fgives zero. The controls caught it.
7. Evidence handling
Evidence was written directly into documentation/audits/evidence-spike-app-update-2026-09-01/ on
DooPlex as each phase produced it — before every revert, including the intermediate ones. The
catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all
happened after their evidence was already off the machine. Nothing was lost and nothing had to be
reproduced.
8. Observations
-
The restore path writes the recovery unit's OLD image pin into the live stack dir, and the catalog syncer overwrites it again within 15 minutes. The overwrite half is measured on demo-hp; that the restore writes to that same path is read at
cmd/controller/main.go:2570, not measured — both halves are graded as such in the row. FILED: R-441 -
remove_hdd_data: trueremoved nothing. 128 MB of app data stayed on the drive while the API returned 200 withhdd_paths_removed:null, hdd_paths_preserved:null. Root cause established with controls:Paths.HDDPathhas no default and demo-hp'scontroller.yamldoes not set it, soParseComposeHDDMountsreturns nil on its first line. FILED: R-442 -
An update can report success over an app it has just broken. HTTP 200 and "updated successfully" while the container was already crash-looping; the truth reached the customer 5m16s later through the dead-app alarm rather than through the update itself. FILED: R-443
-
The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed, and
fstrimfrom inside the unprivileged container is refused;pct fstrimfrom the host reclaimed it. Nothing runs it on the fleet. FILED: R-444 -
The hub kept app telemetry for an app that no longer exists anywhere, and it now sets a fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway. Retained rather than cleared, because the reset is irreversible and on the operator's own surface. FILED: R-445
-
The syncer's debug hash line cannot show what it claims to show.
logFileHashes(internal/sync/sync.go:386) reads the destination after the write, so it printedsrc=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)— the same hash twice, with the word "changed". NOT-A-FINDING: it is DEBUG-only and theUpdated <app>/<file>INFO line immediately above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the next person reading a sync log does not try to learn from those two hashes, which cannot differ.
9. Teardown — all three layers
Layer 1 — the machine. bentopdf restored to its catalog tag v2.8.6, digest
sha256:eaeea1e447205a79…, byte-identical to the run's baseline. The throwaway nextcloud stack
removed via the product's own endpoint; containers, all three volumes and app.yaml gone. The 128 MB
the product failed to remove (observation 2) deleted by hand along with its backup dirs;
find /mnt -iname "*nextcloud*" returns nothing. Five images this run pulled removed by targeted
docker rmi — no prune of any kind was run anywhere. No guest was created; guest 9201 was
hard-reset once by design and returned with all nine apps.
Layer 2 — the host. local-lvm 68.97% → 70.91% during the run → 26.78% after pct fstrim 9201. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to
pre-run values (/ 957 M, /mnt/sys_drive 12 G, hdd_1 5.5 G).
Layer 3 — the hub. This run provisioned NOTHING. No customer record and no appliance record was
created; the customers list is unchanged at five rows, identical to the list read at the start. The
existing demo-hp customer was used. What the run did create is hub events (app_start_failed,
plus deploy/remove for the throwaway app) — retained deliberately, because the event log is an
append-only record and deleting from it to tidy a test damages the surface this project relies on for
history. One residue is retained rather than cleared with the reason and the exact one-line command
recorded in the findings doc §13 (observation 5).
Nothing on demo-felhom, ep0, DooPlex or Peti's box was modified. Peti's box was never contacted.
10. Final verification
felhom-controller: git status --porcelain → empty
HEAD = origin/main = 960d29b0612c
go build ./... → OK
go vet ./... → OK
go test ./... → rc=0, 28 packages, 0 FAIL
The controller tree was left untouched, and that is proven rather than asserted.
11. My own mistakes
- My first customer-page search used a regex where I needed a literal.
grep -o "2.8.6"treats.as a wildcard and reported two hits on a page that contains none. I caught it only because the task requires a control on every such search, and re-ran withgrep -F. The rule earned its place in the same session it was applied. - My first poll for "nextcloud is healthy" matched the wrong container. The break condition
matched
healthyanywhere in the line and fired onnextcloud-redis. Corrected to an exact-name filter. No result depended on it. - I renumbered
STATUS.mdbadly on the first attempt, leaving the list running 1–6, 8, 9 with no item 7, and left the section's own lead line saying one thing was waiting when there were two. Both fixed. This is the ranking-paragraph-goes-stale trap (R-405) in miniature, and it appeared within minutes of my writing about it.