# REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01) **Task class: Spike.** No production code was written in any repo. Output is a findings doc, register rows, two architecture updates and one operator decision. **Findings doc:** `documentation/audits/SPIKE-app-update-2026-09-01.md` **Evidence:** `documentation/audits/evidence-spike-app-update-2026-09-01/` (19 files) --- ## 1. The Phase 1 answer, stated unambiguously **YES — `docker compose up -d` upgrades an app whose compose file has already moved, and the Restart button does it.** Produced by **variants 1a and 1b, and independently by 1c-ii**: - **1a** (target image ABSENT locally): the restart took **18.3 s**, ended with the container on the new tag, and **left the new image in the local store where it had not been** — so `up -d` performed a network pull, in a code path that contains no pull step. - **1b** (target image PRESENT): **0.5 s**, container recreated onto the new tag, image id unchanged. - **1d, the negative control** (file NOT edited): container id, image id, digest and `StartedAt` all identical — `up -d` did not even recreate. **The method can show "no change" when nothing changed.** - **1c-ii** (unattended): the boot reconciler brought a failed app back on the **new** version with nobody pressing anything. **And the answer to 1c is NO, which narrows the exposure:** a hard guest reset upgraded nothing. Docker's `restart: unless-stopped` restored the existing containers on the old image and the reconciler logged its own verdict — `Boot reconciliation: no boot-orphaned apps (nothing to start)`. **Per the task's own rule, no fix is proposed.** The behaviour may have been chosen — `RestartStack` says so in a comment — so it goes to the operator as a decision. --- ## 2. Confirmed baselines actually used | Repo | at §1 of the task | measured at start | end | moved? | |---|---|---|---|---| | felhom-controller | `960d29b0612c` | `960d29b0612c` | `960d29b0612c` | **no — read only** | | felhom.eu | `1d59353df437` | `1d59353df437` | this session's documents | as planned | | app-catalog-felhom.eu | `29edad9c5bf4` | `29edad9c5bf4` | `5d8f25f`, **tree identical to `29edad9c5bf4`** | reverted | | felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched | **None had moved since the task was written.** Live controller on demo-hp: 0.232.0. --- ## 3. Per-phase results, each with its control | Phase | Result | Control | |---|---|---| | **0** | R-438, R-439, R-440 filed, each with rank and owner | ids confirmed free by grepping all three register files (R-437 was highest) | | **1** | **THE GATE: YES.** 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended | **1d** — unchanged file, no change at all | | **2** | Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer | positive (`BentoPDF`, 4 hits) + negative (`zzz-never-present`, 0) on both customer pages | | **3a** | Pull failure: HTTP 500, **app untouched and still running** | re-run as a Restart: also 500, app still ran | | **3b** | **HTTP 200 "update completed" over a crash-looping app**; alarm fired 5m16s later | the alarm is a positive observable (`Event pushed: app_start_failed`), plus the detector heartbeat `1 currently down` | | **4** | demo-hp and demo-felhom: **zero drift by tag**. But **2 of 6 floating pins have already MOVED upstream** | both fully-pinned images (`romm:5.0.0`, `pdo:2.0.5`) were SAME | | **5** | Existing safety dump is **database-only**; DB dumps are 48 KB–395 KB; the **file** half is the cost, and it does not fit for a large app | demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead | | **6** | **App data CANNOT be rolled back.** Old image refuses to start on migrated data | returning to 32.0.9 restored **both** seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked | --- ## 4. The exact symbols that bring an app back — found by reading | path | `file:symbol` | |---|---| | boot reconciler | `felhom-controller/controller/internal/bootrecon/bootrecon.go:269` — `Reconciler.Run` → `StartStack` | | its scheduler | `controller/cmd/controller/main.go:2127` — `runBootReconcile`, called at `:450` | | app-stop guard | `controller/internal/backup/appstop_marker.go:283` — `AppStopGuard.Recover` → `StartStack` | | its hold-aware wrapper | `controller/cmd/controller/main.go:2008` — `gatedAppStopStarter.StartStack` | | drive-return gate | `controller/internal/web/intermediary.go:222` — `Server.restartStacks` → `StartStack` | | guest-boot change | `controller/internal/web/intermediary.go:458` — `Server.processGuestBootChange` | | quiesce restart | `controller/internal/quiesce/quiesce.go:733` — `Loop.restartAll` | | off-site reconstitution | `controller/internal/backup/offbox_reconstitute.go:692` — `Manager.ReconstituteFromOffsite` | Full table of all 13 non-API call sites: findings doc §8. --- ## 5. Register rows opened, updated or re-ranked | row | what | owner | |---|---|---| | **R-438** | opened P1-HIGH, then **updated with the live evidence** (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) | **VIKTOR rules, CC measures** | | **R-439** | opened P3-LOW, then **updated — the severity survives but its stated reason was imprecise** (`isOperationalState` counts `restarting`/`degraded` as operational) | CC | | **R-440** | opened P2-MEDIUM, then **updated: two floating pins have ALREADY moved, with a passing control** | CC | | **R-441** | NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes | CC measures, VIKTOR rules | | **R-442** | NEW — **P1-HIGH**: `remove_hdd_data:true` is inert with no `paths.hdd_path`; 128 MB left, API said neither removed nor preserved | CC | | **R-443** | NEW — the Update button reports success over an app it has broken | CC proposes, VIKTOR rules | | **R-444** | NEW — nothing runs `pct fstrim`; demo-hp's thin pool held ~23.8 GB of freed blocks | CC | | **R-445** | NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation | VIKTOR rules, CC implements | **Nothing was closed and nothing was re-ranked.** The ranking of R-438 relative to existing rows is Viktor's, and I have not moved anything. --- ## 6. Claims in the task that turned out to be wrong, named 1. **"a power cut … is an unattended three-major-version upgrade"** — not as stated. A plain power cut upgraded nothing (measured). The unattended upgrade needs *"and the app did not come back"*. **The exposure is smaller than the operator page claims.** 2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites, nine files. 3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and `runbooks/target-selection.md:161` says *"No access route from DooPlex"*. Two boxes measured live; Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history. 4. **R-439's severity reason** — right conclusion, imprecise reason (see §5). 5. **R-440's "23 pins"** — **exactly right** (79 image lines / 53 apps / 66 distinct; 23 with no patch component). One arguable 24th named rather than rounded away. 6. **The catalog history figures** — spot-checked and **all correct**: 153 commits, 53 apps, 0 files with upgrade metadata, and all four multi-major bumps confirmed by commit hash. 7. **My own method, corrected in-flight:** the first customer-page search used `grep -o "2.8.6"`, whose unescaped `.` produced two false hits; `grep -F` gives zero. The controls caught it. --- ## 7. Evidence handling **Evidence was written directly into `documentation/audits/evidence-spike-app-update-2026-09-01/` on DooPlex as each phase produced it — before every revert, including the intermediate ones.** The catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all happened after their evidence was already off the machine. **Nothing was lost and nothing had to be reproduced.** --- ## 8. Observations 1. **The restore path writes the recovery unit's OLD image pin into the live stack dir, and the catalog syncer overwrites it again within 15 minutes.** The overwrite half is measured on demo-hp; that the restore writes to that same path is read at `cmd/controller/main.go:2570`, not measured — both halves are graded as such in the row. **FILED: R-441** 2. **`remove_hdd_data: true` removed nothing.** 128 MB of app data stayed on the drive while the API returned 200 with `hdd_paths_removed:null, hdd_paths_preserved:null`. Root cause established with controls: `Paths.HDDPath` has no default and demo-hp's `controller.yaml` does not set it, so `ParseComposeHDDMounts` returns nil on its first line. **FILED: R-442** 3. **An update can report success over an app it has just broken.** HTTP 200 and "updated successfully" while the container was already crash-looping; the truth reached the customer 5m16s later through the dead-app alarm rather than through the update itself. **FILED: R-443** 4. **The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed**, and `fstrim` from inside the unprivileged container is refused; `pct fstrim` from the host reclaimed it. Nothing runs it on the fleet. **FILED: R-444** 5. **The hub kept app telemetry for an app that no longer exists anywhere**, and it now sets a fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway. Retained rather than cleared, because the reset is irreversible and on the operator's own surface. **FILED: R-445** 6. **The syncer's debug hash line cannot show what it claims to show.** `logFileHashes` (`internal/sync/sync.go:386`) reads the destination *after* the write, so it printed `src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word "changed". **NOT-A-FINDING:** it is DEBUG-only and the `Updated /` INFO line immediately above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the next person reading a sync log does not try to learn from those two hashes, which cannot differ. --- ## 9. Teardown — all three layers **Layer 1 — the machine.** `bentopdf` restored to its catalog tag `v2.8.6`, digest `sha256:eaeea1e447205a79…`, **byte-identical to the run's baseline**. The throwaway `nextcloud` stack removed via the product's own endpoint; containers, all three volumes and `app.yaml` gone. The 128 MB the product failed to remove (observation 2) deleted by hand along with its backup dirs; `find /mnt -iname "*nextcloud*"` returns nothing. Five images this run pulled removed by **targeted `docker rmi`** — **no `prune` of any kind was run anywhere**. No guest was created; guest 9201 was hard-reset once by design and returned with all nine apps. **Layer 2 — the host.** `local-lvm` 68.97% → 70.91% during the run → **26.78%** after `pct fstrim 9201`. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to pre-run values (`/` 957 M, `/mnt/sys_drive` 12 G, `hdd_1` 5.5 G). **Layer 3 — the hub. This run provisioned NOTHING.** No customer record and no appliance record was created; the customers list is unchanged at five rows, identical to the list read at the start. The existing `demo-hp` customer was used. What the run *did* create is hub **events** (`app_start_failed`, plus deploy/remove for the throwaway app) — **retained deliberately**, because the event log is an append-only record and deleting from it to tidy a test damages the surface this project relies on for history. One residue is **retained rather than cleared** with the reason and the exact one-line command recorded in the findings doc §13 (observation 5). **Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified. Peti's box was never contacted.** --- ## 10. Final verification ``` felhom-controller: git status --porcelain → empty HEAD = origin/main = 960d29b0612c go build ./... → OK go vet ./... → OK go test ./... → rc=0, 28 packages, 0 FAIL ``` **The controller tree was left untouched, and that is proven rather than asserted.** --- ## 11. My own mistakes 1. **My first customer-page search used a regex where I needed a literal.** `grep -o "2.8.6"` treats `.` as a wildcard and reported two hits on a page that contains none. I caught it only because the task requires a control on every such search, and re-ran with `grep -F`. **The rule earned its place in the same session it was applied.** 2. **My first poll for "nextcloud is healthy" matched the wrong container.** The break condition matched `healthy` anywhere in the line and fired on `nextcloud-redis`. Corrected to an exact-name filter. No result depended on it. 3. **I renumbered `STATUS.md` badly on the first attempt**, leaving the list running 1–6, 8, 9 with no item 7, and left the section's own lead line saying one thing was waiting when there were two. Both fixed. **This is the ranking-paragraph-goes-stale trap (R-405) in miniature**, and it appeared within minutes of my writing about it.