Files
felhom.eu/REPORT.md
T
admin 56c7e373a3
gates / gates (push) Successful in 18s
SPIKE: what an app update actually does, and which other paths do it too
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
2026-09-01 21:35:32 +02:00

13 KiB
Raw Blame History

REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01)

Task class: Spike. No production code was written in any repo. Output is a findings doc, register rows, two architecture updates and one operator decision.

Findings doc: documentation/audits/SPIKE-app-update-2026-09-01.md Evidence: documentation/audits/evidence-spike-app-update-2026-09-01/ (19 files)


1. The Phase 1 answer, stated unambiguously

YES — docker compose up -d upgrades an app whose compose file has already moved, and the Restart button does it.

Produced by variants 1a and 1b, and independently by 1c-ii:

  • 1a (target image ABSENT locally): the restart took 18.3 s, ended with the container on the new tag, and left the new image in the local store where it had not been — so up -d performed a network pull, in a code path that contains no pull step.
  • 1b (target image PRESENT): 0.5 s, container recreated onto the new tag, image id unchanged.
  • 1d, the negative control (file NOT edited): container id, image id, digest and StartedAt all identical — up -d did not even recreate. The method can show "no change" when nothing changed.
  • 1c-ii (unattended): the boot reconciler brought a failed app back on the new version with nobody pressing anything.

And the answer to 1c is NO, which narrows the exposure: a hard guest reset upgraded nothing. Docker's restart: unless-stopped restored the existing containers on the old image and the reconciler logged its own verdict — Boot reconciliation: no boot-orphaned apps (nothing to start).

Per the task's own rule, no fix is proposed. The behaviour may have been chosen — RestartStack says so in a comment — so it goes to the operator as a decision.


2. Confirmed baselines actually used

Repo at §1 of the task measured at start end moved?
felhom-controller 960d29b0612c 960d29b0612c 960d29b0612c no — read only
felhom.eu 1d59353df437 1d59353df437 this session's documents as planned
app-catalog-felhom.eu 29edad9c5bf4 29edad9c5bf4 5d8f25f, tree identical to 29edad9c5bf4 reverted
felhom-agent 4586f0f7f6d1 4586f0f7f6d1 4586f0f7f6d1 not touched

None had moved since the task was written. Live controller on demo-hp: 0.232.0.


3. Per-phase results, each with its control

Phase Result Control
0 R-438, R-439, R-440 filed, each with rank and owner ids confirmed free by grepping all three register files (R-437 was highest)
1 THE GATE: YES. 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended 1d — unchanged file, no change at all
2 Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer positive (BentoPDF, 4 hits) + negative (zzz-never-present, 0) on both customer pages
3a Pull failure: HTTP 500, app untouched and still running re-run as a Restart: also 500, app still ran
3b HTTP 200 "update completed" over a crash-looping app; alarm fired 5m16s later the alarm is a positive observable (Event pushed: app_start_failed), plus the detector heartbeat 1 currently down
4 demo-hp and demo-felhom: zero drift by tag. But 2 of 6 floating pins have already MOVED upstream both fully-pinned images (romm:5.0.0, pdo:2.0.5) were SAME
5 Existing safety dump is database-only; DB dumps are 48 KB–395 KB; the file half is the cost, and it does not fit for a large app demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead
6 App data CANNOT be rolled back. Old image refuses to start on migrated data returning to 32.0.9 restored both seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked

4. The exact symbols that bring an app back — found by reading

path file:symbol
boot reconciler felhom-controller/controller/internal/bootrecon/bootrecon.go:269 — Reconciler.Run → StartStack
its scheduler controller/cmd/controller/main.go:2127 — runBootReconcile, called at :450
app-stop guard controller/internal/backup/appstop_marker.go:283 — AppStopGuard.Recover → StartStack
its hold-aware wrapper controller/cmd/controller/main.go:2008 — gatedAppStopStarter.StartStack
drive-return gate controller/internal/web/intermediary.go:222 — Server.restartStacks → StartStack
guest-boot change controller/internal/web/intermediary.go:458 — Server.processGuestBootChange
quiesce restart controller/internal/quiesce/quiesce.go:733 — Loop.restartAll
off-site reconstitution controller/internal/backup/offbox_reconstitute.go:692 — Manager.ReconstituteFromOffsite

Full table of all 13 non-API call sites: findings doc §8.


5. Register rows opened, updated or re-ranked

row what owner
R-438 opened P1-HIGH, then updated with the live evidence (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) VIKTOR rules, CC measures
R-439 opened P3-LOW, then updated — the severity survives but its stated reason was imprecise (isOperationalState counts restarting/degraded as operational) CC
R-440 opened P2-MEDIUM, then updated: two floating pins have ALREADY moved, with a passing control CC
R-441 NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes CC measures, VIKTOR rules
R-442 NEW — P1-HIGH: remove_hdd_data:true is inert with no paths.hdd_path; 128 MB left, API said neither removed nor preserved CC
R-443 NEW — the Update button reports success over an app it has broken CC proposes, VIKTOR rules
R-444 NEW — nothing runs pct fstrim; demo-hp's thin pool held ~23.8 GB of freed blocks CC
R-445 NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation VIKTOR rules, CC implements

Nothing was closed and nothing was re-ranked. The ranking of R-438 relative to existing rows is Viktor's, and I have not moved anything.


6. Claims in the task that turned out to be wrong, named

  1. "a power cut … is an unattended three-major-version upgrade" — not as stated. A plain power cut upgraded nothing (measured). The unattended upgrade needs "and the app did not come back". The exposure is smaller than the operator page claims.
  2. "Five other code paths end in compose up -d" — thirteen non-API call sites, nine files.
  3. "Phase 4 — demo-hp, demo-felhom and Peti's box" — Peti's box is DOWN, not enrolled, and runbooks/target-selection.md:161 says "No access route from DooPlex". Two boxes measured live; Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history.
  4. R-439's severity reason — right conclusion, imprecise reason (see §5).
  5. R-440's "23 pins" — exactly right (79 image lines / 53 apps / 66 distinct; 23 with no patch component). One arguable 24th named rather than rounded away.
  6. The catalog history figures — spot-checked and all correct: 153 commits, 53 apps, 0 files with upgrade metadata, and all four multi-major bumps confirmed by commit hash.
  7. My own method, corrected in-flight: the first customer-page search used grep -o "2.8.6", whose unescaped . produced two false hits; grep -F gives zero. The controls caught it.

7. Evidence handling

Evidence was written directly into documentation/audits/evidence-spike-app-update-2026-09-01/ on DooPlex as each phase produced it — before every revert, including the intermediate ones. The catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all happened after their evidence was already off the machine. Nothing was lost and nothing had to be reproduced.


8. Observations

  1. The restore path writes the recovery unit's OLD image pin into the live stack dir, and the catalog syncer overwrites it again within 15 minutes. The overwrite half is measured on demo-hp; that the restore writes to that same path is read at cmd/controller/main.go:2570, not measured — both halves are graded as such in the row. FILED: R-441

  2. remove_hdd_data: true removed nothing. 128 MB of app data stayed on the drive while the API returned 200 with hdd_paths_removed:null, hdd_paths_preserved:null. Root cause established with controls: Paths.HDDPath has no default and demo-hp's controller.yaml does not set it, so ParseComposeHDDMounts returns nil on its first line. FILED: R-442

  3. An update can report success over an app it has just broken. HTTP 200 and "updated successfully" while the container was already crash-looping; the truth reached the customer 5m16s later through the dead-app alarm rather than through the update itself. FILED: R-443

  4. The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed, and fstrim from inside the unprivileged container is refused; pct fstrim from the host reclaimed it. Nothing runs it on the fleet. FILED: R-444

  5. The hub kept app telemetry for an app that no longer exists anywhere, and it now sets a fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway. Retained rather than cleared, because the reset is irreversible and on the operator's own surface. FILED: R-445

  6. The syncer's debug hash line cannot show what it claims to show. logFileHashes (internal/sync/sync.go:386) reads the destination after the write, so it printed src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed) — the same hash twice, with the word "changed". NOT-A-FINDING: it is DEBUG-only and the Updated <app>/<file> INFO line immediately above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the next person reading a sync log does not try to learn from those two hashes, which cannot differ.


9. Teardown — all three layers

Layer 1 — the machine. bentopdf restored to its catalog tag v2.8.6, digest sha256:eaeea1e447205a79…, byte-identical to the run's baseline. The throwaway nextcloud stack removed via the product's own endpoint; containers, all three volumes and app.yaml gone. The 128 MB the product failed to remove (observation 2) deleted by hand along with its backup dirs; find /mnt -iname "*nextcloud*" returns nothing. Five images this run pulled removed by targeted docker rmi — no prune of any kind was run anywhere. No guest was created; guest 9201 was hard-reset once by design and returned with all nine apps.

Layer 2 — the host. local-lvm 68.97% → 70.91% during the run → 26.78% after pct fstrim 9201. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to pre-run values (/ 957 M, /mnt/sys_drive 12 G, hdd_1 5.5 G).

Layer 3 — the hub. This run provisioned NOTHING. No customer record and no appliance record was created; the customers list is unchanged at five rows, identical to the list read at the start. The existing demo-hp customer was used. What the run did create is hub events (app_start_failed, plus deploy/remove for the throwaway app) — retained deliberately, because the event log is an append-only record and deleting from it to tidy a test damages the surface this project relies on for history. One residue is retained rather than cleared with the reason and the exact one-line command recorded in the findings doc §13 (observation 5).

Nothing on demo-felhom, ep0, DooPlex or Peti's box was modified. Peti's box was never contacted.


10. Final verification

felhom-controller: git status --porcelain → empty
                   HEAD = origin/main = 960d29b0612c
                   go build ./...  → OK
                   go vet ./...    → OK
                   go test ./...   → rc=0, 28 packages, 0 FAIL

The controller tree was left untouched, and that is proven rather than asserted.


11. My own mistakes

  1. My first customer-page search used a regex where I needed a literal. grep -o "2.8.6" treats . as a wildcard and reported two hits on a page that contains none. I caught it only because the task requires a control on every such search, and re-ran with grep -F. The rule earned its place in the same session it was applied.
  2. My first poll for "nextcloud is healthy" matched the wrong container. The break condition matched healthy anywhere in the line and fired on nextcloud-redis. Corrected to an exact-name filter. No result depended on it.
  3. I renumbered STATUS.md badly on the first attempt, leaving the list running 1–6, 8, 9 with no item 7, and left the section's own lead line saying one thing was waiting when there were two. Both fixed. This is the ranking-paragraph-goes-stale trap (R-405) in miniature, and it appeared within minutes of my writing about it.