Files
felhom.eu/REPORT.md
T
admin 56c7e373a3
gates / gates (push) Successful in 18s
SPIKE: what an app update actually does, and which other paths do it too
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
2026-09-01 21:35:32 +02:00

215 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01)
**Task class: Spike.** No production code was written in any repo. Output is a findings doc, register
rows, two architecture updates and one operator decision.
**Findings doc:** `documentation/audits/SPIKE-app-update-2026-09-01.md`
**Evidence:** `documentation/audits/evidence-spike-app-update-2026-09-01/` (19 files)
---
## 1. The Phase 1 answer, stated unambiguously
**YES — `docker compose up -d` upgrades an app whose compose file has already moved, and the Restart
button does it.**
Produced by **variants 1a and 1b, and independently by 1c-ii**:
- **1a** (target image ABSENT locally): the restart took **18.3 s**, ended with the container on the
new tag, and **left the new image in the local store where it had not been** — so `up -d` performed
a network pull, in a code path that contains no pull step.
- **1b** (target image PRESENT): **0.5 s**, container recreated onto the new tag, image id unchanged.
- **1d, the negative control** (file NOT edited): container id, image id, digest and `StartedAt` all
identical — `up -d` did not even recreate. **The method can show "no change" when nothing changed.**
- **1c-ii** (unattended): the boot reconciler brought a failed app back on the **new** version with
nobody pressing anything.
**And the answer to 1c is NO, which narrows the exposure:** a hard guest reset upgraded nothing.
Docker's `restart: unless-stopped` restored the existing containers on the old image and the
reconciler logged its own verdict — `Boot reconciliation: no boot-orphaned apps (nothing to start)`.
**Per the task's own rule, no fix is proposed.** The behaviour may have been chosen — `RestartStack`
says so in a comment — so it goes to the operator as a decision.
---
## 2. Confirmed baselines actually used
| Repo | at §1 of the task | measured at start | end | moved? |
|---|---|---|---|---|
| felhom-controller | `960d29b0612c` | `960d29b0612c` | `960d29b0612c` | **no — read only** |
| felhom.eu | `1d59353df437` | `1d59353df437` | this session's documents | as planned |
| app-catalog-felhom.eu | `29edad9c5bf4` | `29edad9c5bf4` | `5d8f25f`, **tree identical to `29edad9c5bf4`** | reverted |
| felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched |
**None had moved since the task was written.** Live controller on demo-hp: 0.232.0.
---
## 3. Per-phase results, each with its control
| Phase | Result | Control |
|---|---|---|
| **0** | R-438, R-439, R-440 filed, each with rank and owner | ids confirmed free by grepping all three register files (R-437 was highest) |
| **1** | **THE GATE: YES.** 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended | **1d** — unchanged file, no change at all |
| **2** | Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer | positive (`BentoPDF`, 4 hits) + negative (`zzz-never-present`, 0) on both customer pages |
| **3a** | Pull failure: HTTP 500, **app untouched and still running** | re-run as a Restart: also 500, app still ran |
| **3b** | **HTTP 200 "update completed" over a crash-looping app**; alarm fired 5m16s later | the alarm is a positive observable (`Event pushed: app_start_failed`), plus the detector heartbeat `1 currently down` |
| **4** | demo-hp and demo-felhom: **zero drift by tag**. But **2 of 6 floating pins have already MOVED upstream** | both fully-pinned images (`romm:5.0.0`, `pdo:2.0.5`) were SAME |
| **5** | Existing safety dump is **database-only**; DB dumps are 48 KB–395 KB; the **file** half is the cost, and it does not fit for a large app | demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead |
| **6** | **App data CANNOT be rolled back.** Old image refuses to start on migrated data | returning to 32.0.9 restored **both** seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked |
---
## 4. The exact symbols that bring an app back — found by reading
| path | `file:symbol` |
|---|---|
| boot reconciler | `felhom-controller/controller/internal/bootrecon/bootrecon.go:269` — `Reconciler.Run` → `StartStack` |
| its scheduler | `controller/cmd/controller/main.go:2127` — `runBootReconcile`, called at `:450` |
| app-stop guard | `controller/internal/backup/appstop_marker.go:283` — `AppStopGuard.Recover` → `StartStack` |
| its hold-aware wrapper | `controller/cmd/controller/main.go:2008` — `gatedAppStopStarter.StartStack` |
| drive-return gate | `controller/internal/web/intermediary.go:222` — `Server.restartStacks` → `StartStack` |
| guest-boot change | `controller/internal/web/intermediary.go:458` — `Server.processGuestBootChange` |
| quiesce restart | `controller/internal/quiesce/quiesce.go:733` — `Loop.restartAll` |
| off-site reconstitution | `controller/internal/backup/offbox_reconstitute.go:692` — `Manager.ReconstituteFromOffsite` |
Full table of all 13 non-API call sites: findings doc §8.
---
## 5. Register rows opened, updated or re-ranked
| row | what | owner |
|---|---|---|
| **R-438** | opened P1-HIGH, then **updated with the live evidence** (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) | **VIKTOR rules, CC measures** |
| **R-439** | opened P3-LOW, then **updated — the severity survives but its stated reason was imprecise** (`isOperationalState` counts `restarting`/`degraded` as operational) | CC |
| **R-440** | opened P2-MEDIUM, then **updated: two floating pins have ALREADY moved, with a passing control** | CC |
| **R-441** | NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes | CC measures, VIKTOR rules |
| **R-442** | NEW — **P1-HIGH**: `remove_hdd_data:true` is inert with no `paths.hdd_path`; 128 MB left, API said neither removed nor preserved | CC |
| **R-443** | NEW — the Update button reports success over an app it has broken | CC proposes, VIKTOR rules |
| **R-444** | NEW — nothing runs `pct fstrim`; demo-hp's thin pool held ~23.8 GB of freed blocks | CC |
| **R-445** | NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation | VIKTOR rules, CC implements |
**Nothing was closed and nothing was re-ranked.** The ranking of R-438 relative to existing rows is
Viktor's, and I have not moved anything.
---
## 6. Claims in the task that turned out to be wrong, named
1. **"a power cut … is an unattended three-major-version upgrade"** — not as stated. A plain power cut
upgraded nothing (measured). The unattended upgrade needs *"and the app did not come back"*.
**The exposure is smaller than the operator page claims.**
2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites, nine files.
3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and
`runbooks/target-selection.md:161` says *"No access route from DooPlex"*. Two boxes measured live;
Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history.
4. **R-439's severity reason** — right conclusion, imprecise reason (see §5).
5. **R-440's "23 pins"** — **exactly right** (79 image lines / 53 apps / 66 distinct; 23 with no patch
component). One arguable 24th named rather than rounded away.
6. **The catalog history figures** — spot-checked and **all correct**: 153 commits, 53 apps, 0 files
with upgrade metadata, and all four multi-major bumps confirmed by commit hash.
7. **My own method, corrected in-flight:** the first customer-page search used `grep -o "2.8.6"`, whose
unescaped `.` produced two false hits; `grep -F` gives zero. The controls caught it.
---
## 7. Evidence handling
**Evidence was written directly into `documentation/audits/evidence-spike-app-update-2026-09-01/` on
DooPlex as each phase produced it — before every revert, including the intermediate ones.** The
catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all
happened after their evidence was already off the machine. **Nothing was lost and nothing had to be
reproduced.**
---
## 8. Observations
1. **The restore path writes the recovery unit's OLD image pin into the live stack dir, and the
catalog syncer overwrites it again within 15 minutes.** The overwrite half is measured on demo-hp;
that the restore writes to that same path is read at `cmd/controller/main.go:2570`, not measured —
both halves are graded as such in the row. **FILED: R-441**
2. **`remove_hdd_data: true` removed nothing.** 128 MB of app data stayed on the drive while the API
returned 200 with `hdd_paths_removed:null, hdd_paths_preserved:null`. Root cause established with
controls: `Paths.HDDPath` has no default and demo-hp's `controller.yaml` does not set it, so
`ParseComposeHDDMounts` returns nil on its first line. **FILED: R-442**
3. **An update can report success over an app it has just broken.** HTTP 200 and "updated
successfully" while the container was already crash-looping; the truth reached the customer 5m16s
later through the dead-app alarm rather than through the update itself. **FILED: R-443**
4. **The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed**, and `fstrim` from
inside the unprivileged container is refused; `pct fstrim` from the host reclaimed it. Nothing runs
it on the fleet. **FILED: R-444**
5. **The hub kept app telemetry for an app that no longer exists anywhere**, and it now sets a
fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway.
Retained rather than cleared, because the reset is irreversible and on the operator's own surface.
**FILED: R-445**
6. **The syncer's debug hash line cannot show what it claims to show.** `logFileHashes`
(`internal/sync/sync.go:386`) reads the destination *after* the write, so it printed
`src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word
"changed". **NOT-A-FINDING:** it is DEBUG-only and the `Updated <app>/<file>` INFO line immediately
above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the
next person reading a sync log does not try to learn from those two hashes, which cannot differ.
---
## 9. Teardown — all three layers
**Layer 1 — the machine.** `bentopdf` restored to its catalog tag `v2.8.6`, digest
`sha256:eaeea1e447205a79…`, **byte-identical to the run's baseline**. The throwaway `nextcloud` stack
removed via the product's own endpoint; containers, all three volumes and `app.yaml` gone. The 128 MB
the product failed to remove (observation 2) deleted by hand along with its backup dirs;
`find /mnt -iname "*nextcloud*"` returns nothing. Five images this run pulled removed by **targeted
`docker rmi`** — **no `prune` of any kind was run anywhere**. No guest was created; guest 9201 was
hard-reset once by design and returned with all nine apps.
**Layer 2 — the host.** `local-lvm` 68.97% → 70.91% during the run → **26.78%** after `pct fstrim
9201`. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to
pre-run values (`/` 957 M, `/mnt/sys_drive` 12 G, `hdd_1` 5.5 G).
**Layer 3 — the hub. This run provisioned NOTHING.** No customer record and no appliance record was
created; the customers list is unchanged at five rows, identical to the list read at the start. The
existing `demo-hp` customer was used. What the run *did* create is hub **events** (`app_start_failed`,
plus deploy/remove for the throwaway app) — **retained deliberately**, because the event log is an
append-only record and deleting from it to tidy a test damages the surface this project relies on for
history. One residue is **retained rather than cleared** with the reason and the exact one-line command
recorded in the findings doc §13 (observation 5).
**Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified. Peti's box was never contacted.**
---
## 10. Final verification
```
felhom-controller: git status --porcelain → empty
HEAD = origin/main = 960d29b0612c
go build ./... → OK
go vet ./... → OK
go test ./... → rc=0, 28 packages, 0 FAIL
```
**The controller tree was left untouched, and that is proven rather than asserted.**
---
## 11. My own mistakes
1. **My first customer-page search used a regex where I needed a literal.** `grep -o "2.8.6"` treats
`.` as a wildcard and reported two hits on a page that contains none. I caught it only because the
task requires a control on every such search, and re-ran with `grep -F`. **The rule earned its
place in the same session it was applied.**
2. **My first poll for "nextcloud is healthy" matched the wrong container.** The break condition
matched `healthy` anywhere in the line and fired on `nextcloud-redis`. Corrected to an exact-name
filter. No result depended on it.
3. **I renumbered `STATUS.md` badly on the first attempt**, leaving the list running 1–6, 8, 9 with no
item 7, and left the section's own lead line saying one thing was waiting when there were two.
Both fixed. **This is the ranking-paragraph-goes-stale trap (R-405) in miniature**, and it appeared
within minutes of my writing about it.