SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s

THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
This commit is contained in:
2026-09-01 21:35:32 +02:00
parent ac079b8c43
commit 56c7e373a3
43 changed files with 2026 additions and 287 deletions
@@ -99,7 +99,10 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| Scenario | Components | Status | Evidence | Gap / roadmap |
|---|---|---|---|---|
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3` | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. What they do to app DATA is now measured too, and it is a separate row-worth of facts (below).** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3`; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller | **PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
@@ -335,7 +335,7 @@ own; every caller that is not the customer must decide for itself whether the ap
### `sync/`
| File | Class | Reason | Risk |
|---|---|---|---|
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). | clean |
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). **It copies into EVERY stack folder, deployed or not — see "the app-definition seam" below.** | clean |
### `system/` — split per-function (not per-file)
| File | Class | Reason | Risk |
@@ -503,3 +503,77 @@ own; every caller that is not the customer must decide for itself whether the ap
- **S6:** §5(3) self-restore-test → **status-display only**; the agent owns orchestration.
- **Self-update resolved (03 §11):** `updater.go` → **DELETE(→agent)**, `state.go` →
DELETE(obsolete), `version.go` KEEP; §6 + §5(2) updated (bulk = `backup=0` mountpoint recipe).
---
## The app-definition seam — what happens to a deployed app when its compose file is rewritten
> **STATUS: MEASURED BEHAVIOUR, NOT A RECORDED DESIGN.** This section describes what the system was
> observed to do on 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`), with controls. It is
> deliberately **not** marked `[DESIGN]`, because the operator has not ruled on whether the behaviour
> was intended. **Do not read this as an endorsement, and do not spec against it as though it were
> settled.** The open questions are R-438 and R-441.
Until this was measured, no architecture document said what happens here, and the gap itself is
R-438. The three facts below are the ones a reader needs before touching any of it.
### 1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle
`Syncer.copyTemplates` (`sync/sync.go:319`) walks every directory in the catalog cache and copies
`docker-compose.yml` and `.felhom.yml` into the matching stack folder. **There is no test of whether
the app is deployed.** The only guard is a sha256 content compare (`copyIfChanged`) and the only
exclusion is `app.yaml`. Interval is `git.sync_interval`, default `15m` (`config/config.go:351`), plus
one immediate sync at controller start (`sync.go:98`).
**It restarts nothing.** The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing
else. So from the moment it runs, a deployed app's *definition* and its *running containers* disagree,
and they stay that way until something else acts.
### 2. Every lifecycle path resolves that disagreement, silently, by upgrading
`StartStack`, `RestartStack` and `UpdateStack` (`stacks/manager.go:1029 / 1133 / 1170`) all end in
`docker compose up -d`. `up -d` makes the container match the file, **and pulls the image itself if it
is absent** — measured at 18.3 s with a pull versus 0.5 s without, against a negative control
(unchanged file) that did not even recreate the container.
`RestartStack` carries an explicit in-source comment saying this is deliberate — *"so that … any
template changes (new images, healthchecks) are picked up"*. **That is a fact about the source, not an
operator ruling**, and it speaks only for the customer-pressed restart.
**Thirteen non-API call sites across nine files** reach `up -d` without anyone pressing anything. The
full table is §8 of the spike doc; the three that matter most are:
- `bootrecon.Reconciler.Run` → `StartStack` — `bootrecon/bootrecon.go:269`
- `backup.AppStopGuard.Recover` → `StartStack` — `backup/appstop_marker.go:283`
- the drive-return gate, `web.Server.restartStacks` → `StartStack` — `web/intermediary.go:222`
**A plain power cut does NOT trigger this.** Docker's `restart: unless-stopped` restores the existing
containers on the old image, the reconciler finds no orphan, and it logs so
(`no boot-orphaned apps (nothing to start)`). The unattended upgrade needs the narrower precondition
*"and the app did not come back"* — which was measured, and does upgrade.
### 3. Nothing takes a copy first, and the tag cannot be put back afterwards
`UpdateStack` runs `compose pull` then `compose up -d --remove-orphans` and does nothing else — no
dump, no copy, no hold. The R-361 safety dump (`backup/offbox_reconstitute.go:207`) is **database-only**
and is not on this path at all; an app with no database gets nothing from it even on the paths where it
does run.
**And restoring the old tag is not a rollback.** Once an app has migrated its data, the old image
refuses to start — measured on Nextcloud: *"the version of the data (32.0.9.2) is higher than the
docker image version (31.0.14.1) and downgrading is not supported"*. The data itself survives; only the
downgrade is blocked. So the only route back is a **data** restore from a copy taken before the
update.
**And that route fights this seam** (R-441): `stackAdapter.RecreateStackDefinitionFromUnit`
(`cmd/controller/main.go:2570`) writes the recovery unit's captured compose — with the OLD pin — into
the live stack dir, and `copyIfChanged` overwrites it again on the next tick. The overwrite is
measured; that the restore writes to that path is read, not measured.
### What a change here must not break
- The syncer's overwrite is also the **repair** path: a locally-corrupted compose is replaced by the
catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that.
- `up -d`-on-restart is what injects `app.yaml` env into a running stack. Reverting to
`docker compose restart` would silently stop doing that.