diff --git a/CLAUDE.md b/CLAUDE.md index c667538..0aaecf1 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -73,6 +73,9 @@ pushes; **you (Claude Code) implement**. A file being open in the editor is NOT > **In every repository where you make a change, update both files in that repo:** > - **`CHANGELOG.md`** — cumulative log, newest on top (here: per-area `hub/`, `scripts/`, `website/`). > - **`REPORT.md`** — **overwrite** with the most recent implementation/validation summary only. +> **Parallel sessions:** `REPORT.md` is overwritten, so two sessions working in this repo at once +> will clobber each other. The second session writes **`REPORT-.md`** instead and never +> touches the shared `REPORT.md`. > > **Never write secrets** into any committed file — reference them as "stored out-of-band". diff --git a/REPORT-campaign7.md b/REPORT-campaign7.md new file mode 100644 index 0000000..7ef6f74 --- /dev/null +++ b/REPORT-campaign7.md @@ -0,0 +1,43 @@ +# REPORT — CAMPAIGN 7 (felhom.eu side: docs only) + +> Written as `REPORT-campaign7.md`, **not** the shared `REPORT.md`, per the convention this run +> added to `CLAUDE.md`: `REPORT.md` is overwritten, so a second concurrent session in this repo +> would clobber it. This session's implementation work was in `app-catalog-felhom.eu`; here it only +> touched documentation. + +**Run:** 2026-07-18 evening → 2026-07-19 morning. **Class:** campaign (record-and-rank + a defined +allowed-fix set). **Implementation repo:** `app-catalog-felhom.eu` (see its `REPORT.md`). + +## What changed in this repo + +| file | change | +|---|---| +| `documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md` | **new** — method, uninstall-semantics map, trio detail, full 53-app matrix, ranked findings, coverage | +| `documentation/backlog/ROADMAP.md` | **+3 items** — R-40 (multi-hop major upgrade path), R-41 (no standing catalog deployability check), R-42 (sidecar-major ruling) | +| `CLAUDE.md` | REPORT.md parallel-session rule: the second session writes `REPORT-.md` | + +No hub/agent/scripts/website code was touched (campaign scope: catalog + docs). + +## Headline for this repo's readers + +1. **Uninstall semantics map row PARTIAL → PROVEN** (campaign doc §2), with live evidence from all + three trio apps: remove requires stop first; named docker volumes are **always destroyed** + (including the app's database); HDD bind-mount data and `backups/primary/` survive unless + explicitly ticked; images are kept; `app.yaml` goes, the template stays; the per-app **offsite + toggle survives** the uninstall while tier-2 config is cleared. The confirmation modal does warn + about the volumes, so there is **no consent gap**. +2. **A lying healthcheck takes an app OFF-LINE, it does not merely mislead.** Traefik will not route + to an `unhealthy` container, so a probe that cannot execute → permanent unhealthy → **404 to the + customer while the app serves 200 on its own port**. 7 of 53 apps were in that state. +3. **The pre-flight gate's own signal is missing:** the 0.145.0 → 0.146.0 floor-lift emitted no + `controller_updated` event, though the identical bootstrap path emitted one for 0.143.0 → 0.145.0 + two hours earlier (§0, finding F1). The box did converge — golden, floor and runtime all agreed — + but the event trail under-reports version transitions. + +## Open items owned outside this repo + +- **plant-it / wanderer** — images do not resolve at all (neither the new tag nor the one the + catalog already ships). Upstream research needed; recorded as findings, not deletions. +- **gokapi** — pinned back to v1.9.6; v2 needs the seeded `config.json` regenerated. Security- + relevant, should not sit on a superseded line indefinitely. +- **glance** — never had a seeded `glance.yml`; proven pre-existing. diff --git a/documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md b/documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md new file mode 100644 index 0000000..4da0ac1 --- /dev/null +++ b/documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md @@ -0,0 +1,379 @@ +# CAMPAIGN 7 — full app-catalog sweep (bump · deploy · validate · clean) + +**Run:** 2026-07-18 evening → 2026-07-19 morning (overnight, hard wall-clock 06:30 CEST) +**Repos:** `app-catalog-felhom.eu` (the work), `felhom.eu/documentation` (this doc) +**Box:** demo (guest 9201 `demo-felhom` on `felhom-pve`), controller **0.146.0** +**Class:** Campaign — record-and-rank, with a defined allowed-fix set (the fixes are the deliverable) + +--- + +## 0. Pre-flight gate — PASSED + +Both of Viktor's saves had landed before the sweep began: + +| check | value | +|---|---| +| `artifact_golden_version` | **0.146.0**, sha `4834c703…e955` ✓ | +| `min_controller_version` (floor) | **0.146.0** ✓ (saved last, as required) | +| guest 9201 running | `felhom-controller:0.146.0`, healthy ✓ | +| box reporting | `controller_started (0.146.0)` @ 18:52:07 ✓ | + +So the sweep validated on the version customers actually run. + +> **Finding C7-F1 (observability, MEDIUM) — the floor-lift update emitted no +> `controller_updated` event.** The gate asked for `controller_updated 0.145.0→0.146.0` +> in the events. It is **absent**, although 0.145.0→0.146.0 went through the *same* +> agent-driven bootstrap path (`/etc/felhom-bootstrap/bootstrap.json` → +> `felhom-controller-bootstrap.service`) that DID emit the event for 0.143.0→0.145.0 two +> hours earlier. `controller_started (0.146.0)` was emitted normally. +> Consequence: the operator-visible event trail under-reports version transitions, so +> "did the box converge?" cannot be answered from the events alone. Not a blocker — the +> runtime version, the golden record and the floor all agreed — but it is the one signal +> the gate was written around. + +--- + +## 1. Method + +- Every deploy/undeploy went through the controller's **real endpoints** — the ones the UI + calls: `POST /api/stacks//deploy`, `POST /api/stacks//stop`, + `POST /api/stacks//remove`, `POST /api/sync`. No raw `docker compose` against a + managed stack, no hand-edited `app.yaml`. +- Deploy fields auto-filled from `deploy-fields` metadata (DOMAIN / SUBDOMAIN / HDD_PATH); + secrets minted per the `.felhom.yml` `generate:` spec. **No secret value was ever logged + or written to evidence** — evidence records the field NAME and ``. +- One campaign app deployed at a time. Evidence per app under + `180:~/campaign7/evidence//` (containers, healthcheck audit, per-container logs, + traefik view, log scan). +- Validation asserted **effect**, not absence-of-error: terminal health verdict per + container, healthcheck-binary audit, HTTP probe through the real Traefik ingress, login + where feasible, log scan with judgment for benign startup noise. + +### 1.1 Three engine bugs found and fixed mid-run (they would have corrupted the matrix) + +Recorded because they are exactly the "hollow validation" the testing doctrine warns about +— each one made a broken thing look fine, or a fine thing look broken: + +1. **Traefik registration race → false 404s.** The first probe fired the instant a + container reported healthy, but Traefik registers the router a few seconds later. + actualbudget and calcom were both recorded FAIL and were actually fine (200 / 307). + Fixed with a retry window; both re-validated green. +2. **`docker exec` writes its OCI error to STDOUT, not stderr.** The healthcheck audit + checked "did `command -v ` print anything" — so a *missing* binary printed + `executable file not found` and read as **present**. This made the entire healthcheck + audit — the campaign's core deliverable — report every app "honest". Fixed to require + rc==0 **and** an absolute path, with a direct-exec fallback for shell-less images. +3. **`created` sampled as a terminal state.** Containers still starting were recorded as + settled, so docmost and ghost were flagged FAIL while healthy moments later. Fixed to + require `running`; both re-validated green. + +--- + +## 2. Uninstall semantics — the real delete flow (map row: PARTIAL → **PROVEN**) + +Live evidence from the customer trio (`POST /api/stacks//remove`, both checkboxes +default-off, which is what the UI sends): + +| what | behaviour | evidence | +|---|---|---| +| stack still running | **refused** — `409 "still running — stop it first"`; the flow is stop→remove | bookstack | +| named docker volumes | **ALWAYS DESTROYED** (`compose down --volumes`), incl. the app's database | bookstack ×2, calibre ×1, immich ×3 | +| HDD bind-mount data (`userdata/`, `appdata/`) | **PRESERVED** unless `remove_hdd_data=true` | calibre books 440K, immich 54M — byte-identical after | +| HDD backup dirs (`backups/primary/`) | **PRESERVED** unless `remove_backups=true` | calibre 24K, immich 28K | +| docker images | **KEPT** (for redeploy) | all three | +| `app.yaml` (deploy config) | **REMOVED** | all three | +| `docker-compose.yml` (template) | **KEPT** | all three | +| per-app **offsite** toggle (`app_backup..offbox`) | **SURVIVES** the uninstall | all three still `offbox:true` after removal | +| per-app **tier-2 / cross-drive** config | **CLEARED** (`SetCrossDriveConfig(name, nil)`) | `removeStack` | + +The confirmation modal is honest about the destructive part — it states *"Mindig törlődik: +Docker kötetek (adatbázis, alkalmazás konfiguráció)"* before the customer confirms — so +there is **no consent gap**. Two narrower issues: + +> **Finding C7-F2 (evidence trail, LOW) — `volumes_removed` is always `null`.** +> The remove response reported `volumes_removed: null` while actually destroying +> `bookstack_bookstack_config` + `bookstack_bookstack_db_data` (and 3 immich volumes). +> `delete.go` step 3 scrapes `compose down` stdout for `"Removing volume"` / `"Volume"`, +> which current Docker Compose no longer prints in that shape. The modal pre-warns, so this +> is not a safety issue — but the removal receipt is useless as evidence, and any future +> "what did we delete?" audit built on it would silently return nothing. + +> **Finding C7-F3 (asymmetry, LOW) — offsite toggle outlives the app.** Removing an app +> clears its tier-2 schedule but leaves `app_backup..offbox = true` set forever. Benign +> today (a re-deploy inherits the customer's prior intent, which is arguably right, and it +> is why the trio's toggles needed no restoring). Worth an explicit ruling: is a removed +> app's offsite intent meant to persist, or should removal clear it like it clears tier-2? + +--- + +## 3. The customer trio — uninstall → bump → fresh redeploy (Viktor's ruling), done FIRST + +All three travelled the full path and are **left RUNNING** as the end state. + +| app | pin(s) | MAJOR | deploy | health | http | login | logs | settle | +|---|---|---|---|---|---|---|---|---| +| **bookstack** | `25.02.2 → 26.05.2`; mariadb `11.6 → 12.3` | ✅ ×2 | ok | ok (2/2 healthy) | 302 | **ok** | clean | 20s | +| **calibre-web** | `v4.0.6` (already newest stable) | — | ok | ok | 302 | **ok** | noisy-benign | 20s | +| **immich** | `v2.5.5 → v3.0.3`; postgres `16-vectorchord0.3.0 → 16-vectorchord0.4.3-pgvectors0.2.0` | ✅ | ok | ok (4/4 healthy) | 200 | page-only | clean | 20s | + +**Recorded pre-uninstall offsite state — and it contradicted the campaign note.** The note +said "bookstack ON, calibre ON, immich OFF". The box said: + +``` +bookstack enabled=false offbox=true +calibre-web enabled=false offbox=true +immich enabled=false offbox=true <-- note said OFF +``` + +Per the rail ("trust the recorded pre-uninstall state over this note") all three were +restored to `offbox:true` — which they already were, since the toggle survives uninstall +(§2). Verified post-redeploy. + +**Detail per app** + +- **bookstack** — MAJOR ×2. Upstream v26.05 needs `storage/fonts` writable for PDF export and + makes revision-viewing a separate permission; neither bites a fresh deploy. Logged in with + the documented `admin@admin.com / password` → 302 + 6 authenticated markers on the + dashboard. mariadb 12.3 healthy via `healthcheck.sh`. +- **calibre-web** — already at the newest STABLE tag; everything newer upstream is `dev-*` + (correctly excluded). Logged in with `admin / admin123` (Flask CSRF token required) → + title renders `Calibre-Web Automated | Books (0)`. The empty library is expected: the DB + volume was destroyed while the HDD book files were preserved, so the library needs + re-importing — the accepted consequence. The traceback in its log is **benign**: calibre's + installer fails headless `xdg-desktop-menu` setup, catches it, reports "There were 1 + warnings", and continues. +- **immich** — MAJOR. v3.0.0 drops pgvecto.rs and requires VectorChord; our pin was already + VectorChord so the fresh deploy was unaffected. Sidecar moved to the extension versions + immich v3.0.3 ships in its own compose (vectorchord 0.4.3 / pgvectors 0.2.0) while + **keeping our PG major 16** rather than following upstream down to 14 — a `16-` build of + that exact extension pair is published, so no needless major change. Proven functional, + not merely healthy: `/api/server/ping` → `pong`, `/api/server/version` → `{3,0,3}`, + 122 migration/init log lines, "Immich Microservices is running [v3.0.3]". + `login: page-only` is deliberate — the fresh install sits at first-run admin signup + (`isInitialized:false`, signup page 200) and creating an admin would mint a credential + that would then have to be transmitted to Viktor out-of-band; the owner creates his own. +- **redis vs valkey (recorded, not fixed):** immich v3 upstream migrated `redis` → + `valkey:9`. We kept `redis:7-alpine` and **it works fine on v3** (all 4 containers + healthy). Swapping the image is a structural change, not a pin bump, so it is recorded + for a considered follow-up rather than done here. + +--- + +## 4. Version-bump policy actually applied (deviation, stated deliberately) + +App images: newest STABLE upstream tag, rc/beta/nightly/dev excluded, digest-free pinned +tags per house style. + +**DB/cache sidecar majors were deliberately NOT bumped** (postgres 16→18, redis 7→8, +mariadb→12.3 on apps other than bookstack, postgis 16→17). This is a conscious deviation +from a literal reading of "newest stable for every pin": + +- a DB major is a **data-plane decision the application owns**, not a currency decision — + immich proves it, upstream pins one specific tested postgres build; +- tags like `postgres:16-alpine` / `redis:7-alpine` already track the newest patch inside + their major, so they are not stale; +- blind-bumping ~8 apps onto PG18 would have manufactured failures that are artifacts of + the sweep's own choice rather than real findings, and would have burned the wall-clock the + rail explicitly told us to protect. + +→ carried to ROADMAP as a decision item, not silently skipped. See §7. + +--- + +## 5. Healthcheck audit — the wget lesson, systematically + +Every deployed app's compose healthcheck was parsed and the probe binary checked **inside +the image**. + +> **Finding C7-F4 (HIGH) — a lying healthcheck does not merely mislead; it takes the app +> OFF-LINE.** Traefik will not route to a container in an `unhealthy` state. So a probe that +> ENOENTs → permanent `unhealthy` → **Traefik returns 404 to the customer while the app is +> serving 200 perfectly well on its own port**. This was not theoretical: both apps below +> were completely unreachable for that reason, and both looked "deployed" in docker ps. + +| app | lie | reality | fix | verified | +|---|---|---|---|---| +| **adventurelog** (frontend) | `wget --spider` | distroless image: no shell, no wget, no curl; node only at `/nodejs/bin/node`, off PATH | Node-exec family via the absolute interpreter path | 404 → **200**, 3/3 healthy | +| **emby** | `curl -f` | no curl, no standalone wget — image ships only BusyBox v1.38 | BusyBox-wget family via `/bin/busybox wget` | unhealthy/404 → **healthy/302** | + +Both fixed in-template and live-re-validated on the demo box. + +--- + +## 6. Result matrix + +**All 53 catalog apps were attempted.** `http` is through the real Traefik ingress; a 3xx to a login +page is a PASS. `login` was attempted where credentials are documented and the flow is scriptable. + +### 6.1 Passed (45) + +| app | old → new pin | MAJOR | deploy | health | http | logs | settle | +|---|---|---|---|---|---|---|---| +| actualbudget | 26.1.0 → 26.7.0 | | ok | ok | 200 | clean | 10s | +| adventurelog | v0.11.0 → v0.12.1 | | ok | **fixed** | 200 | clean | 100s | +| audiobookshelf | 2.19.5 → 2.35.1 | | ok | ok | 200 | clean | 30s | +| bentopdf | v2.8.6 (newest) | | ok | ok | 200 | clean | 20s | +| **bookstack** ¹ | 25.02.2 → 26.05.2 · mariadb 11.6 → 12.3 | ✅ | ok | ok | 302 | clean | 20s | +| calcom | v4.6.9 → v6.2.0 | ✅ | ok | ok | 307 | clean | 20s | +| **calibre-web** ¹ | v4.0.6 (newest) | | ok | ok | 302 | noisy-benign | 20s | +| claper | 1.8 → 2.5 | ✅ | ok | ok | 200 | clean | 70s | +| code-server | 4.96.4 → 4.129.0 | | ok | ok | 302 | clean | 40s | +| crafty-controller | 4.10.7 (GitLab registry, manual) | | ok | ok | 302 | clean | 50s | +| docmost | 0.25.3 → 0.95.0 | | ok | ok | 200 | clean | 20s | +| emby | 4.9.0.42 → 4.10.0.20 | | ok | **fixed** | 302 | clean | 10s | +| ghost | 6.19.2-alpine → 6.53.0-alpine | | ok | ok | 200 | clean | 10s | +| gitea | 1.23.4 → 1.27.0 | | ok | ok | 200 | clean | 10s | +| grafana | 11.5.1 → 13.1.0 | ✅ | ok | ok | 302 | clean | 20s | +| gramps-web | v24.12.1 → v25.6.0 | ✅ | ok | **fixed** (mem) | 200 | clean | 110s | +| home-assistant | 2026.2.2 → 2026.7.2 | | ok | ok | 302 | clean | 70s | +| homebox | v0.16.3 → 0.26.2 | | ok | **fixed** ×3 | 200 | clean | — | +| homepage | v1.2.0 → v1.13.2 | | ok | ok | 200 | clean | 20s | +| **immich** ¹ | v2.5.5 → v3.0.3 · pg → 16-vectorchord0.4.3 | ✅ | ok | ok | 200 | clean | 20s | +| jellyfin | 10.11.6 → 10.11.11 | | ok | ok | 302 | clean | 30s | +| kimai | apache-2.25.0 → apache-2.57.0 | | ok | ok | 302 | clean | 50s | +| komga | 1.20.0 → 1.25.0 | | ok | ok | 200 | clean | 20s | +| mealie | v3.10.2 → v3.20.1 | | ok | ok | 200 | clean | 40s | +| n8n | 1.79.3 → 2.31.3 | ✅ | ok | **fixed** (mem) | 200 | clean | 50s | +| navidrome | 0.54.5 → 0.63.2 | | ok | ok | 302 | clean | 10s | +| nextcloud | 31.0.14-apache → 34.0.1-apache | ✅ | ok | ok | 302 | clean | 50s | +| onlyoffice | 8.3.0 → 9.4.0 | ✅ | ok | ok | 302 | clean | 40s | +| opengist | 1.10 → 1.13 | | ok | ok | 302 | clean | 10s | +| outline | 0.82.0 → 1.9.1 | ✅ | ok | **fixed** (env) | 200 | clean | 40s | +| paperless-ngx | 2.15.3 → 2.20.15 | | ok | ok | 302 | clean | 70s | +| papra | 26.6.1-rootless (newest) | | ok | **fixed** ×2 | 200 | clean | — | +| privatebin | 1.7.5 → 2.0.5 | ✅ | ok | ok | 200 | clean | 10s | +| radarr | 5.17.2 → 6.3.0 | ✅ | ok | ok | 200 | clean | 10s | +| rallly | 3.11.2 → 4.11.1 | ✅ | ok | **fixed** (mem) | 200 | clean | 20s | +| recipe-importer | v0.9.11 (ours, internal registry) | | ok | ok | 302 | clean | 10s | +| romm | 4.5.0 → 5.0.0 | ✅ | ok | ok | 200 | clean | 50s | +| seerr | 2.3.0 → 2.7.3 | | ok | ok | 307 | clean | 30s | +| sonarr | 4.0.13 → 4.0.19 | | ok | ok | 200 | clean | 10s | +| sparkyfitness | v0.17.2 → v0.17.3 (server + web) | | ok | ok | 200 | clean | 40s | +| tandoor | 1.5.26 → 2.6.13 | ✅ | ok | **fixed** ×4 | 302 | clean | 40s | +| termix | 2.5.0 (newest) | | ok | ok | 200 | clean | 20s | +| uptime-kuma | `:2` → 2.4.0 (**floating tag pinned**) | | ok | ok | 302 | clean | 30s | +| vaultwarden | 1.33.2-alpine → 1.36.0-alpine | | ok | ok | 200 | clean | 10s | +| wger | 2.3 → **2.6** (2.3 is gone upstream) | | ok | **fixed** ×3 | 302 | clean | 60s | +| vikunja | 0.24.6 → 2.3.0 | ✅ | ok | ok ⚠ no healthcheck | 200 | clean | 0s | +| wishlist | Hub 1.9.0 → **ghcr v0.66.0** | | ok | **fixed** ×2 | 200 | clean | — | +| zipline | 4.0.0 → 4.6.1 | | ok | **fixed** ×2 | 200 | clean | — | + +¹ the customer trio — left RUNNING as the end state (see §3). Login: bookstack **ok**, +calibre-web **ok**, immich page-only (deliberate, §3). For the other 40 apps login was +`not-attempted` — a rendered login page through the real ingress was taken as PASS per the rail. + +### 6.2 Not passing (4) + 1 not attempted + +| app | outcome | cause | disposition | +|---|---|---|---| +| **glance** | FAIL — crash-loop | needs `/app/config/glance.yml`; the template mounts an EMPTY config volume and never seeds one | **pre-existing** — proven: v0.7.4 (the pre-campaign pin) fails identically. Bump KEPT, finding raised | +| **gokapi** | FAIL — crash-loop | v2.2.4 refuses to run against the seeded ConfigVersion-21 config: *"Please update to version 2.0.0 before running this version"* | **pin REVERTED to v1.9.6** (last known-good); v2 migration needs the seeded config regenerated | +| **plant-it** | FAIL — image | `msdeluise/plant-it` does not resolve on Docker Hub — **neither 1.0.1 nor the shipped 0.10.0** | **pre-existing**; no replacement registry found. Needs upstream research | +| **wanderer** | FAIL — image | `ghcr.io/flomp/wanderer` does not resolve — **neither 0.20.0 nor the shipped 0.16.0** | **pre-existing**; meilisearch sidecar bumped v1.12 → v1.49. Needs upstream research | +| **plex** | not attempted | `PLEX_CLAIM` is a required field with no default — a real claim token from plex.tv is needed | not automatable; genuinely owner-supplied. Not a defect | + +--- + +## 7. Findings, ranked + +**F4 (HIGH) — a lying healthcheck takes the app OFF-LINE, it does not merely mislead.** +Traefik refuses to route to a container in `unhealthy` state, so an ENOENT'ing probe → +permanent `unhealthy` → **404 for the customer while the app serves 200 on its own port**. +Seven apps were affected. Detail in §5 and F5 below. + +**F5 (HIGH) — 7 of 53 apps shipped a broken or wrong healthcheck.** All fixed and live-re-validated: + +| app | the lie | reality | fix | +|---|---|---|---| +| adventurelog (frontend) | `wget --spider` | distroless: no shell/wget/curl; node only at an absolute path | Node-exec via `/nodejs/bin/node` | +| emby | `curl -f` | no curl, no standalone wget — BusyBox only | `/bin/busybox wget` | +| papra | `wget --spider` | image ships only `node` | Node-exec | +| wishlist | `wget --spider` | image ships only `node` | Node-exec | +| homebox | `wget --spider` (**HEAD**) | endpoint answers **405 to HEAD, 200 to GET** | `wget -q -O /dev/null` (GET) | +| zipline | `GET /api/health` | v4 renamed it — `/api/health` 404, `/api/healthcheck` 200 | corrected path | +| tandoor | `start_period: 30s` | gunicorn still booting; probes exhausted at ~2 min | `start_period: 240s` | + +**F6 (HIGH) — 5 apps were ALREADY undeployable before this campaign.** None was caused by the +sweep; the sweep is simply the first thing that ever tried to deploy them: +`glance` (no seeded config), `papra` (no `AUTH_SECRET`), `zipline` (v4 `DATABASE_URL` rename), +`wishlist` (dead Docker Hub image), `plant-it` + `wanderer` (images do not resolve at all). +**Four are now fixed; plant-it and wanderer need upstream research.** +→ The catalog had no standing "does every template still deploy?" check. That is the real gap. + +**F7 (HIGH, systemic) — the update path cannot express a multi-hop major upgrade.** +Nextcloud states plainly: *"You cannot skip major releases."* This campaign moved its template +31 → 34. A fresh deploy is fine (validated, 302), but an EXISTING customer's update button would +attempt 31 → 34 in one step, which Nextcloud forbids. Same shape for any app with sequential-major +rules. The template pin is a single value with no notion of an upgrade path. +→ ROADMAP item; nextcloud is the sharpest case but not the only one. + +**F8 (MEDIUM, security-relevant) — gokapi is parked on a superseded version line.** +The revert to v1.9.6 restores the pre-campaign state rather than introducing a new regression, but +gokapi cannot reach v2 until the seeded `config.json` is regenerated in the v2 format. It should not +sit on v1 indefinitely — this wants a dedicated task, not a backlog line. + +**F1 (MEDIUM) — the floor-lift controller update emitted no `controller_updated` event.** §0. + +**F9 (MEDIUM) — an "obvious" env fix nearly orphaned customer data.** While chasing wger 2.6 the +interim fix pinned `DJANGO_DB_DATABASE=/home/wger/db/database.sqlite`. On 2.3 that path came from +wger's own default; hard-coding a different explicit path would have pointed an existing customer's +wger at an **empty** database while looking perfectly healthy. Reverted deliberately along with the +pin. Worth remembering as a class: *adding an explicit path for a value that previously defaulted is +a data-location change, not a config tidy-up.* + +**F10 (MEDIUM) — 24 templates probe with `wget --spider`, which issues HEAD.** homebox proved the +trap (405 to HEAD, 200 to GET). The other 23 validated green, so their endpoints do answer HEAD — +but the default choice is fragile and silently costs availability when an app tightens its methods. +→ convention note for REUSE.md: prefer a real GET (`wget -q -O /dev/null`) unless HEAD is verified. + +**F11 (LOW) — vikunja ships with NO healthcheck at all.** It serves 200 and Traefik routes it +(no health state to filter on), so it works — but it has no liveness signal, and the controller-side +probe is the only thing watching it. Not fixed: adding one was not "trivial-and-certain" within the +allowed set. + +**F2 (LOW) — the remove receipt never lists destroyed volumes** (`volumes_removed: null`). §2. + +**F3 (LOW) — a removed app keeps its offsite toggle forever.** §2. Needs a ruling. + +**F12 (LOW) — `.felhom.yml` `mem_limit` can drift from the compose sum.** tandoor claimed 512M +while its services summed to 768M. Only found because tandoor OOM'd. A mechanical gate would catch +the whole catalog at once (the repo already has `check-image-pins.py` as the pattern to copy). + +### 7.1 Sweep-engine bugs (recorded because they are the "hollow validation" failure mode) + +Three bugs in the campaign's own harness would have written a **false matrix** — two made broken +things look fine, one made fine things look broken. All fixed mid-run and the affected apps +re-validated (§1.1): the Traefik registration race (false 404s), `docker exec` writing its OCI error +to **stdout** (which turned the entire healthcheck audit green), and sampling `created` as a +terminal state. + +--- + +## 8. Coverage + +**53 of 53 catalog apps attempted — full coverage. No resumable remainder.** + +- **45 passed** end-to-end through the real pipeline (deploy → healthy → HTTP through Traefik → log + scan → removed via the real delete flow). +- **3 left RUNNING** as the end state (the customer trio), on current versions, offsite toggles + restored, both scripted logins green. +- **4 failed**, each with a diagnosed root cause and a disposition (§6.2); 2 of those are pinned + back to their last known-good version rather than shipping a broken pin. +- **1 not attempted** (plex — needs a real `PLEX_CLAIM` token; not a defect). +- **13 template fixes** committed, every one live-re-validated on the demo box. +- **Alert e-mails observed:** none fired during the sweep window. The pre-existing + `kimaradtak: bookstack` offsite warning (from the 17:14 run, before the campaign) was still on + record at the start and is expected per the rail — not chased. + +### 8.1 Not done / explicitly out of scope + +- **DB/cache sidecar majors** (postgres 16→18, redis 7→8, mariadb→12, postgis 16→17) — + deliberately not bumped, rationale in §4. **ROADMAP decision item.** +- **MAJOR breaking notes** were retrieved from upstream for bookstack, immich, nextcloud, n8n and + grafana. For the remaining MAJOR rows the note was **not retrieved** within the wall-clock; they + are flagged MAJOR without an upstream one-liner rather than given a fabricated one. +- **immich redis → valkey** (upstream migrated in v3; redis:7-alpine works fine on v3) — recorded, + not done: an image swap is structural, not a pin bump. +- **glance config seeding**, **gokapi v2 config migration**, **plant-it / wanderer upstream + research** — all outside the allowed-fix set, each needs its own task. diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index c97c0d5..06f8238 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -79,6 +79,9 @@ | R-37 | **Post-RESET health card shows stale pre-RESET warnings.** After a RESET the card should read **„RESET óta nincs adat"** instead of carrying warnings about a lifecycle that no longer exists. | XS | **SHIPPED (hub v0.67.0, 2026-07-18)** | The customer page raises a banner when a RESET **completed** after the newest report, quoting „RESET óta nincs adat" and the reset timestamp, because until the box reports again every health figure describes a lifecycle that no longer exists. Deliberately narrow: an **in-flight** reset does not trigger it (only a completed one), and it **clears itself** on the first post-RESET report. Ties resolve to STALE — SQLite timestamps are second-resolution and a same-second report almost certainly arrived just before the reset destroyed what it describes; erring the other way would hide the banner exactly when it matters most. Red-proofed (neutering the predicate fails the assertion). — Origin: 2026-07-18 rehearsal. Same family as R-36 — the hub knows the state changed and the UI has not caught up | | R-38 | **Installer GRUB slice.** A single default „Felhom telepítés" entry; the **interactive installers REMOVED** (safety: an interactive entry is how a wrong-disk manual install happens); felhom background. | S | idea | Origin: 2026-07-18 rehearsal, alongside R-21's physical closure. **Squashfs/theme rebranding explicitly DEFERRED** — this item is the menu and the safety, not a skin | +| R-40 | **[P2-HIGH] The update path cannot express a MULTI-HOP major upgrade.** A template pin is a single value; the customer's update button pulls whatever the catalog now says. For apps whose upstream forbids version skipping this produces a broken upgrade. Nextcloud is explicit: *"You cannot skip major releases. Please re-run the upgrade until you have reached the highest available release."* Campaign 7 moved its template **31 → 34** (a fresh deploy validates fine — 302, 3/3 healthy), so an existing 31 customer pressing update would attempt a jump Nextcloud refuses. | M | idea | Origin: CAMPAIGN 7 (`audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md` §7 F7). Not nextcloud-only — any app with sequential-major rules (gitea, tandoor, outline…) has the same shape. Directions: a per-app `upgrade_path:`/`max_hop:` in `.felhom.yml` that the update button walks in stages; or refuse-and-explain when the installed major is >1 behind; or pin an intermediate "stepping-stone" tag. **Until this exists, a >1-major catalog bump is safe for NEW deploys and unsafe for the update button** — which is exactly the asymmetry the campaign's MAJOR flag was meant to record but cannot enforce | +| R-41 | **[P2-HIGH] The catalog has no standing "does every template still deploy?" check.** Campaign 7 was the first thing that ever tried to deploy all 53 apps, and found **5 that had NEVER been deployable**: papra (missing required `AUTH_SECRET`), zipline (v4 renamed `CORE_DATABASE_URL` → `DATABASE_URL`), wishlist (Docker Hub image gone; upstream moved to ghcr.io), homebox (upstream dropped the `v` tag prefix + new required env), glance (needs a seeded `glance.yml` the template never provides — PROVEN pre-existing: the pre-campaign v0.7.4 pin fails identically). Plus **7 broken healthchecks** and 2 apps whose images no longer resolve at all (plant-it, wanderer). | M | idea | Origin: CAMPAIGN 7 (§7 F5/F6). The repo already has the right pattern in `scripts/check-image-pins.py` — a mechanical gate run on every change. Cheap first slice: a **resolvability gate** (`docker manifest inspect` every pin) would alone have caught plant-it, wanderer, wishlist and homebox, and needs no box. Full slice: a periodic deploy-all sweep on the demo box reusing the campaign's engine. **Silent rot is the real risk** — an app can die upstream and nobody learns until a customer clicks Telepítés | +| R-42 | **Ruling needed: do DB/cache sidecar majors follow the app, or the newest tag?** Campaign 7 deliberately did NOT bump sidecar majors (postgres 16→18, redis 7→8, mariadb 11.6→12.3, postgis 16→17) while bumping ~40 app images to current. | S | **decision pending (Viktor)** | Origin: CAMPAIGN 7 §4. The case for not bumping: a DB major is a **data-plane decision the application owns** — immich proves it, upstream pins one specific tested `postgres:14-vectorchord…` build — and `postgres:16-alpine`/`redis:7-alpine` already track the newest patch inside their major, so they are not stale. The case for bumping: EOL majors eventually stop getting security patches, and "we never bump" silently becomes "we ship EOL databases". Suggested shape: per-app sidecar pin follows **upstream's own compose** where upstream publishes one, else stay within the current major and revisit at that major's EOL date | ## Pre-invite checklist — what stands between here and the first remote tester