docs: CAMPAIGN 7 catalog sweep — audit, ROADMAP R-40/41/42, REPORT convention

Adds the CAMPAIGN 7 audit doc (53/53 apps deployed and validated on the demo box
through the real controller endpoints; 45 pass, 13 template fixes, ranked
findings) and the uninstall-semantics map row it proves out: remove requires stop
first, named docker volumes are ALWAYS destroyed including the DB, HDD data and
backups survive unless explicitly ticked, and the per-app offsite toggle outlives
the app while tier-2 config is cleared.

ROADMAP +3:
  R-40 the update path cannot express a multi-hop major upgrade (nextcloud
       forbids skipping majors; the template moved 31 -> 34)
  R-41 no standing 'does every template still deploy?' check -- the sweep was
       the first thing to try, and found 5 apps that had NEVER been deployable
  R-42 ruling needed on DB/cache sidecar majors (deliberately not bumped)

CLAUDE.md: REPORT.md is overwritten, so a second concurrent session in this repo
writes REPORT-<topic>.md instead -- applied to this run's own report.

Campaign 7 catalog sweep.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
This commit is contained in:
2026-07-19 03:24:39 +02:00
parent f5ab70d6e6
commit eee2c85321
4 changed files with 428 additions and 0 deletions
+3
View File
@@ -73,6 +73,9 @@ pushes; **you (Claude Code) implement**. A file being open in the editor is NOT
> **In every repository where you make a change, update both files in that repo:**
> - **`CHANGELOG.md`** — cumulative log, newest on top (here: per-area `hub/`, `scripts/`, `website/`).
> - **`REPORT.md`** — **overwrite** with the most recent implementation/validation summary only.
> **Parallel sessions:** `REPORT.md` is overwritten, so two sessions working in this repo at once
> will clobber each other. The second session writes **`REPORT-<topic>.md`** instead and never
> touches the shared `REPORT.md`.
>
> **Never write secrets** into any committed file — reference them as "stored out-of-band".
+43
View File
@@ -0,0 +1,43 @@
# REPORT — CAMPAIGN 7 (felhom.eu side: docs only)
> Written as `REPORT-campaign7.md`, **not** the shared `REPORT.md`, per the convention this run
> added to `CLAUDE.md`: `REPORT.md` is overwritten, so a second concurrent session in this repo
> would clobber it. This session's implementation work was in `app-catalog-felhom.eu`; here it only
> touched documentation.
**Run:** 2026-07-18 evening → 2026-07-19 morning. **Class:** campaign (record-and-rank + a defined
allowed-fix set). **Implementation repo:** `app-catalog-felhom.eu` (see its `REPORT.md`).
## What changed in this repo
| file | change |
|---|---|
| `documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md` | **new** — method, uninstall-semantics map, trio detail, full 53-app matrix, ranked findings, coverage |
| `documentation/backlog/ROADMAP.md` | **+3 items** — R-40 (multi-hop major upgrade path), R-41 (no standing catalog deployability check), R-42 (sidecar-major ruling) |
| `CLAUDE.md` | REPORT.md parallel-session rule: the second session writes `REPORT-<topic>.md` |
No hub/agent/scripts/website code was touched (campaign scope: catalog + docs).
## Headline for this repo's readers
1. **Uninstall semantics map row PARTIAL → PROVEN** (campaign doc §2), with live evidence from all
three trio apps: remove requires stop first; named docker volumes are **always destroyed**
(including the app's database); HDD bind-mount data and `backups/primary/<app>` survive unless
explicitly ticked; images are kept; `app.yaml` goes, the template stays; the per-app **offsite
toggle survives** the uninstall while tier-2 config is cleared. The confirmation modal does warn
about the volumes, so there is **no consent gap**.
2. **A lying healthcheck takes an app OFF-LINE, it does not merely mislead.** Traefik will not route
to an `unhealthy` container, so a probe that cannot execute → permanent unhealthy → **404 to the
customer while the app serves 200 on its own port**. 7 of 53 apps were in that state.
3. **The pre-flight gate's own signal is missing:** the 0.145.0 → 0.146.0 floor-lift emitted no
`controller_updated` event, though the identical bootstrap path emitted one for 0.143.0 → 0.145.0
two hours earlier (§0, finding F1). The box did converge — golden, floor and runtime all agreed —
but the event trail under-reports version transitions.
## Open items owned outside this repo
- **plant-it / wanderer** — images do not resolve at all (neither the new tag nor the one the
catalog already ships). Upstream research needed; recorded as findings, not deletions.
- **gokapi** — pinned back to v1.9.6; v2 needs the seeded `config.json` regenerated. Security-
relevant, should not sit on a superseded line indefinitely.
- **glance** — never had a seeded `glance.yml`; proven pre-existing.
@@ -0,0 +1,379 @@
# CAMPAIGN 7 — full app-catalog sweep (bump · deploy · validate · clean)
**Run:** 2026-07-18 evening → 2026-07-19 morning (overnight, hard wall-clock 06:30 CEST)
**Repos:** `app-catalog-felhom.eu` (the work), `felhom.eu/documentation` (this doc)
**Box:** demo (guest 9201 `demo-felhom` on `felhom-pve`), controller **0.146.0**
**Class:** Campaign — record-and-rank, with a defined allowed-fix set (the fixes are the deliverable)
---
## 0. Pre-flight gate — PASSED
Both of Viktor's saves had landed before the sweep began:
| check | value |
|---|---|
| `artifact_golden_version` | **0.146.0**, sha `4834c703…e955` ✓ |
| `min_controller_version` (floor) | **0.146.0** ✓ (saved last, as required) |
| guest 9201 running | `felhom-controller:0.146.0`, healthy ✓ |
| box reporting | `controller_started (0.146.0)` @ 18:52:07 ✓ |
So the sweep validated on the version customers actually run.
> **Finding C7-F1 (observability, MEDIUM) — the floor-lift update emitted no
> `controller_updated` event.** The gate asked for `controller_updated 0.145.0→0.146.0`
> in the events. It is **absent**, although 0.145.0→0.146.0 went through the *same*
> agent-driven bootstrap path (`/etc/felhom-bootstrap/bootstrap.json` →
> `felhom-controller-bootstrap.service`) that DID emit the event for 0.143.0→0.145.0 two
> hours earlier. `controller_started (0.146.0)` was emitted normally.
> Consequence: the operator-visible event trail under-reports version transitions, so
> "did the box converge?" cannot be answered from the events alone. Not a blocker — the
> runtime version, the golden record and the floor all agreed — but it is the one signal
> the gate was written around.
---
## 1. Method
- Every deploy/undeploy went through the controller's **real endpoints** — the ones the UI
calls: `POST /api/stacks/<app>/deploy`, `POST /api/stacks/<app>/stop`,
`POST /api/stacks/<app>/remove`, `POST /api/sync`. No raw `docker compose` against a
managed stack, no hand-edited `app.yaml`.
- Deploy fields auto-filled from `deploy-fields` metadata (DOMAIN / SUBDOMAIN / HDD_PATH);
secrets minted per the `.felhom.yml` `generate:` spec. **No secret value was ever logged
or written to evidence** — evidence records the field NAME and `<generated-secret, not
recorded>`.
- One campaign app deployed at a time. Evidence per app under
`180:~/campaign7/evidence/<app>/` (containers, healthcheck audit, per-container logs,
traefik view, log scan).
- Validation asserted **effect**, not absence-of-error: terminal health verdict per
container, healthcheck-binary audit, HTTP probe through the real Traefik ingress, login
where feasible, log scan with judgment for benign startup noise.
### 1.1 Three engine bugs found and fixed mid-run (they would have corrupted the matrix)
Recorded because they are exactly the "hollow validation" the testing doctrine warns about
— each one made a broken thing look fine, or a fine thing look broken:
1. **Traefik registration race → false 404s.** The first probe fired the instant a
container reported healthy, but Traefik registers the router a few seconds later.
actualbudget and calcom were both recorded FAIL and were actually fine (200 / 307).
Fixed with a retry window; both re-validated green.
2. **`docker exec` writes its OCI error to STDOUT, not stderr.** The healthcheck audit
checked "did `command -v <tool>` print anything" — so a *missing* binary printed
`executable file not found` and read as **present**. This made the entire healthcheck
audit — the campaign's core deliverable — report every app "honest". Fixed to require
rc==0 **and** an absolute path, with a direct-exec fallback for shell-less images.
3. **`created` sampled as a terminal state.** Containers still starting were recorded as
settled, so docmost and ghost were flagged FAIL while healthy moments later. Fixed to
require `running`; both re-validated green.
---
## 2. Uninstall semantics — the real delete flow (map row: PARTIAL → **PROVEN**)
Live evidence from the customer trio (`POST /api/stacks/<app>/remove`, both checkboxes
default-off, which is what the UI sends):
| what | behaviour | evidence |
|---|---|---|
| stack still running | **refused**`409 "still running — stop it first"`; the flow is stop→remove | bookstack |
| named docker volumes | **ALWAYS DESTROYED** (`compose down --volumes`), incl. the app's database | bookstack ×2, calibre ×1, immich ×3 |
| HDD bind-mount data (`userdata/`, `appdata/`) | **PRESERVED** unless `remove_hdd_data=true` | calibre books 440K, immich 54M — byte-identical after |
| HDD backup dirs (`backups/primary/<app>`) | **PRESERVED** unless `remove_backups=true` | calibre 24K, immich 28K |
| docker images | **KEPT** (for redeploy) | all three |
| `app.yaml` (deploy config) | **REMOVED** | all three |
| `docker-compose.yml` (template) | **KEPT** | all three |
| per-app **offsite** toggle (`app_backup.<app>.offbox`) | **SURVIVES** the uninstall | all three still `offbox:true` after removal |
| per-app **tier-2 / cross-drive** config | **CLEARED** (`SetCrossDriveConfig(name, nil)`) | `removeStack` |
The confirmation modal is honest about the destructive part — it states *"Mindig törlődik:
Docker kötetek (adatbázis, alkalmazás konfiguráció)"* before the customer confirms — so
there is **no consent gap**. Two narrower issues:
> **Finding C7-F2 (evidence trail, LOW) — `volumes_removed` is always `null`.**
> The remove response reported `volumes_removed: null` while actually destroying
> `bookstack_bookstack_config` + `bookstack_bookstack_db_data` (and 3 immich volumes).
> `delete.go` step 3 scrapes `compose down` stdout for `"Removing volume"` / `"Volume"`,
> which current Docker Compose no longer prints in that shape. The modal pre-warns, so this
> is not a safety issue — but the removal receipt is useless as evidence, and any future
> "what did we delete?" audit built on it would silently return nothing.
> **Finding C7-F3 (asymmetry, LOW) — offsite toggle outlives the app.** Removing an app
> clears its tier-2 schedule but leaves `app_backup.<app>.offbox = true` set forever. Benign
> today (a re-deploy inherits the customer's prior intent, which is arguably right, and it
> is why the trio's toggles needed no restoring). Worth an explicit ruling: is a removed
> app's offsite intent meant to persist, or should removal clear it like it clears tier-2?
---
## 3. The customer trio — uninstall → bump → fresh redeploy (Viktor's ruling), done FIRST
All three travelled the full path and are **left RUNNING** as the end state.
| app | pin(s) | MAJOR | deploy | health | http | login | logs | settle |
|---|---|---|---|---|---|---|---|---|
| **bookstack** | `25.02.2 → 26.05.2`; mariadb `11.6 → 12.3` | ✅ ×2 | ok | ok (2/2 healthy) | 302 | **ok** | clean | 20s |
| **calibre-web** | `v4.0.6` (already newest stable) | — | ok | ok | 302 | **ok** | noisy-benign | 20s |
| **immich** | `v2.5.5 → v3.0.3`; postgres `16-vectorchord0.3.0 → 16-vectorchord0.4.3-pgvectors0.2.0` | ✅ | ok | ok (4/4 healthy) | 200 | page-only | clean | 20s |
**Recorded pre-uninstall offsite state — and it contradicted the campaign note.** The note
said "bookstack ON, calibre ON, immich OFF". The box said:
```
bookstack enabled=false offbox=true
calibre-web enabled=false offbox=true
immich enabled=false offbox=true <-- note said OFF
```
Per the rail ("trust the recorded pre-uninstall state over this note") all three were
restored to `offbox:true` — which they already were, since the toggle survives uninstall
(§2). Verified post-redeploy.
**Detail per app**
- **bookstack** — MAJOR ×2. Upstream v26.05 needs `storage/fonts` writable for PDF export and
makes revision-viewing a separate permission; neither bites a fresh deploy. Logged in with
the documented `admin@admin.com / password` → 302 + 6 authenticated markers on the
dashboard. mariadb 12.3 healthy via `healthcheck.sh`.
- **calibre-web** — already at the newest STABLE tag; everything newer upstream is `dev-*`
(correctly excluded). Logged in with `admin / admin123` (Flask CSRF token required) →
title renders `Calibre-Web Automated | Books (0)`. The empty library is expected: the DB
volume was destroyed while the HDD book files were preserved, so the library needs
re-importing — the accepted consequence. The traceback in its log is **benign**: calibre's
installer fails headless `xdg-desktop-menu` setup, catches it, reports "There were 1
warnings", and continues.
- **immich** — MAJOR. v3.0.0 drops pgvecto.rs and requires VectorChord; our pin was already
VectorChord so the fresh deploy was unaffected. Sidecar moved to the extension versions
immich v3.0.3 ships in its own compose (vectorchord 0.4.3 / pgvectors 0.2.0) while
**keeping our PG major 16** rather than following upstream down to 14 — a `16-` build of
that exact extension pair is published, so no needless major change. Proven functional,
not merely healthy: `/api/server/ping``pong`, `/api/server/version``{3,0,3}`,
122 migration/init log lines, "Immich Microservices is running [v3.0.3]".
`login: page-only` is deliberate — the fresh install sits at first-run admin signup
(`isInitialized:false`, signup page 200) and creating an admin would mint a credential
that would then have to be transmitted to Viktor out-of-band; the owner creates his own.
- **redis vs valkey (recorded, not fixed):** immich v3 upstream migrated `redis`
`valkey:9`. We kept `redis:7-alpine` and **it works fine on v3** (all 4 containers
healthy). Swapping the image is a structural change, not a pin bump, so it is recorded
for a considered follow-up rather than done here.
---
## 4. Version-bump policy actually applied (deviation, stated deliberately)
App images: newest STABLE upstream tag, rc/beta/nightly/dev excluded, digest-free pinned
tags per house style.
**DB/cache sidecar majors were deliberately NOT bumped** (postgres 16→18, redis 7→8,
mariadb→12.3 on apps other than bookstack, postgis 16→17). This is a conscious deviation
from a literal reading of "newest stable for every pin":
- a DB major is a **data-plane decision the application owns**, not a currency decision —
immich proves it, upstream pins one specific tested postgres build;
- tags like `postgres:16-alpine` / `redis:7-alpine` already track the newest patch inside
their major, so they are not stale;
- blind-bumping ~8 apps onto PG18 would have manufactured failures that are artifacts of
the sweep's own choice rather than real findings, and would have burned the wall-clock the
rail explicitly told us to protect.
→ carried to ROADMAP as a decision item, not silently skipped. See §7.
---
## 5. Healthcheck audit — the wget lesson, systematically
Every deployed app's compose healthcheck was parsed and the probe binary checked **inside
the image**.
> **Finding C7-F4 (HIGH) — a lying healthcheck does not merely mislead; it takes the app
> OFF-LINE.** Traefik will not route to a container in an `unhealthy` state. So a probe that
> ENOENTs → permanent `unhealthy` → **Traefik returns 404 to the customer while the app is
> serving 200 perfectly well on its own port**. This was not theoretical: both apps below
> were completely unreachable for that reason, and both looked "deployed" in docker ps.
| app | lie | reality | fix | verified |
|---|---|---|---|---|
| **adventurelog** (frontend) | `wget --spider` | distroless image: no shell, no wget, no curl; node only at `/nodejs/bin/node`, off PATH | Node-exec family via the absolute interpreter path | 404 → **200**, 3/3 healthy |
| **emby** | `curl -f` | no curl, no standalone wget — image ships only BusyBox v1.38 | BusyBox-wget family via `/bin/busybox wget` | unhealthy/404 → **healthy/302** |
Both fixed in-template and live-re-validated on the demo box.
---
## 6. Result matrix
**All 53 catalog apps were attempted.** `http` is through the real Traefik ingress; a 3xx to a login
page is a PASS. `login` was attempted where credentials are documented and the flow is scriptable.
### 6.1 Passed (45)
| app | old → new pin | MAJOR | deploy | health | http | logs | settle |
|---|---|---|---|---|---|---|---|
| actualbudget | 26.1.0 → 26.7.0 | | ok | ok | 200 | clean | 10s |
| adventurelog | v0.11.0 → v0.12.1 | | ok | **fixed** | 200 | clean | 100s |
| audiobookshelf | 2.19.5 → 2.35.1 | | ok | ok | 200 | clean | 30s |
| bentopdf | v2.8.6 (newest) | | ok | ok | 200 | clean | 20s |
| **bookstack** ¹ | 25.02.2 → 26.05.2 · mariadb 11.6 → 12.3 | ✅ | ok | ok | 302 | clean | 20s |
| calcom | v4.6.9 → v6.2.0 | ✅ | ok | ok | 307 | clean | 20s |
| **calibre-web** ¹ | v4.0.6 (newest) | | ok | ok | 302 | noisy-benign | 20s |
| claper | 1.8 → 2.5 | ✅ | ok | ok | 200 | clean | 70s |
| code-server | 4.96.4 → 4.129.0 | | ok | ok | 302 | clean | 40s |
| crafty-controller | 4.10.7 (GitLab registry, manual) | | ok | ok | 302 | clean | 50s |
| docmost | 0.25.3 → 0.95.0 | | ok | ok | 200 | clean | 20s |
| emby | 4.9.0.42 → 4.10.0.20 | | ok | **fixed** | 302 | clean | 10s |
| ghost | 6.19.2-alpine → 6.53.0-alpine | | ok | ok | 200 | clean | 10s |
| gitea | 1.23.4 → 1.27.0 | | ok | ok | 200 | clean | 10s |
| grafana | 11.5.1 → 13.1.0 | ✅ | ok | ok | 302 | clean | 20s |
| gramps-web | v24.12.1 → v25.6.0 | ✅ | ok | **fixed** (mem) | 200 | clean | 110s |
| home-assistant | 2026.2.2 → 2026.7.2 | | ok | ok | 302 | clean | 70s |
| homebox | v0.16.3 → 0.26.2 | | ok | **fixed** ×3 | 200 | clean | — |
| homepage | v1.2.0 → v1.13.2 | | ok | ok | 200 | clean | 20s |
| **immich** ¹ | v2.5.5 → v3.0.3 · pg → 16-vectorchord0.4.3 | ✅ | ok | ok | 200 | clean | 20s |
| jellyfin | 10.11.6 → 10.11.11 | | ok | ok | 302 | clean | 30s |
| kimai | apache-2.25.0 → apache-2.57.0 | | ok | ok | 302 | clean | 50s |
| komga | 1.20.0 → 1.25.0 | | ok | ok | 200 | clean | 20s |
| mealie | v3.10.2 → v3.20.1 | | ok | ok | 200 | clean | 40s |
| n8n | 1.79.3 → 2.31.3 | ✅ | ok | **fixed** (mem) | 200 | clean | 50s |
| navidrome | 0.54.5 → 0.63.2 | | ok | ok | 302 | clean | 10s |
| nextcloud | 31.0.14-apache → 34.0.1-apache | ✅ | ok | ok | 302 | clean | 50s |
| onlyoffice | 8.3.0 → 9.4.0 | ✅ | ok | ok | 302 | clean | 40s |
| opengist | 1.10 → 1.13 | | ok | ok | 302 | clean | 10s |
| outline | 0.82.0 → 1.9.1 | ✅ | ok | **fixed** (env) | 200 | clean | 40s |
| paperless-ngx | 2.15.3 → 2.20.15 | | ok | ok | 302 | clean | 70s |
| papra | 26.6.1-rootless (newest) | | ok | **fixed** ×2 | 200 | clean | — |
| privatebin | 1.7.5 → 2.0.5 | ✅ | ok | ok | 200 | clean | 10s |
| radarr | 5.17.2 → 6.3.0 | ✅ | ok | ok | 200 | clean | 10s |
| rallly | 3.11.2 → 4.11.1 | ✅ | ok | **fixed** (mem) | 200 | clean | 20s |
| recipe-importer | v0.9.11 (ours, internal registry) | | ok | ok | 302 | clean | 10s |
| romm | 4.5.0 → 5.0.0 | ✅ | ok | ok | 200 | clean | 50s |
| seerr | 2.3.0 → 2.7.3 | | ok | ok | 307 | clean | 30s |
| sonarr | 4.0.13 → 4.0.19 | | ok | ok | 200 | clean | 10s |
| sparkyfitness | v0.17.2 → v0.17.3 (server + web) | | ok | ok | 200 | clean | 40s |
| tandoor | 1.5.26 → 2.6.13 | ✅ | ok | **fixed** ×4 | 302 | clean | 40s |
| termix | 2.5.0 (newest) | | ok | ok | 200 | clean | 20s |
| uptime-kuma | `:2` → 2.4.0 (**floating tag pinned**) | | ok | ok | 302 | clean | 30s |
| vaultwarden | 1.33.2-alpine → 1.36.0-alpine | | ok | ok | 200 | clean | 10s |
| wger | 2.3 → **2.6** (2.3 is gone upstream) | | ok | **fixed** ×3 | 302 | clean | 60s |
| vikunja | 0.24.6 → 2.3.0 | ✅ | ok | ok ⚠ no healthcheck | 200 | clean | 0s |
| wishlist | Hub 1.9.0 → **ghcr v0.66.0** | | ok | **fixed** ×2 | 200 | clean | — |
| zipline | 4.0.0 → 4.6.1 | | ok | **fixed** ×2 | 200 | clean | — |
¹ the customer trio — left RUNNING as the end state (see §3). Login: bookstack **ok**,
calibre-web **ok**, immich page-only (deliberate, §3). For the other 40 apps login was
`not-attempted` — a rendered login page through the real ingress was taken as PASS per the rail.
### 6.2 Not passing (4) + 1 not attempted
| app | outcome | cause | disposition |
|---|---|---|---|
| **glance** | FAIL — crash-loop | needs `/app/config/glance.yml`; the template mounts an EMPTY config volume and never seeds one | **pre-existing** — proven: v0.7.4 (the pre-campaign pin) fails identically. Bump KEPT, finding raised |
| **gokapi** | FAIL — crash-loop | v2.2.4 refuses to run against the seeded ConfigVersion-21 config: *"Please update to version 2.0.0 before running this version"* | **pin REVERTED to v1.9.6** (last known-good); v2 migration needs the seeded config regenerated |
| **plant-it** | FAIL — image | `msdeluise/plant-it` does not resolve on Docker Hub — **neither 1.0.1 nor the shipped 0.10.0** | **pre-existing**; no replacement registry found. Needs upstream research |
| **wanderer** | FAIL — image | `ghcr.io/flomp/wanderer` does not resolve — **neither 0.20.0 nor the shipped 0.16.0** | **pre-existing**; meilisearch sidecar bumped v1.12 → v1.49. Needs upstream research |
| **plex** | not attempted | `PLEX_CLAIM` is a required field with no default — a real claim token from plex.tv is needed | not automatable; genuinely owner-supplied. Not a defect |
---
## 7. Findings, ranked
**F4 (HIGH) — a lying healthcheck takes the app OFF-LINE, it does not merely mislead.**
Traefik refuses to route to a container in `unhealthy` state, so an ENOENT'ing probe →
permanent `unhealthy`**404 for the customer while the app serves 200 on its own port**.
Seven apps were affected. Detail in §5 and F5 below.
**F5 (HIGH) — 7 of 53 apps shipped a broken or wrong healthcheck.** All fixed and live-re-validated:
| app | the lie | reality | fix |
|---|---|---|---|
| adventurelog (frontend) | `wget --spider` | distroless: no shell/wget/curl; node only at an absolute path | Node-exec via `/nodejs/bin/node` |
| emby | `curl -f` | no curl, no standalone wget — BusyBox only | `/bin/busybox wget` |
| papra | `wget --spider` | image ships only `node` | Node-exec |
| wishlist | `wget --spider` | image ships only `node` | Node-exec |
| homebox | `wget --spider` (**HEAD**) | endpoint answers **405 to HEAD, 200 to GET** | `wget -q -O /dev/null` (GET) |
| zipline | `GET /api/health` | v4 renamed it — `/api/health` 404, `/api/healthcheck` 200 | corrected path |
| tandoor | `start_period: 30s` | gunicorn still booting; probes exhausted at ~2 min | `start_period: 240s` |
**F6 (HIGH) — 5 apps were ALREADY undeployable before this campaign.** None was caused by the
sweep; the sweep is simply the first thing that ever tried to deploy them:
`glance` (no seeded config), `papra` (no `AUTH_SECRET`), `zipline` (v4 `DATABASE_URL` rename),
`wishlist` (dead Docker Hub image), `plant-it` + `wanderer` (images do not resolve at all).
**Four are now fixed; plant-it and wanderer need upstream research.**
→ The catalog had no standing "does every template still deploy?" check. That is the real gap.
**F7 (HIGH, systemic) — the update path cannot express a multi-hop major upgrade.**
Nextcloud states plainly: *"You cannot skip major releases."* This campaign moved its template
31 → 34. A fresh deploy is fine (validated, 302), but an EXISTING customer's update button would
attempt 31 → 34 in one step, which Nextcloud forbids. Same shape for any app with sequential-major
rules. The template pin is a single value with no notion of an upgrade path.
→ ROADMAP item; nextcloud is the sharpest case but not the only one.
**F8 (MEDIUM, security-relevant) — gokapi is parked on a superseded version line.**
The revert to v1.9.6 restores the pre-campaign state rather than introducing a new regression, but
gokapi cannot reach v2 until the seeded `config.json` is regenerated in the v2 format. It should not
sit on v1 indefinitely — this wants a dedicated task, not a backlog line.
**F1 (MEDIUM) — the floor-lift controller update emitted no `controller_updated` event.** §0.
**F9 (MEDIUM) — an "obvious" env fix nearly orphaned customer data.** While chasing wger 2.6 the
interim fix pinned `DJANGO_DB_DATABASE=/home/wger/db/database.sqlite`. On 2.3 that path came from
wger's own default; hard-coding a different explicit path would have pointed an existing customer's
wger at an **empty** database while looking perfectly healthy. Reverted deliberately along with the
pin. Worth remembering as a class: *adding an explicit path for a value that previously defaulted is
a data-location change, not a config tidy-up.*
**F10 (MEDIUM) — 24 templates probe with `wget --spider`, which issues HEAD.** homebox proved the
trap (405 to HEAD, 200 to GET). The other 23 validated green, so their endpoints do answer HEAD —
but the default choice is fragile and silently costs availability when an app tightens its methods.
→ convention note for REUSE.md: prefer a real GET (`wget -q -O /dev/null`) unless HEAD is verified.
**F11 (LOW) — vikunja ships with NO healthcheck at all.** It serves 200 and Traefik routes it
(no health state to filter on), so it works — but it has no liveness signal, and the controller-side
probe is the only thing watching it. Not fixed: adding one was not "trivial-and-certain" within the
allowed set.
**F2 (LOW) — the remove receipt never lists destroyed volumes** (`volumes_removed: null`). §2.
**F3 (LOW) — a removed app keeps its offsite toggle forever.** §2. Needs a ruling.
**F12 (LOW) — `.felhom.yml` `mem_limit` can drift from the compose sum.** tandoor claimed 512M
while its services summed to 768M. Only found because tandoor OOM'd. A mechanical gate would catch
the whole catalog at once (the repo already has `check-image-pins.py` as the pattern to copy).
### 7.1 Sweep-engine bugs (recorded because they are the "hollow validation" failure mode)
Three bugs in the campaign's own harness would have written a **false matrix** — two made broken
things look fine, one made fine things look broken. All fixed mid-run and the affected apps
re-validated (§1.1): the Traefik registration race (false 404s), `docker exec` writing its OCI error
to **stdout** (which turned the entire healthcheck audit green), and sampling `created` as a
terminal state.
---
## 8. Coverage
**53 of 53 catalog apps attempted — full coverage. No resumable remainder.**
- **45 passed** end-to-end through the real pipeline (deploy → healthy → HTTP through Traefik → log
scan → removed via the real delete flow).
- **3 left RUNNING** as the end state (the customer trio), on current versions, offsite toggles
restored, both scripted logins green.
- **4 failed**, each with a diagnosed root cause and a disposition (§6.2); 2 of those are pinned
back to their last known-good version rather than shipping a broken pin.
- **1 not attempted** (plex — needs a real `PLEX_CLAIM` token; not a defect).
- **13 template fixes** committed, every one live-re-validated on the demo box.
- **Alert e-mails observed:** none fired during the sweep window. The pre-existing
`kimaradtak: bookstack` offsite warning (from the 17:14 run, before the campaign) was still on
record at the start and is expected per the rail — not chased.
### 8.1 Not done / explicitly out of scope
- **DB/cache sidecar majors** (postgres 16→18, redis 7→8, mariadb→12, postgis 16→17) —
deliberately not bumped, rationale in §4. **ROADMAP decision item.**
- **MAJOR breaking notes** were retrieved from upstream for bookstack, immich, nextcloud, n8n and
grafana. For the remaining MAJOR rows the note was **not retrieved** within the wall-clock; they
are flagged MAJOR without an upstream one-liner rather than given a fabricated one.
- **immich redis → valkey** (upstream migrated in v3; redis:7-alpine works fine on v3) — recorded,
not done: an image swap is structural, not a pin bump.
- **glance config seeding**, **gokapi v2 config migration**, **plant-it / wanderer upstream
research** — all outside the allowed-fix set, each needs its own task.
+3
View File
@@ -79,6 +79,9 @@
| R-37 | **Post-RESET health card shows stale pre-RESET warnings.** After a RESET the card should read **„RESET óta nincs adat"** instead of carrying warnings about a lifecycle that no longer exists. | XS | **SHIPPED (hub v0.67.0, 2026-07-18)** | The customer page raises a banner when a RESET **completed** after the newest report, quoting „RESET óta nincs adat" and the reset timestamp, because until the box reports again every health figure describes a lifecycle that no longer exists. Deliberately narrow: an **in-flight** reset does not trigger it (only a completed one), and it **clears itself** on the first post-RESET report. Ties resolve to STALE — SQLite timestamps are second-resolution and a same-second report almost certainly arrived just before the reset destroyed what it describes; erring the other way would hide the banner exactly when it matters most. Red-proofed (neutering the predicate fails the assertion). — Origin: 2026-07-18 rehearsal. Same family as R-36 — the hub knows the state changed and the UI has not caught up |
| R-38 | **Installer GRUB slice.** A single default „Felhom telepítés" entry; the **interactive installers REMOVED** (safety: an interactive entry is how a wrong-disk manual install happens); felhom background. | S | idea | Origin: 2026-07-18 rehearsal, alongside R-21's physical closure. **Squashfs/theme rebranding explicitly DEFERRED** — this item is the menu and the safety, not a skin |
| R-40 | **[P2-HIGH] The update path cannot express a MULTI-HOP major upgrade.** A template pin is a single value; the customer's update button pulls whatever the catalog now says. For apps whose upstream forbids version skipping this produces a broken upgrade. Nextcloud is explicit: *"You cannot skip major releases. Please re-run the upgrade until you have reached the highest available release."* Campaign 7 moved its template **31 → 34** (a fresh deploy validates fine — 302, 3/3 healthy), so an existing 31 customer pressing update would attempt a jump Nextcloud refuses. | M | idea | Origin: CAMPAIGN 7 (`audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md` §7 F7). Not nextcloud-only — any app with sequential-major rules (gitea, tandoor, outline…) has the same shape. Directions: a per-app `upgrade_path:`/`max_hop:` in `.felhom.yml` that the update button walks in stages; or refuse-and-explain when the installed major is >1 behind; or pin an intermediate "stepping-stone" tag. **Until this exists, a >1-major catalog bump is safe for NEW deploys and unsafe for the update button** — which is exactly the asymmetry the campaign's MAJOR flag was meant to record but cannot enforce |
| R-41 | **[P2-HIGH] The catalog has no standing "does every template still deploy?" check.** Campaign 7 was the first thing that ever tried to deploy all 53 apps, and found **5 that had NEVER been deployable**: papra (missing required `AUTH_SECRET`), zipline (v4 renamed `CORE_DATABASE_URL``DATABASE_URL`), wishlist (Docker Hub image gone; upstream moved to ghcr.io), homebox (upstream dropped the `v` tag prefix + new required env), glance (needs a seeded `glance.yml` the template never provides — PROVEN pre-existing: the pre-campaign v0.7.4 pin fails identically). Plus **7 broken healthchecks** and 2 apps whose images no longer resolve at all (plant-it, wanderer). | M | idea | Origin: CAMPAIGN 7 (§7 F5/F6). The repo already has the right pattern in `scripts/check-image-pins.py` — a mechanical gate run on every change. Cheap first slice: a **resolvability gate** (`docker manifest inspect` every pin) would alone have caught plant-it, wanderer, wishlist and homebox, and needs no box. Full slice: a periodic deploy-all sweep on the demo box reusing the campaign's engine. **Silent rot is the real risk** — an app can die upstream and nobody learns until a customer clicks Telepítés |
| R-42 | **Ruling needed: do DB/cache sidecar majors follow the app, or the newest tag?** Campaign 7 deliberately did NOT bump sidecar majors (postgres 16→18, redis 7→8, mariadb 11.6→12.3, postgis 16→17) while bumping ~40 app images to current. | S | **decision pending (Viktor)** | Origin: CAMPAIGN 7 §4. The case for not bumping: a DB major is a **data-plane decision the application owns** — immich proves it, upstream pins one specific tested `postgres:14-vectorchord…` build — and `postgres:16-alpine`/`redis:7-alpine` already track the newest patch inside their major, so they are not stale. The case for bumping: EOL majors eventually stop getting security patches, and "we never bump" silently becomes "we ship EOL databases". Suggested shape: per-app sidecar pin follows **upstream's own compose** where upstream publishes one, else stay within the current major and revisit at that major's EOL date |
## Pre-invite checklist — what stands between here and the first remote tester