Files
felhom.eu/documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md
T
admin eee2c85321 docs: CAMPAIGN 7 catalog sweep — audit, ROADMAP R-40/41/42, REPORT convention
Adds the CAMPAIGN 7 audit doc (53/53 apps deployed and validated on the demo box
through the real controller endpoints; 45 pass, 13 template fixes, ranked
findings) and the uninstall-semantics map row it proves out: remove requires stop
first, named docker volumes are ALWAYS destroyed including the DB, HDD data and
backups survive unless explicitly ticked, and the per-app offsite toggle outlives
the app while tier-2 config is cleared.

ROADMAP +3:
  R-40 the update path cannot express a multi-hop major upgrade (nextcloud
       forbids skipping majors; the template moved 31 -> 34)
  R-41 no standing 'does every template still deploy?' check -- the sweep was
       the first thing to try, and found 5 apps that had NEVER been deployable
  R-42 ruling needed on DB/cache sidecar majors (deliberately not bumped)

CLAUDE.md: REPORT.md is overwritten, so a second concurrent session in this repo
writes REPORT-<topic>.md instead -- applied to this run's own report.

Campaign 7 catalog sweep.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
2026-07-19 03:24:39 +02:00

380 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CAMPAIGN 7 — full app-catalog sweep (bump · deploy · validate · clean)
**Run:** 2026-07-18 evening → 2026-07-19 morning (overnight, hard wall-clock 06:30 CEST)
**Repos:** `app-catalog-felhom.eu` (the work), `felhom.eu/documentation` (this doc)
**Box:** demo (guest 9201 `demo-felhom` on `felhom-pve`), controller **0.146.0**
**Class:** Campaign — record-and-rank, with a defined allowed-fix set (the fixes are the deliverable)
---
## 0. Pre-flight gate — PASSED
Both of Viktor's saves had landed before the sweep began:
| check | value |
|---|---|
| `artifact_golden_version` | **0.146.0**, sha `4834c703…e955` ✓ |
| `min_controller_version` (floor) | **0.146.0** ✓ (saved last, as required) |
| guest 9201 running | `felhom-controller:0.146.0`, healthy ✓ |
| box reporting | `controller_started (0.146.0)` @ 18:52:07 ✓ |
So the sweep validated on the version customers actually run.
> **Finding C7-F1 (observability, MEDIUM) — the floor-lift update emitted no
> `controller_updated` event.** The gate asked for `controller_updated 0.145.0→0.146.0`
> in the events. It is **absent**, although 0.145.0→0.146.0 went through the *same*
> agent-driven bootstrap path (`/etc/felhom-bootstrap/bootstrap.json` →
> `felhom-controller-bootstrap.service`) that DID emit the event for 0.143.0→0.145.0 two
> hours earlier. `controller_started (0.146.0)` was emitted normally.
> Consequence: the operator-visible event trail under-reports version transitions, so
> "did the box converge?" cannot be answered from the events alone. Not a blocker — the
> runtime version, the golden record and the floor all agreed — but it is the one signal
> the gate was written around.
---
## 1. Method
- Every deploy/undeploy went through the controller's **real endpoints** — the ones the UI
calls: `POST /api/stacks/<app>/deploy`, `POST /api/stacks/<app>/stop`,
`POST /api/stacks/<app>/remove`, `POST /api/sync`. No raw `docker compose` against a
managed stack, no hand-edited `app.yaml`.
- Deploy fields auto-filled from `deploy-fields` metadata (DOMAIN / SUBDOMAIN / HDD_PATH);
secrets minted per the `.felhom.yml` `generate:` spec. **No secret value was ever logged
or written to evidence** — evidence records the field NAME and `<generated-secret, not
recorded>`.
- One campaign app deployed at a time. Evidence per app under
`180:~/campaign7/evidence/<app>/` (containers, healthcheck audit, per-container logs,
traefik view, log scan).
- Validation asserted **effect**, not absence-of-error: terminal health verdict per
container, healthcheck-binary audit, HTTP probe through the real Traefik ingress, login
where feasible, log scan with judgment for benign startup noise.
### 1.1 Three engine bugs found and fixed mid-run (they would have corrupted the matrix)
Recorded because they are exactly the "hollow validation" the testing doctrine warns about
— each one made a broken thing look fine, or a fine thing look broken:
1. **Traefik registration race → false 404s.** The first probe fired the instant a
container reported healthy, but Traefik registers the router a few seconds later.
actualbudget and calcom were both recorded FAIL and were actually fine (200 / 307).
Fixed with a retry window; both re-validated green.
2. **`docker exec` writes its OCI error to STDOUT, not stderr.** The healthcheck audit
checked "did `command -v <tool>` print anything" — so a *missing* binary printed
`executable file not found` and read as **present**. This made the entire healthcheck
audit — the campaign's core deliverable — report every app "honest". Fixed to require
rc==0 **and** an absolute path, with a direct-exec fallback for shell-less images.
3. **`created` sampled as a terminal state.** Containers still starting were recorded as
settled, so docmost and ghost were flagged FAIL while healthy moments later. Fixed to
require `running`; both re-validated green.
---
## 2. Uninstall semantics — the real delete flow (map row: PARTIAL → **PROVEN**)
Live evidence from the customer trio (`POST /api/stacks/<app>/remove`, both checkboxes
default-off, which is what the UI sends):
| what | behaviour | evidence |
|---|---|---|
| stack still running | **refused**`409 "still running — stop it first"`; the flow is stop→remove | bookstack |
| named docker volumes | **ALWAYS DESTROYED** (`compose down --volumes`), incl. the app's database | bookstack ×2, calibre ×1, immich ×3 |
| HDD bind-mount data (`userdata/`, `appdata/`) | **PRESERVED** unless `remove_hdd_data=true` | calibre books 440K, immich 54M — byte-identical after |
| HDD backup dirs (`backups/primary/<app>`) | **PRESERVED** unless `remove_backups=true` | calibre 24K, immich 28K |
| docker images | **KEPT** (for redeploy) | all three |
| `app.yaml` (deploy config) | **REMOVED** | all three |
| `docker-compose.yml` (template) | **KEPT** | all three |
| per-app **offsite** toggle (`app_backup.<app>.offbox`) | **SURVIVES** the uninstall | all three still `offbox:true` after removal |
| per-app **tier-2 / cross-drive** config | **CLEARED** (`SetCrossDriveConfig(name, nil)`) | `removeStack` |
The confirmation modal is honest about the destructive part — it states *"Mindig törlődik:
Docker kötetek (adatbázis, alkalmazás konfiguráció)"* before the customer confirms — so
there is **no consent gap**. Two narrower issues:
> **Finding C7-F2 (evidence trail, LOW) — `volumes_removed` is always `null`.**
> The remove response reported `volumes_removed: null` while actually destroying
> `bookstack_bookstack_config` + `bookstack_bookstack_db_data` (and 3 immich volumes).
> `delete.go` step 3 scrapes `compose down` stdout for `"Removing volume"` / `"Volume"`,
> which current Docker Compose no longer prints in that shape. The modal pre-warns, so this
> is not a safety issue — but the removal receipt is useless as evidence, and any future
> "what did we delete?" audit built on it would silently return nothing.
> **Finding C7-F3 (asymmetry, LOW) — offsite toggle outlives the app.** Removing an app
> clears its tier-2 schedule but leaves `app_backup.<app>.offbox = true` set forever. Benign
> today (a re-deploy inherits the customer's prior intent, which is arguably right, and it
> is why the trio's toggles needed no restoring). Worth an explicit ruling: is a removed
> app's offsite intent meant to persist, or should removal clear it like it clears tier-2?
---
## 3. The customer trio — uninstall → bump → fresh redeploy (Viktor's ruling), done FIRST
All three travelled the full path and are **left RUNNING** as the end state.
| app | pin(s) | MAJOR | deploy | health | http | login | logs | settle |
|---|---|---|---|---|---|---|---|---|
| **bookstack** | `25.02.2 → 26.05.2`; mariadb `11.6 → 12.3` | ✅ ×2 | ok | ok (2/2 healthy) | 302 | **ok** | clean | 20s |
| **calibre-web** | `v4.0.6` (already newest stable) | — | ok | ok | 302 | **ok** | noisy-benign | 20s |
| **immich** | `v2.5.5 → v3.0.3`; postgres `16-vectorchord0.3.0 → 16-vectorchord0.4.3-pgvectors0.2.0` | ✅ | ok | ok (4/4 healthy) | 200 | page-only | clean | 20s |
**Recorded pre-uninstall offsite state — and it contradicted the campaign note.** The note
said "bookstack ON, calibre ON, immich OFF". The box said:
```
bookstack enabled=false offbox=true
calibre-web enabled=false offbox=true
immich enabled=false offbox=true <-- note said OFF
```
Per the rail ("trust the recorded pre-uninstall state over this note") all three were
restored to `offbox:true` — which they already were, since the toggle survives uninstall
(§2). Verified post-redeploy.
**Detail per app**
- **bookstack** — MAJOR ×2. Upstream v26.05 needs `storage/fonts` writable for PDF export and
makes revision-viewing a separate permission; neither bites a fresh deploy. Logged in with
the documented `admin@admin.com / password` → 302 + 6 authenticated markers on the
dashboard. mariadb 12.3 healthy via `healthcheck.sh`.
- **calibre-web** — already at the newest STABLE tag; everything newer upstream is `dev-*`
(correctly excluded). Logged in with `admin / admin123` (Flask CSRF token required) →
title renders `Calibre-Web Automated | Books (0)`. The empty library is expected: the DB
volume was destroyed while the HDD book files were preserved, so the library needs
re-importing — the accepted consequence. The traceback in its log is **benign**: calibre's
installer fails headless `xdg-desktop-menu` setup, catches it, reports "There were 1
warnings", and continues.
- **immich** — MAJOR. v3.0.0 drops pgvecto.rs and requires VectorChord; our pin was already
VectorChord so the fresh deploy was unaffected. Sidecar moved to the extension versions
immich v3.0.3 ships in its own compose (vectorchord 0.4.3 / pgvectors 0.2.0) while
**keeping our PG major 16** rather than following upstream down to 14 — a `16-` build of
that exact extension pair is published, so no needless major change. Proven functional,
not merely healthy: `/api/server/ping``pong`, `/api/server/version``{3,0,3}`,
122 migration/init log lines, "Immich Microservices is running [v3.0.3]".
`login: page-only` is deliberate — the fresh install sits at first-run admin signup
(`isInitialized:false`, signup page 200) and creating an admin would mint a credential
that would then have to be transmitted to Viktor out-of-band; the owner creates his own.
- **redis vs valkey (recorded, not fixed):** immich v3 upstream migrated `redis`
`valkey:9`. We kept `redis:7-alpine` and **it works fine on v3** (all 4 containers
healthy). Swapping the image is a structural change, not a pin bump, so it is recorded
for a considered follow-up rather than done here.
---
## 4. Version-bump policy actually applied (deviation, stated deliberately)
App images: newest STABLE upstream tag, rc/beta/nightly/dev excluded, digest-free pinned
tags per house style.
**DB/cache sidecar majors were deliberately NOT bumped** (postgres 16→18, redis 7→8,
mariadb→12.3 on apps other than bookstack, postgis 16→17). This is a conscious deviation
from a literal reading of "newest stable for every pin":
- a DB major is a **data-plane decision the application owns**, not a currency decision —
immich proves it, upstream pins one specific tested postgres build;
- tags like `postgres:16-alpine` / `redis:7-alpine` already track the newest patch inside
their major, so they are not stale;
- blind-bumping ~8 apps onto PG18 would have manufactured failures that are artifacts of
the sweep's own choice rather than real findings, and would have burned the wall-clock the
rail explicitly told us to protect.
→ carried to ROADMAP as a decision item, not silently skipped. See §7.
---
## 5. Healthcheck audit — the wget lesson, systematically
Every deployed app's compose healthcheck was parsed and the probe binary checked **inside
the image**.
> **Finding C7-F4 (HIGH) — a lying healthcheck does not merely mislead; it takes the app
> OFF-LINE.** Traefik will not route to a container in an `unhealthy` state. So a probe that
> ENOENTs → permanent `unhealthy` → **Traefik returns 404 to the customer while the app is
> serving 200 perfectly well on its own port**. This was not theoretical: both apps below
> were completely unreachable for that reason, and both looked "deployed" in docker ps.
| app | lie | reality | fix | verified |
|---|---|---|---|---|
| **adventurelog** (frontend) | `wget --spider` | distroless image: no shell, no wget, no curl; node only at `/nodejs/bin/node`, off PATH | Node-exec family via the absolute interpreter path | 404 → **200**, 3/3 healthy |
| **emby** | `curl -f` | no curl, no standalone wget — image ships only BusyBox v1.38 | BusyBox-wget family via `/bin/busybox wget` | unhealthy/404 → **healthy/302** |
Both fixed in-template and live-re-validated on the demo box.
---
## 6. Result matrix
**All 53 catalog apps were attempted.** `http` is through the real Traefik ingress; a 3xx to a login
page is a PASS. `login` was attempted where credentials are documented and the flow is scriptable.
### 6.1 Passed (45)
| app | old → new pin | MAJOR | deploy | health | http | logs | settle |
|---|---|---|---|---|---|---|---|
| actualbudget | 26.1.0 → 26.7.0 | | ok | ok | 200 | clean | 10s |
| adventurelog | v0.11.0 → v0.12.1 | | ok | **fixed** | 200 | clean | 100s |
| audiobookshelf | 2.19.5 → 2.35.1 | | ok | ok | 200 | clean | 30s |
| bentopdf | v2.8.6 (newest) | | ok | ok | 200 | clean | 20s |
| **bookstack** ¹ | 25.02.2 → 26.05.2 · mariadb 11.6 → 12.3 | ✅ | ok | ok | 302 | clean | 20s |
| calcom | v4.6.9 → v6.2.0 | ✅ | ok | ok | 307 | clean | 20s |
| **calibre-web** ¹ | v4.0.6 (newest) | | ok | ok | 302 | noisy-benign | 20s |
| claper | 1.8 → 2.5 | ✅ | ok | ok | 200 | clean | 70s |
| code-server | 4.96.4 → 4.129.0 | | ok | ok | 302 | clean | 40s |
| crafty-controller | 4.10.7 (GitLab registry, manual) | | ok | ok | 302 | clean | 50s |
| docmost | 0.25.3 → 0.95.0 | | ok | ok | 200 | clean | 20s |
| emby | 4.9.0.42 → 4.10.0.20 | | ok | **fixed** | 302 | clean | 10s |
| ghost | 6.19.2-alpine → 6.53.0-alpine | | ok | ok | 200 | clean | 10s |
| gitea | 1.23.4 → 1.27.0 | | ok | ok | 200 | clean | 10s |
| grafana | 11.5.1 → 13.1.0 | ✅ | ok | ok | 302 | clean | 20s |
| gramps-web | v24.12.1 → v25.6.0 | ✅ | ok | **fixed** (mem) | 200 | clean | 110s |
| home-assistant | 2026.2.2 → 2026.7.2 | | ok | ok | 302 | clean | 70s |
| homebox | v0.16.3 → 0.26.2 | | ok | **fixed** ×3 | 200 | clean | — |
| homepage | v1.2.0 → v1.13.2 | | ok | ok | 200 | clean | 20s |
| **immich** ¹ | v2.5.5 → v3.0.3 · pg → 16-vectorchord0.4.3 | ✅ | ok | ok | 200 | clean | 20s |
| jellyfin | 10.11.6 → 10.11.11 | | ok | ok | 302 | clean | 30s |
| kimai | apache-2.25.0 → apache-2.57.0 | | ok | ok | 302 | clean | 50s |
| komga | 1.20.0 → 1.25.0 | | ok | ok | 200 | clean | 20s |
| mealie | v3.10.2 → v3.20.1 | | ok | ok | 200 | clean | 40s |
| n8n | 1.79.3 → 2.31.3 | ✅ | ok | **fixed** (mem) | 200 | clean | 50s |
| navidrome | 0.54.5 → 0.63.2 | | ok | ok | 302 | clean | 10s |
| nextcloud | 31.0.14-apache → 34.0.1-apache | ✅ | ok | ok | 302 | clean | 50s |
| onlyoffice | 8.3.0 → 9.4.0 | ✅ | ok | ok | 302 | clean | 40s |
| opengist | 1.10 → 1.13 | | ok | ok | 302 | clean | 10s |
| outline | 0.82.0 → 1.9.1 | ✅ | ok | **fixed** (env) | 200 | clean | 40s |
| paperless-ngx | 2.15.3 → 2.20.15 | | ok | ok | 302 | clean | 70s |
| papra | 26.6.1-rootless (newest) | | ok | **fixed** ×2 | 200 | clean | — |
| privatebin | 1.7.5 → 2.0.5 | ✅ | ok | ok | 200 | clean | 10s |
| radarr | 5.17.2 → 6.3.0 | ✅ | ok | ok | 200 | clean | 10s |
| rallly | 3.11.2 → 4.11.1 | ✅ | ok | **fixed** (mem) | 200 | clean | 20s |
| recipe-importer | v0.9.11 (ours, internal registry) | | ok | ok | 302 | clean | 10s |
| romm | 4.5.0 → 5.0.0 | ✅ | ok | ok | 200 | clean | 50s |
| seerr | 2.3.0 → 2.7.3 | | ok | ok | 307 | clean | 30s |
| sonarr | 4.0.13 → 4.0.19 | | ok | ok | 200 | clean | 10s |
| sparkyfitness | v0.17.2 → v0.17.3 (server + web) | | ok | ok | 200 | clean | 40s |
| tandoor | 1.5.26 → 2.6.13 | ✅ | ok | **fixed** ×4 | 302 | clean | 40s |
| termix | 2.5.0 (newest) | | ok | ok | 200 | clean | 20s |
| uptime-kuma | `:2` → 2.4.0 (**floating tag pinned**) | | ok | ok | 302 | clean | 30s |
| vaultwarden | 1.33.2-alpine → 1.36.0-alpine | | ok | ok | 200 | clean | 10s |
| wger | 2.3 → **2.6** (2.3 is gone upstream) | | ok | **fixed** ×3 | 302 | clean | 60s |
| vikunja | 0.24.6 → 2.3.0 | ✅ | ok | ok ⚠ no healthcheck | 200 | clean | 0s |
| wishlist | Hub 1.9.0 → **ghcr v0.66.0** | | ok | **fixed** ×2 | 200 | clean | — |
| zipline | 4.0.0 → 4.6.1 | | ok | **fixed** ×2 | 200 | clean | — |
¹ the customer trio — left RUNNING as the end state (see §3). Login: bookstack **ok**,
calibre-web **ok**, immich page-only (deliberate, §3). For the other 40 apps login was
`not-attempted` — a rendered login page through the real ingress was taken as PASS per the rail.
### 6.2 Not passing (4) + 1 not attempted
| app | outcome | cause | disposition |
|---|---|---|---|
| **glance** | FAIL — crash-loop | needs `/app/config/glance.yml`; the template mounts an EMPTY config volume and never seeds one | **pre-existing** — proven: v0.7.4 (the pre-campaign pin) fails identically. Bump KEPT, finding raised |
| **gokapi** | FAIL — crash-loop | v2.2.4 refuses to run against the seeded ConfigVersion-21 config: *"Please update to version 2.0.0 before running this version"* | **pin REVERTED to v1.9.6** (last known-good); v2 migration needs the seeded config regenerated |
| **plant-it** | FAIL — image | `msdeluise/plant-it` does not resolve on Docker Hub — **neither 1.0.1 nor the shipped 0.10.0** | **pre-existing**; no replacement registry found. Needs upstream research |
| **wanderer** | FAIL — image | `ghcr.io/flomp/wanderer` does not resolve — **neither 0.20.0 nor the shipped 0.16.0** | **pre-existing**; meilisearch sidecar bumped v1.12 → v1.49. Needs upstream research |
| **plex** | not attempted | `PLEX_CLAIM` is a required field with no default — a real claim token from plex.tv is needed | not automatable; genuinely owner-supplied. Not a defect |
---
## 7. Findings, ranked
**F4 (HIGH) — a lying healthcheck takes the app OFF-LINE, it does not merely mislead.**
Traefik refuses to route to a container in `unhealthy` state, so an ENOENT'ing probe →
permanent `unhealthy`**404 for the customer while the app serves 200 on its own port**.
Seven apps were affected. Detail in §5 and F5 below.
**F5 (HIGH) — 7 of 53 apps shipped a broken or wrong healthcheck.** All fixed and live-re-validated:
| app | the lie | reality | fix |
|---|---|---|---|
| adventurelog (frontend) | `wget --spider` | distroless: no shell/wget/curl; node only at an absolute path | Node-exec via `/nodejs/bin/node` |
| emby | `curl -f` | no curl, no standalone wget — BusyBox only | `/bin/busybox wget` |
| papra | `wget --spider` | image ships only `node` | Node-exec |
| wishlist | `wget --spider` | image ships only `node` | Node-exec |
| homebox | `wget --spider` (**HEAD**) | endpoint answers **405 to HEAD, 200 to GET** | `wget -q -O /dev/null` (GET) |
| zipline | `GET /api/health` | v4 renamed it — `/api/health` 404, `/api/healthcheck` 200 | corrected path |
| tandoor | `start_period: 30s` | gunicorn still booting; probes exhausted at ~2 min | `start_period: 240s` |
**F6 (HIGH) — 5 apps were ALREADY undeployable before this campaign.** None was caused by the
sweep; the sweep is simply the first thing that ever tried to deploy them:
`glance` (no seeded config), `papra` (no `AUTH_SECRET`), `zipline` (v4 `DATABASE_URL` rename),
`wishlist` (dead Docker Hub image), `plant-it` + `wanderer` (images do not resolve at all).
**Four are now fixed; plant-it and wanderer need upstream research.**
→ The catalog had no standing "does every template still deploy?" check. That is the real gap.
**F7 (HIGH, systemic) — the update path cannot express a multi-hop major upgrade.**
Nextcloud states plainly: *"You cannot skip major releases."* This campaign moved its template
31 → 34. A fresh deploy is fine (validated, 302), but an EXISTING customer's update button would
attempt 31 → 34 in one step, which Nextcloud forbids. Same shape for any app with sequential-major
rules. The template pin is a single value with no notion of an upgrade path.
→ ROADMAP item; nextcloud is the sharpest case but not the only one.
**F8 (MEDIUM, security-relevant) — gokapi is parked on a superseded version line.**
The revert to v1.9.6 restores the pre-campaign state rather than introducing a new regression, but
gokapi cannot reach v2 until the seeded `config.json` is regenerated in the v2 format. It should not
sit on v1 indefinitely — this wants a dedicated task, not a backlog line.
**F1 (MEDIUM) — the floor-lift controller update emitted no `controller_updated` event.** §0.
**F9 (MEDIUM) — an "obvious" env fix nearly orphaned customer data.** While chasing wger 2.6 the
interim fix pinned `DJANGO_DB_DATABASE=/home/wger/db/database.sqlite`. On 2.3 that path came from
wger's own default; hard-coding a different explicit path would have pointed an existing customer's
wger at an **empty** database while looking perfectly healthy. Reverted deliberately along with the
pin. Worth remembering as a class: *adding an explicit path for a value that previously defaulted is
a data-location change, not a config tidy-up.*
**F10 (MEDIUM) — 24 templates probe with `wget --spider`, which issues HEAD.** homebox proved the
trap (405 to HEAD, 200 to GET). The other 23 validated green, so their endpoints do answer HEAD —
but the default choice is fragile and silently costs availability when an app tightens its methods.
→ convention note for REUSE.md: prefer a real GET (`wget -q -O /dev/null`) unless HEAD is verified.
**F11 (LOW) — vikunja ships with NO healthcheck at all.** It serves 200 and Traefik routes it
(no health state to filter on), so it works — but it has no liveness signal, and the controller-side
probe is the only thing watching it. Not fixed: adding one was not "trivial-and-certain" within the
allowed set.
**F2 (LOW) — the remove receipt never lists destroyed volumes** (`volumes_removed: null`). §2.
**F3 (LOW) — a removed app keeps its offsite toggle forever.** §2. Needs a ruling.
**F12 (LOW) — `.felhom.yml` `mem_limit` can drift from the compose sum.** tandoor claimed 512M
while its services summed to 768M. Only found because tandoor OOM'd. A mechanical gate would catch
the whole catalog at once (the repo already has `check-image-pins.py` as the pattern to copy).
### 7.1 Sweep-engine bugs (recorded because they are the "hollow validation" failure mode)
Three bugs in the campaign's own harness would have written a **false matrix** — two made broken
things look fine, one made fine things look broken. All fixed mid-run and the affected apps
re-validated (§1.1): the Traefik registration race (false 404s), `docker exec` writing its OCI error
to **stdout** (which turned the entire healthcheck audit green), and sampling `created` as a
terminal state.
---
## 8. Coverage
**53 of 53 catalog apps attempted — full coverage. No resumable remainder.**
- **45 passed** end-to-end through the real pipeline (deploy → healthy → HTTP through Traefik → log
scan → removed via the real delete flow).
- **3 left RUNNING** as the end state (the customer trio), on current versions, offsite toggles
restored, both scripted logins green.
- **4 failed**, each with a diagnosed root cause and a disposition (§6.2); 2 of those are pinned
back to their last known-good version rather than shipping a broken pin.
- **1 not attempted** (plex — needs a real `PLEX_CLAIM` token; not a defect).
- **13 template fixes** committed, every one live-re-validated on the demo box.
- **Alert e-mails observed:** none fired during the sweep window. The pre-existing
`kimaradtak: bookstack` offsite warning (from the 17:14 run, before the campaign) was still on
record at the start and is expected per the rail — not chased.
### 8.1 Not done / explicitly out of scope
- **DB/cache sidecar majors** (postgres 16→18, redis 7→8, mariadb→12, postgis 16→17) —
deliberately not bumped, rationale in §4. **ROADMAP decision item.**
- **MAJOR breaking notes** were retrieved from upstream for bookstack, immich, nextcloud, n8n and
grafana. For the remaining MAJOR rows the note was **not retrieved** within the wall-clock; they
are flagged MAJOR without an upstream one-liner rather than given a fabricated one.
- **immich redis → valkey** (upstream migrated in v3; redis:7-alpine works fine on v3) — recorded,
not done: an image swap is structural, not a pin bump.
- **glance config seeding**, **gokapi v2 config migration**, **plant-it / wanderer upstream
research** — all outside the allowed-fix set, each needs its own task.