Files
app-catalog-felhom.eu/REPORT.md
T
2026-07-19 03:22:56 +02:00

3.6 KiB

REPORT — CAMPAIGN 7: full app-catalog sweep (2026-07-18 → 2026-07-19)

Scope: every app under templates/ — bump the pin, deploy it through the real pipeline on the demo box, prove it works, audit its healthcheck, record, remove. Box: demo guest 9201 (demo-felhom on felhom-pve), controller 0.146.0. Full matrix + ranked findings: felhom.eu/documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md

Outcome

count
apps attempted 53 / 53 (full coverage, no remainder)
passed end-to-end 45
still failing 4 (glance, gokapi, plant-it, wanderer)
not automatable 1 (plex — needs a real PLEX_CLAIM token)
template fixes committed + live-re-validated 13
broken healthchecks found and fixed 7
apps that had NEVER been deployable, now fixed 5

Pre-flight gate — PASSED

artifact_golden_version = 0.146.0 (sha 4834c703…e955), min_controller_version = 0.146.0, guest running felhom-controller:0.146.0 healthy and reporting. The sweep validated on the version customers will actually run.

The headline finding

A lying healthcheck does not merely mislead — it takes the app off-line. Traefik refuses to route to a container in unhealthy state, so a probe that cannot execute (missing wget/curl, wrong port, wrong path, HEAD-vs-GET) produces a permanent unhealthy, and the customer gets a 404 while the app is serving 200 perfectly well on its own port. Seven apps were in that state.

Second finding

5 of 53 apps had never been deployable at all — papra, zipline, wishlist, homebox, glance — and none of it was caused by this sweep; the sweep is simply the first thing that ever tried to deploy them. Four are fixed. The catalog has no standing "does every template still deploy?" check, and that absence is the real gap.

End-state (verified against the dashboard's own state, not just docker ps)

  • Deployed customer apps: bookstack, calibre-web, immich — the trio, freshly redeployed on current versions, all containers healthy, offsite toggles restored to their recorded pre-uninstall state (offbox: true on all three).
  • Protected stacks untouched and running: traefik, cloudflared, filebrowser, samba, controller.
  • Orphaned stacks: none. Stuck deploying flags: none.
  • app.yaml present for the trio only; every other campaign app removed via the real delete flow.
  • Image prune run. Docker volume: 179 GB free (pre-campaign 178 GB).
  • No campaign residue in /opt/docker/stacks/ beyond synced templates.

Needs a decision or a follow-up task

  1. Multi-hop major upgrades — nextcloud states "You cannot skip major releases"; the template moved 31 → 34. Fresh deploys are fine, but an existing customer's update button cannot express a staged upgrade. Systemic; ROADMAP.
  2. DB/cache sidecar majors — deliberately not bumped (postgres 16→18, redis 7→8, mariadb→12). Rationale in the campaign doc §4. Wants an explicit ruling.
  3. gokapi v2 config migration — parked on v1.9.6 until the seeded config.json is regenerated. Security-relevant: it should not sit on a superseded line indefinitely.
  4. plant-it / wanderer — images unresolvable; need upstream research (finding, not a deletion).
  5. glance — needs a seeded default glance.yml; it has never started from this catalog.

Secrets

No secret value was logged, printed, or written to evidence at any point. Generated deploy secrets are recorded as <generated-secret, not recorded>; the trio's credentials live in their app.yaml on the box, out-of-band as usual.