Files
app-catalog-felhom.eu/CHANGELOG.md
T

1456 lines
104 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
## docmost converts PostgreSQL 16 → 18 (2026-09-25, `09` §3 decisions 35, 37, 39)
The first PostgreSQL major the catalog moves — ALONE in its commit, through `upgrade-test.py --write-ladder`, which
wrote `engine_conversion {docmost-postgres, postgres, 16, 18}` because BOTH venues converted it: the bench (throwaway
LXC 9401, harness v4 — converted in 11.0 s, check equal over 48 tables, seed read back, 10-minute memory watch, 0
kills; negative control C3 → `failed`) and scratch box 9202 through the product's own conversion (controller
`0.273.0-rc1`: done in 42 s; load failure, unhealthy app and a SIGKILL mid-conversion each undone to 16 with the seed
read back). The data mount moves to `/var/lib/postgresql` (18 refuses the old one). **docmost's own limit 384M →
512M** (decision 39): its memory reached 91 % of 384M under the watch's light load and 80.4 % of 512M — Node sizes its
heap from the limit; 0 kills, 0 restarts in both watches. The superseded 16 step keeps its definition in `steps/`.
Needs controller ≥ 0.273.0 (both demo boxes run 0.274.0); an older box refuses nothing here — it has no conversion and
would HOLD, which is why the floor went first. Evidence: `felhom.eu/documentation/audits/night-2026-09-26/C/`, `…/D/`.
## nextcloud: `after_load:` — its own rescan after a load of kept data (2026-09-25, Part E)
`09` §3 decision 36. **No image moved.** `templates/nextcloud/.felhom.yml` gains `after_load:` (`php occ files:scan
--all` as `www-data` in `nextcloud`), which the controller runs once after "use my kept data" / "Load". Measured on
9202: a file written after the backup is on the drive and invisible to Nextcloud until this rescan (0.6 s, 130 files)
— `felhom.eu/documentation/audits/night-2026-09-26/E/E1-README.md` Q2. A controller without the kept-data release
ignores the key. Gates: `catalog_gates.py --fast nextcloud` all OK.
## The harness converts a PostgreSQL major; the gate lets ONE proven app through (2026-09-25 evening, Part C)
`09` §3 decision 35 (docmost first). **No template moved in this commit.**
- `upgrade-test.py` harness **v4**: a step that moves a PostgreSQL major is CONVERTED on the bench the box's way
(old engine alone → check → `pg_dumpall` with its completion line → volume emptied → new engine alone → the
entrypoint's empty databases dropped, existing roles' `CREATE ROLE` skipped → load with `ON_ERROR_STOP` → check
again + `PG_VERSION`). `render` points a PostgreSQL 18+ data volume at `/var/lib/postgresql` (18 refuses even an
empty volume at `/var/lib/postgresql/data`). The engine probe reads `$PGDATA/PG_VERSION`.
- `--write-ladder` writes `engine_conversion {service, engine, from, to}` only when the bench converted it AND the
box converted it through the product; it moves the data mount line with the image.
- `check-engine-major.py`: a PostgreSQL major passes only with its template's proven, two-venue, marked ladder entry
for that step, as the only image move in its commit. The postgis family is judged now (it was not). Decoys: one
genuine, three look-alikes (no mark, one venue, the mark in a comment), one bundled, postgis — red-proofed.
- `ladder.py` validates the mark; `test_pg_conversion.py` covers the harness pieces that run without Docker.
## Rules: Peti's box retired (2026-09-25)
The unprompted-work rule's fence names DooPlex and ep0 only (operator ruling: Peti's box retired). No template changed.
## Two more apps moved within their major (2026-09-25 night, Part E)
Each proven on BOTH venues before it moved: the bench (throwaway LXC 9401 on demo-hp, harness v3, 10-minute memory
watch; negative control `n8n=alpine:3.20` → `failed`, so the bench measures) and scratch box 9202 through the real
guarded Update on controller v0.271.0 (seed read back before and after). Each written by `upgrade-test.py
--write-ladder`, its own commit, the move gate asked the registry.
| app | step | bench | box | memory peak (own) | commit |
|---|---|---|---|---|---|
| n8n | 2.41.1 → 2.41.2 | proven | proven | 22.4 % | `f14a608` |
| mealie | v3.27.0 → v3.28.0 | proven | proven | 23.1 % | `b996218` |
Published after 02:30 on purpose: no demo box has either app, and a move published before the real night would
have changed what the night tested. Evidence: `felhom.eu/documentation/audits/night-2026-09-25/E/`.
## Every step keeps its own `.felhom.yml` (2026-09-24 night, R-664)
**No `image:` line moved.** Controller v0.269.0 reads them.
- `steps/<step_key(to)>.felhom.yml` beside every step's compose: the box judges a step with ITS probe, memory
request and applied record, not the head's. `ladder.step_meta_file`, `ladder.strip_ladder_block` (a step's copy
carries no ladder).
- **Gate rule 4b:** a step without its `.felhom.yml`, or one without a healthcheck while the template has one, is
refused. Decoys: the file missing; a file with the right NAME and no healthcheck (red-proof: rule off → both pass
wrongly). Suite 86 cases OK.
- **Writer:** `--write-ladder` also keeps the superseded step's `.felhom.yml`. **Backfill** (`steps_meta_backfill.py`,
one-off): 8 files, each from the newest commit whose compose names that step's images.
## Every step keeps its own definition — the box can climb one step at a time (2026-09-24, `09` §6.4 part 5, R-653, R-656, R-663)
**No `image:` line moved.** The box half ships in controller v0.268.0.
- **`steps/<step_key(to)>.yml`** — every `update_ladder:` entry but the newest now carries its own complete
compose file. `ladder.step_key` is sha256 of `to` as canonical JSON, first 16 hex; the controller computes
the same string (`stacks.StepKey`, pinned by one shared value).
- **Gate:** `check-test-record.py` rule 4 — an intermediate step with no file, or a file whose own `image:`
lines are not that step's `to`, is refused. Decoys: no file; the right NAME naming the head's image; the
definition under another name (red-proof: rule removed → all three pass wrongly).
- **Writer:** `upgrade-test.py --write-ladder` keeps the superseded head's compose — fixes included — as its
step file when it appends the next entry (red-proof in `test_ladder_writer.py`).
- **Backfill** (`scripts/steps_backfill.py`, one-off): 8 step files for the 7 apps with more than one entry
(emby, ghost, immich, n8n, navidrome, nextcloud, romm ×2), each from the NEWEST commit whose compose names
that step's images — romm's first step comes from `f4eb94f` (the working 768M / two-worker template), not
from `15f9ebf` (the move that OOM-looped).
- **R-653:** the memory watch's load is `reached` only when at least half its requests got an HTTP answer;
otherwise the edge is `inconclusive`, never `proven`.
- **R-656:** the bench clears the app's own scratch drive folders before FROM and says so (never a bare root,
never outside the scratch roots).
- **R-663:** `test_gate_decoys.py` and `test_ladder_writer.py` had been red since the night of 2026-09-23
(kimai-db and navidrome moved under literals); they now READ the live pins. Suite: 84 cases OK.
## wishlist fits its first boot; uptime-kuma no longer parks on its database wizard (2026-09-23 night, R-612, R-613)
**No `image:` line moved.** Two template fixes, each red-proofed on scratch guest 9202 through the product.
- **wishlist (R-612):** memory 128M → **512M** (`mem_limit` 512M). Measured on the bench: the first-boot
`prisma db seed` peaks at 312–345M and was OOM-killed at 128M on every run (kernel `oom_kill` 3); on
9202 at 128M the kill came after the Role/Group rows this time, so the sign-up lie did NOT reproduce —
the kill is timing-dependent, and at 512M the seed completes (`The seed command has been executed`).
Running, the app uses ~115M — 90 % of the old limit on its own.
- **uptime-kuma (R-613):** `UPTIME_KUMA_DB_TYPE=sqlite`. Before: the box read `running` while the app's own
`/api/entry-page` said `setup-database`. After: `entryPage`, `/metrics` 401 (the real server), the database
in `/app/data` (`kuma.db`, `db-config.json` — inside the backed-up volume). The probe is unchanged.
## The test record: an image move must carry its proof (2026-09-23 night, `09` §6.4 part 4 + the catalog half of part 6)
**No `image:` line moved in this commit.** 21 templates gain an `update_ladder:` (backfill); scripts only otherwise.
- **Format** (`scripts/ladder.py`): `update_ladder:` at the end of `.felhom.yml`, one JSON flow mapping
per line — `from`/`to` per service, `digest` per `to` ref, `verdict` (proven | unrecorded), `tested_at`,
`harness_version`, `evidence`, `memory_peak_pct`, `marks` {files_may_change, needs_person,
memory_tight}, optional `backfilled`. **Spiked live first:** controller v0.266.0 and v0.267.0 on scratch
guest 9202 synced, deployed, probed and badged navidrome with the block exactly as without it.
- **Gates** (rows 8 and 9 of `catalog_gates.py`, both in `--fast`): `check-test-record.py` (static, runs in
CI) and `check-test-record-move.py` (history; the registry is asked ONLY for refs a range moves — an
unreachable registry is INCONCLUSIVE, never a pass). 16 decoy cases in `test_gate_decoys.py` (both
directions); three red-proofs seen failing (bare move, a `failed` entry, a digest mismatch).
- **The writer**: `upgrade-test.py --write-ladder` (bench AND box `proven`, template at FROM, digests
resolved, memory peak as a percent) — `test_ladder_writer.py`. `--move <app> <svc>=<ref>` builds an edge
from the template. Harness v3: the box walk's fixtures run on the bench (`upgrade_boxport.py`,
`upgrade_fixtures_box*.py` ported verbatim — R-462), and a bind-mount tree hash sets `files_may_change`.
- **`scripts/image_digest.py`** resolves a ref's digest (stdlib only); positive control: it equals the
`RepoDigests` Docker recorded for `privatebin/pdo:2.0.6` on 9202.
- **Backfill** (`ladder_backfill.py`): the 21 moves of 2026-09-22, each from the record its commit cited —
**21 proven, 0 unrecorded** (nextcloud's commit cited none; its record `nextcloud-engine-mariadb` was
named and the entry says so). romm carries its memory watch (M1, 80.9 %, `memory_tight`). Digests are
what the registry serves TODAY, and each entry says that.
- `test_catalog_gates.py`: its table test had been stale (asserted 5 gates while there were 7); now 9.
## The upgrade harness watches memory after the readback — the RomM lesson (2026-09-23, R-635/R-462)
**Test code only. No template changed; no `image:` line moved.** `scripts/upgrade-test.py` (harness
version 2) and `scripts/upgrade_fixtures.py`.
**Why.** `romm 5.0.0 -> 5.3.0` was walked, read back and promoted, then OOM-looped on demo-hp for six
hours. A walk lasting minutes measured "the update applied and the data survived", never "the new
version runs".
**What.** After an edge reads back, the harness runs the new version for `--soak` seconds (default
600, `--soak 0` turns it off) under four light callers and samples every 15 s: memory against the
compose limit, **the kernel's own `oom_kill` counter read host-side from the container's cgroup**
(Docker's OOMKilled flag is recorded beside it, never instead — it has read false for real kills on an
LXC guest), and the restart count. A kill or a restart turns `proven` into `failed`; a peak above 80 %
of the limit adds the mark `memory_tight`. The verdict record gains `memory` and `marks`. New: a
`Romm` fixture (RomM's own user API, CSRF cookie and all), edges `M1` (current template) and `M1old`
(the template as promoted, catalog `15f9ebf`: 512M, four workers), and an optional per-edge
`template` directory.
**Red-proof, run on scratch guest 9202:** `M1old` — seeded, migrated, read back, then **OOM-killed at
+76 s** (peak = 100 % of 512 MiB, restarts 0), verdict `failed`. `M1` — `proven`, 0 kills in 608.5 s, peak 81 % of 768 MiB → mark `memory_tight`.
Evidence: `felhom.eu/documentation/audits/update-rulings-2026-09-23/harness/`.
## The probe TARGET is now decidable, and ambiguity is refused rather than guessed (2026-09-22, R-630)
**No `image:` line moved.** `healthcheck.container` added to `paperless-ngx` (`paperless-webserver`)
and `immich` (`immich-server`), and the gate learned the field.
**Why these two.** The controller resolves a probe target by the stack's own name, and neither app
has a container called after its stack: paperless-ngx has `paperless-webserver`/`-postgres`/`-redis`
(nothing even begins with `paperless-ngx`), so **its probe had never run on any box**; immich has
four `immich-*` containers and no exact match, so the old first-prefix rule picked whichever the
container list happened to yield — it could have been the database.
**Why it is a CONVICTION now, not a warning.** A probe that resolves to nothing is not merely
unjudgeable: `verifying` waits on it, so a SUCCESSFUL update of such an app is stopped by
`failAndHold`. Measured on paperless-ngx 2026-09-22 — `failed` at +313.0 s, front door 404 after,
the controller's own words `not healthy within 5m0s (last: no probe container)`.
The gate now resolves the target by the same four rules as `findProbeContainerMeta`: exact stack
name, explicit `container`, a **unique** prefix, else refuse. Eight new decoy cases, including the
no-PyYAML mode CI runs.
## The six versions the twenty-eight proved go live (2026-09-22, R-462)
Operator-approved from `DRILL-the-28-2026-09-22.md`'s promotion list. One commit per app,
`catalog_since: "2026-09-22"` on each, every gate run before each commit.
| app | move |
|---|---|
| emby | `4.10.0.20` -> `4.11.0.1` |
| ghost | `6.53.0-alpine` -> `6.64.0-alpine` |
| immich | `v3.0.3` -> `v3.2.2` |
| radarr | `6.3.0` -> `6.4.4` |
| sonarr | `4.0.19` -> `4.0.20` |
| termix | `2.5.0` -> `2.8.0` |
Each was walked on scratch guest 9202 through the product's own guarded Update, **seeded and read
back through the app's own front door both before and after**, and then restored from its own copy.
Every `to` tag was re-resolved against its registry immediately before the move.
**None of the six is installed on either demo box** (checked, not assumed: demo-hp runs
adventurelog, bentopdf, bookstack, calibre-web, docmost, kimai, opengist, paperless-ngx, privatebin
and romm; demo-felhom runs opengist). So these moves change a badge and nothing else until a
household presses Update.
## romm runs TWO web workers, not four — the actual cause (2026-09-22, R-635)
`WEB_SERVER_CONCURRENCY=2`, with the limit left at the 768M of the previous entry. **No `image:`
line moved.**
**768M was still a guess, and the guess was wrong** — it slowed the kills from ~12/min to ~7/min and
stopped nothing. The number came from measuring instead:
```
pid 2318697 RSS 216 MiB <- warm uvicorn worker
pid 2318717 RSS 215 MiB <- warm uvicorn worker
pid 2319627 RSS 63 MiB <- just restarted after a kill
pid 2319629 RSS 62 MiB <- just restarted after a kill
pid 2315417 RSS 18 MiB <- gunicorn master
```
Four warm workers plus the master reach **~882 MiB** before nginx and the job runner in the same
container, and the cgroup's own `memory.peak` read **exactly 768 MiB** — it hit the new ceiling and
was killed there.
**The lever was in the image all along:** `/init:143` runs
`--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a server default on an appliance
serving one household.** Two measure ~450 MiB and fit 768M with headroom, and they halve the CPU
churn as well — the host's load average had been sitting at **5.2 while otherwise idle**.
**The lesson is not about romm.** A version move is not only an `image:` line: the new version's
*shape* — worker counts, per-worker footprint — has to be measured too, and a walk that lasts
minutes cannot see a ceiling that is reached in two hours.
## romm gets 768M — 5.3.0 does not fit in 512M (2026-09-22, R-635)
**No `image:` line moved, so no `catalog_since` moved.** `romm` 512M -> **768M**; the header and
`.felhom.yml` totals follow (1024M -> 1280M).
**Measured on demo-hp, not reasoned about.** `romm 5.0.0 -> 5.3.0` was promoted that morning after
the edge was proven on the scratch guest, and the guarded Update on demo-hp reached `done` in 74.8 s.
It ran clean for two hours. Then: `OOMKilled: true`, **4,530** `Worker … was sent SIGKILL! Perhaps
out of memory?` in six hours, the container pinned at 457 MiB of its 512 MiB limit, **~500% CPU** in
a permanent restart storm, host load average **5.2 while otherwise idle**.
**Two things this cost that a bigger number alone does not fix, both recorded in R-635.** The update
read `done` and the app read `running` throughout, because nginx answers `GET /` with 200 while the
gunicorn workers behind it are being killed — the probe is right, the port is right, and the green is
still false. And nothing alarmed: OOM signals are invisible inside an LXC guest (R-528), and a
running-but-thrashing app is not "down". **The operator heard the fans. That was the detector.**
768M is a first measured step, not a final answer — the soak that follows this entry is what settles
whether it is enough.
## The probe gate now runs in CI's own degraded mode (2026-09-22, R-618)
**Caught by checking the push's CI run by job id, not by assuming it went green.** Job **877** on
`15d7c2b` read `conclusion: failure`. The CI runner has no PyYAML, the new gate answered
INCONCLUSIVE there, and the runner correctly refuses to call an undetermined result a pass.
A gate that is red on every CI push is bypassed within a week, so it would enforce nothing — the
R-421 shape one level up. Both files it reads use a tiny regular subset of YAML, so it now falls
back to a line reader and says `mode: DEGRADED` when it does. Over all 53 apps the degraded reader
returns **exactly** what the full reader returns: no convictions and the same six warnings.
Pinned by five more decoy cases with PyYAML shadowed out (suite now **56 cases**): the three real
faults re-introduced one at a time are still REFUSED, the fixed tree passes and says it is degraded,
and moving Traefik's label still convicts nothing.
## Fifteen proven versions move to the catalog (2026-09-22, R-462)
Every one was walked on scratch guest 9202 **through the product's own guarded Update**, seeded and
read back through the app's own front door, during the update night of 2026-09-21. One commit per
app; each app's `catalog_since` moves to 2026-09-22 (R-452).
| app | move |
|---|---|
| actualbudget | `26.7.0` -> `26.9.0` |
| audiobookshelf | `2.35.1` -> `2.36.1` |
| bookstack | `26.05.2` -> `26.05.5` |
| docmost | `0.95.0` -> `0.96.0` |
| grafana | `13.1.0` -> `13.2.2` |
| home-assistant | `2026.7.2` -> `2026.9.3` |
| mealie | `v3.20.1` -> `v3.27.0` |
| n8n | `2.31.3` -> `2.40.5` |
| navidrome | `0.63.2` -> `0.64.0` |
| papra | `26.6.1-rootless` -> `26.6.2-rootless` |
| privatebin | `2.0.5` -> `2.0.6` |
| romm | `5.0.0` -> `5.3.0` |
| tandoor | `2.6.13` -> `2.6.15` |
| vikunja | `2.3.0` -> `2.6.0` |
| **nextcloud (ENGINE)** | `mariadb:11.6` -> `mariadb:12.3` |
**tandoor is here because its edge was re-walked TODAY and passed.** On 2026-09-21 the same edge
ended `failed` at +361.9 s with the app stopped — by the wrong health probe, not by the version.
With the probe corrected the identical edge ends **`done` at +41.1 s**, seed read back, both
containers running with zero restarts.
**nextcloud is the one ENGINE move and it is deliberate**: proven end to end in 217.4 s,
`MARIADB_AUTO_UPGRADE=1` already in the template, and `check-engine-major.py` run against that exact
commit ALLOWS it by name under R-469 rather than being assumed to.
Nothing was forced. Every gate was run before each commit, `check-image-resolvable.py` confirms all
18 unique pins still exist upstream, and no move was dropped.
## A gate that refuses a probe the app does not answer (2026-09-22, R-618)
`scripts/check-probe-matches-compose.py`, wired into `catalog_gates.py` as a **`--fast`** gate, so
it runs in the pre-push hook and in CI.
**The oracle was already in the file.** Every template's probed service carries a compose
`healthcheck.test` that dials the app on loopback — `wget http://127.0.0.1:80/accounts/login/`. It
is written by whoever added the template and exercised by docker on every start. The gate compares
the `.felhom.yml` probe against it, statically, with no network and no container. The probed service
is resolved exactly as `findProbeContainer` resolves it (container name equal to the stack name,
else prefix).
**Two verdicts, and the narrower one is a measurement, not a compromise.** A wrong PORT is always
fatal — nothing listens, every dial is refused — so it REFUSES, for every check type. A wrong PATH
is fatal only when the probe can fail on it: `probeHTTP` calls ANY response healthy for type `http`,
and for type `api` with no `expect` block. So a path mismatch REFUSES only with `api` + `expect`,
and WARNS otherwise. Over all 53 apps that is the difference between convicting `zipline` and merely
warning about `home-assistant`, whose badge is right today and would break the day someone adds an
`expect`.
**Six warnings on the current catalog**, none a conviction and none a pass:
| app | why no verdict |
|---|---|
| `paperless-ngx` | **no container_name equals or begins with the stack name** — `findProbeContainer` returns nothing and the stack is skipped, so no probe ever runs and the badge can never go red |
| `vikunja` | the probed service has no compose healthcheck — no oracle |
| `crafty-controller`, `mealie`, `uptime-kuma` | the healthcheck dials through a python/script helper, not a loopback URL — no oracle |
| `home-assistant` | path mismatch on a probe that cannot fail on it (above) |
**Red-proofs, four, plus decoys both ways** (`scripts/test_gate_decoys.py`, now 51 cases): each of
the three real faults re-introduced one at a time and refused; the fixed tree passes; a clean app
given a wrong port refused; an `api`+`expect` probe given a wrong path refused. The decoys move the
LABEL and must not convict — the port in a YAML comment, Traefik's `loadbalancer.server.port`, a
published `ports:` mapping, a NON-probed sidecar's own healthcheck, and a path mismatch on a probe
that cannot fail on it. The gate takes `--root=<dir>` so the suite judges its clone and not the real
repo; without it every case would read identical bytes and pass for nothing.
## Three health probes now dial where the app actually listens (2026-09-22, R-618)
**Templates only, and no `image:` line moved — so no `catalog_since` moves either.**
`.felhom.yml`'s `healthcheck.checks` tells the controller where to knock, and it knocks from INSIDE
the compose network: `<container-name>:<port><path>`. The port is therefore the port the process
listens on inside its container — never the published port and never Traefik's. Three templates
named something else:
| app | was | is | what the app's own compose healthcheck dials |
|---|---|---|---|
| tandoor | port `8080` | port `80` | `http://127.0.0.1:80/accounts/login/` |
| zipline | path `/api/health` | path `/api/healthcheck` | `http://127.0.0.1:3000/api/healthcheck` |
| wger | port `80` | port `8000` | `http://127.0.0.1:8000` |
**Why this is not merely a wrong badge.** The guarded update's `verifying` phase waits on this same
probe, and `failAndHold` runs `compose down` when it never goes green. So a SUCCESSFUL update ended
by STOPPING a working app. Measured on guest 9202, 2026-09-21: tandoor served HTTP 200 on the new
version at four samples across five minutes, docker's own healthcheck green throughout, and the
controller stopped it at +361.9 s.
Proven live on guest 9202 before and after, through the product: at the live pin all three read
`Nem egeszseges` / `Not healthy` on their own app page while their front doors served a real page
(tandoor 200 `Login / Sign In`, zipline 200 `Zipline`, wger 200 `wger Workout Manager`) and docker
reported every container healthy.
## Vikunja fixture: the create verb is PUT, not POST (2026-09-21, R-462)
Test code only, and a one-line correction to the fixture shipped earlier the same night.
Vikunja CREATES a project with **`PUT /api/v1/projects`**. The fixture sent `POST`, which answers
`405 Method Not Allowed` — and a 405 in the middle of a seed reads like a broken app rather than a
wrong verb. Measured on the box before the fixture was ported here; the box-side copy was corrected
at the same time and its edge then walked clean (`vikunja 2.3.0 -> 2.6.0`, proven, 24.6 s).
Kept as its own entry rather than folded into the previous one, because the previous entry claims
these fixtures were proven box-side first — and for this one line that claim was true of the SEED
ROUTE and not of the verb the port carried.
## Upgrade harness: four new fixtures and seven real upstream EDGES (2026-09-21, R-462)
**Test code only. No template changed, and no `image:` line moved.** The update night
(`felhom.eu/documentation/audits/DRILL-update-night-2026-09-21.md`) walked real within-a-major
upstream edges on scratch guest 9202 **through the product's own guarded Update**, each app seeded
and read back through its own front door. This commit brings that work back into the harness so the
same edges can be run here **with their ABORT step**, which the box deliberately does not offer
(`09` §6.1: the box never puts the old version back by itself, because whether the old image starts
on migrated data is per-app and cannot be predicted).
- **`upgrade_fixtures.py` — four new fixtures**: `ActualBudget` (its own bootstrap + login),
`Navidrome` (`/auth/createAdmin` + `/auth/login`), `AudiobookShelf` (`/init` + `/login`) and
`Vikunja` (register → login → create project → authenticated readback). Every one goes in through
the app's OWN interface (R-156) and every one carries its own negative control, run on every
`verify()`: a wrong password, or an id that cannot exist, must NOT read back — so a readback that
has broken into always succeeding fails instead of passing everything.
- **`upgrade-test.py` — seven new edges, `U1`–`U7`**, all REAL upstream moves that existed on
2026-09-21 and that this catalog has NOT made: privatebin 2.0.5→2.0.6, docmost 0.95.0→0.96.0,
bookstack 26.05.2→26.05.5, actualbudget 26.7.0→26.9.0, navidrome 0.63.2→0.64.0,
audiobookshelf 2.35.1→2.36.1, vikunja 2.3.0→2.6.0. Each holds its database engine CONSTANT, per
the standing rule that an engine change gets its own edge.
- **Limitations kept rather than papered over.** `Navidrome` and `AudiobookShelf` seed the DATABASE
half only — the library on the drive is not populated — and both say so in their docstrings, as
`BookStack` already does for its file half (R-460).
**Owed, and stated so it is not mistaken for done: the harness RUNS.** The code is in; the U1–U7
runs, and with them the per-app ABORT answers, have not been performed.
## RULE LIFT: a MariaDB major may cross, as its OWN EDGE; PostgreSQL may not (2026-09-21, R-469 + R-450)
The engine-major rule of 2026-09-13 named its own expiry — *until the Update button takes a verified
backup as its precondition* — precisely so it would be removed deliberately rather than forgotten.
**That condition was met on 2026-09-13**, the same day: update arc Slice 4 shipped (controller
v0.237.0/v0.238.0, any backup tier since v0.239.0). This is the deliberate removal, and it removes
exactly half.
- **LIFTED — the four MariaDB services** (`bookstack-db`, `kimai-db`, `nextcloud-db`, `romm-db`).
They now have both halves: a verified backup in front of the Update, and `MARIADB_AUTO_UPGRADE=1`
on every sidecar (R-459), whose conversion the harness WATCHED run on the E3/E3b edges with the
seeded data read back after.
- **NOT LIFTED — the eleven PostgreSQL services, and MySQL.** Postgres performs no `pg_upgrade` and
REFUSES to start on an older major's datadir (R-463). A backup is a route BACK, not a conversion —
the app simply would not come up. MySQL has nothing measured at all. Both stay refused until R-463
has a scripted `pg_upgrade` edge proven on all eleven. The refusal now says this instead of citing
the shipped R-448.
- **NEW CLAUSE — one edge, one migration (R-450, recorded 2026-09-02 and now enforced).** A MariaDB
major must be the ONLY image move in its template in that commit. bookstack's `0b73e5e` moved the
application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit: two migrations behind one
edge, and an unreadable failure when it breaks. **Within a major is unaffected** and may still ride
with anything.
- **The gate says what it ALLOWED**, by name, rather than passing in silence — a lifted rule that
goes quiet is a lifted rule nobody can audit.
- **Decoys (R-421): two new cases, and two red-proofs each seen to fail.** Removing the own-edge
check lets the bundled kimai case pass; emptying `LIFTED` turns the lone MariaDB major back into a
refusal. The old "a lone MariaDB major is REFUSED" case is now the "…is ALLOWED" case, which is the
lift itself. 40 cases, all green.
`scripts/check-engine-major.py`, `CLAUDE.md`, `scripts/test_gate_decoys.py`. No template moved.
## GATE FIX: the copy gate could not run in CI at all (2026-09-20, R-595)
**Six pushes in a row turned CI red while the local pre-push hook was green**, and each sent an alarm
mail whose own text says what that means: *"If the local pre-push hook was GREEN for this commit, then
CI and the hook disagree — that is a finding about the gates themselves, not about CI, and it
outranks whatever the push was for."* It did.
**The cause, found by contrast rather than by reading a log:** `check-copy-i18n.py` imports PyYAML,
and it is the ONLY gate in this repo that imports anything outside the standard library.
`.gitea/workflows/gates.yml` says in its own header that the runner is *"a host-mode container with
python3 and git and nothing else"*. So the gate raised `ImportError` before it checked anything, and
`catalog_gates.py` exited non-zero — on every push, including the ones that were perfectly clean.
**The fix is a DEGRADED MODE, not a skip.** Without PyYAML the gate now runs the check that needs no
parser and matters most: **every frozen Hungarian string must still occur verbatim in its app's
`.felhom.yml` bytes.** A changed, reworded or deleted Hungarian string fails in CI exactly as it does
locally. It then prints, in full, what it did NOT check — the English block's structure, its language,
its credential tokens and the coverage ratchet — all of which run in the pre-push hook, which is where
they refuse a push anyway. That is the same division `catalog_gates.py` already uses for engine-major
on a shallow clone.
**Measured before it was claimed:** 1 030 of the 1 032 frozen strings appear byte-for-byte in the raw
files. The two that do not are romm help_texts whose YAML source escapes an inner double quote, so the
escaped spelling is accepted too — with that, 1 032 of 1 032 are found, so the degraded check convicts
nothing honest.
**Five new decoy cases**, run with PyYAML shadowed by a module that refuses to import — i.e. what CI
actually executes: a changed Hungarian byte, a deleted Hungarian line and an unfrozen new app each
convict; an untouched tree passes AND says what it did not check; and romm's two escaped-quote
help_texts are still found. 38 cases in total.
## English copy — BATCH 3 of 3: the last 16 apps. THE CATALOG IS TRANSLATED. (2026-09-20, R-560 slice 5)
actualbudget, claper, docmost, emby, gitea, immich, kimai, komga, onlyoffice, opengist, plant-it,
rallly, recipe-importer, seerr, vaultwarden, zipline — **306 strings**. `EN_MISSING_CEILING` 307 → 1.
**1 031 of the catalog's 1 032 customer-facing strings now have an English twin.** The one that does
not is papra's `AUTH_SECRET` description, which is a Hungarian defect (R-593) and is left to fall
back rather than be translated wrongly.
**The gate convicted two of my own sentences, and it was half right.** Vaultwarden's invite step and
its sign-up setting both ended „…can open an account", and the retrieval-promise pattern reads
`can … open` as the claim that sealed backups can be opened. The sentences are about opening an
ACCOUNT, so the conviction was a false positive — but the wording that triggered it was also the
weaker wording, so both now read „can sign up", which is what a household would say. **The gate has
no way to register a legitimate occurrence**, which the shared vocabulary's own design calls for
(`felhom.eu/scripts/customer_copy_vocab.py`: "an occurrence must be REGISTERED with a reason in the
consuming gate's allowlist"). Filed as R-594; rewording was the honest move today, but the first
genuinely true „can be restored" in catalog copy will need the mechanism rather than a rewrite.
**Judgement calls in this batch, recorded rather than hidden:** recipe-importer is a Hungarian-site
importer, so its English names the six Hungarian sites it reads and says so plainly. Vaultwarden's
two select options („Nem – csak meghívással (ajánlott)") become "No - by invitation only
(recommended)" and "Yes - anybody who knows the address can sign up"; the recommendation stays on the
same option.
## English copy — BATCH 2 of 3: 17 apps (2026-09-20, R-560 slice 5)
audiobookshelf, code-server, crafty-controller, glance, gokapi, gramps-web, jellyfin, mealie,
navidrome, plex, sonarr, tandoor, termix, uptime-kuma, vikunja, wger, wishlist — **317 strings**.
`EN_MISSING_CEILING` 624 → 307. No Hungarian byte moved; no image, pin, `catalog_since` or compose
line changed.
**Two lines needed judgement rather than translation, and both are recorded here so the next reader
does not think they are slips.**
- Jellyfin's second first-step reads „Kövesd a beállítás varázslót — **válaszd a magyar nyelvet**".
Told to an English household that is simply wrong advice. The English says "choose your language".
The Hungarian is untouched.
- Mealie's fifth use case reads „Többnyelvű felület — **magyar is elérhető**". In English the
interesting fact is the same one seen from the other side, so it reads "A multilingual interface -
English and Hungarian among them".
Neither adds a promise the Hungarian does not make; both say the same thing to the reader actually
holding the page.
## English copy — BATCH 1 of 3: 17 apps (2026-09-20, R-560 slice 5)
Operator said go on the pilot's tone, so the other fifty follow it. No image, no pin, no
`catalog_since`, no compose line changed; **not one Hungarian byte moved** — the freeze gate proves
it on all 53 apps on every push.
adventurelog, bentopdf, bookstack, calcom, calibre-web, ghost, grafana, home-assistant, homebox,
homepage, n8n, nextcloud, outline, papra, radarr, sparkyfitness, wanderer — **319 strings**.
`EN_MISSING_CEILING` 943 → 624.
**The blocks are GENERATED from a flat `{path: english}` map, not hand-written**, and that is a
decision rather than a convenience: fifty nested blocks whose keys must match the Hungarian side
exactly is fifty chances to mistype an `env_var`, and **a mistyped key is INERT on the box rather
than an error** — the translator would never learn. The generator builds from the same flat paths
the freeze uses, which are derived from the Hungarian file itself, so a key that does not exist on
the Hungarian side cannot be written at all.
**One string was deliberately NOT translated.** papra's `deploy_fields[AUTH_SECRET].description`
reads „Az alkalmazás aldomainje" — the sentence that belongs on `SUBDOMAIN`, sitting on a
session-signing key. A localisation release may not change Hungarian bytes, and translating a wrong
sentence faithfully would ship the error in a second language. So it falls back to the Hungarian,
papra stands at 13/14, and **the ceiling's floor is 1 until R-593 is fixed** — written into the
ceiling's own comment so the last batch does not quietly "fix" it by editing Hungarian.
## English copy — the PILOT: privatebin, paperless-ngx, romm (2026-09-20, R-560 slice 5)
No image, no pin, no `catalog_since`, no compose line changed. Three apps gained an `i18n: en:`
block at the end of their `.felhom.yml`; **not one Hungarian byte moved**, and the freeze gate
proves it on all 53 apps on every push from now on.
89 of the catalog's 1 032 customer-facing strings are now English — privatebin 14, paperless-ngx 33,
romm 42. `EN_MISSING_CEILING` 1032 → 943 in the same commit.
The three were chosen for SHAPE, not for size: between them they exercise every part of the overlay
the controller learned to read in v0.257.0 — a plain app with only a description, tagline and lists
(privatebin); select options, a placeholder and a customer-facing folder label (paperless-ngx); and
the catalog's only `optional_config` block, whose group has no id of its own and is therefore matched
by `match_group`, the Hungarian group name it translates (romm). All three run on the demo box, so
the English pages could be fetched rather than reasoned about.
A box on a controller older than 0.257.0 ignores the block entirely and keeps rendering Hungarian —
which is every box in the fleet until the floor is raised, and is why this push is safe ahead of it.
What a gate CANNOT judge, and is listed per app in `REPORT.md`: whether an app's „first steps" name
the buttons that app's ENGLISH interface actually shows. Those lines are marked unverified and belong
to the slice-6 walk.
## vaultwarden: registration closed by default; paperless-ngx: one worker and room for a batch (2026-09-15, R-512 / R-514 / R-515)
No image or version change — same `vaultwarden/server:1.36.0-alpine` and paperless-ngx 2.20.15; the
templates changed.
- **R-512 — Vaultwarden no longer lets strangers open accounts.** BIGNIGHT: `SIGNUPS_ALLOWED=true`
by default, and the page told the household to close it with a field that is read-only after
install. Measured first (2026-09-15, throwaway on scratch 9202, SMTP off exactly as a box without
app-email): with `SIGNUPS_ALLOWED=false` a stranger's `POST /identity/accounts/register` → **400**;
`POST /admin/invite` with the generated admin token → **200** with no mail needed; the invited
address registers → **200**; a second stranger after the invite → **400**
(`felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/E1-vaultwarden-spike.txt`). So the
invite-first shape needs no controller change. Default is now `false`, the select says what each
choice means, and „Első lépések" is invite-first (admin panel → invite → create account).
- **R-514 — Paperless no longer drops documents on a batch.** BIGNIGHT: 20 uploads → the worker was
OOM-killed inside 768M, 11 failed, 8 stuck, 0 consumed. Measured 2026-09-15 with
`PAPERLESS_TASK_WORKERS=1` / `THREADS_PER_WORKER=1` and no cap: **20/20 consumed**, no OOM,
`memory.peak` **783 294 464 B** — only 2.7 % under the old cap, so one worker alone was not enough.
Now 1 × 1 and `memory: 1280M` (peak × 1.5, rounded up to 256M); `mem_limit` hint 1664M. Caveat
recorded: the test PDFs carried text, so OCR on scanned images may need more — the controller's new
OOM line (controller v0.243.0) makes that visible if it happens.
- **Vaultwarden compose header** no longer tells the reader to "Set SIGNUPS_ALLOWED=false via the
controller" (settings are read-only after install); it describes the invite-first steps. **Proven
live 2026-09-15 on scratch 9202:** deployed from the catalog through the controller API, container
`SIGNUPS_ALLOWED=false`, a stranger's `send-verification-email` through traefik → **400 „Registration
not allowed"**; Paperless deployed the same way: 20 PDFs at once → **20/20 SUCCESS**, `memory.peak`
772 370 432 B under the 1280M cap, no OOM.
- **R-515 — the Paperless card no longer says `admin / admin`.** That login never existed (the
password is generated). `default_creds` removed; first steps point at „Automatikusan generált
értékek", the gokapi wording.
## adventurelog: photos render — the backend's own paths go to the backend (2026-09-13, R-483)
The operator measured it in a browser: on his k3s install a photo uploads and renders; on demo-hp it
uploads and renders as a broken "Uploaded content". The diff against `homelab-manifests`'s
`adventurelog-system/adventurelog.yaml`: its ingress sends `/media`, `/static`, `/admin` and
`/accounts` to the backend and only `/` to the frontend; the catalog template sent every path to the
frontend, which does not serve `/media`, so every photo URL was a 404. Now the backend service carries
its own traefik router for those four prefixes (priority 20; the frontend router is priority 10). **Second cut the same evening:** the router must target port **80** — the backend image runs
nginx in front of gunicorn, and Django hands a photo back as an `X-Accel-Redirect` that only that
nginx serves; on port 8000 the first cut returned a 200 with an empty body (the operator's browser
still showed „Uploaded content"). The k3s service targets port 80 for the same reason. The
frontend's `/api` and `/auth` proxy stays as it is — that is also the k3s shape. Nothing else differed
(`PUBLIC_URL`, `CSRF_TRUSTED_ORIGINS`, `ORIGIN`, `BODY_SIZE_LIMIT` all match). No version moved.
## privatebin 2.0.6 → 2.0.5 — the DRILL bump REVERTED (2026-09-14, BIGNIGHT Phase 4)
Reverts `d5d91e0` in the same phase, as the drill required. The catalog is back on `privatebin/pdo:2.0.5`.
`catalog_since` stays 2026-09-14 because the `catalog-since` gate requires an image move to carry the day of the
move. The drill box (VM 333) keeps 2.0.6 — its guarded Update pinned it — and is torn down at the end of the night.
What the walk proved: the change reached the box on the next 15-minute sync (14 m 25 s after the push), the page
offered „Frissítés elérhető — ma", and the guarded Update finished in 11 s with the data intact.
## privatebin 2.0.5 → 2.0.6 — a DRILL bump, reverted in the same phase (2026-09-14, BIGNIGHT Phase 4)
A real upstream one-step release (`privatebin/pdo:2.0.6`, Docker Hub 2026-08-08), pushed so the guarded
Update could be walked end to end on the big night's test box (no app in the catalog had a newer version).
`catalog_since` set to 2026-09-14 per the rule. **The revert commit follows in the same phase**; nothing
in the fleet is meant to keep 2.0.6 from this commit. Evidence:
`felhom.eu/documentation/audits/evidence-bignight-2026-09-14/phase4/`.
## catalog-since gate: an image move must bump catalog_since (2026-09-13, R-452)
`scripts/check-catalog-since.py`, fifth gate in `catalog_gates.py` (fast, git-history, `--range A..B`
like engine-major): for every compose file changed in the range whose per-service `image:` lines
differ, the app's `.felhom.yml` at the range end must carry a `catalog_since` on or after the day of
the commit that moved them, and not in the future. A date moving in a comment, a README or a
CHANGELOG is not the fact (R-421). Enforced by the pre-push hook; CI's shallow clone skips it out
loud, as it does engine-major. Five decoy cases (`scripts/test_gate_decoys.py`); red-proof: dropping
the date comparison lets the "untouched date" fact through (`felhom.eu` `audits/v0240-2026-09-13/rp-R452.txt`).
## glance: a fresh install lands on a start page instead of a crash loop (2026-09-13, R-473)
The image ships no default config and exits at once without `/app/config/glance.yml`, so every
fresh install crash-looped (measured 2026-09-13 on demo-hp: 13 restarts, the card showing
„restarting"). First boot now seeds a small Hungarian start page — clock, calendar, weather for
Budapest, a link to the box's own dashboard, two news feeds — the way `gokapi` seeds its
`config.json`: an `entrypoint` wrapper that writes the file only when it is absent and then execs
the image's own command (`/app/glance --config /app/config/glance.yml`, read from the image). An
existing config is never touched. The seed was validated with the image's own `config:validate`.
No version moved (v0.8.5 stays).
## adventurelog: Django DEBUG off on the public origin (2026-09-13, R-482)
**Found by the first nightly rotation walk on demo-hp** (`felhom.eu` `audits/nightly-2026-09-13-adventurelog/`):
the backend image defaults `DEBUG` to True, and it served full Django debug pages — settings, paths,
tracebacks — on the internet-facing subdomain (a 500 with a traceback on an unauthenticated POST; a
CSRF page that said *"you have DEBUG = True in your Django settings file"*). Upstream's own compose
sets `DEBUG=False`; this template did not. Now it does: one env line on the backend service, no
version moved. Proven on demo-hp after the sync: the same CSRF failure renders the short production
page (see `audits/v0240-2026-09-13/` in felhom.eu).
**Not measured, recorded:** `wger`, `tandoor` and `paperless-ngx` are Django too. Their images
default DEBUG off as far as their documentation says; each gets measured on its own rotation night.
## Live-test commits for controller slice 4, all reverted the same day (2026-09-13) — NO NET CHANGE
**`templates/glance` and `templates/uptime-kuma` are byte-identical to `3525e35`** (checked with
`git diff 3525e35 HEAD -- templates/glance templates/uptime-kuma` → 0 lines). The commits moved image
tags so the controller's guarded update could be proven live on demo-hp with a throwaway app:
| commit | change | reverted by |
|---|---|---|
| `a1f1c38` | glance v0.8.5 → v0.8.4 | `045495f` — glance **abandoned**: it crash-loops on a fresh install (no `glance.yml` seeded), filed felhom.eu R-473 |
| `01c631d` | uptime-kuma 2.4.0 → 2.3.2 (install from) | `28ce33b` — that revert IS Scenario A's real tag change |
| `29a8cbe` | uptime-kuma → 2.4.999 (does not exist) | `41dd686` — Scenario E |
| `6ce3f65` | uptime-kuma → alpine:3.20 (exits at once) | `6d8c7ab` — Scenario F; **reverted early** (~2 min exposure) once the update had pinned it, after the commit security review flagged that a catalog push is a deploy. No other box had uptime-kuma |
Every commit moved `catalog_since` with the image line, per the catalog rule, and every revert restored it.
## MariaDB finishes its own conversion — `MARIADB_AUTO_UPGRADE=1` on four db services, and an engine-major gate (2026-09-13, R-459 / R-469)
**Templates changed: bookstack, kimai, nextcloud, romm — the db service's `environment:` list only.
No `image:` line moved, so `catalog_since` does NOT move** (the rule ties it to an image change).
**The ruling** (operator, 2026-09-13, on `SPIKE-r459-mariadb-upgrade-2026-09-06.md`): an unconverted
MariaDB datadir is stable but never heals — the engine says `Check required!` on every start forever —
and the conversion costs ~7 s and takes its own system-table backup first. So every `mariadb:` sidecar
now carries `MARIADB_AUTO_UPGRADE=1`; `MARIADB_DISABLE_UPGRADE_BACKUP` stays unset — that backup is the
precaution. The setting is inert until an engine major actually moves.
**Proven by the harness before it shipped** (throwaway LXC 9403 on demo-hp, destroyed after;
evidence `felhom.eu/documentation/audits/r459-close-2026-09-13/harness/`):
| edge | verdict | the engine's own view AFTER (`engine_state_after`) |
|---|---|---|
| C3 (negative control) | **`failed`** — the harness still says no | — |
| E3 (app + engine 11.6 → 12.3) | `proven`, readback after = true | `12.3.3-MariaDB \| This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again. [exit=1]` |
| E3b (engine half alone) | `proven`, readback after = true | same |
The entrypoint, verbatim: `Backing up system database to system_mysql_backup_11.6.2-MariaDB.sql.zst`
→ `Starting mariadb-upgrade` → `Finished mariadb-upgrade` (6 s). `skipped due to $MARIADB_AUTO_UPGRADE`
appears **0** times in either TO log — it appeared on every start before this change. The abort still
starts and serves (spike §5.3), and the abort log still prints `MariaDB upgrade not required`, which
is the R-464 trap and not a soundness claim.
**And a rule with a gate, because the Update button still takes no backup** (R-448 not shipped):
*until Slice 4 ships, no template may move a database-engine image across a MAJOR version* — four
MariaDB, eleven PostgreSQL services, found by image name. `scripts/check-engine-major.py` is the
fourth row of `catalog_gates.py` (`--fast`, git reads only); `.githooks/pre-push` now hands it the
push range. **It needs a parent commit and CI fetches at `--depth 1`** (the R-452 gap, not re-filed),
so on a shallow clone the runner skips it out loud; the hook is where it bites. Red-proof
(`scripts/test_gate_decoys.py`, 7 cases): `mariadb:11.6 → 12.3` and `postgres:16 → 17` REFUSED naming
the rule and its expiry; `11.6 → 11.8` passes; the version moving only in a comment, in kimai's
`serverVersion=` env, in README, or on the app's own image passes. The rule's removal is tracked as
`felhom.eu` R-469 so it is a deliberate act, not a lapse.
## upgrade-test.py records the ENGINE's own view of itself (2026-09-06, R-459) — NOT A RELEASE
**Scripts only. No template changed, and deliberately so** — nothing with `MARIADB_` in it is
committed by this work; that is a fleet-wide decision the operator owns.
The harness returned **`proven`** for edge E3b while MariaDB was logging that the datadir conversion it
requires had been **skipped**. The verdict was not wrong — the app's data did survive, which is what it
asked — but nothing here could see that the engine had been left in a state the engine itself calls
incomplete. It watched the app and the migration log; **neither looks at engine state.**
`engine_state_after` now carries, per database service, the engine's own answer. On the E3b re-run:
```
verdict: proven
engine_state_after.bookstack-db.answer:
"11.6.2-MariaDB| Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required! [exit=0]"
```
**It is reported BESIDE the verdict and never folded into it.** An unconverted datadir is not known to
be a failure — `SPIKE-r459-mariadb-upgrade-2026-09-06.md` measured 5 of 5 restarts with no degradation
— so a verdict that called it `failed` would encode an unproven judgement, which is worse than
reporting a fact and letting a person read both.
**Both engines are covered and they fail differently:** MariaDB starts anyway and skips the conversion
quietly; PostgreSQL refuses to start on a datadir from an older major. "The engine would not start" is
as much an engine-state observation as "the engine says it needs a check".
## upgrade-test.py — a harness that measures whether an upgrade keeps the data (2026-09-06, R-449) — NOT A RELEASE
**No template changed. No image moved. Scripts only.** `scripts/upgrade-test.py` +
`scripts/upgrade_fixtures.py`.
Until today this project had measured **one** app upgrade out of 53 — Nextcloud, by hand, during
`SPIKE-app-update-2026-09-01` §7 — and the whole update arc was designed against that single data
point. A one-off measurement nothing repeats decays into a claim. This is the thing that repeats it.
**Method, per edge:** deploy at the FROM images → seed through the app's **own** interface → prove the
seed reads back (control C1) → swap to the TO images → ask the app for the data **again** → then put
the FROM images back and record what happens.
**Two rules, and the first is the one people get wrong.** Success is an **application-level
readback**, not file identity: `survive2.py`'s sha256+inode rule is right for a redeploy and wrong for
an upgrade, because a migration is *supposed* to rewrite files. And, carried verbatim from that same
script: *nothing is ever seeded into a volume by hand* (R-156) — every seed goes through the app's
HTTP API or its own CLI, and an app with no such route is recorded **`inconclusive`**, never faked.
**Run `C3` first, always.** Its TO image is `alpine:3.20`, which pulls cleanly and exits immediately.
It came back **`failed`** — so the harness can say no, and its greens mean something. If it ever comes
back green, nothing else in the run is evidence.
**First run measured 7 edges across 3 apps** (privatebin, docmost, bookstack) in a throwaway guest on
demo-hp, destroyed afterwards. Findings, including a real defect in this repo's own bookstack
template: `felhom.eu/documentation/audits/SPIKE-upgrade-test-2026-09-06.md`.
## LIVE-TEST for controller v0.235.0 — two pushes, both reverted the same hour (2026-09-06) — NOT A RELEASE
**No template is different after these four commits.** `dc7e548` moved bentopdf's healthcheck
interval 30s → 45s (a NON-image change); `09b4ff5` moved its pin v2.8.6 → v2.8.5; `1798ce6` and
`17cc784` reverted both. The tree is byte-identical to `8220f8d`.
**Why real catalog pushes and not a hand-edited file on the box:** controller v0.235.0 changes what
the CATALOG SYNCER does, so the only faithful test is a real change travelling the real 15-minute
cycle — the same method the 2026-09-01 spike used.
**What they measured, live on demo-hp:**
- the non-image change **reached** the pinned app on the normal cycle (`[INFO] [sync] Updated
bentopdf/docker-compose.yml`, 08:01:51Z) with the image and the container untouched — *fixes flow*;
- the image change **did not** (08:20:29Z): the live compose file still named v2.8.6 while the catalog
offered v2.8.5, and a `POST /api/stacks/bentopdf/restart` afterwards took **0.1 s**, did not recreate
the container, and **never pulled v2.8.5** — against the spike's measurement of **18.3 s with a
pull** for the identical sequence before the change.
**Why bentopdf:** deployed on demo-hp only, file-based, with no database and no volume, so no data
anywhere could be touched. Evidence:
`felhom.eu/documentation/tests/VALIDATION-update-slice3-2026-09-06.md`.
## catalog_since on all 53 apps — how long has a newer pin been sitting here? (2026-09-02, update arc slice 2)
**One new optional key in every `.felhom.yml`, no compose file changed, no image moved.**
`catalog_since: "YYYY-MM-DD"` records the date THIS repo last changed that app's pinned images. The
controller (v0.233.0) compares what a box is ACTUALLY running against what the template now pins and
renders one Hungarian badge — *"Naprakész"* or *"Frissítés elérhető — 45 napja"*. **No version number
is shown to the customer anywhere** (operator ruling, 2026-09-02: a household cannot act on
`26.05.2`), so there is deliberately no `version:` key to keep in sync alongside this one.
**Backfilled from this repo's own git history**, not typed by hand: for each app, the newest commit
whose set of `image:` values differs from its parent's. Four anchors confirmed against the spike's
own numbers — nextcloud `5e2c1ae`, grafana `b789acc`, calcom `147cee7`, vikunja `3fa63cd`, all
2026-07-18 — and six apps re-checked against `git log` by hand (bentopdf, plex, crafty-controller,
homebox, bookstack, wanderer).
**Two commits were EXCLUDED BY HASH and the reason is the point:** `214d448` and `30bd892`, the
bentopdf pin move and its same-hour revert, are recorded above as a spike MEASUREMENT and not a
release. Counting them would have dated bentopdf 2026-09-02 for a pin that has not actually moved
since `71828a8` (2026-07-12), which is when `:latest` became `v2.8.6`.
Range: 2026-02-15 (plex, still on its original pin) to 2026-07-21. Thirty-eight of the 53 sit on
2026-07-18, the catalog-wide bump.
**The rule this creates, now in `CLAUDE.md`:** any commit that changes an `image:` line must set that
app's `catalog_since` to the same day. **There is no gate enforcing it yet** — the gates runner
fetches at `--depth 1` and has no parent commit to diff against — and that gap is filed as a register
row rather than left implicit.
Gates after the change: `image-pins` **OK**. `image-resolvable` and `volume-persistence` came back
INCONCLUSIVE for environmental reasons unrelated to this change (Docker Hub throttled 6 of 65
unauthenticated manifest lookups; the volume prober's own canary needs a scratch Docker host).
## SPIKE measurement — bentopdf pin moved and reverted the same hour (2026-09-01, R-438) — NOT A RELEASE
**No template is different after this pair of commits.** `214d448` moved
`templates/bentopdf/docker-compose.yml` from `v2.8.6` to `v2.8.5`; `30bd892` reverted it. Both are on
`main` deliberately, because the measurement needed a REAL catalog change travelling the real
15-minute sync — a hand-edit on the box would have proved nothing about the syncer.
**What it measured, live on demo-hp:** the sync at 17:45:17Z rewrote the DEPLOYED app's
`docker-compose.yml` to `v2.8.5` while its container went on running `v2.8.6`, and nothing told the
customer. Then a boot reconciliation started the app on `v2.8.5` with nobody pressing anything.
**Why bentopdf:** it is deployed on demo-hp only (demo-felhom runs opengist alone; Peti's box is down
with no enrolled host), and it is file-based with no database and no volume, so no data anywhere could
be touched. Evidence and the full findings:
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
absent). **All four found by accident.** The gates enforce everything else here and were the one part
nothing had checked.
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
proves the cause was the listing rather than the decoy.
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
**In this repo:** no gate changed, and that is the result. `image-pins` was decoyed and is SOUND —
a real untagged `image:` line is convicted; the `x-image:` attempt was WITHDRAWN as illegitimate
because `x-` fields are inert in Compose, so the label had no fact behind it either way.
`image-resolvable` and `volume-persistence` need a container runtime and are named in the
decoy-coverage exemption list (R-426) as UNTESTED, not as sound.
## docs — the "CI is still owed" claim was stale; corrected (2026-08-06, R-229 part 2) — no version bump
**One sentence, no code.** This file asserted that continuous integration was still owed
(`felhom.eu` `OPEN-ITEMS.md` R-168). **R-168 was CLOSED on 2026-08-02** — a Gitea Actions runner
re-runs each repo's gate entry point on every push and emails the operator on failure. Found while
confirming this session's own push by run ID, which is the check that caught it.
The same stale sentence was in four instruction files across all four repos and is corrected in all
four. In `felhom-agent/CLAUDE.md` it **contradicted the same file's release section**, which already
said R-168 mails the failure — a contradiction inside one instruction file, which is the exact class
the R-229 work exists to find.
### papra — the volume is mounted where the app actually writes (2026-08-03, R-156, last leg)
**The third and last of the three apps that kept their data where backups never looked.** papra
mounted `papra_data:/app/data` while the application writes to `/app/app-data`, so its database sat
in the container's **writable layer**: lost on redeploy, and tarred nightly as an empty directory
while the healthcheck stayed green.
**Decided from the IMAGE, not the README.** `docker inspect ghcr.io/papra-hq/papra:26.6.1-rootless`
gives `WORKDIR=/app` and all three data paths under `./app-data` — `DATABASE_URL=file:./app-data/db/db.sqlite`,
`DOCUMENT_STORAGE_FILESYSTEM_ROOT=./app-data/documents`, `PAPRA_CONFIG_DIR=./app-data` — and
**`/app/data` does not exist in the image at all**.
**Why the mount moved rather than the app being reconfigured.** Pointing all three env vars at
`/app/data` would have worked, but it enumerates data paths: a fourth one added upstream escapes to
the writable layer again, silently, which is this defect re-armed. Mounting the app's own data ROOT
captures every current and future path by construction.
**Precondition checked, not inherited:** papra is deployed nowhere — `docker ps -a` (including
stopped) on both demo guests, plus the hub fleet view showing two enrolled hosts and zero papra
references. Both boxes were wiped and rebuilt on 3 August, so the 2 August evidence was re-measured.
**Proven by the runtime gate, in both directions.** `check-volume-persistence.py papra` → **CLEAN**,
with its self-test passing on the same run. Red-proof: reverting the mount to `/app/data` → **BROKEN**
with the exact R-156 evidence (`DATA in the writable layer at /app/app-data/db`,
`declared volume /app/data is EMPTY`). Full `catalog_gates.py papra`: all three gates OK.
**Note for the next run of that gate:** it needs **root** (it reads `/var/lib/docker/volumes`, mode
`drwx--x---`; as a normal user its own canary fails UNDETERMINED and it correctly refuses a verdict),
and it should be scoped to the app touched — unscoped it deploys all 53 templates.
## CI — the static catalog gate runs on every push (2026-08-02, R-168)
**No version bump, no build, no deploy** — this adds a workflow file only. Stated explicitly so the
omission reads as a decision rather than a miss.
**`.gitea/workflows/gates.yml` (new).** Triggers on `push`, `runs-on: felhom-gates`, obtains the
source with a shallow `git fetch` of the **exact pushed SHA** from the in-cluster Gitea Service, and
runs this repo's entry point with `--fast` — nothing else. **No `uses:` step anywhere**: JavaScript
actions need a node runtime the host-mode runner does not have, and probe P3 measured a plain
`git fetch` as sufficient. No `|| true`; the entry point's exit code IS the job's result.
**It REPORTS, it cannot REFUSE**, and the workflow header says so: this repo pushes straight to
`main` with no pull request, so there is no merge for a status check to stand at. The refusing half
is `.githooks/pre-push`, which is per-clone and `--no-verify`-able; this half notices when that was
skipped. Making CI blocking needs branch protection plus a PR workflow → felhom.eu `OPEN-ITEMS.md`
R-169, an operator decision.
**A failed run emails the operator** via Resend and prints the provider's accepted id, because probe
P5 measured that Gitea itself sends nothing at all on a failed run. Demonstrated end to end on a real
red run (`RESEND-ACCEPTED id=…`), not assumed. Full detail:
`felhom.eu/documentation/audits/SPIKE-ci-runner-2026-08-02.md`.
**`--fast` only, and that is the point.** `check-image-pins.py` runs; `check-image-resolvable.py`
(network) and `check-volume-persistence.py` (Docker, minutes per app) do **not**. CI that pulls 53
images on every push gets disabled, and the bypass becomes the habit. Measured in the first run:
`image-pin gate OK — 53 templates, 0 unpinned images`, with both runtime gates announced as skipped
and their own output absent from the log. They remain deliberate periodic runs.
No sibling clone is needed here — unlike the controller and the agent, `catalog_gates --fast` does
not invoke the shared reuse checker.
## 2026-08-02 — `--fast` for the pre-push hook (no version: this repo carries none)
**`scripts/catalog_gates.py --fast`** selects only gates that touch no network and no container
runtime. Today that is gate 1, `check-image-pins.py`. `check-image-resolvable.py` (network) and
`check-volume-persistence.py` (Docker, minutes per app) are **not** in it, and the skip is
**announced**, with the reason and with what still owes a periodic run — a silently narrowed run
reads as "covered everything" when it did not. Default behaviour with no flag is unchanged.
**Why the runtime gates are never in a hook.** A push that pulls images and starts containers gets
bypassed within a week, and the bypass becomes the habit. They stay deliberate periodic runs: the
start of a catalog campaign, before a publish train that vouches the catalog, and whenever a
template's `volumes:` block or image tag changes — on a scratch host, never a customer box.
**`.githooks/pre-push` (new)** runs `catalog_gates.py --fast` and refuses the push. It is per-clone
(`git config core.hooksPath .githooks`) and `git push --no-verify` bypasses it on purpose; both
limits are written into the hook. This is R-161's convention half made automatic-ish; the
unbypassable half is CI, now tracked as `felhom.eu` `OPEN-ITEMS.md` **R-168**.
**`scripts/test_catalog_gates.py` (new, 5 tests)** pins `--fast`'s CONTENT, not just its exit code:
the static gate's own stdout must appear (an inert runner prints the summary while calling nothing),
the runtime gates' must not, the skip must be announced, and the no-flag path must still select all
three. Red-proofed with an inert `run_gate`.
## 2026-08-02 — one entry point for the catalog's gates (R-161 ruling)
`scripts/catalog_gates.py` runs all three gates — image-pins, image-resolvable, volume-persistence —
and exits non-zero if any fails. Mandated in `CLAUDE.md` the way `felhom.eu/scripts/site_gates.py` is:
**run it after any template change**, naming the app(s) you touched.
**Operator ruling, recorded because the alternatives were rejected for measured reasons.**
Controller-side enforcement at template load was rejected: such a check can only read the file, and a
static audit of all 53 templates reports the catalog clean **including papra** — it would pass on the
exact defect it exists to catch. CI was rejected for now: neither repo has any, and there are no users
yet. What was chosen copies the shape that demonstrably works here — of this project's gates, the only
ones that ever get run are the ones with a single entry point named in a CLAUDE.md; `site_gates.py` is
run, and R-29's three orphaned gates are named nowhere and have stopped nothing.
Behaviour: `0` all clean · `1` convicted · `2` UNDETERMINED, **never a pass**; a conviction outranks an
undetermined result in the summary so the reader knows which they have. Gate output is streamed, not
captured — a runner that swallows diagnostics makes a conviction unreadable. Scoping passes app names
through to the two gates that accept them; with no names the runtime gate deploys every template and
belongs on a scratch host.
**R-161 stays OPEN at reduced scope:** this is convention, run by a person. Real automatic enforcement
is owed when a second person touches templates.
Verified: `image-pins` passes standalone (53 templates, 0 unpinned); the unknown-option path exits 2;
the aggregation was unit-checked over five gate-code combinations. **The runtime leg was deliberately
NOT executed on DooPlex** — it deploys templates via `docker compose`, and DooPlex is the recovery
chain; it belongs on a scratch host.
## 2026-08-02 — persistence sweep: does every app's data land in a folder the template preserves?
Campaign 10's R-156 found papra writing its database into the container's writable layer while the
volume the template preserves stayed empty — so its backup completed, verified, and contained
nothing. papra was never the point: **nothing anywhere checked that the folder a template preserves
is the folder the app writes to**, across 53 templates. All 53 have now been measured live.
**Result: 43 CLEAN · 3 BROKEN · 7 UNDETERMINED.** Full report and per-app evidence:
`audits/persistence-sweep-2026-08-02/`.
**New gate — `scripts/check-volume-persistence.py`, the third and the only RUNTIME one.**
The two image gates are static, and **this defect class is invisible to static analysis** — measured,
not assumed: a static audit of all 53 composes (every declared volume attached, no anonymous mounts,
no stray host binds) reports the catalog clean *and reports papra clean*. papra's compose is
well-formed; only its behaviour is wrong. So the gate deploys each template, exercises it into
writing data, and compares where the data landed with what is mounted. Exit **0** all clean /
**1 REFUSED** / **2** undecided. `UNDETERMINED` is exit 2 and is never a pass.
It **refuses to report at all** unless it has just re-proven itself in both directions against two
canary templates built from a purpose-made image reproducing papra's ownership shape — the pair
differ only in which path the volume mounts at, so every run carries a live demonstration of R-156
and of its fix. A detector that flags nothing turns an unexamined catalog into a documented-clean one.
41 fixture tests (`scripts/test_check_volume_persistence.py`, no Docker) driving `check()` — the
function `__main__` calls — plus `rollup_diff`/`classify`. Every rule red-proofed.
**Two templates FIXED** (neither deployed anywhere in the fleet, so no data was stranded):
- **`gramps-web`** — mounted `/app/data`, `/app/media`, `/tmp`, and **`/app/data` is a path the
application never writes**. Its accounts database (`GRAMPSWEB_USER_DB_URI` → `/app/users`) and
**its family tree** (`GRAMPS_DATABASE_PATH` → `/root/.gramps/grampsdb`) both landed in the
container's writable layer: destroyed by any redeploy, absent from every backup, while
`gramps_data` was tarred nightly as an empty directory. Now persists the eight paths the image's
own environment names, matching upstream's reference compose.
- **`wishlist`** — mounted `wishlist_data:/data`, another path the app never writes. `prod.db` went
into the **anonymous** volume docker creates for the image's `VOLUME /usr/src/app/data` directive.
Anonymous volumes are absent from `ResolveDockerVolumeNames`, so `DumpAppVolumes` never backs them
up, and `compose down` + `up` orphans them — a store that survives a restart, loses on redeploy and
is never in a backup. Now mounts `/usr/src/app/data` and `/usr/src/app/uploads` per upstream.
Every corrected path is confirmed by **two independent sources** — the shipped image's own
environment/`Config.Volumes`, and upstream's reference compose — never inferred from a directory name.
**`papra` is NOT fixed — referred to the operator.** The one-line fix is prepared and proven, but
papra is live on one box, and changing the mount target makes the next `compose up -d` recreate the
container and destroy the writable layer its documents currently live in. That data is already on
borrowed time, but the fix is what *schedules* the loss. See the report §6.1 for which box it is,
how far that was determined, and the two options. No migration was written.
**7 UNDETERMINED, counted separately and never folded into CLEAN** — `bentopdf` (stateless by
design), `uptime-kuma` / `privatebin` / `recipe-importer` (write nothing until a user completes
setup), `glance` (crash-loops for want of a seeded config — pre-existing, Campaign 7 §6.2),
`plant-it` (image does not resolve; `lifecycle: abandoned`), `wanderer` (unhealthy).
`CLAUDE.md` and `REUSE.md` updated with the gate and the traps it encodes.
## 2026-07-21 (later) — app lifecycle replaces the `retired/` directory move
**The `retired/` mechanism shipped earlier today was wrong and is withdrawn.** Moving a template out
of `templates/` does un-offer it — but it also makes the controller's orphan detector see the
template as GONE for anyone already running the app, flagging their working install `Elavult` and
offering a Törlés button. Withdrawing an app must never take a working app away from a customer.
Replaced by an optional top-level `lifecycle:` field in `.felhom.yml` (controller v0.158.0):
- `available` — default. Absent or empty means this, so all existing templates are unchanged.
- `hidden` — not offered for new installs; nothing shown to anyone already running it.
- `abandoned` — not offered for new installs, and every box already running it shows a permanent
„Nem karbantartott" badge plus a notice that updates and security fixes will no longer arrive.
Deployed instances keep full function in every state; the controller refuses a deploy of a
non-available template server-side. An unknown value degrades to `available` with one WARN.
- **`plant-it` returns to `templates/`** with `lifecycle: abandoned` — the first user of the
mechanism, and the case that motivated it. Its compose is deliberately unchanged: it pins
`msdeluise/plant-it:0.10.0`, a repository that does not exist (the real one is `-server`), and the
app is not installable, so rewriting it would imply it is. `retired/` is removed.
- **The resolvability gate is now lifecycle-aware.** Non-available apps are skipped by default and
REPORTED, not silently dropped; `--all` includes them. An abandoned app's dead image is the
expected end state, not a finding — counting it would leave the gate permanently red for something
nobody intends to fix, and a gate that is always red is a gate nobody reads. 6 new fixture tests
(19 total), including one asserting an all-skipped run is a pass rather than an error.
Catalog is back to **53 apps** (52 offered + plant-it abandoned).
## 2026-07-21 — catalog honesty: wanderer re-pinned, plant-it retired, and a standing rot gate (R-41 slice 1)
Campaign 7 left two apps sitting behind a working "Telepítés" button with images that did not
resolve at all, recorded as findings rather than fixed. Both are now diagnosed rather than hidden,
and the class of defect gets a gate so it cannot recur silently.
**wanderer — RE-PINNED. The project is alive; the template was pointing at a ghost.**
`ghcr.io/flomp/wanderer:0.16.0` does not resolve because upstream did three things at once: split
the app into two images, moved registry, and renamed the GitHub org (Flomp → open-wanderer). Current
shape, taken from upstream's own compose at tag v0.20.0 (2026-07-07):
- `flomp/wanderer-web:v0.20.0` — the SvelteKit web app, port 3000, `curl` on PATH.
- `flomp/wanderer-db:v0.20.0` — PocketBase, port 8090. Built FROM `scratch`: no shell, no package
manager, a static curl baked in at `/curl` — hence the absolute-path healthcheck.
- `getmeili/meilisearch:v1.36.0` — still a required sidecar; both other services wait on its health.
**Pinned DOWN from the v1.49 Campaign 7 had set**, per the R-42 ruling: a sidecar pin follows the
app template's own proposed pin, never the newest tag independently.
- **New required volume** `/data/plugins` on the db — v0.20.0 moved the Strava/Komoot/Hammerhead
integrations into a WASM plugin sandbox that lives there.
- **New: a second hostname** (`SUBDOMAIN_DB`, default `hike-db`). `PUBLIC_POCKETBASE_URL` is a
browser-side variable — the user's browser talks to PocketBase directly, so it cannot be an
internal address. Upstream's own proxy example uses two hostnames for the same reason.
- New generated secret `POCKETBASE_ENCRYPTION_KEY` (`hex:16` → exactly the 32 characters upstream
requires). `mem_limit` 384M → 1024M, matching the sum of the three services.
**plant-it — RETIRED to `retired/plant-it/` (operator ruling 2026-07-21).** The pin was only
slightly wrong — the repository is `msdeluise/plant-it-server`, and `0.10.0` was the right version —
but correcting the name would have been the wrong fix. Upstream has **discontinued self-hosting**:
`backend/` and `deployment/` are deleted from `main`, the project is now an Android app on
F-Droid/Obtainium, and the last server image was pushed **2024-12-10** (a security-frozen Spring
Boot 3.4.0). It also requires **MySQL 8.0 + Redis**, which the template never had — its header
claimed "Database: None (file-based)", which was never true. Ruling: do not ship unmaintained
software to customers. Retirement is reversible (`git mv retired/plant-it templates/plant-it`);
nothing is deleted. Catalog is now **52 apps**.
**`scripts/check-image-resolvable.py` — R-41 slice 1: the standing rot gate.** `check-image-pins.py`
is syntactic and proves only that a template pins *something* concrete; it cannot see that the thing
is gone. This resolves every unique pin with `docker manifest inspect`, one image at a time, and
exits 0 / 1 (GONE) / 2 (inconclusive). Two traps are encoded in it, both observed live during this
change:
- `docker manifest inspect` prints `toomanyrequests: …` and **still exits 0** — the same
exits-0-on-failure shape as the ISO tooling's `validate-answer`, so stderr is checked even on rc=0.
- The inverse, which the first full sweep actually did: it called **24 of 65 pins dead**, including
`postgres:16-alpine` and `redis:7-alpine`, purely because Docker Hub throttled it partway through.
Ambiguity now resolves to INCONCLUSIVE, never to an accusation — a gate that cries wolf gets
ignored, and then it protects nothing.
14 fixture tests (`scripts/test_check_image_resolvable.py`), no network — the resolver is injected.
## 2026-07-19 — docs: workspace-root pointer follows the CC move to DooPlex
**Docs only, no template change.** Claude Code now runs on DooPlex (192.168.0.180, Debian 13)
instead of the Windows workstation. `CLAUDE.md`'s cross-repo pointer becomes
`/mnt/5_hdd/felhom.eu/git/CLAUDE.md`. This repo carried **no other** environment-specific content —
it was the only one of the four that needed nothing else.
## 2026-07-19 — CAMPAIGN 7: full catalog sweep (53/53 apps deployed + validated on the demo box)
Every app in the catalog was bumped to its newest stable upstream tag where one existed, then
**actually deployed** through the controller's real endpoints on the demo box (controller 0.146.0),
validated (all containers healthy, HTTP through the real Traefik ingress, log scan), and removed
again through the real delete flow. Full evidence + result matrix:
`felhom.eu/documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md`.
**Result: 45 apps pass end-to-end, 4 do not, 1 is not automatable (plex needs a real PLEX_CLAIM).**
**Version bumps** — ~40 templates moved to current upstream, 15 of them across a major
(bookstack 25.02→26.05, immich v2→v3, calcom v4→v6, nextcloud 31→34, grafana 11→13, n8n 1→2,
outline 0.82→1.9, vikunja 0.24→2.3, tandoor 1→2, romm 4→5, radarr 5→6, privatebin 1→2,
onlyoffice 8→9, claper 1→2, gramps-web v24→v25). `uptime-kuma` moved off the floating `:2` tag
to `2.4.0`. **DB/cache sidecar majors were deliberately NOT bumped** — rationale in the campaign
doc §4 (a DB major is the application's decision, and `postgres:16-alpine` already tracks 16.x).
**13 template fixes, every one live-re-validated:**
- **7 broken healthchecks.** This is not cosmetic: Traefik will not route to an `unhealthy`
container, so a probe that cannot run makes the app return **404 to the customer while it serves
200 on its own port**. adventurelog (wget in a distroless image → Node-exec at an absolute path),
emby (curl absent, BusyBox only), papra + wishlist (node-only images), homebox (`--spider` sends
HEAD, endpoint answers 405 to HEAD / 200 to GET), zipline (v4 renamed `/api/health` →
`/api/healthcheck`), tandoor (`start_period` too short for gunicorn).
- **5 apps that had NEVER been deployable** and were fixed: papra (missing required `AUTH_SECRET`,
now a generated `data_key` secret), zipline (v4 `CORE_DATABASE_URL` → `DATABASE_URL`), wishlist
(dead Docker Hub image → followed upstream to `ghcr.io/cmintey/wishlist:v0.66.0`), homebox
(upstream dropped the `v` tag prefix + new required `HBOX_AUTH_API_KEY_PEPPER`), wger (2.6 needs
the full `DJANGO_DB_*` set and listens on :8000, not :80 — the Traefik port was wrong too).
- **4 memory/OOM corrections proven by a live OOM:** gramps-web 384M→1024M, n8n 512M→1536M
(V8 heap), rallly 256M→768M, tandoor 512M→1024M (+ its `mem_limit` sum was already wrong).
- **gokapi reverted v2.2.4 → v1.9.6**: v2 refuses to run against the seeded ConfigVersion-21
config and demands an intermediate v2.0.0 pass, even on a fresh deploy. Shipping it would have
broken every new gokapi deploy. Needs a dedicated v2 config-migration task.
**Still failing (recorded, not fixed):** `glance` (needs a seeded `glance.yml`; PROVEN pre-existing —
the pre-campaign v0.7.4 pin fails identically), `gokapi` (above), `plant-it` and `wanderer`
(their images do not resolve at all — neither the new tag nor the one the catalog already shipped).
## 2026-07-14 — backup classification `backup:` blocks for the 13 bind-bearing apps (controller v0.132.0)
Adds the referential-coupling `backup:` classification block to every catalog app that binds
`${HDD_PATH}`/`${USERDATA_PATH}` (13 apps: immich, paperless-ngx, nextcloud, calibre-web,
audiobookshelf, komga, navidrome, radarr, sonarr, emby, jellyfin, plex, romm). Each block lists its
`userdata:`/`hdd:` binds with a `class ∈ {mandatory, optional, excluded}` (COUPLED /
DECOUPLED-precious / DECOUPLED-bulk); classes are operator-ruled (Viktor, 2026-07-14) + spike SQ2.
Requires **controller v0.132.0**, which parses + validates these blocks (Task 2 of the
backup-classification-redesign arc, `felhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md`).
The classification is **INERT** — no backup tier changes behavior yet; Task 3 (tier policy engine)
and Task 4 (manual `.fab` UI) consume it. All 13 blocks were verified against the shipped controller
parser: parse-clean, every bind resolves `explicit` to its ruled class (zero validation errors).
Notes: audiobookshelf `media/audiobooks` = **optional** (consistency with komga/romm curated media;
the spike proposed excluded — PENDING a Viktor veto). radarr/sonarr `downloads` = excluded (transient
cross-app queue). emby/jellyfin/plex `media` = excluded (`:ro` readers; state in volumes). The 42
volume-only apps get no block (classification moot — state rides in the recovery unit's volume dumps).
## 2026-07-12 — image pinning sweep: `:latest` eliminated from all templates (5 pins) + standing gate
A catalog sweep found 5/53 templates with unpinned images. Beyond version discipline, `:latest`
breaks restore fidelity: the controller's recovery-unit `ImagePins` pins the *tag*, so restoring a
`:latest` app re-pulls whatever `:latest` means at restore time — potentially schema-incompatible
with the data being restored. Rule applied: a deployed app pins to the digest it is RUNNING
(pin ≠ upgrade); undeployed apps pin to the verified upstream stable. All five pins are
digest-identical to what `:latest` resolved to on 2026-07-12 — a pure no-op for running apps.
| App | Old | New | Evidence |
|-----|-----|-----|----------|
| bentopdf | `ghcr.io/alam00000/bentopdf:latest` | `:v2.8.6` | digest == latest (`eaeea1e4…`); undeployed |
| calibre-web | `crocodilestick/calibre-web-automated:latest` | `:v4.0.6` | digest == RUNNING image on demo 9201 (`c31a738b…`) |
| papra | `ghcr.io/papra-hq/papra:latest` | `:26.6.1-rootless` | latest == the -rootless variant (`a7a42e22…`); `-root` differs — variant preserved |
| recipe-importer | `gitea.dooplex.hu/admin/recipe-importer:latest` | `:v0.9.11` | tag pre-existed in registry, digest == latest (`f3cb617c…`) — no retag needed |
| termix | `ghcr.io/lukegus/termix:latest` | `:2.5.0` | digest == latest == release-2.5.0 (`4d337131…`); undeployed |
- New rerunnable gate `scripts/check-image-pins.py`: fails on `:latest`/`dev`/`nightly`/`edge`/
`main`/`master` AND on untagged image refs (implicit :latest); `@sha256:` digests count as pinned.
Red-proofed both shapes (revert→exit 1→restore).
- Standing rule added to `CLAUDE.md` (never :latest / untagged; deployed apps pin to running digest).
- `templates.json` carries no image strings (legacy metadata only) — untouched.
- Fleet caveat: non-deployment of bentopdf/papra/termix verified on demo 9201 only; felhotest
unreachable + Peti's box offline at sweep time (operator approved proceeding — pins are
digest-equal to latest, so worst case equals the status quo).
## 2026-07-06 — healthcheck sweep: `localhost` → `127.0.0.1` across all 48 templates
Escalation of the re-run vaultwarden observation
(`felhom.eu/documentation/audits/RERUN-p1p3-2026-07-06.md`) from an instance to a **class**: 48/53
templates used `localhost` in their docker healthcheck `test:` line. BusyBox `wget` (and the node /
python / curl one-shot forms, incl. mealie's `socket.create_connection`) resolve `localhost`→IPv6
`::1` with no cross-address-family fallback, so an IPv4-only-binding app reads docker-`unhealthy`
while fully serving. Mechanical sweep `localhost`→`127.0.0.1`, scoped strictly to the healthcheck
`test:` lines (diff-reviewed: no app env/config/label line changed; `.felhom.yml` files were already
clean). Industry practice — never `localhost` in container healthchecks. New REUSE.md convention row.
## 2026-07-06 — vaultwarden F1 fix: _ENABLE_SMTP boot-gate (campaign finding, pilot-blocking)
The no-mercy campaign (felhom.eu `audits/CAMPAIGN-nomercy-2026-07-06.md`, finding F1) proved that a
FRESH vaultwarden deploy with app-email off — the default state — crash-loops: the template always
defines `SMTP_HOST=${SMTP_HOST:-}` / `SMTP_FROM=${SMTP_FROM:-}`, and vaultwarden treats a
defined-but-EMPTY env var as "set", so its config validation (`smtp_host.is_some() ==
smtp_from.is_empty()`) errors out and the process exits. The old comment ("empty SMTP_HOST = mail
stays disabled") was wrong for this image. Empirically proven on the pinned
`vaultwarden/server:1.33.2-alpine` (probe P1: defined-empty pair → exact campaign error, exit 12;
P2: `_ENABLE_SMTP=false` + same empty pair → boots; P3: `_ENABLE_SMTP=true` + host+from → boots).
Fix: gate the whole SMTP group with vaultwarden's own `_ENABLE_SMTP` flag — compose default
`false` (validation skipped, mail off, clean boot), flipped to `"true"` by the app-email injection
via `smtp_mapping.extra` (no controller change needed — `extra` already rides `smtpEnv`). The ON
path is byte-identical to the previously send-tested state plus the flag.
Sweep note (no edits): the other five smtp-mapped templates (calcom, gitea, mealie, nextcloud,
rallly) are boot-proven tolerant of defined-empty mail env — all ran healthy as fresh email-off
deploys during the campaign; gitea's `GITEA__mailer__SMTP_ADDR=${...:-}` pattern likewise.
Vaultwarden was the only strict image. New REUSE.md trap row: strict images need an enable-flag
gated `false` in compose + `"true"` in `smtp_mapping.extra`; boot-prove fresh email-off deploys.
## 2026-07-03 — sparkyfitness FINALIZED + live-validated (both VERIFY markers resolved); REUSE probe-naming row
The first worked example of the new `felhom-app-catalog` skill (felhom.eu). Both
`VERIFY-BEFORE-FINALIZE` healthcheck guesses resolved by inspecting the real images on the demo box:
frontend (Alpine/nginx) HAS BusyBox wget → drafted `wget --spider :80/` probe confirmed + kept;
server HAS node v24.17.0 → node-exec `:3010/api/health` probe confirmed (path proven live:
`{"status":"UP"}`). Frontend `container_name` renamed → `sparkyfitness` (= the stack name): the
controller-side probe dials the exact-name container, fallback is the FIRST prefix match (could be
the DB) — new REUSE.md §2 "Probe-container naming" row records the convention (verified in
felhom-controller healthprobe.go). Mem-sum comment added (512+1024+256 = 1792M, value unchanged).
Live-validated on demo via the real dashboard UI (sync + Frissítés): 3/3 containers healthy,
controller probe `healthy: true` (http :80 → 200), `sparky.demo-felhom.eu` 200 via Traefik;
data_key secrets untouched (server/db containers not recreated). Kept deployed.
## 2026-07-03 — docs: CLAUDE.md light expansion
The minimal REUSE-rollout stub expanded to a proper (still ~30-line) CLAUDE.md: what the repo is
(one dir per app, two template files, Hungarian customer text), the push-to-main = deploy contract
(controller sync ≤15 min / manual trigger), legacy `templates.json` warning, and pointers
(REUSE.md, README format spec, the `felhom-build-deploy` skill). No template changes.
## 2026-07-03 — docs: REUSE.md introduced
Cross-repo reuse-map rollout (docs-only). New `REUSE.md`: catalog conventions verified against all
53 apps — the canonical example app (paperless-ngx), `.felhom.yml` required fields, healthcheck
family per image type (BusyBox wget / curl / Node / Python / DB sidecars), memory-limit convention,
new-app checklist, and traps (gokapi entrypoint hack, legacy templates.json). Known README drift
recorded in §6 (NOT fixed). Also a minimal `CLAUDE.md` carrying the REUSE.md pointer + maintenance
rule (full CLAUDE.md is a separate task).
## 2026-06-29 — App-email: calcom + nextcloud (tls_mode=plaintext :2526 + nextcloud split-From)
- **nextcloud** — `smtp_mapping` with `tls_mode: plaintext` (controller injects port 2526, the plaintext-only
listener) + **split From** (`from_var=MAIL_FROM_ADDRESS` + `from_domain_var=MAIL_DOMAIN` → nextcloud@felhom.eu).
Compose references the injected `${SMTP_*}`/`${MAIL_*}`. Live-confirmed: real password-reset delivered via
plaintext :2526 (Symfony Mailer never attempted STARTTLS).
- **calcom** — `smtp_mapping` with `tls_mode: plaintext` (EMAIL_SERVER_HOST/PORT, EMAIL_FROM=calcom@felhom.eu).
**Plus three pre-existing template fixes** (calcom never deployed before — the image pin was invalid):
(1) image `v4.8.7`→`v4.6.9` (the pinned tag has no published image); (2) added required `DATABASE_DIRECT_URL`
(Prisma `migrate deploy` fails without it → incomplete schema → 500s); (3) healthcheck `/api/health`→
`/api/auth/providers` (the old path 404s in v4.x → container stayed unhealthy → Traefik wouldn't route).
- Both apps point at the controller's `:2526` plaintext-only listener because their SMTP clients
opportunistically STARTTLS-upgrade and can't skip the self-signed cert — the listener simply doesn't offer
STARTTLS, so they stay plaintext (accepted on the single-tenant app bridge).
## 2026-06-29 — App-email rollout: gitea + rallly (calcom/nextcloud/immich = findings)
- **gitea 1.23.4** — added `smtp_mapping` (STARTTLS via `GITEA__mailer__PROTOCOL=smtp+starttls` +
`FORCE_TRUST_SERVER_CERT=true` to trust the shim's self-signed cert; single `GITEA__mailer__FROM`). Compose
references the injected `GITEA__mailer__*` keys; env applied every boot.
- **rallly** — added `smtp_mapping` (Nodemailer STARTTLS, `SMTP_SECURE=false` + `SMTP_REJECT_UNAUTHORIZED=false`
to accept the self-signed cert; single `NOREPLY_EMAIL`). **Also fixed three pre-existing template bugs** that
made rallly undeployable (never caught because the bad pin never ran): (1) image pin `3.12.1` doesn't exist →
`3.11.2`; (2) healthcheck used `wget`, absent from the rallly image (exit 127) → container unhealthy →
**Traefik wouldn't route it** → replaced with a Node http check; (3) added required `SUPPORT_EMAIL` + a valid
`NOREPLY_EMAIL` default (rallly refuses to boot without them).
- **Both gitea and rallly send-tested live** end-to-end (app → shim → hub → Resend): gitea password-reset
(From `gitea@felhom.eu`) and rallly registration code (From `rallly@felhom.eu`) both delivered.
- **NOT wired — reported as findings** (`felhom.eu/documentation/audits/FINDING-app-email-rollout-2026-06-29.md`):
- **cal.com v4.8.7** — hard-codes TLS `rejectUnauthorized:true` with no override; opportunistic STARTTLS
against the self-signed shim fails. Needs a non-STARTTLS-advertising plaintext listener (mechanism change).
- **nextcloud 31** — no cert-skip env (same opportunistic-STARTTLS gap) **and** a split From
(`MAIL_FROM_ADDRESS`+`MAIL_DOMAIN`) the single-`from_var` mapping can't express.
- **immich v2.5.5** — no SMTP env vars at all; config is admin-UI/DB or an `IMMICH_CONFIG_FILE` JSON. Does not
fit env-injection; left for a future config-file-injection mechanism (or manual admin-UI setup).
## 2026-06-29 — App-email: smtp_mapping for Vaultwarden + Mealie
- Added the `smtp_mapping` block to `templates/vaultwarden/.felhom.yml` and `templates/mealie/.felhom.yml`,
enabling managed outbound email (app → in-controller shim → hub → Resend) for the two spike-proven apps
(`SPIKE-smtp-app-relay-2026-06-28`). The controller injects `SMTP_*` at deploy/redeploy when app-email is
on (global + per-app); the From address is `<app>@felhom.eu`. SMTP auth creds are intentionally left unset
(the shim accepts no-auth on the Docker network).
- **Vaultwarden:** STARTTLS (`SMTP_SECURITY=starttls`) + `SMTP_ACCEPT_INVALID_CERTS/HOSTNAMES=true` to
accept the shim's self-signed cert.
- **Mealie:** plaintext (`SMTP_AUTH_STRATEGY=NONE`) on :2525 — Mealie has no accept-invalid-cert option, so
STARTTLS to a self-signed shim would fail; plaintext to the Docker-network-only shim is the spike-validated
mode.
- Both `docker-compose.yml` files now reference the injected `${SMTP_*}` keys (with harmless defaults) so the
values reach the container; empty `SMTP_HOST` keeps mail disabled when the toggle is off.
- Documented the `smtp_mapping` pattern in `README.md` so further apps are easy adds.
## 2026-06-28 — Add SparkyFitness (v0.17.2) — nutrition/workout tracker
- New app `templates/sparkyfitness/{docker-compose.yml,.felhom.yml}`: a self-hosted nutrition/calorie +
workout/weight tracker (alternative to wger). Three containers — nginx **frontend** (SPA :80, the sole
Traefik ingress, proxies `/api`+`/uploads` internally) + Node **server** (:3010) + dedicated
**postgres:15-alpine**. Server + DB stay on the internal network with no Traefik labels.
- **Native email/password auth** (no OIDC/Authentik — that's DooPlex-specific); subdomain `sparky`
(deliberately ≠ wger's `fitness` to avoid a Host() collision). `pi_compatible: false`, `needs_hdd: false`.
- **Two DB roles**: `sparky` (POSTGRES superuser, runs init/migrations) + `sparkyapp` (limited app role the
server auto-creates on first boot) — separate `DB_PASSWORD`/`APP_DB_PASSWORD`. `PGDATA` in a `pgdata`
subdir of the named volume. Four auto-generated, `locked_after_deploy` secrets; `API_ENCRYPTION_KEY` +
`BETTER_AUTH_SECRET` carry `data_key: true` (restore recovers, never regenerates — both are 64-char hex).
- Transcribed from the validated k3s manifest `homelab-manifests/workout-system/sparkyfitness.yaml`
(pinned image tags, two-DB-role model, never-change crypto keys, `/api/health`, pg15 + PGDATA subdir).
- **Image-probe findings (build server, v0.17.2):** server keeps the `node -e` `/api/health` probe (node
present); frontend keeps the `wget --spider` probe (both `wget` and `curl` present). No probe changes needed.
- **Live-validated on guest 9201 (controller v0.87.0):** synced via "Sablonok frissítése"; deployed through
the real dashboard flow (Domain auto, Subdomain `sparky`, 4 secrets auto-gen). All 3 containers healthy;
server log shows clean migrations + `sparkyapp` role created + RLS applied, no crash loop, no uploads
EACCES; `GET /api/health` through the public edge returns `{"status":"UP"}`; login/register page serves
over a valid TLS cert at `https://sparky.demo-felhom.eu`.
## 2026-06-26 — crafty-controller: image bump 4.4.8→4.10.7 + publish Java port range + connection guidance
- **Image bump** `crafty-4:4.4.8` → `4.10.7` (latest stable; 4.10.8/4.11.0 don't exist in the registry).
6 minor versions of fixes incl. security CVEs. **Java 25 verified present** in 4.10.7
(`/usr/lib/jvm/java-25-openjdk-amd64`, default `java -version` = openjdk 25.0.3; 8/11/17/21 also
available) — so the latest-Minecraft (`26.x`, needs Java 25) blocker is resolved. Healthcheck + Traefik
https-backend labels unchanged (Crafty still serves HTTPS on 8443).
- **Published the Java game-port range** `25565-25575:25565-25575` (TCP, 11 ports = up to 11 Java
servers; first server 25565, rest 25566–25575). No `network_mode: host` (would break Traefik routing).
Bedrock UDP 19132 intentionally out of scope.
- **App-page guidance** (`.felhom.yml` first_steps + prerequisites): how to set the server port within
25565–25575, how to connect on the LAN (manual IP:port — "scan for LAN" won't auto-list), and that
internet access needs operator port-forwarding. (Static text — can't show the live LAN IP.)
- **Live-verified on guest 9201:** 4.10.7 healthy; public URL 302; the guest's bridged LAN IP
`192.168.0.121` reaches the real Crafty "test" server on `25565` (TCP OPEN + Minecraft SLP handshake
returns JSON status); `:25575` reachable, `:25600` closed (negative control). In-place upgrade preserved
the admin, the operator's configured MFA, and the test server.
- **Correction (earlier draft was wrong):** an earlier note here claimed the upgrade "locked out the
admin (TOTP)." That was a misdiagnosis — the `totp_data` row + recovery codes were **operator-configured
MFA**, so the 401 on a password-only login was correct behaviour, NOT an upgrade bug. There is **no
upgrade regression**; the bump preserves data and MFA correctly.
## 2026-06-26 — crafty-controller: seed a felhom-generated admin password (replaces Crafty's ugly random one)
- **crafty-controller**: instead of reading Crafty's auto-generated (long, symbol-laden) random admin
password, we now **seed** a clean felhom-generated one — same pattern as gokapi, so initial passwords are
consistent across the catalog.
- Crafty's image ships `app/config_original/default.json = {"username":"admin","password":"crafty"}`;
"crafty" is 6 chars < Crafty's 8-char minimum, so Crafty rejected it and generated a random password.
- New `CRAFTY_PASSWORD` deploy field (`type: password`, `generate: password:24`, locked after deploy —
mirrors gokapi's `GOKAPI_PASSWORD`). The compose **entrypoint** overwrites the `default.json` template
with this password before the launcher runs; on fresh install Crafty creates the `admin` user with it.
- `initial_credentials.file` repointed `default-creds.txt` → `default.json` (same json/username/password
keys), so the controller's app-page "Kezdeti belépési adatok" card shows the **seeded** password — the
customer sees the same value at deploy time and on the app page.
- Catalog-only change (reuses felhom-controller v0.84.0's initial_credentials reader + the gokapi-style
seed). Requires a fresh install to take effect (the seed is only read on first run).
## 2026-06-26 — crafty-controller: surface the auto-generated initial admin password on the app page
- **crafty-controller**: Crafty writes a random admin password to `/crafty/app/config/default-creds.txt`
at first boot (its built-in default is rejected as "too short"). Customers had to read the container
logs to find it. Added an `initial_credentials` block (new general felhom-controller v0.84.0 mechanism):
`file` + `format: json` + `username_key`/`password_key` + a `note`. The controller reads the file live
from the container and shows username + password (masked, reveal/copy) on the app's page under "Kezdeti
belépési adatok". Requires felhom-controller ≥ v0.84.0.
- Updated `first_steps` to point at the app page for the initial login instead of "find it in the logs".
## 2026-06-26 — crafty-controller: Traefik https backend + scoped skip-verify (fixes 502)
- **crafty-controller**: the healthcheck fix un-withheld the Traefik route, exposing a pre-existing
**502** — Traefik proxied `http://…:8443` to Crafty's **HTTPS-only** self-signed backend (Crafty serves
no plain-HTTP panel; `:8000` only redirects). Added two service labels:
- `loadbalancer.server.scheme=https` — Traefik now speaks HTTPS to the backend.
- `loadbalancer.serverstransport=insecure-skip-verify@file` — references the **named**
serversTransport defined in the controller-managed Traefik dynamic config (felhom-controller v0.83.0),
which skips verifying Crafty's per-container self-signed cert. Verification stays ON for every other
backend (scoped Option B; no global `insecureSkipVerify`). The `@file` suffix is the cross-provider
reference from the docker provider to the file-provider transport.
- Requires felhom-controller ≥ v0.83.0 (which renders the `insecure-skip-verify` transport). `port=8443`
and the router/tls labels are unchanged.
## 2026-06-26 — crafty-controller healthcheck fix (curl-absent + http-vs-TLS probe)
- **crafty-controller**: container was permanently `unhealthy` → route withheld (`routeUnpublished`).
Two independent healthcheck root causes, both fixed in one change:
- **Docker healthcheck** ran `curl -fk https://localhost:8443`, but the `crafty-4:4.4.8` image has
**no `curl` and no `wget`** (`exec: "curl": not found`, FailingStreak 150). Replaced with a
dependency-free **python3 TLS-socket** liveness probe (`/usr/bin/python3` is present): completes a
TLS handshake to `127.0.0.1:8443` (unverified context mirrors the old `-k`; Crafty's cert is
self-signed). `start_period` 30s → 60s for cold-boot headroom (cert gen + migrations).
- **Controller-side probe** (`.felhom.yml healthcheck.checks`) was `type: http` against Crafty's
**TLS-only** 8443 → `probeHTTP` sent plaintext HTTP, got a TLS record → `HealthProbe.Healthy=false`,
which `manager.go` re-applies to override Docker's verdict back to `unhealthy`. Changed `http` →
`tcp` (`probeTCP` dial succeeds against a TLS listener). Both layers had to change together.
- Live-validated on guest 9201 (`demo-felhom`): synced → recreated via the update path → Docker
`State.Health: healthy` (ExitCode 0), `health_probe.healthy: true` (tcp :8443, 5ms), http-vs-TLS
WARNs stopped, stable green 3+ min, Traefik now **publishes** the route (`crafty-controller@docker`).
- **Known follow-up (separate, out of this fix's scope):** the public URL still returns **502** — a
distinct pre-existing bug the un-withheld route exposed: Traefik proxies `http://…:8443` to Crafty's
HTTPS-only backend. Needs a Traefik HTTPS-backend + self-signed `serversTransport`
(`insecureSkipVerify`) in the controller-generated Traefik config — tracked separately.
## 2026-06-23 — gokapi: index redirect + admin username display
- **gokapi**: seed `RedirectUrl` repointed from Gokapi's GitHub default → `https://${SUBDOMAIN}.${DOMAIN}/admin`.
Gokapi's bare root `/` redirects to `RedirectUrl`; the controller's "Megnyitás" link is always the bare
subdomain root, so it was landing on Gokapi's GitHub instead of the app. Now `/` → `/admin` → login.
Applied to the live demo (config.json edit + restart) and the seed (future deploys).
- **gokapi**: added `app_info.default_creds` ("Felhasználó: admin · jelszó a Beállítások oldalon") so the
app-info page shows the initial admin user like other apps; fixed `first_steps` (no more setup wizard).
## 2026-06-23 — gokapi reproducible headless setup (fixes public "maintenance mode")
- **gokapi**: was stuck in "maintenance mode" on the public URL since first deploy — Gokapi's one-time
`/setup` wizard was never completed, and (verified against the docs + the v1.9.6 binary) **no Gokapi
version supports env-var headless setup** for admin credentials. Worse, the unconfigured `/setup` was
publicly reachable = an unauthenticated admin-takeover window.
- Fix: the compose `entrypoint` now seeds a `config.json` on first boot (admin user, this app's public
URL `https://${SUBDOMAIN}.${DOMAIN}/`, local storage, Encryption Level 0 so it restarts without a
prompt) with `Password`/`SaltAdmin`/`SaltFiles` cleared, then runs Gokapi's documented
`--deployment-password` one-shot to set the **felhom-generated** admin password **before** the
server starts serving. The admin account is claimed at first boot → `/setup` is never exposed.
- `.felhom.yml`: new `GOKAPI_PASSWORD` deploy field (`type: password`, `generate: password:24`,
shown to the customer, locked after deploy). Admin username is `admin`.
- Seed is pinned to Gokapi **v1.9.6** (`ConfigVersion 21`) — re-capture the seed if the image is bumped.
- Live-validated on guest 9201: fresh remove+redeploy → headless auto-config, public login works, no
maintenance page, admin claimed at first boot (browser-verified login).
## 2026-06-22 — gitea healthcheck fix (unattended test campaign)
- **gitea**: healthcheck probe repointed `/api/v1/version` → `/api/healthz` (docker HC + controller
`.felhom.yml` probe), `start_period` 30s → 90s.
- Surfaced during the Phase-2 deploy sweep: a fresh gitea reported `unhealthy` because
`/api/v1/version` returns 404 until the install wizard / INSTALL_LOCK completes, while the
container was serving fine on :3000 (`/api/healthz` → 200). Same class as the komga fix.
## 2026-06-22 — komga healthcheck fix (unattended test campaign)
- **komga**: healthcheck probe repointed `/api/v1/actuator/health` → `/actuator/health`.
- Root cause: komga's Spring Boot actuator endpoint is served unauthenticated at `/actuator/health`
(HTTP 200), while everything under the `/api/v1` prefix is auth-gated — so the old probe got
HTTP 401, `curl -f` exited 22, and the container reported `unhealthy` despite serving normally
on :25600. Diagnosed live on guest 9201 (probe matrix: `/`, `/actuator/health`,
`/api/v1/oauth2/providers`, `/login` all 200; `/api/v1/actuator/health` → 401).
- The `gotson/komga:1.20.0` image ships `curl` (verified), so the probe tool is unchanged.