3525e355a1
gates / gates (push) Successful in 0s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
793 lines
59 KiB
Markdown
793 lines
59 KiB
Markdown
## MariaDB finishes its own conversion — `MARIADB_AUTO_UPGRADE=1` on four db services, and an engine-major gate (2026-09-13, R-459 / R-469)
|
||
|
||
**Templates changed: bookstack, kimai, nextcloud, romm — the db service's `environment:` list only.
|
||
No `image:` line moved, so `catalog_since` does NOT move** (the rule ties it to an image change).
|
||
|
||
**The ruling** (operator, 2026-09-13, on `SPIKE-r459-mariadb-upgrade-2026-09-06.md`): an unconverted
|
||
MariaDB datadir is stable but never heals — the engine says `Check required!` on every start forever —
|
||
and the conversion costs ~7 s and takes its own system-table backup first. So every `mariadb:` sidecar
|
||
now carries `MARIADB_AUTO_UPGRADE=1`; `MARIADB_DISABLE_UPGRADE_BACKUP` stays unset — that backup is the
|
||
precaution. The setting is inert until an engine major actually moves.
|
||
|
||
**Proven by the harness before it shipped** (throwaway LXC 9403 on demo-hp, destroyed after;
|
||
evidence `felhom.eu/documentation/audits/r459-close-2026-09-13/harness/`):
|
||
|
||
| edge | verdict | the engine's own view AFTER (`engine_state_after`) |
|
||
|---|---|---|
|
||
| C3 (negative control) | **`failed`** — the harness still says no | — |
|
||
| E3 (app + engine 11.6 → 12.3) | `proven`, readback after = true | `12.3.3-MariaDB \| This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again. [exit=1]` |
|
||
| E3b (engine half alone) | `proven`, readback after = true | same |
|
||
|
||
The entrypoint, verbatim: `Backing up system database to system_mysql_backup_11.6.2-MariaDB.sql.zst`
|
||
→ `Starting mariadb-upgrade` → `Finished mariadb-upgrade` (6 s). `skipped due to $MARIADB_AUTO_UPGRADE`
|
||
appears **0** times in either TO log — it appeared on every start before this change. The abort still
|
||
starts and serves (spike §5.3), and the abort log still prints `MariaDB upgrade not required`, which
|
||
is the R-464 trap and not a soundness claim.
|
||
|
||
**And a rule with a gate, because the Update button still takes no backup** (R-448 not shipped):
|
||
*until Slice 4 ships, no template may move a database-engine image across a MAJOR version* — four
|
||
MariaDB, eleven PostgreSQL services, found by image name. `scripts/check-engine-major.py` is the
|
||
fourth row of `catalog_gates.py` (`--fast`, git reads only); `.githooks/pre-push` now hands it the
|
||
push range. **It needs a parent commit and CI fetches at `--depth 1`** (the R-452 gap, not re-filed),
|
||
so on a shallow clone the runner skips it out loud; the hook is where it bites. Red-proof
|
||
(`scripts/test_gate_decoys.py`, 7 cases): `mariadb:11.6 → 12.3` and `postgres:16 → 17` REFUSED naming
|
||
the rule and its expiry; `11.6 → 11.8` passes; the version moving only in a comment, in kimai's
|
||
`serverVersion=` env, in README, or on the app's own image passes. The rule's removal is tracked as
|
||
`felhom.eu` R-469 so it is a deliberate act, not a lapse.
|
||
|
||
## upgrade-test.py records the ENGINE's own view of itself (2026-09-06, R-459) — NOT A RELEASE
|
||
|
||
**Scripts only. No template changed, and deliberately so** — nothing with `MARIADB_` in it is
|
||
committed by this work; that is a fleet-wide decision the operator owns.
|
||
|
||
The harness returned **`proven`** for edge E3b while MariaDB was logging that the datadir conversion it
|
||
requires had been **skipped**. The verdict was not wrong — the app's data did survive, which is what it
|
||
asked — but nothing here could see that the engine had been left in a state the engine itself calls
|
||
incomplete. It watched the app and the migration log; **neither looks at engine state.**
|
||
|
||
`engine_state_after` now carries, per database service, the engine's own answer. On the E3b re-run:
|
||
|
||
```
|
||
verdict: proven
|
||
engine_state_after.bookstack-db.answer:
|
||
"11.6.2-MariaDB| Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required! [exit=0]"
|
||
```
|
||
|
||
**It is reported BESIDE the verdict and never folded into it.** An unconverted datadir is not known to
|
||
be a failure — `SPIKE-r459-mariadb-upgrade-2026-09-06.md` measured 5 of 5 restarts with no degradation
|
||
— so a verdict that called it `failed` would encode an unproven judgement, which is worse than
|
||
reporting a fact and letting a person read both.
|
||
|
||
**Both engines are covered and they fail differently:** MariaDB starts anyway and skips the conversion
|
||
quietly; PostgreSQL refuses to start on a datadir from an older major. "The engine would not start" is
|
||
as much an engine-state observation as "the engine says it needs a check".
|
||
|
||
## upgrade-test.py — a harness that measures whether an upgrade keeps the data (2026-09-06, R-449) — NOT A RELEASE
|
||
|
||
**No template changed. No image moved. Scripts only.** `scripts/upgrade-test.py` +
|
||
`scripts/upgrade_fixtures.py`.
|
||
|
||
Until today this project had measured **one** app upgrade out of 53 — Nextcloud, by hand, during
|
||
`SPIKE-app-update-2026-09-01` §7 — and the whole update arc was designed against that single data
|
||
point. A one-off measurement nothing repeats decays into a claim. This is the thing that repeats it.
|
||
|
||
**Method, per edge:** deploy at the FROM images → seed through the app's **own** interface → prove the
|
||
seed reads back (control C1) → swap to the TO images → ask the app for the data **again** → then put
|
||
the FROM images back and record what happens.
|
||
|
||
**Two rules, and the first is the one people get wrong.** Success is an **application-level
|
||
readback**, not file identity: `survive2.py`'s sha256+inode rule is right for a redeploy and wrong for
|
||
an upgrade, because a migration is *supposed* to rewrite files. And, carried verbatim from that same
|
||
script: *nothing is ever seeded into a volume by hand* (R-156) — every seed goes through the app's
|
||
HTTP API or its own CLI, and an app with no such route is recorded **`inconclusive`**, never faked.
|
||
|
||
**Run `C3` first, always.** Its TO image is `alpine:3.20`, which pulls cleanly and exits immediately.
|
||
It came back **`failed`** — so the harness can say no, and its greens mean something. If it ever comes
|
||
back green, nothing else in the run is evidence.
|
||
|
||
**First run measured 7 edges across 3 apps** (privatebin, docmost, bookstack) in a throwaway guest on
|
||
demo-hp, destroyed afterwards. Findings, including a real defect in this repo's own bookstack
|
||
template: `felhom.eu/documentation/audits/SPIKE-upgrade-test-2026-09-06.md`.
|
||
|
||
## LIVE-TEST for controller v0.235.0 — two pushes, both reverted the same hour (2026-09-06) — NOT A RELEASE
|
||
|
||
**No template is different after these four commits.** `dc7e548` moved bentopdf's healthcheck
|
||
interval 30s → 45s (a NON-image change); `09b4ff5` moved its pin v2.8.6 → v2.8.5; `1798ce6` and
|
||
`17cc784` reverted both. The tree is byte-identical to `8220f8d`.
|
||
|
||
**Why real catalog pushes and not a hand-edited file on the box:** controller v0.235.0 changes what
|
||
the CATALOG SYNCER does, so the only faithful test is a real change travelling the real 15-minute
|
||
cycle — the same method the 2026-09-01 spike used.
|
||
|
||
**What they measured, live on demo-hp:**
|
||
|
||
- the non-image change **reached** the pinned app on the normal cycle (`[INFO] [sync] Updated
|
||
bentopdf/docker-compose.yml`, 08:01:51Z) with the image and the container untouched — *fixes flow*;
|
||
- the image change **did not** (08:20:29Z): the live compose file still named v2.8.6 while the catalog
|
||
offered v2.8.5, and a `POST /api/stacks/bentopdf/restart` afterwards took **0.1 s**, did not recreate
|
||
the container, and **never pulled v2.8.5** — against the spike's measurement of **18.3 s with a
|
||
pull** for the identical sequence before the change.
|
||
|
||
**Why bentopdf:** deployed on demo-hp only, file-based, with no database and no volume, so no data
|
||
anywhere could be touched. Evidence:
|
||
`felhom.eu/documentation/tests/VALIDATION-update-slice3-2026-09-06.md`.
|
||
|
||
## catalog_since on all 53 apps — how long has a newer pin been sitting here? (2026-09-02, update arc slice 2)
|
||
|
||
**One new optional key in every `.felhom.yml`, no compose file changed, no image moved.**
|
||
|
||
`catalog_since: "YYYY-MM-DD"` records the date THIS repo last changed that app's pinned images. The
|
||
controller (v0.233.0) compares what a box is ACTUALLY running against what the template now pins and
|
||
renders one Hungarian badge — *"Naprakész"* or *"Frissítés elérhető — 45 napja"*. **No version number
|
||
is shown to the customer anywhere** (operator ruling, 2026-09-02: a household cannot act on
|
||
`26.05.2`), so there is deliberately no `version:` key to keep in sync alongside this one.
|
||
|
||
**Backfilled from this repo's own git history**, not typed by hand: for each app, the newest commit
|
||
whose set of `image:` values differs from its parent's. Four anchors confirmed against the spike's
|
||
own numbers — nextcloud `5e2c1ae`, grafana `b789acc`, calcom `147cee7`, vikunja `3fa63cd`, all
|
||
2026-07-18 — and six apps re-checked against `git log` by hand (bentopdf, plex, crafty-controller,
|
||
homebox, bookstack, wanderer).
|
||
|
||
**Two commits were EXCLUDED BY HASH and the reason is the point:** `214d448` and `30bd892`, the
|
||
bentopdf pin move and its same-hour revert, are recorded above as a spike MEASUREMENT and not a
|
||
release. Counting them would have dated bentopdf 2026-09-02 for a pin that has not actually moved
|
||
since `71828a8` (2026-07-12), which is when `:latest` became `v2.8.6`.
|
||
|
||
Range: 2026-02-15 (plex, still on its original pin) to 2026-07-21. Thirty-eight of the 53 sit on
|
||
2026-07-18, the catalog-wide bump.
|
||
|
||
**The rule this creates, now in `CLAUDE.md`:** any commit that changes an `image:` line must set that
|
||
app's `catalog_since` to the same day. **There is no gate enforcing it yet** — the gates runner
|
||
fetches at `--depth 1` and has no parent commit to diff against — and that gap is filed as a register
|
||
row rather than left implicit.
|
||
|
||
Gates after the change: `image-pins` **OK**. `image-resolvable` and `volume-persistence` came back
|
||
INCONCLUSIVE for environmental reasons unrelated to this change (Docker Hub throttled 6 of 65
|
||
unauthenticated manifest lookups; the volume prober's own canary needs a scratch Docker host).
|
||
|
||
## SPIKE measurement — bentopdf pin moved and reverted the same hour (2026-09-01, R-438) — NOT A RELEASE
|
||
|
||
**No template is different after this pair of commits.** `214d448` moved
|
||
`templates/bentopdf/docker-compose.yml` from `v2.8.6` to `v2.8.5`; `30bd892` reverted it. Both are on
|
||
`main` deliberately, because the measurement needed a REAL catalog change travelling the real
|
||
15-minute sync — a hand-edit on the box would have proved nothing about the syncer.
|
||
|
||
**What it measured, live on demo-hp:** the sync at 17:45:17Z rewrote the DEPLOYED app's
|
||
`docker-compose.yml` to `v2.8.5` while its container went on running `v2.8.6`, and nothing told the
|
||
customer. Then a boot reconciliation started the app on `v2.8.5` with nobody pressing anything.
|
||
|
||
**Why bentopdf:** it is deployed on demo-hp only (demo-felhom runs opengist alone; Peti's box is down
|
||
with no enrolled host), and it is file-based with no database and no volume, so no data anywhere could
|
||
be touched. Evidence and the full findings:
|
||
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
|
||
|
||
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
|
||
|
||
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
|
||
|
||
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
|
||
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
|
||
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
|
||
absent). **All four found by accident.** The gates enforce everything else here and were the one part
|
||
nothing had checked.
|
||
|
||
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
|
||
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
|
||
|
||
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
|
||
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
|
||
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
|
||
proves the cause was the listing rather than the decoy.
|
||
|
||
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
|
||
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
|
||
|
||
**In this repo:** no gate changed, and that is the result. `image-pins` was decoyed and is SOUND —
|
||
a real untagged `image:` line is convicted; the `x-image:` attempt was WITHDRAWN as illegitimate
|
||
because `x-` fields are inert in Compose, so the label had no fact behind it either way.
|
||
`image-resolvable` and `volume-persistence` need a container runtime and are named in the
|
||
decoy-coverage exemption list (R-426) as UNTESTED, not as sound.
|
||
|
||
## docs — the "CI is still owed" claim was stale; corrected (2026-08-06, R-229 part 2) — no version bump
|
||
|
||
**One sentence, no code.** This file asserted that continuous integration was still owed
|
||
(`felhom.eu` `OPEN-ITEMS.md` R-168). **R-168 was CLOSED on 2026-08-02** — a Gitea Actions runner
|
||
re-runs each repo's gate entry point on every push and emails the operator on failure. Found while
|
||
confirming this session's own push by run ID, which is the check that caught it.
|
||
|
||
The same stale sentence was in four instruction files across all four repos and is corrected in all
|
||
four. In `felhom-agent/CLAUDE.md` it **contradicted the same file's release section**, which already
|
||
said R-168 mails the failure — a contradiction inside one instruction file, which is the exact class
|
||
the R-229 work exists to find.
|
||
|
||
|
||
### papra — the volume is mounted where the app actually writes (2026-08-03, R-156, last leg)
|
||
|
||
**The third and last of the three apps that kept their data where backups never looked.** papra
|
||
mounted `papra_data:/app/data` while the application writes to `/app/app-data`, so its database sat
|
||
in the container's **writable layer**: lost on redeploy, and tarred nightly as an empty directory
|
||
while the healthcheck stayed green.
|
||
|
||
**Decided from the IMAGE, not the README.** `docker inspect ghcr.io/papra-hq/papra:26.6.1-rootless`
|
||
gives `WORKDIR=/app` and all three data paths under `./app-data` — `DATABASE_URL=file:./app-data/db/db.sqlite`,
|
||
`DOCUMENT_STORAGE_FILESYSTEM_ROOT=./app-data/documents`, `PAPRA_CONFIG_DIR=./app-data` — and
|
||
**`/app/data` does not exist in the image at all**.
|
||
|
||
**Why the mount moved rather than the app being reconfigured.** Pointing all three env vars at
|
||
`/app/data` would have worked, but it enumerates data paths: a fourth one added upstream escapes to
|
||
the writable layer again, silently, which is this defect re-armed. Mounting the app's own data ROOT
|
||
captures every current and future path by construction.
|
||
|
||
**Precondition checked, not inherited:** papra is deployed nowhere — `docker ps -a` (including
|
||
stopped) on both demo guests, plus the hub fleet view showing two enrolled hosts and zero papra
|
||
references. Both boxes were wiped and rebuilt on 3 August, so the 2 August evidence was re-measured.
|
||
|
||
**Proven by the runtime gate, in both directions.** `check-volume-persistence.py papra` → **CLEAN**,
|
||
with its self-test passing on the same run. Red-proof: reverting the mount to `/app/data` → **BROKEN**
|
||
with the exact R-156 evidence (`DATA in the writable layer at /app/app-data/db`,
|
||
`declared volume /app/data is EMPTY`). Full `catalog_gates.py papra`: all three gates OK.
|
||
|
||
**Note for the next run of that gate:** it needs **root** (it reads `/var/lib/docker/volumes`, mode
|
||
`drwx--x---`; as a normal user its own canary fails UNDETERMINED and it correctly refuses a verdict),
|
||
and it should be scoped to the app touched — unscoped it deploys all 53 templates.
|
||
|
||
## CI — the static catalog gate runs on every push (2026-08-02, R-168)
|
||
|
||
**No version bump, no build, no deploy** — this adds a workflow file only. Stated explicitly so the
|
||
omission reads as a decision rather than a miss.
|
||
|
||
**`.gitea/workflows/gates.yml` (new).** Triggers on `push`, `runs-on: felhom-gates`, obtains the
|
||
source with a shallow `git fetch` of the **exact pushed SHA** from the in-cluster Gitea Service, and
|
||
runs this repo's entry point with `--fast` — nothing else. **No `uses:` step anywhere**: JavaScript
|
||
actions need a node runtime the host-mode runner does not have, and probe P3 measured a plain
|
||
`git fetch` as sufficient. No `|| true`; the entry point's exit code IS the job's result.
|
||
|
||
**It REPORTS, it cannot REFUSE**, and the workflow header says so: this repo pushes straight to
|
||
`main` with no pull request, so there is no merge for a status check to stand at. The refusing half
|
||
is `.githooks/pre-push`, which is per-clone and `--no-verify`-able; this half notices when that was
|
||
skipped. Making CI blocking needs branch protection plus a PR workflow → felhom.eu `OPEN-ITEMS.md`
|
||
R-169, an operator decision.
|
||
|
||
**A failed run emails the operator** via Resend and prints the provider's accepted id, because probe
|
||
P5 measured that Gitea itself sends nothing at all on a failed run. Demonstrated end to end on a real
|
||
red run (`RESEND-ACCEPTED id=…`), not assumed. Full detail:
|
||
`felhom.eu/documentation/audits/SPIKE-ci-runner-2026-08-02.md`.
|
||
|
||
**`--fast` only, and that is the point.** `check-image-pins.py` runs; `check-image-resolvable.py`
|
||
(network) and `check-volume-persistence.py` (Docker, minutes per app) do **not**. CI that pulls 53
|
||
images on every push gets disabled, and the bypass becomes the habit. Measured in the first run:
|
||
`image-pin gate OK — 53 templates, 0 unpinned images`, with both runtime gates announced as skipped
|
||
and their own output absent from the log. They remain deliberate periodic runs.
|
||
|
||
No sibling clone is needed here — unlike the controller and the agent, `catalog_gates --fast` does
|
||
not invoke the shared reuse checker.
|
||
|
||
## 2026-08-02 — `--fast` for the pre-push hook (no version: this repo carries none)
|
||
|
||
**`scripts/catalog_gates.py --fast`** selects only gates that touch no network and no container
|
||
runtime. Today that is gate 1, `check-image-pins.py`. `check-image-resolvable.py` (network) and
|
||
`check-volume-persistence.py` (Docker, minutes per app) are **not** in it, and the skip is
|
||
**announced**, with the reason and with what still owes a periodic run — a silently narrowed run
|
||
reads as "covered everything" when it did not. Default behaviour with no flag is unchanged.
|
||
|
||
**Why the runtime gates are never in a hook.** A push that pulls images and starts containers gets
|
||
bypassed within a week, and the bypass becomes the habit. They stay deliberate periodic runs: the
|
||
start of a catalog campaign, before a publish train that vouches the catalog, and whenever a
|
||
template's `volumes:` block or image tag changes — on a scratch host, never a customer box.
|
||
|
||
**`.githooks/pre-push` (new)** runs `catalog_gates.py --fast` and refuses the push. It is per-clone
|
||
(`git config core.hooksPath .githooks`) and `git push --no-verify` bypasses it on purpose; both
|
||
limits are written into the hook. This is R-161's convention half made automatic-ish; the
|
||
unbypassable half is CI, now tracked as `felhom.eu` `OPEN-ITEMS.md` **R-168**.
|
||
|
||
**`scripts/test_catalog_gates.py` (new, 5 tests)** pins `--fast`'s CONTENT, not just its exit code:
|
||
the static gate's own stdout must appear (an inert runner prints the summary while calling nothing),
|
||
the runtime gates' must not, the skip must be announced, and the no-flag path must still select all
|
||
three. Red-proofed with an inert `run_gate`.
|
||
|
||
## 2026-08-02 — one entry point for the catalog's gates (R-161 ruling)
|
||
|
||
`scripts/catalog_gates.py` runs all three gates — image-pins, image-resolvable, volume-persistence —
|
||
and exits non-zero if any fails. Mandated in `CLAUDE.md` the way `felhom.eu/scripts/site_gates.py` is:
|
||
**run it after any template change**, naming the app(s) you touched.
|
||
|
||
**Operator ruling, recorded because the alternatives were rejected for measured reasons.**
|
||
Controller-side enforcement at template load was rejected: such a check can only read the file, and a
|
||
static audit of all 53 templates reports the catalog clean **including papra** — it would pass on the
|
||
exact defect it exists to catch. CI was rejected for now: neither repo has any, and there are no users
|
||
yet. What was chosen copies the shape that demonstrably works here — of this project's gates, the only
|
||
ones that ever get run are the ones with a single entry point named in a CLAUDE.md; `site_gates.py` is
|
||
run, and R-29's three orphaned gates are named nowhere and have stopped nothing.
|
||
|
||
Behaviour: `0` all clean · `1` convicted · `2` UNDETERMINED, **never a pass**; a conviction outranks an
|
||
undetermined result in the summary so the reader knows which they have. Gate output is streamed, not
|
||
captured — a runner that swallows diagnostics makes a conviction unreadable. Scoping passes app names
|
||
through to the two gates that accept them; with no names the runtime gate deploys every template and
|
||
belongs on a scratch host.
|
||
|
||
**R-161 stays OPEN at reduced scope:** this is convention, run by a person. Real automatic enforcement
|
||
is owed when a second person touches templates.
|
||
|
||
Verified: `image-pins` passes standalone (53 templates, 0 unpinned); the unknown-option path exits 2;
|
||
the aggregation was unit-checked over five gate-code combinations. **The runtime leg was deliberately
|
||
NOT executed on DooPlex** — it deploys templates via `docker compose`, and DooPlex is the recovery
|
||
chain; it belongs on a scratch host.
|
||
|
||
## 2026-08-02 — persistence sweep: does every app's data land in a folder the template preserves?
|
||
|
||
Campaign 10's R-156 found papra writing its database into the container's writable layer while the
|
||
volume the template preserves stayed empty — so its backup completed, verified, and contained
|
||
nothing. papra was never the point: **nothing anywhere checked that the folder a template preserves
|
||
is the folder the app writes to**, across 53 templates. All 53 have now been measured live.
|
||
|
||
**Result: 43 CLEAN · 3 BROKEN · 7 UNDETERMINED.** Full report and per-app evidence:
|
||
`audits/persistence-sweep-2026-08-02/`.
|
||
|
||
**New gate — `scripts/check-volume-persistence.py`, the third and the only RUNTIME one.**
|
||
The two image gates are static, and **this defect class is invisible to static analysis** — measured,
|
||
not assumed: a static audit of all 53 composes (every declared volume attached, no anonymous mounts,
|
||
no stray host binds) reports the catalog clean *and reports papra clean*. papra's compose is
|
||
well-formed; only its behaviour is wrong. So the gate deploys each template, exercises it into
|
||
writing data, and compares where the data landed with what is mounted. Exit **0** all clean /
|
||
**1 REFUSED** / **2** undecided. `UNDETERMINED` is exit 2 and is never a pass.
|
||
|
||
It **refuses to report at all** unless it has just re-proven itself in both directions against two
|
||
canary templates built from a purpose-made image reproducing papra's ownership shape — the pair
|
||
differ only in which path the volume mounts at, so every run carries a live demonstration of R-156
|
||
and of its fix. A detector that flags nothing turns an unexamined catalog into a documented-clean one.
|
||
|
||
41 fixture tests (`scripts/test_check_volume_persistence.py`, no Docker) driving `check()` — the
|
||
function `__main__` calls — plus `rollup_diff`/`classify`. Every rule red-proofed.
|
||
|
||
**Two templates FIXED** (neither deployed anywhere in the fleet, so no data was stranded):
|
||
|
||
- **`gramps-web`** — mounted `/app/data`, `/app/media`, `/tmp`, and **`/app/data` is a path the
|
||
application never writes**. Its accounts database (`GRAMPSWEB_USER_DB_URI` → `/app/users`) and
|
||
**its family tree** (`GRAMPS_DATABASE_PATH` → `/root/.gramps/grampsdb`) both landed in the
|
||
container's writable layer: destroyed by any redeploy, absent from every backup, while
|
||
`gramps_data` was tarred nightly as an empty directory. Now persists the eight paths the image's
|
||
own environment names, matching upstream's reference compose.
|
||
- **`wishlist`** — mounted `wishlist_data:/data`, another path the app never writes. `prod.db` went
|
||
into the **anonymous** volume docker creates for the image's `VOLUME /usr/src/app/data` directive.
|
||
Anonymous volumes are absent from `ResolveDockerVolumeNames`, so `DumpAppVolumes` never backs them
|
||
up, and `compose down` + `up` orphans them — a store that survives a restart, loses on redeploy and
|
||
is never in a backup. Now mounts `/usr/src/app/data` and `/usr/src/app/uploads` per upstream.
|
||
|
||
Every corrected path is confirmed by **two independent sources** — the shipped image's own
|
||
environment/`Config.Volumes`, and upstream's reference compose — never inferred from a directory name.
|
||
|
||
**`papra` is NOT fixed — referred to the operator.** The one-line fix is prepared and proven, but
|
||
papra is live on one box, and changing the mount target makes the next `compose up -d` recreate the
|
||
container and destroy the writable layer its documents currently live in. That data is already on
|
||
borrowed time, but the fix is what *schedules* the loss. See the report §6.1 for which box it is,
|
||
how far that was determined, and the two options. No migration was written.
|
||
|
||
**7 UNDETERMINED, counted separately and never folded into CLEAN** — `bentopdf` (stateless by
|
||
design), `uptime-kuma` / `privatebin` / `recipe-importer` (write nothing until a user completes
|
||
setup), `glance` (crash-loops for want of a seeded config — pre-existing, Campaign 7 §6.2),
|
||
`plant-it` (image does not resolve; `lifecycle: abandoned`), `wanderer` (unhealthy).
|
||
|
||
`CLAUDE.md` and `REUSE.md` updated with the gate and the traps it encodes.
|
||
|
||
## 2026-07-21 (later) — app lifecycle replaces the `retired/` directory move
|
||
|
||
**The `retired/` mechanism shipped earlier today was wrong and is withdrawn.** Moving a template out
|
||
of `templates/` does un-offer it — but it also makes the controller's orphan detector see the
|
||
template as GONE for anyone already running the app, flagging their working install `Elavult` and
|
||
offering a Törlés button. Withdrawing an app must never take a working app away from a customer.
|
||
|
||
Replaced by an optional top-level `lifecycle:` field in `.felhom.yml` (controller v0.158.0):
|
||
|
||
- `available` — default. Absent or empty means this, so all existing templates are unchanged.
|
||
- `hidden` — not offered for new installs; nothing shown to anyone already running it.
|
||
- `abandoned` — not offered for new installs, and every box already running it shows a permanent
|
||
„Nem karbantartott" badge plus a notice that updates and security fixes will no longer arrive.
|
||
|
||
Deployed instances keep full function in every state; the controller refuses a deploy of a
|
||
non-available template server-side. An unknown value degrades to `available` with one WARN.
|
||
|
||
- **`plant-it` returns to `templates/`** with `lifecycle: abandoned` — the first user of the
|
||
mechanism, and the case that motivated it. Its compose is deliberately unchanged: it pins
|
||
`msdeluise/plant-it:0.10.0`, a repository that does not exist (the real one is `-server`), and the
|
||
app is not installable, so rewriting it would imply it is. `retired/` is removed.
|
||
- **The resolvability gate is now lifecycle-aware.** Non-available apps are skipped by default and
|
||
REPORTED, not silently dropped; `--all` includes them. An abandoned app's dead image is the
|
||
expected end state, not a finding — counting it would leave the gate permanently red for something
|
||
nobody intends to fix, and a gate that is always red is a gate nobody reads. 6 new fixture tests
|
||
(19 total), including one asserting an all-skipped run is a pass rather than an error.
|
||
|
||
Catalog is back to **53 apps** (52 offered + plant-it abandoned).
|
||
|
||
## 2026-07-21 — catalog honesty: wanderer re-pinned, plant-it retired, and a standing rot gate (R-41 slice 1)
|
||
|
||
Campaign 7 left two apps sitting behind a working "Telepítés" button with images that did not
|
||
resolve at all, recorded as findings rather than fixed. Both are now diagnosed rather than hidden,
|
||
and the class of defect gets a gate so it cannot recur silently.
|
||
|
||
**wanderer — RE-PINNED. The project is alive; the template was pointing at a ghost.**
|
||
`ghcr.io/flomp/wanderer:0.16.0` does not resolve because upstream did three things at once: split
|
||
the app into two images, moved registry, and renamed the GitHub org (Flomp → open-wanderer). Current
|
||
shape, taken from upstream's own compose at tag v0.20.0 (2026-07-07):
|
||
|
||
- `flomp/wanderer-web:v0.20.0` — the SvelteKit web app, port 3000, `curl` on PATH.
|
||
- `flomp/wanderer-db:v0.20.0` — PocketBase, port 8090. Built FROM `scratch`: no shell, no package
|
||
manager, a static curl baked in at `/curl` — hence the absolute-path healthcheck.
|
||
- `getmeili/meilisearch:v1.36.0` — still a required sidecar; both other services wait on its health.
|
||
**Pinned DOWN from the v1.49 Campaign 7 had set**, per the R-42 ruling: a sidecar pin follows the
|
||
app template's own proposed pin, never the newest tag independently.
|
||
- **New required volume** `/data/plugins` on the db — v0.20.0 moved the Strava/Komoot/Hammerhead
|
||
integrations into a WASM plugin sandbox that lives there.
|
||
- **New: a second hostname** (`SUBDOMAIN_DB`, default `hike-db`). `PUBLIC_POCKETBASE_URL` is a
|
||
browser-side variable — the user's browser talks to PocketBase directly, so it cannot be an
|
||
internal address. Upstream's own proxy example uses two hostnames for the same reason.
|
||
- New generated secret `POCKETBASE_ENCRYPTION_KEY` (`hex:16` → exactly the 32 characters upstream
|
||
requires). `mem_limit` 384M → 1024M, matching the sum of the three services.
|
||
|
||
**plant-it — RETIRED to `retired/plant-it/` (operator ruling 2026-07-21).** The pin was only
|
||
slightly wrong — the repository is `msdeluise/plant-it-server`, and `0.10.0` was the right version —
|
||
but correcting the name would have been the wrong fix. Upstream has **discontinued self-hosting**:
|
||
`backend/` and `deployment/` are deleted from `main`, the project is now an Android app on
|
||
F-Droid/Obtainium, and the last server image was pushed **2024-12-10** (a security-frozen Spring
|
||
Boot 3.4.0). It also requires **MySQL 8.0 + Redis**, which the template never had — its header
|
||
claimed "Database: None (file-based)", which was never true. Ruling: do not ship unmaintained
|
||
software to customers. Retirement is reversible (`git mv retired/plant-it templates/plant-it`);
|
||
nothing is deleted. Catalog is now **52 apps**.
|
||
|
||
**`scripts/check-image-resolvable.py` — R-41 slice 1: the standing rot gate.** `check-image-pins.py`
|
||
is syntactic and proves only that a template pins *something* concrete; it cannot see that the thing
|
||
is gone. This resolves every unique pin with `docker manifest inspect`, one image at a time, and
|
||
exits 0 / 1 (GONE) / 2 (inconclusive). Two traps are encoded in it, both observed live during this
|
||
change:
|
||
|
||
- `docker manifest inspect` prints `toomanyrequests: …` and **still exits 0** — the same
|
||
exits-0-on-failure shape as the ISO tooling's `validate-answer`, so stderr is checked even on rc=0.
|
||
- The inverse, which the first full sweep actually did: it called **24 of 65 pins dead**, including
|
||
`postgres:16-alpine` and `redis:7-alpine`, purely because Docker Hub throttled it partway through.
|
||
Ambiguity now resolves to INCONCLUSIVE, never to an accusation — a gate that cries wolf gets
|
||
ignored, and then it protects nothing.
|
||
|
||
14 fixture tests (`scripts/test_check_image_resolvable.py`), no network — the resolver is injected.
|
||
|
||
## 2026-07-19 — docs: workspace-root pointer follows the CC move to DooPlex
|
||
|
||
**Docs only, no template change.** Claude Code now runs on DooPlex (192.168.0.180, Debian 13)
|
||
instead of the Windows workstation. `CLAUDE.md`'s cross-repo pointer becomes
|
||
`/mnt/5_hdd/felhom.eu/git/CLAUDE.md`. This repo carried **no other** environment-specific content —
|
||
it was the only one of the four that needed nothing else.
|
||
|
||
## 2026-07-19 — CAMPAIGN 7: full catalog sweep (53/53 apps deployed + validated on the demo box)
|
||
|
||
Every app in the catalog was bumped to its newest stable upstream tag where one existed, then
|
||
**actually deployed** through the controller's real endpoints on the demo box (controller 0.146.0),
|
||
validated (all containers healthy, HTTP through the real Traefik ingress, log scan), and removed
|
||
again through the real delete flow. Full evidence + result matrix:
|
||
`felhom.eu/documentation/audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md`.
|
||
|
||
**Result: 45 apps pass end-to-end, 4 do not, 1 is not automatable (plex needs a real PLEX_CLAIM).**
|
||
|
||
**Version bumps** — ~40 templates moved to current upstream, 15 of them across a major
|
||
(bookstack 25.02→26.05, immich v2→v3, calcom v4→v6, nextcloud 31→34, grafana 11→13, n8n 1→2,
|
||
outline 0.82→1.9, vikunja 0.24→2.3, tandoor 1→2, romm 4→5, radarr 5→6, privatebin 1→2,
|
||
onlyoffice 8→9, claper 1→2, gramps-web v24→v25). `uptime-kuma` moved off the floating `:2` tag
|
||
to `2.4.0`. **DB/cache sidecar majors were deliberately NOT bumped** — rationale in the campaign
|
||
doc §4 (a DB major is the application's decision, and `postgres:16-alpine` already tracks 16.x).
|
||
|
||
**13 template fixes, every one live-re-validated:**
|
||
|
||
- **7 broken healthchecks.** This is not cosmetic: Traefik will not route to an `unhealthy`
|
||
container, so a probe that cannot run makes the app return **404 to the customer while it serves
|
||
200 on its own port**. adventurelog (wget in a distroless image → Node-exec at an absolute path),
|
||
emby (curl absent, BusyBox only), papra + wishlist (node-only images), homebox (`--spider` sends
|
||
HEAD, endpoint answers 405 to HEAD / 200 to GET), zipline (v4 renamed `/api/health` →
|
||
`/api/healthcheck`), tandoor (`start_period` too short for gunicorn).
|
||
- **5 apps that had NEVER been deployable** and were fixed: papra (missing required `AUTH_SECRET`,
|
||
now a generated `data_key` secret), zipline (v4 `CORE_DATABASE_URL` → `DATABASE_URL`), wishlist
|
||
(dead Docker Hub image → followed upstream to `ghcr.io/cmintey/wishlist:v0.66.0`), homebox
|
||
(upstream dropped the `v` tag prefix + new required `HBOX_AUTH_API_KEY_PEPPER`), wger (2.6 needs
|
||
the full `DJANGO_DB_*` set and listens on :8000, not :80 — the Traefik port was wrong too).
|
||
- **4 memory/OOM corrections proven by a live OOM:** gramps-web 384M→1024M, n8n 512M→1536M
|
||
(V8 heap), rallly 256M→768M, tandoor 512M→1024M (+ its `mem_limit` sum was already wrong).
|
||
- **gokapi reverted v2.2.4 → v1.9.6**: v2 refuses to run against the seeded ConfigVersion-21
|
||
config and demands an intermediate v2.0.0 pass, even on a fresh deploy. Shipping it would have
|
||
broken every new gokapi deploy. Needs a dedicated v2 config-migration task.
|
||
|
||
**Still failing (recorded, not fixed):** `glance` (needs a seeded `glance.yml`; PROVEN pre-existing —
|
||
the pre-campaign v0.7.4 pin fails identically), `gokapi` (above), `plant-it` and `wanderer`
|
||
(their images do not resolve at all — neither the new tag nor the one the catalog already shipped).
|
||
|
||
## 2026-07-14 — backup classification `backup:` blocks for the 13 bind-bearing apps (controller v0.132.0)
|
||
|
||
Adds the referential-coupling `backup:` classification block to every catalog app that binds
|
||
`${HDD_PATH}`/`${USERDATA_PATH}` (13 apps: immich, paperless-ngx, nextcloud, calibre-web,
|
||
audiobookshelf, komga, navidrome, radarr, sonarr, emby, jellyfin, plex, romm). Each block lists its
|
||
`userdata:`/`hdd:` binds with a `class ∈ {mandatory, optional, excluded}` (COUPLED /
|
||
DECOUPLED-precious / DECOUPLED-bulk); classes are operator-ruled (Viktor, 2026-07-14) + spike SQ2.
|
||
|
||
Requires **controller v0.132.0**, which parses + validates these blocks (Task 2 of the
|
||
backup-classification-redesign arc, `felhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md`).
|
||
The classification is **INERT** — no backup tier changes behavior yet; Task 3 (tier policy engine)
|
||
and Task 4 (manual `.fab` UI) consume it. All 13 blocks were verified against the shipped controller
|
||
parser: parse-clean, every bind resolves `explicit` to its ruled class (zero validation errors).
|
||
|
||
Notes: audiobookshelf `media/audiobooks` = **optional** (consistency with komga/romm curated media;
|
||
the spike proposed excluded — PENDING a Viktor veto). radarr/sonarr `downloads` = excluded (transient
|
||
cross-app queue). emby/jellyfin/plex `media` = excluded (`:ro` readers; state in volumes). The 42
|
||
volume-only apps get no block (classification moot — state rides in the recovery unit's volume dumps).
|
||
|
||
## 2026-07-12 — image pinning sweep: `:latest` eliminated from all templates (5 pins) + standing gate
|
||
|
||
A catalog sweep found 5/53 templates with unpinned images. Beyond version discipline, `:latest`
|
||
breaks restore fidelity: the controller's recovery-unit `ImagePins` pins the *tag*, so restoring a
|
||
`:latest` app re-pulls whatever `:latest` means at restore time — potentially schema-incompatible
|
||
with the data being restored. Rule applied: a deployed app pins to the digest it is RUNNING
|
||
(pin ≠ upgrade); undeployed apps pin to the verified upstream stable. All five pins are
|
||
digest-identical to what `:latest` resolved to on 2026-07-12 — a pure no-op for running apps.
|
||
|
||
| App | Old | New | Evidence |
|
||
|-----|-----|-----|----------|
|
||
| bentopdf | `ghcr.io/alam00000/bentopdf:latest` | `:v2.8.6` | digest == latest (`eaeea1e4…`); undeployed |
|
||
| calibre-web | `crocodilestick/calibre-web-automated:latest` | `:v4.0.6` | digest == RUNNING image on demo 9201 (`c31a738b…`) |
|
||
| papra | `ghcr.io/papra-hq/papra:latest` | `:26.6.1-rootless` | latest == the -rootless variant (`a7a42e22…`); `-root` differs — variant preserved |
|
||
| recipe-importer | `gitea.dooplex.hu/admin/recipe-importer:latest` | `:v0.9.11` | tag pre-existed in registry, digest == latest (`f3cb617c…`) — no retag needed |
|
||
| termix | `ghcr.io/lukegus/termix:latest` | `:2.5.0` | digest == latest == release-2.5.0 (`4d337131…`); undeployed |
|
||
|
||
- New rerunnable gate `scripts/check-image-pins.py`: fails on `:latest`/`dev`/`nightly`/`edge`/
|
||
`main`/`master` AND on untagged image refs (implicit :latest); `@sha256:` digests count as pinned.
|
||
Red-proofed both shapes (revert→exit 1→restore).
|
||
- Standing rule added to `CLAUDE.md` (never :latest / untagged; deployed apps pin to running digest).
|
||
- `templates.json` carries no image strings (legacy metadata only) — untouched.
|
||
- Fleet caveat: non-deployment of bentopdf/papra/termix verified on demo 9201 only; felhotest
|
||
unreachable + Peti's box offline at sweep time (operator approved proceeding — pins are
|
||
digest-equal to latest, so worst case equals the status quo).
|
||
|
||
## 2026-07-06 — healthcheck sweep: `localhost` → `127.0.0.1` across all 48 templates
|
||
|
||
Escalation of the re-run vaultwarden observation
|
||
(`felhom.eu/documentation/audits/RERUN-p1p3-2026-07-06.md`) from an instance to a **class**: 48/53
|
||
templates used `localhost` in their docker healthcheck `test:` line. BusyBox `wget` (and the node /
|
||
python / curl one-shot forms, incl. mealie's `socket.create_connection`) resolve `localhost`→IPv6
|
||
`::1` with no cross-address-family fallback, so an IPv4-only-binding app reads docker-`unhealthy`
|
||
while fully serving. Mechanical sweep `localhost`→`127.0.0.1`, scoped strictly to the healthcheck
|
||
`test:` lines (diff-reviewed: no app env/config/label line changed; `.felhom.yml` files were already
|
||
clean). Industry practice — never `localhost` in container healthchecks. New REUSE.md convention row.
|
||
|
||
## 2026-07-06 — vaultwarden F1 fix: _ENABLE_SMTP boot-gate (campaign finding, pilot-blocking)
|
||
|
||
The no-mercy campaign (felhom.eu `audits/CAMPAIGN-nomercy-2026-07-06.md`, finding F1) proved that a
|
||
FRESH vaultwarden deploy with app-email off — the default state — crash-loops: the template always
|
||
defines `SMTP_HOST=${SMTP_HOST:-}` / `SMTP_FROM=${SMTP_FROM:-}`, and vaultwarden treats a
|
||
defined-but-EMPTY env var as "set", so its config validation (`smtp_host.is_some() ==
|
||
smtp_from.is_empty()`) errors out and the process exits. The old comment ("empty SMTP_HOST = mail
|
||
stays disabled") was wrong for this image. Empirically proven on the pinned
|
||
`vaultwarden/server:1.33.2-alpine` (probe P1: defined-empty pair → exact campaign error, exit 12;
|
||
P2: `_ENABLE_SMTP=false` + same empty pair → boots; P3: `_ENABLE_SMTP=true` + host+from → boots).
|
||
|
||
Fix: gate the whole SMTP group with vaultwarden's own `_ENABLE_SMTP` flag — compose default
|
||
`false` (validation skipped, mail off, clean boot), flipped to `"true"` by the app-email injection
|
||
via `smtp_mapping.extra` (no controller change needed — `extra` already rides `smtpEnv`). The ON
|
||
path is byte-identical to the previously send-tested state plus the flag.
|
||
|
||
Sweep note (no edits): the other five smtp-mapped templates (calcom, gitea, mealie, nextcloud,
|
||
rallly) are boot-proven tolerant of defined-empty mail env — all ran healthy as fresh email-off
|
||
deploys during the campaign; gitea's `GITEA__mailer__SMTP_ADDR=${...:-}` pattern likewise.
|
||
Vaultwarden was the only strict image. New REUSE.md trap row: strict images need an enable-flag
|
||
gated `false` in compose + `"true"` in `smtp_mapping.extra`; boot-prove fresh email-off deploys.
|
||
|
||
## 2026-07-03 — sparkyfitness FINALIZED + live-validated (both VERIFY markers resolved); REUSE probe-naming row
|
||
|
||
The first worked example of the new `felhom-app-catalog` skill (felhom.eu). Both
|
||
`VERIFY-BEFORE-FINALIZE` healthcheck guesses resolved by inspecting the real images on the demo box:
|
||
frontend (Alpine/nginx) HAS BusyBox wget → drafted `wget --spider :80/` probe confirmed + kept;
|
||
server HAS node v24.17.0 → node-exec `:3010/api/health` probe confirmed (path proven live:
|
||
`{"status":"UP"}`). Frontend `container_name` renamed → `sparkyfitness` (= the stack name): the
|
||
controller-side probe dials the exact-name container, fallback is the FIRST prefix match (could be
|
||
the DB) — new REUSE.md §2 "Probe-container naming" row records the convention (verified in
|
||
felhom-controller healthprobe.go). Mem-sum comment added (512+1024+256 = 1792M, value unchanged).
|
||
Live-validated on demo via the real dashboard UI (sync + Frissítés): 3/3 containers healthy,
|
||
controller probe `healthy: true` (http :80 → 200), `sparky.demo-felhom.eu` 200 via Traefik;
|
||
data_key secrets untouched (server/db containers not recreated). Kept deployed.
|
||
|
||
## 2026-07-03 — docs: CLAUDE.md light expansion
|
||
|
||
The minimal REUSE-rollout stub expanded to a proper (still ~30-line) CLAUDE.md: what the repo is
|
||
(one dir per app, two template files, Hungarian customer text), the push-to-main = deploy contract
|
||
(controller sync ≤15 min / manual trigger), legacy `templates.json` warning, and pointers
|
||
(REUSE.md, README format spec, the `felhom-build-deploy` skill). No template changes.
|
||
|
||
## 2026-07-03 — docs: REUSE.md introduced
|
||
|
||
Cross-repo reuse-map rollout (docs-only). New `REUSE.md`: catalog conventions verified against all
|
||
53 apps — the canonical example app (paperless-ngx), `.felhom.yml` required fields, healthcheck
|
||
family per image type (BusyBox wget / curl / Node / Python / DB sidecars), memory-limit convention,
|
||
new-app checklist, and traps (gokapi entrypoint hack, legacy templates.json). Known README drift
|
||
recorded in §6 (NOT fixed). Also a minimal `CLAUDE.md` carrying the REUSE.md pointer + maintenance
|
||
rule (full CLAUDE.md is a separate task).
|
||
|
||
## 2026-06-29 — App-email: calcom + nextcloud (tls_mode=plaintext :2526 + nextcloud split-From)
|
||
- **nextcloud** — `smtp_mapping` with `tls_mode: plaintext` (controller injects port 2526, the plaintext-only
|
||
listener) + **split From** (`from_var=MAIL_FROM_ADDRESS` + `from_domain_var=MAIL_DOMAIN` → nextcloud@felhom.eu).
|
||
Compose references the injected `${SMTP_*}`/`${MAIL_*}`. Live-confirmed: real password-reset delivered via
|
||
plaintext :2526 (Symfony Mailer never attempted STARTTLS).
|
||
- **calcom** — `smtp_mapping` with `tls_mode: plaintext` (EMAIL_SERVER_HOST/PORT, EMAIL_FROM=calcom@felhom.eu).
|
||
**Plus three pre-existing template fixes** (calcom never deployed before — the image pin was invalid):
|
||
(1) image `v4.8.7`→`v4.6.9` (the pinned tag has no published image); (2) added required `DATABASE_DIRECT_URL`
|
||
(Prisma `migrate deploy` fails without it → incomplete schema → 500s); (3) healthcheck `/api/health`→
|
||
`/api/auth/providers` (the old path 404s in v4.x → container stayed unhealthy → Traefik wouldn't route).
|
||
- Both apps point at the controller's `:2526` plaintext-only listener because their SMTP clients
|
||
opportunistically STARTTLS-upgrade and can't skip the self-signed cert — the listener simply doesn't offer
|
||
STARTTLS, so they stay plaintext (accepted on the single-tenant app bridge).
|
||
|
||
## 2026-06-29 — App-email rollout: gitea + rallly (calcom/nextcloud/immich = findings)
|
||
- **gitea 1.23.4** — added `smtp_mapping` (STARTTLS via `GITEA__mailer__PROTOCOL=smtp+starttls` +
|
||
`FORCE_TRUST_SERVER_CERT=true` to trust the shim's self-signed cert; single `GITEA__mailer__FROM`). Compose
|
||
references the injected `GITEA__mailer__*` keys; env applied every boot.
|
||
- **rallly** — added `smtp_mapping` (Nodemailer STARTTLS, `SMTP_SECURE=false` + `SMTP_REJECT_UNAUTHORIZED=false`
|
||
to accept the self-signed cert; single `NOREPLY_EMAIL`). **Also fixed three pre-existing template bugs** that
|
||
made rallly undeployable (never caught because the bad pin never ran): (1) image pin `3.12.1` doesn't exist →
|
||
`3.11.2`; (2) healthcheck used `wget`, absent from the rallly image (exit 127) → container unhealthy →
|
||
**Traefik wouldn't route it** → replaced with a Node http check; (3) added required `SUPPORT_EMAIL` + a valid
|
||
`NOREPLY_EMAIL` default (rallly refuses to boot without them).
|
||
- **Both gitea and rallly send-tested live** end-to-end (app → shim → hub → Resend): gitea password-reset
|
||
(From `gitea@felhom.eu`) and rallly registration code (From `rallly@felhom.eu`) both delivered.
|
||
- **NOT wired — reported as findings** (`felhom.eu/documentation/audits/FINDING-app-email-rollout-2026-06-29.md`):
|
||
- **cal.com v4.8.7** — hard-codes TLS `rejectUnauthorized:true` with no override; opportunistic STARTTLS
|
||
against the self-signed shim fails. Needs a non-STARTTLS-advertising plaintext listener (mechanism change).
|
||
- **nextcloud 31** — no cert-skip env (same opportunistic-STARTTLS gap) **and** a split From
|
||
(`MAIL_FROM_ADDRESS`+`MAIL_DOMAIN`) the single-`from_var` mapping can't express.
|
||
- **immich v2.5.5** — no SMTP env vars at all; config is admin-UI/DB or an `IMMICH_CONFIG_FILE` JSON. Does not
|
||
fit env-injection; left for a future config-file-injection mechanism (or manual admin-UI setup).
|
||
|
||
## 2026-06-29 — App-email: smtp_mapping for Vaultwarden + Mealie
|
||
- Added the `smtp_mapping` block to `templates/vaultwarden/.felhom.yml` and `templates/mealie/.felhom.yml`,
|
||
enabling managed outbound email (app → in-controller shim → hub → Resend) for the two spike-proven apps
|
||
(`SPIKE-smtp-app-relay-2026-06-28`). The controller injects `SMTP_*` at deploy/redeploy when app-email is
|
||
on (global + per-app); the From address is `<app>@felhom.eu`. SMTP auth creds are intentionally left unset
|
||
(the shim accepts no-auth on the Docker network).
|
||
- **Vaultwarden:** STARTTLS (`SMTP_SECURITY=starttls`) + `SMTP_ACCEPT_INVALID_CERTS/HOSTNAMES=true` to
|
||
accept the shim's self-signed cert.
|
||
- **Mealie:** plaintext (`SMTP_AUTH_STRATEGY=NONE`) on :2525 — Mealie has no accept-invalid-cert option, so
|
||
STARTTLS to a self-signed shim would fail; plaintext to the Docker-network-only shim is the spike-validated
|
||
mode.
|
||
- Both `docker-compose.yml` files now reference the injected `${SMTP_*}` keys (with harmless defaults) so the
|
||
values reach the container; empty `SMTP_HOST` keeps mail disabled when the toggle is off.
|
||
- Documented the `smtp_mapping` pattern in `README.md` so further apps are easy adds.
|
||
|
||
## 2026-06-28 — Add SparkyFitness (v0.17.2) — nutrition/workout tracker
|
||
- New app `templates/sparkyfitness/{docker-compose.yml,.felhom.yml}`: a self-hosted nutrition/calorie +
|
||
workout/weight tracker (alternative to wger). Three containers — nginx **frontend** (SPA :80, the sole
|
||
Traefik ingress, proxies `/api`+`/uploads` internally) + Node **server** (:3010) + dedicated
|
||
**postgres:15-alpine**. Server + DB stay on the internal network with no Traefik labels.
|
||
- **Native email/password auth** (no OIDC/Authentik — that's DooPlex-specific); subdomain `sparky`
|
||
(deliberately ≠ wger's `fitness` to avoid a Host() collision). `pi_compatible: false`, `needs_hdd: false`.
|
||
- **Two DB roles**: `sparky` (POSTGRES superuser, runs init/migrations) + `sparkyapp` (limited app role the
|
||
server auto-creates on first boot) — separate `DB_PASSWORD`/`APP_DB_PASSWORD`. `PGDATA` in a `pgdata`
|
||
subdir of the named volume. Four auto-generated, `locked_after_deploy` secrets; `API_ENCRYPTION_KEY` +
|
||
`BETTER_AUTH_SECRET` carry `data_key: true` (restore recovers, never regenerates — both are 64-char hex).
|
||
- Transcribed from the validated k3s manifest `homelab-manifests/workout-system/sparkyfitness.yaml`
|
||
(pinned image tags, two-DB-role model, never-change crypto keys, `/api/health`, pg15 + PGDATA subdir).
|
||
- **Image-probe findings (build server, v0.17.2):** server keeps the `node -e` `/api/health` probe (node
|
||
present); frontend keeps the `wget --spider` probe (both `wget` and `curl` present). No probe changes needed.
|
||
- **Live-validated on guest 9201 (controller v0.87.0):** synced via "Sablonok frissítése"; deployed through
|
||
the real dashboard flow (Domain auto, Subdomain `sparky`, 4 secrets auto-gen). All 3 containers healthy;
|
||
server log shows clean migrations + `sparkyapp` role created + RLS applied, no crash loop, no uploads
|
||
EACCES; `GET /api/health` through the public edge returns `{"status":"UP"}`; login/register page serves
|
||
over a valid TLS cert at `https://sparky.demo-felhom.eu`.
|
||
|
||
## 2026-06-26 — crafty-controller: image bump 4.4.8→4.10.7 + publish Java port range + connection guidance
|
||
- **Image bump** `crafty-4:4.4.8` → `4.10.7` (latest stable; 4.10.8/4.11.0 don't exist in the registry).
|
||
6 minor versions of fixes incl. security CVEs. **Java 25 verified present** in 4.10.7
|
||
(`/usr/lib/jvm/java-25-openjdk-amd64`, default `java -version` = openjdk 25.0.3; 8/11/17/21 also
|
||
available) — so the latest-Minecraft (`26.x`, needs Java 25) blocker is resolved. Healthcheck + Traefik
|
||
https-backend labels unchanged (Crafty still serves HTTPS on 8443).
|
||
- **Published the Java game-port range** `25565-25575:25565-25575` (TCP, 11 ports = up to 11 Java
|
||
servers; first server 25565, rest 25566–25575). No `network_mode: host` (would break Traefik routing).
|
||
Bedrock UDP 19132 intentionally out of scope.
|
||
- **App-page guidance** (`.felhom.yml` first_steps + prerequisites): how to set the server port within
|
||
25565–25575, how to connect on the LAN (manual IP:port — "scan for LAN" won't auto-list), and that
|
||
internet access needs operator port-forwarding. (Static text — can't show the live LAN IP.)
|
||
- **Live-verified on guest 9201:** 4.10.7 healthy; public URL 302; the guest's bridged LAN IP
|
||
`192.168.0.121` reaches the real Crafty "test" server on `25565` (TCP OPEN + Minecraft SLP handshake
|
||
returns JSON status); `:25575` reachable, `:25600` closed (negative control). In-place upgrade preserved
|
||
the admin, the operator's configured MFA, and the test server.
|
||
- **Correction (earlier draft was wrong):** an earlier note here claimed the upgrade "locked out the
|
||
admin (TOTP)." That was a misdiagnosis — the `totp_data` row + recovery codes were **operator-configured
|
||
MFA**, so the 401 on a password-only login was correct behaviour, NOT an upgrade bug. There is **no
|
||
upgrade regression**; the bump preserves data and MFA correctly.
|
||
|
||
## 2026-06-26 — crafty-controller: seed a felhom-generated admin password (replaces Crafty's ugly random one)
|
||
- **crafty-controller**: instead of reading Crafty's auto-generated (long, symbol-laden) random admin
|
||
password, we now **seed** a clean felhom-generated one — same pattern as gokapi, so initial passwords are
|
||
consistent across the catalog.
|
||
- Crafty's image ships `app/config_original/default.json = {"username":"admin","password":"crafty"}`;
|
||
"crafty" is 6 chars < Crafty's 8-char minimum, so Crafty rejected it and generated a random password.
|
||
- New `CRAFTY_PASSWORD` deploy field (`type: password`, `generate: password:24`, locked after deploy —
|
||
mirrors gokapi's `GOKAPI_PASSWORD`). The compose **entrypoint** overwrites the `default.json` template
|
||
with this password before the launcher runs; on fresh install Crafty creates the `admin` user with it.
|
||
- `initial_credentials.file` repointed `default-creds.txt` → `default.json` (same json/username/password
|
||
keys), so the controller's app-page "Kezdeti belépési adatok" card shows the **seeded** password — the
|
||
customer sees the same value at deploy time and on the app page.
|
||
- Catalog-only change (reuses felhom-controller v0.84.0's initial_credentials reader + the gokapi-style
|
||
seed). Requires a fresh install to take effect (the seed is only read on first run).
|
||
|
||
## 2026-06-26 — crafty-controller: surface the auto-generated initial admin password on the app page
|
||
- **crafty-controller**: Crafty writes a random admin password to `/crafty/app/config/default-creds.txt`
|
||
at first boot (its built-in default is rejected as "too short"). Customers had to read the container
|
||
logs to find it. Added an `initial_credentials` block (new general felhom-controller v0.84.0 mechanism):
|
||
`file` + `format: json` + `username_key`/`password_key` + a `note`. The controller reads the file live
|
||
from the container and shows username + password (masked, reveal/copy) on the app's page under "Kezdeti
|
||
belépési adatok". Requires felhom-controller ≥ v0.84.0.
|
||
- Updated `first_steps` to point at the app page for the initial login instead of "find it in the logs".
|
||
|
||
## 2026-06-26 — crafty-controller: Traefik https backend + scoped skip-verify (fixes 502)
|
||
- **crafty-controller**: the healthcheck fix un-withheld the Traefik route, exposing a pre-existing
|
||
**502** — Traefik proxied `http://…:8443` to Crafty's **HTTPS-only** self-signed backend (Crafty serves
|
||
no plain-HTTP panel; `:8000` only redirects). Added two service labels:
|
||
- `loadbalancer.server.scheme=https` — Traefik now speaks HTTPS to the backend.
|
||
- `loadbalancer.serverstransport=insecure-skip-verify@file` — references the **named**
|
||
serversTransport defined in the controller-managed Traefik dynamic config (felhom-controller v0.83.0),
|
||
which skips verifying Crafty's per-container self-signed cert. Verification stays ON for every other
|
||
backend (scoped Option B; no global `insecureSkipVerify`). The `@file` suffix is the cross-provider
|
||
reference from the docker provider to the file-provider transport.
|
||
- Requires felhom-controller ≥ v0.83.0 (which renders the `insecure-skip-verify` transport). `port=8443`
|
||
and the router/tls labels are unchanged.
|
||
|
||
## 2026-06-26 — crafty-controller healthcheck fix (curl-absent + http-vs-TLS probe)
|
||
- **crafty-controller**: container was permanently `unhealthy` → route withheld (`routeUnpublished`).
|
||
Two independent healthcheck root causes, both fixed in one change:
|
||
- **Docker healthcheck** ran `curl -fk https://localhost:8443`, but the `crafty-4:4.4.8` image has
|
||
**no `curl` and no `wget`** (`exec: "curl": not found`, FailingStreak 150). Replaced with a
|
||
dependency-free **python3 TLS-socket** liveness probe (`/usr/bin/python3` is present): completes a
|
||
TLS handshake to `127.0.0.1:8443` (unverified context mirrors the old `-k`; Crafty's cert is
|
||
self-signed). `start_period` 30s → 60s for cold-boot headroom (cert gen + migrations).
|
||
- **Controller-side probe** (`.felhom.yml healthcheck.checks`) was `type: http` against Crafty's
|
||
**TLS-only** 8443 → `probeHTTP` sent plaintext HTTP, got a TLS record → `HealthProbe.Healthy=false`,
|
||
which `manager.go` re-applies to override Docker's verdict back to `unhealthy`. Changed `http` →
|
||
`tcp` (`probeTCP` dial succeeds against a TLS listener). Both layers had to change together.
|
||
- Live-validated on guest 9201 (`demo-felhom`): synced → recreated via the update path → Docker
|
||
`State.Health: healthy` (ExitCode 0), `health_probe.healthy: true` (tcp :8443, 5ms), http-vs-TLS
|
||
WARNs stopped, stable green 3+ min, Traefik now **publishes** the route (`crafty-controller@docker`).
|
||
- **Known follow-up (separate, out of this fix's scope):** the public URL still returns **502** — a
|
||
distinct pre-existing bug the un-withheld route exposed: Traefik proxies `http://…:8443` to Crafty's
|
||
HTTPS-only backend. Needs a Traefik HTTPS-backend + self-signed `serversTransport`
|
||
(`insecureSkipVerify`) in the controller-generated Traefik config — tracked separately.
|
||
|
||
## 2026-06-23 — gokapi: index redirect + admin username display
|
||
- **gokapi**: seed `RedirectUrl` repointed from Gokapi's GitHub default → `https://${SUBDOMAIN}.${DOMAIN}/admin`.
|
||
Gokapi's bare root `/` redirects to `RedirectUrl`; the controller's "Megnyitás" link is always the bare
|
||
subdomain root, so it was landing on Gokapi's GitHub instead of the app. Now `/` → `/admin` → login.
|
||
Applied to the live demo (config.json edit + restart) and the seed (future deploys).
|
||
- **gokapi**: added `app_info.default_creds` ("Felhasználó: admin · jelszó a Beállítások oldalon") so the
|
||
app-info page shows the initial admin user like other apps; fixed `first_steps` (no more setup wizard).
|
||
|
||
## 2026-06-23 — gokapi reproducible headless setup (fixes public "maintenance mode")
|
||
- **gokapi**: was stuck in "maintenance mode" on the public URL since first deploy — Gokapi's one-time
|
||
`/setup` wizard was never completed, and (verified against the docs + the v1.9.6 binary) **no Gokapi
|
||
version supports env-var headless setup** for admin credentials. Worse, the unconfigured `/setup` was
|
||
publicly reachable = an unauthenticated admin-takeover window.
|
||
- Fix: the compose `entrypoint` now seeds a `config.json` on first boot (admin user, this app's public
|
||
URL `https://${SUBDOMAIN}.${DOMAIN}/`, local storage, Encryption Level 0 so it restarts without a
|
||
prompt) with `Password`/`SaltAdmin`/`SaltFiles` cleared, then runs Gokapi's documented
|
||
`--deployment-password` one-shot to set the **felhom-generated** admin password **before** the
|
||
server starts serving. The admin account is claimed at first boot → `/setup` is never exposed.
|
||
- `.felhom.yml`: new `GOKAPI_PASSWORD` deploy field (`type: password`, `generate: password:24`,
|
||
shown to the customer, locked after deploy). Admin username is `admin`.
|
||
- Seed is pinned to Gokapi **v1.9.6** (`ConfigVersion 21`) — re-capture the seed if the image is bumped.
|
||
- Live-validated on guest 9201: fresh remove+redeploy → headless auto-config, public login works, no
|
||
maintenance page, admin claimed at first boot (browser-verified login).
|
||
|
||
## 2026-06-22 — gitea healthcheck fix (unattended test campaign)
|
||
- **gitea**: healthcheck probe repointed `/api/v1/version` → `/api/healthz` (docker HC + controller
|
||
`.felhom.yml` probe), `start_period` 30s → 90s.
|
||
- Surfaced during the Phase-2 deploy sweep: a fresh gitea reported `unhealthy` because
|
||
`/api/v1/version` returns 404 until the install wizard / INSTALL_LOCK completes, while the
|
||
container was serving fine on :3000 (`/api/healthz` → 200). Same class as the komga fix.
|
||
|
||
## 2026-06-22 — komga healthcheck fix (unattended test campaign)
|
||
- **komga**: healthcheck probe repointed `/api/v1/actuator/health` → `/actuator/health`.
|
||
- Root cause: komga's Spring Boot actuator endpoint is served unauthenticated at `/actuator/health`
|
||
(HTTP 200), while everything under the `/api/v1` prefix is auth-gated — so the old probe got
|
||
HTTP 401, `curl -f` exited 22, and the container reported `unhealthy` despite serving normally
|
||
on :25600. Diagnosed live on guest 9201 (probe matrix: `/`, `/actuator/health`,
|
||
`/api/v1/oauth2/providers`, `/login` all 200; `/api/v1/actuator/health` → 401).
|
||
- The `gotson/komga:1.20.0` image ships `curl` (verified), so the probe tool is unchanged.
|