R-449. Until today one upgrade out of 53 had ever been measured - Nextcloud, by hand, in a spike - and the whole update arc was designed against that single data point. Per edge: deploy at FROM, seed through the app's OWN interface, prove the seed reads back, swap to TO, ask the app for the data again, then put the FROM images back and record what happens - verbatim, and never called a rollback. Success is an application-level readback, not file identity: survive2.py's sha256+inode rule is right for a redeploy and wrong for an upgrade, because a migration is supposed to rewrite files. And nothing is ever seeded by hand (R-156) - an app with no non-browser route is recorded inconclusive, never faked. C3 is a negative control whose TO image exits immediately, and it must be run first: it came back failed, which is what makes the greens mean anything. The bookstack fixture uses artisan for both halves and carries its own negative control on every call, because the obvious HTTP-login readback cannot work: the template's https APP_URL makes the session cookies secure, so curl over http gets 419 on every login and it looks exactly like a wrong password.
8.9 KiB
CLAUDE.md — app-catalog-felhom.eu
Loads when Claude Code touches this repo. Current state:
CONTEXT.md+CHANGELOG.mdtop. Cross-repo orientation: workspace-root/mnt/5_hdd/felhom.eu/git/CLAUDE.md.
What this repo is
The Felhom app catalog: one directory per app under templates/<app>/, each holding exactly
docker-compose.yml + .felhom.yml (deploy fields, resources, healthcheck probe, app_info — all
customer-facing text in Hungarian). The felhom-controller git-syncs these to every customer box;
.felhom.yml drives the deploy wizard. templates.json + scripts/generate-customer.sh are LEGACY
(Portainer-era) — new apps don't touch them.
Deploy contract
Push to main = deploy. The controller's sync picks changes up within 15 minutes (or trigger via
the dashboard "Sablonok frissítése" button / POST /api/sync). Only the two template files sync;
deployed app.yaml (customer secrets) is never overwritten. Full deploy details: the
felhom-build-deploy skill.
Conventions
- See
REUSE.mdbefore adding or editing an app — canonical example app (paperless-ngx), required.felhom.ymlfields, healthcheck family per image type, memory-limit rules, traps. - Update
REUSE.mdin the same commit that changes a catalog-wide convention. README.mdis the format spec — update its app tables when adding an app.- Update
CHANGELOG.md(newest on top) and overwriteREPORT.mdwith every pushed change. - No secrets in any committed file; secrets are generated at deploy time via
deploy_fieldsgenerate:specs. - Run
python3 scripts/catalog_gates.py <app>after ANY template change — it is the ONE entry point and runs all three gates below, exiting non-zero if any fails. Name the app(s) you touched and it scopes the two gates that accept scoping, which is fast; with no names the runtime gate deploys every template, so that form belongs on a scratch host, never a customer box. Exit: 0 all clean · 1 convicted · 2 UNDETERMINED, which is never a pass. Why a runner and not four separate invocations (operator ruling 2026-08-02, R-161): of this project's gates, the only ones that ever get run are the ones with a single entry point named in a CLAUDE.md —felhom.eu/scripts/site_gates.pyis run, and R-29's three orphans are named nowhere and have stopped nothing. Controller-side enforcement was rejected because a check at template load can only read the file, and a static audit of all 53 templates reports the catalog clean including papra — it would pass on the exact defect it exists to catch. CI was rejected for now: neither repo has any, and there are no users yet. R-161 stays open at reduced scope — this is convention, run by a person; real automatic enforcement is owed when a second person touches templates. Update 2026-08-02:.githooks/pre-pushnow runscatalog_gates.py --faston every push, which is gate 1 (check-image-pins.py) only — the other two need network and a container runtime and take minutes per app, and a push that pulls images and starts containers gets bypassed within a week, after which the bypass is the habit. They stay deliberate periodic runs. The hook is per-clone (git config core.hooksPath .githooks) andgit push --no-verifybypasses it, which is why gate 1 alone cannot be the whole story — CI re-runs the entry point on every push and emails on failure (felhom.euOPEN-ITEMS.mdR-168, CLOSED 2026-08-02), which is what notices a bypass. - Never
:latestor untagged images in templates — pin a concrete version tag; an app deployed anywhere in the fleet is pinned to the digest it is currently running (a pin must never cause a version jump). Digest pins (@sha256:) also count. Gate:python scripts/check-image-pins.py(run after any compose change; exit 1 on any floating/missing tag). - Any commit that changes an
image:line MUST set that app'scatalog_sinceto the same day..felhom.ymlcarriescatalog_since: "YYYY-MM-DD"— the date THIS repo last moved that app's pinned images. It is not a version and not an upstream release date; the controller uses it, and only it, to tell a customer "Frissítés elérhető — 45 napja". No version number is ever shown to the customer, so there is deliberately noversion:key to keep in sync alongside it. A stalecatalog_sinceunder-reports how long a box has been behind, which is the one number the badge exists to give. Absent, empty, malformed or FUTURE-dated all degrade to a badge with no age and one WARN in the controller log — never a broken template. There is no gate for this yet and that is a known gap, filed as a register row: the gates runner fetches at--depth 1and has no parent commit to diff animage:line against, so a drift gate needs a deeper fetch. Backfilled for all 53 apps from git history on 2026-09-02. - A pinned tag can still rot away upstream — the pin gate is syntactic and cannot see that.
Second gate:
python3 scripts/check-image-resolvable.py(exit 0 resolve / 1 GONE / 2 inconclusive), run at the start of every catalog campaign and before any publish train that vouches the catalog. Needs network +docker; unauthenticated Docker Hub throttles a full sweep, sodocker loginfirst or expect exit 2. It reports a throttle as INCONCLUSIVE, never as a dead image. - A well-formed template can still be un-upgradable, and no gate can see that either. Fourth
instrument, and the only one that needs REAL DATA:
python3 scripts/upgrade-test.py <edge>deploys an app at a FROM image set, seeds through the app's own interface, swaps to a TO set, and asks the app for the data back. Success is an application-level readback, never file identity — a migration is supposed to rewrite files, so the persistence sweep's sha256+inode rule would fail every correct upgrade. It also records what the ABORT does (putting the old images back), verbatim. RunC3first, every time: it is a negative control whose TO image exits immediately, and if it does not come backfailedthe harness is not measuring anything. Fixtures live inscripts/upgrade_fixtures.py; an app with no non-browser seed route is recordedinconclusive, never faked. Needs Docker + real images + minutes per edge, on a scratch host, never a customer box. First run:felhom.eu/documentation/audits/SPIKE-upgrade-test-2026-09-06.md. - A well-formed template can still preserve the wrong folder — and no static check can see it.
Third gate, the only RUNTIME one:
python3 scripts/check-volume-persistence.py(0 all clean / 1 REFUSED / 2 undecided). It deploys each template, exercises it into writing data, and compares where the data landed against what the compose mounts. Needs Docker + network and minutes per app, so it is periodic like the resolvability gate — run it whenever a template'svolumes:block or image tag changes, and at the start of every catalog campaign, on a scratch host, never a customer box. It refuses to report at all unless it has just re-proven itself in both directions against two canary templates. Fixture tests (no Docker):python3 scripts/test_check_volume_persistence.py. UNDETERMINED is exit 2 and is never a pass — an app that wrote nothing has not been shown to be correct. Why it exists: papra mountedpapra_data:/app/datawhile the app wrote its database to/app/app-data/db/, so its backup completed, verified, and contained an empty directory (R-156, Campaign 10). - Taking an app out of circulation — use
lifecycle:, never a directory move..felhom.ymlgains an optionallifecycle:field:available(default; absent/empty means this),hidden(not offered for new installs, no explanation owed),abandoned(upstream stopped developing it — not offered for new installs, and every box already running it shows a permanent "Nem karbantartott" notice). Deployed instances keep working in full either way — the state affects what is OFFERED, never what already runs, and the controller REFUSES a deploy of a non-available template server-side. An unknown value degrades toavailablewith one WARN, so a typo can never brick a template. This supersedes the short-livedretired/directory move, which was wrong: removing a template orphans every customer already running it.
A gate ships with a decoy test that has been seen to fail (R-421). A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. scripts/decoy_coverage_gate.py refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: documentation/audits/AUDIT-gate-decoys-2026-09-01.md and
felhom-controller/.claude/rules/gates.md. Scope is a fact too — prefer os.walk over
os.listdir, and a glob over a hand-maintained list.