The rule (CLAUDE.md, operator ruling 2026-09-13): until the Update button takes a verified backup
as its precondition, no template may move a mariadb:/postgres: image across a major version. Four
MariaDB and eleven PostgreSQL services; the gate finds them by image name, not by a list.
scripts/check-engine-major.py — fast (git reads only), diffs each changed template's per-service
image: line between the two ends of the push range, refuses a major move naming the rule and its
expiry (R-448). Fourth row of catalog_gates.py; .githooks/pre-push now hands the push range
through as --range=<remote sha>..<local sha>.
HONEST LIMIT: it needs a parent commit and CI fetches at --depth 1 (the R-452 gap, not re-filed),
so on a shallow clone the runner SKIPS it out loud instead of reddening every CI push. The hook,
which has the full clone, is where it bites.
Red-proof (scripts/test_gate_decoys.py, 7 cases, all seen to judge correctly): mariadb 11.6->12.3
REFUSED, postgres 16->17 REFUSED, mariadb:lts INCONCLUSIVE; 11.6->11.8 PASSES; the major moving
only in a comment / kimai's serverVersion env / README / the app's own image PASSES. COVERS literal
registered for felhom.eu's decoy_coverage_gate (which now reads 1 covered, 3 exempt, 0 unaccounted).
test_catalog_gates.py pins the four-gate table and the announced shallow-clone skip.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
bookstack-db, kimai-db, nextcloud-db, romm-db each gain `MARIADB_AUTO_UPGRADE=1` in the db
service's environment list. Operator ruling 2026-09-13 on the measurement in
felhom.eu/documentation/audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md: an unconverted datadir
is stable but never heals; the conversion costs ~7 s and the engine backs its system tables up
first. MARIADB_DISABLE_UPGRADE_BACKUP is deliberately left UNSET — that backup is the precaution.
NO `image:` line changed, so `catalog_since` does NOT move — the CLAUDE.md rule ties it to an
image change and this is not one. Do not "fix" that.
The setting is inert until an engine major actually moves, and none may until Slice 4 (R-448)
ships — see the engine-major rule in CLAUDE.md and scripts/check-engine-major.py (next commit).
The eleven PostgreSQL templates are untouched: R-463 is a different engine and a different
measurement.
REUSE.md: one convention row for the MariaDB sidecar env, same commit.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-459. The harness returned proven for E3b while MariaDB was logging that the
conversion it requires had been skipped. The verdict was right - the app's data
survived, which is what it asked - but the harness watched the app and the
migration log, and neither looks at engine state.
engine_state_after now carries each database service's own answer: MariaDB's
datadir version plus 'mariadb-upgrade --check-if-upgrade-is-needed', and
PostgreSQL's PG_VERSION.
It sits BESIDE the verdict and is never folded into it. An unconverted datadir is
not known to be a failure - 5 of 5 restarts showed no degradation - so a verdict
that called it failed would encode an unproven judgement, which is worse than
reporting a fact and letting a person read both.
No template changed. Nothing with MARIADB_ in it is committed by this work: that
is a fleet-wide decision the operator owns, and it affects four apps.
R-449. Until today one upgrade out of 53 had ever been measured - Nextcloud, by
hand, in a spike - and the whole update arc was designed against that single data
point.
Per edge: deploy at FROM, seed through the app's OWN interface, prove the seed
reads back, swap to TO, ask the app for the data again, then put the FROM images
back and record what happens - verbatim, and never called a rollback.
Success is an application-level readback, not file identity: survive2.py's
sha256+inode rule is right for a redeploy and wrong for an upgrade, because a
migration is supposed to rewrite files. And nothing is ever seeded by hand (R-156)
- an app with no non-browser route is recorded inconclusive, never faked.
C3 is a negative control whose TO image exits immediately, and it must be run
first: it came back failed, which is what makes the greens mean anything.
The bookstack fixture uses artisan for both halves and carries its own negative
control on every call, because the obvious HTTP-login readback cannot work: the
template's https APP_URL makes the session cookies secure, so curl over http gets
419 on every login and it looks exactly like a wrong password.
An IMAGE change, pushed to prove that a pinned app does NOT receive it — the whole
point of slice 3. bentopdf is deployed on demo-hp only and is file-based with no
database and no volume, so no data anywhere can be touched. The revert follows in
the same session.
A NON-image template change, pushed to prove that a pinned app still receives
template fixes on the normal 15-minute cycle. No image pin is touched. The revert
commit follows in the same session.
Backfilled from this repo's own git history (newest commit whose image: set differs
from its parent's), excluding the two bentopdf SPIKE commits by hash — they moved a
pin and reverted it in the same hour and are a measurement, not a release.
Four anchors confirmed (nextcloud 5e2c1ae, grafana b789acc, calcom 147cee7,
vikunja 3fa63cd, all 2026-07-18); six apps re-checked against git log by hand.
Consumed by felhom-controller v0.233.0 to render 'Frissitesi elerheto - N napja'.
No version number is ever shown to the customer, so there is no version: key.
Rule added to CLAUDE.md; the missing drift gate is filed as a register row (the
gates runner fetches at --depth 1 and has no parent to diff an image: line against).
Phase 2 of SPIKE-app-update-2026-09-01: measure live whether the 15-minute
catalog sync rewrites a DEPLOYED app's docker-compose.yml on demo-hp while its
running container keeps the old image (R-438).
bentopdf is deployed on demo-hp ONLY (demo-felhom runs opengist alone; Peti's
box is down with no access route), it is file-based with no database and no
volume, so the change cannot touch customer data anywhere.
This commit is reverted as soon as the measurement is taken.
All 29 gate scripts across the four repos were read and DECOYED - the label constructed without the
fact, the gate run, the verdict recorded. 16 were fooled. None of them were in this repo.
A decoy that nobody would write proves nothing, so the attempts that turned out illegitimate were
WITHDRAWN rather than counted. Both of this repo were withdrawn, and both are named in the audit.
The gates here that could not be given a plausible decoy are listed BY NAME in
felhom.eu/scripts/decoy_coverage_gate.py EXEMPT (R-426) as UNTESTED - not as sound. A gate nobody
tried to fool is UNKNOWN, and calling it sound would be the same confident guess this sweep exists
to find.
Survey table: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
Corrected in all four instruction files across all four repos. Found while confirming this
session own push by run ID, which is precisely the check that catches it.
In felhom-agent/CLAUDE.md the sentence contradicted the same file release section, which
already said R-168 mails the failure -- a contradiction inside one instruction file, the exact
class the R-229 work exists to find.
REPORT.md deliberately NOT overwritten in the sibling repos: a one-line docs correction must not
destroy the record of their last real implementation.
The workspace root is already documented (workspace-CLAUDE.md, the workspace-root
CLAUDE.md 'stay inside it') and work drifted into a home directory anyway. A rule
that has failed once as a reminder is not fixed by writing it down again, so it is
now asserted where it can bite.
A push is the right trigger: throwaway clones under /tmp for probes and red-proofs
never push, so nothing legitimate breaks. Symlinks are resolved on both sides; an
absent workspace root SKIPS the check rather than failing it, so this cannot brick
a legitimate clone on another machine. The only bypass is the documented
--no-verify, whose use is already reportable.
Identical in all four repos.
papra mounted papra_data:/app/data while the application writes to
/app/app-data, so its database sat in the container's writable layer: lost on
redeploy, and tarred nightly as an empty directory while the healthcheck stayed
green. Last of the three apps R-156 convicted.
Decided from the IMAGE, not the README. docker inspect of
ghcr.io/papra-hq/papra:26.6.1-rootless gives WORKDIR=/app and all three data
paths under ./app-data (DATABASE_URL, DOCUMENT_STORAGE_FILESYSTEM_ROOT,
PAPRA_CONFIG_DIR) — and /app/data does not exist in the image at all.
Reconfiguring the app to write to /app/data was available and deliberately not
taken: it enumerates data paths, so a fourth added upstream would escape to the
writable layer again, silently — this defect re-armed. Mounting the app's own
data ROOT captures every current and future path by construction.
Precondition checked rather than inherited: docker ps -a (including stopped) on
BOTH demo guests, plus the hub fleet view (two enrolled hosts, zero papra) —
both boxes were wiped and rebuilt today, so the 2 August evidence was re-measured.
Proven by the runtime gate in both directions: CLEAN with the self-test passing,
and BROKEN when the mount is reverted, with the exact R-156 evidence. Full
catalog_gates.py papra: all three gates OK.
--fast only: check-image-pins runs; the network and container-runtime gates do NOT. CI that
pulls 53 images on every push gets disabled, and they remain deliberate periodic runs.
No sibling clone needed here — unlike the controller and the agent, catalog_gates --fast
does not invoke the shared reuse checker. No uses: step, no version bump.
--fast selects only gates that touch no network and no container runtime: gate 1
(check-image-pins) runs, image-resolvable and volume-persistence do NOT. Default behaviour with
no flag is unchanged. The skip is ANNOUNCED with the reason and with what still owes a periodic
run — a silently narrowed run reads as 'covered everything' when it did not.
Why the runtime gates are never in a hook: a push that pulls images and starts containers gets
bypassed within a week, and the bypass becomes the habit. They stay deliberate periodic runs at
the start of a catalog campaign, before a publish train, and when a template's volumes: block or
image tag changes — on a scratch host, never a customer box.
.githooks/pre-push runs catalog_gates.py --fast and refuses the push. Per-clone and
--no-verify-able, both stated in the hook itself.
test_catalog_gates.py pins --fast's CONTENT, not just its exit code: the runtime gates must not
run, the skip must be announced, and the no-flag path must still select all three. Red-proofed:
an inert run_gate turns it red.
scripts/catalog_gates.py runs all three gates - image-pins, image-resolvable,
volume-persistence - and exits non-zero if any fails. Mandated in CLAUDE.md the way
felhom.eu/scripts/site_gates.py is: run it after any template change, naming the
app(s) you touched.
Operator ruling, recorded because both alternatives were rejected for measured
reasons. Controller-side enforcement at template load was rejected because such a
check can only read the file, and a static audit of all 53 templates reports the
catalog clean INCLUDING papra - it would pass on the exact defect it exists to
catch; the property is decidable only at runtime. CI was rejected for now: neither
repo has any, and there are no users yet. What was chosen copies the shape that
demonstrably works here - of this project's gates, the only ones that ever get run
are the ones with a single entry point named in a CLAUDE.md; site_gates.py is run,
and R-29's three orphans are named nowhere and have stopped nothing.
Behaviour: 0 all clean / 1 convicted / 2 UNDETERMINED, never a pass; a conviction
outranks an undetermined result so the reader knows which they have. Gate output is
streamed, not captured. App names scope the two gates that accept scoping; with no
names the runtime gate deploys every template and belongs on a scratch host.
Adding a fourth gate means one line in GATES.
R-161 stays OPEN at reduced scope: this is convention, run by a person. Real
automatic enforcement is owed when a second person touches templates.
Verified: image-pins passes standalone (53 templates, 0 unpinned), the
unknown-option path exits 2, and the aggregation was unit-checked over five
gate-code combinations. The runtime leg was deliberately NOT executed - it deploys
templates via docker compose and DooPlex is the recovery chain - so the runner's
end-to-end invocation of that third gate is inferred, not measured, and is flagged
in REPORT.md to be closed on a scratch host at the next campaign.
REPORT.md overwritten per convention; the persistence sweep's report is preserved
at audits/persistence-sweep-2026-08-02/ and pointed to from the new one.
The register grep that put these at R-158..R-161 was true when run and stale within hours: the
parallel session pushed SPIKE-recovery-unit-space-2026-08-02.md and CAMPAIGN-10-closeout-2026-08-02.md
mid-run, both using R-158 for an unrelated finding, and neither files it — ROADMAP.md and
OPEN-ITEMS.md still stop at R-155.
So two sessions minted the same number for different findings on the same day, which is the exact
failure §8.0 was already documenting about R-154/R-155 and R-156/R-157. Now recorded with itself as
the third instance. R-156, R-157 and R-158 are all live in audit documents and none is filed.
Layer 1: guest 9301 destroyed, vm-9301-disk-0 removed, pct list clean.
Layer 2: local-lvm available returned to 258 702 410 KiB — EXACTLY the pre-run baseline (29.27%).
Layer 3: no hub record was ever created (the scratch guest ran no controller and was never
enrolled, by design); confirmed absent from /configs and /hosts after teardown.
papra referral updated: c10-soak was torn down by the other session mid-run and papra disappeared
from hub telemetry with grafana, homebox and rallly — Campaign 10's four discriminator apps (§A4).
So papra's one deployment was on c10-soak. The fix is STILL not applied: that is absence evidence,
demo-hp is fenced, and applying it wrongly is irreversible while leaving it is a one-line push.
The single command that settles it is recorded.
Campaign 10's R-156 found papra writing its database into the container's writable layer while the
volume the template preserves stayed empty — a backup that completes, verifies, and contains
nothing. papra was never the point: nothing anywhere checked that the folder a template preserves
is the folder the app writes to. All 53 templates have now been measured live.
43 CLEAN / 3 BROKEN / 7 UNDETERMINED. UNDETERMINED is counted separately, each with its reason,
and never folded into CLEAN.
FIXED (neither app is deployed anywhere, so nothing was stranded):
- gramps-web mounted /app/data, /app/media, /tmp — and /app/data is a path the application never
writes. Its accounts database and ITS FAMILY TREE both landed in the writable layer while
gramps_data was tarred nightly as an empty directory. Now persists the eight paths the image's
own environment names, matching upstream's reference compose. Proven: users.sqlite and the
family-tree files survive a redeploy byte-identical, same inode.
- wishlist mounted wishlist_data:/data, another path the app never writes; prod.db landed in the
ANONYMOUS volume from the image's VOLUME directive — absent from ResolveDockerVolumeNames, so
never backed up, and orphaned by a redeploy. Now mounts /usr/src/app/data + /usr/src/app/uploads.
Proven: prod.db byte-identical, same inode, across a redeploy.
Every corrected path confirmed by two independent sources — the shipped image's own
environment/Config.Volumes and upstream's reference compose — never inferred from a directory name.
papra is NOT fixed. It is live on one box, and changing the mount target makes the next compose up
recreate the container and destroy the writable layer its documents live in. The fix is prepared
and proven in the scratch guest (current: db.sqlite differs after a redeploy, so a real account
created via the API is lost; fixed: byte-identical, it survives). Referred to the operator with the
two options; no migration written.
NEW GATE scripts/check-volume-persistence.py — the third catalog gate and the only RUNTIME one.
This class is invisible to static analysis, measured not assumed: a static audit of all 53 composes
reports the catalog clean AND reports papra clean. Exit 0 clean / 1 REFUSED / 2 undecided. It
refuses to report at all unless it has just re-proven itself in both directions against two canary
templates that differ only in which path the volume mounts at, so every run carries a live
demonstration of R-156 and of its fix. No docker exec anywhere (Campaign 7 §1.1). 44 fixture tests
driving check(), the function __main__ calls; every rule red-proofed.
Enforcement is convention, not CI — this repo has no CI. Stated plainly in the report; raising it
is proposed as R-160.
Report, per-app evidence, proofs and proposed register entries (R-158..R-161, NOT filed — felhom.eu
is fenced this session): audits/persistence-sweep-2026-08-02/
paperless-ngx and calibre-web are the only two catalog apps with a drop-zone,
and each has exactly ONE ingest bind. Both move from ${USERDATA_PATH}/import/<app>
to ${IMPORT_PATH}/<app> — the canonical root on the system drive — so a
multi-drive box has one drop-zone instead of one real folder plus a dead
lookalike on every other drive (and import/* is class: excluded, so files
stranded in a dead one would never be backed up either).
The matching backup: entries move to the new `import:` list IN THE SAME COMMIT.
This is not cosmetic: ValidateBackupSpec rejects an entry matching no compose
bind, and the rejection is WHOLE-BLOCK, so a stale `userdata: import/paperless`
would have discarded paperless's `hdd: appdata/paperless/media class: mandatory`
too and silently degraded the customer's document originals to legacy handling.
Both classes stay `excluded` — the move must not change data handling.
New data_paths: blocks on paperless-ngx, calibre-web and romm — role + Hungarian
label over paths that already exist as compose binds. Covers all three roles and
the multi-entry case. Requires controller v0.172.0 (deployed to both demo boxes
before this push, since ${IMPORT_PATH} is unset on older controllers).
Storage-layout header comments updated in both composes — they are the only
in-repo documentation of the layout.
Moving a template out of templates/ un-offers it but also makes the
controller's orphan detector see it as GONE for anyone already running the
app - flagging their working install Elavult with a Torles button. Withdrawing
an app must never take a working app away from a customer.
Optional lifecycle: available|hidden|abandoned in .felhom.yml instead.
plant-it returns to templates/ as the first abandoned app; retired/ removed.
Resolvability gate skips (and reports) non-available apps.
wanderer: ghcr.io/flomp/wanderer:0.16.0 is a ghost - upstream split the app
into web+db images, moved registry and renamed the org. Restructured to
upstream's own v0.20.0 compose (3 services, new /data/plugins volume, second
public hostname for PocketBase, meilisearch pinned DOWN to upstream's v1.36.0
per the R-42 ruling).
plant-it: retired. The repo name was wrong (plant-it-server) but upstream has
DELETED self-hosting; last server image is 2024-12-10 and it needs MySQL+Redis
the template never had. Moved to retired/ rather than deleted - reversible.
R-41 slice 1: check-image-resolvable.py. Encodes two traps - manifest inspect
exits 0 while printing toomanyrequests, and the inverse, where the first sweep
called 24 of 65 pins dead because Hub throttled it. Ambiguity is INCONCLUSIVE,
never an accusation.
With 2.6 pulling and booting, wger still sat unhealthy: the healthcheck got
wget rc=4 (network failure, not an HTTP error) because nothing listens on :80.
wger's own log says it plainly -- 'Using django's development server on port
8000...' -- and the image declares ExposedPorts {"8000/tcp"}.
Corrected both the healthcheck target and the Traefik loadbalancer port, which
was pointing at :80 as well (so routing would have failed even once healthy).
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Correcting my own earlier revert. Reverting to 2.3 was WRONG: wger/server:2.3 is
no longer published on Docker Hub (only 2.4, 2.5, 2.6 resolve), so that pin could
not be pulled at all -- the deploy was accepted and no container was ever created.
An unpullable pin is strictly worse than the problem it was meant to avoid.
Every available version (2.4/2.5/2.6, verified) reads the whole DJANGO_DB_* set
unconditionally, even with the sqlite engine. So supplying it is not a hack
around one version -- it is now the only way to run wger at all. Verified live:
2.6 with the full set boots and applies its migrations ('Applying auth.0001_initial
... OK').
USER/PASSWORD/HOST/PORT are ignored by the sqlite backend but must be present.
DJANGO_DB_DATABASE points into the existing wger_data volume (/home/wger/db),
which is where wger's own default sqlite file lived -- so an existing install is
not pointed at an empty database somewhere else.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
With the DATABASE_URL fix zipline started cleanly ('server started
hostname=0.0.0.0 port=3000') but stayed unhealthy. Probing the running container
from inside:
/api/health -> 404
/api/healthcheck -> 200
v4 renamed the endpoint; the template still probed the v3 path.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
After the memory fix tandoor still never went healthy. The probe output was
'wget: can't connect to remote host (127.0.0.1): Connection refused' while the
app log was still printing 'Booting worker' / 'Running django-vite' -- i.e. the
app had simply not finished starting. With start_period 30s and retries 3 the
healthcheck gave up at roughly two minutes.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
With the image and the pepper fixed, homebox still sat unhealthy. The probe's own
exit code was 8 (server error response) and homebox's log shows exactly why:
method=HEAD path=/api/v1/status status=405
method=GET path=/api/v1/status status=200
'wget --spider' issues a HEAD request; homebox's status endpoint only implements
GET. Switched to 'wget -q -O /dev/null' (a real GET).
NOTE: 24 templates use --spider. The rest validated green, so their endpoints do
answer HEAD -- but this is a latent trap worth a convention note (see the
campaign doc).
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
wger 2.6 crash-loops on a fresh deploy. Its settings/main.py reads the whole
DJANGO_DB_* set unconditionally -- even when the engine is sqlite:
ImproperlyConfigured: Set the DJANGO_DB_DATABASE environment variable
...then, once that was supplied:
ImproperlyConfigured: Set the DJANGO_DB_USER environment variable
Satisfying it means either stuffing in dummy USER/PASSWORD/HOST/PORT values that
the sqlite backend ignores, or giving wger a real Postgres sidecar. The first is
a hack, the second is compose restructuring -- both outside this campaign's
allowed-fix set.
Reverted to the previously shipped 2.3. NOTE: the interim DJANGO_DB_DATABASE
line added earlier in this campaign is reverted TOO, deliberately -- on 2.3 the
sqlite path came from wger's own default, and pinning a different explicit path
would have pointed an existing customer's wger at an empty database.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Both crashes above were masking a second defect: once each app actually stayed
up, it sat permanently unhealthy because its healthcheck called wget, and
neither image has wget or curl -- only node. Traefik does not route to an
unhealthy container, so both would have served 404 to the customer regardless.
Switched both to the Node-exec family (http.get, exit non-zero on 5xx/error).
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
zipline crash-looped 15 times on a fresh deploy:
ERROR config::readDbVars] No database environment variables found
(DATABASE_URL or all of [DATABASE_USERNAME, DATABASE_PASSWORD,
DATABASE_HOST, DATABASE_PORT, DATABASE_NAME]), exiting...
In Zipline v4 the database variable is DATABASE_URL with NO prefix, while
CORE_SECRET kept its prefix (confirmed against the upstream config docs) -- so
only the DB variable was renamed here.
The template was already pinned to the v4 line (4.0.0) before this campaign, and
v4 has always required DATABASE_URL, so zipline has been undeployable from the
catalog for the whole v4 series. PRE-EXISTING defect, surfaced by the sweep.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
wishlist deployed but never created a container: the pull fails because
cmintey/wishlist:1.9.0 no longer exists on Docker Hub. Upstream publishes to
ghcr.io/cmintey/wishlist, where the current release is v0.66.0 (2026-07-08,
confirmed against the GitHub releases API).
Note the version string goes '1.9.0' -> 'v0.66.0'. That is NOT a downgrade: the
old Hub pin used a scheme upstream does not publish under, so the catalog was
pinned to an image reference that has no counterpart in the real release stream.
This was a PRE-EXISTING defect -- wishlist was undeployable before this campaign
too (it was not version-bumped by the sweep).
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
wger 2.6 crash-looped 15 times on a fresh deploy:
django.core.exceptions.ImproperlyConfigured:
Set the DJANGO_DB_DATABASE environment variable
The template already selected the sqlite3 engine, but 2.6 no longer supplies a
default database path -- it must be given explicitly even for sqlite. Pointed at
the existing wger_data volume (/home/wger/db) so the database survives redeploys.
Env var proven wrong by the deploy.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
tandoor 2.6.13 never reached healthy; gunicorn workers were SIGKILLed in a loop:
[ERROR] Worker (pid:366) was sent SIGKILL! Perhaps out of memory?
[ERROR] Worker (pid:367) was sent SIGKILL! Perhaps out of memory?
Separately, .felhom.yml mem_limit was already inconsistent with the compose file:
it claimed 512M while the services summed to 512+256 = 768M. Per the REUSE.md
rule mem_limit is the SUM of the compose limits, so it is now 1024+256 = 1280M.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
rallly 4.11.1 (Next.js 16) booted and was immediately killed by the cgroup OOM
killer, 14 times in 420s:
▲ Next.js 16.2.6
- Local: http://localhost:3000
✓ Ready in 0ms
Killed
Migrations applied fine; it simply cannot live in 256M any more. App limit
256M -> 768M; sidecar Postgres left at 256M; .felhom.yml mem_limit updated to
the new sum (768+256 = 1024M) per the REUSE.md rule.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
With the tag fixed so the image actually pulls, homebox 0.26.2 then panicked on
every start (16 restarts):
panic: auth.api_key_pepper must be set to at least 32 bytes; generate with
`openssl rand -base64 48` and provide via HBOX_AUTH_API_KEY_PEPPER.
Rotating it invalidates all issued API keys
Added as a generated base64key:48 secret (matching upstream's own suggestion)
and marked data_key: true -- rotating it invalidates every issued API key, so
restore must recover the original rather than mint a new one.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
papra was NOT version-bumped by this campaign -- it crash-looped at its existing
26.6.1-rootless pin, so this is a PRE-EXISTING catalog defect: papra has never
been deployable from this template.
Invalid configuration: In production, the auth secret must not be the default
one. Please set a secure auth secret using the AUTH_SECRET environment
variable.
Added as a generated hex:32 secret and marked data_key: true -- it signs
sessions, so regenerating it on restore would invalidate every login.
Campaign 7 catalog sweep.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE