- A RE-TEST entry: from == to, digest = the registry's new digest, digest_from = the tested one, box_evidence.
ladder.check_entry refuses one with no new digest, no digest_from or no box proof; check-test-record rule 2b
ties digest_from to the previous entry's digest; check-test-record-move now judges re-tests too (they change
.felhom.yml only — the gate looked at compose moves alone) and refuses a digest the registry no longer serves.
Decoys: 8 cases in test_gate_decoys.py, seen red with the rules switched off.
- upgrade-test.py --retest <app> [svc]: FROM the ladder head's tested digest TO the registry's current one, the
full method; --write-ladder writes a re-test entry (plain refs + digest_from), refusing without the box venue or
when the registry moved again. Writer tests, red-proofed.
- scripts/retest-floating.py — ONE command: --dry-run lists, --engines-only is the ruled start; bench, box
(retest_box.py on 9202 via the drill catalog), writer, gates, one commit per app. box_walk.py moves the box
client into the catalog. Run today: nothing to re-test on the database/redis lines.
- End to end on 9202 (drill): docmost at the OLD redis digest, the re-test, "run tonight's chain now" -> the leg
pressed it, the new digest runs, read back, badge current.
- Also: upgrade-test.py BENCH_ENV_OVERRIDES (R-739, wanderer's DB address on the bench, recorded per verdict);
test_gate_decoys.py read kimai's tag and date from the clone (red on main since kimai moved).
Evidence: felhom.eu/documentation/audits/night-rulings-2026-09-30/A/
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- upgrade-test.py --restep <verdict> --definition <dir> --catalog <c> --evidence <rel>: the only way a
superseded step's own files (steps/<key>.yml + .felhom.yml) change after the fact (Part D: immich's
step 0b8272068aab36bf still pins 512M). Refuses a non-proven or OOM-killing re-proof, the head entry,
and a definition whose images are not the step's; leaves the ladder entry untouched (digests, box
evidence, the box's failed-step fingerprint). Two tests; the gate accepts the result.
- zipline's verify waited on `/` for 200/302/307; on 9202 it answers 301 and the readback returned
False with no line — it now waits on /api/healthcheck like its seed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- fixtures (upgrade_fixtures_box.py): calibre-web (Upload form -> OPDS readback + the served EPUB),
wger (web login -> weight entry API), crafty-controller (API v2 roles), uptime-kuma (socket.io
polling: setup, login, addStatusPage -> public /api/status-page/<slug>); gitea posts its own
first-run installer form (R-624's fixable case) and keeps the admin CLI for an installed one.
calibre-web and wger run the template's own after_install on the bench, which has none.
- R-735: the bench's `password:N:special` now has the controller's shape (randomWithSpecial);
test seen failing first (length 32, no special), then 45/45.
- upgrade_boxport / the memory watch: a backend traefik reaches over https (loadbalancer.server.scheme)
is reached over https on the bench too (crafty-controller).
- upgrade-test: files-before/after-detail.json and `files_changed_detail` NAME the files behind a
files_may_change mark (R-734's method, now in the harness); test ChangedFiles.
- test_catalog_gates: the gate count was stale (9, the runner has 10) and red on main; now 10.
Evidence: felhom.eu/documentation/audits/more-night-apps-2026-09-30/
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- memory watch: load no longer follows redirects to the unresolvable test
domain (every request had been 'err' on box-fixture apps), and sends the
app's own Host; the app's own memory (anon) is sampled beside the cgroup
peak, and memory_tight reads anon where measured (09 decision 22, CC).
- ladder writer: memory_peak_pct = anon (else cgroup peak), with
memory_basis and memory_cgroup_peak_pct beside it.
- fixtures: opengist 1.15 serves under /-/ and marks its cookie Secure
(readback = the account's own page + a never-created user 404); komga's
user endpoint is /api/v2/users/me; a wishlist fixture (form actions).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
update_ladder: in .felhom.yml, one JSON entry per line (spiked live on
controller v0.266.0 and v0.267.0 first). Two gates: check-test-record.py
(static, CI too) and check-test-record-move.py (history + registry for
moved refs only). 16 decoys, 3 red-proofs. The ONLY writer is
upgrade-test.py --write-ladder (bench AND box proven, digests resolved).
Harness v3: box fixtures on the bench, files_may_change.
Backfill: the 21 moves of 2026-09-22, 21 proven from their records.
No image: line moved.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
After an edge reads back, --soak seconds (default 600) of light load while
the kernel's own oom_kill counter is read host-side from the container's
cgroup. A kill or restart turns proven into failed; a peak over 80% of the
limit adds the memory_tight mark. New Romm fixture; edges M1 / M1old.
Red-proof on scratch 9202: M1old (template as promoted, 512M, 4 workers)
OOM-killed at +76 s -> failed. M1 (current, 768M, 2 workers) proven, 0
kills in 608.5 s, peak 81% -> memory_tight.
Test code only; no template changed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Test code only — no template changed and no image: line moved.
The update night walked real within-a-major upstream edges on scratch guest 9202 through the
product's own guarded Update, against a PRIVATE DRILL CATALOG; the live catalog was never
touched. This brings the expensive half of that work — the seed routes — back into the harness
so the same edges can be run here WITH their ABORT step, which the box deliberately does not
offer (09 6.1: whether the old image starts on migrated data is per-app and unpredictable).
- upgrade_fixtures.py: ActualBudget, Navidrome, AudiobookShelf, Vikunja. Each seeds through the
app's OWN interface (R-156); each carries a negative control run on every verify(), so a
readback that has broken into always succeeding fails instead of passing everything.
- upgrade-test.py: edges U1..U7, all real upstream moves existing 2026-09-21 that this catalog
has NOT made, each holding its database engine constant.
- Limitations kept: Navidrome and AudiobookShelf seed the DATABASE half only, and say so.
OWED, stated so it is not mistaken for done: the U1..U7 harness RUNS, and with them the per-app
ABORT answers. The code is in; the runs are not.
Gates: catalog_gates.py --fast — image-pins, engine-major, catalog-since, copy-i18n all OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-459. The harness returned proven for E3b while MariaDB was logging that the
conversion it requires had been skipped. The verdict was right - the app's data
survived, which is what it asked - but the harness watched the app and the
migration log, and neither looks at engine state.
engine_state_after now carries each database service's own answer: MariaDB's
datadir version plus 'mariadb-upgrade --check-if-upgrade-is-needed', and
PostgreSQL's PG_VERSION.
It sits BESIDE the verdict and is never folded into it. An unconverted datadir is
not known to be a failure - 5 of 5 restarts showed no degradation - so a verdict
that called it failed would encode an unproven judgement, which is worse than
reporting a fact and letting a person read both.
No template changed. Nothing with MARIADB_ in it is committed by this work: that
is a fleet-wide decision the operator owns, and it affects four apps.
R-449. Until today one upgrade out of 53 had ever been measured - Nextcloud, by
hand, in a spike - and the whole update arc was designed against that single data
point.
Per edge: deploy at FROM, seed through the app's OWN interface, prove the seed
reads back, swap to TO, ask the app for the data again, then put the FROM images
back and record what happens - verbatim, and never called a rollback.
Success is an application-level readback, not file identity: survive2.py's
sha256+inode rule is right for a redeploy and wrong for an upgrade, because a
migration is supposed to rewrite files. And nothing is ever seeded by hand (R-156)
- an app with no non-browser route is recorded inconclusive, never faked.
C3 is a negative control whose TO image exits immediately, and it must be run
first: it came back failed, which is what makes the greens mean anything.
The bookstack fixture uses artisan for both halves and carries its own negative
control on every call, because the obvious HTTP-login readback cannot work: the
template's https APP_URL makes the session cookies secure, so curl over http gets
419 on every login and it looks exactly like a wrong password.