The operator's request of 2026-10-01 recorded in CONTEXT with the reviewer's defaults (new apps only; the 53 get a
read-only gap page). The pilot: the draft caught R-752 and R-755, missed R-737 and (for a new app) R-738; the
sharpened and new rows then found R-762 (wger serves no static files or photos), R-763 (strangers sign up, guest
accounts), R-764 (no mail). Also R-758 (8 mem_limit under the sum), R-759 (wger's open record rows), R-760
(vikunja healthcheck), R-761 (logo comment). R-755 note. Register 392 -> 399.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Decision 52: catalog 6a3ead9 (re-test entries, gates, decoys, the monthly command); proven end to end on 9202
through the leg; runbook monthly-floating-retest.md; nothing to re-test on the engine lines today.
- Decision 53 + R-741: controller v0.284.2 (0.284.0/0.284.1 never floored — two wiring faults found live on 9202);
floor 0.284.2; one-time sweep 9202 26.6 -> 5.7 GB, demo-hp 24.3 -> 13.5 GB; the install hold proven as a stranger.
- Golden 0.284.2 baked, round-trip identical, vouched (agent 0.138.0, min_agent 0.131.0); the gate prints OK.
- Rows 377 -> 383: opened R-743..R-748, closed R-736, R-737, R-740, R-741, R-748; narrowed R-739, R-698, R-446.
- register_shape_gate: a lettered id (R-88a) is a row too (R-748), with a decoy seen red.
Evidence: documentation/audits/night-rulings-2026-09-30/, documentation/tests/golden-0.284.2-2026-09-30/.
Report: REPORT-night-rulings-2026-09-30.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Twelve steps published on both venues (catalog): first ladders for calibre-web, gitea, wger,
crafty-controller, uptime-kuma, zipline (two steps); within-major emby, ghost, home-assistant,
outline, rallly. immich's step v3.0.3 -> v3.2.2 re-proven at 768M (Part D).
- Part C: the night leg skips a digest-only change AND the catalog never records a same-tag re-test,
so a same-name upstream fix reaches no box. Row R-740; the decision in STATUS; `09` decision 30
carries a dated note (the decision itself unchanged).
- Rows: 369 -> 377. Opened R-735..R-742; closed R-735, R-738, R-742; narrowed R-462, R-624, R-446,
R-440, R-734, R-732. The currency audit gains §1b (and corrects its 31+1 to 30+2).
Evidence: documentation/audits/more-night-apps-2026-09-30/. Report: REPORT-more-night-apps-2026-09-30.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- R-732: the first-start geodata import runs up to 9 concurrent 5000-row INSERTs; the database needs
~400 MB anon + ~170 MB touched shared_buffers (the image's FIXED 512MB, not host-RAM sizing). 512M fits
only with swap (bench swap 0: 61-104 kills; 9202 swap 512 MiB: survived by swapping). Controls: swap
alone, limit alone flip it; shared_buffers 128MB alone does not. Catalog 56c4888: v3.2.4 + 768M,
proven with swap off on both venues. audits/immich-first-start-2026-09-30/A-cause.md.
- R-730: scripts/iso/build-felhom-iso.sh refuses an uncommitted/untracked/unpushed tree (no bypass),
records repo-commit from the gate and iso-v<version>; test iso/test/clean-tree.sh, red-proof run
(status check removed -> 2 of 4 cases fail -> restored).
- R-731 narrowed (gitea 28.0.0 GA; mariadb 13.0 a short-term Rolling line). R-676 note.
- New rows R-733 (bench has no swap, boxes 512 MiB), R-734 (immich .immich markers -> files_may_change).
- STATUS: the golden line corrected (no bake is due; 0.283.1 is the newest release). Register 364 -> 366.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- R-463 CLOSED: rallly, outline, sparkyfitness -> PostgreSQL 18 (catalog 25ffd89 / aeb0cd6 / 1666572),
each proven on the bench and on 9202 through the guarded Update with one undo case; zipline,
adventurelog, immich stay on 16 by 09 decision 42's rule (upstream runs 16 / 16 / 14).
- Part F: bookstack, kimai, audiobookshelf, n8n, navidrome, grafana, komga moved on both venues;
immich not (R-732: its first start OOM-killed its database on the bench).
- R-687 item 4 PROVEN LIVE on demo-hp: two deferrals while the leg stepped two apps, the whole-guest
backup on the first poll after, success; config + window put back and read back.
- Catalog currency audit (Part D): 25/53 behind inside a major, 19 across; night-updatable 28 -> 31 (+1).
- R-446 and R-440 narrowed (measured on demo-hp); R-624 corrected (outline, rallly, zipline have routes);
R-548 note (demo-hp's local tier refused for space since 09-27).
- New rows: R-730 (the ISO 1.29.0 build commit cannot be proven -> no tag; installer-v* is the script's
line), R-731 (tag-shape switches), R-732. Register 361 -> 364.
- 09 §3: the 2026-09-30 operator notes (by day; Tester-2 pre-checks done; decision 52 not needed, not
recorded); §6.4 dated currency note. STATUS, CONTEXT, REPORT-pg-last-six-2026-09-30.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Audit doc, STATUS one sentence, capability-map first-hour row (walk re-proven, day-one
off-site sentence narrowed: R-720/R-726/R-727), teardown in four layers, R-600 measured again.
Secret scan over all audits: 0 hits, positive control 1.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
the backup, after it and seconds before the press read back; cut-off copy
held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
ruled; R-646 opened. STATUS asks the floor question.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
(PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
back with data written before AND after the backup. The product's loader
cannot do it: over a migrated PG database it fails on the new tables'
foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.
No product code. Live catalog untouched; 9202 back on it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §3: decisions 11 (one window = a leg of the backup chain), 12 (automatic,
per-box switch on by default), 13 (the test decides, not the tag - replaces
decision 3's "never across a major"), 14 (the ladder), 15 (the box undoes a
failed update - replaces §6.1's no-auto-undo), 16 (Postgres majors converted
by the box), 17 (digests), 18 (fleet view, later).
- 09 §3b marked ANSWERED with a pointer per question; kept as the reasoning.
- 09 §6.2 rewritten to the ruled shape; §6.1 abort paragraph and §4 point at 15.
- Register: R-450, R-451, R-446, R-463 cite the decisions.
- STATUS: the seven questions no longer wait on the operator.
Documents only. No product code.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Both demo boxes took it in under 20 seconds and read healthy. drill-r50 is DOWN and correctly HELD:
its agent is 0.129.0, below the declared MinAgent, so the hub refuses to hand it a controller it
cannot run (R-472's guard, working).
Boxes below the floor: 5 -> 3. The three that remain are drill-r50 and the two guests the hub does
not hear from, including scratch 9202 which runs hub.enabled: false.
Read back from the hub rather than from the POST's own answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.
B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.
C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".
F (R-614): phase done before the remove, no phase at all after redeploying the same name.
Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.
09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits
app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the
OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard.
CUSTOMER skipped, correctly.
I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from
a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape
again, from the same hand, four days later.
R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and
emits once per container lifetime, so 4530 worker kills over six hours produced exactly one
warning-level mail - indistinguishable from one transient kill. The magnitude was already collected
(App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event.
Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found because the operator heard the fans. romm 5.3.0 was promoted that morning; the update read
`done` and the app ran clean for two hours, then OOM-crash-looped for six - 4530 worker SIGKILLs,
~500% CPU, host load 5.2 while otherwise idle, and nothing alarmed.
Raising the limit to 768M was still a guess and fixed nothing (memory.peak hit exactly 768 MiB).
Measured instead: ~216 MiB per warm uvicorn worker, so the image's default of 4 workers needs
~882 MiB. /init reads WEB_SERVER_CONCURRENCY; set to 2.
Proven under load, not just at idle: 26,645 requests over 300 s, memory 416-614 MiB against 768,
trending down, zero SIGKILLs, OOMKilled false. Idle CPU 500% -> 1.64%.
The first soak measured nothing - it was pointed at the scratch-guest subdomain, every request
404'd at traefik in 9 ms, and the counter reported 14,026 successes. Positive and negative controls
are now asserted before any load is driven.
Carried into R-462: `proven` has meant "the update applied and the data survived", not "the new
version runs".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven,
5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the
right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven
live for the first time. Each app also got the half the update night skipped: a restore from its own
copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2.
R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it
waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy.
The controller's own words: "not healthy within 5m0s (last: no probe container)".
R-633 opened: a remove sent during a restore reports success and leaves a container restarting with
a live public route. The product already refuses that clash for update and for restore, naming the
blocker; remove has no such guard.
R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is
then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others.
R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness -
named, with what each cost. No product code. The live catalog's image: lines are byte-identical to
the start of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part 1: tandoor/zipline/wger probes corrected in the catalog and red-proofed live on 9202 in both
directions - "Nem egeszseges" with the front door serving 200, then "Fut" after the real sync with
no redeploy. tandoor's failed edge re-walked: done at +41.1s where it was failed at +361.9s.
Part 2: fifteen proven versions on the live catalog, one commit per app; the guarded Update pressed
on four apps on demo-hp, all four done.
Opened: R-630 (paperless-ngx's probe has never run on any box - a silent absence, worse than the
wrong probe that was found in one night), R-631 (five templates no static rule can judge),
R-632 (28 of 53 templates never deployed by any drill). Closed: R-618.
Register 318 -> 321. No product code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.
WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.
THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.
WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.
TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.
Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.
Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.
Gates: repo_gates.py --fast, all 15 OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all
down or blocked; both demo boxes SERVED.
Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and
two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live
compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s
after the cut decision and the seeded data read back intact — so the branch that is
one step from old-binary-on-migrated-database is now evidence, not argument.
Instrument limit stated: `starting` lasts under a second; all three landed in
`verifying`, which RecoverUpdates handles in the same branch.
Part 3 (R-611) — the night the previous session skipped without saying so. An app
updated with nobody pressing anything; a terminally-refused app was pressed exactly
once and never again over three passes. The unattended HOLD was NOT produced: the
within-a-major rule correctly refused the broken edge before it was attempted, so
Q4 still rests on the attended hold from slice 4. Said plainly rather than implied.
Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh
install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard),
R-614 (stale update phase survives a redeploy). R-520's pointer corrected.
Catalog: two drill pairs, both reverted; every image line byte-identical to
ff9717d3. The alpine:3.20 negative control a security review flagged is cleared.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS