Audit doc, STATUS one sentence, capability-map first-hour row (walk re-proven, day-one
off-site sentence narrowed: R-720/R-726/R-727), teardown in four layers, R-600 measured again.
Secret scan over all audits: 0 hits, positive control 1.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
the backup, after it and seconds before the press read back; cut-off copy
held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
ruled; R-646 opened. STATUS asks the floor question.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the
full-system backup waits for the update leg, inside its window; built later).
- Bake-off on 9202, docmost / romm / vikunja: both methods pass every case;
the folder copy wins because an app with no database server gets no dump,
so dump-and-load would need the folder copy anyway. 1-5 s extra downtime,
~420 MB/s, disk = the volumes.
- R-645 filed: lifting an update hold by hand lets the recovery unit be
re-captured with the failed definition within seconds.
Documents and evidence only; product code follows in the controller.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
(PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
back with data written before AND after the backup. The product's loader
cannot do it: over a migrated PG database it fails on the new tables'
foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.
No product code. Live catalog untouched; 9202 back on it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §3: decisions 11 (one window = a leg of the backup chain), 12 (automatic,
per-box switch on by default), 13 (the test decides, not the tag - replaces
decision 3's "never across a major"), 14 (the ladder), 15 (the box undoes a
failed update - replaces §6.1's no-auto-undo), 16 (Postgres majors converted
by the box), 17 (digests), 18 (fleet view, later).
- 09 §3b marked ANSWERED with a pointer per question; kept as the reasoning.
- 09 §6.2 rewritten to the ruled shape; §6.1 abort paragraph and §4 point at 15.
- Register: R-450, R-451, R-446, R-463 cite the decisions.
- STATUS: the seven questions no longer wait on the operator.
Documents only. No product code.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.
B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.
C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".
F (R-614): phase done before the remove, no phase at all after redeploying the same name.
Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.
09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits
app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the
OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard.
CUSTOMER skipped, correctly.
I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from
a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape
again, from the same hand, four days later.
R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and
emits once per container lifetime, so 4530 worker kills over six hours produced exactly one
warning-level mail - indistinguishable from one transient kill. The magnitude was already collected
(App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event.
Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found because the operator heard the fans. romm 5.3.0 was promoted that morning; the update read
`done` and the app ran clean for two hours, then OOM-crash-looped for six - 4530 worker SIGKILLs,
~500% CPU, host load 5.2 while otherwise idle, and nothing alarmed.
Raising the limit to 768M was still a guess and fixed nothing (memory.peak hit exactly 768 MiB).
Measured instead: ~216 MiB per warm uvicorn worker, so the image's default of 4 workers needs
~882 MiB. /init reads WEB_SERVER_CONCURRENCY; set to 2.
Proven under load, not just at idle: 26,645 requests over 300 s, memory 416-614 MiB against 768,
trending down, zero SIGKILLs, OOMKilled false. Idle CPU 500% -> 1.64%.
The first soak measured nothing - it was pointed at the scratch-guest subdomain, every request
404'd at traefik in 9 ms, and the counter reported 14,026 successes. Positive and negative controls
are now asserted before any load is driven.
Carried into R-462: `proven` has meant "the update applied and the data survived", not "the new
version runs".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS