The records the catalog's ladder entries cite (apps/<app>/bench and
apps/<app>/verdict.json), the drill tools, the spike, R-626's run,
R-612/R-613 red-proofs, the chaos schedule (seed 20260923) drawn
BEFORE round 1. Docs and register follow at the end of the night.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Allowlisted, not operator-only, seeded for new households and added
once (add-only) to every existing enabled_events row. mail.event entries
in hu and en name the app from details.stack_name. Per-app cooldown on
both the operator and the household leg.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
the backup, after it and seconds before the press read back; cut-off copy
held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
ruled; R-646 opened. STATUS asks the floor question.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the
full-system backup waits for the update leg, inside its window; built later).
- Bake-off on 9202, docmost / romm / vikunja: both methods pass every case;
the folder copy wins because an app with no database server gets no dump,
so dump-and-load would need the folder copy anyway. 1-5 s extra downtime,
~420 MB/s, disk = the volumes.
- R-645 filed: lifting an update hold by hand lets the recovery unit be
re-captured with the failed definition within seconds.
Documents and evidence only; product code follows in the controller.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
(PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
back with data written before AND after the backup. The product's loader
cannot do it: over a migrated PG database it fails on the new tables'
foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.
No product code. Live catalog untouched; 9202 back on it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §3: decisions 11 (one window = a leg of the backup chain), 12 (automatic,
per-box switch on by default), 13 (the test decides, not the tag - replaces
decision 3's "never across a major"), 14 (the ladder), 15 (the box undoes a
failed update - replaces §6.1's no-auto-undo), 16 (Postgres majors converted
by the box), 17 (digests), 18 (fleet view, later).
- 09 §3b marked ANSWERED with a pointer per question; kept as the reasoning.
- 09 §6.2 rewritten to the ruled shape; §6.1 abort paragraph and §4 point at 15.
- Register: R-450, R-451, R-446, R-463 cite the decisions.
- STATUS: the seven questions no longer wait on the operator.
Documents only. No product code.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Both demo boxes took it in under 20 seconds and read healthy. drill-r50 is DOWN and correctly HELD:
its agent is 0.129.0, below the declared MinAgent, so the hub refuses to hand it a controller it
cannot run (R-472's guard, working).
Boxes below the floor: 5 -> 3. The three that remain are drill-r50 and the two guests the hub does
not hear from, including scratch 9202 which runs hub.enabled: false.
Read back from the hub rather than from the POST's own answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.
B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.
C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".
F (R-614): phase done before the remove, no phase at all after redeploying the same name.
Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.
09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Both from their own records, not re-run.
calibre-web: the flash_error refusal is in its log; that walk predated the refusal-capture code.
The report's own prose already said so while the table disagreed.
calcom: the restore was accepted and the app read running; the classifier then saw "starting", a
settling state it did not list beside running/unhealthy, and fell through to failed. The brief
supposed a read-back artefact - that is wrong, and restore_state_seen says so.
Restores correctly refused: 2 -> 3.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits
app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the
OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard.
CUSTOMER skipped, correctly.
I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from
a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape
again, from the same hand, four days later.
R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and
emits once per container lifetime, so 4530 worker kills over six hours produced exactly one
warning-level mail - indistinguishable from one transient kill. The magnitude was already collected
(App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event.
Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS