Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.5 KiB
DRILL — night 2026-09-24 (run 2026-09-24 from 11:07 CEST): the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; automatic updates spiked; a chaos hour
Written in parts. The "Not done, or changed" section and the results are filled at the end; the Part E schedule below was written and committed BEFORE round 1.
Part E — the schedule (committed before round 1)
Drawn by night-2026-09-24/tools/chaos_schedule.py, seed 20260924 (constraints: every action at least
once; controller killed at most 3 times, ≥ 20 min apart; Start-after-unhealthy after a stop; Remove-on-held
after a hold). Guest 9202 (scratch), controller v0.269.1, drill catalog.
seed 20260924, 12 rounds
| round | action | accident |
|---|---|---|
| 1 | oom_app | power_cut |
| 2 | install | docker_restart |
| 3 | cutoff_undo_file_app | power_cut |
| 4 | failing_step | controller_killed |
| 5 | remove_held | none |
| 6 | failing_step | whole_box_backup |
| 7 | crash_loop | whole_box_backup |
| 8 | oom_app | disk_to_floor_plus_1g |
| 9 | remove_held | controller_killed |
| 10 | ladder_step | docker_restart |
| 11 | cutoff_undo_file_app | disk_to_floor_plus_1g |
| 12 | start_after_unhealthy | none |
The app each round acts on — fixed here, not drawn (an app's state after round N decides what N+1 can do):
| round | app | why |
|---|---|---|
| 1 oom_app | chaosoom (drill-only: alpine, 32M cap, a 200 MB eater in a loop) |
out-of-memory storm → the box must stop it |
| 2 install | n8n at its ladder's FIRST from (2 steps behind), seeded through its front door |
gives round 10 a step |
| 3 cutoff_undo_file_app | nextcloud (the file app), a failing step (redis 7-alpine → 7.4.0-alpine, probe 8999), the NEWEST undo copy's marker cut at verifying → the hold must name the second drive → its whole restore, then the seeds + files read back | decision 26 under a power cut |
| 4 failing_step | wishlist v0.67.1 → :latest with probe 8999 → undo |
the undo under a controller kill |
| 5 remove_held | gokapi (held since A3: the box stopped its crash loop) — Remove with data | Remove on a held app |
| 6 failing_step | navidrome 0.64.1 → :latest with probe 8999 → undo |
the undo under a backup run |
| 7 crash_loop | chaoscrash (drill-only: alpine that exits after 3 s) |
the box must stop it |
| 8 oom_app | chaosoomb (a second eater) |
the OOM stop under a full disk |
| 9 remove_held | chaoscrash (held since round 7) |
Remove on a held app under a controller kill |
| 10 ladder_step | n8n, one step (2 → 1) | a good step under a docker restart |
| 11 cutoff_undo_file_app | nextcloud again | decision 26 with the disk 1 GB above the floor (a disk refusal is a valid outcome) |
| 12 start_after_unhealthy | chaosoomb (held since round 8) — Start → one more try → the repeat stop says support is told |
decision 28's second half |
Changed before round 1: the accident "a whole-box backup started on 9202" is the box's own full backup run
(POST /api/backup/run, as on 2026-09-23), not a vzdump: demo-hp's backup storage had ~4 GB free and the host
root was at 90 %, so a vzdump of 9202 could have filled the HOST (demo-hp also had R-672's full pool today).
How accidents are fired: the runner arms the round and waits; the session fires each host-level accident
as its own visible command (power cut = pct stop 9202; pct start 9202, kill -9 of the controller PID,
systemctl restart docker in the guest, a fallocate to leave 3 GB free on /var/lib/docker).