# DRILL — night 2026-09-24 (run 2026-09-24 from 11:07 CEST): the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; automatic updates spiked; a chaos hour > Written in parts. The "Not done, or changed" section and the results are filled at the end; **the > Part E schedule below was written and committed BEFORE round 1.** ## Part E — the schedule (committed before round 1) Drawn by `night-2026-09-24/tools/chaos_schedule.py`, seed **20260924** (constraints: every action at least once; controller killed at most 3 times, ≥ 20 min apart; Start-after-unhealthy after a stop; Remove-on-held after a hold). Guest **9202** (scratch), controller v0.269.1, drill catalog. seed 20260924, 12 rounds | round | action | accident | |---|---|---| | 1 | oom_app | power_cut | | 2 | install | docker_restart | | 3 | cutoff_undo_file_app | power_cut | | 4 | failing_step | controller_killed | | 5 | remove_held | none | | 6 | failing_step | whole_box_backup | | 7 | crash_loop | whole_box_backup | | 8 | oom_app | disk_to_floor_plus_1g | | 9 | remove_held | controller_killed | | 10 | ladder_step | docker_restart | | 11 | cutoff_undo_file_app | disk_to_floor_plus_1g | | 12 | start_after_unhealthy | none | **The app each round acts on — fixed here, not drawn** (an app's state after round N decides what N+1 can do): | round | app | why | |---|---|---| | 1 oom_app | `chaosoom` (drill-only: alpine, 32M cap, a 200 MB eater in a loop) | out-of-memory storm → the box must stop it | | 2 install | n8n at its ladder's FIRST `from` (2 steps behind), seeded through its front door | gives round 10 a step | | 3 cutoff_undo_file_app | nextcloud (the file app), a failing step (redis 7-alpine → 7.4.0-alpine, probe 8999), the NEWEST undo copy's marker cut at verifying → the hold must name the second drive → its whole restore, then the seeds + files read back | decision 26 under a power cut | | 4 failing_step | wishlist v0.67.1 → `:latest` with probe 8999 → undo | the undo under a controller kill | | 5 remove_held | gokapi (held since A3: the box stopped its crash loop) — Remove with data | Remove on a held app | | 6 failing_step | navidrome 0.64.1 → `:latest` with probe 8999 → undo | the undo under a backup run | | 7 crash_loop | `chaoscrash` (drill-only: alpine that exits after 3 s) | the box must stop it | | 8 oom_app | `chaosoomb` (a second eater) | the OOM stop under a full disk | | 9 remove_held | `chaoscrash` (held since round 7) | Remove on a held app under a controller kill | | 10 ladder_step | n8n, one step (2 → 1) | a good step under a docker restart | | 11 cutoff_undo_file_app | nextcloud again | decision 26 with the disk 1 GB above the floor (a `disk` refusal is a valid outcome) | | 12 start_after_unhealthy | `chaosoomb` (held since round 8) — Start → one more try → the repeat stop says support is told | decision 28's second half | **Changed before round 1:** the accident "a whole-box backup started on 9202" is the box's own full backup run (`POST /api/backup/run`, as on 2026-09-23), not a vzdump: demo-hp's backup storage had ~4 GB free and the host root was at 90 %, so a vzdump of 9202 could have filled the HOST (demo-hp also had R-672's full pool today). **How accidents are fired:** the runner arms the round and waits; the session fires each host-level accident as its own visible command (power cut = `pct stop 9202; pct start 9202`, kill -9 of the controller PID, `systemctl restart docker` in the guest, a `fallocate` to leave 3 GB free on `/var/lib/docker`).