3208cb2083
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
51 lines
3.5 KiB
Markdown
51 lines
3.5 KiB
Markdown
# DRILL — night 2026-09-24 (run 2026-09-24 from 11:07 CEST): the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; automatic updates spiked; a chaos hour
|
|
|
|
> Written in parts. The "Not done, or changed" section and the results are filled at the end; **the
|
|
> Part E schedule below was written and committed BEFORE round 1.**
|
|
|
|
## Part E — the schedule (committed before round 1)
|
|
|
|
Drawn by `night-2026-09-24/tools/chaos_schedule.py`, seed **20260924** (constraints: every action at least
|
|
once; controller killed at most 3 times, ≥ 20 min apart; Start-after-unhealthy after a stop; Remove-on-held
|
|
after a hold). Guest **9202** (scratch), controller v0.269.1, drill catalog.
|
|
|
|
seed 20260924, 12 rounds
|
|
| round | action | accident |
|
|
|---|---|---|
|
|
| 1 | oom_app | power_cut |
|
|
| 2 | install | docker_restart |
|
|
| 3 | cutoff_undo_file_app | power_cut |
|
|
| 4 | failing_step | controller_killed |
|
|
| 5 | remove_held | none |
|
|
| 6 | failing_step | whole_box_backup |
|
|
| 7 | crash_loop | whole_box_backup |
|
|
| 8 | oom_app | disk_to_floor_plus_1g |
|
|
| 9 | remove_held | controller_killed |
|
|
| 10 | ladder_step | docker_restart |
|
|
| 11 | cutoff_undo_file_app | disk_to_floor_plus_1g |
|
|
| 12 | start_after_unhealthy | none |
|
|
|
|
**The app each round acts on — fixed here, not drawn** (an app's state after round N decides what N+1 can do):
|
|
|
|
| round | app | why |
|
|
|---|---|---|
|
|
| 1 oom_app | `chaosoom` (drill-only: alpine, 32M cap, a 200 MB eater in a loop) | out-of-memory storm → the box must stop it |
|
|
| 2 install | n8n at its ladder's FIRST `from` (2 steps behind), seeded through its front door | gives round 10 a step |
|
|
| 3 cutoff_undo_file_app | nextcloud (the file app), a failing step (redis 7-alpine → 7.4.0-alpine, probe 8999), the NEWEST undo copy's marker cut at verifying → the hold must name the second drive → its whole restore, then the seeds + files read back | decision 26 under a power cut |
|
|
| 4 failing_step | wishlist v0.67.1 → `:latest` with probe 8999 → undo | the undo under a controller kill |
|
|
| 5 remove_held | gokapi (held since A3: the box stopped its crash loop) — Remove with data | Remove on a held app |
|
|
| 6 failing_step | navidrome 0.64.1 → `:latest` with probe 8999 → undo | the undo under a backup run |
|
|
| 7 crash_loop | `chaoscrash` (drill-only: alpine that exits after 3 s) | the box must stop it |
|
|
| 8 oom_app | `chaosoomb` (a second eater) | the OOM stop under a full disk |
|
|
| 9 remove_held | `chaoscrash` (held since round 7) | Remove on a held app under a controller kill |
|
|
| 10 ladder_step | n8n, one step (2 → 1) | a good step under a docker restart |
|
|
| 11 cutoff_undo_file_app | nextcloud again | decision 26 with the disk 1 GB above the floor (a `disk` refusal is a valid outcome) |
|
|
| 12 start_after_unhealthy | `chaosoomb` (held since round 8) — Start → one more try → the repeat stop says support is told | decision 28's second half |
|
|
|
|
**Changed before round 1:** the accident "a whole-box backup started on 9202" is the box's own full backup run
|
|
(`POST /api/backup/run`, as on 2026-09-23), not a vzdump: demo-hp's backup storage had ~4 GB free and the host
|
|
root was at 90 %, so a vzdump of 9202 could have filled the HOST (demo-hp also had R-672's full pool today).
|
|
**How accidents are fired:** the runner arms the round and waits; the session fires each host-level accident
|
|
as its own visible command (power cut = `pct stop 9202; pct start 9202`, kill -9 of the controller PID,
|
|
`systemctl restart docker` in the guest, a `fallocate` to leave 3 GB free on `/var/lib/docker`).
|