Files
felhom.eu/documentation/audits/DRILL-night-2026-09-24.md
T
admin 3208cb2083
gates / gates (push) Successful in 24s
night 2026-09-24: Part E schedule + pairing committed before round 1
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 13:34:56 +02:00

51 lines
3.5 KiB
Markdown

# DRILL — night 2026-09-24 (run 2026-09-24 from 11:07 CEST): the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; automatic updates spiked; a chaos hour
> Written in parts. The "Not done, or changed" section and the results are filled at the end; **the
> Part E schedule below was written and committed BEFORE round 1.**
## Part E — the schedule (committed before round 1)
Drawn by `night-2026-09-24/tools/chaos_schedule.py`, seed **20260924** (constraints: every action at least
once; controller killed at most 3 times, ≥ 20 min apart; Start-after-unhealthy after a stop; Remove-on-held
after a hold). Guest **9202** (scratch), controller v0.269.1, drill catalog.
seed 20260924, 12 rounds
| round | action | accident |
|---|---|---|
| 1 | oom_app | power_cut |
| 2 | install | docker_restart |
| 3 | cutoff_undo_file_app | power_cut |
| 4 | failing_step | controller_killed |
| 5 | remove_held | none |
| 6 | failing_step | whole_box_backup |
| 7 | crash_loop | whole_box_backup |
| 8 | oom_app | disk_to_floor_plus_1g |
| 9 | remove_held | controller_killed |
| 10 | ladder_step | docker_restart |
| 11 | cutoff_undo_file_app | disk_to_floor_plus_1g |
| 12 | start_after_unhealthy | none |
**The app each round acts on — fixed here, not drawn** (an app's state after round N decides what N+1 can do):
| round | app | why |
|---|---|---|
| 1 oom_app | `chaosoom` (drill-only: alpine, 32M cap, a 200 MB eater in a loop) | out-of-memory storm → the box must stop it |
| 2 install | n8n at its ladder's FIRST `from` (2 steps behind), seeded through its front door | gives round 10 a step |
| 3 cutoff_undo_file_app | nextcloud (the file app), a failing step (redis 7-alpine → 7.4.0-alpine, probe 8999), the NEWEST undo copy's marker cut at verifying → the hold must name the second drive → its whole restore, then the seeds + files read back | decision 26 under a power cut |
| 4 failing_step | wishlist v0.67.1 → `:latest` with probe 8999 → undo | the undo under a controller kill |
| 5 remove_held | gokapi (held since A3: the box stopped its crash loop) — Remove with data | Remove on a held app |
| 6 failing_step | navidrome 0.64.1 → `:latest` with probe 8999 → undo | the undo under a backup run |
| 7 crash_loop | `chaoscrash` (drill-only: alpine that exits after 3 s) | the box must stop it |
| 8 oom_app | `chaosoomb` (a second eater) | the OOM stop under a full disk |
| 9 remove_held | `chaoscrash` (held since round 7) | Remove on a held app under a controller kill |
| 10 ladder_step | n8n, one step (2 → 1) | a good step under a docker restart |
| 11 cutoff_undo_file_app | nextcloud again | decision 26 with the disk 1 GB above the floor (a `disk` refusal is a valid outcome) |
| 12 start_after_unhealthy | `chaosoomb` (held since round 8) — Start → one more try → the repeat stop says support is told | decision 28's second half |
**Changed before round 1:** the accident "a whole-box backup started on 9202" is the box's own full backup run
(`POST /api/backup/run`, as on 2026-09-23), not a vzdump: demo-hp's backup storage had ~4 GB free and the host
root was at 90 %, so a vzdump of 9202 could have filled the HOST (demo-hp also had R-672's full pool today).
**How accidents are fired:** the runner arms the round and waits; the session fires each host-level accident
as its own visible command (power cut = `pct stop 9202; pct start 9202`, kill -9 of the controller PID,
`systemctl restart docker` in the guest, a `fallocate` to leave 3 GB free on `/var/lib/docker`).