Files
felhom.eu/documentation/audits/DRILL-night-2026-09-24.md
T
admin 3208cb2083
gates / gates (push) Successful in 24s
night 2026-09-24: Part E schedule + pairing committed before round 1
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 13:34:56 +02:00

3.5 KiB

DRILL — night 2026-09-24 (run 2026-09-24 from 11:07 CEST): the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; automatic updates spiked; a chaos hour

Written in parts. The "Not done, or changed" section and the results are filled at the end; the Part E schedule below was written and committed BEFORE round 1.

Part E — the schedule (committed before round 1)

Drawn by night-2026-09-24/tools/chaos_schedule.py, seed 20260924 (constraints: every action at least once; controller killed at most 3 times, ≥ 20 min apart; Start-after-unhealthy after a stop; Remove-on-held after a hold). Guest 9202 (scratch), controller v0.269.1, drill catalog.

seed 20260924, 12 rounds

round action accident
1 oom_app power_cut
2 install docker_restart
3 cutoff_undo_file_app power_cut
4 failing_step controller_killed
5 remove_held none
6 failing_step whole_box_backup
7 crash_loop whole_box_backup
8 oom_app disk_to_floor_plus_1g
9 remove_held controller_killed
10 ladder_step docker_restart
11 cutoff_undo_file_app disk_to_floor_plus_1g
12 start_after_unhealthy none

The app each round acts on — fixed here, not drawn (an app's state after round N decides what N+1 can do):

round app why
1 oom_app chaosoom (drill-only: alpine, 32M cap, a 200 MB eater in a loop) out-of-memory storm → the box must stop it
2 install n8n at its ladder's FIRST from (2 steps behind), seeded through its front door gives round 10 a step
3 cutoff_undo_file_app nextcloud (the file app), a failing step (redis 7-alpine → 7.4.0-alpine, probe 8999), the NEWEST undo copy's marker cut at verifying → the hold must name the second drive → its whole restore, then the seeds + files read back decision 26 under a power cut
4 failing_step wishlist v0.67.1 → :latest with probe 8999 → undo the undo under a controller kill
5 remove_held gokapi (held since A3: the box stopped its crash loop) — Remove with data Remove on a held app
6 failing_step navidrome 0.64.1 → :latest with probe 8999 → undo the undo under a backup run
7 crash_loop chaoscrash (drill-only: alpine that exits after 3 s) the box must stop it
8 oom_app chaosoomb (a second eater) the OOM stop under a full disk
9 remove_held chaoscrash (held since round 7) Remove on a held app under a controller kill
10 ladder_step n8n, one step (2 → 1) a good step under a docker restart
11 cutoff_undo_file_app nextcloud again decision 26 with the disk 1 GB above the floor (a disk refusal is a valid outcome)
12 start_after_unhealthy chaosoomb (held since round 8) — Start → one more try → the repeat stop says support is told decision 28's second half

Changed before round 1: the accident "a whole-box backup started on 9202" is the box's own full backup run (POST /api/backup/run, as on 2026-09-23), not a vzdump: demo-hp's backup storage had ~4 GB free and the host root was at 90 %, so a vzdump of 9202 could have filled the HOST (demo-hp also had R-672's full pool today). How accidents are fired: the runner arms the round and waits; the session fires each host-level accident as its own visible command (power cut = pct stop 9202; pct start 9202, kill -9 of the controller PID, systemctl restart docker in the guest, a fallocate to leave 3 GB free on /var/lib/docker).