CHAOS NIGHT: household repaired, headroom checked, round 2 armed
gates / gates (push) Successful in 21s

The five apps my disk-full burst broke are removed with their data and
redeployed cleanly, one at a time. Two distinct faults of mine, with different
cures, and separating them is what made either fixable:
  - image layers written while the pool was full -> "invalid ELF header",
    exit 127; cured by dropping the image so compose re-pulls
  - my re-seed's FRESH database passwords over volumes initialised with the
    first set -> Postgres auth_failed / MariaDB "Access denied"; cured only by
    removing the app with its data and deploying once

gokapi proves they are different: a new image left it Restarting(1), a new
database made it healthy.

Headroom measured so the disk-full failure cannot quietly repeat: pool 39% of
75.8G, docker filesystem 29%, data drive 1%, 4.3G guest RAM free.

Round 2's script hardened before it runs unattended: its "steady" test now
compares against the count it measured itself in the same round, instead of a
hardcoded 24 that the rebuilt household might never reach.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 23:20:27 +02:00
parent a1a57ea6d5
commit c3e1986aaf
4 changed files with 76 additions and 0 deletions
@@ -173,6 +173,39 @@ container; identical byte-for-byte to a no-such-host control) and a crash-loopin
layers corrupted while my thin pool stood at 100 %; „invalid ELF header", repaired by a re-pull).
Detail in `evidence-chaos-night-2026-09-17/round-1-notes.txt`.
### Phase 0 postscript — repairing my own damage, and three conclusions I had to retract
Nine of the first ten deploys failed because I fired twelve at once onto a thin pool carved from a
32 GB disk, and the pool hit 100 %. Repairing that took the rest of Phase 0 and produced **two
distinct faults of mine, with different cures**, which only separating them made fixable:
| fault | symptom | cure |
|---|---|---|
| image layers written while the pool was full | `php: … libxml2.so.2: **invalid ELF header**`, exit 127 crash loop | drop the image, let compose pull it again |
| my re-seed generated **fresh database passwords** over volumes already initialised with the first set | Postgres `auth_failed`, MariaDB „Access denied for user … (using password: YES)" | remove the app **with its data**, deploy once with one consistent secret set |
A fresh image did not fix gokapi and a fresh database did — that is the evidence the two faults are
different things rather than one.
**Three conclusions I wrote and then had to retract, each corrected where it stood:**
1. „the five 404s were my mistimed sweep" — wrong for four of them. A **negative control** (a
no-such-host request) returned the identical 404 of 19 bytes, and a positive control returned
200/1200 bytes: traefik simply has **no route to an unhealthy container**.
2. „nextcloud is repaired" — wrong. The re-pull fixed the crash, and the app still could not reach its
database. The container reported **`healthy`** throughout, because the image's own healthcheck asks
whether Apache answers, not whether the application works.
3. „all the broken apps are corrupt layers" — wrong. bookstack logged a clean startup, gokapi logged
nothing at all, immich showed a Postgres auth failure.
I also nearly filed a defect against the drive gate, which was working and logging at DEBUG while I
read a settings snapshot inside its 30-second tick. **An absent log line is not evidence.**
None of this is a product defect and none of it is filed as one. What the product did throughout was
correct and legible: it refused to route to unhealthy containers, `app_start_failed` named the app
that was down, `backup_run_failures` said „1 of 12 … nextcloud", the memory guard refused the
eleventh and twelfth installs with both numbers quoted, and `storage_fill_critical` fired at 100 %.
### Rounds 2-12
PENDING
@@ -124,3 +124,22 @@ The bar appeared while the off-site copy was paused, told the household exactly
disappeared **for good** when they did it — on a box that installed itself from the published ISO,
with no one setting the scene for the test. „védi"/„védené" are both 0 on /backups/apps for now
because no class-A app is installed yet; that sentence is checked again once they are.
## Headroom before the rounds — so the disk-full failure cannot quietly repeat
Measured after all the repairs and re-pulls (2026-09-16 ~21:20Z):
LVM thin pool `data` 75.81 g, **39.15 %** used (was 11.80 g at 100 % when it wedged)
guest / 32 G, 944 M used, 4 %
guest /var/lib/docker 69 G, 19 G used, **29 %** (all twelve apps' images)
data drive /mnt/felhom-drives/hdd_1 98 G, 361 M used, 1 %
guest memory 6144 MB total, **4373 MB available**
drill host /mnt/hdd_1 938 G, 57 G used, 7 % (834 G free)
Every number that mattered when the box wedged now has room: the pool that filled is at two fifths,
the docker filesystem that held the corrupt layers is at under a third, and the guest has more than
4 GB of memory free with all twelve apps running. The eleven remaining rounds include image pulls
(none, as it happens, since the catalog bump was not pushed) and repeated app restarts, and none of
them can exhaust this.
This check exists because the failure it guards against was not predicted — it was discovered by the
box remounting itself read-only. Measuring the headroom is cheaper than meeting the wall again.
@@ -180,3 +180,19 @@ spacing a 23:45 start still reaches round 12 by about 04:20.
The queue behind the current job: **nextcloud, immich and bookstack** get the same clean
remove-and-redeploy. They are done one at a time, never two repair jobs at once on the same box —
concurrent stop/start on one controller is what produced half of tonight's confusion already.
## The diagnosis is confirmed by the cure
gokapi had already had a fresh image pulled and was STILL `Restarting (1)`. After a clean removal
(with data) and a fresh deploy with one consistent set of secrets:
gokapi **Up 47 seconds (healthy)**
paperless-ngx INSTALLED at 2026-09-16T21:17:15Z
paperless-postgres Up 19 seconds (healthy)
paperless-redis Up 19 seconds (healthy)
A new image did not fix it and a new database did — which is the difference between the corrupt-layer
fault (nextcloud's exit 127) and the regenerated-password fault (everything else). Two faults, both
mine, with different cures, and it took separating them to fix either.
Household loop through all of this: unit active, **4 sampled apps since the classifier fix, 0
failures** — the box kept serving its other apps while five were being torn down and rebuilt.
@@ -64,3 +64,11 @@ was written. That is the documented rebuild behaviour, surfaced honestly with th
2026-09-16T21:16:35Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:16:55Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:17:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:17:35Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:17:55Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:18:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:18:35Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:18:55Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:19:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:19:35Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:19:56Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":