2026-09-16T21:29:45Z ================ ROUND 3 : use bookstack, while: disk-95-full ================
2026-09-16T21:29:48Z --- BEFORE --- containers=26  wiki=200  status=200  paste=200
2026-09-16T21:29:48Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:29:48Z --- ACTION: use on bookstack ---
2026-09-16T21:29:48Z     wiki read 1 -> 200
2026-09-16T21:29:49Z     wiki read 2 -> 200
2026-09-16T21:29:49Z     wiki read 3 -> 200
2026-09-16T21:29:49Z --- ACCIDENT: disk-95-full (injected after the action started) ---

## MID-ACCIDENT measurements, taken while the disk was held full (21:31:39Z)
    guest /                      32G, **29G used, 1.5G free, 96%**   <- the app's view: full
    the fill file                /var/tmp/.chaosfill, **28G**
    the shared LVM thin pool     <75.81g, **39.69%**                 <- essentially unchanged
    writable?                    `touch` succeeded: **WRITE OK**     <- no read-only wedge

**The important detail, so this round is not over-read later: the POOL was never stressed.**
`fallocate` reserves filesystem blocks without writing them, so the guest's ext4 believes it is 96 %
full while the thin pool underneath never allocated the extents. The accident is therefore faithful
at the level the ACCIDENT IS ABOUT — an application meeting a full disk — and is **not** a repeat of
the pool exhaustion that wedged the box in Phase 0. My pool guard computed a cap and, as it turns
out, the cap was never the binding constraint.

Recorded because the opposite conclusion is easy to reach from the words „disk 95 % full" alone, and
this project has a standing lesson about measurements that look like something they are not.

## Household loop attribution for this round
The loop shows 2 `UNREACHABLE` lines at **21:27:57Z**, which is BEFORE this round began (21:29:45Z):
they are the tail of round 2's power-cut recovery, caught at the moment the loop was reinstalled as
a persistent unit while the box was still coming back. **They belong to round 2, not round 3**, and
round 2's household measure remains „not collected" because the loop was dead for the cut itself.
From 21:29:57Z the loop reads normally again (`paste read ok http=301`).

## The alarm the ladder PREDICTED would not fire, and did not (checked at 21:32:17Z)
With the guest's root filesystem held at **96 %**, the newest event in the hub is still
`controller_started` from 21:28. **No `disk_critical`, no `disk_warning`, no `health_degraded`.**

This was predicted before the round from `08-alarm-ladder.md` and the controller's own fill-watch:
the fill check runs **daily at 03:30**, plus once about **90 seconds after a controller start**. A
ten-minute window therefore contains no check at all unless a controller restart happens to land
inside it.

The timing here is sharper than the general rule, and it is worth the detail:
    21:26:01Z  power restored (round 2's accident)
    21:28      `controller_started` — so its opportunistic fill check ran at about **21:29:30**
    21:29:49Z  my fill began
    21:31:39Z  disk measured at 96 %, pool 39.69 %, still writable
So the single check this box would have made in the whole window ran roughly **twenty seconds before
the disk filled**, and the next one is not due until 03:30.

**What this means, stated carefully:** a disk that fills and empties between checks is invisible to
the alarm ladder. That is by design rather than a defect — the fill-watch is a daily sweep, not a
monitor — but it is the honest answer to „would the household be told?" for a transient full disk:
**no, unless the controller happens to restart while it is full.**

Re-checked at the END of the ten-minute window before this is called final (below), because an alarm
arriving late is a different finding from an alarm never arriving.

## INTERIM reading at 21:34:23Z — five minutes into the full disk (NOT the end-of-window check)
I nearly mislabelled this one. The fill began at **21:29:49Z**, so the ten-minute hold runs to
**~21:39:49Z**; at 21:34:23Z the window was only half over. Recorded as interim, and the
end-of-window check remains owed rather than quietly satisfied by this reading.

    guest /        32G, **29G used, 1.5G free, 96%**   — still full, fill file still present (28G)
    thin pool      **39.69%**                          — unchanged, as expected for fallocate
    writable?      **WRITE OK**
    containers     **26 running**
    alarms         newest event still `controller_started` 21:28 — **nothing new fired**

**What that already establishes, independent of how the window ends:** the household's twelve apps
keep running and serving with the system disk at 96 % full, and after five minutes nobody has been
told anything. The apps do not fall over; the silence is the finding.

## A third mistimed reading — and why it matters more than it looks
At 21:36:00Z I took what I labelled „the end-of-window alarm check". The fill window runs
**21:29:49Z -> ~21:39:49Z**, so that was again mid-window, the third time tonight I have estimated
the clock instead of reading it (the earlier two: „the window ends ~23:39" and the 21:34 reading).

The readings themselves are consistent and unchanged — newest event still `controller_started` at
21:28, **no `disk_critical`, no `disk_warning`, no `health_degraded`** — so no conclusion is altered.
What is wrong is the LABEL, and the label is the whole point here: „nothing has fired yet, five
minutes in" and „nothing fired in the whole ten-minute window" are different findings, and only the
second one answers „would the household be told?".

**Corrected practice for the rest of the night:** the end-of-window check is taken when the round's
own runner reports completion, not when I think the clock has moved far enough. The runner knows when
it released the fill; I was guessing.

This is the seventh item on tonight's list of my own instrument errors, and it has the same shape as
the rest — **a confident label on a measurement whose precondition was not met.**
    2026-09-16T21:29:49Z ACCIDENT=disk-95-full round=3
    2026-09-16T21:29:49Z filling the customer guest's SYSTEM disk to ~95 % — with a POOL GUARD
    2026-09-16T21:29:51Z     pool now: 39% used; guest / has 29352 MB free
    /dev/mapper/pve-vm--9201--disk--0   32G  944M   29G   4% /
    /dev/mapper/pve-vm--9201--disk--0   32G   29G  1.5G  96% /
    2026-09-16T21:29:55Z     pool after the fill: 39.69
    2026-09-16T21:29:55Z full — holding 10 minutes
    /dev/mapper/pve-vm--9201--disk--0   32G  944M   29G   4% /
    2026-09-16T21:39:56Z freed
    /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh: line 106: unexpected EOF while looking for matching `''
2026-09-16T21:39:56Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T21:39:58Z     t+609s containers=26 (before 26)
2026-09-16T21:39:58Z     STEADY after 609s
2026-09-16T21:39:59Z     front doors: wiki=200  status=200  paste=200  wiki=200
2026-09-16T21:39:59Z     household lines this round: 12  failures: 0
2026-09-16T21:39:59Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
  | Sep 16 21:20 | warning | app_oom | Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította | controller
  | Sep 16 21:19 | info | app_deployed | Alkalmazás telepítve: Immich | controller
  | Sep 16 21:19 | info | app_deploy_started | Alkalmazás telepítése elindult: Immich | controller
  | Sep 16 21:19 | info | app_removed | Alkalmazás eltávolítva: immich | controller
2026-09-16T21:40:00Z ================ END ROUND 3 ================

## ROUND 3 — the five things
**action:** `use` bookstack (three reads through its own front door)
**accident:** the guest's system disk filled to ~96 % and held for ten minutes

1. **What the customer saw.** Nothing. `wiki` answered 200 before, during and after; `status` and
   `paste` likewise. No banner, no warning, no mail — the household was never told the disk was full.
2. **What the box did by itself.** Kept all twelve apps running on a 96 %-full root filesystem and
   released the space cleanly when the fill was removed:
       21:29:55Z  32G, **29G used, 1.5G free, 96%** — held
       21:39:56Z  32G, **944M used, 29G free, 4%** — freed
   The shared thin pool never moved (**39.69 %** throughout), and the filesystem stayed writable.
3. **Time to steady.** The box never left steady — 26 containers before, **26 after**, none restarted.
   (The runner's „STEADY after 609s" is simply the elapsed window, not a recovery time.)
4. **Alarm fired / true?** **None fired.** Checked twice independently — the runner's own dump at
   21:39:59Z and a separate query at 21:40:13Z, after the fill was released. Newest event in both is
   still `controller_started` from 21:28.
5. **Should have fired and did not.** `disk_critical` is defined for ≥95 % used, and the disk sat at
   **96 % for ten minutes**. It did not fire — **but that is the ladder working as designed, not a
   miss**: the fill-watch is a daily sweep (03:30) plus one check ~90 s after a controller start, so a
   transient full disk between sweeps is invisible. Predicted from the ladder before the round, and
   confirmed. The honest answer to „would the household be told?" is **no**, unless the controller
   happens to restart while the disk is full.

**Household loop:** **12 operations in the window, 0 failures.** The household kept reading its apps
normally throughout — which is the other half of the finding: nothing broke, and nothing was said.

**Known pre-existing condition, unchanged:** immich still down (`app_oom`, its Postgres killed by the
memory limit). Not caused by this round.

## An error in MY injector, diagnosed rather than waved away
The run printed:
    inject.sh: line 106: unexpected EOF while looking for matching `''
yet `bash -n inject.sh` passes cleanly. The fault is not local syntax: my `G()` helper runs
`pct exec 9201 -- $*`, which FLATTENS its argument, so the nested `bash -c 'rm -f …; df -h …'` had its
quoting re-parsed by the remote shell and fell apart there.

**The accident itself completed regardless**, and that was verified on the box rather than taken from
the script's own account: the fill file was gone, the filesystem back to 4 %, the pool unmoved at
39.69 %, 26 containers running and a real write succeeding. So the error was in the reporting layer,
not in the experiment — the same distinction I have had to draw repeatedly tonight.

Fixed by splitting it into two plain `G` calls with no nested quotes. Round 4's accident
(`docker kill cloudflared`) and round 5's (`systemctl restart docker`) contain no nested quoting and
were unaffected, which is why round 4 was allowed to start before this was tidied.
