--- the new disk as the box sees it ---
sda                             32G disk
sdb                            100G disk
sdc                             64G disk
sda                             32G disk
sdb                            100G disk
sdc                             64G disk
--- extend the thin pool onto it ---
  pvcreate ok
  vgextend ok
  WARNING: Set activation/thin_pool_autoextend_threshold below 100 to trigger automatic extension of thin pools before they get full.
  Logical volume pve/data successfully resized.
--- reclaim what MY failed pulls left behind (dangling layers only) ---
Total reclaimed space: 0B
--- after ---
    LV             LSize   Data% 
    data           <75.81g 15.57 
    root           <13.81g       
    swap            <3.88g       
    vm-9201-disk-0  32.00g 5.24  
    vm-9201-disk-1  70.00g 14.47 
TYPE            TOTAL     ACTIVE    SIZE      RECLAIMABLE
Images          13        5         3.966GB   2.998GB (75%)
Containers      5         5         54.78kB   0B (0%)

## Recovery from MY OWN harness damage — what was changed, and why each change is a fixture change
The box was left with a 100 %-full thin pool and nine failed installs. Two fixture changes were made,
both to the drill VM, neither to the product:

  1. **A third disk (64 G) was attached to the VM and the LVM thin pool extended onto it.**
     before: `data` 11.80 g, Data% **100.00**
     after:  `data` 75.81 g, Data% **15.57**
     The 32 G system disk was simply too small for a twelve-app household once thin-provisioning
     over-subscribed a 32 G rootfs and a 70 G data volume onto an 11.8 G pool.

  2. **The customer guest's RAM was raised 4096 -> 6144 MB.** The box has 8 GB and the guest had
     half of it; the controller's memory guard counts COMMITTED memory, so twelve apps against a
     3712 MB usable budget cannot fit however they are ordered.

`docker image prune -f` reclaimed **0 B** — the 3 GB `docker system df` calls "reclaimable" are
layers still referenced by the 13 pulled images, not dangling ones. Recorded because "75 %
reclaimable" reads like free space and is not.

**These are changes to the drill's own fixture, made in Phase 0 (setup), and they are not part of any
round's measurement.** The re-seed that follows installs the apps ONE AT A TIME, waiting for each to
reach `deployed: true`, which is what a household does and what the earlier parallel burst was not.
=== BEFORE: how is the guest rootfs mounted? ===
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16,emergency_ro)
/dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker type ext4 (rw,relatime,stripe=16,emergency_ro)
=== restart the guest so ext4 remounts clean (the pool now has room) ===
=== AFTER: mount state ===
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16)
=== PROOF: can it actually write? (not an assumption) ===
  WRITE OK
=== containers back? ===
cloudflared Up 45 seconds
felhom-controller Up 44 seconds (healthy)
filebrowser Up 43 seconds (healthy)
privatebin Up 45 seconds (healthy)
traefik Up 45 seconds
=== pool ===
  data           <75.81g 15.64 
  vm-9201-disk-0  32.00g 5.24  
  vm-9201-disk-1  70.00g 14.54 

## A filled thin pool wedges the guest READ-ONLY, and adding space does not un-wedge it
Worth writing down beyond tonight, because the second half surprised me.

When the pool hit 100 %, both of the guest's ext4 filesystems remounted themselves with
**`emergency_ro`** — visible in the mount flags, not only in dmesg:

  BEFORE:
    /dev/mapper/pve-vm--9201--disk--0 on /                type ext4 (rw,relatime,stripe=16,**emergency_ro**)
    /dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker  type ext4 (rw,relatime,stripe=16,**emergency_ro**)

The symptom this produced was NOT „no space left": every deploy was refused with
  „saving app config: writing /opt/docker/stacks/<app>/app.yaml.tmp: **read-only file system**"
which reads like a permissions problem and is nothing of the kind. Note the flags still say `rw` —
`emergency_ro` sits beside it, so a careless glance at `mount` says the filesystem is writable.

**Extending the thin pool from 11.8 G to 75.8 G did not clear it.** The space was there (Data% fell
to 15.6 %) and every write still failed. It took a guest restart for ext4 to mount clean:

  AFTER:  /dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16)   — no emergency_ro
  PROOF:  `touch /opt/docker/stacks/.rwtest` -> **WRITE OK** (a real write, not an inference from flags)
  all five containers back in ~45 s; pool 15.64 %

The proof line matters: „the flags look right" and „the filesystem accepts a write" are different
claims, and only the second one is the thing that was broken.

## CORRECTION — the drive gate was NOT stuck. I was reading a stale snapshot and a silent log.
For about four minutes I believed I had found a defect: the data drive was mounted (`df` showed
/dev/sdb, 98 G, on both the box and inside the guest) while the controller still recorded
`"disconnected": true, "stopped_stacks": ["immich","jellyfin","nextcloud","paperless-ngx"]`, and no
`storage_reconnected` event had appeared. I was about to file it.

**It was self-healing and it healed.** The controller's own DEBUG ring shows:
    [gate] drive RETURNED /mnt/felhom-drives/hdd_1 — re-attached + restarted gate-stopped apps
    Event pushed: storage_reconnected (info) — Meghajtó újra csatlakoztatva: Adatlemez
and the current state reads:
    disconnected=None   stopped_stacks=None
    hub event, Sep 16 20:44  info  storage_reconnected  „Meghajtó újra csatlakoztatva: Adatlemez"

**Why I nearly got it wrong — two instrument faults at once:**
  1. `driveGateLoop` runs on a **30-second ticker**; I read the settings file inside that window and
     treated one sample as a settled state.
  2. The gate's lines are **DEBUG**, so `docker logs` showed nothing, and I read that silence as
     „the gate never ran". An absent log line is not evidence — the debug ring had the lines all
     along (`/api/debug/logs?level=DEBUG`).
And a third, smaller one: my first attempt to read the ring parsed the JSON wrongly and reported
„total ring entries: 0" for a 29 503-byte response, which looked like confirmation of the silence.

Recorded in full because the wrong version of this paragraph would have been a filed P-row against a
mechanism that works.

## Tonight's own release, proven through its WHOLE lifecycle on this box (R-543)
    while escrow_state=pending : the bar was on every page (measured earlier in Phase 0)
    after the ceremony (escrow_state=**escrowed**), the same four pages:
        /dashboard     escrow-bar-hits=0
        /launcher      escrow-bar-hits=0
        /backups/apps  escrow-bar-hits=0
        /storage       escrow-bar-hits=0
The bar appeared while the off-site copy was paused, told the household exactly what to do, and
disappeared **for good** when they did it — on a box that installed itself from the published ISO,
with no one setting the scene for the test. „védi"/„védené" are both 0 on /backups/apps for now
because no class-A app is installed yet; that sentence is checked again once they are.

## Headroom before the rounds — so the disk-full failure cannot quietly repeat
Measured after all the repairs and re-pulls (2026-09-16 ~21:20Z):

    LVM thin pool `data`   75.81 g, **39.15 %** used      (was 11.80 g at 100 % when it wedged)
    guest /                32 G, 944 M used, 4 %
    guest /var/lib/docker  69 G, 19 G used, **29 %**      (all twelve apps' images)
    data drive /mnt/felhom-drives/hdd_1   98 G, 361 M used, 1 %
    guest memory           6144 MB total, **4373 MB available**
    drill host /mnt/hdd_1  938 G, 57 G used, 7 % (834 G free)

Every number that mattered when the box wedged now has room: the pool that filled is at two fifths,
the docker filesystem that held the corrupt layers is at under a third, and the guest has more than
4 GB of memory free with all twelve apps running. The eleven remaining rounds include image pulls
(none, as it happens, since the catalog bump was not pushed) and repeated app restarts, and none of
them can exhaust this.

This check exists because the failure it guards against was not predicted — it was discovered by the
box remounting itself read-only. Measuring the headroom is cheaper than meeting the wall again.

## The household, verified after the repairs (2026-09-16 21:21Z)
    containers: **26 running, 26 total** — nothing stopped, nothing restarting
    the box's own record: all **twelve** apps `deployed: true`
    auth-failure lines on the rebuilt DB-backed apps:
        bookstack 0 · nextcloud 0 · paperless-webserver 0 · immich-server 0
    front doors from DooPlex: **10 of 12 answer 200**
        paste · share · recipes · inventory · status · travel · paperless · cloud · media · vault
        still 404: wiki (bookstack `unhealthy` — no traefik route) and photos (immich, seconds old)

The two cures are both confirmed by the front door, not by a badge: `cloud` (nextcloud) went from a
crash loop to **200**, and `share` (gokapi) from `Restarting (1)` to **200**. Zero auth failures
anywhere is the other half — the regenerated-password fault is gone, not merely quieter.

## My SIXTH slip tonight, same shape as the others
The first attempt at this verification ran `docker …` over SSH **on the VM** instead of inside the
customer guest — I dropped the `pct exec 9201 --` wrapper. Every line came back as
    bash: docker: command not found
    bookstack   absent   auth-failure lines: 0
i.e. a tidy table of „absent" and „0" that reads exactly like „nothing is wrong". The numbers were
produced by a shell that had no docker at all.

That is now six errors of mine tonight sharing one shape: **a command whose precondition failed, still
printing a confident answer.** („all twelve deploys ACCEPTED" · „login ok (csrf 0)" · `head -12` that
hid a disk · a `pkill` that killed itself · the JSON parser that reported 0 entries for a 29 KB ring ·
and this.) The discipline that caught every one of them was the same: **ask the box directly, with a
control, before believing my own reporting layer.**

## bookstack: not the database, not the image — the app itself returns 500
Measured rather than assumed, because two wrong explanations had already been written tonight:

    healthcheck definition : ["CMD","curl","-f","http://127.0.0.1:80"], interval 30s, retries 3
    health status          : **unhealthy**, failing streak **5**
    healthcheck exit code  : **22** — curl's „the server returned an HTTP error", i.e. something IS
                             listening and answering, with a failure status
    direct request to the container, bypassing traefik: **HTTP 500**
    its log                : migrations run to completion („… DONE" for each) then
                             „[custom-init] No custom files found" / „[ls.io-init] done."
    auth-failure lines     : **0** — the database is fine

So: the container is up, the database is reachable, migrations applied, and the web application
answers 500. Traefik then refuses it a route, which is why `wiki` reads 404 at the front door — the
same „no route to an unhealthy container" shape established earlier with controls.

**The likely cause is mine, and it is checked before it is acted on:** I passed
`APP_KEY=base64:` + 32 random *alphanumerics*. Laravel expects `base64:` followed by a base64-encoded
**32-byte** key; 32 alphanumerics decode to something else entirely, and a wrong-length key is a 500
on every request. The repair script decodes the stored key and prints its true length before
redeploying, so the hypothesis is confirmed or refuted in the evidence rather than assumed — and the
fresh deploy uses `base64:$(openssl rand -base64 32)`, which is the correct shape.

## bookstack: the hypothesis was CONFIRMED by decoding, then cured
The repair script decoded the stored key before touching anything:
    prefix ok: **False**   body length: **96**
    **NOT valid base64** („Only base64 data is allowed") -> Laravel cannot use it -> HTTP 500
So the key I supplied was not merely the wrong length, it was not a usable key at all. After a clean
removal and a deploy with `base64:$(openssl rand -base64 32)`:
    t+15s  bookstack Created
    t+30s  bookstack Up 13 seconds (health: starting)
    t+45s  bookstack **Up 28 seconds (healthy)**
Seventh error of mine tonight, and the first one I confirmed with a measurement before acting on it
rather than after.

## immich: diagnosed, and deliberately NOT chased further
    Error: write **CONNECTION_CLOSED immich-postgres:5432**
    „Unable to initialize reverse geocoding: Error: write CONNECTION_CLOSED immich-postgres:5432"
    „Metadata service init failed" -> „microservices worker exited with code 1" -> restart
    **RestartCount=12**, ExitCode=0; immich-postgres, -redis and -machine-learning all `healthy`
The database connection is dropped during immich's reverse-geocoding import — the heaviest thing it
does at first start — on a guest with 6 GB of RAM running eleven other apps.

**Decision: immich is left broken and recorded, not repaired.** Checked against the drawn schedule
first: the twelve rounds act on adventurelog, gokapi, bookstack, mealie, privatebin, nextcloud,
uptime-kuma and paperless-ngx — **no round acts on immich**. Spending more of a finite night on the
one app the schedule never touches would cost rounds that were drawn before the night began. It is
named as a known pre-existing condition in every round from here on, so no later failure can be
quietly attributed to an accident when it was already broken.

**The household the rounds actually run against: 11 of 12 apps healthy and serving, immich
crash-looping.** That is the honest pre-state, written down before round 2 rather than after it.
