# DRILL — R-361, and the two loose ends v0.220.2 left (2026-08-22 → 23) **Subject:** `demo-hp` (Tier 0), guest 9201. Controller **v0.220.2 → v0.221.0 → v0.221.1**. **Method:** endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex. **Subjects FOUND, not rebuilt:** `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), plus `privatebin`, `kimai`, `calibre-web`, `opengist`, `paperless-ngx`, `romm`. ## R-361 — the whole thing, in one comparison The canonical dump's sha256, before and after a restore: | app | before | after | |---|---|---| | `docmost` | `5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed` | **identical** | | `bookstack` | `7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b` | **identical** | **Before the fix, on the same box:** neither app had a canonical dump at all — only `pre-restore-*` copies. Every safety dump had overwritten the app's real backup and then moved it away. **The dangerous lookalike, named:** a test asserting "the pre-restore file exists" passes just as well when the app's own backup was destroyed. Only the canonical dump's **bytes** convict. ## Part 3 — the measurement that cancelled Part 2 The runbook's reading was that a HELD app raises a dead-app banner and a customer e-mail. **It does not.** Measured on the shipped v0.220.2, hold created 21:11:19Z, scans every 30 s running over it: `docmost` aggregated to **`unhealthy`**, `IsDownState` is `{stopped, exited, degraded}`, and the heartbeat read **`0 currently down`** at scans 600 and 620. **Part 2 was dropped in full** — 2.1, 2.2 and 2.3 — because 2.2 and 2.3 existed only to make 2.1 safe and complete. **Why the reading was wrong:** its first half was right (a held app is not `StateStopped`); its second half assumed the remaining state would be a fault. `aggregateState` checks `unhealthy > 0` **before** the mixed-case degraded branch, and `unhealthy` is deliberately not a down state. **The positive control, and two live attempts that failed.** `docker stop privatebin` → `stopped`, whitelisted by design. `docker stop bookstack-db` → `degraded` for a moment, then `unhealthy`. Neither put a stack in a lasting down state, and an absent alarm from a detector never shown working proves nothing. The control that works is at the layer the detector lives in: `classifyRunStates` is pure, and it raises the banner for `degraded`/`exited` while staying silent for the states measured live. ## Part 4 — the double-failure path, on what actually ships **The trigger, chosen deliberately:** the undo copy is **lost between being written and being needed** — a drive that goes away, a filesystem that goes read-only, an external cleanup. It is realistic, it is one of the only two ways a rollback can fail, and the product must not assume the file it wrote twenty seconds ago is still there. Verified on 0.221.1: replay failed → rollback failed → **app held**; the customer's start button refused with a reason and a route; a **full controller restart** left it stopped (`Recover()` and the boot sweep both honoured the hold); the operator route listed it, cleared it, and after the restart the command names, the app started. **The customer message, verbatim, 369 bytes:** > A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi > állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai > ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: > pre-restore-20260822T222028Z-docmost-postgres.sql Hex in `16-part4-message.txt`. **No engine output** — R-381 holds. **But the last clause is false**, and that is now **R-383**: it says the previous state's backup exists while naming the very file whose absence caused the failure. **An accident worth keeping:** the first trigger attempt removed the `.tmp` instead of the final file — a side-effect of R-361's own fix, which now writes `.tmp`. The safety dump failed and the **fail-closed refusal fired with the app untouched**: *"a visszaállítás nem indult el"*. Not the case being tested, but a free confirmation that no-undo-means-no-restore still holds. ## Findings | id | what | |---|---| | **R-361** | CLOSED — shipped v0.221.0/.1, proven by the sha256 comparison above | | **R-383** | NEW — the double-failure message names an undo copy that is not there | | **R-384** | NEW — an app whose database has died reads `unhealthy` and raises no dead-app alarm | | *held-app alarm* | **no row opened** — it does not fire; the negative is in the capability map | ## Two defects found in this session's own work 1. **A red-proof passed, twice over.** The behavioural tests inject the dump seam, so a mutation *inside* `DumpOneTo` was invisible to them; and Part 1.3 initially had no test at all. Guards were added at the layer each defect lives in, and both mutations then convicted. 2. **One change made another unreachable.** Excluding the undo copies from `db_dumps` made that list stable, which let `CaptureRecoveryUnit`'s already-current early return fire — and the prune sat after it. **Four copies on disk against a cap of three, counted on the box minutes later.** Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live. **And the lost-update window recorded in v0.220.2 was observed, not just reasoned:** clearing a hold and restarting in one breath lost the clear. Doing it with a verification between the steps held. ## Teardown — three layers 1. **Nothing was provisioned.** Every app was already on the box; all 20 containers healthy at the end. `docmost` still holds its 4 pages and 1 user; both canonical dumps present. 2. **No storage added.** The undo copies are capped at 3 per app and the cap was verified live. 3. **No hub-side record created.** No customer, no appliance. **No app is left held** — the hold created for Part 4 was cleared and the app restarted. ## Deliberately left open R-102 (the Tier-2 unit mirror read by nothing) and R-359 (no readability check on the off-site store). Untouched here.