Files
felhom.eu/documentation/audits/DRILL-r361-2026-08-22/README.md
T
admin 1eb64bec51
gates / gates (push) Successful in 17s
R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an
invariant the code did not have, for four months - and a [DESIGN] on the db_dumps
decision INCLUDING the trap it created: a stable list lets the already-current
early return fire, so per-capture housekeeping must sit above it.

00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a
held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which
IsDownState excludes. Measured on the shipped build with the scans demonstrably
running over it. No suppression was built and no row opened.

R-383: the double-failure message names an undo copy that is not there - R-361's
own class, one surface over, observed on both 0.220.2 and 0.221.1.
R-384: an app whose database has died reads unhealthy and raises no alarm.

R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes.

Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate
blocked this push and that block is not circular, so it was satisfied rather than
bypassed - no --no-verify anywhere in this session.
2026-08-23 00:33:12 +02:00

104 lines
6.1 KiB
Markdown

# DRILL — R-361, and the two loose ends v0.220.2 left (2026-08-22 → 23)
**Subject:** `demo-hp` (Tier 0), guest 9201. Controller **v0.220.2 → v0.221.0 → v0.221.1**.
**Method:** endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex.
**Subjects FOUND, not rebuilt:** `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), plus
`privatebin`, `kimai`, `calibre-web`, `opengist`, `paperless-ngx`, `romm`.
## R-361 — the whole thing, in one comparison
The canonical dump's sha256, before and after a restore:
| app | before | after |
|---|---|---|
| `docmost` | `5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed` | **identical** |
| `bookstack` | `7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b` | **identical** |
**Before the fix, on the same box:** neither app had a canonical dump at all — only `pre-restore-*`
copies. Every safety dump had overwritten the app's real backup and then moved it away.
**The dangerous lookalike, named:** a test asserting "the pre-restore file exists" passes just as well
when the app's own backup was destroyed. Only the canonical dump's **bytes** convict.
## Part 3 — the measurement that cancelled Part 2
The runbook's reading was that a HELD app raises a dead-app banner and a customer e-mail. **It does
not.** Measured on the shipped v0.220.2, hold created 21:11:19Z, scans every 30 s running over it:
`docmost` aggregated to **`unhealthy`**, `IsDownState` is `{stopped, exited, degraded}`, and the
heartbeat read **`0 currently down`** at scans 600 and 620. **Part 2 was dropped in full** — 2.1, 2.2
and 2.3 — because 2.2 and 2.3 existed only to make 2.1 safe and complete.
**Why the reading was wrong:** its first half was right (a held app is not `StateStopped`); its second
half assumed the remaining state would be a fault. `aggregateState` checks `unhealthy > 0` **before**
the mixed-case degraded branch, and `unhealthy` is deliberately not a down state.
**The positive control, and two live attempts that failed.** `docker stop privatebin` → `stopped`,
whitelisted by design. `docker stop bookstack-db` → `degraded` for a moment, then `unhealthy`. Neither
put a stack in a lasting down state, and an absent alarm from a detector never shown working proves
nothing. The control that works is at the layer the detector lives in: `classifyRunStates` is pure,
and it raises the banner for `degraded`/`exited` while staying silent for the states measured live.
## Part 4 — the double-failure path, on what actually ships
**The trigger, chosen deliberately:** the undo copy is **lost between being written and being needed**
— a drive that goes away, a filesystem that goes read-only, an external cleanup. It is realistic, it
is one of the only two ways a rollback can fail, and the product must not assume the file it wrote
twenty seconds ago is still there.
Verified on 0.221.1: replay failed → rollback failed → **app held**; the customer's start button
refused with a reason and a route; a **full controller restart** left it stopped (`Recover()` and the
boot sweep both honoured the hold); the operator route listed it, cleared it, and after the restart
the command names, the app started.
**The customer message, verbatim, 369 bytes:**
> A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi
> állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai
> ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan:
> pre-restore-20260822T222028Z-docmost-postgres.sql
Hex in `16-part4-message.txt`. **No engine output** — R-381 holds. **But the last clause is false**,
and that is now **R-383**: it says the previous state's backup exists while naming the very file whose
absence caused the failure.
**An accident worth keeping:** the first trigger attempt removed the `.tmp` instead of the final file
— a side-effect of R-361's own fix, which now writes `<final>.tmp`. The safety dump failed and the
**fail-closed refusal fired with the app untouched**: *"a visszaállítás nem indult el"*. Not the case
being tested, but a free confirmation that no-undo-means-no-restore still holds.
## Findings
| id | what |
|---|---|
| **R-361** | CLOSED — shipped v0.221.0/.1, proven by the sha256 comparison above |
| **R-383** | NEW — the double-failure message names an undo copy that is not there |
| **R-384** | NEW — an app whose database has died reads `unhealthy` and raises no dead-app alarm |
| *held-app alarm* | **no row opened** — it does not fire; the negative is in the capability map |
## Two defects found in this session's own work
1. **A red-proof passed, twice over.** The behavioural tests inject the dump seam, so a mutation
*inside* `DumpOneTo` was invisible to them; and Part 1.3 initially had no test at all. Guards were
added at the layer each defect lives in, and both mutations then convicted.
2. **One change made another unreachable.** Excluding the undo copies from `db_dumps` made that list
stable, which let `CaptureRecoveryUnit`'s already-current early return fire — and the prune sat
after it. **Four copies on disk against a cap of three, counted on the box minutes later.** Fixed in
v0.221.1; the cap now holds at 3 on both apps, verified live.
**And the lost-update window recorded in v0.220.2 was observed, not just reasoned:** clearing a hold
and restarting in one breath lost the clear. Doing it with a verification between the steps held.
## Teardown — three layers
1. **Nothing was provisioned.** Every app was already on the box; all 20 containers healthy at the end.
`docmost` still holds its 4 pages and 1 user; both canonical dumps present.
2. **No storage added.** The undo copies are capped at 3 per app and the cap was verified live.
3. **No hub-side record created.** No customer, no appliance.
**No app is left held** — the hold created for Part 4 was cleared and the app restarted.
## Deliberately left open
R-102 (the Tier-2 unit mirror read by nothing) and R-359 (no readability check on the off-site store).
Untouched here.