Files
felhom.eu/documentation/audits/DRILL-r361-2026-08-22
admin 1eb64bec51
gates / gates (push) Successful in 17s
R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an
invariant the code did not have, for four months - and a [DESIGN] on the db_dumps
decision INCLUDING the trap it created: a stable list lets the already-current
early return fire, so per-capture housekeeping must sit above it.

00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a
held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which
IsDownState excludes. Measured on the shipped build with the scans demonstrably
running over it. No suppression was built and no row opened.

R-383: the double-failure message names an undo copy that is not there - R-361's
own class, one surface over, observed on both 0.220.2 and 0.221.1.
R-384: an app whose database has died reads unhealthy and raises no alarm.

R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes.

Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate
blocked this push and that block is not circular, so it was satisfied rather than
bypassed - no --no-verify anywhere in this session.
2026-08-23 00:33:12 +02:00
..

DRILL — R-361, and the two loose ends v0.220.2 left (2026-08-22 → 23)

Subject: demo-hp (Tier 0), guest 9201. Controller v0.220.2 → v0.221.0 → v0.221.1. Method: endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex. Subjects FOUND, not rebuilt: docmost (Postgres 16) and bookstack (MariaDB 12.3), plus privatebin, kimai, calibre-web, opengist, paperless-ngx, romm.

R-361 — the whole thing, in one comparison

The canonical dump's sha256, before and after a restore:

app before after
docmost 5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed identical
bookstack 7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b identical

Before the fix, on the same box: neither app had a canonical dump at all — only pre-restore-* copies. Every safety dump had overwritten the app's real backup and then moved it away.

The dangerous lookalike, named: a test asserting "the pre-restore file exists" passes just as well when the app's own backup was destroyed. Only the canonical dump's bytes convict.

Part 3 — the measurement that cancelled Part 2

The runbook's reading was that a HELD app raises a dead-app banner and a customer e-mail. It does not. Measured on the shipped v0.220.2, hold created 21:11:19Z, scans every 30 s running over it: docmost aggregated to unhealthy, IsDownState is {stopped, exited, degraded}, and the heartbeat read 0 currently down at scans 600 and 620. Part 2 was dropped in full — 2.1, 2.2 and 2.3 — because 2.2 and 2.3 existed only to make 2.1 safe and complete.

Why the reading was wrong: its first half was right (a held app is not StateStopped); its second half assumed the remaining state would be a fault. aggregateState checks unhealthy > 0 before the mixed-case degraded branch, and unhealthy is deliberately not a down state.

The positive control, and two live attempts that failed. docker stop privatebin → stopped, whitelisted by design. docker stop bookstack-db → degraded for a moment, then unhealthy. Neither put a stack in a lasting down state, and an absent alarm from a detector never shown working proves nothing. The control that works is at the layer the detector lives in: classifyRunStates is pure, and it raises the banner for degraded/exited while staying silent for the states measured live.

Part 4 — the double-failure path, on what actually ships

The trigger, chosen deliberately: the undo copy is lost between being written and being needed — a drive that goes away, a filesystem that goes read-only, an external cleanup. It is realistic, it is one of the only two ways a rollback can fail, and the product must not assume the file it wrote twenty seconds ago is still there.

Verified on 0.221.1: replay failed → rollback failed → app held; the customer's start button refused with a reason and a route; a full controller restart left it stopped (Recover() and the boot sweep both honoured the hold); the operator route listed it, cleared it, and after the restart the command names, the app started.

The customer message, verbatim, 369 bytes:

A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-20260822T222028Z-docmost-postgres.sql

Hex in 16-part4-message.txt. No engine output — R-381 holds. But the last clause is false, and that is now R-383: it says the previous state's backup exists while naming the very file whose absence caused the failure.

An accident worth keeping: the first trigger attempt removed the .tmp instead of the final file — a side-effect of R-361's own fix, which now writes <final>.tmp. The safety dump failed and the fail-closed refusal fired with the app untouched: "a visszaállítás nem indult el". Not the case being tested, but a free confirmation that no-undo-means-no-restore still holds.

Findings

id what
R-361 CLOSED — shipped v0.221.0/.1, proven by the sha256 comparison above
R-383 NEW — the double-failure message names an undo copy that is not there
R-384 NEW — an app whose database has died reads unhealthy and raises no dead-app alarm
held-app alarm no row opened — it does not fire; the negative is in the capability map

Two defects found in this session's own work

  1. A red-proof passed, twice over. The behavioural tests inject the dump seam, so a mutation inside DumpOneTo was invisible to them; and Part 1.3 initially had no test at all. Guards were added at the layer each defect lives in, and both mutations then convicted.
  2. One change made another unreachable. Excluding the undo copies from db_dumps made that list stable, which let CaptureRecoveryUnit's already-current early return fire — and the prune sat after it. Four copies on disk against a cap of three, counted on the box minutes later. Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live.

And the lost-update window recorded in v0.220.2 was observed, not just reasoned: clearing a hold and restarting in one breath lost the clear. Doing it with a verification between the steps held.

Teardown — three layers

  1. Nothing was provisioned. Every app was already on the box; all 20 containers healthy at the end. docmost still holds its 4 pages and 1 user; both canonical dumps present.
  2. No storage added. The undo copies are capped at 3 per app and the cap was verified live.
  3. No hub-side record created. No customer, no appliance.

No app is left held — the hold created for Part 4 was cleared and the app restarted.

Deliberately left open

R-102 (the Tier-2 unit mirror read by nothing) and R-359 (no readability check on the off-site store). Untouched here.