07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an invariant the code did not have, for four months - and a [DESIGN] on the db_dumps decision INCLUDING the trap it created: a stable list lets the already-current early return fire, so per-capture housekeeping must sit above it. 00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which IsDownState excludes. Measured on the shipped build with the scans demonstrably running over it. No suppression was built and no row opened. R-383: the double-failure message names an undo copy that is not there - R-361's own class, one surface over, observed on both 0.220.2 and 0.221.1. R-384: an app whose database has died reads unhealthy and raises no alarm. R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes. Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate blocked this push and that block is not circular, so it was satisfied rather than bypassed - no --no-verify anywhere in this session.
DRILL — R-361, and the two loose ends v0.220.2 left (2026-08-22 → 23)
Subject: demo-hp (Tier 0), guest 9201. Controller v0.220.2 → v0.221.0 → v0.221.1.
Method: endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex.
Subjects FOUND, not rebuilt: docmost (Postgres 16) and bookstack (MariaDB 12.3), plus
privatebin, kimai, calibre-web, opengist, paperless-ngx, romm.
R-361 — the whole thing, in one comparison
The canonical dump's sha256, before and after a restore:
| app | before | after |
|---|---|---|
docmost |
5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed |
identical |
bookstack |
7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b |
identical |
Before the fix, on the same box: neither app had a canonical dump at all — only pre-restore-*
copies. Every safety dump had overwritten the app's real backup and then moved it away.
The dangerous lookalike, named: a test asserting "the pre-restore file exists" passes just as well when the app's own backup was destroyed. Only the canonical dump's bytes convict.
Part 3 — the measurement that cancelled Part 2
The runbook's reading was that a HELD app raises a dead-app banner and a customer e-mail. It does
not. Measured on the shipped v0.220.2, hold created 21:11:19Z, scans every 30 s running over it:
docmost aggregated to unhealthy, IsDownState is {stopped, exited, degraded}, and the
heartbeat read 0 currently down at scans 600 and 620. Part 2 was dropped in full — 2.1, 2.2
and 2.3 — because 2.2 and 2.3 existed only to make 2.1 safe and complete.
Why the reading was wrong: its first half was right (a held app is not StateStopped); its second
half assumed the remaining state would be a fault. aggregateState checks unhealthy > 0 before
the mixed-case degraded branch, and unhealthy is deliberately not a down state.
The positive control, and two live attempts that failed. docker stop privatebin → stopped,
whitelisted by design. docker stop bookstack-db → degraded for a moment, then unhealthy. Neither
put a stack in a lasting down state, and an absent alarm from a detector never shown working proves
nothing. The control that works is at the layer the detector lives in: classifyRunStates is pure,
and it raises the banner for degraded/exited while staying silent for the states measured live.
Part 4 — the double-failure path, on what actually ships
The trigger, chosen deliberately: the undo copy is lost between being written and being needed — a drive that goes away, a filesystem that goes read-only, an external cleanup. It is realistic, it is one of the only two ways a rollback can fail, and the product must not assume the file it wrote twenty seconds ago is still there.
Verified on 0.221.1: replay failed → rollback failed → app held; the customer's start button
refused with a reason and a route; a full controller restart left it stopped (Recover() and the
boot sweep both honoured the hold); the operator route listed it, cleared it, and after the restart
the command names, the app started.
The customer message, verbatim, 369 bytes:
A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-20260822T222028Z-docmost-postgres.sql
Hex in 16-part4-message.txt. No engine output — R-381 holds. But the last clause is false,
and that is now R-383: it says the previous state's backup exists while naming the very file whose
absence caused the failure.
An accident worth keeping: the first trigger attempt removed the .tmp instead of the final file
— a side-effect of R-361's own fix, which now writes <final>.tmp. The safety dump failed and the
fail-closed refusal fired with the app untouched: "a visszaállítás nem indult el". Not the case
being tested, but a free confirmation that no-undo-means-no-restore still holds.
Findings
| id | what |
|---|---|
| R-361 | CLOSED — shipped v0.221.0/.1, proven by the sha256 comparison above |
| R-383 | NEW — the double-failure message names an undo copy that is not there |
| R-384 | NEW — an app whose database has died reads unhealthy and raises no dead-app alarm |
| held-app alarm | no row opened — it does not fire; the negative is in the capability map |
Two defects found in this session's own work
- A red-proof passed, twice over. The behavioural tests inject the dump seam, so a mutation
inside
DumpOneTowas invisible to them; and Part 1.3 initially had no test at all. Guards were added at the layer each defect lives in, and both mutations then convicted. - One change made another unreachable. Excluding the undo copies from
db_dumpsmade that list stable, which letCaptureRecoveryUnit's already-current early return fire — and the prune sat after it. Four copies on disk against a cap of three, counted on the box minutes later. Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live.
And the lost-update window recorded in v0.220.2 was observed, not just reasoned: clearing a hold and restarting in one breath lost the clear. Doing it with a verification between the steps held.
Teardown — three layers
- Nothing was provisioned. Every app was already on the box; all 20 containers healthy at the end.
docmoststill holds its 4 pages and 1 user; both canonical dumps present. - No storage added. The undo copies are capped at 3 per app and the cap was verified live.
- No hub-side record created. No customer, no appliance.
No app is left held — the hold created for Part 4 was cleared and the app restarted.
Deliberately left open
R-102 (the Tier-2 unit mirror read by nothing) and R-359 (no readability check on the off-site store). Untouched here.