372
P. M. Vassiliev et al.
These “better equations” are always context-dependent; minimally, they depend
on the way that the compound structure is represented and on the computational
method for predictive relation. It was shown as early as 1972 that it is impossible
to choose a single regularity that provides an adequate description of a predicted
biological activity when you are faced with a number of QSAR models with comparable accuracy [18]. On the other hand, if one simultaneously uses several equations
for prediction and the calculated estimates of the activity of a certain compound
coincide, the prediction error is considerably lower. Mathematical tools for object
classification that simultaneously use several decision rules began to be developed
heavily starting in the 1990s [65]. Within the framework of in silico drug discovery,
this approach is referred to as a consensus, ensemble, or committee approach by
different authors, and methods using these approaches were first used in the 2000s
[10]. At present, the consensus approach to prediction based on the synthesis of data
obtained from several QSAR dependencies is at the peak of popularity [42].
The process of selecting the so-called significant variables is an indispensable
stage of virtually all QSAR analysis methods [9, 27, 33, 72]. Despite its apparent
clarity and attractiveness, this approach often results in the generation of artifact
dependencies, and this effect has been noted by several authors on many occasions
[18, 39]. The equations calculated in this manner, though simple and clear at first
glance, do not give a fair representation of the individual features of the predicted
compounds; they are therefore only marginally suitable for the design of intriguing,
highly active substances that show mostly nonstandard characteristics, which distinguishes them from other compounds. The transformation of a primary parameter
set into latent variables, as in PLS-regression [24] or in multilayer artificial neural
networks [40], does not settle the issue because primary variable weighting (the
determination of significance) is also performed in these cases. On the other hand,
the mathematical methods themselves often contain limitations that preclude the
simultaneous use of a large number of variables in model construction. Thus, there
should be no less than three observations for each variable in a regression analysis
[22], and no less than two observations in artificial neural network modeling [9].
QSAR approaches that consider the use of all of the available variables in making
decision rules (for example, the support vector machine (SVM) method) have only
recently appeared [91] in conjunction with kernel function use [13, 58].
The concept of significant variables presumes their independence. The methods
for restoring empirical regularities that are used in QSAR are also intended to be
only used in a Euclidean space [1]. However, all of the parameters of a compound
description are generated from the same object (the chemical structure of the compound), which is why all of the obtained variables are always inter-dependent. Furthermore, if the feature space is also nonlinear, then we must at least face the task of
adapting the existing QSAR methods to such spaces.
Taking into account all of the peculiarities of the chemical-biological universum discussed above, and in order to overcome the deficiencies of the existing
QSAR systems, a complex methodology to predict the properties of organic compounds was developed [93, 96, 102, 116] and served as a base for the development
of IT Microcosm [45, 98, 103, 105,] in the form of a software package [113]. This
Précédent

- 381/556

Suivant