377
12 Consensus Drug Design Using IT Microcosm
of real-world spaces representing real-world objects where there are no rough or
postulative assumptions about the linearity or orthogonality of these spaces; the
spaces are not divided into “informative” or “noninformative” subspaces. An object
description, as a totality of the values of all of the variables in a mega-dimensional
space, is context-independent because it makes no arbitrary assumptions about the
properties of the space nor does it include the results of any preliminary analyses in
any manner, which provides for a high adequacy of this description.
Complementarity of Decision Rules in the Computer Prediction of the
Chemical Properties of a Compound. As a rule, making a prediction algorithm for
a high-complexity dynamic chemical system using a single mathematical method
is based on a series of assumptions that do not prove to be true. When comparing prediction relationships that are comparable in their accuracy, one cannot unambiguously determine which of them are more adequate. These contradictions in
predicting the activity of the same compound can be resolved by applying several
approaches that differ in their mathematical content as much as possible. For each
level of description and each parameter group, several methods are used to calculate
several classification rules including all of the variables of the given local space. An
integral multimodel decision rule is developed by generalizing the obtained spectrum of intermediate prediction estimates. This suggests that the final estimate that
is calculated in this way is a reliable indicator of the actual activity of the predicted
compound; it also gives an adequate idea of the specifics of its behavior in the given
biological system.
The generally accepted classification system [1] that is applied to the prediction
of biological activity consists of the following stages:
1. developing a training set from active and inactive compounds;
2. building a primary space that describes the compound structure according to one
of the classical QSAR paradigms on the basis of a simple, well-known model;
3. selecting “significant” variables using a procedure that is correlated with the
model; and
4. calculating several activity-structure relationships using the “significant” variables set and choosing the most appropriate one.
This scheme of decision rule development is context-dependent; thus, “the most
appropriate” QSAR regularities built from the “significant” variables are mostly
artificial dependencies.
IT Microcosm produces context-independent classification rules; it has the following features [105] that distinguish it from the classical scheme above:
1. When the training set is developed, no primary space for the description is built;
rather, working models of generalized patterns of active/inactive compounds
in extra-large dimensions are developed without defining the “significant” and
“insignificant” variables. The parametric description of models and generalized
patterns is context-independent from the methods of data analysis because it is
extremely redundant and does not suggest any procedure for the detection of
“informative” features.
Précédent

- 386/556

Suivant