396
P. M. Vassiliev et al.
distributed version with limited functionality, for the prediction of individual organic compound activity.
12.3 Predictive Power of IT Microcosm
The adequacy, validity and accuracy of IT Microcosm were analyzed by testing the
training sets for the structure and pharmacological activity of structurally diverse
and structurally similar chemical compounds. It is well known that lead generation
methods are intended to represent searches for structures with high novelty, parent
compounds of new chemical classes with the desired pharmacological activity, and
new chemical entities; these methods should provide for an adequate prediction of
activity in structurally diverse compounds. Lead optimization methods are aimed at
revealing the most active compounds in a series of structurally similar compounds;
these approaches are expected to produce stable results and to predict the pharmacological activity of one class of chemical derivatives. The theoretical concepts
of IT Microcosm are of a universal nature; therefore, the technology permits an
equally successful prediction of the presence and extent of pharmacological activity in both structurally diverse and structurally similar compounds, including chiral
compounds.
12.3.1 Structurally Diverse Compounds
The training sets were constructed from available reference literature and information from the internet. The sets include structures of compounds that reliably show
the predicted activity and compounds that reliably show no activity. The total size of
databases for 34 types of activity is 10,703 structures of known drugs and biologically active substances. For 19/34 activity types, the active compounds were divided
into highly active and moderately active classes using the expert method. The selection, verification and primary processing of information in relation to structure and
activity were performed by competent experts: chemists and doctoral-level pharmacologists with considerable experience in the corresponding fields. The method of
training set construction is discussed in detail in references [105, 109].
The sizes of the training sets ranged from 30 to 1140 compounds. Depending on
the type of activity, the indices of chemical diversity varied within the following
limits: the dimensionality of the object domain description ranged from 1798 to
22,461 variables, and the mean number of unique features per compound ranged
from 14 to 149 QL descriptors.
A summary of the analysis of the adequate decision rules when predicting the
presence/absence of an activity in structurally diverse compounds is shown in
Table 12.3. A decision rule was deemed to be adequate if the values of all of the
prediction accuracy indices F 0 , F a and F n in all testing methods were at least 60 %,
which corresponds to a confidence level of p ≥ 0.9 with a training set size of N ≥ 30.
P. M. Vassiliev et al.
distributed version with limited functionality, for the prediction of individual organic compound activity.
12.3 Predictive Power of IT Microcosm
The adequacy, validity and accuracy of IT Microcosm were analyzed by testing the
training sets for the structure and pharmacological activity of structurally diverse
and structurally similar chemical compounds. It is well known that lead generation
methods are intended to represent searches for structures with high novelty, parent
compounds of new chemical classes with the desired pharmacological activity, and
new chemical entities; these methods should provide for an adequate prediction of
activity in structurally diverse compounds. Lead optimization methods are aimed at
revealing the most active compounds in a series of structurally similar compounds;
these approaches are expected to produce stable results and to predict the pharmacological activity of one class of chemical derivatives. The theoretical concepts
of IT Microcosm are of a universal nature; therefore, the technology permits an
equally successful prediction of the presence and extent of pharmacological activity in both structurally diverse and structurally similar compounds, including chiral
compounds.
12.3.1 Structurally Diverse Compounds
The training sets were constructed from available reference literature and information from the internet. The sets include structures of compounds that reliably show
the predicted activity and compounds that reliably show no activity. The total size of
databases for 34 types of activity is 10,703 structures of known drugs and biologically active substances. For 19/34 activity types, the active compounds were divided
into highly active and moderately active classes using the expert method. The selection, verification and primary processing of information in relation to structure and
activity were performed by competent experts: chemists and doctoral-level pharmacologists with considerable experience in the corresponding fields. The method of
training set construction is discussed in detail in references [105, 109].
The sizes of the training sets ranged from 30 to 1140 compounds. Depending on
the type of activity, the indices of chemical diversity varied within the following
limits: the dimensionality of the object domain description ranged from 1798 to
22,461 variables, and the mean number of unique features per compound ranged
from 14 to 149 QL descriptors.
A summary of the analysis of the adequate decision rules when predicting the
presence/absence of an activity in structurally diverse compounds is shown in
Table 12.3. A decision rule was deemed to be adequate if the values of all of the
prediction accuracy indices F 0 , F a and F n in all testing methods were at least 60 %,
which corresponds to a confidence level of p ≥ 0.9 with a training set size of N ≥ 30.
