390
P. M. Vassiliev et al.
The risk strategy is also based on selecting the method of superior accuracy out
of four methods according to the results of a leave-one-out cross-validation of the
training set, but it is performed for each one of 11 levels of QL representation separately, with consideration to the descriptor type. This strategy implements a model
of supremum consensus. There is no final voting procedure because each one of 44
prediction sets is treated as an independent information space.
Within the risk strategy, the classification metric and membership function for
the predicted compound C are calculated from the formula pairs (1, 4), (5, 7), (8,
11), (14, 16) corresponding to the selected prediction method and QL description
level.
The risk strategy permits the indirect consideration of the peculiarities associated
with possible effect mechanisms, assigns more weight to the novelty of the predicted structure than the two other strategies, and offers a broad range of extrapolation
possibilities. However, this strategy requires great caution because there is typically
a high probability of erroneous results.
12.2.5 Evaluation of Prediction Accuracy
The accuracy of the consensus prediction regularities obtained by these three strategies is evaluated according to four indicators of the recognizing and predicting
abilities of the integral decision rule, i.e., the results of self-prediction, leave-oneout cross-validation, split-half cross-validation, and double leave-one-out crossvalidation.
Self-Prediction. The activity of each one of N compounds in the training set is
calculated without any changes in the QL matrix or recalculation of the decision
rules.
Leave-One-Out Cross-Validation. Each compound in the training set is in turn
excluded from the QL matrix. In the changed set, new decision rules are calculated
for N1 compounds, and the excluded compound is used as an independent testing
object. The procedure is repeated N times.
Split-Half Cross-Validation. The working QL matrix and new decision rules are
calculated with the odd-numbered compounds in the training set, and the even compounds serve as an independent testing set. A reverse procedure is then carried out,
and the classification regularities are calculated with the even compounds, while the
odd ones serve as a testing set. The results of testing are averaged out.
Double Leave-One-Out Cross-Validation. One compound is excluded from the
training set. The procedure of leave-one-out validation is performed with the remaining N1 compounds. On the basis of the leave-one-out validation results, the
parameters of the final decision rule are calculated for the Bayesian binary classifier. This decision rule is used to classify the excluded compound, and the procedure
is repeated N times. This testing method is only used in the normal strategy.
P. M. Vassiliev et al.
The risk strategy is also based on selecting the method of superior accuracy out
of four methods according to the results of a leave-one-out cross-validation of the
training set, but it is performed for each one of 11 levels of QL representation separately, with consideration to the descriptor type. This strategy implements a model
of supremum consensus. There is no final voting procedure because each one of 44
prediction sets is treated as an independent information space.
Within the risk strategy, the classification metric and membership function for
the predicted compound C are calculated from the formula pairs (1, 4), (5, 7), (8,
11), (14, 16) corresponding to the selected prediction method and QL description
level.
The risk strategy permits the indirect consideration of the peculiarities associated
with possible effect mechanisms, assigns more weight to the novelty of the predicted structure than the two other strategies, and offers a broad range of extrapolation
possibilities. However, this strategy requires great caution because there is typically
a high probability of erroneous results.
12.2.5 Evaluation of Prediction Accuracy
The accuracy of the consensus prediction regularities obtained by these three strategies is evaluated according to four indicators of the recognizing and predicting
abilities of the integral decision rule, i.e., the results of self-prediction, leave-oneout cross-validation, split-half cross-validation, and double leave-one-out crossvalidation.
Self-Prediction. The activity of each one of N compounds in the training set is
calculated without any changes in the QL matrix or recalculation of the decision
rules.
Leave-One-Out Cross-Validation. Each compound in the training set is in turn
excluded from the QL matrix. In the changed set, new decision rules are calculated
for N1 compounds, and the excluded compound is used as an independent testing
object. The procedure is repeated N times.
Split-Half Cross-Validation. The working QL matrix and new decision rules are
calculated with the odd-numbered compounds in the training set, and the even compounds serve as an independent testing set. A reverse procedure is then carried out,
and the classification regularities are calculated with the even compounds, while the
odd ones serve as a testing set. The results of testing are averaged out.
Double Leave-One-Out Cross-Validation. One compound is excluded from the
training set. The procedure of leave-one-out validation is performed with the remaining N1 compounds. On the basis of the leave-one-out validation results, the
parameters of the final decision rule are calculated for the Bayesian binary classifier. This decision rule is used to classify the excluded compound, and the procedure
is repeated N times. This testing method is only used in the normal strategy.
