In order to use this rule-based system for activity prediction, a set of bioactive
compounds with known activities (e.g. experimentally determined activities) have
to be collected (from the literature or an experimental laboratory). A training set is
then formed by picking up compounds from the data set suitably to train the system
to learn the structural requirement for a compound to be active. A fewer number of
compounds are also kept for testing purposes (test set). Once the training is done,
activity predictions for both training set compounds (retrofit studies) and test set
compounds are carried out. For predicting the activities of the test set compounds,
the D
À4 index values for the test set compounds are computed. If the system is
found to produce high (acceptable) percentage of correct activity predictions for
both the training set and the test set compounds along with none or very few
(acceptable) wrong activity predictions, it may be regarded as standardized for the
prediction of activity of chemical compounds for the biological endpoint for which
the system is standardized.
2.3 Training Set–Test Set Split
It is always important that a suitable training set be obtained from a data set of
bioactive compounds such that the structural characteristics of the compounds,
present in the data set, is reflected in the training set, and the learning of the (expert)
system/prediction tool is as adequate as possible for getting useful activity predictions by the method used in this purpose. In general, researchers look for the
diversity present in the structures in creating a training set from a given data set.
Presumably, some intuition or expertise of the drug designer/medicinal chemist
may be required to do that or some mathematical diversity analysis may be carried
out in obtaining a suitable training set. However, it appears that generating a large
number (e.g. 1000) of training set–test set splits (combinations) and reporting the
successful predictions of all or some (e.g. top 20, 25) of the best-predicting splits
for a given data set of bioactive compounds would be a very straightforward and
useful approach for identifying a suitable training set. Having obtained various top
performing splits, one can select a suitable split that gives high percentage of
successful predictions for both training set and test set and obtains activity prediction for the compounds present in both the sets. Although such splits have been
used [24, 25] for evaluating the performance of vertex indices and a rule-based
method for activity prediction [18, 19] considering small and large data sets, no
algorithm is available to report the activity predictions for different splits. We have
incorporated this algorithm in the program for reporting the outcome of activity
predictions for different splits so that one can consider a suitable split for further
work such as structure generation. This can be done for both quantitative data and
qualitative data (active–inactive type). It may also be noted that the computer
program can be used for the identification of training set–test set splits and activity
predictions by considering both hydrogen-filled (H-filled) and hydrogen-suppressed
(H-suppressed) molecular graphs of the compounds under consideration.
78
Md.I. H. Rizvi et al.
compounds with known activities (e.g. experimentally determined activities) have
to be collected (from the literature or an experimental laboratory). A training set is
then formed by picking up compounds from the data set suitably to train the system
to learn the structural requirement for a compound to be active. A fewer number of
compounds are also kept for testing purposes (test set). Once the training is done,
activity predictions for both training set compounds (retrofit studies) and test set
compounds are carried out. For predicting the activities of the test set compounds,
the D
À4 index values for the test set compounds are computed. If the system is
found to produce high (acceptable) percentage of correct activity predictions for
both the training set and the test set compounds along with none or very few
(acceptable) wrong activity predictions, it may be regarded as standardized for the
prediction of activity of chemical compounds for the biological endpoint for which
the system is standardized.
2.3 Training Set–Test Set Split
It is always important that a suitable training set be obtained from a data set of
bioactive compounds such that the structural characteristics of the compounds,
present in the data set, is reflected in the training set, and the learning of the (expert)
system/prediction tool is as adequate as possible for getting useful activity predictions by the method used in this purpose. In general, researchers look for the
diversity present in the structures in creating a training set from a given data set.
Presumably, some intuition or expertise of the drug designer/medicinal chemist
may be required to do that or some mathematical diversity analysis may be carried
out in obtaining a suitable training set. However, it appears that generating a large
number (e.g. 1000) of training set–test set splits (combinations) and reporting the
successful predictions of all or some (e.g. top 20, 25) of the best-predicting splits
for a given data set of bioactive compounds would be a very straightforward and
useful approach for identifying a suitable training set. Having obtained various top
performing splits, one can select a suitable split that gives high percentage of
successful predictions for both training set and test set and obtains activity prediction for the compounds present in both the sets. Although such splits have been
used [24, 25] for evaluating the performance of vertex indices and a rule-based
method for activity prediction [18, 19] considering small and large data sets, no
algorithm is available to report the activity predictions for different splits. We have
incorporated this algorithm in the program for reporting the outcome of activity
predictions for different splits so that one can consider a suitable split for further
work such as structure generation. This can be done for both quantitative data and
qualitative data (active–inactive type). It may also be noted that the computer
program can be used for the identification of training set–test set splits and activity
predictions by considering both hydrogen-filled (H-filled) and hydrogen-suppressed
(H-suppressed) molecular graphs of the compounds under consideration.
78
Md.I. H. Rizvi et al.
