training. This section, therefore, contains the results of the performance of training
set identification and activity prediction. We have also reported here the results of
structure generation for some of the NA series compounds in the same way as it has
been done for the barbiturate series. For identifying a suitable training set–test set
combination for the purpose of identifying a suitable training set that can produce
high percentage of successful activity predictions, the program generates 1000 such
combinations. The program has the option of getting the output on the basis of best
test set predictions (starting from no misprediction) and best training set predictions.
It has been observed that there are combinations where no mispredictions are found
for the training set although there are 2 or more mispredictions for the test sets. On
the other hand, there are combinations where there is one misprediction each for
both the training set and the test set and it seems quite reasonable to consider such a
balanced combination for activity prediction of newly generated compounds. We
have reported here the activity predictions and MPS values of such a balanced
outcome in Table 5 for the nucleoside analogues (NA) considered for the present
study given in Table 4. The structural information of the compounds has been taken
from the corresponding MOL files.
Activity Prediction for Nucleoside Analogues
For carrying out activity prediction and prioritization studies for NA series of
compounds, we have used training set–test set split algorithm and the prediction
results for split that has given one misprediction each for the training set and the test
set are reported here.
It can be seen that for this NA series, activities of 92.86% (13 out of 14) of the
training set compounds and 83.33% (5 out of 6) of the test set compounds have
been predicted correctly, compound no. 10 of the training set and compound no.
13 of the test set being the lone mispredictions in each case. It is interesting to note
that in both the cases the inactive compounds have been predicted to be active
which may be regarded as an important factor in situations where a drug designer
Table 4 A series of 20 nucleoside analogues
a considered for the present study
Compound Name
Compound Name
1.
3′-deoxyadenosine
11.
2′-deoxyinosine
2.
2′-deoxycytidine
12.
2′,3′-dideoxythymidine
3.
2′-deoxyadenosine
13.
2′,3′-dideoxyuridine
4.
2′,3′-dideoxyadenosine
14.
2′,3′,5′-trideoxyadenosine
5.
2′,3′-dideoxycytidine
15.
3′-amino-2′,3′-dideoxycytidine
6.
3′-fluoro-2′,3′-dideoxythymidine
16.
3′-amino-2′,3′-dideoxyadenosine
7.
3′-azido-2′,3′-dideoxythymidine
17.
2′-deoxyguanosine
8.
2′,3′-dideoxyinosine
18.
3′-azido-2′,3′-dideoxyadenosine
9.
2′,3′-dideoxyguanosine
19.
3′-azido-2′,3′-dideoxycytidine
10.
5′-iodo-2′-deoxycytidine
20.
3′-azido-3′-deoxyadenosine
a Data were taken from Raychaudhury et al. [20, 21]
Combinatorial Drug Discovery from Activity-Related Substructure …
95
set identification and activity prediction. We have also reported here the results of
structure generation for some of the NA series compounds in the same way as it has
been done for the barbiturate series. For identifying a suitable training set–test set
combination for the purpose of identifying a suitable training set that can produce
high percentage of successful activity predictions, the program generates 1000 such
combinations. The program has the option of getting the output on the basis of best
test set predictions (starting from no misprediction) and best training set predictions.
It has been observed that there are combinations where no mispredictions are found
for the training set although there are 2 or more mispredictions for the test sets. On
the other hand, there are combinations where there is one misprediction each for
both the training set and the test set and it seems quite reasonable to consider such a
balanced combination for activity prediction of newly generated compounds. We
have reported here the activity predictions and MPS values of such a balanced
outcome in Table 5 for the nucleoside analogues (NA) considered for the present
study given in Table 4. The structural information of the compounds has been taken
from the corresponding MOL files.
Activity Prediction for Nucleoside Analogues
For carrying out activity prediction and prioritization studies for NA series of
compounds, we have used training set–test set split algorithm and the prediction
results for split that has given one misprediction each for the training set and the test
set are reported here.
It can be seen that for this NA series, activities of 92.86% (13 out of 14) of the
training set compounds and 83.33% (5 out of 6) of the test set compounds have
been predicted correctly, compound no. 10 of the training set and compound no.
13 of the test set being the lone mispredictions in each case. It is interesting to note
that in both the cases the inactive compounds have been predicted to be active
which may be regarded as an important factor in situations where a drug designer
Table 4 A series of 20 nucleoside analogues
a considered for the present study
Compound Name
Compound Name
1.
3′-deoxyadenosine
11.
2′-deoxyinosine
2.
2′-deoxycytidine
12.
2′,3′-dideoxythymidine
3.
2′-deoxyadenosine
13.
2′,3′-dideoxyuridine
4.
2′,3′-dideoxyadenosine
14.
2′,3′,5′-trideoxyadenosine
5.
2′,3′-dideoxycytidine
15.
3′-amino-2′,3′-dideoxycytidine
6.
3′-fluoro-2′,3′-dideoxythymidine
16.
3′-amino-2′,3′-dideoxyadenosine
7.
3′-azido-2′,3′-dideoxythymidine
17.
2′-deoxyguanosine
8.
2′,3′-dideoxyinosine
18.
3′-azido-2′,3′-dideoxyadenosine
9.
2′,3′-dideoxyguanosine
19.
3′-azido-2′,3′-dideoxycytidine
10.
5′-iodo-2′-deoxycytidine
20.
3′-azido-3′-deoxyadenosine
a Data were taken from Raychaudhury et al. [20, 21]
Combinatorial Drug Discovery from Activity-Related Substructure …
95
