2.4 Compound Prioritization
The present method [15] also contains a section that can be used for prioritization
of potentially active compounds. This may be particularly useful for screening few
highly active compounds from a big database, e.g. from a set of combinatorially
generated compounds (described in the next section). This method is based on
some of the characteristics of active and inactive ranges found in the ordering of
vertex index values. Therefore, one has to look into some details of such ranges. In
doing that, two factors may be given special attention—(1) the number of vertex
index values in an active range (active range length: ARL); (2) the number of
compounds contributing to form the range (active range weight: ARW). By
applying one’s intuition too, it becomes apparent that a joint effect of these two
factors may help prioritize predicted active compounds. Therefore, we first propose
a measure, active range value (ARV), as the algebraic sum of ARL and ARW values
given by:
ARV ¼ ARL þ ARW
ð
Þ
ð 1Þ
Clearly, a range larger in length and contributed by more number of compounds in
forming the range would have higher ARV value. We define such a range of higher
ARV value a “STRONGER” range compared to those which have lower ARV
values. Now, let us assume that M out of N vertices of a molecular graph
G (representing a chemical compound) have fallen in different active ranges. If the
vertices are denoted by v 1 ; v 2 ; . . .; v M , one would get M number of ARV measures as
ARV v 1
ð Þ; ARV v 2
ð Þ; . . .; ARV v M
ð Þ. In order to get a measure of the contribution of
the vertices falling in different active ranges (i.e. contribution of activity-related
vertices), we further propose a molecular activity index (MAI) as:
MAI G
ð Þ ¼
X M
i¼1
ARV v i
ð Þ
ð2Þ
It may also be noted that while considering the length of an active range and the
number of compounds contributing to form the range, some single values that come
from both active and inactive compounds are taken into account since they are part
of the active range according to the second rule of range selection mentioned
earlier.
At the same time, there is a possibility that some of the vertex indices of
molecular graph G may fall in inactive ranges too (the second rule for activity
prediction) and that may be considered to pose a negative effect on the activity of
the compound. For the prediction purpose, therefore, vertices falling in inactive
ranges have to be considered. For doing that, let us assume that M
0 vertices of G,
viz. u 1 ; u 2 ; . . .; u M 0 fall in inactive ranges. We, thus, propose a measure, molecular
de-activity index (MDI) for G and it may be defined as:
Combinatorial Drug Discovery from Activity-Related Substructure …
79
The present method [15] also contains a section that can be used for prioritization
of potentially active compounds. This may be particularly useful for screening few
highly active compounds from a big database, e.g. from a set of combinatorially
generated compounds (described in the next section). This method is based on
some of the characteristics of active and inactive ranges found in the ordering of
vertex index values. Therefore, one has to look into some details of such ranges. In
doing that, two factors may be given special attention—(1) the number of vertex
index values in an active range (active range length: ARL); (2) the number of
compounds contributing to form the range (active range weight: ARW). By
applying one’s intuition too, it becomes apparent that a joint effect of these two
factors may help prioritize predicted active compounds. Therefore, we first propose
a measure, active range value (ARV), as the algebraic sum of ARL and ARW values
given by:
ARV ¼ ARL þ ARW
ð
Þ
ð 1Þ
Clearly, a range larger in length and contributed by more number of compounds in
forming the range would have higher ARV value. We define such a range of higher
ARV value a “STRONGER” range compared to those which have lower ARV
values. Now, let us assume that M out of N vertices of a molecular graph
G (representing a chemical compound) have fallen in different active ranges. If the
vertices are denoted by v 1 ; v 2 ; . . .; v M , one would get M number of ARV measures as
ARV v 1
ð Þ; ARV v 2
ð Þ; . . .; ARV v M
ð Þ. In order to get a measure of the contribution of
the vertices falling in different active ranges (i.e. contribution of activity-related
vertices), we further propose a molecular activity index (MAI) as:
MAI G
ð Þ ¼
X M
i¼1
ARV v i
ð Þ
ð2Þ
It may also be noted that while considering the length of an active range and the
number of compounds contributing to form the range, some single values that come
from both active and inactive compounds are taken into account since they are part
of the active range according to the second rule of range selection mentioned
earlier.
At the same time, there is a possibility that some of the vertex indices of
molecular graph G may fall in inactive ranges too (the second rule for activity
prediction) and that may be considered to pose a negative effect on the activity of
the compound. For the prediction purpose, therefore, vertices falling in inactive
ranges have to be considered. For doing that, let us assume that M
0 vertices of G,
viz. u 1 ; u 2 ; . . .; u M 0 fall in inactive ranges. We, thus, propose a measure, molecular
de-activity index (MDI) for G and it may be defined as:
Combinatorial Drug Discovery from Activity-Related Substructure …
79
