386
P. M. Vassiliev et al.
The nearest neighbor method is a local geometric nonparametric method; it is
useful for the calculation of a piecewise linear separating function [32]. The basic
algorithm was described in reference [34] and adapted to extra-large dimensional
spaces as described in [105, 109]. The classification metric is the distance in the mdimensional feature space between the object of interest and the object of class k.
Compound C belongs to the class where its nearest neighbor is located.
The squared Euclidean distance in the space of QL descriptors of i-type from the
predicted compound C to each compound H l in the training set is
(12.8)
where c ij are coordinates of compound C for descriptor ij; and
h ijl are coordinates of compound H l for descriptor ij.
Therefore, in the space of i-type descriptors, the distance from compound C to
its nearest neighbor of class k is
(12.9)
Compound C is defined as active for a descriptor of i-type, if
(12.10)
otherwise it is classified as inactive.
In the nearest neighbor method, the membership function of compound C belonging to activity class k for descriptor of i-type is
(12.11)
The local distribution method is one combination method using the geometric local nonparametric method in parallel to a probabilistic central parametric method
for decision rule construction. The algorithm was first described in [109] and later
modified as described in [105]. Two metrics serve as classification metrics: the
similarity coefficient of the features of the object to be predicted and class k objects
in m-dimensional space, and the probability that the object of interest belongs to
the subclass of similar objects in class k. Compound C is assigned to the class with
the greatest local probability that the compound belongs to the structurally similar
subclass.
The similarity coefficient of the predicted compound C and each compound H l
in the training set in the space of i-type descriptors is
2
2
1
,
1,...
( )
, ,
(
)
i
l
d
i
l
ij
ijl
j
j C H
D
N
H
h
l
c
=
∈ ∪
=
=
-
∑
D
D H
ik
l
N
H k
i
l
l
2
1
2
= =
∈
min{ ( )}
D
D
ia
in
2
2
≤
;
2
2
2
.
(
) 1
2
ik
i
ia
in
D
Fb C k
D D
β
β
+
∈ = -
+
+ ⋅
P. M. Vassiliev et al.
The nearest neighbor method is a local geometric nonparametric method; it is
useful for the calculation of a piecewise linear separating function [32]. The basic
algorithm was described in reference [34] and adapted to extra-large dimensional
spaces as described in [105, 109]. The classification metric is the distance in the mdimensional feature space between the object of interest and the object of class k.
Compound C belongs to the class where its nearest neighbor is located.
The squared Euclidean distance in the space of QL descriptors of i-type from the
predicted compound C to each compound H l in the training set is
(12.8)
where c ij are coordinates of compound C for descriptor ij; and
h ijl are coordinates of compound H l for descriptor ij.
Therefore, in the space of i-type descriptors, the distance from compound C to
its nearest neighbor of class k is
(12.9)
Compound C is defined as active for a descriptor of i-type, if
(12.10)
otherwise it is classified as inactive.
In the nearest neighbor method, the membership function of compound C belonging to activity class k for descriptor of i-type is
(12.11)
The local distribution method is one combination method using the geometric local nonparametric method in parallel to a probabilistic central parametric method
for decision rule construction. The algorithm was first described in [109] and later
modified as described in [105]. Two metrics serve as classification metrics: the
similarity coefficient of the features of the object to be predicted and class k objects
in m-dimensional space, and the probability that the object of interest belongs to
the subclass of similar objects in class k. Compound C is assigned to the class with
the greatest local probability that the compound belongs to the structurally similar
subclass.
The similarity coefficient of the predicted compound C and each compound H l
in the training set in the space of i-type descriptors is
2
2
1
,
1,...
( )
, ,
(
)
i
l
d
i
l
ij
ijl
j
j C H
D
N
H
h
l
c
=
∈ ∪
=
=
-
∑
D
D H
ik
l
N
H k
i
l
l
2
1
2
= =
∈
min{ ( )}
D
D
ia
in
2
2
≤
;
2
2
2
.
(
) 1
2
ik
i
ia
in
D
Fb C k
D D
β
β
+
∈ = -
+
+ ⋅
