130
G. Idakwo et al.
intrinsically perform a nonlinear mapping of the input space to a feature space followed by performing linear PCA in this feature space. KPCA generated vectors have
been used to train SVM models [59], and it was shown that KPCA is efficient over a
wide range of virtual screening dataset inputs using MACCS and ECFP fingerprints.
It was also observed that the KPCA embedding largely depended on the properties
of the underlying representation as its performance on the ECFP fingerprint varied
with the hashing employed.
7.3.2 Autoencoder
Autoencoders [63, 64] are unsupervised neural networks with an odd number of
hidden layers that can be applied for nonlinear feature extraction. They employ
the backpropagation algorithm to try to create a set of output values which are
equal to the input by minimizing the error between the output and the input layer.
The network architecture can be designed such that the middle layer is smaller,
i.e., has fewer nodes than the input and output layers (Fig. 7.2). In that case, the
network is forced to learn a compact representation (embedding) of the input data
[65]. In an early work, Hinton et al. [17] demonstrated that autoencoders generated
embeddings of images that were used to reconstruct images. A major drawback
of autoencoders is that physical meaning for theoretical insight will be lost. They
are also complex to train because they typically require a large amount of training
data and a search through many possible hyperparameter values. Blaschke et al.
[66] employed generative autoencoders to design new molecules in silico based on
Fig. 7.2 An autoencoder indicating the reduced dimension in the middle layer
Précédent

- 143/416

Suivant