7 A Review of Feature Reduction Methods …
131
the recreated output layer. Burgoon [67] used autoencoders to screen chemicals for
potential estrogenic activity by projecting the two neurons in the middle layer into a
Cartesian plane. The application of autoencoders for toxicity prediction has not been
widely reported, especially for feature extraction. This provides an opportunity for
a future area of research.
7.3.3 Linear Discriminant Analysis
Like PCA, linear discriminant analysis (LDA) [65, 68] is a linear transformation
technique commonly used for dimensionality reduction. However, LDA is supervised
since the discrimination power of the features is taken into consideration. LDA
computes an optimal transformation (projection) of the input data on to a line such
that classes are separated as clusters. The goal of the projection is to ensure maximum
class discrimination by minimizing the within-class distance while maximizing the
between-class distance [26]. A weakness of LDA is that if the distribution of a
dataset is significantly non-Gaussian, the LDA projections will not be able to preserve
any complex structure of the data [69]. Thus, the resulting features may not have
good discriminative power. Features extracted with LDA were used by Ren et al.
[70] in a stepwise forward manner from a combined pool of experimental data, and
chemical structure-based descriptors were employed for predicting aquatic toxicity
mode of action. In this work, logistic regression was shown to have a better predictive
performance than LDA using the extracted features, with a 7.3% improvement over
previously reported classification rates.
In addition to the above-mentioned nonlinear dimensionality reduction techniques, there are also spectral and manifold learning methods, such as t-distributed
Stochastic Neighbor Embedding (t-SNE) [71], multi-dimensional scaling (MDS)
[72], spectral embedding [73], and isomap [74]. Manifold learning, a class of unsupervised nonlinear algorithms, assumes that the dimensionality of a datasets is only
artificially high and thus attempts to uncover the intrinsic low dimensionality. Typically, these algorithms work by computing the similarities between points to find a
nearest-neighbor, and then an eigen problem for embedding high-dimensional points
into a lower dimensional space [75].
7.4 Miscellaneous
7.4.1 Feature Stability
It is common to use the performance of a model as the metric to evaluate the suitability
of a feature reduction algorithm. Therefore, it is an obvious choice to optimize the
selection process to obtain the best prediction power possible. However, the stability
131
the recreated output layer. Burgoon [67] used autoencoders to screen chemicals for
potential estrogenic activity by projecting the two neurons in the middle layer into a
Cartesian plane. The application of autoencoders for toxicity prediction has not been
widely reported, especially for feature extraction. This provides an opportunity for
a future area of research.
7.3.3 Linear Discriminant Analysis
Like PCA, linear discriminant analysis (LDA) [65, 68] is a linear transformation
technique commonly used for dimensionality reduction. However, LDA is supervised
since the discrimination power of the features is taken into consideration. LDA
computes an optimal transformation (projection) of the input data on to a line such
that classes are separated as clusters. The goal of the projection is to ensure maximum
class discrimination by minimizing the within-class distance while maximizing the
between-class distance [26]. A weakness of LDA is that if the distribution of a
dataset is significantly non-Gaussian, the LDA projections will not be able to preserve
any complex structure of the data [69]. Thus, the resulting features may not have
good discriminative power. Features extracted with LDA were used by Ren et al.
[70] in a stepwise forward manner from a combined pool of experimental data, and
chemical structure-based descriptors were employed for predicting aquatic toxicity
mode of action. In this work, logistic regression was shown to have a better predictive
performance than LDA using the extracted features, with a 7.3% improvement over
previously reported classification rates.
In addition to the above-mentioned nonlinear dimensionality reduction techniques, there are also spectral and manifold learning methods, such as t-distributed
Stochastic Neighbor Embedding (t-SNE) [71], multi-dimensional scaling (MDS)
[72], spectral embedding [73], and isomap [74]. Manifold learning, a class of unsupervised nonlinear algorithms, assumes that the dimensionality of a datasets is only
artificially high and thus attempts to uncover the intrinsic low dimensionality. Typically, these algorithms work by computing the similarities between points to find a
nearest-neighbor, and then an eigen problem for embedding high-dimensional points
into a lower dimensional space [75].
7.4 Miscellaneous
7.4.1 Feature Stability
It is common to use the performance of a model as the metric to evaluate the suitability
of a feature reduction algorithm. Therefore, it is an obvious choice to optimize the
selection process to obtain the best prediction power possible. However, the stability
