7 A Review of Feature Reduction Methods …
129
A mathematical description of feature extraction is as follows: Considering a
descriptor space, x ∈ R
n , find a mapping y = f (x) to obtain transformed feature
vector y, where y ∈ R
k and k < n. The vector y should preserve the majority of
molecular information in R
n . The goal is to achieve a reduction in dimension without
negatively impacting the prediction performance. An optimal mapping, y = f (x),
is one that minimizes the prediction error.
Feature extraction transforms the initial feature space to a new, lower dimension
feature space by combining the features in the original space. As a result, it is difficult
to associate the new features with the old. Further analysis, such as feature importance
explanation, becomes very difficult as there is no physical meaning for the newly
mapped features that are obtained from feature extraction. Here we discuss some
commonly used feature extraction techniques.
7.3.1 Principal Component Analysis
Principal component analysis (PCA) is a multivariate, nonparametric method
employed for dimensionality reduction [56, 57]. It works by performing a linear
combination of the features, also referred to as the principal components, to achieve
the maximum variance. At its core, PCA is centered on determining the eigenvectors of the input data’s covariance matrix. This linear transformation can minimize
redundancy and reduce the number of features, which increases the information in
the resulting features. Each of the resulting features, called principal components,
is a combination of several original features. These principal components are also
highly uncorrelated because the first principal component accounts for as much of
the variability in the data as possible, and each succeeding component accounts
for as much of the remaining variability as possible [26]. A detailed discussion on
the different applications of PCA in SAR modeling was provided in [57]. Klepsch
et al. [58] applied PCA to a curated P-glycoprotein inhibitors data set of 1608 compounds, where the first two principal components were reported to explain 71.7%
of the variance in the dataset. This approach was applied to classification and an
analysis into the effect of the initial descriptors on these two components showed
that hydrophobic information, such as the number of aromatic bonds and the partition coefficient, was the major contributor to the principal components. According
to [59], 2-aryl-1,3,4-Thiadiazole derivatives were classified into distinct clusters of
active or inactive molecules when PCA was performed instead of using all of the
descriptors calculated.
Considering that principal components are combinations of the original features,
all the original features are still available within the components. This is useful for
interpretation of models because knowing the original features that contribute to a
component can reveal the types of features that are closely related. A key challenge
with PCA is that it is unable to handle data with complicated structures that may not
be represented in a linear subspace [60]. Kernel PCA (KPCA) [61, 62] was designed
to serve as the nonlinear form of PCA. KPCA is based on kernel functions that
129
A mathematical description of feature extraction is as follows: Considering a
descriptor space, x ∈ R
n , find a mapping y = f (x) to obtain transformed feature
vector y, where y ∈ R
k and k < n. The vector y should preserve the majority of
molecular information in R
n . The goal is to achieve a reduction in dimension without
negatively impacting the prediction performance. An optimal mapping, y = f (x),
is one that minimizes the prediction error.
Feature extraction transforms the initial feature space to a new, lower dimension
feature space by combining the features in the original space. As a result, it is difficult
to associate the new features with the old. Further analysis, such as feature importance
explanation, becomes very difficult as there is no physical meaning for the newly
mapped features that are obtained from feature extraction. Here we discuss some
commonly used feature extraction techniques.
7.3.1 Principal Component Analysis
Principal component analysis (PCA) is a multivariate, nonparametric method
employed for dimensionality reduction [56, 57]. It works by performing a linear
combination of the features, also referred to as the principal components, to achieve
the maximum variance. At its core, PCA is centered on determining the eigenvectors of the input data’s covariance matrix. This linear transformation can minimize
redundancy and reduce the number of features, which increases the information in
the resulting features. Each of the resulting features, called principal components,
is a combination of several original features. These principal components are also
highly uncorrelated because the first principal component accounts for as much of
the variability in the data as possible, and each succeeding component accounts
for as much of the remaining variability as possible [26]. A detailed discussion on
the different applications of PCA in SAR modeling was provided in [57]. Klepsch
et al. [58] applied PCA to a curated P-glycoprotein inhibitors data set of 1608 compounds, where the first two principal components were reported to explain 71.7%
of the variance in the dataset. This approach was applied to classification and an
analysis into the effect of the initial descriptors on these two components showed
that hydrophobic information, such as the number of aromatic bonds and the partition coefficient, was the major contributor to the principal components. According
to [59], 2-aryl-1,3,4-Thiadiazole derivatives were classified into distinct clusters of
active or inactive molecules when PCA was performed instead of using all of the
descriptors calculated.
Considering that principal components are combinations of the original features,
all the original features are still available within the components. This is useful for
interpretation of models because knowing the original features that contribute to a
component can reveal the types of features that are closely related. A key challenge
with PCA is that it is unable to handle data with complicated structures that may not
be represented in a linear subspace [60]. Kernel PCA (KPCA) [61, 62] was designed
to serve as the nonlinear form of PCA. KPCA is based on kernel functions that
