122
G. Idakwo et al.
for a given library, each descriptor adds a dimension to the n-dimensional chemical
space. Every molecule in the library is assigned a coordinate depending on its values
for all the descriptors. A reduction in the dimensionality of the chemical space correlates with an increasing similarity between molecules. This is important because the
underlying assumption in SAR modeling posits that molecules with similar structures
should have similar activity [20, 21]. Thus, one of the most important tasks prior to
modeling is dimension reduction focused on keeping the most important and relevant descriptors with the maximum amount of biologically meaningful information
required for predicting the desired toxicity end point. Shen et al. [13] demonstrated
the usefulness of feature selection for toxicity prediction, particularly for interpreting
the role of the features. By reducing the feature space, they were able to pinpoint
MolRef and AlogP as the most important descriptors for predicting the toxicity of
aromatic compounds.
In simple terms, dimensionality reduction is considered desirable for activity
prediction modeling for the following reasons [22]:
(i) Employing fewer descriptors means that the model can focus on important
information for establishing a relationship, thus improving prediction accuracy
and reducing overfitting (Models with many features enjoy more discriminating
power during training but are often not generalizable).
(ii) As the number of features decreases, interpretability of certain models
increases.
(iii) Computational costs reduce significantly as the complexity of many learning
algorithms is greater than linear [19, 23].
(iv) Elimination of irrelevant descriptors can help remove activity cliffs [7].
(v) Machine learning algorithms are statistical in nature; hence, they suffer from
the “curse of dimensionality”, which is common with optimization problems
as described by Bellman [24].
As the dimensionality increases, the amount of data needed to develop generalizable models increases exponentially [25, 26]. SAR data rarely have an abundance
of labeled molecules and, as such, the final model and resulting toxicity prediction
will benefit from a reduction in dimension as a smaller dimension means fewer samples will be required during training. The optimal subset of a feature space is one
which has the least number of dimensions yet offers the best learning accuracy [26].
Two techniques used to alleviate the challenges of high dimension in SAR datasets
include feature selection and feature extraction.
In this review, we discuss different methods for both feature selection and feature
extraction techniques, as well as their applications in SAR modeling. In the next two
sections, we discuss feature selection and feature extraction methods consecutively.
In the last section, we highlight important aspects that must be considered while
attempting feature space reduction, such as the stability and validation of the methods.
G. Idakwo et al.
for a given library, each descriptor adds a dimension to the n-dimensional chemical
space. Every molecule in the library is assigned a coordinate depending on its values
for all the descriptors. A reduction in the dimensionality of the chemical space correlates with an increasing similarity between molecules. This is important because the
underlying assumption in SAR modeling posits that molecules with similar structures
should have similar activity [20, 21]. Thus, one of the most important tasks prior to
modeling is dimension reduction focused on keeping the most important and relevant descriptors with the maximum amount of biologically meaningful information
required for predicting the desired toxicity end point. Shen et al. [13] demonstrated
the usefulness of feature selection for toxicity prediction, particularly for interpreting
the role of the features. By reducing the feature space, they were able to pinpoint
MolRef and AlogP as the most important descriptors for predicting the toxicity of
aromatic compounds.
In simple terms, dimensionality reduction is considered desirable for activity
prediction modeling for the following reasons [22]:
(i) Employing fewer descriptors means that the model can focus on important
information for establishing a relationship, thus improving prediction accuracy
and reducing overfitting (Models with many features enjoy more discriminating
power during training but are often not generalizable).
(ii) As the number of features decreases, interpretability of certain models
increases.
(iii) Computational costs reduce significantly as the complexity of many learning
algorithms is greater than linear [19, 23].
(iv) Elimination of irrelevant descriptors can help remove activity cliffs [7].
(v) Machine learning algorithms are statistical in nature; hence, they suffer from
the “curse of dimensionality”, which is common with optimization problems
as described by Bellman [24].
As the dimensionality increases, the amount of data needed to develop generalizable models increases exponentially [25, 26]. SAR data rarely have an abundance
of labeled molecules and, as such, the final model and resulting toxicity prediction
will benefit from a reduction in dimension as a smaller dimension means fewer samples will be required during training. The optimal subset of a feature space is one
which has the least number of dimensions yet offers the best learning accuracy [26].
Two techniques used to alleviate the challenges of high dimension in SAR datasets
include feature selection and feature extraction.
In this review, we discuss different methods for both feature selection and feature
extraction techniques, as well as their applications in SAR modeling. In the next two
sections, we discuss feature selection and feature extraction methods consecutively.
In the last section, we highlight important aspects that must be considered while
attempting feature space reduction, such as the stability and validation of the methods.
