128
G. Idakwo et al.
the literature, the most reported is the combination of filter and wrapper methods.
Their use has been widely reported for biomedical data [35]. Hsu et al. [49] separately filtered two sets of features using F-score or information gain as the filtering
criterion. The resulting features were combined and further treated with wrappers
(Fig. 7.1d). They reported improved predictions in comparison with using filters
alone and a decreased computational time compared to using wrappers only. Reddy
et al. [53] applied a hybrid GA-based descriptor optimization technique for consistently selecting descriptor subsets that represented the whole initial descriptor space.
The weights of the selected subsets were analyzed to understand the contribution
of each feature to the prediction of HIV protease inhibitors, revealing the role of
hydrophobic interactions. This implies the interpretability of the method.
Ensemble methods represent the application of a feature selection method on
different subsets of features obtained by using subsampling strategies like bootstrapping. The resulting features from each of the subsets are aggregated using mean,
weights, or simple linear aggregation [38, 39] (Fig. 7.1e). This method is often
used to deal with the challenges of perturbation and instability experienced by most
feature selection methods. Seijo-Pardo et al. [39] provided an in-depth discussion
of ensemble methods of feature selection. Dutta et al. [54] proposed an ensemble
descriptor selection that searches for descriptor subsets using a genetic algorithm
whose objective function is a linear combination of the root-mean-square deviation
(RMSE) of all the models in the ensemble. They reported an improvement and found
that the resulting model had good performance on the PDGFR and COX-2 datasets.
A 96% reduction in noise and an improvement in performance was reported by Zhu
et al. [55], using a recursive random forest to rule out a quarter of the least important
descriptors at each iteration. This performed better than the least absolute shrinkage
and selection operator (LASSO). The authors highlighted that the difference between
the prediction performance of random forest and LASSO mainly resulted from the
use of variables selected by different strategies, rather than from differences between
the learning algorithms.
We have summarized the characteristics, strengths, and weaknesses of the five
classes of feature selection methods described above in Table 7.1 in order to assist
a user in choosing the appropriate tool based on user-specific requirements and/or
goals.
7.3 Feature Extraction
The algorithms employed for mathematical representation of molecular descriptors
and fingerprints are independent of the size of molecules, allowing the generation
of a fixed length set of descriptors for every molecule regardless of size [7]. The
generation of fixed length vectors can introduce redundant descriptors for certain
molecules within a library. An optimized feature set achieved by feature extraction
can minimize redundancy, noise, correlation between descriptors, and consequently
generate classifiers with improved prediction accuracy [20].
G. Idakwo et al.
the literature, the most reported is the combination of filter and wrapper methods.
Their use has been widely reported for biomedical data [35]. Hsu et al. [49] separately filtered two sets of features using F-score or information gain as the filtering
criterion. The resulting features were combined and further treated with wrappers
(Fig. 7.1d). They reported improved predictions in comparison with using filters
alone and a decreased computational time compared to using wrappers only. Reddy
et al. [53] applied a hybrid GA-based descriptor optimization technique for consistently selecting descriptor subsets that represented the whole initial descriptor space.
The weights of the selected subsets were analyzed to understand the contribution
of each feature to the prediction of HIV protease inhibitors, revealing the role of
hydrophobic interactions. This implies the interpretability of the method.
Ensemble methods represent the application of a feature selection method on
different subsets of features obtained by using subsampling strategies like bootstrapping. The resulting features from each of the subsets are aggregated using mean,
weights, or simple linear aggregation [38, 39] (Fig. 7.1e). This method is often
used to deal with the challenges of perturbation and instability experienced by most
feature selection methods. Seijo-Pardo et al. [39] provided an in-depth discussion
of ensemble methods of feature selection. Dutta et al. [54] proposed an ensemble
descriptor selection that searches for descriptor subsets using a genetic algorithm
whose objective function is a linear combination of the root-mean-square deviation
(RMSE) of all the models in the ensemble. They reported an improvement and found
that the resulting model had good performance on the PDGFR and COX-2 datasets.
A 96% reduction in noise and an improvement in performance was reported by Zhu
et al. [55], using a recursive random forest to rule out a quarter of the least important
descriptors at each iteration. This performed better than the least absolute shrinkage
and selection operator (LASSO). The authors highlighted that the difference between
the prediction performance of random forest and LASSO mainly resulted from the
use of variables selected by different strategies, rather than from differences between
the learning algorithms.
We have summarized the characteristics, strengths, and weaknesses of the five
classes of feature selection methods described above in Table 7.1 in order to assist
a user in choosing the appropriate tool based on user-specific requirements and/or
goals.
7.3 Feature Extraction
The algorithms employed for mathematical representation of molecular descriptors
and fingerprints are independent of the size of molecules, allowing the generation
of a fixed length set of descriptors for every molecule regardless of size [7]. The
generation of fixed length vectors can introduce redundant descriptors for certain
molecules within a library. An optimized feature set achieved by feature extraction
can minimize redundancy, noise, correlation between descriptors, and consequently
generate classifiers with improved prediction accuracy [20].
