126
G. Idakwo et al.
in local optima. The use of GA to preselect descriptor subsets for SAR modeling
of artificial and real data was shown to be successful in [13] where 2D descriptors were employed to discriminate between active and inactive compounds. Particle
swarm optimization (PSO) [47] and ant colony optimization (ACO) [50] algorithms
may also be employed for heuristic subset search. For instance, it has been shown
that the ACO algorithm is a useful method for selecting descriptors for predicting
cyclooxygenase inhibitors [50].
7.2.3 Embedded
Embedded feature selection methods incorporate feature selection into the model
training process. Embedded feature learning, much like wrapper methods, takes the
potential dependencies among features into consideration while being more computationally efficient and less prone to overfitting as compared to wrappers [18, 27, 28,
41]. A common embedded feature selection algorithm is random forest. A random
forest is an ensemble of learners with a built-in mechanism for feature selection, such
as ID3 and C4.5 [28, 51]. Base learners, i.e., decision trees, look at each feature in
the feature space individually and assign importance to them based on how well they
contribute to the model attaining an optimal fit. Features with the lowest importance
are discarded, and the forest with the least number of features and highest predictive performance is selected [28] (Fig. 7.1c). Using the top 20 molecular descriptors
from the random forest predictor importance method, Newby et al. [44] obtained
more accurate decision tree classification models in most cases, compared to the use
of filter methods such as information gain, chi-square, and greedy search.
Pruning is another embedded feature selection approach that has been applied to
neural networks as well as classical learning algorithms, specifically support vector
machines (SVMs) [25]. For instance, SVM-recursive feature elimination (SVMRFE) begins with all the features and recursively removes features that do not contribute positively to the model’s predictive accuracy. To determine the optimal number of features for an RFE-based model, cross-validation is used to evaluate and
select the subset with the best performance. Hence, RFE can select the best features
for a specific learning algorithm. RFE is considered to be computationally expensive
as it traverses through all the features one after the other [41]. Weighted Kernels [49]
and regularization methods [52], like Lasso, Ridge and Elastic net, have also gained
prominence.
7.2.4 Hybrid and Ensemble Feature Selection
Hybrid methods for feature selection involve combining at least two different methods and applying them, usually in succession. Hybrid methods attempt to take advantage of the benefits of the constituent methods while leveraging their strengths. In
Précédent

- 139/416

Suivant