132
G. Idakwo et al.
or degree of variance of feature selection methods becomes a crucial challenge when
the task at hand goes beyond optimizing prediction accuracy to include improving
interpretability. A simple scenario may be the case for using substructure-based
descriptors for SAR modeling. It is common to consider a substructure that is very
relevant for prediction as a major contributor to the activity of that molecule, implying
a potential research target. However, many feature selection algorithms tend to be
unstable and would yield a different subset if a little perturbation is applied (i.e.,
when new training samples are added or when some training samples are removed).
If every perturbation results in wide variation in the selected subset, then it is difficult
to conclude that a feature may be important to the molecule’s activity.
Kalousis et al. [76] defined the stability of a feature selection algorithm as “the
robustness of the feature subset the algorithm produces in the presence of perturbations in training sets drawn from the same generating distribution.” Essentially,
stability quantifies how different training sets affect the variation in the selected feature subset. Hence, a similarity measure is often employed to measure the stability of
feature selection algorithms. A reliable algorithm should produce the same or similar
subset for any perturbations in the training data. Alelyani et al. [77] performed experiments to investigate the causes of instability and reported that dimension, sample
size, and the distribution of the training data influenced stability. Larger sample size
translated to improved stability, while larger dimensions caused negative effects.
Thus, researchers should pay attention to the characteristics of a training dataset.
Certain algorithms are also more prone to instability than others. ReliefF-based feature selection is affected by the order of samples in a training set, while stochastic
search algorithms like GA that use random initialization parameters tend to yield
subsets that are unstable [78, 79]. Various metrics for measuring stability have been
proposed [78]. To overcome the stability challenge, it has been suggested to employ
ensemble selection algorithms based on the technicalities of the selection algorithm
in use [78, 80, 81]. Some of these algorithms include Bootstrap sampling, random
data partitioning, parameter randomization, or the combination of several of these.
Developing algorithms for feature selection that are stable and possess high predictive power is still an open and challenging area. SAR-based toxicity prediction
stands to gain a lot from such techniques that can improve speed and accuracy of
predictions for regulatory as well as lead optimization purposes.
7.4.2 Validation of Feature Selection
In selecting the optimal feature subset, it is common to evaluate the performance of
a learner based on its prediction error. A very common and overlooked mistake is
to select features using the entire dataset as a preprocessing step. While this appears
to be obviously wrong, it has been reported that many researchers, especially in the
biomedical fields, continue to make this mistake and successfully publish in topranking journals [82, 83]. If a test set is to be used to evaluate the performance of a
feature set, it must not be involved in the feature selection step as that will result in a
Précédent

- 145/416

Suivant