7 A Review of Feature Reduction Methods …
133
selection bias that will yield overly optimistic performance estimates. This is because
the features used will have an unfair advantage since they were chosen based on all of
the samples. As a result, the model would have gained insight into the features which
are more important in the test set. This challenge is more common with wrapper
methods [83].
In many practical cases of SAR-based toxicity modeling, there are rarely a large
number of compounds across the different end points to be predicted. This makes it
difficult to set aside a reasonable batch of data for evaluation purposes. Methods such
as cross-validation and bootstrap sampling can be used to avoid sampling bias [34,
82, 83]. Cross-validation techniques like leave-one-out cross-validation (LOOCV)
and the k-fold method were suggested. Feature selection is to be done in the inner
loop of the cross-validation procedure; hence, the algorithm takes the following form
for a k-fold technique [82]:
(i) Randomly shuffle the data set.
(ii) Randomly split the dataset into K folds.
(iii) For each fold k = 1, 2,…, K.
a. Perform feature selection to obtain an optimal subset with good univariate
correlation with the desired end point using all the data except the kth fold.
b. Use the selected features and build a multivariate model with all data except
the kth fold.
c. Perform an evaluation using the kth fold.
(iv) Aggregate the performance across all K folds to get an unbiased evaluation.
7.5 Summary
QSAR-based predictive toxicity modeling methods are faced with input spaces of
thousands of features. To improve the ability of a learner to find a generalizable
relationship between molecular descriptors and the toxicity end point of interest, it
is expedient to provide the learning algorithm with the minimum number of descriptors while ensuring that the resulting model is interpretable and computationally
inexpensive to build. The relevance of a descriptor is assessed by its ability to discriminate between classes in qualitative classification or its correlation to a scalar in
quantitative prediction.
In this review, we have discussed different feature selection and extraction methods applicable to SAR-based toxicity modeling. The strengths and weaknesses of
each method are highlighted. The choice of which to use should largely depend on
the available dataset, and we suggest beginning a new task with a few baseline performance values from a number of methods since no single approach is universally
superior. Where the importance of descriptors is sought, feature selection methods
such as filter, wrapper, embedded or their combinations (hybrid and ensemble) may
apply. Feature extraction methods transform the features into a lower dimension while
Précédent

- 146/416

Suivant