272
F. Firouzi et al.
Feature selection techniques can also be classified as follows:
• Filter methods: Filter method is typically used as a preprocessing phase. Filter
methods are mostly univariate and non-iterative. Filter tries to assess the
predictive power of each feature. To do so, several statistical techniques can be
utilized to be able to compute the score (power) of the feature that demonstrates
its “level of relationship/correlation” with the output. Some famous examples
are chi-squared, F score, information gain, ANOVA, regression, and Pearson
correlation.
• Wrapper methods: Wrapper methods address the feature selection problem
similar to a search problem. These methods are called wrapper because they
wrap a machine learning (e.g., classification) inside the feature selection process.
Wrapper methods can be implemented in several ways, including:
– Forward Selection: Forward selection is an iterative method which starts by
an empty set. Then, we need to execute the machine learning model for each
feature to find the strongest one that results in the best performance. In the
next iteration, the selected feature from the previous step is combined with all
other features one by one to find the best pair of features leading to the highest
performance. We keep these two features and move to the next iteration. In
the next iteration, we try to find the best three features, and so on until the
specified number of features are selected.
– Backward Elimination: This is also an iterative approach. We start with all
features, and in each iteration, we remove/delete one of the features that does
not have a significant impact on the performance of the machine learning
model. This process is iteratively performed until a stopping criterion is
reached.
• Embedded methods: Embedded methods are implemented using those machine
learning techniques that have built-in feature selection abilities. In other words,
feature selection is integrated/embedded as part of the learning algorithm.
Regularization methods, which we discussed before, are one of the most common
approaches in this regard. These methods find the appropriate features by adding
some constraints into the optimization and cost function of the machine learning.
Lasso and elastic net regressions are examples of embedded techniques.
5.3.1.1 Chi-Square Test
The chi-square test (also called the chi-square test of independence, and chi
sounds like “Hi” but with a “K”) is used to study the significant correlation
between two categorical variables. Note that you cannot use chi-square to compare
continuous variables or a categorical with a continuous variable. Let us explain the
fundamentals of this method with a simple example. Suppose we observed 100
people to see who is interested in IoT and who is interested in Arts. Therefore,
we have one categorical feature (independent variable) which shows the gender (i.e.,
Précédent

- 278/647

Suivant