5 Machine Learning for IoT
271
more accurate model while requiring fewer feature (input). This is accomplished
by removing the irrelevant or redundant information from the input feature set.
For example, in a supervised learning problem (either classification or regression),
although there could be a large number of available features in the input dataset, only
a subset of those features is relevant to the learning task. In this situation, incorporating all of the features may result in a risk of overfitting and high computational
cost. Feature selection algorithms allow us to overcome this challenge.
An irrelevant feature is defined as a feature that contains no useful information
regarding the problem (output variables) and is not capable of describing the
relationship in data. Irrelevant features can also adversely impact the performance
of the model. Note that there is a possibility to convert an irrelevant feature to a
relevant one by combining it with some other features. For example, to approximate
an XOR function by a machine learning algorithm, a single input is irrelevant, but as
combined with the other input, their combination can be used to produce the output
of the XOR function. This case is called feature interaction. In the case of feature
interaction, which means that there are multiple interacting features, the impact
of individual features on the output is not significant. However, they may show a
correlation to the target variable when considered in combination.
Another important problem is the presence of highly correlated features. In this
case, any individual feature may provide similar performance to the correlated feature subset. These correlated features are also called redundant features. Typically,
not much additional information from this type of features can be provided to
achieve a better machine learning model.
The main advantages and benefits of feature selection are listed below:
• Overfitting reduction: Less redundant data helps reduce noises and thus generates
a more accurate output.
• Accuracy improvement: Less irrelevant data contributes to more accurate model.
• Training time reduction: Fewer data accelerates the algorithm training process.
• Fewer attributes: Feature selection results in a simpler model that requires less
explanation.
5.3.1 Feature Selection Techniques
In general, there are two categories of feature selection techniques:
• Univariate method: Input variables are processed one by one to calculate their
relationship with the output, and then the most powerful input variables (i.e.,
those inputs with the highest correlations with output) are selected. This approach
works well in practice but may also fail because it does not take into account the
intercorrelations among input variables and the impacts of inputs on each other.
• Multivariate method: The whole group of variables is processed together.
Although this approach is more efficient than the previous one, it is more complex
and requires more computational resources.
271
more accurate model while requiring fewer feature (input). This is accomplished
by removing the irrelevant or redundant information from the input feature set.
For example, in a supervised learning problem (either classification or regression),
although there could be a large number of available features in the input dataset, only
a subset of those features is relevant to the learning task. In this situation, incorporating all of the features may result in a risk of overfitting and high computational
cost. Feature selection algorithms allow us to overcome this challenge.
An irrelevant feature is defined as a feature that contains no useful information
regarding the problem (output variables) and is not capable of describing the
relationship in data. Irrelevant features can also adversely impact the performance
of the model. Note that there is a possibility to convert an irrelevant feature to a
relevant one by combining it with some other features. For example, to approximate
an XOR function by a machine learning algorithm, a single input is irrelevant, but as
combined with the other input, their combination can be used to produce the output
of the XOR function. This case is called feature interaction. In the case of feature
interaction, which means that there are multiple interacting features, the impact
of individual features on the output is not significant. However, they may show a
correlation to the target variable when considered in combination.
Another important problem is the presence of highly correlated features. In this
case, any individual feature may provide similar performance to the correlated feature subset. These correlated features are also called redundant features. Typically,
not much additional information from this type of features can be provided to
achieve a better machine learning model.
The main advantages and benefits of feature selection are listed below:
• Overfitting reduction: Less redundant data helps reduce noises and thus generates
a more accurate output.
• Accuracy improvement: Less irrelevant data contributes to more accurate model.
• Training time reduction: Fewer data accelerates the algorithm training process.
• Fewer attributes: Feature selection results in a simpler model that requires less
explanation.
5.3.1 Feature Selection Techniques
In general, there are two categories of feature selection techniques:
• Univariate method: Input variables are processed one by one to calculate their
relationship with the output, and then the most powerful input variables (i.e.,
those inputs with the highest correlations with output) are selected. This approach
works well in practice but may also fail because it does not take into account the
intercorrelations among input variables and the impacts of inputs on each other.
• Multivariate method: The whole group of variables is processed together.
Although this approach is more efficient than the previous one, it is more complex
and requires more computational resources.
