5 Machine Learning for IoT
259
Machine Type
Machine A
Machine B
Machine A
Machine C
Machine A
Machine Type
1
2
1
3
1
A
B
C
1
0
0
0
1
0
1
0
0
0
0
1
1
0
0
Categorical Feature
Integer Encoding
One-hot Encoding
Fig. 5.12 A simple example of integer encoding and one-hot encoding
different categories (i.e., Machine A, Machine B, and Machine C). In 0nehot encoding approach, we generate one boolean column for each category.
Therefore, in our example, we have three columns and only one of these
columns could have the value 1 for each observation (see Fig. 5.12).
4. Normalizing Data: In this step, we rescale all the features/attributes to have a
common scale (usually into the range 0 to 1). For example, suppose we have
a dataset containing two features, temperature and voltage. In this dataset, the
temperature ranges from −30 to 120, whereas the voltage ranges from 0 to 12.
Therefore, the temperature is about ten times larger than the voltage. In this case,
we may consider rescaling those features into the range 0–1.
5. Standardizing Data: In this step, we rescale features so that they have a mean
value of 0 and a standard deviation of 1. In this case, assume that our data has a
Gaussian (bell curve) distribution.
6. Splitting the Dataset into Training and Test Set: In every machine learning
project, there are two well-known techniques to split the dataset for training
and evaluating a model, namely, hold-out and cross-validation (k-fold crossvalidation).
(a) Hold-out: In this technique, the original dataset is divided into three
groups: training dataset, validation dataset, and test (hold-out) dataset (see
Fig. 5.13). The training dataset is the majority of data and typically contains
60–90% of the original set. The validation set is a subset of the training data,
and it is used to evaluate the performance of the model during the training
phase. In other words, the data of the validation test are not used for the
training of the model, but instead, can be used to tune the hyperparameters
of the model. The validation set is very useful to tackle some important
problems in machine learning such as overfitting (it will be discussed in the
next sections). The test dataset is just utilized to evaluate how well the trained
259
Machine Type
Machine A
Machine B
Machine A
Machine C
Machine A
Machine Type
1
2
1
3
1
A
B
C
1
0
0
0
1
0
1
0
0
0
0
1
1
0
0
Categorical Feature
Integer Encoding
One-hot Encoding
Fig. 5.12 A simple example of integer encoding and one-hot encoding
different categories (i.e., Machine A, Machine B, and Machine C). In 0nehot encoding approach, we generate one boolean column for each category.
Therefore, in our example, we have three columns and only one of these
columns could have the value 1 for each observation (see Fig. 5.12).
4. Normalizing Data: In this step, we rescale all the features/attributes to have a
common scale (usually into the range 0 to 1). For example, suppose we have
a dataset containing two features, temperature and voltage. In this dataset, the
temperature ranges from −30 to 120, whereas the voltage ranges from 0 to 12.
Therefore, the temperature is about ten times larger than the voltage. In this case,
we may consider rescaling those features into the range 0–1.
5. Standardizing Data: In this step, we rescale features so that they have a mean
value of 0 and a standard deviation of 1. In this case, assume that our data has a
Gaussian (bell curve) distribution.
6. Splitting the Dataset into Training and Test Set: In every machine learning
project, there are two well-known techniques to split the dataset for training
and evaluating a model, namely, hold-out and cross-validation (k-fold crossvalidation).
(a) Hold-out: In this technique, the original dataset is divided into three
groups: training dataset, validation dataset, and test (hold-out) dataset (see
Fig. 5.13). The training dataset is the majority of data and typically contains
60–90% of the original set. The validation set is a subset of the training data,
and it is used to evaluate the performance of the model during the training
phase. In other words, the data of the validation test are not used for the
training of the model, but instead, can be used to tune the hyperparameters
of the model. The validation set is very useful to tackle some important
problems in machine learning such as overfitting (it will be discussed in the
next sections). The test dataset is just utilized to evaluate how well the trained
