258
F. Firouzi et al.
(a) Each row of the matrix corresponds to one single observation (it is also called
a sample or a data point or a case).
(b) Each column represents a feature (also known as an attribute or a vector).
A feature is an individual measurable property or attribute of a phenomenon
being observed. In supervised techniques, features are also called independent variables.
(c) In supervised machine learning techniques, there is one column corresponding to “label” (also known as “response,” “output,” or “dependent variable”).
2. Handling Missing Values: Possible variations include ‘NaN’, ‘NA’, ‘None’, ‘’,
‘?’ and others (see Fig. 5.11). Missing values can occur both in numerical features
and categorical features:
(a) Numerical Features: Depending on the nature of the problem, you may
consider using one of the following techniques to address missing values:
(i) Ignore those rows with a missing value.
(ii) Use the mean/median/mode value of the feature for those missing
values.
(iii) Use the previous/next values to replace missing values.
(iv) Predict the missing values; for example, curve-fitting and regression
algorithms can be used to find the missing values.
(b) Categorical Features: Note that a categorical feature is a variable that can
take a limited number of possible labeled values. For example, a color feature
with the values “Red” and “Blue.” Dealing with categorical features is tricky.
In general, you might be able to utilize one of the following techniques
depending on the project’s constraints:
(i) Ignore those rows with a missing value.
(ii) Replace the missing value with the most frequent category.
(iii) Use the previous/next values to replace missing values.
(iv) Predict the missing values.
3. Handling Categorical Values: Some machine learning techniques (e.g., decision
tree) can directly use categorical features. However, to be able to use categorical
features in other machine learning algorithms, typically, we need to use an
encoding technique to convert them to numerical features. The most wellknown encoding techniques are integer encoding and one-hot encoding (see
Fig. 5.12).
(a) Integer Encoding (Label Encoding): In this approach, a unique integer
number is assigned to each category.
(b) One-Hot Encoding: Integer encoding can result in poor models because
machine learning algorithms may consider some kind of order between
categories (e.g., 0 < 1 < 2). To tackle this issue, the one-hot encoding
method can be applied to the feature. Let us explain the idea of onehot encoding by a simple example. Suppose we have a feature with three
F. Firouzi et al.
(a) Each row of the matrix corresponds to one single observation (it is also called
a sample or a data point or a case).
(b) Each column represents a feature (also known as an attribute or a vector).
A feature is an individual measurable property or attribute of a phenomenon
being observed. In supervised techniques, features are also called independent variables.
(c) In supervised machine learning techniques, there is one column corresponding to “label” (also known as “response,” “output,” or “dependent variable”).
2. Handling Missing Values: Possible variations include ‘NaN’, ‘NA’, ‘None’, ‘’,
‘?’ and others (see Fig. 5.11). Missing values can occur both in numerical features
and categorical features:
(a) Numerical Features: Depending on the nature of the problem, you may
consider using one of the following techniques to address missing values:
(i) Ignore those rows with a missing value.
(ii) Use the mean/median/mode value of the feature for those missing
values.
(iii) Use the previous/next values to replace missing values.
(iv) Predict the missing values; for example, curve-fitting and regression
algorithms can be used to find the missing values.
(b) Categorical Features: Note that a categorical feature is a variable that can
take a limited number of possible labeled values. For example, a color feature
with the values “Red” and “Blue.” Dealing with categorical features is tricky.
In general, you might be able to utilize one of the following techniques
depending on the project’s constraints:
(i) Ignore those rows with a missing value.
(ii) Replace the missing value with the most frequent category.
(iii) Use the previous/next values to replace missing values.
(iv) Predict the missing values.
3. Handling Categorical Values: Some machine learning techniques (e.g., decision
tree) can directly use categorical features. However, to be able to use categorical
features in other machine learning algorithms, typically, we need to use an
encoding technique to convert them to numerical features. The most wellknown encoding techniques are integer encoding and one-hot encoding (see
Fig. 5.12).
(a) Integer Encoding (Label Encoding): In this approach, a unique integer
number is assigned to each category.
(b) One-Hot Encoding: Integer encoding can result in poor models because
machine learning algorithms may consider some kind of order between
categories (e.g., 0 < 1 < 2). To tackle this issue, the one-hot encoding
method can be applied to the feature. Let us explain the idea of onehot encoding by a simple example. Suppose we have a feature with three
