276
F. Firouzi et al.
Note that MI is zero, when there is not any correlation between X and Y, meaning
that they are statistically independent variables. The maximum of MI happens when
Y is completely dependent on X.
5.3.2 Feature Extraction
Feature extraction is slightly different from feature selection. While the latter selects
a subset of the original input variables, the former generates some new variables
(features) from the original ones. Principal component analysis (PCA), linear
discriminant analysis (LDA), and spectral transformations (such as Fourier and
wavelet transforms) are among the most well-known feature extraction techniques.
Feature extraction improves the performance of the machine learning model by
incorporating new relevant features.
5.4 Classification
Classification is another form of supervised machine learning, in which the output
(target of the dependent parameter) labels are categorical. In other words, the output
is classified into different groups. The goal of classification is to train and create a
model (called classifier) based on the training dataset, which is then able to classify
(i.e., predict the label or class) unseen data. Before presenting the details of different
classification models and approaches, it is crucial to know about performance
metrics, which is the first step for constructing classification models.
5.4.1 Measuring Performance for Classification Problems
Performance metrics are one of the key aspects of every machine learning projects,
which determine how the performance of the algorithm is measured and is compared
with other algorithms. Throughout this section, we explain each performance metric
via a simple classification problem. We would like to predict if a given person has
cancer (true/positive class) or not (false/negative class).
5.4.1.1 Confusion Matrix (Error Matrix)
The confusion matrix is a table to present the performance of a classification model.
For a binary classifier, the confusion matrix is a 2 × 2 matrix similar to Fig. 5.28.
In our example (prediction of cancer), the confusion matrix has two dimensions,
namely, the actual dimension and predicted dimension. The actual dimension has
F. Firouzi et al.
Note that MI is zero, when there is not any correlation between X and Y, meaning
that they are statistically independent variables. The maximum of MI happens when
Y is completely dependent on X.
5.3.2 Feature Extraction
Feature extraction is slightly different from feature selection. While the latter selects
a subset of the original input variables, the former generates some new variables
(features) from the original ones. Principal component analysis (PCA), linear
discriminant analysis (LDA), and spectral transformations (such as Fourier and
wavelet transforms) are among the most well-known feature extraction techniques.
Feature extraction improves the performance of the machine learning model by
incorporating new relevant features.
5.4 Classification
Classification is another form of supervised machine learning, in which the output
(target of the dependent parameter) labels are categorical. In other words, the output
is classified into different groups. The goal of classification is to train and create a
model (called classifier) based on the training dataset, which is then able to classify
(i.e., predict the label or class) unseen data. Before presenting the details of different
classification models and approaches, it is crucial to know about performance
metrics, which is the first step for constructing classification models.
5.4.1 Measuring Performance for Classification Problems
Performance metrics are one of the key aspects of every machine learning projects,
which determine how the performance of the algorithm is measured and is compared
with other algorithms. Throughout this section, we explain each performance metric
via a simple classification problem. We would like to predict if a given person has
cancer (true/positive class) or not (false/negative class).
5.4.1.1 Confusion Matrix (Error Matrix)
The confusion matrix is a table to present the performance of a classification model.
For a binary classifier, the confusion matrix is a 2 × 2 matrix similar to Fig. 5.28.
In our example (prediction of cancer), the confusion matrix has two dimensions,
namely, the actual dimension and predicted dimension. The actual dimension has
