4 Experimental Results
In this paper, the experimental effects of the cardiovascular diseases’ diagnosis and the
following algorithms LR, AdaBoostM1, MOEFC, FURIA, GFS-LB and FH-GBML
are examined in this phase with the use of Keel and Weka tools. Meanwhile, machine
learning algorithm efficiency is derived using values like True Positive (TP), True
Negative (TN), False Positive (FP) and False Negative (FN). These measures are used
for the calculation of the sensitivity, specificity, accuracy and error rate.
Sensitivity Recall
ð
Þor True positive rate TPR
ð
Þ ¼ TP= TP þ FN
ð
Þ :
ð1Þ
Specificity ¼ TN= TN þ FP
ð
Þ
ð 2Þ
Accuracy ACC
ð
Þ ¼ TP þ TN
ð
Þ = TP þ TN þ FP þ FN
ð
Þ :
ð3Þ
Error rate ¼ FP þ FN
ð
Þ = P þ N
ð
Þ:
ð4Þ
4.1 Evaluation of Results
Setting up the Experiment under WEKA Software. In our experiment, the problem
has been transformed into binary classification with 0 presents absence and 1 presence
of heart disease. For this, Table 2 shows the results obtained by binary classification
and 10-fold cross-validation. The highest accuracy 80.20 is gained by majority voting,
while LR obtained lowest accuracy and AdaBoostM1 has the highest accuracy when
applied without ensemble.
Setting Up the Experiment under KEEL Software. Our purpose is to make a
comparison of three methods that belong to different ML techniques. In this step, we
have used a GFS-LogitBoost-C classifier with a previous pre-processing stage of
prototype selection guided by a Generational Genetic Algorithm for Feature Selection
(GGA-FS) model. We have also used a FURIA classifier with a previous preprocessing
stage of replacing missing values guided by a KNN-MV (K-Nearest Neighbor Imputation) algorithm as well as prototype feature selection guided by SSGA-Integer-knnFS (Steady-state GA with integer coding scheme for wrapper feature selection with KNN) and an FH-GBML that uses a Generational Genetic Algorithm for Feature
Table 2. Multi-class classification results by 10-fold cross-validation
Algorithm
Sensitivity Specificity Accuracy
MOEFC
79.96
75.44
79.42
LR
78.22
71.34
78.77
AdaBoostM1 80.11
75.40
80.01
Vote
84.76
74.82
80.20
304
F. Z. Abdeldjouad et al.
In this paper, the experimental effects of the cardiovascular diseases’ diagnosis and the
following algorithms LR, AdaBoostM1, MOEFC, FURIA, GFS-LB and FH-GBML
are examined in this phase with the use of Keel and Weka tools. Meanwhile, machine
learning algorithm efficiency is derived using values like True Positive (TP), True
Negative (TN), False Positive (FP) and False Negative (FN). These measures are used
for the calculation of the sensitivity, specificity, accuracy and error rate.
Sensitivity Recall
ð
Þor True positive rate TPR
ð
Þ ¼ TP= TP þ FN
ð
Þ :
ð1Þ
Specificity ¼ TN= TN þ FP
ð
Þ
ð 2Þ
Accuracy ACC
ð
Þ ¼ TP þ TN
ð
Þ = TP þ TN þ FP þ FN
ð
Þ :
ð3Þ
Error rate ¼ FP þ FN
ð
Þ = P þ N
ð
Þ:
ð4Þ
4.1 Evaluation of Results
Setting up the Experiment under WEKA Software. In our experiment, the problem
has been transformed into binary classification with 0 presents absence and 1 presence
of heart disease. For this, Table 2 shows the results obtained by binary classification
and 10-fold cross-validation. The highest accuracy 80.20 is gained by majority voting,
while LR obtained lowest accuracy and AdaBoostM1 has the highest accuracy when
applied without ensemble.
Setting Up the Experiment under KEEL Software. Our purpose is to make a
comparison of three methods that belong to different ML techniques. In this step, we
have used a GFS-LogitBoost-C classifier with a previous pre-processing stage of
prototype selection guided by a Generational Genetic Algorithm for Feature Selection
(GGA-FS) model. We have also used a FURIA classifier with a previous preprocessing
stage of replacing missing values guided by a KNN-MV (K-Nearest Neighbor Imputation) algorithm as well as prototype feature selection guided by SSGA-Integer-knnFS (Steady-state GA with integer coding scheme for wrapper feature selection with KNN) and an FH-GBML that uses a Generational Genetic Algorithm for Feature
Table 2. Multi-class classification results by 10-fold cross-validation
Algorithm
Sensitivity Specificity Accuracy
MOEFC
79.96
75.44
79.42
LR
78.22
71.34
78.77
AdaBoostM1 80.11
75.40
80.01
Vote
84.76
74.82
80.20
304
F. Z. Abdeldjouad et al.
