3.2 Data Pre-processing
In medical informatics, the diagnosis of diseases becomes quicker and easier if data is
free from missing, redundant and irrelevant data. In this study and after collection of
various records, we begin the preprocessing process. The dataset contains a total of 303
patients records, where 7 records are with some missing values. Those 7 records have
been removed from the dataset and the remaining 296 records are used in the process.
3.3 Feature Selection
Feature selection is a process of selecting a relevant feature of original features
according to definite condition. Further, feature collection algorithms intended with
different evaluation criteria mostly fall into three categories: the filter, wrapper, and
hybrid models [10]. In our work, we used only the wrapper method under Keel tool. As
per our objective, from among the 14 attributes of the dataset, two attributes pertaining
to age and sex are used to identify the personal information of the patient. The
remaining 12 attributes are considered important as they contain vital clinical records.
3.4 Feature Extraction
Feature extraction is a process that extracts a subset of new features from the original
set by means of some functional mapping. In order to meet the goal of the work, we
used PCA as one of the most widely used dimensionality reduction technique for the
medical applications under Weka tool, where the extracted information is represented
by a set of new variables, termed components or features. With PCA, we reduced the
attributes number to 6 which contributes more towards the diagnosis of the CVD.
3.5 Classification Algorithms
Under Weka tool, different predictive algorithms were chosen to build the first model,
namely: Multi-Objective Evolutionary Fuzzy Classifier (MOEFC), Logistic Regression
(LR), Adaptive Boosting (AdaBoostM1), while Genetic Fuzzy System-LogitBoost
(GFS-LB), Fuzzy Unordered Rule Induction Algorithm (FURIA) and Fuzzy Hybrid
Genetic Based Machine Learning (FH-GBML) were used under Keel tool to build the
second model. Therefore, we selected the best model in order to achieve the highest
possible performance on medical datasets and allow effective data classification.
3.6 Test Model
In the second stage, we tested our selected model only when the model is completely
trained. Its accuracy on the test data gives a realistic estimate of the model performance
on completely unseen patient data and confirms the actual predictive power of the
model.
A Hybrid Approach for Heart Disease Diagnosis and Prediction
303
Précédent

- 308/446

Suivant