Machine Learning Classification Models
with SPD/ED Dataset: Comparative Study
of Abstract Versus Full Article Approach
Mayara Khadhraoui
1,2(&) , Hatem Bellaaj
2(&) ,
Mehdi Ben Ammar
3(&) , Habib Hamam
4(&) ,
and Mohamed Jmaiel
1,2(&)
1 ENIS, ReDCAD Laboratory, University of Sfax, B.P. 1173, Sfax, Tunisia
khadhraouimayara@gmail.com
2 Digital Research Center of Sfax, 3021 Sfax, Tunisia
{hatem.bellaaj,mohamed.jmaiel}@redcad.org
3 Solutions Galore Inc., Moncton, NB, Canada
mehdi.benammar@gmail.com
4 Faculty of Engineering, Moncton University, Moncton, NB, Canada
habib.hamam@gmail.com
Abstract. In response to the researchers need in the bio-medical domain, we
opted for automating the bibliographic research stage. In this context, several
classification models of supervised machine learning are used. Namely the
SVM, Random Forest, Decision Tree, KNN, and Gradient Boosting. In this
paper, we conduct a comparative study between experimental results of full
article classification and abstract classification approaches. Furthermore, we
evaluate our results by using evaluation metrics such as accuracy, precision,
recall and F1-score. We observe that the abstract approach outperforms the full
article approach in terms of learning time and efficiency.
Keywords: Text classification Á Data mining Á Supervised machine learning Á
Medical informatics Á Public health
1 Introduction
In the vast field of artificial intelligence, machine learning is called upon to play a
central role allowing machines to learn automatically in the context of scientific
research. In fact, the field of scientific research seems to be a challenging task and can
generate difficulties for researchers. In this paper, we are interested in the epidemiological research domain. Here is a list of some of today’s challenges; 1) Research in
medicine requires an efficient working methodology to better attain pertinent results,
confirm/affirm or complete a hypothesis or theory, evaluate a procedure or a program,
minimize bias, etc., 2) Medical researchers face challenges in epidemiological research,
such as the choice of population, sample size, time of study, and target knowledge
base; the selection of reference subjects; the required budget, data collection, 3)
Developing coherent epidemiological research requires the integration of knowledge
and skill, 4) Based on the results of [2], one of the major challenges of this specified
© The Author(s) 2020
M. Jmaiel et al. (Eds.): ICOST 2020, LNCS 12157, pp. 348–356, 2020.
https://doi.org/10.1007/978-3-030-51517-1_31
with SPD/ED Dataset: Comparative Study
of Abstract Versus Full Article Approach
Mayara Khadhraoui
1,2(&) , Hatem Bellaaj
2(&) ,
Mehdi Ben Ammar
3(&) , Habib Hamam
4(&) ,
and Mohamed Jmaiel
1,2(&)
1 ENIS, ReDCAD Laboratory, University of Sfax, B.P. 1173, Sfax, Tunisia
khadhraouimayara@gmail.com
2 Digital Research Center of Sfax, 3021 Sfax, Tunisia
{hatem.bellaaj,mohamed.jmaiel}@redcad.org
3 Solutions Galore Inc., Moncton, NB, Canada
mehdi.benammar@gmail.com
4 Faculty of Engineering, Moncton University, Moncton, NB, Canada
habib.hamam@gmail.com
Abstract. In response to the researchers need in the bio-medical domain, we
opted for automating the bibliographic research stage. In this context, several
classification models of supervised machine learning are used. Namely the
SVM, Random Forest, Decision Tree, KNN, and Gradient Boosting. In this
paper, we conduct a comparative study between experimental results of full
article classification and abstract classification approaches. Furthermore, we
evaluate our results by using evaluation metrics such as accuracy, precision,
recall and F1-score. We observe that the abstract approach outperforms the full
article approach in terms of learning time and efficiency.
Keywords: Text classification Á Data mining Á Supervised machine learning Á
Medical informatics Á Public health
1 Introduction
In the vast field of artificial intelligence, machine learning is called upon to play a
central role allowing machines to learn automatically in the context of scientific
research. In fact, the field of scientific research seems to be a challenging task and can
generate difficulties for researchers. In this paper, we are interested in the epidemiological research domain. Here is a list of some of today’s challenges; 1) Research in
medicine requires an efficient working methodology to better attain pertinent results,
confirm/affirm or complete a hypothesis or theory, evaluate a procedure or a program,
minimize bias, etc., 2) Medical researchers face challenges in epidemiological research,
such as the choice of population, sample size, time of study, and target knowledge
base; the selection of reference subjects; the required budget, data collection, 3)
Developing coherent epidemiological research requires the integration of knowledge
and skill, 4) Based on the results of [2], one of the major challenges of this specified
© The Author(s) 2020
M. Jmaiel et al. (Eds.): ICOST 2020, LNCS 12157, pp. 348–356, 2020.
https://doi.org/10.1007/978-3-030-51517-1_31
