domain is the literature review task which should be exhaustive. In this paper, we focus
on the last aforementioned challenge. To overcome this problem, machine learning
techniques and algorithms are recommended. In our work, we are concerned with
several machine learning methods namely the Support Vector Machine (SVM),
K Nearest Neighbor (KNN), Gradient Boosting (GB), Random Forest (RF), Decision
Tree (DT), Multi-Nominal Naive Bayes (MNB) and Logistic Regression (LR).
The originality of our work lies in the creation of a new public database SPD/ED in
the biomedical domain based on title, abstract, keywords and full scientific papers. Our
database is a collection of several scientific papers classified into four different categories according to the taxonomy of the epidemiological studies (Analytic, descriptive,
Meta-Analysis and Others) [5]. Based on the aforementioned machine learning
methods, we will conduct a comparative study between the text classification task
based on the abstract versus the full article.
The paper is organized as follows. Section 2 explains basic concepts of machine
learning methods. Related work is discussed in Sect. 3. In Sect. 4, we present our
method. Section 5 discusses the experimental results. Section 6 concludes the paper
and outlines areas for future research.
2 Machine Learning
In our work, we are interested on text classification, defined as the process of associating a category (or class) with free text, based on the information it contains, is an
important element of information retrieval systems. In our work, we deal with the text
classification challenge and accuracy problem. In fact, the main challenge consists in,
for each new entry, being able to determine to which category this entry belongs.
Associating a class with free text is a costly and difficult task, therefore the automation
of this task has become a challenge for the scientific community. To help the scientific
community the task of Text classification is assisted by the machine learning.
2.1 Different Types of Approaches
There are several Machine learning methods: supervised, unsupervised, reinforcement
and semi-supervised learning. In our work, we are interested on the supervised
learning.
2.2 Machine Learning Algorithms
The objective of machine learning is to recognize among data structures that are
difficult to detect manually. From these structures, we seek to classify new textual data.
In our work, we focus on the classification of scientific papers in the epidemiological
domain based on the taxonomy of the epidemiological studies.
2.2.1 Decision Tree (DT)
Decision trees are classification rules which base their decisions on a series of tests
associated with a set of attributes. These tests are organized in a tree structure. The
Machine Learning Classification Models with SPD/ED Dataset
349
on the last aforementioned challenge. To overcome this problem, machine learning
techniques and algorithms are recommended. In our work, we are concerned with
several machine learning methods namely the Support Vector Machine (SVM),
K Nearest Neighbor (KNN), Gradient Boosting (GB), Random Forest (RF), Decision
Tree (DT), Multi-Nominal Naive Bayes (MNB) and Logistic Regression (LR).
The originality of our work lies in the creation of a new public database SPD/ED in
the biomedical domain based on title, abstract, keywords and full scientific papers. Our
database is a collection of several scientific papers classified into four different categories according to the taxonomy of the epidemiological studies (Analytic, descriptive,
Meta-Analysis and Others) [5]. Based on the aforementioned machine learning
methods, we will conduct a comparative study between the text classification task
based on the abstract versus the full article.
The paper is organized as follows. Section 2 explains basic concepts of machine
learning methods. Related work is discussed in Sect. 3. In Sect. 4, we present our
method. Section 5 discusses the experimental results. Section 6 concludes the paper
and outlines areas for future research.
2 Machine Learning
In our work, we are interested on text classification, defined as the process of associating a category (or class) with free text, based on the information it contains, is an
important element of information retrieval systems. In our work, we deal with the text
classification challenge and accuracy problem. In fact, the main challenge consists in,
for each new entry, being able to determine to which category this entry belongs.
Associating a class with free text is a costly and difficult task, therefore the automation
of this task has become a challenge for the scientific community. To help the scientific
community the task of Text classification is assisted by the machine learning.
2.1 Different Types of Approaches
There are several Machine learning methods: supervised, unsupervised, reinforcement
and semi-supervised learning. In our work, we are interested on the supervised
learning.
2.2 Machine Learning Algorithms
The objective of machine learning is to recognize among data structures that are
difficult to detect manually. From these structures, we seek to classify new textual data.
In our work, we focus on the classification of scientific papers in the epidemiological
domain based on the taxonomy of the epidemiological studies.
2.2.1 Decision Tree (DT)
Decision trees are classification rules which base their decisions on a series of tests
associated with a set of attributes. These tests are organized in a tree structure. The
Machine Learning Classification Models with SPD/ED Dataset
349
