3 Related Work
In reference [5], the authors presented the various classic and new techniques for
classifying texts: the preprocessing of documents such as tokenization, the removal of
stop words, stemming; Lemmatizing, machine learning algorithms for document
modeling; representation of document characteristics, optimal data representation;
learning based on machine learning classifiers; measuring the performance of the
classification model based on evaluation methods and performance metrics.
The authors of reference [5], presented five classifiers (SVM, NB, KNN, Decision
Tree and Decision Table) with three different versions of the database. In addition,
accuracy and scalability are calculated to evaluate and examine the advantages and
disadvantages of them for Arabic TC based on the efficient tools of machine learning
(Weka and RapidMiner).
In reference [4], the authors summarized the eminent multi-class classifiers, based
on the literature, in order to apply them to evaluate on a new benchmark dataset of
Vietnamese News (VNNews-01). In the data collect process, they are referred to more
than thirty Vietnamese online newspaper websites and grouped into twenty-five categories. They added that their work might promote the text mining research in Vietnam.
Some authors present a comparative study of three machine learning algorithm to
do the task of classifying human facial expression. Then, they analyzed the main
performance. In the experimental study process, they introduced 23 variables calculated
from the distance of facial features as the input in the classification phase. As output,
they defined seven categories, such as: angry, disgust, fear, happy, neutral, sad, and
surprise. As experimental results, they recorded 75.15% of K-Nearest Neighbor
(KNN)’s accuracy, 80% for Support Vector Machine (SVM), and 76.97% for Random
Forest algorithm. As for the result using the largest amount of data, the accuracy is
98.85% for KNN, 90% for SVM, and 98.85% for Random Forest algorithm [6].
4 Method
4.1 Data Collection
In our work, we focused on supervised learning. To do so, we collected a set of labeled
scientific articles from different scientific journals including Science direct, PubMed,
Google scholar, etc. In addition, scientific articles were classified in 4 different predefined classes related to the taxonomy of epidemiological study, including Descriptive, Analytic, Meta-Analysis and Experimental. The several categories’ definitions are
presented in Table 1.
Data collection was performed on the basis of two different approaches. The first
approach is only interested in the Abstract part. We notice that, in the field of epidemiology, the Abstract part is composed of different parts in particular Aim/Introduction/
Purpose, Methods, Results and discussion and conclusion. The first approach reveals a
first database made up of 300 abstracts per category. The second approach is to collect the
full article without omitting any section from Abstract to the references. This exercise led
us to the creation of a second extended database of 300 articles by category. Figures 1 and
2 exemplify the distribution of scientific papers according to their categories.
Machine Learning Classification Models with SPD/ED Dataset
351
In reference [5], the authors presented the various classic and new techniques for
classifying texts: the preprocessing of documents such as tokenization, the removal of
stop words, stemming; Lemmatizing, machine learning algorithms for document
modeling; representation of document characteristics, optimal data representation;
learning based on machine learning classifiers; measuring the performance of the
classification model based on evaluation methods and performance metrics.
The authors of reference [5], presented five classifiers (SVM, NB, KNN, Decision
Tree and Decision Table) with three different versions of the database. In addition,
accuracy and scalability are calculated to evaluate and examine the advantages and
disadvantages of them for Arabic TC based on the efficient tools of machine learning
(Weka and RapidMiner).
In reference [4], the authors summarized the eminent multi-class classifiers, based
on the literature, in order to apply them to evaluate on a new benchmark dataset of
Vietnamese News (VNNews-01). In the data collect process, they are referred to more
than thirty Vietnamese online newspaper websites and grouped into twenty-five categories. They added that their work might promote the text mining research in Vietnam.
Some authors present a comparative study of three machine learning algorithm to
do the task of classifying human facial expression. Then, they analyzed the main
performance. In the experimental study process, they introduced 23 variables calculated
from the distance of facial features as the input in the classification phase. As output,
they defined seven categories, such as: angry, disgust, fear, happy, neutral, sad, and
surprise. As experimental results, they recorded 75.15% of K-Nearest Neighbor
(KNN)’s accuracy, 80% for Support Vector Machine (SVM), and 76.97% for Random
Forest algorithm. As for the result using the largest amount of data, the accuracy is
98.85% for KNN, 90% for SVM, and 98.85% for Random Forest algorithm [6].
4 Method
4.1 Data Collection
In our work, we focused on supervised learning. To do so, we collected a set of labeled
scientific articles from different scientific journals including Science direct, PubMed,
Google scholar, etc. In addition, scientific articles were classified in 4 different predefined classes related to the taxonomy of epidemiological study, including Descriptive, Analytic, Meta-Analysis and Experimental. The several categories’ definitions are
presented in Table 1.
Data collection was performed on the basis of two different approaches. The first
approach is only interested in the Abstract part. We notice that, in the field of epidemiology, the Abstract part is composed of different parts in particular Aim/Introduction/
Purpose, Methods, Results and discussion and conclusion. The first approach reveals a
first database made up of 300 abstracts per category. The second approach is to collect the
full article without omitting any section from Abstract to the references. This exercise led
us to the creation of a second extended database of 300 articles by category. Figures 1 and
2 exemplify the distribution of scientific papers according to their categories.
Machine Learning Classification Models with SPD/ED Dataset
351
