5 Results
This section summarizes the experimental results obtained using our Dataset SPD/ED
in two version, the extended and the closed one. In fact, we used several machine
learning classifiers, aforementioned detailed. In each approach, the dataset is divided
into train and test dataset with the ratio of 25% of test data and 75% of training data.
Both train and test data need to be preprocessed and converted into feature vectors.
As depicted in Table 2, we present a comparative table of two approaches proposed
at the level of this paper.
From Table 2, we can see that each algorithm shows high performance. In another
side, we can see, from Table 2, that using the abstract (Aims, Methods, Results and
Conclusion) only from the whole paper, is more fruitful and efficient in terms of
accuracy. Then, we conclude that SVM, RF and GB are more accurate than the others
used algorithms.
Based on the experimental results, we first concluded that the best scores obtained
are justified by the relevant choice of scientific papers in the learning phase. Second,
we can see, according to Table 2, that the training data process with the Abstract
approach is more efficient and fruitful in terms of accuracy than the Full paper
approach. We recorded that SVM, KNN, RF and MN Naïve Bayes present motivating
performances.
We explored both methods: Machine learning and Deep learning. In contrast to
image classification, we observed in our specific application of text mining that the
time-consuming process of deep learning does not outperform machine learning. At the
contrary, in some cases, machine learning produces significantly better results. We
believe that, in our experience, deep learning does not provide efficient results given the
size of our database.
Table 2. Comparative table of machine learning algorithms based on our database in the case of
300 Full papers and 300 abstracts.
Machine
learning
methods
300 full papers
300 abstracts
Accuracy Precision Recall F1-score Accuracy Precision Recall F1-score
SVM
80%
75%
81% 78%
81%
74%
72% 73%
KNN
62%
53%
58% 55%
65%
49%
56% 52%
RF
81%
85%
72% 78%
83%
81%
70% 75%
MN_NB
74%
82%
58% 68%
79%
83%
58% 69%
DT
86%
79%
86% 82%
75%
65%
74% 69%
GB
78%
81%
78% 79%
81%
75%
69% 71%
354
M. Khadhraoui et al.
Précédent

- 359/446

Suivant