(2) The impact of data set ratio on event recognition
This experiment used 4000 data as the training set, 200 data as the test set, and the ratio
of the amount of emergency data to non-emergency data was 1:1.
There are two important indicators for the model test, namely TP and TN. TP
stands for the numbers that test and true values are both positive. TN represents the
numbers that test value is negative but true value is positive. In this test result, the
number of TP is 95, but the number of TN is only 72, then it is suspected that the effect
of the model is related to the ratio and number of training data sets. Because the amount
of negative sample data for training is too small, the characteristics of the model for
negative samples are not completely extracted. Then we increase the number of training
for negative samples by 1000 each time. The test results are as follows.
Table 2 shows that when the data ratio is 2:3, the accuracy rate reaches 90%, and
when the positive and negative ratios are too large, the feature extraction will be
disordered, resulting in a decrease in accuracy.
(3) The influence of word vector dimension on event recognition
The dimension of the word vector represents the characteristics of the word. Generally
speaking, the more features, the more accurately distinguish the word from the word,
but as the dimension increases, the relationship between words will also fade. Too high
dimension will fade the relationship between words. If the dimension is too low, the
word can not be distinguished, so the choice of the dimension of the word vector
depends on your actual application scenario.
Here, the dimensions of the word vector are 50, 70, 90, 100, 110, 130, and 150. The
accuracy, recall, and F1 values are as in Fig. 5.
Table 1. Different model performance verification results
Model Accuracy (%) Recall (%) F1 (%)
RT
53.50
52.29
52.89
RR
57.50
63.64
60.41
GT
75.50
68.35
71.74
GR
77
70.45
73.60
LT
57
64
60.30
LR
83.50
77.24
80.25
Table 2. Comparison of different data ratios
Data ratio Accuracy (%) Recall (%) F1 (%)
1:1
83.5
77.24
80.25
2:3
90
92.55
91.26
1:2
88.5
88.89
88.69
102
H. He et al.
Précédent

- 114/679

Suivant