3.3 Incorporating Sentimental Polarity into CNN-LSTM
The specific workflow of CNN-LSTM rumor detection model integrating sentimental
polarity is as follows:
(1) Construct microblog corpus, and the data set is divided into rumor microblog and
normal microblog.
(2) Cleaning data; filtering garbage comments and word segmentation; and training
word vector model to vectorize microblog text.
(3) Sentimental analysis of microblog text. The formula for calculating the sentimental variance of microblog comments is as follows:
sentiment variance ¼
1
n
X n
i¼1
text sentiment À s
ð
Þ
2
ð2Þ
(4) Constructing the sentimental feature set of microblog comments and the sentimental difference is calculated according to the sentimental intensity and variance
of microblog comments. The formula for calculating sentimental differences is as
follows:
sentiment diff ¼ a text sentiment À a
j
j þ b sentiment variance À b
j
j
ð3Þ
(5) Training a rumor detection model based on CNN-LSTM, and predicting the text
as the probability of rumors (and non-rumors).
(6) When the sentimental difference and the predicted results of CNN-LSTM are
substituted into the rumor calculation formula, if the calculation results are larger
than the threshold value, then microblog is a rumor; otherwise, it is not a rumor.
Rumor predict represents rumor detection value, k and c are weighting parameters,
and the values are 0.4 and 0.6. Then the rumor calculation formula is as follows:
rumor predict ¼ k à sentiment diff þ c à cnn À lstm predict
ð4Þ
Rumor detection threshold is set to K. If rumor predict > K, it is judged as a rumor;
otherwise, it is judged as a non-rumor.
The model framework is shown in Fig. 2.
4 Experiments and Analysis
4.1 Data Set and Spam Comment Filtering
Directly crawl the fake Weibo information published by Weibo Community Management Center, and collect the name of the rumor microblog reporter, the name of the
publisher, and the comments and so on. Then crawl the same number of normal
microblogs to ensure the quality of information by limiting the information published
by at least blue V users. Finally, 10,229 rumors and 10,200 non-rumors were collected
as rumor training sets, coae2014 task 4 sentimental data set as sentimental training sets.
Microblog Rumor Detection Based on Comment Sentiment and CNN-LSTM
151
Précédent

- 163/679

Suivant