2 Related Works
Nowadays, most researchers regard rumor detection as text classification [1] and design
different features in the framework of classification, mainly considering the features of
text content, communication structure, and credibility.
The first method detects rumors based on text features of microblog. Qazvinian
et al. [2] detect rumors on Bayesian classifier by selecting Twitter text features, user
features, and related character label features. Sun et al. [3] have studied the major
rumors on the microblog platform. They extract features from the microblog text,
microblog video pictures, and microblog users and use machine learning algorithms
such as Bayesian and decision tree to carry out experiments.
The second method detects rumors based on the characteristics of communication
structure. Takahashi et al. [4] collected comments from users on Twitter after the
tsunami in Japan and found that we can find features from the growth of rumor
microblogs and forwarding rate. Nourbakhsh et al. [5] collected hundreds of rumors
and studied their semantic features. At the same time, the influence of different user
roles on rumor communication structure was studied. Mendoza et al. [6] studied the
structure of rumors and user behavior under the topic of “Chile earthquake” on Twitter.
They found that rumors are more intense and faster than normal news discussions.
The third method detects rumors based on credibility. Gupta et al. [7] studied hot
events and found that there was more false information about hot events on Twitter.
Suzuki et al. [8] designed a method to evaluate the credibility of Twitter information
and judged the credibility of Twitter information by the retention rate of Twitter content
forwarding. If the original Twitter message remains almost unchanged after forwarding, the credibility will be higher; on the contrary, it will be lower.
However, up to now, the traditional machine learning method [9] relies on feature
engineering, which is time-consuming and laborious. Common in-depth learning
model also has good performance, but it cannot learn deeper features, such as the
sentimental characteristics of Weibo. In order to solve the above problems, combining
CNN and LSTM to extract the basic features of rumors and utilizing the sentimental
analysis technology to mine sentimental features can detect rumor effectively [10].
3 Rumor Detection Based on Comment Sentiment
and CNN-LSTM
3.1 CNN-LSTM
The CNN-LSTM is a fusion model, which combines the advantages of the two models.
The corresponding vectors are generated by word vector technology after text data
pretreatment. First, the convolution layer of CNN [11] is used to extract local features,
then the pooling layer is used to reduce the parameter dimension, then the LSTM [12]
layer is entered, and then the full connection layer and Softmax layer are used to realize
the classification output. The structure of the CNN-LSTM model is shown in Fig. 1.
Microblog Rumor Detection Based on Comment Sentiment and CNN-LSTM
149
Nowadays, most researchers regard rumor detection as text classification [1] and design
different features in the framework of classification, mainly considering the features of
text content, communication structure, and credibility.
The first method detects rumors based on text features of microblog. Qazvinian
et al. [2] detect rumors on Bayesian classifier by selecting Twitter text features, user
features, and related character label features. Sun et al. [3] have studied the major
rumors on the microblog platform. They extract features from the microblog text,
microblog video pictures, and microblog users and use machine learning algorithms
such as Bayesian and decision tree to carry out experiments.
The second method detects rumors based on the characteristics of communication
structure. Takahashi et al. [4] collected comments from users on Twitter after the
tsunami in Japan and found that we can find features from the growth of rumor
microblogs and forwarding rate. Nourbakhsh et al. [5] collected hundreds of rumors
and studied their semantic features. At the same time, the influence of different user
roles on rumor communication structure was studied. Mendoza et al. [6] studied the
structure of rumors and user behavior under the topic of “Chile earthquake” on Twitter.
They found that rumors are more intense and faster than normal news discussions.
The third method detects rumors based on credibility. Gupta et al. [7] studied hot
events and found that there was more false information about hot events on Twitter.
Suzuki et al. [8] designed a method to evaluate the credibility of Twitter information
and judged the credibility of Twitter information by the retention rate of Twitter content
forwarding. If the original Twitter message remains almost unchanged after forwarding, the credibility will be higher; on the contrary, it will be lower.
However, up to now, the traditional machine learning method [9] relies on feature
engineering, which is time-consuming and laborious. Common in-depth learning
model also has good performance, but it cannot learn deeper features, such as the
sentimental characteristics of Weibo. In order to solve the above problems, combining
CNN and LSTM to extract the basic features of rumors and utilizing the sentimental
analysis technology to mine sentimental features can detect rumor effectively [10].
3 Rumor Detection Based on Comment Sentiment
and CNN-LSTM
3.1 CNN-LSTM
The CNN-LSTM is a fusion model, which combines the advantages of the two models.
The corresponding vectors are generated by word vector technology after text data
pretreatment. First, the convolution layer of CNN [11] is used to extract local features,
then the pooling layer is used to reduce the parameter dimension, then the LSTM [12]
layer is entered, and then the full connection layer and Softmax layer are used to realize
the classification output. The structure of the CNN-LSTM model is shown in Fig. 1.
Microblog Rumor Detection Based on Comment Sentiment and CNN-LSTM
149
