Data Stream Adaptive Partitioning of Sliding
Window Based on Gaussian Restricted
Boltzmann Machine
Wei Wang
(&) and Mengjun Zhang
Tianjin Key Laboratory of Wireless Mobile Communications and Power
Transmission, Tianjin Normal University, Tianjin 300387, China
weiwang@tjnu.edu.cn
Abstract. In this paper, the adaptive partitioning problem of data stream under
sliding window is discussed. Gaussian restricted Boltzmann machine (GRBM)
model supporting decimal input is proposed, which can be trained through
iteration for data reconstruction subsequently. At the same time, a data stream
adaptive block algorithm based on Kullback–Leibler divergence (KL distance)
is proposed to compare the probability distribution difference in the sliding
window. Then, obtain the predicted value by the distribution of the previous
data and determine whether the KL distance is within the confidence interval, so
as to realize the adaptive adjustment of the sliding window, and the divided of
data stream.
Keywords: Data stream Á Gaussian–Bernoulli restricted Boltzmann machine
model Á Kullback–Leibler divergence
1 Introduction
Data stream classification is a core problem in data mining due to its characteristics of
continuous, real-time arrival, large amount, unrestricted, and unpredictable. The sliding
window was adopted to ensure that the current data are the latest valid data in the data
stream. However, the choice of window size will seriously affect the results of the
experiment, for that too small window selection will lead to incomplete information
collection while too large window selection will affect the efficiency of learning. So,
how to choose the length of the window is a problem worth studying owing to
appropriate length that can effectively improve the possibility and efficiency of the
algorithm, reduce energy consumption, and can obviously save memory. Sliding
window is used by Bifet in [1] which size is not fixed but recalculated online according
to the rate of change observed from the data of the window itself. Judgement mechanisms of false positives and false negatives were used while it is inefficient in time and
memory. The sliding window can also be incorporated with different prediction
mechanisms such as [2], making them suitable for online prediction by adapting the
number of the samples.
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 220–228, 2020.
https://doi.org/10.1007/978-981-15-0187-6_25
Précédent

- 232/679

Suivant