Imbalanced Data Classification with Deep
Support Vector Machines
Li Zhang, Wei Wang
(&) , Mengjun Zhang, and Zhixiong Wang
Tianjin Key Laboratory of Wireless Mobile Communications and Power
Transmission, Tianjin Normal University, Tianjin 300387, China
weiwang@tjnu.edu.cn
Abstract. In recent years, deep learning has become increasingly popular in
various fields. However, the performance of deep learning on imbalanced data
has not been examined. The imbalanced data is a special problem in target
detection and classification task, where the number of one class is less than the
other classes. This paper focuses on evaluating the performance of the deep
support vector machine (DSVM) algorithm in dealing with imbalanced human
target detection datasets. Furthermore, we optimize the parameters of the DSVM
algorithm to obtain better detection performance. It is compared with the stacked
auto-encoder (SAE) and the support vector machine (SVM) algorithm. Finally,
numerical experimental results show that the DSVM algorithm can effectively
capture the minority class.
Keywords: Deep support vector machines Á Imbalanced data Á
Human target detection
1 Introduction
The skewed distribution of data samples between different classes is a common phenomenon in many real-world classification issues, such as spam review detection [1],
insurance fraud detection [2], web author identification [3], wilt disease classification
[4] and multimedia concept detection [5]. In this paper, we focus on binary classes
classification issues for imbalanced datasets, where the classes that contain a small
number of instances are named the minority class, while another dominant instance
space is named the majority class. The imbalanced datasets have reduced the performance of existing learning methods and posed a relatively new challenge to them for
the hardness to learn the minority instance.
Confronted with the issue of imbalanced learning, plenty of approaches have been
proposed to solve the imbalanced distribution. The existing methods can be mainly
divided into two categories, one is the data processing level, and the other is the
classification algorithms level. Sampling strategies are often used at the data processing
level to provide a balanced class distribution. In [6], a new over-sampling technique
called Density-Based Synthetic Minority Over-sampling Technique (DBSMOTE) was
proposed, and the method was designed to over-sample an arbitrarily shaped cluster.
He et al. [7] applied two ways for reducing the bias and shifting the classification
decision boundary by the Adaptive Synthetic Sampling approach. In [8], two kinds of
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 87–95, 2020.
https://doi.org/10.1007/978-981-15-0187-6_10
Support Vector Machines
Li Zhang, Wei Wang
(&) , Mengjun Zhang, and Zhixiong Wang
Tianjin Key Laboratory of Wireless Mobile Communications and Power
Transmission, Tianjin Normal University, Tianjin 300387, China
weiwang@tjnu.edu.cn
Abstract. In recent years, deep learning has become increasingly popular in
various fields. However, the performance of deep learning on imbalanced data
has not been examined. The imbalanced data is a special problem in target
detection and classification task, where the number of one class is less than the
other classes. This paper focuses on evaluating the performance of the deep
support vector machine (DSVM) algorithm in dealing with imbalanced human
target detection datasets. Furthermore, we optimize the parameters of the DSVM
algorithm to obtain better detection performance. It is compared with the stacked
auto-encoder (SAE) and the support vector machine (SVM) algorithm. Finally,
numerical experimental results show that the DSVM algorithm can effectively
capture the minority class.
Keywords: Deep support vector machines Á Imbalanced data Á
Human target detection
1 Introduction
The skewed distribution of data samples between different classes is a common phenomenon in many real-world classification issues, such as spam review detection [1],
insurance fraud detection [2], web author identification [3], wilt disease classification
[4] and multimedia concept detection [5]. In this paper, we focus on binary classes
classification issues for imbalanced datasets, where the classes that contain a small
number of instances are named the minority class, while another dominant instance
space is named the majority class. The imbalanced datasets have reduced the performance of existing learning methods and posed a relatively new challenge to them for
the hardness to learn the minority instance.
Confronted with the issue of imbalanced learning, plenty of approaches have been
proposed to solve the imbalanced distribution. The existing methods can be mainly
divided into two categories, one is the data processing level, and the other is the
classification algorithms level. Sampling strategies are often used at the data processing
level to provide a balanced class distribution. In [6], a new over-sampling technique
called Density-Based Synthetic Minority Over-sampling Technique (DBSMOTE) was
proposed, and the method was designed to over-sample an arbitrarily shaped cluster.
He et al. [7] applied two ways for reducing the bias and shifting the classification
decision boundary by the Adaptive Synthetic Sampling approach. In [8], two kinds of
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 87–95, 2020.
https://doi.org/10.1007/978-981-15-0187-6_10
