A Guideline for Object Detection Using
Convolutional Neural Networks
Xingguo Zhang, Guoyue Chen
(&) , Kazuki Saruta, and Yuki Terata
Akita Prefectural University, 84-4 Aza Ebinokuchi Tsuchiya, Yurihonjo City
015-0055, Japan
chen@akita-pu.ac.jp
Abstract. The main purpose of object detection is to detect and locate specific
targets from images. The traditional detection methods are usually complex and
require prior knowledge of the detection target. In this paper, we will introduce
how to use convolutional neural networks to perform object detection from
image. This is one of the important areas of computer vision. In order to build up
to object detection, we first learn about how we can get the object localization or
landmark by a neural network. And then I will give the detail of sliding windows detection algorithm and introduce how to use the convolutional implementation of sliding windows to speed up the process. Then we will introduce
the transfer learning and how to prepare your own learning data for training
networks.
Keywords: Object detection Á Bounding box Á Transfer learning
1 Introduction
The main purpose of object detection is to detect and locate specific targets from
images. The traditional detection model usually represents the target object by handcraft features and then predicts the category and location by classifier [1]. These
methods are usually intuitive and easy to understand. However, the design of these
methods is often complex and requires prior knowledge of the detection target. In
addition, it is highly dependent on specific tasks and has poor portability. Once the
detection target changes significantly, it is necessary to redesign the algorithm.
In recent years, with the improvement of hardware and algorithm, convolutional
neural networks (ConvNet) have achieved great success in image classification, which
led researchers to study its effects in other areas of computer vision. The early algorithm based on deep learning is generally divided into three steps, selecting the object
candidate region (proposal), extracting the features of the candidate region, and finally
putting the extracted features into the classifier to predict the object category. In 2014,
Girshick designed the R-CNN model [2] based on ConvNet. The mean average precision (mAP) of this model in the object detection task of PASCAL VOC [3] was
62.4%, which was nearly 20% higher than the traditional algorithm. He et al. proposed
Spatial Pyramid Pooling Net (SPP net) [4], which only performed convolution operation on the whole picture once and added pyramid pooling layer after convolution
layer, so as to fix the feature map to the required size. This greatly saves time, reducing
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 157–164, 2020.
https://doi.org/10.1007/978-981-15-0187-6_18
Convolutional Neural Networks
Xingguo Zhang, Guoyue Chen
(&) , Kazuki Saruta, and Yuki Terata
Akita Prefectural University, 84-4 Aza Ebinokuchi Tsuchiya, Yurihonjo City
015-0055, Japan
chen@akita-pu.ac.jp
Abstract. The main purpose of object detection is to detect and locate specific
targets from images. The traditional detection methods are usually complex and
require prior knowledge of the detection target. In this paper, we will introduce
how to use convolutional neural networks to perform object detection from
image. This is one of the important areas of computer vision. In order to build up
to object detection, we first learn about how we can get the object localization or
landmark by a neural network. And then I will give the detail of sliding windows detection algorithm and introduce how to use the convolutional implementation of sliding windows to speed up the process. Then we will introduce
the transfer learning and how to prepare your own learning data for training
networks.
Keywords: Object detection Á Bounding box Á Transfer learning
1 Introduction
The main purpose of object detection is to detect and locate specific targets from
images. The traditional detection model usually represents the target object by handcraft features and then predicts the category and location by classifier [1]. These
methods are usually intuitive and easy to understand. However, the design of these
methods is often complex and requires prior knowledge of the detection target. In
addition, it is highly dependent on specific tasks and has poor portability. Once the
detection target changes significantly, it is necessary to redesign the algorithm.
In recent years, with the improvement of hardware and algorithm, convolutional
neural networks (ConvNet) have achieved great success in image classification, which
led researchers to study its effects in other areas of computer vision. The early algorithm based on deep learning is generally divided into three steps, selecting the object
candidate region (proposal), extracting the features of the candidate region, and finally
putting the extracted features into the classifier to predict the object category. In 2014,
Girshick designed the R-CNN model [2] based on ConvNet. The mean average precision (mAP) of this model in the object detection task of PASCAL VOC [3] was
62.4%, which was nearly 20% higher than the traditional algorithm. He et al. proposed
Spatial Pyramid Pooling Net (SPP net) [4], which only performed convolution operation on the whole picture once and added pyramid pooling layer after convolution
layer, so as to fix the feature map to the required size. This greatly saves time, reducing
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 157–164, 2020.
https://doi.org/10.1007/978-981-15-0187-6_18
