the processing time of single-frame image to 2.5 s. However, SPP net failed to optimize the hard disk storage space of R-CNN. To address the issue, Girshic proposed the
Fast R-CNN [5].
Fast R-CNN proposed a multitask loss function, which adds the loss of target
positioning to the traditional loss function to correct the position information. After
that, Ren et al. designed the structure of Faster R-CNN [6] and used the ConvNet to
generate candidate regions directly. Faster R-CNN realizes end-to-end training and
realizes real-time detection. On this basis, some improved methods such as regionbased fully convolutional networks (R-FCN) [7], Mask R-CNN [8] have been proposed
successively. However, most of these methods are based on the three-step strategy of
candidate region selection, feature extraction, and classification.
After that, some researchers proposed the regression-based object detection
method, which was represented by YOLO [9, 10], SSD [11], etc. These methods only
use one ConvNet for detection and use the regression method to correct the object
location, making the detection speed much faster than the candidate region-based
detection methods such as Faster R-CNN. However, their disadvantages are large
localization error, which is mainly because it is very difficult to get an accurate location
by regression directly in the absence of proposal regions. In the training, there will be a
wide range jitter of the bounding box and the losses function is hard to converge. In
addition, the detection performance often becomes poor when multiple adjacent small
objects appear. The rest of this paper, I will focus on the key techniques used in object
detection using ConvNet.
2 Convolutional Implementation of Sliding Windows
In this section, we will introduce to build a convolutional neural network to implementation of sliding windows. The traditional sliding window method is very slow due
to a lot of repeated convolution computation. In this part, let us see how we can
improve the processing speed by using the convolutional implementation of sliding
windows.
Firstly, let us see how we can transform fully connected layers into convolutional
layers. For illustrative purposes, let us see a simple network as shown in Fig. 1a. Our
object detection algorithm inputs 14 Â 14 Â 3 images, then 16 of 3 Â 3 filters are
used to map the inputs images from 14 Â 14 Â 3 to 10 Â 10 Â 16, then does a 2 Â 2
max pooling to reduce it to 4 Â 4 Â 16, then has a fully connected layer to connect
512 units, and then finally, outputs y using a softmax layer. Here, we assume that there
are four categories, which are pedestrian, car, traffic light, and background. Then y
have four units, corresponding to the crossed probabilities of the four classes that the
softmax unit is classifying among. And the four classes could be pedestrian, car, traffic
lights, or background.
158
X. Zhang et al.
Fast R-CNN [5].
Fast R-CNN proposed a multitask loss function, which adds the loss of target
positioning to the traditional loss function to correct the position information. After
that, Ren et al. designed the structure of Faster R-CNN [6] and used the ConvNet to
generate candidate regions directly. Faster R-CNN realizes end-to-end training and
realizes real-time detection. On this basis, some improved methods such as regionbased fully convolutional networks (R-FCN) [7], Mask R-CNN [8] have been proposed
successively. However, most of these methods are based on the three-step strategy of
candidate region selection, feature extraction, and classification.
After that, some researchers proposed the regression-based object detection
method, which was represented by YOLO [9, 10], SSD [11], etc. These methods only
use one ConvNet for detection and use the regression method to correct the object
location, making the detection speed much faster than the candidate region-based
detection methods such as Faster R-CNN. However, their disadvantages are large
localization error, which is mainly because it is very difficult to get an accurate location
by regression directly in the absence of proposal regions. In the training, there will be a
wide range jitter of the bounding box and the losses function is hard to converge. In
addition, the detection performance often becomes poor when multiple adjacent small
objects appear. The rest of this paper, I will focus on the key techniques used in object
detection using ConvNet.
2 Convolutional Implementation of Sliding Windows
In this section, we will introduce to build a convolutional neural network to implementation of sliding windows. The traditional sliding window method is very slow due
to a lot of repeated convolution computation. In this part, let us see how we can
improve the processing speed by using the convolutional implementation of sliding
windows.
Firstly, let us see how we can transform fully connected layers into convolutional
layers. For illustrative purposes, let us see a simple network as shown in Fig. 1a. Our
object detection algorithm inputs 14 Â 14 Â 3 images, then 16 of 3 Â 3 filters are
used to map the inputs images from 14 Â 14 Â 3 to 10 Â 10 Â 16, then does a 2 Â 2
max pooling to reduce it to 4 Â 4 Â 16, then has a fully connected layer to connect
512 units, and then finally, outputs y using a softmax layer. Here, we assume that there
are four categories, which are pedestrian, car, traffic light, and background. Then y
have four units, corresponding to the crossed probabilities of the four classes that the
softmax unit is classifying among. And the four classes could be pedestrian, car, traffic
lights, or background.
158
X. Zhang et al.
