Now, let us show you how these fully connected layers can be turned into convolutional layers. As shown in Fig. 1b, the first few layers of the ConvNet have the
same structure. After that, to implement the convolutional layer, we use 512 of
5 Â 5 Â 16 filters to do the convolution, and the output dimension is going to be
1 Â 1 Â 512. In mathematically, this is the same as a fully connected layer, because
each of these 512 nodes has a filter of dimension 5 Â 5 Â 16, and so each of those 512
values is some arbitrary linear function of these 5 Â 5 Â 16 activations from the
previous layer.
Finally, we are going to use four of 1 Â 1 Â 512 filters by a softmax activation to
get a 1 Â 1 Â 4 volume as the output of this network. So, this shows how you can take
these fully connected layers and implement those using convolutional layers, and these
fully connected layers are now implemented as 1 Â 1 Â 512 and 1 Â 1 Â 4 volumes.
When the size of the test image we inputted changed to 16 Â 16 Â 3, assume our
trained detection window is still 14 Â 14 (shown in Fig. 2a). So in the original sliding
windows algorithm, you might want to input the first 14 Â 14 regions into a ConvNet
and run that once to generate a classification 0 or 1. Then slide the window to the right
by a stride = 2 pixels to get the second rectangular area, and run the whole ConvNet for
this window to get another label 0 or 1. Then repeat this process by slide the window
until get the output of lower right window. You will find for this small input image, we
run this ConvNet from above four times in order to get four labels. But we can see that
many of the operations by these four ConvNet are highly duplicated. So what the
convolutional implementation of sliding windows does is it allows these four forward
passes of the ConvNet to share a lot of computation.
As shown in Fig. 2b, you can take the ConvNet and just run it by the same
5 Â 5 Â 16 filters with same parameters, and you can get a 12 Â 12 Â 16 output
volume, and then do the max pool same as before, get a 6 Â 6 Â 16 output, run
Fig. 1. Convolutional implementation for fully connected layers
A Guideline for Object Detection Using Convolutional …
159
Précédent

- 171/679

Suivant