through your same 400 of 5 Â 5 filters to get a 2 Â 2 Â 400 volume output. So, now
instead of a 1 Â 1 Â 400 volume, you can get a 2 Â 2 Â 400 volume. Do that one
more time, now you are left with a 2 Â 2 Â 4 output volume instead of 1 Â 1 Â 4.
So, in this final output layer, which turns out the upper left 1 Â 1 Â 4 subset gives you
the result of running in the upper left corner 14 Â 14 regions of input image, the upper
right 1 Â 1 Â 4 volume gives you the upper right result and same argument for the
lower left and right 1 Â 1 Â 4 volume. If you step through all the steps of the calculation, you will find that the value of each cell is the same as if you cut out each
region and input it to ConvNet.
So, what this convolutional implementation does is you need not run forward
propagation on four subsets of the input image independently. We only need to do
forward propagation once for the entire image and share a lot of the computation in the
regions of the image that is common. You can refer to [12] to get more details about
this part.
3 Transfer Learning
When your network is ready, prepare training data for learning. The process of optimizing the neural network is to find the value of each layer of parameters that minimize
the output of the loss function. But if you do a little bit of calculation, you will find that
there are huge parameters in a neural network that needs to be trained, i. e, a mediumsized vgg-16 network [13] has at least 1.5 Â 10
7 parameters.
If you have a new network that needs to be trained from scratch, then you have to
prepare a lot of training samples, such as a database like ImageNet [14], MS COCO
Fig. 2. Object detection by convolutional implementation
160
X. Zhang et al.
instead of a 1 Â 1 Â 400 volume, you can get a 2 Â 2 Â 400 volume. Do that one
more time, now you are left with a 2 Â 2 Â 4 output volume instead of 1 Â 1 Â 4.
So, in this final output layer, which turns out the upper left 1 Â 1 Â 4 subset gives you
the result of running in the upper left corner 14 Â 14 regions of input image, the upper
right 1 Â 1 Â 4 volume gives you the upper right result and same argument for the
lower left and right 1 Â 1 Â 4 volume. If you step through all the steps of the calculation, you will find that the value of each cell is the same as if you cut out each
region and input it to ConvNet.
So, what this convolutional implementation does is you need not run forward
propagation on four subsets of the input image independently. We only need to do
forward propagation once for the entire image and share a lot of the computation in the
regions of the image that is common. You can refer to [12] to get more details about
this part.
3 Transfer Learning
When your network is ready, prepare training data for learning. The process of optimizing the neural network is to find the value of each layer of parameters that minimize
the output of the loss function. But if you do a little bit of calculation, you will find that
there are huge parameters in a neural network that needs to be trained, i. e, a mediumsized vgg-16 network [13] has at least 1.5 Â 10
7 parameters.
If you have a new network that needs to be trained from scratch, then you have to
prepare a lot of training samples, such as a database like ImageNet [14], MS COCO
Fig. 2. Object detection by convolutional implementation
160
X. Zhang et al.
