be generated one *.xml file for each image in the images directory. These will be used
to train the new object detection classifier.
5 Conclusion
Object detection has been developing rapidly in the field of computer vision and has
repeatedly created amazing achievements. The ConvNet has a strong versatility and
portability, and it is widely used in robot, autonomous driving, medical assistance, etc.
This chapter focuses on several core problems in object detection: object localization,
sliding windows object detection, and transfer learning. In addition, we also introduced
how to classify and locate cells in medical image by ConvNet.
In summary, the success of ConvNet in recent years mainly depends on three
pillars: data, model, and calculation power. A large amount of manually annotated data
makes it possible to conduct supervised training. A deeper and larger model improves
the recognition ability of the neural network. The combination with GPU and the rapid
development of computer hardware makes large-scale training become time-saving and
effective. However, the research on ConvNet is just beginning, and many aspects need
further study.
At present, ConvNet needs training samples of tens of thousands or even millions
of levels, and the training process for such a large number of samples is also extremely
long. And in many areas, it is very expensive to obtain large numbers of precisely
labeled samples. How to generate a good neural network from a small amount of data
will be the future research direction. How to get satisfactory performance by train a
small amount of dataset will be the future research direction.
In addition, the trend is that the deeper the network, the better the performance of
ConvNet, and some networks even reach thousands of layers. However, as the network
deepens, overfitting and gradient disappeared become more serious. Although the
research such as residual network [18], etc., is devoted to solving such problems, the
large model also limits the application of ConvNet on common devices, especially
mobile devices. The ConvNet needs to optimize the structural design to find more
efficient neurons and structural units.
References
1. Wang X, Han TX, Yan S (2009) An HOG-LBP human detector with partial occlusion
handling. In: 2009 IEEE 12th international conference on computer vision, Sept 2009,
pp 32–39
2. Girshick R, Donahue J, Darrell T, Malik J (2014) Rich feature hierarchies for accurate object
detection and semantic segmentation. In: Proceedings of the IEEE computer society
conference on computer vision and pattern recognition, pp 580–587
3. Everingham M, Gool L, Williams CKI, Winn J, Zisserman A (2009) The pascal visual object
classes (VOC) challenge. Int J Comput Vis 88(2):303–338
4. He K, Zhang X, Ren S, Sun J (2015) Spatial pyramid pooling in deep convolutional
networks for visual recognition. IEEE Trans Pattern Anal Mach Intell 346–361
A Guideline for Object Detection Using Convolutional …
163
to train the new object detection classifier.
5 Conclusion
Object detection has been developing rapidly in the field of computer vision and has
repeatedly created amazing achievements. The ConvNet has a strong versatility and
portability, and it is widely used in robot, autonomous driving, medical assistance, etc.
This chapter focuses on several core problems in object detection: object localization,
sliding windows object detection, and transfer learning. In addition, we also introduced
how to classify and locate cells in medical image by ConvNet.
In summary, the success of ConvNet in recent years mainly depends on three
pillars: data, model, and calculation power. A large amount of manually annotated data
makes it possible to conduct supervised training. A deeper and larger model improves
the recognition ability of the neural network. The combination with GPU and the rapid
development of computer hardware makes large-scale training become time-saving and
effective. However, the research on ConvNet is just beginning, and many aspects need
further study.
At present, ConvNet needs training samples of tens of thousands or even millions
of levels, and the training process for such a large number of samples is also extremely
long. And in many areas, it is very expensive to obtain large numbers of precisely
labeled samples. How to generate a good neural network from a small amount of data
will be the future research direction. How to get satisfactory performance by train a
small amount of dataset will be the future research direction.
In addition, the trend is that the deeper the network, the better the performance of
ConvNet, and some networks even reach thousands of layers. However, as the network
deepens, overfitting and gradient disappeared become more serious. Although the
research such as residual network [18], etc., is devoted to solving such problems, the
large model also limits the application of ConvNet on common devices, especially
mobile devices. The ConvNet needs to optimize the structural design to find more
efficient neurons and structural units.
References
1. Wang X, Han TX, Yan S (2009) An HOG-LBP human detector with partial occlusion
handling. In: 2009 IEEE 12th international conference on computer vision, Sept 2009,
pp 32–39
2. Girshick R, Donahue J, Darrell T, Malik J (2014) Rich feature hierarchies for accurate object
detection and semantic segmentation. In: Proceedings of the IEEE computer society
conference on computer vision and pattern recognition, pp 580–587
3. Everingham M, Gool L, Williams CKI, Winn J, Zisserman A (2009) The pascal visual object
classes (VOC) challenge. Int J Comput Vis 88(2):303–338
4. He K, Zhang X, Ren S, Sun J (2015) Spatial pyramid pooling in deep convolutional
networks for visual recognition. IEEE Trans Pattern Anal Mach Intell 346–361
A Guideline for Object Detection Using Convolutional …
163
