All in all, the previous approaches have a common need to extract suitable handcraft features beforehand, which is a challenging task. Thanks to the development of
neural networks, the new deep neural network approaches provide a hopeful way to
solve the problems. However, the main concerns of these popular networks are
detecting relatively large objects in the images. Their abilities are limited when they are
employed to detect small objects. More specific network structures should be designed
for different applications, such as hand detection. Le et al. [8], Hoang Ngan Le et al. [9]
solve the problem by synchronizing the global and the local context features using
Faster R-CNN for semantic detection. Ding et al. [10] adopt multi-scale CNN networks
to detect hand. Their region proposals are generated at multiple scales; then the feature
maps extracted from different layers are fused to get the hand bounding boxes. Deng
et al. [11] design a network to detect hand region and hand in-plane rotation firstly and
then give the final hand detection results by feature sharing. Although these RPN-based
methods achieve better hand detection results than traditional models, they need
applying hundreds of times per region sub-network, which is time consuming. Different
from the above-mentioned networks, our new hand detection network is efficient, as
well as better or comparative detection accuracy.
3 Methods
3.1 Overview of the Approach
The object of our work is to detect the hand and report their positions from RGB
images. Since the hand sizes, shapes, and appearances exist in large variants, feature
maps from different levels should be integrated to preserve both the global and local
cues. So, we use multi-scale structure to extract feature maps with different resolution,
then gradually merge these feature maps to form identifiable features. Supervised by
the object region confidence and the coordinates of the bounding boxes, the network
can predict the hand region accurately. The whole network model is shown in Fig. 1.
The model is lightweighted, pixel-wise scoring, no time-consuming candidate proposal
step, and very efficient.
3.2 Network Structure
For object detection, region proposal-based CNN is an important branch. The classical
and representative networks are R-CNN and its variants. R-CNN applies deep convolutional layers as feature extractor to classify given region proposals. It is tailored by
support vector machines (SVM) for object detection and bounding box regression. RCNN can achieve accurate results, but is quite time consuming. Fast R-CNN and Faster
R-CNN accelerate the detections by feature sharing, ROI pooling, and special region
proposal network, but they still lack high efficiency due to the dependence on the
external region proposal methods. Moreover, they are difficult to detect relatively small
hands in cars because of the ROI pooling.
To improve efficiency, in our model, as shown in Fig. 1, the region proposal
network is discarded. The object region is determined by the confidence of every pixel
22
Y. Li et al.
neural networks, the new deep neural network approaches provide a hopeful way to
solve the problems. However, the main concerns of these popular networks are
detecting relatively large objects in the images. Their abilities are limited when they are
employed to detect small objects. More specific network structures should be designed
for different applications, such as hand detection. Le et al. [8], Hoang Ngan Le et al. [9]
solve the problem by synchronizing the global and the local context features using
Faster R-CNN for semantic detection. Ding et al. [10] adopt multi-scale CNN networks
to detect hand. Their region proposals are generated at multiple scales; then the feature
maps extracted from different layers are fused to get the hand bounding boxes. Deng
et al. [11] design a network to detect hand region and hand in-plane rotation firstly and
then give the final hand detection results by feature sharing. Although these RPN-based
methods achieve better hand detection results than traditional models, they need
applying hundreds of times per region sub-network, which is time consuming. Different
from the above-mentioned networks, our new hand detection network is efficient, as
well as better or comparative detection accuracy.
3 Methods
3.1 Overview of the Approach
The object of our work is to detect the hand and report their positions from RGB
images. Since the hand sizes, shapes, and appearances exist in large variants, feature
maps from different levels should be integrated to preserve both the global and local
cues. So, we use multi-scale structure to extract feature maps with different resolution,
then gradually merge these feature maps to form identifiable features. Supervised by
the object region confidence and the coordinates of the bounding boxes, the network
can predict the hand region accurately. The whole network model is shown in Fig. 1.
The model is lightweighted, pixel-wise scoring, no time-consuming candidate proposal
step, and very efficient.
3.2 Network Structure
For object detection, region proposal-based CNN is an important branch. The classical
and representative networks are R-CNN and its variants. R-CNN applies deep convolutional layers as feature extractor to classify given region proposals. It is tailored by
support vector machines (SVM) for object detection and bounding box regression. RCNN can achieve accurate results, but is quite time consuming. Fast R-CNN and Faster
R-CNN accelerate the detections by feature sharing, ROI pooling, and special region
proposal network, but they still lack high efficiency due to the dependence on the
external region proposal methods. Moreover, they are difficult to detect relatively small
hands in cars because of the ROI pooling.
To improve efficiency, in our model, as shown in Fig. 1, the region proposal
network is discarded. The object region is determined by the confidence of every pixel
22
Y. Li et al.
