Hand Detection Based on Multi-scale Fully
Convolutional Networks
Yibo Li, MingMing Shi, and Xiangbo Lin
(&)
Faculty of Electronic Information and Electrical Engineering,
Dalian University of Technology, Dalian, China
linxbo@dlut.edu.cn
Abstract. Accurate hand detection is a challenging task because of large
variations of hand images in real-world scenarios. We present a simple yet
powerful multi-scale fully convolutional network structure that yields fast and
accurate hand detection on challenging VIVA dataset. The proposed model
directly detects and locates hands in driver’s cab of various size, shape,
appearance, and illumination in full images without time-consuming region
proposal step. The simple model with the well-designed loss functions promotes
the proposed approach to achieve very good hand detection results.
Keywords: Hand detection Á Convolutional neural network Á
Multi-scale fusion
1 Introduction
Hand detection from RGB images is a fundamental task in a lot of applications, such as
virtual reality, human activity monitoring, and robotic behavior demonstration. Though
many progresses have been made [1–3], they are still far from mature due to plenty of
difficulties in practice. Hands in an image often occupy small regions with resolutions
not high enough. Often their detection is seriously influenced by highly occlusions,
varied shape and appearances, different lighting conditions, and viewpoints. Due to the
less accurate detection performance, labor consuming manual detection has to be used
in many practical scenes. In this paper, we focus on hand detection problem in driver’s
cab environment by exploring more efficient and accurate hand detection method.
Traditional methods use handcraft features to detect human hands in the images.
These approaches suffer from limited discriminative capabilities and will not get reliable detections in complex situations. Along with the success of deep learning in the
computer vision field, convolutional neural networks (CNN) quickly spread in multiple
research areas. Many powerful object detection CNN frameworks are proposed, such as
variants of R-CNNs and YOLOs. However, these CNN networks are usually effective
to detect relatively large objects on several public datasets with limited category of
objects. This paper focuses on hand detection, and a specific network is proposed using
multi-scale fully convolutional structure. Both global and local context information are
efficiently synchronized and simultaneously represent the hand features. In addition to
geometric coordinates of the hand bounding box, a ROI region-related score map is
used to form the loss function. A balanced cross-entropy loss is used to facilitate a
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 20–27, 2020.
https://doi.org/10.1007/978-981-15-0187-6_3
Convolutional Networks
Yibo Li, MingMing Shi, and Xiangbo Lin
(&)
Faculty of Electronic Information and Electrical Engineering,
Dalian University of Technology, Dalian, China
linxbo@dlut.edu.cn
Abstract. Accurate hand detection is a challenging task because of large
variations of hand images in real-world scenarios. We present a simple yet
powerful multi-scale fully convolutional network structure that yields fast and
accurate hand detection on challenging VIVA dataset. The proposed model
directly detects and locates hands in driver’s cab of various size, shape,
appearance, and illumination in full images without time-consuming region
proposal step. The simple model with the well-designed loss functions promotes
the proposed approach to achieve very good hand detection results.
Keywords: Hand detection Á Convolutional neural network Á
Multi-scale fusion
1 Introduction
Hand detection from RGB images is a fundamental task in a lot of applications, such as
virtual reality, human activity monitoring, and robotic behavior demonstration. Though
many progresses have been made [1–3], they are still far from mature due to plenty of
difficulties in practice. Hands in an image often occupy small regions with resolutions
not high enough. Often their detection is seriously influenced by highly occlusions,
varied shape and appearances, different lighting conditions, and viewpoints. Due to the
less accurate detection performance, labor consuming manual detection has to be used
in many practical scenes. In this paper, we focus on hand detection problem in driver’s
cab environment by exploring more efficient and accurate hand detection method.
Traditional methods use handcraft features to detect human hands in the images.
These approaches suffer from limited discriminative capabilities and will not get reliable detections in complex situations. Along with the success of deep learning in the
computer vision field, convolutional neural networks (CNN) quickly spread in multiple
research areas. Many powerful object detection CNN frameworks are proposed, such as
variants of R-CNNs and YOLOs. However, these CNN networks are usually effective
to detect relatively large objects on several public datasets with limited category of
objects. This paper focuses on hand detection, and a specific network is proposed using
multi-scale fully convolutional structure. Both global and local context information are
efficiently synchronized and simultaneously represent the hand features. In addition to
geometric coordinates of the hand bounding box, a ROI region-related score map is
used to form the loss function. A balanced cross-entropy loss is used to facilitate a
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 20–27, 2020.
https://doi.org/10.1007/978-981-15-0187-6_3
