AP ¼
TP
TP þ FP
ð5Þ
AR ¼
TP
TP þ FN
ð6Þ
F ¼
2 Ã AP Ã AR
AP þ AR
ð7Þ
The proposed approach is compared with some state-of-the-art methods [9, 10, 13,
14] on VIVA dataset. The results are shown in Table 1. Our detection is faster than the
other methods, the AP value is the highest, and the AR value is the second best.
5 Conclusions
A hand detector from full images using a single neural network is proposed in this
paper. By incorporating proper loss functions and pixel-wise scoring, the proposed
model can detect human hand under different conditions, such as occlusions, varied
illuminations, varied hand pose, shape, and size. The proposed model is simple,
accurate, and efficient, which has been proved by the experimental results on challenging VIVA dataset.
References
1. The Vision for Intelligent Vehicles and Applications (VIVA) Challenge, Laboratory for
Intelligent and Safe Automobiles, UCSD. http://cvrr.ucsd.edu/vivachallenge/
2. Mittal A, Zisserman A et al (2011) Hand detection using multiple proposals. In: Hoey J,
McKenna S, Trucco E (eds) Proceeding of the British machine vision conference, vol 75.
BMVA Press, pp 1–11
3. Ohn-Bar E, Martin S et al (2014) Head, eye, and hand patterns for driver activity
recognition. In: The 22nd international conference on pattern recognition (ICPR). IEEE
Press, Stockholm, pp 660–665
4. Li C, Kitani K (2013) Pixel-level hand detection in ego-centric videos. In: Proceedings of
computer vision and pattern recognition (CVPR). IEEE Press, Portland, pp 3570–3577
Table 1. Quantitative comparisons of different methods on VIVA dataset
Methods
AP (%)
AR (%)
F
FPS
Ours
98.3
86.7
92.1
42
[9]
94.8
74.7
83.6
4.7
[10]
93.5
91.4
92.4
5
[13]
73.3
69.9
71.6
–
[14]
65.1
47.1
54.7
–
26
Y. Li et al.
TP
TP þ FP
ð5Þ
AR ¼
TP
TP þ FN
ð6Þ
F ¼
2 Ã AP Ã AR
AP þ AR
ð7Þ
The proposed approach is compared with some state-of-the-art methods [9, 10, 13,
14] on VIVA dataset. The results are shown in Table 1. Our detection is faster than the
other methods, the AP value is the highest, and the AR value is the second best.
5 Conclusions
A hand detector from full images using a single neural network is proposed in this
paper. By incorporating proper loss functions and pixel-wise scoring, the proposed
model can detect human hand under different conditions, such as occlusions, varied
illuminations, varied hand pose, shape, and size. The proposed model is simple,
accurate, and efficient, which has been proved by the experimental results on challenging VIVA dataset.
References
1. The Vision for Intelligent Vehicles and Applications (VIVA) Challenge, Laboratory for
Intelligent and Safe Automobiles, UCSD. http://cvrr.ucsd.edu/vivachallenge/
2. Mittal A, Zisserman A et al (2011) Hand detection using multiple proposals. In: Hoey J,
McKenna S, Trucco E (eds) Proceeding of the British machine vision conference, vol 75.
BMVA Press, pp 1–11
3. Ohn-Bar E, Martin S et al (2014) Head, eye, and hand patterns for driver activity
recognition. In: The 22nd international conference on pattern recognition (ICPR). IEEE
Press, Stockholm, pp 660–665
4. Li C, Kitani K (2013) Pixel-level hand detection in ego-centric videos. In: Proceedings of
computer vision and pattern recognition (CVPR). IEEE Press, Portland, pp 3570–3577
Table 1. Quantitative comparisons of different methods on VIVA dataset
Methods
AP (%)
AR (%)
F
FPS
Ours
98.3
86.7
92.1
42
[9]
94.8
74.7
83.6
4.7
[10]
93.5
91.4
92.4
5
[13]
73.3
69.9
71.6
–
[14]
65.1
47.1
54.7
–
26
Y. Li et al.
