4 Experiments and Results
4.1 Dataset
The dataset is from the VIVA Challenge, which consists of 54 videos from seven
viewpoints with annotations of hands of drivers and passengers using 2D bounding
boxes. These videos were collected in naturalistic driving environments with illumination changes, cluttered backgrounds, quick hand motions, and common occlusion.
All the images are divided into 5500 training and 5500 testing data for standard
evaluation.
4.2 Data Augmentation
Rich data is better for training a good model in deep neural networks. To get more
training data, we did data augmentation using rotation, translation, Gaussian blurring,
and sharping operations.
– Augmentation rule 1: The ratio of brightness enhancement is (1.2–1.5), the scaling
factor is (0.7–1.5), and the upper limit of the translation is 40 pixels along x axis and
60 pixels along y axis.
– Augmentation rule 2: Boundary clipping range is (0–16) pixels. Randomly select
50% images and did horizontally flip.
– Augmentation rule 3: Vertically flip the images with added Gaussian blurring. The
kernel size is (0–3.0).
– Augmentation rule 4: Rotate the images randomly. The largest rotation angular is
45°. Add Gaussian white noise with standard deviation 0.2. Randomly select 50%
images and do the sharping process.
After data augmentation, the total amount of training data is 22,000. They are split
into training subset and validation subset. The ratio is 9:1.
4.3 Label Generation
Since the score of each pixel is computed in the output part, the original labels in the
bounding boxes should be further processed to fulfill the requirements. The original
bounding boxes are zoomed out lightly to get much tighter object regions. The labels
are set to be 1 for those pixels within the zoomed bounding boxes and 0 for other
pixels.
4.4 Evaluation Criteria
The proposed model is evaluated on VIVA dataset using several criteria.
The visual inspections are listed in Fig. 2. It shows good hand detection results in
some typical instances, such as varied illuminations, different hand shapes and sizes,
and different hand numbers in one image.
24
Y. Li et al.
Précédent

- 36/679

Suivant