62
S. Hu et al.
Fig. 2. Attention model implementation process left: merge feature map right: generate
weight map.
the merged score maps across all the scales. To the optimization problem of
finding, the minimum value in the BCE loss is solved by Adam optimizer and
back-propagation process. At the same time, the experimental data is divided
into three different proportions to obverse the effect of convergence speed. Adam
with mini-batch is used for training. We set the initial learning rate of 0.001, and
epoch is 100. The learning rate is multiplied by 0.1 after 2000 iterations. We use
a momentum of 0.9 and weight decay of 0.0005. Due to hardware limitation, the
mini-batch size id set to 5 when training the basic network and 2 when training
the network with attention model.
4.2 Network Architectures
Our network is based on the publicly available model: FCN-8s [2] with VGG16
as a backbone network and U-net [14]. They all have proven effective in semantic segmentation. FCN-8s use all of the pre-trained convolutional layer weights
from VGG-16 as pre-trained weights; for U-net network parameters, Gaussian
initialization is used to initialize parameters.
4.3 Datasets
The experiments were conducted using images acquired by the Inria Aerial Image
Labeling Dataset, which covers multiple urban areas from densely populated
areas to alpine towns, with high spatial resolution. Therefore, in aerial image
labeling, the goal is to classify each pixel as building marking class (foreground)
or non-building marking (background). The training set includes 5 regions, each
region contains 36 tiles, numbered 1–36. According to the datasets, we remove
the first five images of every location from the training set. To submit results,
Précédent

- 74/679

Suivant