Segmentation of Aerial Image with Multi-scale Feature . . .
59
– Unbalanced distribution—The building categories are unevenly distributed,
some are densely or sparsely distributed, and there are no buildings in the
image.
– Small size—In airborne imagery, the size of some buildings compared to other
objects in the image is quite small. In some of the case, building just consists
of only a few pixels.
– Shadow—Shadow creates a different illumination over buildings causing
changes in their appearance. This reason, like the occlusion, could reduce the
accuracy of automatic building labeling algorithms, especially deep learning
methods which need a lot of training samples.
– Complex background—Structures such as load resemble with high similar
building labeling.
These challenges raise the level of difficulty when it comes to image segmentation of aerial building image positioning and detection. Even for well-performing
models, there is a high degree of uncertainty in the segmentation results [1]. The
validity of the second part will be affected by the accuracy of image segmentation. Therefore, it is meaningful to use the state-of-the-art model to improve the
accuracy.
2 Related Work
Semantic pixel-wise segmentation is always an active topic of research. In 2014,
Berkeley [2] proposed a fully convolutional neural networks (FCNs) for segmentation. Based on the convolutional neural network, FCN changes the fully connected layer to 1 × 1 convolution layer. The success of FCN for semantic segmentation has more recently led researchers to exploit feature learning capabilities
for segmentation. Badrinarayanan et al. [3] purposed SegNet, a typical encoder–
decoder structure, achieving high scores for road scene understanding which is
efficient both in terms of memory and computational time. A series of networks
named DeepLab also used the convolution layer with dilation. DeepLab V1 [4]
used deep convolutional neural networks (DCNNs) and fully connected conditional random field (fcCRF) to solve the problems. DeepLab V2 [5] added atrous
spatial pyramid pooling (ASPP) based on V1, which enables segmentation on
multiple scales. DeepLab V3 [6] proposed a module that consists of atrous convolution with various rates and batch normalization layers and experimented
with laying out the modules in cascade or in parallel.
It is known that multi-scale feature is useful for computer vision task [7].
Farabet et al. [8] employed a Laplacian pyramid and share-net, passed each
scale through a shared network, and fused the features of all scales. Eigen and
Fergus [9] fed input images of three scales to DCNNs. The DCNNs at different
scales have different structures, and this model required two-step process.
In computer vision, attention models are widely used in image classification
[10] and object detection [11]. Chen et al. [12] proposed a mechanism of attention,
which combine attention model and multi-scale. They learn an attention model
Précédent

- 71/679

Suivant