60
S. Hu et al.
Fig. 1. Introduction of model.
that softly weights the features from different input scales when predicting the
semantic label of a pixel. The final output of the model is produced by the
weighted sum of score map across all the scales.
Here, we concentrate on building marking pixel-wise semantic segmentation
using high-resolution aerial images, and our proposal is based on combing FCNs
with multi-scale and attention model for this segmentation task. We use two
different depths of baseline networks in experiment and combining them with
models of multi-scale features and attention mechanisms as our new segmentation model.
3 Attention Model for Scales
There is a common method, which is share-net, resizes the input image to several
scales, passes through a deep network sharing weights, and then computes the
final prediction based on the fusion of the resulting multi-scale features. As
shown in Fig. 1, we resize the input image to three scales and pass them through
the same network to obtain score maps of different scales (the output of the last
layer before softmax). Lastly, the fusion feature map feeds into attention model
to generate weight map and multiplies the weight map with the fused feature
map to get the final feature map. Herein, the attention model used here allows
us to judge the importance of features at different locations and scales to achieve
better segmentation.
3.1 Attention Mechanism
Attention mechanism developed from human visual research: Different attentions
of the different parts are different, when focusing on a certain target or scene.
Similarly, the most relevant parts of the statement and description change as the
S. Hu et al.
Fig. 1. Introduction of model.
that softly weights the features from different input scales when predicting the
semantic label of a pixel. The final output of the model is produced by the
weighted sum of score map across all the scales.
Here, we concentrate on building marking pixel-wise semantic segmentation
using high-resolution aerial images, and our proposal is based on combing FCNs
with multi-scale and attention model for this segmentation task. We use two
different depths of baseline networks in experiment and combining them with
models of multi-scale features and attention mechanisms as our new segmentation model.
3 Attention Model for Scales
There is a common method, which is share-net, resizes the input image to several
scales, passes through a deep network sharing weights, and then computes the
final prediction based on the fusion of the resulting multi-scale features. As
shown in Fig. 1, we resize the input image to three scales and pass them through
the same network to obtain score maps of different scales (the output of the last
layer before softmax). Lastly, the fusion feature map feeds into attention model
to generate weight map and multiplies the weight map with the fused feature
map to get the final feature map. Herein, the attention model used here allows
us to judge the importance of features at different locations and scales to achieve
better segmentation.
3.1 Attention Mechanism
Attention mechanism developed from human visual research: Different attentions
of the different parts are different, when focusing on a certain target or scene.
Similarly, the most relevant parts of the statement and description change as the
