Segmentation of Aerial Image with Multi-scale Feature . . .
61
description changes. There are two types of attention methods: soft attention
and hard attention, and its output vector distribution is soft and another one is
hot, which will affect the selection of context directly. So attention model can
improve network performance by focusing on the most relevant features as need.
It is different from the attention model applied in two-dimensional and time [13],
and attention model combined with multi-scale features is applied to semantic
segmentation to improve the performance.
3.2 Attention Model for Scales
Herein, the attention model used is based on multi-scale features, which learns to
softly weights for each scale and pixel. And this attention model is differentiable,
so it trains end-to-end. As shown in Fig. 2, suppose an input image is resized
to several scales s ∈ {1, . . . , S}. Each scale is passed through the FCN (weights
are shared across all scales) and produces a score map f
s
i,c for scale, where i
ranges over all the spatial positions, and c ∈ {1, . . . , C} where C is the number
of classes of interest. The score map f
s
i,c is resized to the same resolution by
bilinear interpolation. g i,c is the weights sum of the score maps at (i, c) for all
scales, so
g i,c =
s
s=1
ω
s
i • f
s
i,c
(1)
The weight ω
s
i is computed by
ω
s
i =
exp(h
s
i )
s
t=1 exp(h t
i )
(2)
where h
s
i is the score map produced by attention model in position i for scale s.
The proposed attention model consists of two layers: the first layer is convolution
operation that has 512 filters with kernel size 3 × 3; the second layer has S filter
with kernel size 1 × 1, where S is the number of scales employed. The weight
ω
s
i produced by attention model reflects the importance of feature at position
i for scale s. Note in formulation 1, the average-pooling and max-pooling are
two special cases: the weights ω
s
i will be replaced by 1/S for average-pooling;
while the summation becomes the max operation and ω
s
i = 1∀s, i in the case
of max-pooling. The attention model computes a soft weight for each scale and
position, and it allows the gradient to be backpropagated through. Therefore,
attention model can be trained with FCN part end-to-end and implementing the
model adaptively to find the best weights on scales.
4 Experiments
4.1 Extra Supervision
In our experiments, we learn the network parameters by comparing the final
output of the model with the corresponding ground truth for each image at the
pixel-level. The final output is produced by performing a softmax operation on
61
description changes. There are two types of attention methods: soft attention
and hard attention, and its output vector distribution is soft and another one is
hot, which will affect the selection of context directly. So attention model can
improve network performance by focusing on the most relevant features as need.
It is different from the attention model applied in two-dimensional and time [13],
and attention model combined with multi-scale features is applied to semantic
segmentation to improve the performance.
3.2 Attention Model for Scales
Herein, the attention model used is based on multi-scale features, which learns to
softly weights for each scale and pixel. And this attention model is differentiable,
so it trains end-to-end. As shown in Fig. 2, suppose an input image is resized
to several scales s ∈ {1, . . . , S}. Each scale is passed through the FCN (weights
are shared across all scales) and produces a score map f
s
i,c for scale, where i
ranges over all the spatial positions, and c ∈ {1, . . . , C} where C is the number
of classes of interest. The score map f
s
i,c is resized to the same resolution by
bilinear interpolation. g i,c is the weights sum of the score maps at (i, c) for all
scales, so
g i,c =
s
s=1
ω
s
i • f
s
i,c
(1)
The weight ω
s
i is computed by
ω
s
i =
exp(h
s
i )
s
t=1 exp(h t
i )
(2)
where h
s
i is the score map produced by attention model in position i for scale s.
The proposed attention model consists of two layers: the first layer is convolution
operation that has 512 filters with kernel size 3 × 3; the second layer has S filter
with kernel size 1 × 1, where S is the number of scales employed. The weight
ω
s
i produced by attention model reflects the importance of feature at position
i for scale s. Note in formulation 1, the average-pooling and max-pooling are
two special cases: the weights ω
s
i will be replaced by 1/S for average-pooling;
while the summation becomes the max operation and ω
s
i = 1∀s, i in the case
of max-pooling. The attention model computes a soft weight for each scale and
position, and it allows the gradient to be backpropagated through. Therefore,
attention model can be trained with FCN part end-to-end and implementing the
model adaptively to find the best weights on scales.
4 Experiments
4.1 Extra Supervision
In our experiments, we learn the network parameters by comparing the final
output of the model with the corresponding ground truth for each image at the
pixel-level. The final output is produced by performing a softmax operation on
