Segmentation of Aerial Image with Multi-scale Feature . . .
65
5 Conclusion
In this work, we combine multi-scale features and attention mechanisms to
achieve aerial image segmentation with high accuracy and robustness. In this
method, multi-scale features are used, and the attention model is combined with
FCN-8s and U-net, respectively, to make the model adaptive, to find the optimal
weight on the scale, and to achieve end-to-end training, which compared with
our FCNs based-line network. The experimental results show that the multiscale features are better than the single-scale segmentation; attention model can
add additional supervision for better model performance by generating a soft
weight at different positions of each scales. Therefore, the segmentation network
model combining with multi-scale features and attention mechanism is applied
to aerial image labeling, which can effectively improve the segmentation effect.
References
1. Lagrange A, Le Saux B, Beaupere A et al (2015) Benchmarking classification of
earth-observation data: from learning explicit features to convolutional networks.
In: Geoscience and remote sensing symposium (IGARSS). IEEE, Milan, pp 4173–
4176
2. Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic
segmentation. In: Proceedings of the IEEE conference on computer vision and
pattern recognition. IEEE, Boston, pp 3431–3440
3. Badrinarayanan V, Kendall A, Cipolla R (2015) Segnet: a deep convolutional encoder-decoder architecture for image segmentation. arXiv preprint
arXiv:1511.00561
4. Chen LC, Papandreou G, Kokkinos I et al (2014) Semantic image segmentation with deep convolutional nets and fully connected CRFs. arXiv preprint
arXiv:1412.7062
5. Chen LC, Papandreou G, Kokkinos I et al (2018) Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected
CRFs. IEEE Trans Pattern Anal Mach Intell 40(4):834–848
6. Chen LC, Papandreou G, Schroff F et al (2017) Rethinking atrous convolution for
semantic image segmentation. arXiv preprint arXiv:1706.05587
7. Arbelaez P, Maire M, Fowlkes C et al (2011) Contour detection and hierarchical
image segmentation. IEEE Trans Pattern Anal Mach Intell 33(5):898–916
8. Pinheiro PH, Collobert R (2013) Recurrent convolutional neural networks for scene
parsing, 2. arXiv preprint arXiv:1306.2795
9. Eigen D, Fergus R (2015) Predicting depth, surface normals and semantic labels
with a common multi-scale convolutional architecture. In: Proceedings of the IEEE
international conference on computer vision. IEEE, Santiago, pp 2650–2658
10. Cao C, Liu X, Yang Y et al (2015) Look and think twice: capturing top-down visual
attention with feedback convolutional neural networks. In: Proceedings of the IEEE
international conference on computer vision. IEEE, Santiago, pp 2956–2964
11. Ba J, Mnih V, Kavukcuoglu K (2014) Multiple object recognition with visual
attention. arXiv preprint arXiv:1412.7755
12. Chen LC, Yang Y, Wang J et al (2016) Attention to scale: scale-aware semantic
image segmentation. In: Proceedings of the IEEE conference on computer vision
and pattern recognition. IEEE, Las Vegas, pp 3640–3649
65
5 Conclusion
In this work, we combine multi-scale features and attention mechanisms to
achieve aerial image segmentation with high accuracy and robustness. In this
method, multi-scale features are used, and the attention model is combined with
FCN-8s and U-net, respectively, to make the model adaptive, to find the optimal
weight on the scale, and to achieve end-to-end training, which compared with
our FCNs based-line network. The experimental results show that the multiscale features are better than the single-scale segmentation; attention model can
add additional supervision for better model performance by generating a soft
weight at different positions of each scales. Therefore, the segmentation network
model combining with multi-scale features and attention mechanism is applied
to aerial image labeling, which can effectively improve the segmentation effect.
References
1. Lagrange A, Le Saux B, Beaupere A et al (2015) Benchmarking classification of
earth-observation data: from learning explicit features to convolutional networks.
In: Geoscience and remote sensing symposium (IGARSS). IEEE, Milan, pp 4173–
4176
2. Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic
segmentation. In: Proceedings of the IEEE conference on computer vision and
pattern recognition. IEEE, Boston, pp 3431–3440
3. Badrinarayanan V, Kendall A, Cipolla R (2015) Segnet: a deep convolutional encoder-decoder architecture for image segmentation. arXiv preprint
arXiv:1511.00561
4. Chen LC, Papandreou G, Kokkinos I et al (2014) Semantic image segmentation with deep convolutional nets and fully connected CRFs. arXiv preprint
arXiv:1412.7062
5. Chen LC, Papandreou G, Kokkinos I et al (2018) Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected
CRFs. IEEE Trans Pattern Anal Mach Intell 40(4):834–848
6. Chen LC, Papandreou G, Schroff F et al (2017) Rethinking atrous convolution for
semantic image segmentation. arXiv preprint arXiv:1706.05587
7. Arbelaez P, Maire M, Fowlkes C et al (2011) Contour detection and hierarchical
image segmentation. IEEE Trans Pattern Anal Mach Intell 33(5):898–916
8. Pinheiro PH, Collobert R (2013) Recurrent convolutional neural networks for scene
parsing, 2. arXiv preprint arXiv:1306.2795
9. Eigen D, Fergus R (2015) Predicting depth, surface normals and semantic labels
with a common multi-scale convolutional architecture. In: Proceedings of the IEEE
international conference on computer vision. IEEE, Santiago, pp 2650–2658
10. Cao C, Liu X, Yang Y et al (2015) Look and think twice: capturing top-down visual
attention with feedback convolutional neural networks. In: Proceedings of the IEEE
international conference on computer vision. IEEE, Santiago, pp 2956–2964
11. Ba J, Mnih V, Kavukcuoglu K (2014) Multiple object recognition with visual
attention. arXiv preprint arXiv:1412.7755
12. Chen LC, Yang Y, Wang J et al (2016) Attention to scale: scale-aware semantic
image segmentation. In: Proceedings of the IEEE conference on computer vision
and pattern recognition. IEEE, Las Vegas, pp 3640–3649
