Segmentation of Aerial Image with Multi-scale Feature . . .
63
use the exact same file names as the input color images, and output 0/255
8-bit single-channel files in the validation set. The training set and validation
set are divided into three different proportions for training, 31:5, 26:5, and 21:5.
Different proportions of datasets can affect model convergence. Since the dataset
is a high-resolution image, the image is cut into 512 × 512 size. When using the
attention model for training, three scales are required and resize the input images
to add extra scales: 256 × 256, 1024 × 1024.
4.4 Evaluation
Concerning the evaluation criteria, we use the Intersection over Union (IoU) of
positive (building) class, which are widely used in semantic segmentation task.
In these metrics, n ij is the pixel number belonging to class i which has been
predicted as class j and n cl stand for the number of classed with t i =
j n ij
representing the total number of pixels belonging to class i. We use the dice
similarity coefficient also due to the heavy unbalance in the dataset. The number
of pixels belonging to each class does have an effect on these two criteria. X and
Y represent prediction and ground truth, respectively. The criteria are derived
as follows:
IoU of positive class:
IoU =
area (C) ∩ area (G)
area (C) ∪ area (G)
(3)
Dice similarity coefficient:
1
n cl
i
n ii
t i +
j n ji − n ii
(4)
4.5 Results
From Table 1, for the dataset with ratio of 31:5, the network combines multi-scale
features and attention mechanism. The IoU and dice coefficients of the deeper
neural network U-net are higher (IoU is 0.11 higher and dice coefficient is 0.9
higher), due to deeper networks which are able to extract higher-level features.
By comparing FCN-32s and U-net combining multi-scale features and attention
mechanisms respectively, found that our model, IoU and Dice coefficient are
improved, and the deeper network U-net segmentation effect is better (IoU is
0.784 Dice coefficient is 0.879). The results show that the segmentation performance of semantic segmentation model combined with multi-scale features and
attention mechanism is improved.
The result of the segmentation is shown in Fig. 3. All images are the same
size: 512 × 512, from top to bottom: aerial image, basic network result, the combined attention model result, and ground truth; from left to right, the training
results of the model in the dataset at different proportion (the ratio is from small
to large). The segmentation results prove that the network with multi-scale and
Précédent

- 75/679

Suivant