Global Descriptors of Convolution Neural Networks . . .
5
3.2 Experimental Introduction
For input data, 75% images in each class serve as the training set and the
remaining serve as the testing set. In experiments, the global feature descriptors are extracted from two pre-trained CNN models consisting of VGGNet-16
and ResNet-50, which are trained on the ImageNet dataset. The dimensions of
global feature descriptors are 1 × 1000. The random forest is applied for training
and classification. The confusion matrices performed the result of our proposed
method; the training time is analyzed on the different methods to train.
3.3 Discuss Different CNN Models with PCA
The comparison between two types of CNN-based features classification is shown
in Tables 1 and 2. The result of the classification on VGGNet-16-based and
ResNet-50-based features without PCA transformation is displayed in Table 1
and the classification details of two types of CNN-based features with PCA
transformation are performed in Table 2. The ratio of PCA transformation is set
as 95%.The comparison in tables demonstrates that VGGNet-16 performs better
classification effect. VGGNet-16-based features are better discrimination, especially in some similar categories, such as “beach,” “parking lot,” and “runway”,
which is because the initial parameters in VGGNet-16 are much richer.
From two tables, we can see that using PCA transformation performs greatly
for both types pre-trained CNN models, especially for the ResNet-50, which
has improved about 8%. In general, PCA transformation help to improve the
description ability of the global features.
3.4 Analysis of the Proposed Method
In this section, we research the confusion matrix on the 21-classes public remote
dataset. Figure 2 describes the confusion matrix of proposed method, which
includes feature fusion and PCA transformation. The confusion matrix is analyzed via using 25% of the training dataset. As confusion matrix is shown, the
entry in the ith row and jth column means the rate of RSI belongs to the ith class
and classifies to jth class. The classification results are proposed as percentages.
In Fig. 2, the average accuracies of classification for proposed method are
86.78%. It performs great ability (the accuracy of classification ≥90%) on the
classes, such as “airplane,” “agricultural,” “baseballdiamond,” “chaparral,” “forest,” “golf course,” “harbor,” “overpass,” “tennis court,” “river.” The features
extracted from these RSI include considerable information, which helps classify
correctly.
Moreover, some classes obtain poor classification effect, such as “dense residential,” “medium residential.” This is the reason that high dimensional features
of these classes are too similar to distinguish. In addition, several classes include
plenty of building elements, which results an error decision.
Table 3 presents the comparison between state-of-art methods and our proposed method. These existing methods are detailed in [8,12–14]. As we can see
5
3.2 Experimental Introduction
For input data, 75% images in each class serve as the training set and the
remaining serve as the testing set. In experiments, the global feature descriptors are extracted from two pre-trained CNN models consisting of VGGNet-16
and ResNet-50, which are trained on the ImageNet dataset. The dimensions of
global feature descriptors are 1 × 1000. The random forest is applied for training
and classification. The confusion matrices performed the result of our proposed
method; the training time is analyzed on the different methods to train.
3.3 Discuss Different CNN Models with PCA
The comparison between two types of CNN-based features classification is shown
in Tables 1 and 2. The result of the classification on VGGNet-16-based and
ResNet-50-based features without PCA transformation is displayed in Table 1
and the classification details of two types of CNN-based features with PCA
transformation are performed in Table 2. The ratio of PCA transformation is set
as 95%.The comparison in tables demonstrates that VGGNet-16 performs better
classification effect. VGGNet-16-based features are better discrimination, especially in some similar categories, such as “beach,” “parking lot,” and “runway”,
which is because the initial parameters in VGGNet-16 are much richer.
From two tables, we can see that using PCA transformation performs greatly
for both types pre-trained CNN models, especially for the ResNet-50, which
has improved about 8%. In general, PCA transformation help to improve the
description ability of the global features.
3.4 Analysis of the Proposed Method
In this section, we research the confusion matrix on the 21-classes public remote
dataset. Figure 2 describes the confusion matrix of proposed method, which
includes feature fusion and PCA transformation. The confusion matrix is analyzed via using 25% of the training dataset. As confusion matrix is shown, the
entry in the ith row and jth column means the rate of RSI belongs to the ith class
and classifies to jth class. The classification results are proposed as percentages.
In Fig. 2, the average accuracies of classification for proposed method are
86.78%. It performs great ability (the accuracy of classification ≥90%) on the
classes, such as “airplane,” “agricultural,” “baseballdiamond,” “chaparral,” “forest,” “golf course,” “harbor,” “overpass,” “tennis court,” “river.” The features
extracted from these RSI include considerable information, which helps classify
correctly.
Moreover, some classes obtain poor classification effect, such as “dense residential,” “medium residential.” This is the reason that high dimensional features
of these classes are too similar to distinguish. In addition, several classes include
plenty of building elements, which results an error decision.
Table 3 presents the comparison between state-of-art methods and our proposed method. These existing methods are detailed in [8,12–14]. As we can see
