124
V. Tra and J.-M. Kim
which n optimal fault-features are determined according to n highest eigenvalues of
a covariance matrix defined by the original feature space.
Table 11.3 shows the average classification accuracy of the k-NN classifier as its
input is feature-sets selected by different feature-analysis schemes. To increase the
reliability of the experimental results, we utilize k-cv scheme with k = 3 which means
the initial dataset is divided into three folds. For each iteration, one of the three folds
is reserved as the test dataset and the other two folds are exploited as the training
dataset. The corresponding predicted classification accuracy is determined by using
(11.7). Repeating the process until each fold is reserved as either the training dataset
or the test dataset at least once. The average classification performance (ACA) of
the model is determined by averaging three predicted accuracies according to three
iterations in the process.
From Table 11.3, it is obvious that for both the datasets, the average classification
accuracy of the classifier associated with our new feature-selection scheme is higher
than the figures for two rivals. More specifically, in terms of dataset 1, the feature-set
yielded by the proposed feature-selection scheme helps the k-NN classifier achieve
the highest accuracy percentage, at 98.33%, much higher the figure of 77.7% for PCA,
and the figure of 80.49% for ICA. The results are also similar in terms of dataset
2. While the figure for the proposed model is 97.22%, the accuracy percentages of
those utilizing PCA and ICA are lower, at about 82%.
The reason behind this result is that although component analysis-based methods
have shown their effectiveness in many areas, there is no rule to determine the ideal
number of principal components. The wrong decision would deteriorate the quality
of a feature-set. Also, due to the inability of measuring inter-category separability,
the feature-sets yielded by two unsupervised approaches PCA and ICA are unable
to represent samples in different categories separately. As a result, the models using
these feature-sets have lower classification accuracy than those using the feature-set
produced by supervised methods such as the proposed GA-based feature-selection
scheme.
Table 11.3 Average
classification accuracies of
proposed model and other
rival models
Datasets
Methodologies
ACA (%)
Dataset 1
PCA
77.7
ICA
80.4
Proposed method
98.3
Dataset 2
PCA
81.3
ICA
82.4
Proposed method
97.2
Précédent

- 142/567

Suivant