11 Improving Bearing Diagnostic Performance …
121
11.2.3.2 Determining the Most Discriminative Fault Features Using
GA-Based Feature Analysis and k-NN-Based Classifier
The quality of candidate feature-sets is assessed by using the GA-based featureselection scheme as described in Sect. 2.2. In order to reduce the variability as
well as increase the robustness of the optimal solution determined by the scheme,
the study adopts k-fold cross-validation (k-cv) to the original dataset (i.e., k = 2).
This means the dataset is divided into two parts or sub-datasets. One sub-dataset is
used to determine the most discriminatory feature-set or the optimal solution and
another sub-dataset is utilized to validate the fitness of the solution via calculating
the classification accuracy of the k-NN classifier. This process is repeated until each
sub-dataset is reserved as either the training dataset or the test dataset at least once.
Correspondingly, k candidate feature-sets are yielded for each iteration.
To search for exhaustly possible solutions, we repeat the above procedure
N interations times (i.e., N interations = 15). It means after N interations , the total of
N interations × k(i.e., 30) optimal feature-sets are determined. In order to pick up
the best among them, a decision rule is proposed in which the decision is made based
on 2 criteria. The first one is the estimated diagnostic performance of a solution and
the second is the occurring frequency of the candidates. The most optimal featureset is the one that helps the k-NN classifier achieve the highest performance and
improve the model’s robustness. In this study, the estimated classification accuracy,
or predicted classification accuracy (CA), is determined as follows:
C A =
L N T P
N samples
× 100(%)
(11.7)
where L is the number of categories, N T P is the number of samples in the test dataset
belonging to the category i that are exactly classified to be the category i, and N samples
is the total number of samples in the test dataset.
11.3 Experimental Results and Discussion
11.3.1 Training Dataset and Test Dataset Configuration
This study uses datasets made of acoustic emission (AE) signals that are obtained in
the conditions of a free-defect (BFD) and incipient single defects at different bearing
elements i.e., a crack on the outer raceway (BCO), crack on the inner raceway (BCI),
and crack on the roller (BCR). Two datasets, according to two different rotating
speeds (i.e., 300 revolutions per minute (r/min) and 500 (r/m)), are used to validate
the effectiveness of the proposed method. The total number of samples in each dataset
is N Classes × N Signals or 360 samples, where N Classes is the number of defect types
(N Classes = 4), and N Signals is the number of AE signal segments collected under
Précédent

- 139/567

Suivant