122
V. Tra and J.-M. Kim
Table 11.1 Description of the datasets for two different crack sizes
Dataset with single and
compound seeded
bearing failures
Revolutions per minute
(r/min)
Crack size
Length (mm) Width (mm) Depth (mm)
Dataset 1 Training set 300 RPM
3
0.35
0.30
Test set
300 RPM
0.35
0.30
Dataset 2 Training set 500 RPM
3
0.35
0.30
Test set
500 RPM
0.35
0.30
each bearing condition (N Signals = 90). Each segmented signal is 5 s in length. The
datasets are presented in detail as given in Table 11.1.
11.3.2 Most Optimal Fault-Feature Sets
To select the most optimal feature-set, we first adopt the GA-based feature-selection
scheme to find potential feature-sets. By using the k-cv with k is 2 and run the
scheme 15 times (i.e., N interations = 15 and k = 2 for k-cv), the total of 30 featuresets are yielded. Following this, the decision rule is applied via two steps. First, the
classification accuracies of the k-NN classifier, according to different feature-sets,
are estimated using 11.7). Second, the occurring frequency of each feature-set in the
30 original solutions is considered to decide which feature-set will be chosen as the
best solution. Figure 11.4 illustrates the predictive accuracies of the k-NN classifier,
according to the 30 candidate solutions searched by the proposed feature-selection
scheme.
Overall, it is clear that the k-NN classifier’s performance is heavily dependent on
the choice of an input feature-set. Also, the model’s robustness is improved with some
feature-sets and vice versa. More specifically, as using dataset 1 for validation, the
k-NN classifier achieves the highest predicted accuracy values with the input featureset of f 5 , f 7 , and f 10 , at approximately 98%. This feature-set also helps the model
to be robust as its classification performance is stable over different testing times.
Whereas, the predictive capacity of the model tends to be deteriorated as using some
other input-feature sets. The reason behind this result is that the ability to represent
samples of different classes in the form of separate clusters depends on the choice
of an input feature-set. In terms of dataset 2, the k-NN classifier’s classification
accuracies associated with different input feature-sets also shows a similar trend.
The occurring frequency of the feature-set of f 2 , f 10 , and f 13 is 4 and its according
predicted accuracy is about 95%, relatively higher than most of the counterparts’
figures. For the evidence described above, it is obvious that the optimal feature-set
has a significant impact on the classifier’s performance. Therefore, selecting the best
solution among potential candidates must be the priority of the model’s designers.
V. Tra and J.-M. Kim
Table 11.1 Description of the datasets for two different crack sizes
Dataset with single and
compound seeded
bearing failures
Revolutions per minute
(r/min)
Crack size
Length (mm) Width (mm) Depth (mm)
Dataset 1 Training set 300 RPM
3
0.35
0.30
Test set
300 RPM
0.35
0.30
Dataset 2 Training set 500 RPM
3
0.35
0.30
Test set
500 RPM
0.35
0.30
each bearing condition (N Signals = 90). Each segmented signal is 5 s in length. The
datasets are presented in detail as given in Table 11.1.
11.3.2 Most Optimal Fault-Feature Sets
To select the most optimal feature-set, we first adopt the GA-based feature-selection
scheme to find potential feature-sets. By using the k-cv with k is 2 and run the
scheme 15 times (i.e., N interations = 15 and k = 2 for k-cv), the total of 30 featuresets are yielded. Following this, the decision rule is applied via two steps. First, the
classification accuracies of the k-NN classifier, according to different feature-sets,
are estimated using 11.7). Second, the occurring frequency of each feature-set in the
30 original solutions is considered to decide which feature-set will be chosen as the
best solution. Figure 11.4 illustrates the predictive accuracies of the k-NN classifier,
according to the 30 candidate solutions searched by the proposed feature-selection
scheme.
Overall, it is clear that the k-NN classifier’s performance is heavily dependent on
the choice of an input feature-set. Also, the model’s robustness is improved with some
feature-sets and vice versa. More specifically, as using dataset 1 for validation, the
k-NN classifier achieves the highest predicted accuracy values with the input featureset of f 5 , f 7 , and f 10 , at approximately 98%. This feature-set also helps the model
to be robust as its classification performance is stable over different testing times.
Whereas, the predictive capacity of the model tends to be deteriorated as using some
other input-feature sets. The reason behind this result is that the ability to represent
samples of different classes in the form of separate clusters depends on the choice
of an input feature-set. In terms of dataset 2, the k-NN classifier’s classification
accuracies associated with different input feature-sets also shows a similar trend.
The occurring frequency of the feature-set of f 2 , f 10 , and f 13 is 4 and its according
predicted accuracy is about 95%, relatively higher than most of the counterparts’
figures. For the evidence described above, it is obvious that the optimal feature-set
has a significant impact on the classifier’s performance. Therefore, selecting the best
solution among potential candidates must be the priority of the model’s designers.
