118
V. Tra and J.-M. Kim
in the frequency domain, including frequency center (FC), RMS frequency (RMSF),
root variance frequency (RVF), kurtosis frequency (KF), entropy frequency (EF),
sum square frequency (SSF) [2]; and 16 features which are the relative energy in a
wavelet packet node (REWPN) and the entropy in a wavelet packet node (EWPN)
[7]. However, because the number of features in a pool is quite large and redundant,
it is essential to have a proper feature selection technique so that we can withdraw
redundant features but still retain informative features.
11.2.2 GA-Based Fault Feature Selection
The genetic algorithm (GA), which is the hypothesis of Darwinian about natural
selection and genetic nature in biological systems, has been a notable algorithm in
generating optimal solutions for search problems. Determining the most optimal
features among numerous options is a representation of search problems, in which
a set of potential features (i.e., a feature vector) is considered as an individual in a
population. In order to apply the genetic algorithm in our feature-selection scheme,
each candidate fault-feature in the feature-pool is treated as a gene in a chromosome.
The value of a gene is 1 if the corresponding feature is considered to be useful and to
need to be conserved. In contrast, if a specific feature is determined redundant, the
corresponding gene is assigned a value of 0. In order to search exhaustively potential
solutions, a series of iterative generational computations are implemented. For each
iteration, new offspring is randomly generated by the process of crossover or mutation
to create the successive population. Since the genetic algorithm utilizes an objective
function or fitness function to evaluate the quality of candidate solutions, so it is
crucial to design a fitness function that matches well to the problem. In this study,
the Gaussian radial basis function (RBF), which is able to measure the similarity of
two input samples, is used as a kernel of the GA-based feature-selection scheme.
k(x, z) = exp(−γ ||x − z||),
(11.1)
where γ = 1
2σ
2 and σ is an adjustable parameter. This parameter is selected via
experimental tests and is carefully tuned into the scheme. From (11.1), it makes sense
that the within-category RBF value and between-category RBF value of two input
samples x and z belonging to the same category are expected as (11.2) and (11.3),
respectively
k(x, z) ≈ 1, ∀x, z ∈ C i , ∀i = 1, 2, . . . , L ,
(11.2)
k(x, z) ≈ 0, ∀x ∈ C i , ∀z ∈ C j , ∀i, j = 1, 2, . . . , L , i = j,
(11.3)
where C i is a set of samples in the category i, i = 1, 2, . . . , L , and L is the number
of categories. Figures 11.2 and 11.3 illustrate the within-category RBF values and
Précédent

- 136/567

Suivant