Feature Extraction from Hyperspectral Data Using ICA
201
may not be error-free, so the quality of the training data may be poor in that
it may not accurately characterize the classes. This would decrease the quality
of both feature extraction and classification (Swain and King 1973).
Alternatively, when no prior information about the classes is available,
one can perform unsupervised feature extraction. In this case, the statistics
and distance measures between classes cannot be computed or estimated and
the main goal of feature extraction shifts to reduction of data redundancy.
The narrow bandwidths associated with the spectral bands of hyperspectral
data lead to correlation between the adjacent bands resulting in a relatively
high level of redundancy in the data (Richards and Jia 1999). Based on this
observation, one can simply proceed to perform feature selection by analyzing
the correlation matrix and selecting only a few bands from each group of
highly correlated bands. A better approach is to transform the data such
that the resulting features are decorrelated. The variance of the individual
components is considered to be an indicator of information content; large
values suggest high levels of information and low values indicate the presence
of mostly noise. Based on this, only the features with high variance are selected
for further processing (Richards and Jia 1999).
Both PCA and ICA are multivariate methods that, given a random vector,
proceed in such a way that the resulting components have increased class
separability. In case of PCA, the separability is achieved through decorrelation
whereas in ICA, it is via independence (Lee 1998).
When considering hyperspectral data, each band corresponds to a component of random pixel vectors and constitutes a feature of the data. In this
context, the image cube resulting after PCA processing has the bands decorrelated and sorted according to their variance. This indicates that most of the
information can be retrieved from the first few features and that most of the last
features have close to zero variance (indicating a lack of information). Based
on this observation, PCA is frequently used for feature extraction by dropping
the lowest variance components (Richards and Jia 1999).
Decorrelation based feature extraction has been observed to be less efficient
when dealing with small classes (i.e. classes having small spatial extents) in
the image. In other words, small classes are sparsely represented in the data
(i.e. small classes contain very few pixels). Due to their size, these classes tend
to have little influence on the band variance leading to the possibility of being
discarded in the lower variance bands. In the context of target detection, loss
of information regarding small targets (that correspond to small classes in the
image) affects the accuracy of the feature extraction algorithms (Achalakul
and Taylor 2000; Tu et al. 2001). Therefore, there is no guarantee that band
reduction using PCA will correctly preserve the entire information content of
the image cube (Richards and Jia 1999; Tu et al. 2001).
The PCA model assumes that each original band is a linear mixture of
unknown uncorrelated bands, and proceeds to recover them such that their
variance is maximized. In ICA, the assumption is that each original band is,
a linear mixture of unknown independent bands (Robila et al. 2000). The goal
is to find the unmixing matrix and to perform the inversion, in order to recover
the independent bands (Robila and Varshney 2002). In the case of ICA based
Précédent

- 208/327

Suivant