Independent Component Analysis
113
As mentioned earlier, when some of the resulting components are nonGaussian in nature, kurtosis can be used to identify the non-Gaussian independent components from the ICA solution. Since a linear mixture of independent Gaussian components is also Gaussian, its kurtosis will be zero. Selecting
the components with non-zero kurtosis values is equivalent to identifying the
non-Gaussian independent components.
If the assumptions for ICA are satisfied, a relatively unique ICA solution can
be obtained. The term "relative" indicates that multiplying a component with
a scalar or permuting the order of the components preserves independence.
Therefore, the solution is, in fact, identifiable up to scaling and multiplication
(Hyvarinen et al. 2001; Lee 1998).
In practice, it is unlikely that the observed data are characterized as a linear
mixture of components that are independent in that they satisfy (4.2). This may
be due to the fact that the components are not independent, that there is noise
or interference present and that the data available are of discrete nature. Even in
these situations, the ICA algorithm provides a solution where the components
are "as independent as possible", i. e., dependence among them is as small as
possible.
4.3
ICA Algorithms
We proceed to solve the ICA problem stated in the previous section. The first
step is to decorrelate and sphere the observed data, thereby eliminating the
first and second order statistics (i. e. mean and standard deviation). The next
step is to determine the matrix A such that its inverse allows the recovery of
the independent components. A is usually obtained through gradient based
techniques that optimize a certain cost function c(·):
dc(A,x)
L1A =
.
dA
(4.4)
The cost function is designed such that its optimization corresponds to
achieving independence. One direct approach for the solution of the ICA
problem is to consider c(·) as the mutual information of the components of
A -1 x. This method is described in detail in Sect. 4.3.2. A second approach is
based on the observation that projecting the data x into components, which are
decorrelated and are as non-Gaussian as possible, also leads to independence
(see Sect. 4.3.3) Both methods can be shown to be equivalent and choice of
one over the other, or modification of either of them depends on the specific
problem or the preference of the user.
4.3.1
Preprocessing using PCA
Whitening (decor relation and sphering) of the data is usually performed before
the application of I CA as a mean to reduce the effect of first and second order
113
As mentioned earlier, when some of the resulting components are nonGaussian in nature, kurtosis can be used to identify the non-Gaussian independent components from the ICA solution. Since a linear mixture of independent Gaussian components is also Gaussian, its kurtosis will be zero. Selecting
the components with non-zero kurtosis values is equivalent to identifying the
non-Gaussian independent components.
If the assumptions for ICA are satisfied, a relatively unique ICA solution can
be obtained. The term "relative" indicates that multiplying a component with
a scalar or permuting the order of the components preserves independence.
Therefore, the solution is, in fact, identifiable up to scaling and multiplication
(Hyvarinen et al. 2001; Lee 1998).
In practice, it is unlikely that the observed data are characterized as a linear
mixture of components that are independent in that they satisfy (4.2). This may
be due to the fact that the components are not independent, that there is noise
or interference present and that the data available are of discrete nature. Even in
these situations, the ICA algorithm provides a solution where the components
are "as independent as possible", i. e., dependence among them is as small as
possible.
4.3
ICA Algorithms
We proceed to solve the ICA problem stated in the previous section. The first
step is to decorrelate and sphere the observed data, thereby eliminating the
first and second order statistics (i. e. mean and standard deviation). The next
step is to determine the matrix A such that its inverse allows the recovery of
the independent components. A is usually obtained through gradient based
techniques that optimize a certain cost function c(·):
dc(A,x)
L1A =
.
dA
(4.4)
The cost function is designed such that its optimization corresponds to
achieving independence. One direct approach for the solution of the ICA
problem is to consider c(·) as the mutual information of the components of
A -1 x. This method is described in detail in Sect. 4.3.2. A second approach is
based on the observation that projecting the data x into components, which are
decorrelated and are as non-Gaussian as possible, also leads to independence
(see Sect. 4.3.3) Both methods can be shown to be equivalent and choice of
one over the other, or modification of either of them depends on the specific
problem or the preference of the user.
4.3.1
Preprocessing using PCA
Whitening (decor relation and sphering) of the data is usually performed before
the application of I CA as a mean to reduce the effect of first and second order
