Independent Component Analysis
115
minimize the dependence between the components of the data. In the case of
PCA, dependence minimization is achieved when the covariance between the
components is zero, whereas ICA considers the components to be independent
when the probability density function of the vector can be expressed as the
product of the probability density functions of the components (see (4.2)).
When the data is Gaussian, decorrelation is equivalent to independence, and
thus PCA can be considered to be equivalent to ICA. However, when the data
is non-Gaussian, PCA does not fully achieve independence. This is because, in
this case, dependence in the data needs to be characterized through third and
fourth order statistical measures, which is not the case with PCA that depends
on second order statistics only.
Nevertheless, decorrelation prior to ICA processing allows ICA algorithms to
achieve faster convergence. In addition, the data is also sphered. This transform
is given by (Hyvarinen et al. 2001):
(4.11)
Once the data is transformed, we are ready to proceed with the ICA solution.
4.3.2
Information Minimization Solution for leA
For an m dimensional random vector u, the mutual information of its components is defined as:
( 4.12)
The mutual information is the most natural way of assessing statistical independence of random variables (Hyvarinen et al. 2001). When the components
are independent, the ratio inside the logarithm in (4.12) reduces to one, and
thus the mutual information becomes zero. It should be pointed out that 1(·) is
lower bounded by zero. It can also be shown that when the mutual information
is zero, the components of u are independent (Lee 1998). This property of the
mutual information is used here to achieve independence of components.
The mutual information can also be defined based on the entropy. For
a random variable u, the entropy of u is defined as:
H(u) = -E {log (p(u))} .
(4.13)
Similarly, for a random vector u, the joint entropy of u is defined as:
H(u) = -E {log (p(u))} .
( 4.14)
115
minimize the dependence between the components of the data. In the case of
PCA, dependence minimization is achieved when the covariance between the
components is zero, whereas ICA considers the components to be independent
when the probability density function of the vector can be expressed as the
product of the probability density functions of the components (see (4.2)).
When the data is Gaussian, decorrelation is equivalent to independence, and
thus PCA can be considered to be equivalent to ICA. However, when the data
is non-Gaussian, PCA does not fully achieve independence. This is because, in
this case, dependence in the data needs to be characterized through third and
fourth order statistical measures, which is not the case with PCA that depends
on second order statistics only.
Nevertheless, decorrelation prior to ICA processing allows ICA algorithms to
achieve faster convergence. In addition, the data is also sphered. This transform
is given by (Hyvarinen et al. 2001):
(4.11)
Once the data is transformed, we are ready to proceed with the ICA solution.
4.3.2
Information Minimization Solution for leA
For an m dimensional random vector u, the mutual information of its components is defined as:
( 4.12)
The mutual information is the most natural way of assessing statistical independence of random variables (Hyvarinen et al. 2001). When the components
are independent, the ratio inside the logarithm in (4.12) reduces to one, and
thus the mutual information becomes zero. It should be pointed out that 1(·) is
lower bounded by zero. It can also be shown that when the mutual information
is zero, the components of u are independent (Lee 1998). This property of the
mutual information is used here to achieve independence of components.
The mutual information can also be defined based on the entropy. For
a random variable u, the entropy of u is defined as:
H(u) = -E {log (p(u))} .
(4.13)
Similarly, for a random vector u, the joint entropy of u is defined as:
H(u) = -E {log (p(u))} .
( 4.14)
