Independent Component Analysis
119
process (Lee 1998). The MMI algorithm for ICA will repeatedly perform an
update of the matrix W:
W=W+k1W,
(4.32)
where the update step is given by (4.31) and k is an update coefficient « 1)
that controls the convergence speed.
The above update procedure does not restrict the search space to vectors of
uncorrelated components. This contradicts our earlier statement that ICA is
an extension of decorrelation methods (like PCA). However, if the components
of u are uncorrelated, since x is also uncorrelated (after PCA processing) we
have:
L = W L W T = WIW T = WW T •
( 4.33)
u
x
Therefore, for the resulting components to be decorrelated, the matrix W
needs to be orthogonal. It makes sense to restrict the optimization space
only to orthogonal matrices, by adding an additional step in which W is
orthogonalized after each iteration:
(4.34)
In Hyvarinen et al. (2001), a direct solution to the orthogonalization is given
by:
(4.35)
where Cw and Aware the eigenvalue and eigenvector matrices respectively
for WW T •
Results available in the literature indicate that full independence is not
always achieved by the ICA solution obtained by the above algorithm even
under ideal conditions (when the components are independent and there is
no noise) (Hyvarinen et al. 2001; Lee 1998). This is because, similar to all
other gradient-based algorithms, the MMI algorithm may converge to a local
minimum. To reduce the likelihood of this, remedial methods (for example, the
use of variable update coefficients) can be employed. This way, the likelihood
of achieving full independence is increased (Hyvarinen et al. 2001).
Suitable selection of the nonlinear functions gi(') is essential to obtain an
accurate solution from the algorithm since they need to approximate the probability density functions of the components of u. In Cochocki and Amari (2002),
it was suggested to model the probability density functions of the components
by a weighted sum of parametric logistic functions. However, the estimation
of the weights proved to be computationally expensive and the accuracy of
the results did not always improve (Lee 1998). On the other hand, experimental results have indicated that even with a single fixed nonlinearity, the
119
process (Lee 1998). The MMI algorithm for ICA will repeatedly perform an
update of the matrix W:
W=W+k1W,
(4.32)
where the update step is given by (4.31) and k is an update coefficient « 1)
that controls the convergence speed.
The above update procedure does not restrict the search space to vectors of
uncorrelated components. This contradicts our earlier statement that ICA is
an extension of decorrelation methods (like PCA). However, if the components
of u are uncorrelated, since x is also uncorrelated (after PCA processing) we
have:
L = W L W T = WIW T = WW T •
( 4.33)
u
x
Therefore, for the resulting components to be decorrelated, the matrix W
needs to be orthogonal. It makes sense to restrict the optimization space
only to orthogonal matrices, by adding an additional step in which W is
orthogonalized after each iteration:
(4.34)
In Hyvarinen et al. (2001), a direct solution to the orthogonalization is given
by:
(4.35)
where Cw and Aware the eigenvalue and eigenvector matrices respectively
for WW T •
Results available in the literature indicate that full independence is not
always achieved by the ICA solution obtained by the above algorithm even
under ideal conditions (when the components are independent and there is
no noise) (Hyvarinen et al. 2001; Lee 1998). This is because, similar to all
other gradient-based algorithms, the MMI algorithm may converge to a local
minimum. To reduce the likelihood of this, remedial methods (for example, the
use of variable update coefficients) can be employed. This way, the likelihood
of achieving full independence is increased (Hyvarinen et al. 2001).
Suitable selection of the nonlinear functions gi(') is essential to obtain an
accurate solution from the algorithm since they need to approximate the probability density functions of the components of u. In Cochocki and Amari (2002),
it was suggested to model the probability density functions of the components
by a weighted sum of parametric logistic functions. However, the estimation
of the weights proved to be computationally expensive and the accuracy of
the results did not always improve (Lee 1998). On the other hand, experimental results have indicated that even with a single fixed nonlinearity, the
