114
4: Stefan A. Robila
statistics, and to speed up the convergence process. A random vector is said to
be white if the mean of each component is zero and the covariance matrix is
the identity matrix. Among the various choices for preprocessing techniques,
principal component analysis (PCA) (Hyvarinen et al. 2001; Chang et al. 2002)
is widely used. This method has the additional advantage of eliminating dependence up to the second order between the components (i. e. the components
are decorrelated).
PCA attempts to solve the following problem. For the multidimensional
random vector x find a linear transform W so that the components obtained
are uncorrelated (Richards and Jia 1999):
y=Wx
such that
Ly = E {(y- E {y}) (y- E {y})T}
(4.5)
(4.6)
is diagonal. The vector of expected values E{y} and the covariance matrix Iy
can be expressed in terms of the vector of expected values and the covariance
matrix for x:
E {y} = E {Wx} = WE{x} ,
Ly=WLx WT .
(4.7)
(4.8)
If Ax is the matrix denoting the normalized eigenvectors for the covariance
matrix Ix, we have (Rencher 1995):
(4.9)
where ex is the corresponding diagonal eigenvalue matrix. Since in (4.6), Iy is
required to be diagonal, we note that W = AI leads to the PCA solution:
-AT
y- xx.
(4.lO)
The components of yare called principal components and the eigenvector
matrix Ax is called the principal component transform (PCT). The eigenvalues of the covariance matrix for x correspond to the variances of the principal
components. When these eigenvalues are arranged in a decreasing order (along
with the corresponding permutation of the eigenvectors), we get the components of y sorted in the decreasing order of their variance. High variance is
usually associated with high signal to noise ratio. Therefore, it is useful to
consider the highest variance components for further processing as they are
expected to contain most of the information. For this reason, PCA has been
considered as a possible tool for data compression or feature extraction (by
dropping the lowest variance components).
It is also interesting to note that ICA is frequently compared with PCA. The
similarity between the two originates from their goals - that is they both try to
4: Stefan A. Robila
statistics, and to speed up the convergence process. A random vector is said to
be white if the mean of each component is zero and the covariance matrix is
the identity matrix. Among the various choices for preprocessing techniques,
principal component analysis (PCA) (Hyvarinen et al. 2001; Chang et al. 2002)
is widely used. This method has the additional advantage of eliminating dependence up to the second order between the components (i. e. the components
are decorrelated).
PCA attempts to solve the following problem. For the multidimensional
random vector x find a linear transform W so that the components obtained
are uncorrelated (Richards and Jia 1999):
y=Wx
such that
Ly = E {(y- E {y}) (y- E {y})T}
(4.5)
(4.6)
is diagonal. The vector of expected values E{y} and the covariance matrix Iy
can be expressed in terms of the vector of expected values and the covariance
matrix for x:
E {y} = E {Wx} = WE{x} ,
Ly=WLx WT .
(4.7)
(4.8)
If Ax is the matrix denoting the normalized eigenvectors for the covariance
matrix Ix, we have (Rencher 1995):
(4.9)
where ex is the corresponding diagonal eigenvalue matrix. Since in (4.6), Iy is
required to be diagonal, we note that W = AI leads to the PCA solution:
-AT
y- xx.
(4.lO)
The components of yare called principal components and the eigenvector
matrix Ax is called the principal component transform (PCT). The eigenvalues of the covariance matrix for x correspond to the variances of the principal
components. When these eigenvalues are arranged in a decreasing order (along
with the corresponding permutation of the eigenvectors), we get the components of y sorted in the decreasing order of their variance. High variance is
usually associated with high signal to noise ratio. Therefore, it is useful to
consider the highest variance components for further processing as they are
expected to contain most of the information. For this reason, PCA has been
considered as a possible tool for data compression or feature extraction (by
dropping the lowest variance components).
It is also interesting to note that ICA is frequently compared with PCA. The
similarity between the two originates from their goals - that is they both try to
