150
K. M. Sørensen et al.
where x m is the mth column of X. The correct form for PCA is then written:
X = X + T · P
T
+ E
(7.19)
Mean centering is often considered together with its analog in correlation analysis,
namely autoscaling. In spectral data, each wavelength is expressed in absorbance (or
reflectance) units and, thus, is approximately equal in scale and variance, which in
turn can be weighed equal in the analysis. It is thus sensible to only apply mean
centering in spectroscopic analysis. However, if the variables were composed from
different measurement types, with different units, they will be weighed unequally in
the analysis. Autoscaling seeks to rectify this by scaling each variable to have unit
variance [11], where the mth column of X is mean centered and normalized by the
standard deviation of the mth column:
ˆ
x m =
x m − x m
std(x m )
(7.20)
Due to the orthogonality constraint imposed in the model, PCA has a simple
and unambiguous solution that can be calculated rapidly using, e.g., singular value
decomposition (SVD). PCA is thus present in all commercial chemometric software packages due to its extraordinary robust data reduction and data summarizing
capabilities.
7.4.2 Explained Variance
It is very useful to be able to quantify how much of the information in the data that
a given component—or the residual—is describing. In this respect, each set of t’s
and p’s is called a component; thus, T is of size n by f and P is of size m by f . The
graphical representation of a PCA with f components is identical to Fig. 7.2. The
total variance of X is explained by the sum of the individual components (plus the
residual):
X = T · P
T
+ E = t 1 · p
T
1 + t 2 · p
T
2 + · · · + t f · p
T
f + E
(7.21)
where t f is the f th column vector of T and p f is the f th column vector of P.
The sum of squares of X (size n × m) is defined as the summation of the squared
of each value of X:
SSQ(X) =
n
m
X
2
n,m
(7.22)
Similarly, the sum of squares can be calculated for an individual component f :
K. M. Sørensen et al.
where x m is the mth column of X. The correct form for PCA is then written:
X = X + T · P
T
+ E
(7.19)
Mean centering is often considered together with its analog in correlation analysis,
namely autoscaling. In spectral data, each wavelength is expressed in absorbance (or
reflectance) units and, thus, is approximately equal in scale and variance, which in
turn can be weighed equal in the analysis. It is thus sensible to only apply mean
centering in spectroscopic analysis. However, if the variables were composed from
different measurement types, with different units, they will be weighed unequally in
the analysis. Autoscaling seeks to rectify this by scaling each variable to have unit
variance [11], where the mth column of X is mean centered and normalized by the
standard deviation of the mth column:
ˆ
x m =
x m − x m
std(x m )
(7.20)
Due to the orthogonality constraint imposed in the model, PCA has a simple
and unambiguous solution that can be calculated rapidly using, e.g., singular value
decomposition (SVD). PCA is thus present in all commercial chemometric software packages due to its extraordinary robust data reduction and data summarizing
capabilities.
7.4.2 Explained Variance
It is very useful to be able to quantify how much of the information in the data that
a given component—or the residual—is describing. In this respect, each set of t’s
and p’s is called a component; thus, T is of size n by f and P is of size m by f . The
graphical representation of a PCA with f components is identical to Fig. 7.2. The
total variance of X is explained by the sum of the individual components (plus the
residual):
X = T · P
T
+ E = t 1 · p
T
1 + t 2 · p
T
2 + · · · + t f · p
T
f + E
(7.21)
where t f is the f th column vector of T and p f is the f th column vector of P.
The sum of squares of X (size n × m) is defined as the summation of the squared
of each value of X:
SSQ(X) =
n
m
X
2
n,m
(7.22)
Similarly, the sum of squares can be calculated for an individual component f :
