The graphs may be mirror images of those obtained with vegan. This is
unimportant since the choice of the sign of the principal components, made within
the eigen() function used in PCA functions, is arbitrary.
5.3.6 Imputation of Missing Values in PCA
In an ideal world, data would be perfect, with clear signal, as little noise as possible
and, above all, without missing values. Unfortunately, it happens sometimes that for
various reasons, in an otherwise good and interesting data set, values are missing.
This poses problems. PCA cannot be computed on an incomplete data matrix. What
should one do: delete the affected row(s) or column(s) and thereby loose precious
existing data? Or, to avoid this, replace (“impute” in statistical language) the missing
value(s) by some meaningful substitute, e.g. the mean of the corresponding variable
(s) or an estimate obtained by regression analyses involving the other variables?
Attractive as these two solutions appear, they do not consider the uncertainty
attached to the estimation of the missing values and their use in further analyses.
As a consequence, standard errors of estimators computed on the imputed data set
are underestimated, leading to invalid tests if some are performed using these data.
To address these shortcomings, Josse and Husson (2012) proposed “a regularized
iterative PCA algorithm to provide point estimates of the principal axes and
components (. . .)”. The PCA is computed iteratively, axis by axis, in a completely
different way as the one used in vegan. When missing values are present, they are
first replaced by the mean of the corresponding variable. Then a PCA is computed,
and the missing values receive new estimates based on the result. The procedure is
repeated until convergence, which corresponds to the point where the imputed
values become stable. Complex procedures are implemented (1) to avoid overfitting
of the imputation model (i.e., the problem of having too many parameters to adjust
with respect to the amount of data, a situation occurring when many values are
missing), and (2) to overcome the problem of the underestimation of the variance of
the imputed data.
There is an important caveat here: the PCA solutions obtained with imputed
values estimated using different numbers of PCA axes are not nested, i.e., a solution
based on s axes is not identical to the first s axes of a solution with (s + 1) or (s + 2)
axes.. Consequently, it is important to make an appropriate a priori decision about
the desired number of axes. The decision may be empirical (for example, use the
axes with sum of eigenvalues larger than the broken stick model prediction), or it can
be informed by a cross-validation procedure on the data themselves, where data are
deleted, reconstructed in PCAs with 1 to n À 1 dimensions, and the reconstruction
error is computed. One retains the number of dimensions minimizing the reconstruction error.
5.3 Principal Component Analysis (PCA)
171
unimportant since the choice of the sign of the principal components, made within
the eigen() function used in PCA functions, is arbitrary.
5.3.6 Imputation of Missing Values in PCA
In an ideal world, data would be perfect, with clear signal, as little noise as possible
and, above all, without missing values. Unfortunately, it happens sometimes that for
various reasons, in an otherwise good and interesting data set, values are missing.
This poses problems. PCA cannot be computed on an incomplete data matrix. What
should one do: delete the affected row(s) or column(s) and thereby loose precious
existing data? Or, to avoid this, replace (“impute” in statistical language) the missing
value(s) by some meaningful substitute, e.g. the mean of the corresponding variable
(s) or an estimate obtained by regression analyses involving the other variables?
Attractive as these two solutions appear, they do not consider the uncertainty
attached to the estimation of the missing values and their use in further analyses.
As a consequence, standard errors of estimators computed on the imputed data set
are underestimated, leading to invalid tests if some are performed using these data.
To address these shortcomings, Josse and Husson (2012) proposed “a regularized
iterative PCA algorithm to provide point estimates of the principal axes and
components (. . .)”. The PCA is computed iteratively, axis by axis, in a completely
different way as the one used in vegan. When missing values are present, they are
first replaced by the mean of the corresponding variable. Then a PCA is computed,
and the missing values receive new estimates based on the result. The procedure is
repeated until convergence, which corresponds to the point where the imputed
values become stable. Complex procedures are implemented (1) to avoid overfitting
of the imputation model (i.e., the problem of having too many parameters to adjust
with respect to the amount of data, a situation occurring when many values are
missing), and (2) to overcome the problem of the underestimation of the variance of
the imputed data.
There is an important caveat here: the PCA solutions obtained with imputed
values estimated using different numbers of PCA axes are not nested, i.e., a solution
based on s axes is not identical to the first s axes of a solution with (s + 1) or (s + 2)
axes.. Consequently, it is important to make an appropriate a priori decision about
the desired number of axes. The decision may be empirical (for example, use the
axes with sum of eigenvalues larger than the broken stick model prediction), or it can
be informed by a cross-validation procedure on the data themselves, where data are
deleted, reconstructed in PCAs with 1 to n À 1 dimensions, and the reconstruction
error is computed. One retains the number of dimensions minimizing the reconstruction error.
5.3 Principal Component Analysis (PCA)
171
