Husson and Josse have grouped all necessary functions to perform iterative
imputation of missing values in PCA in a package called missMDA. The function
performing the imputation is imputePCA(). We shall now experiment it on the
Doubs environmental data in two runs, one with only 3 missing values (i.e. 1% of the
319 values of the table) and the other with 32 (10%) missing values.
Our first experiment consists in removing three values selected to represent
various situations. The first value will be close to the mean of a variable with a
reasonably symmetrical distribution. pH has mean ¼ 8.04. Let us delete the value in
site 2 (8.0). The second will be close to the mean of a highly asymmetrical variable,
pho, with mean 0.57. Let us delete the value in site 19 (row 18, 0.60). The third will
also be removed from an asymmetrical variable, but far from its mean. bod has a
mean of 5.01. Let us delete the value of site 23 (row 22, 16.4, the second highest
value). How does imputePCA() perform?
# Imputation of missing values in PCA - 1: 3 missing values
# Replacement of 3 selected values by NA
env.miss3 <- env
env.miss3[2, 5] <- NA
# pH
env.miss3[18, 7] <- NA
# pho
env.miss3[22, 11] <- NA # dbo
# New means of the involved variables (without missing values)
mean(env.miss3[, 5], na.rm = TRUE)
mean(env.miss3[, 7], na.rm = TRUE)
mean(env.miss3[, 11], na.rm = TRUE)
# Imputation
env.imp <- imputePCA(env.miss3)
# Imputed values
env.imp$completeObs[2, 5]
# Original value: 8.0
env.imp$completeObs[18, 7]
# Original value: 0.60
env.imp$completeObs[22, 11] # Original value: 16.4
# PCA on the imputed data
env.imp3 <- env.imp$completeObs
env.imp3.pca <- rda(env.imp3, scale = TRUE)
# Procrustes comparison of original PCA and PCA on imputed data
pca.proc <- procrustes(env.pca, env.imp3.pca, scaling = 1)
The only good imputation is the one for pH, which falls right on the original
value. Remember that pH is symmetrically distributed in the data. The two other
imputations fall far away from the original values. In both cases the variables have
strongly asymmetrical distributions. Obviously, this has more importance than the
172
5 Unconstrained Ordination
Précédent

- 185/444

Suivant