The methods that will be presented in this chapter are:
• Principal component analysis (PCA): the main eigenvector-based method.
Analysis performed on raw, quantitative data. Preserves the Euclidean distance
among sites in scaling 1, and the Mahalanobis distance in scaling 2. Scaling is
explained in Sect. 5.3.2.2.
• Correspondence analysis (CA): works on data that must be frequencies or
frequency-like, dimensionally homogeneous, and non-negative. Preserves the χ
2
distance among rows (in scaling 1) or columns (in scaling 2). In ecology,
almost exclusively used to analyse community composition data.
• Multiple correspondence analysis MCA: ordination of a table of categorical
variables, i.e. a data frame where all variables are factors.
• Principal coordinate analysis (PCoA): devoted to the ordination of dissimilarity
matrices, most often in the Q mode, instead of site-by-variables tables. Hence,
great flexibility in the choice of association measures (Chap. 3).
• Nonmetric multidimensional scaling (NMDS): unlike the three others, this is
not an eigenvector-based method. NMDS tries to represent the set of objects
along a predetermined number of axes while preserving the ordering relationships
among them. Operates from a dissimilarity matrix.
PCoA and NMDS can produce ordinations from any square dissimilarity matrix,
which have been created in (transformed to) class “dist” in R.
5.3 Principal Component Analysis (PCA)
5.3.1 Overview
Imagine a data set whose variables are normally distributed. This data set will be said to
show a multinormal distribution. The first principal axis (or principal-component axis)
of a PCA of this data set is the straight line that goes through the greatest dimension of
the concentration ellipsoid describing this multinormal distribution. The following
axes, which are orthogonal to one another and successively shorter, go through the
following greatest dimensions of the ellipsoid (Legendre and Legendre 2012). One can
extract a maximum of p principal axes from a data set containing p variables.
Stated otherwise, PCA carries out a rotation of the original system of axes defined
by the variables, such that the successive new axes (called principal components) are
orthogonal to one another, and correspond to the successive dimensions of maximum variance of the scatter of points. The principal components give the positions
of the objects in the new system of coordinates. PCA works on a dispersion matrix S,
i.e. an association matrix among variables containing the variances and covariances
of the variables (when these are dimensionally homogeneous), or the correlations
computed from dimensionally heterogeneous variables. It is exclusively devoted to
the analysis of quantitative variables. The dissimilarity preserved is the Euclidean
distance and the relationships detected are linear. Therefore, it is not generally
appropriate to the analysis of raw species abundance data. These can, however, be
subjected to PCA after an appropriate pre-transformation (Sects. 3.5 and 5.3.3).
5.3 Principal Component Analysis (PCA)
153
• Principal component analysis (PCA): the main eigenvector-based method.
Analysis performed on raw, quantitative data. Preserves the Euclidean distance
among sites in scaling 1, and the Mahalanobis distance in scaling 2. Scaling is
explained in Sect. 5.3.2.2.
• Correspondence analysis (CA): works on data that must be frequencies or
frequency-like, dimensionally homogeneous, and non-negative. Preserves the χ
2
distance among rows (in scaling 1) or columns (in scaling 2). In ecology,
almost exclusively used to analyse community composition data.
• Multiple correspondence analysis MCA: ordination of a table of categorical
variables, i.e. a data frame where all variables are factors.
• Principal coordinate analysis (PCoA): devoted to the ordination of dissimilarity
matrices, most often in the Q mode, instead of site-by-variables tables. Hence,
great flexibility in the choice of association measures (Chap. 3).
• Nonmetric multidimensional scaling (NMDS): unlike the three others, this is
not an eigenvector-based method. NMDS tries to represent the set of objects
along a predetermined number of axes while preserving the ordering relationships
among them. Operates from a dissimilarity matrix.
PCoA and NMDS can produce ordinations from any square dissimilarity matrix,
which have been created in (transformed to) class “dist” in R.
5.3 Principal Component Analysis (PCA)
5.3.1 Overview
Imagine a data set whose variables are normally distributed. This data set will be said to
show a multinormal distribution. The first principal axis (or principal-component axis)
of a PCA of this data set is the straight line that goes through the greatest dimension of
the concentration ellipsoid describing this multinormal distribution. The following
axes, which are orthogonal to one another and successively shorter, go through the
following greatest dimensions of the ellipsoid (Legendre and Legendre 2012). One can
extract a maximum of p principal axes from a data set containing p variables.
Stated otherwise, PCA carries out a rotation of the original system of axes defined
by the variables, such that the successive new axes (called principal components) are
orthogonal to one another, and correspond to the successive dimensions of maximum variance of the scatter of points. The principal components give the positions
of the objects in the new system of coordinates. PCA works on a dispersion matrix S,
i.e. an association matrix among variables containing the variances and covariances
of the variables (when these are dimensionally homogeneous), or the correlations
computed from dimensionally heterogeneous variables. It is exclusively devoted to
the analysis of quantitative variables. The dissimilarity preserved is the Euclidean
distance and the relationships detected are linear. Therefore, it is not generally
appropriate to the analysis of raw species abundance data. These can, however, be
subjected to PCA after an appropriate pre-transformation (Sects. 3.5 and 5.3.3).
5.3 Principal Component Analysis (PCA)
153
