5.4 Correspondence Analysis (CA)
5.4.1 Introduction
For a long time, CA has been one of the favourite tools for the analysis of species
presence-absence or abundance data. The raw data are first transformed into a matrix
Q of cell-by-cell contributions to the Pearson χ
2 statistic, and the resulting table is
submitted to a singular value decomposition to compute its eigenvalues (which are
the squares of the singular values) and eigenvectors. The result is an ordination in
which the χ
2 distance (D 16 ) is preserved among sites instead of the Euclidean
distance D 1 . The χ
2 distance is not influenced by double zeros; it is an asymmetrical
D function, as shown in Sect. 3.3.1. Therefore, CA is a method adapted to the
analysis of species abundance data without pre-transformation. Note that the data
submitted to CA must be frequencies or frequency-like, dimensionally homogeneous and non-negative; that is the case of species counts, biomasses, or presenceabsence data.
For technical reasons due to the implicit centring of the frequencies in the
calculation of the
Q matrix, CA ordination produces one axis fewer than min[n,
p]. As in PCA, the orthogonal axes are ranked in decreasing order of the variation
they represent, but instead of the total variance of the data, the variation is measured
by a quantity called the total inertia (sum of squares of all values in matrix
Q, see
Legendre and Legendre 2012, under Eq. 9.25). Individual eigenvalues are always
smaller than 1. To know the amount of variation represented along an axis, one
divides the eigenvalue of this axis by the total inertia of the species data matrix.
In CA, both the objects and the species are generally represented as points in the
same joint plot. As in PCA, two scalings of the results are available; they are most
useful in ecology. In the explanation that follows, remember the objects (sites) of the
data matrix are the rows and the species are the columns:
• CA scaling 1: rows (sites) are at the centroids of columns (species). This scaling
is the most appropriate if one is primarily interested in the ordination of objects
(sites). In the multidimensional space, the χ
2 distance is preserved among objects.
Interpretation: (1) the distances among objects in the reduced space approximate
their χ
2 distances. Thus, object points that are close to one another are likely to be
fairly similar in their species relative frequencies. (2) Any object found near the
point representing a species is likely to contain a high contribution of that species.
For presence-absence data, the object is more likely to possess the state “1” for
that species.
• CA scaling 2: columns (species) are at the centroids of rows (sites). This scaling
is the most appropriate if one is primarily interested in the ordination of species.
In the multidimensional space, the χ
2 distance is preserved among species.
Interpretation: (1) the distances among species in the reduced space approximate
their χ
2 distances. Thus, species points that are close to one another are likely to
have fairly similar relative frequencies along the objects. (2) Any species that lies
5.4 Correspondence Analysis (CA)
175
Précédent

- 188/444

Suivant