Judging by these plots of bivariate relationships, would you favour the use of
Kendall’s tau or Pearson’s r?
3.4.4 R Mode: Binary Data (Other than Species Abundance Data)
The simplest way of comparing pairs of binary variables is to compute a matrix of
Pearson’s r. In that case Pearson’s r is called the point correlation coefficient or
Pearson’s φ. That coefficient is closely related to the chi-square statistic for 2 Â 2
tables without correction for continuity: χ
2
¼ nφ
2 where n is the number of objects.
3.5 Pre-transformations for Species Data
In Sect. 3.2.2, we explained why species abundance data should be treated in a
special way, avoiding the use of double zeros as indications of resemblance among
sites; linear methods, which explicitly or implicitly use the Euclidean distance
among sites or the covariance or correlation among variables, are therefore not
appropriate for such data. Unfortunately, many of the most powerful statistical
tools available to ecologists, like ANOVA, k-means partitioning (see Chap. 4),
principal component analysis (PCA, see Chap. 5) and redundancy analysis (RDA,
see Chap. 6) are methods of linear analysis. Consequently, these methods were more
or less “forbidden” to species data until Legendre and Gallagher (2001) showed that
several asymmetrical association measures (i.e., measures that are appropriate to
species data) can be obtained in two computation steps: a transformation of the raw
data followed by the calculation of the Euclidean distance. These two steps reconstruct the asymmetrical distance among sites, therefore allowing the use of all linear
methods of analysis with species data.
As will be seen in the following chapters, in many cases one simply has to apply
the pre-transformation to the species data, and then feed these to the linear methods
of data analysis: PCA, RDA, k-means, and so on.
Legendre and Gallagher proposed five pre-transformations of the species data
9 .
Four of them are available in vegan as arguments of the function decostand():
profiles of relative abundances by site ("total")
10 , site normalization, also called
the chord transformation ("normalize"), Hellinger transformation
("hellinger"), and chi-square double standardization ("chi.square"). We
can add the log-chord transformation to this list. See Sects. 2.2.4 and 3.3.1 for
examples. All these transformations express the data as relative abundances per sites
9 These authors proposed two forms of chi-square transformation. These forms are closely related,
so that implementing only one is sufficient for data analysis.
10 Note that Legendre and De Cáceres (2013) have shown that, contrary to the other transformations,
the distance between species profiles lacks important properties to study beta diversity and should
therefore be avoided in this wide context.
3.5 Pre-transformations for Species Data
55
Kendall’s tau or Pearson’s r?
3.4.4 R Mode: Binary Data (Other than Species Abundance Data)
The simplest way of comparing pairs of binary variables is to compute a matrix of
Pearson’s r. In that case Pearson’s r is called the point correlation coefficient or
Pearson’s φ. That coefficient is closely related to the chi-square statistic for 2 Â 2
tables without correction for continuity: χ
2
¼ nφ
2 where n is the number of objects.
3.5 Pre-transformations for Species Data
In Sect. 3.2.2, we explained why species abundance data should be treated in a
special way, avoiding the use of double zeros as indications of resemblance among
sites; linear methods, which explicitly or implicitly use the Euclidean distance
among sites or the covariance or correlation among variables, are therefore not
appropriate for such data. Unfortunately, many of the most powerful statistical
tools available to ecologists, like ANOVA, k-means partitioning (see Chap. 4),
principal component analysis (PCA, see Chap. 5) and redundancy analysis (RDA,
see Chap. 6) are methods of linear analysis. Consequently, these methods were more
or less “forbidden” to species data until Legendre and Gallagher (2001) showed that
several asymmetrical association measures (i.e., measures that are appropriate to
species data) can be obtained in two computation steps: a transformation of the raw
data followed by the calculation of the Euclidean distance. These two steps reconstruct the asymmetrical distance among sites, therefore allowing the use of all linear
methods of analysis with species data.
As will be seen in the following chapters, in many cases one simply has to apply
the pre-transformation to the species data, and then feed these to the linear methods
of data analysis: PCA, RDA, k-means, and so on.
Legendre and Gallagher proposed five pre-transformations of the species data
9 .
Four of them are available in vegan as arguments of the function decostand():
profiles of relative abundances by site ("total")
10 , site normalization, also called
the chord transformation ("normalize"), Hellinger transformation
("hellinger"), and chi-square double standardization ("chi.square"). We
can add the log-chord transformation to this list. See Sects. 2.2.4 and 3.3.1 for
examples. All these transformations express the data as relative abundances per sites
9 These authors proposed two forms of chi-square transformation. These forms are closely related,
so that implementing only one is sufficient for data analysis.
10 Note that Legendre and De Cáceres (2013) have shown that, contrary to the other transformations,
the distance between species profiles lacks important properties to study beta diversity and should
therefore be avoided in this wide context.
3.5 Pre-transformations for Species Data
55
