• Are you comparing objects (Q-mode) or variables (R-mode analysis)?
• Are you dealing with species data (which usually leads to asymmetrical coefficients) or other types of variables (symmetrical coefficients)?
• Are your data binary (binary coefficients), quantitative (quantitative coefficients)
or mixed or of other types (e.g. ordinal; special coefficients)?
In the following sections, you will explore various possibilities. In most cases,
more than one association measure is available to study a given problem.
3.3 Q Mode: Computing Dissimilarity Matrices Among
Objects
In the Q mode, we will use six packages: stats (included in the standard installation of R), vegan, ade4, adespatial, cluster and FD. Note that this list of
packages is by far not exhaustive, but it should satisfy the needs of most ecologists.
Although the literature provides similarity as well as dissimilarity measures, in
R all similarity measures are converted to dissimilarities to compute a square matrix
of class "dist" in which the diagonal (distance between each object and itself) is
0 and can be ignored. The conversion formula varies with the package used, and this
is not without consequences:
• in stats, FD and vegan, the conversion from similarities S to dissimilarities
D is done by computing D ¼ 1 – S;
• in ade4, it is done as
ffiffiffiffiffiffiffiffiffiffiffi
1 À S
p
. This allows several indices to become Euclidean
5 ,
a geometric property that will be useful in some analyses, e.g. in principal
coordinate analysis (see Chap. 5). We will come to it when it becomes relevant.
Dissimilarity matrices computed by other packages using coefficients that are not
Euclidean can often be made Euclidean by computing D2 <- sqrt(D);
• in adespatial, most similarity coefficients are converted as D ¼ 1 – S, the
exceptions being three classical indices for presence-absence data, namely
Jaccard, Sørensen and Ochiai (see below), which are converted as
ffiffiffiffiffiffiffiffiffiffiffi
1 À S
p
: For
all computed coefficients, function dist.ldc() of that package produces
messages indicating if the selected coefficient is Euclidean or non-Euclidean, or
if sqrt(D) would be Euclidean;
• in cluster, all available measures are dissimilarities, so no conversion has to
be made;
5 Metric dissimilarity measures, which are also called distances, share four properties: minimum
0, positiveness, symmetry and triangle inequality. Furthermore, the points can be represented in a
Euclidean space. It may happen, though, that some dissimilarity matrices are metric (triangle
inequality respected) but non Euclidean (all points cannot be represented in a Euclidean space).
See Legendre and Legendre (2012) p. 500.
38
3 Association Measures and Matrices
• Are you dealing with species data (which usually leads to asymmetrical coefficients) or other types of variables (symmetrical coefficients)?
• Are your data binary (binary coefficients), quantitative (quantitative coefficients)
or mixed or of other types (e.g. ordinal; special coefficients)?
In the following sections, you will explore various possibilities. In most cases,
more than one association measure is available to study a given problem.
3.3 Q Mode: Computing Dissimilarity Matrices Among
Objects
In the Q mode, we will use six packages: stats (included in the standard installation of R), vegan, ade4, adespatial, cluster and FD. Note that this list of
packages is by far not exhaustive, but it should satisfy the needs of most ecologists.
Although the literature provides similarity as well as dissimilarity measures, in
R all similarity measures are converted to dissimilarities to compute a square matrix
of class "dist" in which the diagonal (distance between each object and itself) is
0 and can be ignored. The conversion formula varies with the package used, and this
is not without consequences:
• in stats, FD and vegan, the conversion from similarities S to dissimilarities
D is done by computing D ¼ 1 – S;
• in ade4, it is done as
ffiffiffiffiffiffiffiffiffiffiffi
1 À S
p
. This allows several indices to become Euclidean
5 ,
a geometric property that will be useful in some analyses, e.g. in principal
coordinate analysis (see Chap. 5). We will come to it when it becomes relevant.
Dissimilarity matrices computed by other packages using coefficients that are not
Euclidean can often be made Euclidean by computing D2 <- sqrt(D);
• in adespatial, most similarity coefficients are converted as D ¼ 1 – S, the
exceptions being three classical indices for presence-absence data, namely
Jaccard, Sørensen and Ochiai (see below), which are converted as
ffiffiffiffiffiffiffiffiffiffiffi
1 À S
p
: For
all computed coefficients, function dist.ldc() of that package produces
messages indicating if the selected coefficient is Euclidean or non-Euclidean, or
if sqrt(D) would be Euclidean;
• in cluster, all available measures are dissimilarities, so no conversion has to
be made;
5 Metric dissimilarity measures, which are also called distances, share four properties: minimum
0, positiveness, symmetry and triangle inequality. Furthermore, the points can be represented in a
Euclidean space. It may happen, though, that some dissimilarity matrices are metric (triangle
inequality respected) but non Euclidean (all points cannot be represented in a Euclidean space).
See Legendre and Legendre (2012) p. 500.
38
3 Association Measures and Matrices
