3.3.2 Q Mode: Binary (Presence-Absence) Species Data
When the only data available are binary, or when the abundances are irrelevant, or,
sometimes, when the data table contains quantitative values of uncertain or unequal
quality, the analyses are done on presence-absence (1–0) data.
The Doubs fish dataset is quantitative, but for the sake of the example we shall
compute binary coefficients from these data (using the appropriate arguments when
needed): all values larger than 0 will be given a value equal to 1. The exercise
consists in computing several dissimilarity matrices based on appropriate similarity
coefficients: the Jaccard (S 7 ) and Sørensen (S 8 ) similarities. For each pair of sites, the
Jaccard similarity is the ratio between the number of double 1’s and the number of
species, excluding the species represented by double zeros in the pair of objects
considered. Therefore, a Jaccard similarity of 0.25 means that 25% of the total
number of species observed at two sites were present in both sites and 75% in one
site only. The Jaccard dissimilarity computed in R is either (1–0.25) or
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
1 À 0:25
p
depending on the package used. The Sørensen similarity (S 8 ) gives double weight to
the number of double 1’s; its reciprocal (complement to 1) is equivalent to the
percentage difference (aka Bray-Curtis) dissimilarity computed on species presenceabsence data.
A further interesting relationship is that the Ochiai similarity (S 14 ), which is also
appropriate for species presence-absence data, is related to the chord, Hellinger and
log-chord distances. Computing either one of these distances on presence-absence
data, followed by division by
ffiffi ffi
2
p
, produces
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
1 À Ochiai similarity
p
. This reasoning
shows that the chord, Hellinger and log-chord transformations are meaningful for
species presence-absence data. This is also the case for the chi-square transformation: the Euclidean distance computed on data transformed in that way produces the
chi-square distance, which is appropriate for both quantitative and presenceabsence data.
## Q-mode dissimilarity measures for binary data
# Jaccard dissimilarity matrix using function vegdist()
spe.dj <- vegdist(spe, "jac", binary = TRUE)
head(spe.dj)
head(sqrt(spe.dj))
# Jaccard dissimilarity matrix using function dist()
spe.dj2 <- dist(spe, "binary")
head(spe.dj2)
# Jaccard dissimilarity matrix using function dist.binary()
spe.dj3 <- dist.binary(spe, method = 1)
head(spe.dj3)
# Sorensen dissimilarity matrix using function dist.ldc()
spe.ds <- dist.ldc(spe, "sorensen")
42
3 Association Measures and Matrices
When the only data available are binary, or when the abundances are irrelevant, or,
sometimes, when the data table contains quantitative values of uncertain or unequal
quality, the analyses are done on presence-absence (1–0) data.
The Doubs fish dataset is quantitative, but for the sake of the example we shall
compute binary coefficients from these data (using the appropriate arguments when
needed): all values larger than 0 will be given a value equal to 1. The exercise
consists in computing several dissimilarity matrices based on appropriate similarity
coefficients: the Jaccard (S 7 ) and Sørensen (S 8 ) similarities. For each pair of sites, the
Jaccard similarity is the ratio between the number of double 1’s and the number of
species, excluding the species represented by double zeros in the pair of objects
considered. Therefore, a Jaccard similarity of 0.25 means that 25% of the total
number of species observed at two sites were present in both sites and 75% in one
site only. The Jaccard dissimilarity computed in R is either (1–0.25) or
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
1 À 0:25
p
depending on the package used. The Sørensen similarity (S 8 ) gives double weight to
the number of double 1’s; its reciprocal (complement to 1) is equivalent to the
percentage difference (aka Bray-Curtis) dissimilarity computed on species presenceabsence data.
A further interesting relationship is that the Ochiai similarity (S 14 ), which is also
appropriate for species presence-absence data, is related to the chord, Hellinger and
log-chord distances. Computing either one of these distances on presence-absence
data, followed by division by
ffiffi ffi
2
p
, produces
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
1 À Ochiai similarity
p
. This reasoning
shows that the chord, Hellinger and log-chord transformations are meaningful for
species presence-absence data. This is also the case for the chi-square transformation: the Euclidean distance computed on data transformed in that way produces the
chi-square distance, which is appropriate for both quantitative and presenceabsence data.
## Q-mode dissimilarity measures for binary data
# Jaccard dissimilarity matrix using function vegdist()
spe.dj <- vegdist(spe, "jac", binary = TRUE)
head(spe.dj)
head(sqrt(spe.dj))
# Jaccard dissimilarity matrix using function dist()
spe.dj2 <- dist(spe, "binary")
head(spe.dj2)
# Jaccard dissimilarity matrix using function dist.binary()
spe.dj3 <- dist.binary(spe, method = 1)
head(spe.dj3)
# Sorensen dissimilarity matrix using function dist.ldc()
spe.ds <- dist.ldc(spe, "sorensen")
42
3 Association Measures and Matrices
