The groups of sites can be defined in various ways. A simple, although circular
way is to use the result of a cluster analysis based on the species data. The indicator
species are then simply the most prominent members of these groups. Of course,
one cannot rely on statistical tests to identify the indicator species in this case, since
the classification and the indicator species do not derive from independent data.
Another approach, which is conceptually and statistically more fruitful, consists in
clustering the sites on the basis of independent data (environmental variables, for
instance). The indicator species can then be considered indicator in the true sense of
the word, i.e. species closely related to the ecological conditions of their group. The
statistical significance of the indicator values (i.e., the probability of obtaining by
chance as high an indicator value as observed) is assessed by means of a permutation test.
The Dufrêne and Legendre index is available in function indval()of package
labdsv. Let us apply it to our fish data. As an example of a search for indicator
species related to one explanatory variable, we could for instance divide the data set
into groups of contiguous sites on the basis of the variable dfs (distance from the
source), acting as a surrogate of the overall river gradient conditions, and look for
indicator species in the groups.
# Divide the sites into 4 groups depending on the distance from
# the source of the river
dfs.D1 <- dist(data.frame(dfs = env[, 1],
row.names = rownames(env)))
dfsD1.kmeans <- kmeans(dfs.D1, centers = 4, nstart = 100)
# Cluster delimitation and numbering
dfsD1.kmeans$cluster
# Cluster labels are in arbitrary order
# See Hint below
grps <- rep(1:4, c(8, 10, 6, 5))
# Indicator species for this typology of the sites
(iva <- indval(spe, grps, numitr = 10000))
Hint BEWARE: in the output table, the cluster numbering corresponds to the cluster
numbers fed to the program, which do not necessarily follow a meaningful order.
The k-means analysis produces arbitrary group labels in the form of numbers. The
column headers of the indval() output are the ones provided to the program,
but the columns have been reordered according to the labels. Consequently, the
groups do not follow the order along the river. This is why we have manually
constructed an object (grps) with sequential group numbers.
4.11 Indicator Species
121
way is to use the result of a cluster analysis based on the species data. The indicator
species are then simply the most prominent members of these groups. Of course,
one cannot rely on statistical tests to identify the indicator species in this case, since
the classification and the indicator species do not derive from independent data.
Another approach, which is conceptually and statistically more fruitful, consists in
clustering the sites on the basis of independent data (environmental variables, for
instance). The indicator species can then be considered indicator in the true sense of
the word, i.e. species closely related to the ecological conditions of their group. The
statistical significance of the indicator values (i.e., the probability of obtaining by
chance as high an indicator value as observed) is assessed by means of a permutation test.
The Dufrêne and Legendre index is available in function indval()of package
labdsv. Let us apply it to our fish data. As an example of a search for indicator
species related to one explanatory variable, we could for instance divide the data set
into groups of contiguous sites on the basis of the variable dfs (distance from the
source), acting as a surrogate of the overall river gradient conditions, and look for
indicator species in the groups.
# Divide the sites into 4 groups depending on the distance from
# the source of the river
dfs.D1 <- dist(data.frame(dfs = env[, 1],
row.names = rownames(env)))
dfsD1.kmeans <- kmeans(dfs.D1, centers = 4, nstart = 100)
# Cluster delimitation and numbering
dfsD1.kmeans$cluster
# Cluster labels are in arbitrary order
# See Hint below
grps <- rep(1:4, c(8, 10, 6, 5))
# Indicator species for this typology of the sites
(iva <- indval(spe, grps, numitr = 10000))
Hint BEWARE: in the output table, the cluster numbering corresponds to the cluster
numbers fed to the program, which do not necessarily follow a meaningful order.
The k-means analysis produces arbitrary group labels in the form of numbers. The
column headers of the indval() output are the ones provided to the program,
but the columns have been reordered according to the labels. Consequently, the
groups do not follow the order along the river. This is why we have manually
constructed an object (grps) with sequential group numbers.
4.11 Indicator Species
121
