THE USE O F STATISTICS IN PHYTOSOCIOLOQY
79
general consensus of opinion and confine our attention to methods of
hierarchical classification as the primary approach.
A second decision to be made concerns the use of “subdivisive” or
“agglomerative” methods. The former begin with the whole population
of sites and divide it successively into smaller groups, each group being
examined independently for possible further subdivision as it is
extracted; the latter begin at the bottom and combine the individual
sites which are most alike until all individuals are eventually united
in a single population. Subdivisive methods thus concentrate essentially on differences, while agglomerative methods seek similarities. Two
points are of particular interest here. First, it should be noted that
subdivisive methods start from maximal information obtained over the
whole population, while agglomerative techniques start from single
units of minimal information. Secondly, subdivisive methods can be
terminated at any convenient level, while agglomerative methods require the whole analysis to be completed before the large-scale divisions
at the top of the hierarchy can be obtained. In general, therefore,
subdivisive methods are to be preferred on theoretical grounds, although
the actual calculations involved frequently require more computing
time.
A further choice lies between the use of “polythetic” or ‘‘monothetic” methods. Polythetic methods employ a combination of characters
to form the groups, while monothetic methods use only a single character
for each division; monothetic methods can only be subdivisive, but polythetic methods can be either subdivisive or agglomerative. The theoretical advantage of polythetic systems is that the classification obtained
is usually more stable and, by its nature, more informative; against this,
monothetic methods usually involve much less computation.
In addition to the above, there are yet further decisions to be made
on more subtle statistical considerations. These have been set out fully
by Williams and Dale in the parallel paper already mentioned, and need
only very brief consideration here. The &st concerns the question of
“internal weighting”. We have already (p. 71) condemned the practice
of external weighting on subjective estimations of the “most important”
species, but the problem of internal weighting of those species found to
be most informative during the course of the analysis itself is less easily
resolved. Given that the primary requirement is to find maximum
similarities between closely related sites, and maximum dissimilarities
between the groups, the analysis can frequently be made more powerful
by internal weighting. The principle involved is roughly as follows.
While the population of sites as a whole may possess a multiplicity of
attributes, any single site may in fact possess very few, so that the site
could contain very little information in its own right. By considering
79
general consensus of opinion and confine our attention to methods of
hierarchical classification as the primary approach.
A second decision to be made concerns the use of “subdivisive” or
“agglomerative” methods. The former begin with the whole population
of sites and divide it successively into smaller groups, each group being
examined independently for possible further subdivision as it is
extracted; the latter begin at the bottom and combine the individual
sites which are most alike until all individuals are eventually united
in a single population. Subdivisive methods thus concentrate essentially on differences, while agglomerative methods seek similarities. Two
points are of particular interest here. First, it should be noted that
subdivisive methods start from maximal information obtained over the
whole population, while agglomerative techniques start from single
units of minimal information. Secondly, subdivisive methods can be
terminated at any convenient level, while agglomerative methods require the whole analysis to be completed before the large-scale divisions
at the top of the hierarchy can be obtained. In general, therefore,
subdivisive methods are to be preferred on theoretical grounds, although
the actual calculations involved frequently require more computing
time.
A further choice lies between the use of “polythetic” or ‘‘monothetic” methods. Polythetic methods employ a combination of characters
to form the groups, while monothetic methods use only a single character
for each division; monothetic methods can only be subdivisive, but polythetic methods can be either subdivisive or agglomerative. The theoretical advantage of polythetic systems is that the classification obtained
is usually more stable and, by its nature, more informative; against this,
monothetic methods usually involve much less computation.
In addition to the above, there are yet further decisions to be made
on more subtle statistical considerations. These have been set out fully
by Williams and Dale in the parallel paper already mentioned, and need
only very brief consideration here. The &st concerns the question of
“internal weighting”. We have already (p. 71) condemned the practice
of external weighting on subjective estimations of the “most important”
species, but the problem of internal weighting of those species found to
be most informative during the course of the analysis itself is less easily
resolved. Given that the primary requirement is to find maximum
similarities between closely related sites, and maximum dissimilarities
between the groups, the analysis can frequently be made more powerful
by internal weighting. The principle involved is roughly as follows.
While the population of sites as a whole may possess a multiplicity of
attributes, any single site may in fact possess very few, so that the site
could contain very little information in its own right. By considering
