92
D.E. Soltis and P.S. Soltis
the best-resolved and best-supported topology to date for angiosperms, with virtually all major clades, as well as the spine of the tree, well supported. These developments indicate that the phylogenetic analysis of large data sets is not only feasible,
but relatively straightforward.
1 Introduction
A firm understanding of phylogenetic relationships is obviously central to elucidating many questions involving evolution; a solid phylogenetic underpinning is also
critical for a better understanding of biodiversity. Assessing phylogenetic relationships in many groups will require the compilation and phylogenetic analysis of
data sets (DNA sequences and/or nonmolecular traits) for numerous taxa. This is
particularly true, for example, in many large groups of organisms for which relationships have remained obscure, sometimes despite decades of study. Obvious examples include fungi, bacteria, insects, green plants, and large subclades within the
green plants such as land plants, ferns, and angiosperms. However, the need for
phylogenetic analysis of large data sets is not restricted to higher taxonomic groups,
but may involve any portion of the taxonomic hierarchy, extending to the population level, or even to the level of "strains" of bacteria or viruses.
For convenience, we will arbitrarily define a large data set as one having over
150 placeholders. Although the phylogenetic analysis of large data sets often is
central to understanding relationships within many groups, the feasibility of analysis of these data sets has been much debated (reviewed in Hillis 1996; Graybeal
1998; Soltis et al. 1998). For example, large data sets are problematic in parsimony
searches because of the enormous number of trees possible for a large number of
terminal taxa-the number of potential solutions increases logarithmically as taxa
are added (Felsenstein 1978a). Hence, for just 20 taxa the number of possible rooted
trees (8.87 x 10 23 ) slightly exceeds Avogadro's number-that is, roughly one mole
of trees. For large data sets involving hundreds of taxa, the number of possible trees
likely exceeds the number of atoms in the universe (Hillis, pers. comm.). A basic
question, therefore, is how can we possibly examine a universe of possible trees,
and be reasonably certain of the solution that we select?
Early simulation studies also suggested that the phylogenetic analysis of large
data sets was impractical or impossible. For example, in some instances (those
involving extreme branch-length heterogeneity), the correct reconstruction of phylogeny for only four taxa requires over 10,000 base pairs (bp) of sequence data
(Hillis et al. 1994). Such problems and complexity with only four taxa prompted
some to propose that large phylogenetic problems be broken into a series of smaller
problems (e.g., Kim 1996; Mishler 1994; Rice et al. 1997), with one extreme view
being to break large data sets into a large number of four-taxon problems (Graur et
al. 1996).
Another problem posed by large data sets involves the assessment of support for
individual clades. Two commonly used procedures for estimating branch support
Précédent

- 102/321

Suivant