6. Phylogenetic Analyses of Large Data Sets
93
are the bootstrap (Felsenstein 1985) and Bremer support or decay index (Bremer
1988). Both are time-consuming, however, and impractical for large data sets given
software presently available.
Early parsimony analyses of large data sets seemed to bear out, in part, the dire
predictions of computational difficulty for large data sets (e.g., Felsenstein 1978;
Graur et al. 1996; Hillis et al. 1994). For example, parsimony searches of a large
data set of 500 rbcL sequences did not swap to completion (Chase et al. 1993), even
after large investments of computer time. Rice et al. (1997) devoted "approximately
11.6 months of CPU time" using three Sun workstations in a reanalysis of the rbcL
data set of Chase et aI., yet none of their searches swapped to completion. Similarly, Soltis et al. (1997a) invested over two years of computer time in analyses of a
228-taxon data set of 18S rDNA sequences for angiosperms; again, no searches
swapped to completion. Savolainen et al. (in press) encountered similar difficulties
in phylogenetic analyses of a large data set of hundreds of atpB sequences representing the angiosperms.
Significantly, however, recent empirical and simulation studies suggest that
large data sets are much more tractable than thought only a few years ago. Furthermore, much of what systematists have learned about "hands-on" analyses of large
data sets has been garnered via the study of several large DNA sequence data sets
compiled for angiosperms. Herein we attempt to review some of the approaches
that can be used in the analysis of large data sets. We rely heavily on work recently
completed using the three large DNA sequence data sets compiled for angiosperms
to exemplify several of the possible solutions to the analytical problems posed by
large data sets. We highlight here three general approaches (that are not mutually
exclusive) to the analysis of large data sets encompassing hundreds of taxa: 1) the
addition of taxa and characters; 2) the use of "fast" or "quick" searches such as the
fast bootstrap and parsimony jackknife; and 3) the application of computer programs such as NONA and the RATCHET to facilitate faster searches of tree space.
We also discuss two other approaches to the analysis of large data sets: 4) the
supertree method, and 5) compartmentalization.
2 Background: Angiosperm Data Sets
Three large DNA sequence data sets have been assembled for angiosperms: 18S
rDNA (1,855bp), rbcL (1,428bp), and atpB (1,450bp). From a historical perspective (reviewed in Soltis et al. 1997b; Chase and Albert 1998; Chase and Cox 1998),
rbcL was the first gene to be sequenced widely in angiosperms (Chase et al. 1993),
followed by 18S rDNA (Soltis et al. 1997a), and most recently atpB (Savolainen et
al., in press). As noted, parsimony searches of these individual data sets did not
swap to completion due to their large size. Nonetheless, the topologies obtained
based on phylogenetic analysis of these individual genes were strikingly similar,
prompting investigators to suggest that these genes were tracking the same
organismal phylogeny and that analysis of large data sets might be a productive
Précédent

- 103/321

Suivant