94
D.E. Soltis and P.S. Soltis
exercise (e.g., Chase and Albert 1998; Chase and Cox 1998; and Soltis et ai. 1997b,
1998).
The initial data sets compiled for 18S rDNA, rbeL, and atpB did not "match"
completely in terms of the exemplars used. However, via careful planning and collaboration, these three genes have now been sequenced for a nearly identical suite
of taxa (often using the same DNA) for 560 angiosperms, as well as seven outgroups.
This 567-taxon data set has been analyzed via several approaches: parsimony as
implemented by PAUP*4.0, the ratchet approach of Nixon (unpubI.), and the parsimony jackknife method of Farris et ai. (1996) (Fig. 1).
Several initial combinations of data sets contributed greatly to our understanding of how to solve large phylogenetic problems. These include the combination of
18S rDNA and rbeL sequences (Soltis et ai. 1997b) and rbeL and atpB (Savolainen
et aI., in press), and 193 taxa for 18S rDNA, atpB, and rbeL (Soltis et ai. 1998).
Other noteworthy data sets and combined analyses include the rbeL and
nonmolecular data sets compiled and analyzed by Nandi et ai. (1998). Below we
summarize some of the approaches used in the analysis of large data sets in general, with an emphasis on recent results for angiosperms.
3 Add Taxa and Add Characters
Both simulation and empirical studies indicate that adding taxa and increasing the
number of characters (base pairs in this case) in phylogenetic analyses may not only
increase the accuracy of the estimated trees, but also reduce the computational difficulty. These results are in contrast to earlier suggestions that large data sets may
be extremely complex and difficult to analyze phylogenetically.
Several simulation studies have revealed the importance of adding taxa. Using
the 228-taxon 18S rDNA data set for angiosperms and shortest trees obtained (Soltis
et al. 1997a) as the basis for simulation studies, Hillis (1996) suggested that the
phylogenetic analysis of large data sets may be more tractable than he and coworkers had previously suggested based on simulations involving only four taxa (e.g.,
Hillis et al. 1994). The more recent simulation studies of Graybeal (1998) further
support this contention. These studies suggest that adding taxa, while perhaps
counterintuitive, made phylogenetic analyses more straightforward, apparently because the addition of taxa breaks up long branches and disperses homoplasy.
Empirical studies further demonstrate the importance of adding both taxa and
characters in the phylogenetic analysis of large data sets. Soltis et al. (1998) conducted heuristic parsimony searches and fast bootstrap analyses on separate and
combined DNA data sets for 190 angiosperms and three outgroups; separate data
sets of 18S rDNA, rbeL and atpB sequences were combined into a single matrix
4,733 bp in length. Analyses of the combined 193-taxon data set revealed great
improvements in computer run times compared to the analyses of the separate data
sets and the data sets combined in pairs. Six analyses of the combined three-gene
data set were conducted, and in all cases TBR branch swapping was completed, on
D.E. Soltis and P.S. Soltis
exercise (e.g., Chase and Albert 1998; Chase and Cox 1998; and Soltis et ai. 1997b,
1998).
The initial data sets compiled for 18S rDNA, rbeL, and atpB did not "match"
completely in terms of the exemplars used. However, via careful planning and collaboration, these three genes have now been sequenced for a nearly identical suite
of taxa (often using the same DNA) for 560 angiosperms, as well as seven outgroups.
This 567-taxon data set has been analyzed via several approaches: parsimony as
implemented by PAUP*4.0, the ratchet approach of Nixon (unpubI.), and the parsimony jackknife method of Farris et ai. (1996) (Fig. 1).
Several initial combinations of data sets contributed greatly to our understanding of how to solve large phylogenetic problems. These include the combination of
18S rDNA and rbeL sequences (Soltis et ai. 1997b) and rbeL and atpB (Savolainen
et aI., in press), and 193 taxa for 18S rDNA, atpB, and rbeL (Soltis et ai. 1998).
Other noteworthy data sets and combined analyses include the rbeL and
nonmolecular data sets compiled and analyzed by Nandi et ai. (1998). Below we
summarize some of the approaches used in the analysis of large data sets in general, with an emphasis on recent results for angiosperms.
3 Add Taxa and Add Characters
Both simulation and empirical studies indicate that adding taxa and increasing the
number of characters (base pairs in this case) in phylogenetic analyses may not only
increase the accuracy of the estimated trees, but also reduce the computational difficulty. These results are in contrast to earlier suggestions that large data sets may
be extremely complex and difficult to analyze phylogenetically.
Several simulation studies have revealed the importance of adding taxa. Using
the 228-taxon 18S rDNA data set for angiosperms and shortest trees obtained (Soltis
et al. 1997a) as the basis for simulation studies, Hillis (1996) suggested that the
phylogenetic analysis of large data sets may be more tractable than he and coworkers had previously suggested based on simulations involving only four taxa (e.g.,
Hillis et al. 1994). The more recent simulation studies of Graybeal (1998) further
support this contention. These studies suggest that adding taxa, while perhaps
counterintuitive, made phylogenetic analyses more straightforward, apparently because the addition of taxa breaks up long branches and disperses homoplasy.
Empirical studies further demonstrate the importance of adding both taxa and
characters in the phylogenetic analysis of large data sets. Soltis et al. (1998) conducted heuristic parsimony searches and fast bootstrap analyses on separate and
combined DNA data sets for 190 angiosperms and three outgroups; separate data
sets of 18S rDNA, rbeL and atpB sequences were combined into a single matrix
4,733 bp in length. Analyses of the combined 193-taxon data set revealed great
improvements in computer run times compared to the analyses of the separate data
sets and the data sets combined in pairs. Six analyses of the combined three-gene
data set were conducted, and in all cases TBR branch swapping was completed, on
