6. Phylogenetic Analyses of Large Data Sets
97
the 500-sequence rbcL data set for angiosperms and other seed plants (Chase and
Albert 1998; Kallersjo et aI., in press). Internal support and resolution are also
much higher than achieved for large data sets of 18S rDNA alone (Soltis et al.
1997a) or for atpB and atpB plus rbcL (Savolainen et aI., in press).
4 Quick Searches for Well-Supported Clades
A significant problem with the analysis of large data sets is the difficulty in assessing internal support for clades. As reviewed earlier, large data sets are not amenable to standard bootstrapping (Felsenstein 1985) and decay or Bremer support
analysis (Bremer 1988). However, programs that conduct fast bootstrap and fast
jackknife analyses are available and were designed with large data sets in mind.
The fast jackknife (or parsimony jackknife) was recommended by Farris et al. (1996).
A fast jackknife program is also available on PAUP*4.0, and the two methods yield
similar values (Mort et aI., in press).
Several studies illustrate the value of conducting relatively quick searches and
"saving" only those clades above some threshhold of internal support. Below a
minimal threshhold (e.g., a bootstrap or jackknife value of 50%), confidence in a
clade is low or nonexistant. Hence, why conduct extensive, time-consuming parsimony searches looking for shorter and shorter trees when continued branch-swapping does not result in increased support for weakly-supported clades? The wellsupported clades appear relatively quickly in the analysis of big data sets - continued branch-swapping involving poorly-supported branches does not suddenly result in strongly-supported clades. Lengthy parsimony analyses may, in fact, be a
waste of time. This is exemplified well by the reanalysis of the 500-sequence rbcL
matrix (Chase et al. 1993) by Rice et al. (1997). The initial searches of this large
data set by Chase et al. (1993) did not swap to completion; furthermore, the authors
realized trees shorter than those that they had found existed (reviewed in Chase and
Albert 1998). Lengthy reanalysis of this data set (Rice et al. 1997) did find shorter
trees than those reported by Chase et al. (1993); however, these searches did not
result in a significantly different topology from that provided initially by Chase et
al. (1993).
The parsimony jackknife method (Farris 1996; Farris et al. 1996) is well suited
for the "quick" analysis of large data sets and has been applied to a data set of 2538
rbcL sequences (Kallersjo et aI., in press). The approach uses jackknifing (resampling
of characters without replacement and stepwise addition of taxa without branchswapping). A consensus of the trees generated by the jackknife replications depicts
only well-supported clades (usually with support of 50% or greater); thus, the approach provides a rapid assessment of branch support. A second version of the J ac
program is now available (Farris et aI., in prep); it has a faster tree-building algorithm, and also allows branch swapping to be conducted at each replication. With
the second version, it is also possible to perform several random-addition sequences
per replicate. The number of replications performed is typically 500 or 1000.
Précédent

- 107/321

Suivant