98
D.E. Soltis and P.S. Soltis
Several studies have employed the parsimony jackknife approach and illustrate
its utility, as well as its speed. Soltis et al. (1997a) used the original lac program
without branch swapping in their analysis of 228 angiosperm 18S rDNA sequences.
The parsimony jackknife analysis of this data set (based on 1000 replicates) was
completed in only 949 seconds. An extremely promising example of the applicability of the parsimony jackknife is that of Kallersjo et al. (in press), who analyzed
2538 rbcL sequences representing a broad taxonomic range, from cyanobacteria to
flowering plants. Kallersjo et al. used both the original and newer versions of the
lac program, with surprising results - numerous (a total of 1400) clades with lac
values 2: 50% were retrieved, including major clades such as green plants, land
plants, angiosperms, and eudicots. With regard to the two lac programs Farris and
coworkers have provided, Kallersjo et al. (in press) found that while branch-swapping improved both resolution and the support for some clades, most of the groups
were recovered by the original program (which is faster and simpler). The original
lac program required 356 seconds per replicate; the new version with branchswapping required approximately 1304.5 seconds per replicate. The initial analysis
of the 2538-sequence matrix using the original lac program with 1000 replicates
required 99 hours on a 133 MHz Pentium computer. Because the newer lac program has a stop-restart function (which was employed by Kallersjo et al., in press),
the amount of time needed to complete the analysis using the new lac program was
not precisely known (Kallersjo et aI., in press), but can be estimated to be roughly
360 hours for 1000 replicates. Soltis et at. (submitted) have applied the new lac
program (with the help of S. Farris) to the 567-taxon, three-gene angiosperm data
set; 1000 replicates were completed in 60.63 hours. Application of the parsimony
jackknife approach to this 567-taxon data set yielded a topology with few nodes
lacking support 2: 50%; even the spine of the tree is generally well supported (Fig.
1). In this case, one could legitimately ask if lengthy parsimony searches are even
necessary.
5 Developments in Parsimony Analysis
The numerous improvements and new options available in PAUP* 4.0 have been a
great asset to those interested in analyzing large data sets. One obvious improvement is the faster speed with which PAUP* 4.0 (Swofford 1998) conducts heuristic
parsimony searches compared to version 3.1. Similarly, improvements in the maximum likelihood (ML) program as implemented in PAUP* 4.0 have made it possible to analyze much larger data sets than before using this approach. Although
ML analyses are now possible with data sets up to 50-60 taxa, these data sets are
not technically "large" as considered for this review (2: 150 taxa). ML analysis is
still not feasible with truly large data sets.
Algorithms for finding shortest trees such as Hennig 86, (Farris 1988), NONA
(Goloboff 1993), and PAUP (Swofford 1993) were developed for working on what
we consider here to be small data sets. Perhaps one of the most important develop-
D.E. Soltis and P.S. Soltis
Several studies have employed the parsimony jackknife approach and illustrate
its utility, as well as its speed. Soltis et al. (1997a) used the original lac program
without branch swapping in their analysis of 228 angiosperm 18S rDNA sequences.
The parsimony jackknife analysis of this data set (based on 1000 replicates) was
completed in only 949 seconds. An extremely promising example of the applicability of the parsimony jackknife is that of Kallersjo et al. (in press), who analyzed
2538 rbcL sequences representing a broad taxonomic range, from cyanobacteria to
flowering plants. Kallersjo et al. used both the original and newer versions of the
lac program, with surprising results - numerous (a total of 1400) clades with lac
values 2: 50% were retrieved, including major clades such as green plants, land
plants, angiosperms, and eudicots. With regard to the two lac programs Farris and
coworkers have provided, Kallersjo et al. (in press) found that while branch-swapping improved both resolution and the support for some clades, most of the groups
were recovered by the original program (which is faster and simpler). The original
lac program required 356 seconds per replicate; the new version with branchswapping required approximately 1304.5 seconds per replicate. The initial analysis
of the 2538-sequence matrix using the original lac program with 1000 replicates
required 99 hours on a 133 MHz Pentium computer. Because the newer lac program has a stop-restart function (which was employed by Kallersjo et al., in press),
the amount of time needed to complete the analysis using the new lac program was
not precisely known (Kallersjo et aI., in press), but can be estimated to be roughly
360 hours for 1000 replicates. Soltis et at. (submitted) have applied the new lac
program (with the help of S. Farris) to the 567-taxon, three-gene angiosperm data
set; 1000 replicates were completed in 60.63 hours. Application of the parsimony
jackknife approach to this 567-taxon data set yielded a topology with few nodes
lacking support 2: 50%; even the spine of the tree is generally well supported (Fig.
1). In this case, one could legitimately ask if lengthy parsimony searches are even
necessary.
5 Developments in Parsimony Analysis
The numerous improvements and new options available in PAUP* 4.0 have been a
great asset to those interested in analyzing large data sets. One obvious improvement is the faster speed with which PAUP* 4.0 (Swofford 1998) conducts heuristic
parsimony searches compared to version 3.1. Similarly, improvements in the maximum likelihood (ML) program as implemented in PAUP* 4.0 have made it possible to analyze much larger data sets than before using this approach. Although
ML analyses are now possible with data sets up to 50-60 taxa, these data sets are
not technically "large" as considered for this review (2: 150 taxa). ML analysis is
still not feasible with truly large data sets.
Algorithms for finding shortest trees such as Hennig 86, (Farris 1988), NONA
(Goloboff 1993), and PAUP (Swofford 1993) were developed for working on what
we consider here to be small data sets. Perhaps one of the most important develop-
