4 Phylogeny of Animals
127
Another approach used was to to employ alternative marker genes such as the
large subunit of ribosomal RNA (LSU or 28S), often in combination with other
genes such as SSU (Mallatt and Winchell 2002, Winchell et al. 2002), elongation factor I alpha (EFIa) (Littlewood et al. 2001), heat-shock proteins (Borchiellini
et al. 1998) or sodium-potassium ATPase beta-subunit (Anderson et al. 2004). This
approach has provided valuable insights into the phylogeny of metazoans, but as
a whole the trees based on each of the individual genes exhibited very incongruent
topologies. The inference of an improved metazoan tree based on nuclear genes thus
requires that this problem of incongruence be settled.
4.3 The Power and Pitfalls of Phylogenomics
Genome is the far side of all known organisms and represent inestimable source of
phylogenetic information. Indeed, one can consider that morphological features are
often shaped by strong evolutionary pressure whereas reliable arguments indicate
that genome evolution is an essentially neutralistic process (see Kimura 1983, Lynch
2007). Two main classes of phylogenetic evidence could be drawn from genomic
data: the first and most quantitative consists of inferring trees from the sequences of
large sets of nuclear genes sampled in the genomic data. A second, more qualitative,
type of evidence consists of detecting discrete molecular signatures by surveying
whole genome features (Philippe et al. 2005a).
The first approach is based on the analysis of primary sequences from a large set
of genes. This approach attempts to overcome the incongruence problem observed
when the topologies obtained from independent marker genes are compared (cf.
supra). It has been proposed that the concatenation of a large number of nuclear
genes enables the detection of the bona fide species tree among the alternative
topologies that are retrieved for different genes (Rokas et al. 2003). However, the
large number of positions considered tends to increase systematic biases, which
could lead to well-supported but inaccurate phylogenies (Jeffroy et al. 2006). Such
systematic biases are mainly related to the heterogeneity of evolutionary rates
and sequence composition, resulting in differences between taxa and sites of the
alignment (Philippe et al. 2005a). It has been proposed that such biases could
be limited by using improved models of substitution and by increasing the taxon
sampling (Delsuc et al. 2005). For instance, the long-branch attraction problems
could be reduced by using the CAT model that has recently been developed to
account for the heterogeneity of sites along the alignment (Lartillot and Philippe
2004).
Two strategies can be employed to analyze such gene-rich datasets: the supertree
or the supermatrix (Delsuc et al. 2005). The supertree approach infers an independent tree for each marker gene, with the subsequent computation of a supertree
that summarizes all the independent topologies recovered. This approach limits
127
Another approach used was to to employ alternative marker genes such as the
large subunit of ribosomal RNA (LSU or 28S), often in combination with other
genes such as SSU (Mallatt and Winchell 2002, Winchell et al. 2002), elongation factor I alpha (EFIa) (Littlewood et al. 2001), heat-shock proteins (Borchiellini
et al. 1998) or sodium-potassium ATPase beta-subunit (Anderson et al. 2004). This
approach has provided valuable insights into the phylogeny of metazoans, but as
a whole the trees based on each of the individual genes exhibited very incongruent
topologies. The inference of an improved metazoan tree based on nuclear genes thus
requires that this problem of incongruence be settled.
4.3 The Power and Pitfalls of Phylogenomics
Genome is the far side of all known organisms and represent inestimable source of
phylogenetic information. Indeed, one can consider that morphological features are
often shaped by strong evolutionary pressure whereas reliable arguments indicate
that genome evolution is an essentially neutralistic process (see Kimura 1983, Lynch
2007). Two main classes of phylogenetic evidence could be drawn from genomic
data: the first and most quantitative consists of inferring trees from the sequences of
large sets of nuclear genes sampled in the genomic data. A second, more qualitative,
type of evidence consists of detecting discrete molecular signatures by surveying
whole genome features (Philippe et al. 2005a).
The first approach is based on the analysis of primary sequences from a large set
of genes. This approach attempts to overcome the incongruence problem observed
when the topologies obtained from independent marker genes are compared (cf.
supra). It has been proposed that the concatenation of a large number of nuclear
genes enables the detection of the bona fide species tree among the alternative
topologies that are retrieved for different genes (Rokas et al. 2003). However, the
large number of positions considered tends to increase systematic biases, which
could lead to well-supported but inaccurate phylogenies (Jeffroy et al. 2006). Such
systematic biases are mainly related to the heterogeneity of evolutionary rates
and sequence composition, resulting in differences between taxa and sites of the
alignment (Philippe et al. 2005a). It has been proposed that such biases could
be limited by using improved models of substitution and by increasing the taxon
sampling (Delsuc et al. 2005). For instance, the long-branch attraction problems
could be reduced by using the CAT model that has recently been developed to
account for the heterogeneity of sites along the alignment (Lartillot and Philippe
2004).
Two strategies can be employed to analyze such gene-rich datasets: the supertree
or the supermatrix (Delsuc et al. 2005). The supertree approach infers an independent tree for each marker gene, with the subsequent computation of a supertree
that summarizes all the independent topologies recovered. This approach limits
