6 Genomics of Marine Algae
189
sea and, in a later study, from a transect that extended along the eastern border of
the North Atlantic through the Panama canal to the South Pacific (the Sorcerer II
Global Ocean Survey). In total, more than 9.3 million sequencing reads were carried out on 2–6 kbp DNA fragments cloned from these samples, generating more
than 6.3 billion base pairs of sequence. This approach provides a wealth of information about the diversity of organisms present in a particular environment. For
example, the Sargasso Sea samples alone were estimated to contain samples from
1,800 species including 148 previously unknown bacterial phylotypes. The interest
of high-throughput sequencing is not limited to the investigation of species diversity,
however. The sequence data also furnishes information at the gene level, leading to
the discovery of new genes and providing an overview of the ensemble of genes
present in a community. This global collection of genes, often described as the
metagenome, can provide important insights, for example, into the different functional pathways that are represented and relative importance of each pathway in a
particular environment.
The studies carried out by Craig Venter and colleagues focused on the bacterial components of the ecosystems studied. To do this they employed filters which
both eliminated smaller entities, such as viruses, and larger organisms, such as large
eukaryotic cells (Venter et al. 2004, Rusch et al. 2007). The samples did however contain a significant proportion of picoeukaryotes, which have approximately
the same size as bacteria (2.8% of the sequences in the Sorcerer II study). These
sequences represent species that are broadly dispersed throughout the eukaryotic
tree of life (Piganeau et al. 2008). The studies have therefore provided a considerable amount of sequence data and information about planktonic algae. Perhaps
more importantly, however, these two analyses clearly show the potential of highthroughput sequencing as a means to explore the diversity of marine algae and the
ecosystems in which they live. Moreover, samples at other filter sizes were collected
during these surveys so it will be possible to carry out a more complete analysis of
the algal component of these ecosystems at a later date.
As with all novel approaches, high-throughput environmental sequencing also
creates new problems to be solved. In particular, the assembly of sequence data from
ecosystems with a high level of biodiversity such as the plankton can pose a serious
challenge because, despite the large number of sequences generated, the depth of
sequencing is much less than when the sequencing is focused on a single genome.
In this respect, more directed, organism-based approaches can provide information that is highly complementary to that generated by environmental sequencing.
For example, two complete genomes have been published for prasinophytes of the
genus Ostreococcus since the data from the Sargasso Sea survey was made available. These genome sequences have been used to recover related sequences from
the Sargasso Sea dataset and it has been possible to show that at least two species of
Ostreococcus were represented in these samples (Piganeau and Moreau 2007). The
two types of information are highly complementary, the genome sequences allowing
the recovery of species-specific data from the environmental dataset and the environmental dataset providing additional information about both the ecology of the
organism that has been sequenced and about the molecular evolution of its genome,
based on sequence comparisons.
189
sea and, in a later study, from a transect that extended along the eastern border of
the North Atlantic through the Panama canal to the South Pacific (the Sorcerer II
Global Ocean Survey). In total, more than 9.3 million sequencing reads were carried out on 2–6 kbp DNA fragments cloned from these samples, generating more
than 6.3 billion base pairs of sequence. This approach provides a wealth of information about the diversity of organisms present in a particular environment. For
example, the Sargasso Sea samples alone were estimated to contain samples from
1,800 species including 148 previously unknown bacterial phylotypes. The interest
of high-throughput sequencing is not limited to the investigation of species diversity,
however. The sequence data also furnishes information at the gene level, leading to
the discovery of new genes and providing an overview of the ensemble of genes
present in a community. This global collection of genes, often described as the
metagenome, can provide important insights, for example, into the different functional pathways that are represented and relative importance of each pathway in a
particular environment.
The studies carried out by Craig Venter and colleagues focused on the bacterial components of the ecosystems studied. To do this they employed filters which
both eliminated smaller entities, such as viruses, and larger organisms, such as large
eukaryotic cells (Venter et al. 2004, Rusch et al. 2007). The samples did however contain a significant proportion of picoeukaryotes, which have approximately
the same size as bacteria (2.8% of the sequences in the Sorcerer II study). These
sequences represent species that are broadly dispersed throughout the eukaryotic
tree of life (Piganeau et al. 2008). The studies have therefore provided a considerable amount of sequence data and information about planktonic algae. Perhaps
more importantly, however, these two analyses clearly show the potential of highthroughput sequencing as a means to explore the diversity of marine algae and the
ecosystems in which they live. Moreover, samples at other filter sizes were collected
during these surveys so it will be possible to carry out a more complete analysis of
the algal component of these ecosystems at a later date.
As with all novel approaches, high-throughput environmental sequencing also
creates new problems to be solved. In particular, the assembly of sequence data from
ecosystems with a high level of biodiversity such as the plankton can pose a serious
challenge because, despite the large number of sequences generated, the depth of
sequencing is much less than when the sequencing is focused on a single genome.
In this respect, more directed, organism-based approaches can provide information that is highly complementary to that generated by environmental sequencing.
For example, two complete genomes have been published for prasinophytes of the
genus Ostreococcus since the data from the Sargasso Sea survey was made available. These genome sequences have been used to recover related sequences from
the Sargasso Sea dataset and it has been possible to show that at least two species of
Ostreococcus were represented in these samples (Piganeau and Moreau 2007). The
two types of information are highly complementary, the genome sequences allowing
the recovery of species-specific data from the environmental dataset and the environmental dataset providing additional information about both the ecology of the
organism that has been sequenced and about the molecular evolution of its genome,
based on sequence comparisons.
