Bioinformatic Techniques on Marine Genomics 10.6 Large-Scale Sequence Analysis 299
Part B | 10.6
rine microbiota have attempted to describe: who is
there?; what are they doing?; who is doing what?, and
what evolutionary processes determine these parameters? Recent conceptual and technological advances
have allowed the increased use of sequence-guided
metagenomic investigations of the marine environment.
Recent advances in high-throughput sequencing technology and lower cost sequencing technologies have
made random shotgun sequencing of environmental
DNA economically feasible. The majority of metagenomic investigations to date have employed whole
metagenome shotgun sequencing approaches for the
cloning and sequencing of microbial DNA from marine environments. This involves the generation of
small insert DNA clone libraries, and their subsequent
analysis using Sanger dideoxy sequencing, generating
sequences that can then be used to query databases,
whereby phylogeny can be inferred or putative functional genes identified. Where putative protein encoding sequences are detected proximal to phylogenetic
marker genes on a single cloned fragment, function
can be linked to taxonomy. This approach can give
read lengths ranging from 600900 bp in length, which
can be extended through the entire fosmid clone. The
contigs generated using this approach give a fuller context for any given gene detected, thus giving a greater
probability of reconstructing the metabolic pathways
of individual members within a particular marine microbial consortium. While the benefits of the technology both from a sequence data generation and cost
standpoint (2540 Mb of DNA sequence per run at
an accuracy of 98%) are obvious, the limitations
of lower read lengths of between 200300 bp can be
problematic [10.23]. Read lengths achievable using pyrosequencing are currently around 450 bp, and this is
likely to continue to improve, which should improve the
amount of useful data generated from pyrosequencing
of metagenomics samples (Fig. 10.3).
10.6 Large-Scale Sequence Analysis
With the advent of the omics era, more and more
metagenomic and metatranscriptomic datasets are being generated from marine environments. While the
computational needs for single-genome analysis are
solved and can be performed on today’s commodity
hardware, the new large-scale metagenomic sequencing
projects which generate 20003000 genome equivalents of sequence information per project bring new
challenges. On the one hand, the challenge is the sheer
amount of sequence data per metagenome. On the other
hand, all metagenome sequences are mere fragments
of unknown organismal origin. Taken together, these
challenges demand further development of software for
assembly, gene calling, and annotation. Recently, several new and dedicated data processing and database
resources have emerged to address the current need
for large-scale metagenomic data analysis and management, for example, the community cyber infrastructure for advanced marine microbial ecology research
and analysis (CAMERA) and the meta genomics-rapid
annotation using subsystems technology (MG-RAST)
platform [10.24]. However, just a simple automatic annotation based on the basic local alignment search
tool (BLAST) sequence similarity searches poses a severe computational bottleneck. Consider the following
example: in November 2009, the MG-RAST server processed 278 metagenomes with an average of 33 Mbp
of sequence data per project. On a single high-end
server (Xeon E5540, 8 cpu’s, 2:53 GHz, 16 GB ram)
a complete BLAST analysis of a single project against
HITS
Metagenomic
library
Functional screen
for enzyme
activity
Homology screen –
e.g. DNA
hybridization,
bioinformatics
DNA sequence
analysis – whole
metagenome or
PCR amplicon
Environmetal sample
Sequence, expression, and functional analysis
DNA isolation
Fig. 10.4 Enzyme discovery from metagenomes: functional and
sequence-based approaches
Part B | 10.6
rine microbiota have attempted to describe: who is
there?; what are they doing?; who is doing what?, and
what evolutionary processes determine these parameters? Recent conceptual and technological advances
have allowed the increased use of sequence-guided
metagenomic investigations of the marine environment.
Recent advances in high-throughput sequencing technology and lower cost sequencing technologies have
made random shotgun sequencing of environmental
DNA economically feasible. The majority of metagenomic investigations to date have employed whole
metagenome shotgun sequencing approaches for the
cloning and sequencing of microbial DNA from marine environments. This involves the generation of
small insert DNA clone libraries, and their subsequent
analysis using Sanger dideoxy sequencing, generating
sequences that can then be used to query databases,
whereby phylogeny can be inferred or putative functional genes identified. Where putative protein encoding sequences are detected proximal to phylogenetic
marker genes on a single cloned fragment, function
can be linked to taxonomy. This approach can give
read lengths ranging from 600900 bp in length, which
can be extended through the entire fosmid clone. The
contigs generated using this approach give a fuller context for any given gene detected, thus giving a greater
probability of reconstructing the metabolic pathways
of individual members within a particular marine microbial consortium. While the benefits of the technology both from a sequence data generation and cost
standpoint (2540 Mb of DNA sequence per run at
an accuracy of 98%) are obvious, the limitations
of lower read lengths of between 200300 bp can be
problematic [10.23]. Read lengths achievable using pyrosequencing are currently around 450 bp, and this is
likely to continue to improve, which should improve the
amount of useful data generated from pyrosequencing
of metagenomics samples (Fig. 10.3).
10.6 Large-Scale Sequence Analysis
With the advent of the omics era, more and more
metagenomic and metatranscriptomic datasets are being generated from marine environments. While the
computational needs for single-genome analysis are
solved and can be performed on today’s commodity
hardware, the new large-scale metagenomic sequencing
projects which generate 20003000 genome equivalents of sequence information per project bring new
challenges. On the one hand, the challenge is the sheer
amount of sequence data per metagenome. On the other
hand, all metagenome sequences are mere fragments
of unknown organismal origin. Taken together, these
challenges demand further development of software for
assembly, gene calling, and annotation. Recently, several new and dedicated data processing and database
resources have emerged to address the current need
for large-scale metagenomic data analysis and management, for example, the community cyber infrastructure for advanced marine microbial ecology research
and analysis (CAMERA) and the meta genomics-rapid
annotation using subsystems technology (MG-RAST)
platform [10.24]. However, just a simple automatic annotation based on the basic local alignment search
tool (BLAST) sequence similarity searches poses a severe computational bottleneck. Consider the following
example: in November 2009, the MG-RAST server processed 278 metagenomes with an average of 33 Mbp
of sequence data per project. On a single high-end
server (Xeon E5540, 8 cpu’s, 2:53 GHz, 16 GB ram)
a complete BLAST analysis of a single project against
HITS
Metagenomic
library
Functional screen
for enzyme
activity
Homology screen –
e.g. DNA
hybridization,
bioinformatics
DNA sequence
analysis – whole
metagenome or
PCR amplicon
Environmetal sample
Sequence, expression, and functional analysis
DNA isolation
Fig. 10.4 Enzyme discovery from metagenomes: functional and
sequence-based approaches
