2 Metagenome Analysis
51
sufficient to back up genome sequencing projects (Goldberg et al. 2006), and for resequencing projects (Margulies et al. 2005). It was also used to obtain insights into
the metabolic equipment of single uncultured TM7 cells (Marcy et al. 2007) and
of about 600 cells of a single filament of uncultured marine Beggiatoa (Mussmann
et al. 2007). Pyrosequencing technology was also a great step forward in the analyses of the metabolic capabilities of microbial communities and the differences
between microbial assemblages present in habitats with distinctly different biogeochemical parameters (Angly et al. 2006, Edwards et al. 2006, Biddle et al. 2008,
Dinsdale et al. 2008). However, the phylogenetic affiliation of identified genes, a
binning of sequences belonging to distinct microbial species or clades and therefore
a reliable assignment of functions to distinct microbial groups in a mixed microbial
community was difficult due to short read lengths (Frias-Lopez et al. 2008). With
the 454 GS FLX system read length improved to 200–300 bp and more than 400 bp
reads can be achieved with the new 454 GS FLX Titanium system. Another issue
related to this technique is the higher error rate of pyrosequencing due to the poor
resolution of homopolymer stretches in template sequences (Margulies et al. 2005,
Huse et al. 2007, Wheeler et al. 2008). However, this can be markedly improved
by rigorous exclusion of low quality reads from the dataset (Margulies et al. 2005,
Huse et al. 2007, Wheeler et al. 2008).
2.4 Bioinformatic Challenges in Metagenome Analysis
The typical workflow for metagenomic data analysis follows the well established
analysis schema for full or draft genome sequencing projects (Glöckner and
Meyerdierks 2006, Stothard and Wishart 2006). After the raw reads have been
obtained, either from large or short insert size clone libraries, assembly is usually the
first step in data processing. If longer contigs (continuous sequences) or scaffolds
(contigs that still contain sequencing gaps) can be successfully established, gene
calling and subsequent annotation is performed to gain insights into the phylogenetic and functional diversity of the sample. To get an overview of the functional and
metabolic capacities represented in the sample, the protein coding genes are often
mapped against Subsystems (Overbeek et al. 2005), the Kyoto Encyclopaedia of
Gene and Genomes (KEGG) (Kanehisa et al. 2004) and the Clusters of Orthologous
Groups of proteins (COGs) (Tatusov et al. 1997). Comparative metagenomics
(Tringe et al. 2005, DeLong et al. 2006) and functional metagenomics approaches
such as metatranscriptomics (Poretsky et al. 2005, Frias-Lopez et al. 2008) and
metaproteomics (Ram et al. 2005, Wilmes and Bond 2006) are currently emerging
techniques that aim at obtaining a more dynamic understanding of the differences
and adaptations of the organisms to their environment.
Although processing metagenomic sequence data may seem to be straightforward a priori, this process is far from trivial in practise. In fact it is the same as trying
to reconstruct a puzzle with millions of pieces where most of them show a similar
colour and texture. When highly diverse environments, such as marine ecosystems,
are being studied and especially when the data produced corresponds to shallow
51
sufficient to back up genome sequencing projects (Goldberg et al. 2006), and for resequencing projects (Margulies et al. 2005). It was also used to obtain insights into
the metabolic equipment of single uncultured TM7 cells (Marcy et al. 2007) and
of about 600 cells of a single filament of uncultured marine Beggiatoa (Mussmann
et al. 2007). Pyrosequencing technology was also a great step forward in the analyses of the metabolic capabilities of microbial communities and the differences
between microbial assemblages present in habitats with distinctly different biogeochemical parameters (Angly et al. 2006, Edwards et al. 2006, Biddle et al. 2008,
Dinsdale et al. 2008). However, the phylogenetic affiliation of identified genes, a
binning of sequences belonging to distinct microbial species or clades and therefore
a reliable assignment of functions to distinct microbial groups in a mixed microbial
community was difficult due to short read lengths (Frias-Lopez et al. 2008). With
the 454 GS FLX system read length improved to 200–300 bp and more than 400 bp
reads can be achieved with the new 454 GS FLX Titanium system. Another issue
related to this technique is the higher error rate of pyrosequencing due to the poor
resolution of homopolymer stretches in template sequences (Margulies et al. 2005,
Huse et al. 2007, Wheeler et al. 2008). However, this can be markedly improved
by rigorous exclusion of low quality reads from the dataset (Margulies et al. 2005,
Huse et al. 2007, Wheeler et al. 2008).
2.4 Bioinformatic Challenges in Metagenome Analysis
The typical workflow for metagenomic data analysis follows the well established
analysis schema for full or draft genome sequencing projects (Glöckner and
Meyerdierks 2006, Stothard and Wishart 2006). After the raw reads have been
obtained, either from large or short insert size clone libraries, assembly is usually the
first step in data processing. If longer contigs (continuous sequences) or scaffolds
(contigs that still contain sequencing gaps) can be successfully established, gene
calling and subsequent annotation is performed to gain insights into the phylogenetic and functional diversity of the sample. To get an overview of the functional and
metabolic capacities represented in the sample, the protein coding genes are often
mapped against Subsystems (Overbeek et al. 2005), the Kyoto Encyclopaedia of
Gene and Genomes (KEGG) (Kanehisa et al. 2004) and the Clusters of Orthologous
Groups of proteins (COGs) (Tatusov et al. 1997). Comparative metagenomics
(Tringe et al. 2005, DeLong et al. 2006) and functional metagenomics approaches
such as metatranscriptomics (Poretsky et al. 2005, Frias-Lopez et al. 2008) and
metaproteomics (Ram et al. 2005, Wilmes and Bond 2006) are currently emerging
techniques that aim at obtaining a more dynamic understanding of the differences
and adaptations of the organisms to their environment.
Although processing metagenomic sequence data may seem to be straightforward a priori, this process is far from trivial in practise. In fact it is the same as trying
to reconstruct a puzzle with millions of pieces where most of them show a similar
colour and texture. When highly diverse environments, such as marine ecosystems,
are being studied and especially when the data produced corresponds to shallow
