2 Metagenome Analysis
59
A commonly used alternative to show an overview of the taxonomic composition of the metagenome, is the analysis of the taxonomic affiliation of the best hit
by searching for similarities of the contigs or genes to the UniProt or nr databases
(Treusch et al. 2004, DeLong et al. 2006, Turnbaugh et al. 2006). Due to misclassifications of the sequences in the databases this is often not very accurate and curated
genome databases like the genomesDB provided by GenBank and implemented in
the JCoast system (Richter et al. 2008) are preferable for taxonomic breakdowns.
A recently introduced alternative is based on the mapping of single-copy or equalcopy marker genes onto a reference species phylogeny (Ciccarelli et al. 2006, von
Mering et al. 2007). The advantage of the system is that 31 phylogenetic marker
genes can be used, compared to only two available markers when rRNA sequences
are applied. It is important to note that the coverage of the phylogenetic protein
marker databases is still rather small (only around 600 per gene) when compared
to the more than 800,000 publicly available rRNA genes. Nevertheless, the two
methods provide different views of the data and should therefore be regarded as
complementary (Raes et al. 2007).
2.4.7.2 Functional Diversity
As described above, functional diversity should be primarily accessed and described
by classical annotation and metabolic reconstruction approaches involving manual
curation. Unfortunately, the lack of common standards means that there is a significant difference in the quality of data handling and interpretation and this can hamper
comparative metagenomic analyses (Raes et al. 2007). Furthermore, the incredible
amount of data produced by metagenomic sequencing campaigns often renders time
consuming approaches unrealistic. In consequence, the analysis is often limited to
the determination of general statistical descriptors such as over- and underrepresentation of genes (Tringe et al. 2005) or the richness, membership and structure
of microbial communities (Schloss and Handelsman 2008). These approaches are
often the only available option for obtaining comparative insights into the functional adaptations of the populations, especially when only shallow or short read
length sequencing data is available for several sampling sites. Successful examples
of this sort of analysis include the investigation of stratified microbial assemblages
in the North Pacific Subtropical Gyre (DeLong et al. 2006), the comparison of
mesocosms amended with DMSP (Mou et al. 2008) and the investigation of several viral communities (Edwards and Rohwer 2005, Angly et al. 2006, Culley et al.
2006).
In summary, for high diversity environments a “gene-centric” approach invoking
the descriptors mentioned above, plus additional ones like differences in functional categories, is useful to obtain a better view of the ecology and function of
the microbial communities. With low diversity communities, where only a few
dominating species are present, the classical “genome-centric” approach is often
preferable because this provides more detailed information. The assignment of the
assembled contigs to organism bins combined with subsequent annotation usually enables the reconstruction of individual metabolic properties and leads to
59
A commonly used alternative to show an overview of the taxonomic composition of the metagenome, is the analysis of the taxonomic affiliation of the best hit
by searching for similarities of the contigs or genes to the UniProt or nr databases
(Treusch et al. 2004, DeLong et al. 2006, Turnbaugh et al. 2006). Due to misclassifications of the sequences in the databases this is often not very accurate and curated
genome databases like the genomesDB provided by GenBank and implemented in
the JCoast system (Richter et al. 2008) are preferable for taxonomic breakdowns.
A recently introduced alternative is based on the mapping of single-copy or equalcopy marker genes onto a reference species phylogeny (Ciccarelli et al. 2006, von
Mering et al. 2007). The advantage of the system is that 31 phylogenetic marker
genes can be used, compared to only two available markers when rRNA sequences
are applied. It is important to note that the coverage of the phylogenetic protein
marker databases is still rather small (only around 600 per gene) when compared
to the more than 800,000 publicly available rRNA genes. Nevertheless, the two
methods provide different views of the data and should therefore be regarded as
complementary (Raes et al. 2007).
2.4.7.2 Functional Diversity
As described above, functional diversity should be primarily accessed and described
by classical annotation and metabolic reconstruction approaches involving manual
curation. Unfortunately, the lack of common standards means that there is a significant difference in the quality of data handling and interpretation and this can hamper
comparative metagenomic analyses (Raes et al. 2007). Furthermore, the incredible
amount of data produced by metagenomic sequencing campaigns often renders time
consuming approaches unrealistic. In consequence, the analysis is often limited to
the determination of general statistical descriptors such as over- and underrepresentation of genes (Tringe et al. 2005) or the richness, membership and structure
of microbial communities (Schloss and Handelsman 2008). These approaches are
often the only available option for obtaining comparative insights into the functional adaptations of the populations, especially when only shallow or short read
length sequencing data is available for several sampling sites. Successful examples
of this sort of analysis include the investigation of stratified microbial assemblages
in the North Pacific Subtropical Gyre (DeLong et al. 2006), the comparison of
mesocosms amended with DMSP (Mou et al. 2008) and the investigation of several viral communities (Edwards and Rohwer 2005, Angly et al. 2006, Culley et al.
2006).
In summary, for high diversity environments a “gene-centric” approach invoking
the descriptors mentioned above, plus additional ones like differences in functional categories, is useful to obtain a better view of the ecology and function of
the microbial communities. With low diversity communities, where only a few
dominating species are present, the classical “genome-centric” approach is often
preferable because this provides more detailed information. The assignment of the
assembled contigs to organism bins combined with subsequent annotation usually enables the reconstruction of individual metabolic properties and leads to
