46
Taxonomic binning is a core problem in metagenomics in which DNA reads are
classified into taxonomic groups termed bins. To date, several binning algorithms
are capable of classifying sequences into pre-defined bins based on either sequence
or DNA signature similarity. Sequence similarity-based algorithms attempt to align
the metagenomic data to a database of fully or partially sequenced genomes (e.g.,
the MEGAN software). The approach may provide accurate results when the binned
sequence has close relative(s) in the genome database, but its accuracy and sensitivity drop significantly drop when this is not the case. DNA signature-based approaches
use oligonucleotide frequencies to generate profile vectors for all sequences in the
metagenome. Once generated, these vectors are compared to a set of predefined
profile vectors representing known genomes and assigned to the closest genome
(e.g., the PhyloPythia algorithm) (McHardy et al. 2007). These algorithms tend to
be more accurate but may still fail for short read lengths. Most existing algorithms
are, in fact, classification algorithms that require a “training set” consisting of
sequenced genomes. These algorithms are likely to fail for sequences with no close
relatives in the training set.
With the advent of metagenomics and computational biology, scientists all over
the world have generated huge amounts of genomic data in the form of reads and
contigs. Under the situation a variety of the bioinformatics tools has been developed
with time to process, analyze, interpret, and annotate these huge amounts of metagenomics data. Furthermore, the precision and accuracy of the bioinformatic tools
also created a big question mark; however, it has been managed with newly developed technologies. Under the topic of annotation, RAST (rapid annotation by subsystem technology) is a most commonly used server (more than 20,000 global
users) for metagenomic data annotation and annotated approximately 70,000–
80,000 (Overbeek et al. 2013). More detail about the server is well explained by
Aziz et al. (2008). In the process of annotation, generally feature- and functionbased analysis of coding regions has been carried out with the help of the various
bioinformatics tools such as the gene finder, FragGeneScan (Thomas et al. 2012).
In Fig. 4.4, functional abundance in a particular metagenomic dataset is illustrated. Pathway reconstruction is a related problem in which common cellular or
physiological processes in a genome or a metagenome are determined without estimating their abundance (Filippo et al. 2012). Figure 4.5 provides an example of
pathway mapping through the KeggMapper tool of the MG-RAST server. Functional
analysis at the pathway level is mainly used for two purposes: computation of pathway relative abundance, and pathway content comparison. Computing relative
abundance of pathways within a single sample provides an overall view of the
environment and was used in many studies and platforms. Comparing pathways’
abundance between samples makes it possible to identify pathways that are enriched
within one of the environments with respect to the other. Derivatives of pathway
content comparison may be used for clustering functionally similar environments
using metrics over pathway abundance vectors (Filippo et al. 2012). Some limitations, such as less-developed bioinformatics tools and assembly of a thousand
sequence reads of the metagenomics analysis, still exist so that complex microbial
communities may not appear in any assembly. High-throughput sequencing tech4 Single-Cell Genomics and Metagenomics for Microbial Diversity Analysis
Taxonomic binning is a core problem in metagenomics in which DNA reads are
classified into taxonomic groups termed bins. To date, several binning algorithms
are capable of classifying sequences into pre-defined bins based on either sequence
or DNA signature similarity. Sequence similarity-based algorithms attempt to align
the metagenomic data to a database of fully or partially sequenced genomes (e.g.,
the MEGAN software). The approach may provide accurate results when the binned
sequence has close relative(s) in the genome database, but its accuracy and sensitivity drop significantly drop when this is not the case. DNA signature-based approaches
use oligonucleotide frequencies to generate profile vectors for all sequences in the
metagenome. Once generated, these vectors are compared to a set of predefined
profile vectors representing known genomes and assigned to the closest genome
(e.g., the PhyloPythia algorithm) (McHardy et al. 2007). These algorithms tend to
be more accurate but may still fail for short read lengths. Most existing algorithms
are, in fact, classification algorithms that require a “training set” consisting of
sequenced genomes. These algorithms are likely to fail for sequences with no close
relatives in the training set.
With the advent of metagenomics and computational biology, scientists all over
the world have generated huge amounts of genomic data in the form of reads and
contigs. Under the situation a variety of the bioinformatics tools has been developed
with time to process, analyze, interpret, and annotate these huge amounts of metagenomics data. Furthermore, the precision and accuracy of the bioinformatic tools
also created a big question mark; however, it has been managed with newly developed technologies. Under the topic of annotation, RAST (rapid annotation by subsystem technology) is a most commonly used server (more than 20,000 global
users) for metagenomic data annotation and annotated approximately 70,000–
80,000 (Overbeek et al. 2013). More detail about the server is well explained by
Aziz et al. (2008). In the process of annotation, generally feature- and functionbased analysis of coding regions has been carried out with the help of the various
bioinformatics tools such as the gene finder, FragGeneScan (Thomas et al. 2012).
In Fig. 4.4, functional abundance in a particular metagenomic dataset is illustrated. Pathway reconstruction is a related problem in which common cellular or
physiological processes in a genome or a metagenome are determined without estimating their abundance (Filippo et al. 2012). Figure 4.5 provides an example of
pathway mapping through the KeggMapper tool of the MG-RAST server. Functional
analysis at the pathway level is mainly used for two purposes: computation of pathway relative abundance, and pathway content comparison. Computing relative
abundance of pathways within a single sample provides an overall view of the
environment and was used in many studies and platforms. Comparing pathways’
abundance between samples makes it possible to identify pathways that are enriched
within one of the environments with respect to the other. Derivatives of pathway
content comparison may be used for clustering functionally similar environments
using metrics over pathway abundance vectors (Filippo et al. 2012). Some limitations, such as less-developed bioinformatics tools and assembly of a thousand
sequence reads of the metagenomics analysis, still exist so that complex microbial
communities may not appear in any assembly. High-throughput sequencing tech4 Single-Cell Genomics and Metagenomics for Microbial Diversity Analysis
