44
4.2.2.5 Community-Level Analysis
Community-level analysis is another structural and functional issue of interest in
answering questions related to microbial communities as a whole, rather than specific species of functional systems (Zarraonaindia et al. 2013). Community makeup
is usually studied by analyzing phylogenetic marker genes, most notably the 16S
SSU and 23S LSU rRNA coding genes. Species are identified on the basis of
sequence alignment against databases such as RDP, and species richness is estimated based on statistical tools such as rarefaction curves (Sharon 2010; Guo et al.
2016). The functional attributes of the microbes are estimated based on the gene
contents of the metagenome and are built by estimating the relative abundance of
each gene or pathway (Kunin et al. 2008; Carr and Borenstein 2014). Identification
of genes and pathways is usually done by aligning all sequences in the metagenome
to databases such as the NCBI Clusters of Orthologs (COG) or the KyotoEncyclopedia developed for the Genes and Genomes (Randle-Boggis et al. 2016).
Once the number of reads carrying each gene or pathway is determined, it is possible to use these numbers for estimating the comparative quantitative analysis of
each function (Filippo et al. 2012).
4.2.3 Metagenomics Data Analysis: Available Tools
and Techniques
A number of databases have enormous capability to store and analyze the metagenomic data. MG-RAST, IMG/M, and CAMERA represent three well-known
metagenomics servers. MG-RAST (Keegan et al. 2016) provides a data repository
along with a full-fledged platform for analysis with multiple comparisons. The EBI
Metagenomics service has facilitated the users with an automated pipeline for the
analysis and storage of metagenomic data and allows detailed analysis of the taxonomic relationships along with the functional and metabolic potential of any sample
(Hunter et al. 2014b). Genomes-OnLine-Database (GOLD) renders knowledge
concerning finalized and in progress microbial mega-projects across the globe
(Mukherjee et al. 2017). Both IMG/M (Chen et al. 2017) and MG-RAST are wellorganized platforms that allow users to insert their metagenomic data, compare the
data with available databases, metagenomes, without requiring the end-user-based
raw data. CAMERA (Seshadri et al. 2007) presents more flexible annotation schema
through the user’s needs to understand the data annotation and analytical pipelines
sufficient for their interpretation. MEGAN is one more tool applied for analyzing
annotated data obtained from BLAST analysis under functional or phylogenic study
(Huson et al. 2007).
Many reference databases such as KEGG (Kanehisa and Goto 2000), eggNOG
(Muller et al. 2010), COG (Tatusov et al. 2003), PFAM (Finn et al. 2014), and
TIGRFAM (Haft et al. 2003) provide overall functional physiology of the metagenome. However, none of the reference databases covers entire functions of a par4 Single-Cell Genomics and Metagenomics for Microbial Diversity Analysis
4.2.2.5 Community-Level Analysis
Community-level analysis is another structural and functional issue of interest in
answering questions related to microbial communities as a whole, rather than specific species of functional systems (Zarraonaindia et al. 2013). Community makeup
is usually studied by analyzing phylogenetic marker genes, most notably the 16S
SSU and 23S LSU rRNA coding genes. Species are identified on the basis of
sequence alignment against databases such as RDP, and species richness is estimated based on statistical tools such as rarefaction curves (Sharon 2010; Guo et al.
2016). The functional attributes of the microbes are estimated based on the gene
contents of the metagenome and are built by estimating the relative abundance of
each gene or pathway (Kunin et al. 2008; Carr and Borenstein 2014). Identification
of genes and pathways is usually done by aligning all sequences in the metagenome
to databases such as the NCBI Clusters of Orthologs (COG) or the KyotoEncyclopedia developed for the Genes and Genomes (Randle-Boggis et al. 2016).
Once the number of reads carrying each gene or pathway is determined, it is possible to use these numbers for estimating the comparative quantitative analysis of
each function (Filippo et al. 2012).
4.2.3 Metagenomics Data Analysis: Available Tools
and Techniques
A number of databases have enormous capability to store and analyze the metagenomic data. MG-RAST, IMG/M, and CAMERA represent three well-known
metagenomics servers. MG-RAST (Keegan et al. 2016) provides a data repository
along with a full-fledged platform for analysis with multiple comparisons. The EBI
Metagenomics service has facilitated the users with an automated pipeline for the
analysis and storage of metagenomic data and allows detailed analysis of the taxonomic relationships along with the functional and metabolic potential of any sample
(Hunter et al. 2014b). Genomes-OnLine-Database (GOLD) renders knowledge
concerning finalized and in progress microbial mega-projects across the globe
(Mukherjee et al. 2017). Both IMG/M (Chen et al. 2017) and MG-RAST are wellorganized platforms that allow users to insert their metagenomic data, compare the
data with available databases, metagenomes, without requiring the end-user-based
raw data. CAMERA (Seshadri et al. 2007) presents more flexible annotation schema
through the user’s needs to understand the data annotation and analytical pipelines
sufficient for their interpretation. MEGAN is one more tool applied for analyzing
annotated data obtained from BLAST analysis under functional or phylogenic study
(Huson et al. 2007).
Many reference databases such as KEGG (Kanehisa and Goto 2000), eggNOG
(Muller et al. 2010), COG (Tatusov et al. 2003), PFAM (Finn et al. 2014), and
TIGRFAM (Haft et al. 2003) provide overall functional physiology of the metagenome. However, none of the reference databases covers entire functions of a par4 Single-Cell Genomics and Metagenomics for Microbial Diversity Analysis
