41
4.2.2 Metagenomics Data Analysis: Approaches
and Challenges
The success of metagenomics is completely dependent on the high-throughput techniques for the processing of DNA from different environments and their sequence
analysis after running on high-end sequencers. Furthermore, analyzing millions or
trillions of the reads and their assembly to achieve a complete genome is a really
challenging task (Aguiar-Pulido et al. 2016). Metagenomic projects face several
challenges during the gathered data analysis, such as assembly, phylogenetic analysis, taxonomic binning, gene calling, and community-level analysis. Metagenomics
invites many interesting computational problems such as gene prediction, sequence
classification, and clustering, genome assembly, statistical comparison, functional
annotation, microbial interactions modelling, etc. Data coverage is again a major
challenge in the metagenomic study because of the poor availability of computational methods to estimate the magnitude of coverage (Rodriguezr and Konstantinidis
2014). Some of the major computational challenges include the assembly of the
whole data, phylogenetic surveys, gene finding, and comparative metagenomic
analysis for the metabolic pathways (Wooley and Ye 2009; Sharpton 2014; Mendoza
et al. 2015; Filippo et al. 2012). In general, there are two approaches of computational metagenomics data processing and interpretation: (i) statistical analysis of
functional metagenomic data and (ii) categorization of genes and from millions of
the metagenomic reads (Sharon 2010). Metagenomic sequence data can inform us
not only about the structural microbial communities in the soil or other habitats but
can also assign functions to the microbes inhabiting the different habitats (Thomas
et al. 2012; Sharpton 2014; Zhou et al. 2015).
4.2.2.1 Assembly
Assembly is mostly relevant to Sanger sequencing data wherein read length usually
approaches 1000 bps. It is usually carried out using a single genome assembler. The
use of comparative analysis with respect to template genomes may increase confidence in the final result. The transformation from the assembly of a single genome
to the assembly of many genomes at once is not trivial and raises several issues. The
presence of conserved regions in several different organisms is likely to harm the
assembly because of the assembler’s inability to differentiate closely related
stretches. Such stretches are usually considered as being repetitive and are ignored.
The same problem is caused by the presence of several copies of a region that
belongs to an abundant genome, in particular when the region is highly polymorphic. This problem may be treated manually (Kunin et al. 2008; Miller et al. 2010;
Teeling and Glöckner 2012). There exist certain tools for estimating the amount of
gaps in single genome assembly for a certain amount of coverage. Constructing
such models for multiple genomes is much more difficult and requires estimation of
the diversity of organisms in the tested environment (Hooper et al. 2010; Ekblom
4.2 Metagenomics
4.2.2 Metagenomics Data Analysis: Approaches
and Challenges
The success of metagenomics is completely dependent on the high-throughput techniques for the processing of DNA from different environments and their sequence
analysis after running on high-end sequencers. Furthermore, analyzing millions or
trillions of the reads and their assembly to achieve a complete genome is a really
challenging task (Aguiar-Pulido et al. 2016). Metagenomic projects face several
challenges during the gathered data analysis, such as assembly, phylogenetic analysis, taxonomic binning, gene calling, and community-level analysis. Metagenomics
invites many interesting computational problems such as gene prediction, sequence
classification, and clustering, genome assembly, statistical comparison, functional
annotation, microbial interactions modelling, etc. Data coverage is again a major
challenge in the metagenomic study because of the poor availability of computational methods to estimate the magnitude of coverage (Rodriguezr and Konstantinidis
2014). Some of the major computational challenges include the assembly of the
whole data, phylogenetic surveys, gene finding, and comparative metagenomic
analysis for the metabolic pathways (Wooley and Ye 2009; Sharpton 2014; Mendoza
et al. 2015; Filippo et al. 2012). In general, there are two approaches of computational metagenomics data processing and interpretation: (i) statistical analysis of
functional metagenomic data and (ii) categorization of genes and from millions of
the metagenomic reads (Sharon 2010). Metagenomic sequence data can inform us
not only about the structural microbial communities in the soil or other habitats but
can also assign functions to the microbes inhabiting the different habitats (Thomas
et al. 2012; Sharpton 2014; Zhou et al. 2015).
4.2.2.1 Assembly
Assembly is mostly relevant to Sanger sequencing data wherein read length usually
approaches 1000 bps. It is usually carried out using a single genome assembler. The
use of comparative analysis with respect to template genomes may increase confidence in the final result. The transformation from the assembly of a single genome
to the assembly of many genomes at once is not trivial and raises several issues. The
presence of conserved regions in several different organisms is likely to harm the
assembly because of the assembler’s inability to differentiate closely related
stretches. Such stretches are usually considered as being repetitive and are ignored.
The same problem is caused by the presence of several copies of a region that
belongs to an abundant genome, in particular when the region is highly polymorphic. This problem may be treated manually (Kunin et al. 2008; Miller et al. 2010;
Teeling and Glöckner 2012). There exist certain tools for estimating the amount of
gaps in single genome assembly for a certain amount of coverage. Constructing
such models for multiple genomes is much more difficult and requires estimation of
the diversity of organisms in the tested environment (Hooper et al. 2010; Ekblom
4.2 Metagenomics
