42
and Wolf 2014; Sims et al. 2014). Assembly of next-generation sequencing (NGS)
data is rarely done because of the short read length. Current assemblers for NGS
data such as Velvet may handle de novo assembly of microbial genomes but are not
suitable for eukaryotic genomes as well as metagenomes. Strategies and assemblers
for NGS metagenomic data are, therefore, of great interest for the metagenomics
community (Miller et al. 2010; Henson et al. 2012).
4.2.2.2 Phylogenetic Analysis
Phylogenetic analysis from the metagenomic data is again a challenge because the
data are composed of short sequences from many organisms, raising several issues
related to phylogenetic analysis, with respect to both traditional and metagenomicspecific analysis (Sharpton 2014; Kembel et al. 2011; Zhou et al. 2015). Locating an
organism whose genome is known on the phylogenetic tree with respect to other
known organisms is traditionally carried out using phylogenetic informative markers, primarily rRNA-encoding genes. More than 1,400,000 microbial 16S small
subunit (SSU) rRNA coding genes can be found in such global databases as the
Ribosomal-Database-Project for phylogenetic classification of environmental data
(Rajendhran and Gunasekaran 2011; Cole et al. 2009; Sharon 2010). Other phylogenetic marker genes such as the recA and HSP70 are also used (Wu et al. 2014; Wu
and Scott 2012). However, most reads and scaffolds in metagenomic projects obviously do not contain phylogenetic marker genes, which raises the need for other
methods that may be based on oligonucleotide frequencies or other properties of the
genome. Methods that make no use of universal genes may be useful also for phylogenetic analysis of viruses, in which no universal genes are available (Pride and
Schoenfeld 2008; Iwasaki et al. 2013; Sharpton 2014; Darling et al. 2014). Even
when reads with phylogenetic marker genes are available, the task of reconstructing
a phylogenetic tree may not be trivial because only gene fragments are available. As
a first step in the process, a multiple sequence alignment (MSA) of all sequences
may be determined. Next, the MSA may be used for constructing a distance matrix
for obtaining a phylogenetic tree (White et  al. 2010). The fragmented nature of
metagenomic data usually results in partial sequences of the same region, which
makes it impossible to use current algorithms for MSA on such data. This problem
may require a new type of algorithm for aligning and scoring multiple sequences
(Sharpton 2014).
4.2.2.3 Taxonomic Binning
Taxonomic binning is another problem in metagenomics analysis. Sequence binning is a process for the separation of genome sequences into taxon-specific groups.
A binning step may be part of the assembly process of metagenomic data or may be
used for separating the genomes of a few members to study the biological processes.
The two main approaches with respect to this problem employ comparative and
4 Single-Cell Genomics and Metagenomics for Microbial Diversity Analysis
Précédent

- 58/118

Suivant