43
nucleotide frequencies techniques, the latter being more successful (Thomas et al.
2012; Dröge and McHardy 2012; Sedlar et al. 2017). The TETRA sequence analysis software and its underlying similarity measure between sequences use a representation of tetranucleotide frequencies for a sequence, and the Pearson correlation
coefficient as a similarity measure between sequences (Teeling et  al. 2004). The
PhyloPythia classification algorithm employs a support vector machine (SVM) for
classifying sequences to known taxa, based on their nucleotide frequencies (Higashi
et al. 2012).
4.2.2.4 Gene Calling
Gene calling is another challenge to be deciphered. Several good solutions have
been proposed in reference to the issue of gene finding in fully sequenced
genomes. Algorithms based on the Hidden Markov-Model (HMM) and
Interpolated-Markov- Model (IMM) use sequence statistics for distinguishing
between exonic and intronic regions. These algorithms are trained on sequences
from the target genome; an initial set of genes is obtained by comparative methods. Approaches including search for translating gene segments (ORFs) by considering start and stop codons or aligning the genome against databases of known
genes or proteins are well recognized (Salzberg et al. 1999; Mathé et al. 2002;
Zhang 2002; Wang et al. 2004). All aforementioned approaches may encounter
difficulties that result from the fragmented nature of the sequences in the projects.
A large portion, sometimes up to 50%, of all the reads in metagenomic projects
cannot be assembled with the rest of the sequences being assembled into short
contigs of a few Kbps (Charuvaka and Rangwala 2011; Lin and Liao 2016).
Considering the length of reads (400–1000 bps for Sanger and 454 sequencing,
less than 100 for Solexa) and the fact that unicellular organism genes are usually
small, devoid of introns and located approximately every 1000  bps, probably
valuable information is present on most reads and contigs (Sharon 2010). Sequence
statistics-based algorithms may encounter difficulties in analyzing such fragmented data as they depend on whole-genome statistics. Comparative-based
methods are also expected to encounter difficulties from the many gene fractions
(Ekblom and Wolf 2014). Certain approaches involve a clustering algorithm in
which similar sequences in Global Ocean Sample (GOS) are clustered together on
the basis of sequence similarity. The primary input for the algorithm is the pairwise sequence similarity between all sequences in the database computed using
BLAST search. These similarity scores are used both for removing redundancies
and also for the construction of core sets, which contain highly conserved
sequences. Next, close core sets are unified based on profile–profile comparison.
Last, the profiles of the resulting sets are used for sequence recruitment using PSIBLAST.  Clusters containing sequences that are similar to annotated sequences
may be assigned predicted functionality; however, approximately 25% of the
clusters in GOS do not have any known homologue (Rusch et al. 2007; Yooseph
et al. 2007; Li et al. 2012).
4.2 Metagenomics
Précédent

- 59/118

Suivant