54
A. Meyerdierks and F.O. Glöckner
consequently accumulating in complex environmental samples, could be an initial
round of conservative assembly combined with an unsupervised pre-binning of the
resulting contigs and scaffolds. These “seed” bins can then be used as templates for
further sequence correlation and clustering of the single reads or merged mate pairs,
or to train systems like Phylophytia for classification. However, this will only work
if sequencing depth is sufficient to produce the initial contigs, or if sequences from
large insert constructs are available.
2.4.2 Gene Prediction
Classical gene predictors like ZCurve (Guo et al. 2003), Glimmer (Delcher et al.
2007) or a combination of several gene finders (Glöckner et al. 2003) work reasonably well for high quality contigs from assembled shotgun sequences or large insert
constructs. If only a small number of contigs needs to be analysed a slight overpredition of 10–20%, typical of intrinsic gene predictors like ZCurve or Glimmer,
can be accepted and these can later be eliminated by manual curation. Gene predictors using extrinsic strategies such as similarity based searches against existing
genomic and metagenomic databases to delineate coding regions (Badger and Olsen
1999, Krause et al. 2006) require more time and tend to underpredict genes due to a
lack of homologous information in the databases. In the case of large metagenomic
datasets, processing time and potential loss of information can be a severe limitation.
Furthermore, metagenomic gene predictors have to cope with (1) fragmented genes,
(2) low sequence quality leading to frameshifts and (3) high phylogenetic diversity,
which limits the initial training step of intrinsic gene finders. MetaGene, a recently
developed genefinder for the analysis of metagenomic fragments, is able to cope
with most of these limitations. MetaGene utilises di-codon frequencies estimated
by the GC content of a given sequence. This estimation, in combination with measures of the length distribution of ORFs, the distance from the leftmost start codons
and the orientation and distance of neighbouring ORFs, extracted by statistical analysis of around 130 bacterial and archaeal genomes, are used for gene prediction
(Noguchi et al. 2006). The system is quite fast and, based on our experiences,
currently provides the best results when tested on a broad range of metagenomic
assemblies from prokaryotes.
2.4.3 Functional Annotation
Functional annotation can be regarded as the most important step in the process
of analysing genomic fragments obtained from metagenomic studies. It is at this
stage of the process that the investigator obtains a substantial insight into the abundance of the genetic potential available in the environmental sample being studied.
Annotation should be carried out with care since poor annotations will – like the
proverbial first ice crystal – start a snowball effect by continuous error propagation
Précédent

- 69/410

Suivant