332
V. Mittard-Runte et al.
encoding genes in prokaryotic genomes can therefore be regarded as a two class
classification problem: Discriminating coding sequences from the majority of noncoding ORFs, which correspond to genomic regions that are not transcribed and
translated in vivo.
For prokaryotes, many different gene finding software tools have become available that use either intrinsic or a combination of intrinsic and similarity based
methods to automatically discover genes from raw genomic sequences. While these
methods in general achieve impressive accuracy values, the correct identification
of short genes, genes with atypical sequences properties and translation start sites
are still challenging tasks. Short genes are more difficult to identify because their
sequences carry less information that can be evaluated for identification (Skovgaard
et al. 2001, Larsen and Krogh 2003, Ou et al. 2004). Moreover, genes from the same
organism may exhibit divergent sequence properties, owing to expression dependent
codon usage (McHardy et al. 2004b), leading/lagging strand-related biases (Lafay
et al. 1999), or the transfer of genes between different bacterial species, called horizontal gene transfer (Smith et al. 1992). In such cases, intrinsic methods may have
difficulties in identifying the different gene classes. This issue has been addressed
by the inclusion of an additional model for genes with “atypical” sequence composition or by the unsupervised discovery of CDS classes prior to the prediction phase
(Lukashin and Borodovsky 1998, Krause et al. 2007).
Among the most widely used intrinsic gene finding programs are Glimmer
(Delcher et al. 1999, 2007) and Genemark (Lukashin and Borodovsky 1998).
Glimmer is fast, easy to install locally, and detects about 99% of “certain” genes
with known functions on average. However, Glimmer-2 has been reported to produce a rather high number of false positive predictions (McHardy et al. 2004a,
Krause et al. 2007). In the latest version (Glimmer-3) this has been addressed by
selecting a set of highest-scoring predictions consistent with the maximal allowed
overlap during a post-processing phase. Genemark, which uses a hidden Markov
model (Besemer et al. 2001, Besemer and Borodovsky 2005) also achieves a high
prediction accuracy, with a sensitivity of up to 99% and a specificity 2 of about 93%
(Delcher et al. 2007). The program can easily be executed via a web-interface.
Similarity based methods in general are much slower and more difficult to install
than intrinsic approaches but more reliable predictions are obtained. Moreover, in
a combined approach, similarity supported predictions provide a reliable initial
training set to train a genome-specific intrinsic model. For example, by combining a Pfam-based search for conserved protein family members with an intrinsic
approach, the gene finder GISMO is able to identify up to 99% of genes with known
functions (Krause et al. 2007). GISMO also produces highly reliable predictions
with a specificity of more than 94% and achieves high accuracy levels for the identification of short genes and for finding genes in GC-rich genomes. For intrinsic
2 Specificity measures the reliability of the predictions. It is defined as the fraction of correct gene
predictions, i.e. the fraction of predicted genes that corresponds to known genes.
Précédent

- 343/410

Suivant