336
V. Mittard-Runte et al.
Several eukaryotic gene finding algorithms are able to exploit different types
of evidence. Programs like N-SCAN, N-SCAN_EST (Wei and Brent 2006),
TWINSCAN, AUGUSTUS, and GENIE (Reese et al. 2000) employ HMMs to
incorporate intrinsic and extrinsic information into a single probabilistic model. The
eukaryotic gene finder EUGENE (Schiex et al. 2001), which has been used for the
annotation of Arabidopsis thaliana, employs a comparable approach. In EUGENE,
intrinsic and extrinsic information is combined in a weighted directed acyclic
graph, the most likely gene structure of the analysed DNA sequence is determined
by the shortest path in that graph. Other programs, such as JIGSAW (Allen and
Salzberg 2005), EXONHUNTER (Brejova et al. 2005), GLEAN (Elsik et al. 2007),
EXOGEAN (Djebali et al. 2006), EVIDENCEMODELER (Haas et al. 2008), and
AGUSTUS-any (Stanke et al. 2006), mimic the human gene identification process
by evaluating evidence from diverse sources.
The accuracy of eukaryotic gene finding methods in general increases as the
number and size of introns decreases. Accuracy levels of up to 70% are achieved
for the de novo prediction of complete gene structures in compact genomes using
comparative methods. For genomes of mammals with high proportions of intergenic
regions, complex gene structures and alternative splice sites, usually a considerably
lower accuracy is obtained.
In 2005 the EGASP project was launched to evaluate the accuracy of existing gene finding methods on the ENCODE regions of the human genome (Guigo
and Reese 2005, Guigo et al. 2006). The best programs correctly predicted at
least one transcript for almost 70% of the annotated genes. However, when alternative splicing variants were taken into account, the accuracy dropped to values
between approximately 40 and 50%. At the coding exon level, the best evaluated
methods achieved a sensitivity and specificity of more than 80%, and close to
90% at the coding nucleotide level. Programs relying on diverse information, in
particular sequences of gene transcripts or proteins, were the most accurate, followed by methods relying on sequence comparisons across two or more genomes.
Algorithms evaluating only intrinsic sequence features achieved the lowest accuracy. The performance for predicting non-coding exons was low for almost all
evaluated programs.
To conclude, the best contemporary eukaryotic gene finding methods
achieve good accuracy levels in general for the identification of exons,
but the prediction of complete gene structures in complex genomes is still
highly challenging. To obtain predictions of the highest possible quality it is
recommended to combine programs that detect repetitive regions and pseudogenes with methods that rely on a mapping of gene transcripts and dual,
multi-genome, and intrinsic gene finders. Several popular gene finding algorithms can easily be run via a public web-interface, including: AUGUSTUS
(http://augustus.gobics.de/), GENEMARK (http://exon.gatech.edu/GeneMark/),
TWINSCAN/NSCAN (http://mblab.wustl.edu/software/twinscan/), GLIMMERM
(http://www.tigr.org/tdb/glimmerm/glmr_form.html) and GLIMMERHMM (http://
nbc9.biologie.uni-kl.de/framed/left/menu/auto/right/glimmerhmm/).
Précédent

- 347/410

Suivant