9 Genomic Techniques and How to Apply Them to Marine Questions
335
At present, the most reliable gene structure predictions are obtained by mapping
the sequence of gene transcripts onto the genome of interest. Many different programs have been devised for this task, including EST_GENOME (Mott 1997), AAT
(Huang et al. 1997), SIM4 (Florea et al. 1998), GENESEQUER (Usuka et al. 2000),
BLAT (Kent 2002), and GMAP (Wu and Watanabe 2005). The main disadvantage
of this approach is that genes that are expressed at a low level or those expressed
only in specific tissues or under specific conditions are likely to be missed.
Dual or multi-genome gene finding algorithms, such as SGP2 (Parra et al. 2003),
SLAM (Alexandersson et al. 2003), TWAIN (Majoros et al. 2005), N-SCAN (Gross
and Brent 2006), and TWINCAN (Korf et al. 2001) analyse the pattern of sequence
conservation between two or more genomes of evolutionary related organisms.
In cases where gene transcripts have not been sequenced, but genomes of close
relatives are available this class of programs can provide accurate exon predictions. However, dual and multi-genome gene finding methods may achieve only
a modest accuracy for the prediction of complete gene structures, miss genes without sequence homologies, and frequently mistake pseudogenes as functional. The
latter problem has recently been addressed by combining a comparative method
with a pseudogene detection using PPFINDER (van Baren and Brent 2006). If the
sequence of a closely related genome is not available, gene finding methods using
alignments with known proteins can be used to reliably identify exon locations
(Huang et al. 1997, Slater and Birney 2005). The popular GENEWISE (Birney et al.
2004) for example employs a Hidden Markov model that incorporates a model for
the alignment of protein sequences to a genome and a model of eukaryotic gene
structure.
To complement alignment based approaches, intrinsic methods allow the identification of genes based on the evaluation of compositional sequence properties.
The majority of these programs, such as GENSCAN (Burge and Karlin 1997),
GENEMARK.hmm (Lomsadze et al. 2005), GLIMMERHMM (Majoros et al.
2004), FGENESH (Salamov and Solovyev 2000) and AUGUSTUS (Stanke and
Waack 2003), employ HMMs or generalized HMMs (GHMMs) to partition genomic
sequences into introns, exons, and intergenic regions. By using HMMs or GHMMs
with states for introns, exons, intergenic regions, start and stop codons, splice sites,
and polyadenylation signals, diverse sequence features can be combined into a
coherent, probabilistic model. Before intrinsic methods can be applied to novel
genomes, they need to be trained so that they learn the genome specific compositional sequences properties, such as codon usage and splice site patterns. While most
intrinsic methods require supervised training on known genes, SNAP (Korf 2004),
and GENEMARK.hmm ES (Lomsadze et al. 2005) are able to discover sequence
properties of novel genomes in an unsupervised manner. For many of the already
sequenced eukaryotic genomes, pre-trained intrinsic models are publicly available.
In general, intrinsic gene finding programs perform fairly well when the identification of exons is considered, but they are still far from perfect when it comes to
reconstructing the exon-intron structures of complete genes. Furthermore, owing to
the high number of false positives produced, predictions that are not supported by
external evidence, such as sequence similarities, are unreliable.
Précédent

- 346/410

Suivant