9 Genomic Techniques and How to Apply Them to Marine Questions
331
Within Marine Genomics Europe (MGE), SAMS was successfully used to analyse 44 EST projects, including ESTs from the following species: Fucus serratus and
Fucus vesiculosus, Dicentrarchus labrax (Sea bass), Emiliania huxleyi, or Balanus
amphitrite.
9.3.2 Gene Prediction
The prediction of tRNA, rRNA, and protein encoding genes from raw genomic
sequences is one of the first essential steps during the annotation of newly sequenced
genomes. This section focuses on the computational strategies and existing software for the automated identification of protein encoding genes (coding sequences,
CDSs) in genome sequences. For gene prediction, two different approaches are
applied: intrinsic and similarity based methods. Intrinsic methods analyse sequence
properties of genomes to discriminate between coding and non-coding regions.
These methods exploit the different compositional properties of coding and noncoding sequences, mainly caused by a bias in codon usage in CDSs, which optimizes
the translation efficiency in protein biosynthesis (Gouy and Gautier 1982).
Intrinsic gene finding methods frequently employ a statistical model representing the frequencies of short oligonucleotides 1 in coding and non-coding sequences.
Generative models with Markov properties (Durbin et al. 1998), such as fixed-order
Markov chains on nucleotides (Delcher et al. 1999, Larsen and Krogh 2003) or
codons (Badger and Olsen 1999), are often applied to represent the sequence composition. Additional features, such as ribosome binding sites or overlaps between
adjacent genes, can also be integrated into a probabilistic framework, if hidden
Markov models (HMMs) are used to describe the context of a gene (Delcher et al.
1999).
Similarity based methods predict genes by searching for stretches of DNA sharing a significant similarity to reference sequences, including genomes from close
relatives, phylogenetically related proteins, or gene transcripts. Several similarity
based approaches discriminate conserved coding sequences from conserved noncoding regions based on their synonymous substitution rate (Badger and Olsen
1999, Moore and Lake 2003, Nekrutenko et al. 2003). This is based on the fact
that, to maintain the amino acid sequence of an encoded protein, coding sequences
show a much higher number of synonymous mutations, i.e. mutations that do not
modify the encoded amino acid, than conserved non-coding regions.
9.3.2.1 Gene Finding in Prokaryotes
In prokaryotes (Bacteria and Archaea), protein encoding genes are open reading
frames (ORFs), which are sequences of codons beginning with a start codon, ending
with a stop codon, and without an internal stop codon. The task of predicting protein
1 Usually frequencies of oligonucleotides of length between 3 and 12 bp are modelled.
331
Within Marine Genomics Europe (MGE), SAMS was successfully used to analyse 44 EST projects, including ESTs from the following species: Fucus serratus and
Fucus vesiculosus, Dicentrarchus labrax (Sea bass), Emiliania huxleyi, or Balanus
amphitrite.
9.3.2 Gene Prediction
The prediction of tRNA, rRNA, and protein encoding genes from raw genomic
sequences is one of the first essential steps during the annotation of newly sequenced
genomes. This section focuses on the computational strategies and existing software for the automated identification of protein encoding genes (coding sequences,
CDSs) in genome sequences. For gene prediction, two different approaches are
applied: intrinsic and similarity based methods. Intrinsic methods analyse sequence
properties of genomes to discriminate between coding and non-coding regions.
These methods exploit the different compositional properties of coding and noncoding sequences, mainly caused by a bias in codon usage in CDSs, which optimizes
the translation efficiency in protein biosynthesis (Gouy and Gautier 1982).
Intrinsic gene finding methods frequently employ a statistical model representing the frequencies of short oligonucleotides 1 in coding and non-coding sequences.
Generative models with Markov properties (Durbin et al. 1998), such as fixed-order
Markov chains on nucleotides (Delcher et al. 1999, Larsen and Krogh 2003) or
codons (Badger and Olsen 1999), are often applied to represent the sequence composition. Additional features, such as ribosome binding sites or overlaps between
adjacent genes, can also be integrated into a probabilistic framework, if hidden
Markov models (HMMs) are used to describe the context of a gene (Delcher et al.
1999).
Similarity based methods predict genes by searching for stretches of DNA sharing a significant similarity to reference sequences, including genomes from close
relatives, phylogenetically related proteins, or gene transcripts. Several similarity
based approaches discriminate conserved coding sequences from conserved noncoding regions based on their synonymous substitution rate (Badger and Olsen
1999, Moore and Lake 2003, Nekrutenko et al. 2003). This is based on the fact
that, to maintain the amino acid sequence of an encoded protein, coding sequences
show a much higher number of synonymous mutations, i.e. mutations that do not
modify the encoded amino acid, than conserved non-coding regions.
9.3.2.1 Gene Finding in Prokaryotes
In prokaryotes (Bacteria and Archaea), protein encoding genes are open reading
frames (ORFs), which are sequences of codons beginning with a start codon, ending
with a stop codon, and without an internal stop codon. The task of predicting protein
1 Usually frequencies of oligonucleotides of length between 3 and 12 bp are modelled.
