9 Genomic Techniques and How to Apply Them to Marine Questions
337
9.3.3 Genome Annotation and Beyond
With either gene prediction or EST assembly being finished, the next step in working
with genomic sequences is the prediction and analysis of the functions and roles a
gene or assembled EST might have. Almost all in-silico methods are based on the
correlation between gene sequences, their corresponding amino acid sequence, its
two- and three-dimensional structure and the function the protein may carry out.
Since a large number of sequences from different sources are already known and
their function has been identified and described in scientific publications, searching
for similar sequences has proven to be a good method for analysing novel genes and
genomes.
9.3.3.1 Introduction to Sequence Similarity
Comparing sequences requires an exact measure for the level of similarity and the
degree of difference between them; without this measure a computer is unable to
automate the processing. In Bioinformatics “alignments” are used to describe the
similarity of sequences. They represent the necessary steps to convert one sequence
into another sequence. Each single step may either be (i) match or mismatch of an
amino acid or nucleotide base, (ii) insertion of an amino acid or nucleotide base,
(iii) deletion of an amino acid or nucleotide base.
Combined with a metric that scores the steps in an alignment, e.g. by putting a
penalty on insertions and deletions and rewarding preservation, the best or “optimal”
alignment may be deduced. In the simplest implementation all possible alignments
are generated and scored to find the optimal one; unfortunately the number of alignments grows exponentially with the length of the sequences, tripling the number of
alignments for each nucleotide base or amino acid added to the process.
Thus a software implementing algorithms for sequence similarity has to filter out
most alignments and reduce them to the best one, depending on how the quality of
alignments is estimated. This may involve the use of simple scores for the different
steps up to highly sophisticated models that include mutation events, recombination
and other events that change sequences. In 1970 Needleman and Wunsch (1970)
described a dynamic programming algorithm whose processing time grows in relation to the product of the length of the two sequences, thus making the alignment of
two large sequences feasible. These “global” alignments work well if the sequences
are well conserved. In 1981 Smith and Waterman presented a similar approach for
“local” alignments that works well if parts of the sequences are not conserved (Smith
and Waterman 1981).
With the amount of published sequences growing, efforts were made to collect them in central repositories and make them available to the community. While
the first release of “GenBank” in 1982 contained about 600 entries, the repository
size has grown in an exponential manner, and in 2008 it now contains 80,000,000
DNA sequences. The first widely used program to search the repositories for similar sequences was “FASTA”, released by Pearson and Lipman (1988). This program
uses special optimizations of the Needleman-Wunsch algorithm that made working
Précédent

- 348/410

Suivant