338
V. Mittard-Runte et al.
with large sequence repositories feasible. In 1990, Altschul et al. published the first
release of “BLAST”, the basic local alignment search tool. BLAST implements a
fast algorithm to find similar sequences in repositories based on a local alignment
(Altschul et al. 1990).
Although other methods for searching for similar sequences in large repositories
exist, FASTA and BLAST are still the most well known examples for sequence
similarity based tools, and almost every large genome sequencing and annotation
effort is based on them.
9.3.3.2 From Gene Annotation to Genome Annotation
Using the various bioinformatics tools available, the analysis of a novel gene and the
prediction of its function based on similarity are quite easy. Many institutes provide
services for analysing single genes with BLAST or other tools, with easy to use web
interfaces and nice graphical post processing of results.
Things change dramatically if a gene set of a complete genome or a collection of
ESTs has to be analysed. Running all the necessary prediction tools by hand is not
feasible if several thousand sequences need to be processed. Data management and
automation are the keys to successfully annotating a complete genome or a large
set of EST sequences. Several integrated systems exist that aid biologists to carry
out the annotation process, manage data, handle the various tools and present the
results in a well-arranged manner. Based on the stored information most of the systems also offer higher-level functionality like metabolic reconstruction, integration
of experimental data and comparisons with other genomes.
An easy to use example for a genome annotation system is “Artemis”, developed
at the Sanger Centre and published in 2000 (Rutherford et al. 2000). It offers visualisation and annotation of local sequence data, provided in simple sequence files,
tabular files or structured files that already contain information about genes and their
annotation. Installation is simply a matter of downloading and unpacking a file. It
also allows the integration of various function prediction tools and adaptation to
the local computer systems. Due to the focus on local files Artemis does not allow
multiple users to work on the same genome at the same time.
The multiple user problem is addressed by annotation systems like GenDB
(Meyer et al. 2003) and SAMS (see also Section 9.3.1.b). Sequences and information about genes, tool results and annotations are stored in a central relational
database. Access is provided by a web interface, making a locally installed software unnecessary. Support from large computer clusters and pipeline processing
make GenDB and SAMS powerful tools, and a well defined and open programming interface allows developers to add functionality with ease. Both systems
support complete automated annotation of sequence data and may also guide manual
annotation by a group of users.
Most large sequencing and bioinformatics centres have developed their own processing pipelines and are offering other institutes access to their systems. Much of
this development has been carried out in the context of large eukaryotic genome
projects. The Ensembl system is a good example of this kind of system (Flicek
V. Mittard-Runte et al.
with large sequence repositories feasible. In 1990, Altschul et al. published the first
release of “BLAST”, the basic local alignment search tool. BLAST implements a
fast algorithm to find similar sequences in repositories based on a local alignment
(Altschul et al. 1990).
Although other methods for searching for similar sequences in large repositories
exist, FASTA and BLAST are still the most well known examples for sequence
similarity based tools, and almost every large genome sequencing and annotation
effort is based on them.
9.3.3.2 From Gene Annotation to Genome Annotation
Using the various bioinformatics tools available, the analysis of a novel gene and the
prediction of its function based on similarity are quite easy. Many institutes provide
services for analysing single genes with BLAST or other tools, with easy to use web
interfaces and nice graphical post processing of results.
Things change dramatically if a gene set of a complete genome or a collection of
ESTs has to be analysed. Running all the necessary prediction tools by hand is not
feasible if several thousand sequences need to be processed. Data management and
automation are the keys to successfully annotating a complete genome or a large
set of EST sequences. Several integrated systems exist that aid biologists to carry
out the annotation process, manage data, handle the various tools and present the
results in a well-arranged manner. Based on the stored information most of the systems also offer higher-level functionality like metabolic reconstruction, integration
of experimental data and comparisons with other genomes.
An easy to use example for a genome annotation system is “Artemis”, developed
at the Sanger Centre and published in 2000 (Rutherford et al. 2000). It offers visualisation and annotation of local sequence data, provided in simple sequence files,
tabular files or structured files that already contain information about genes and their
annotation. Installation is simply a matter of downloading and unpacking a file. It
also allows the integration of various function prediction tools and adaptation to
the local computer systems. Due to the focus on local files Artemis does not allow
multiple users to work on the same genome at the same time.
The multiple user problem is addressed by annotation systems like GenDB
(Meyer et al. 2003) and SAMS (see also Section 9.3.1.b). Sequences and information about genes, tool results and annotations are stored in a central relational
database. Access is provided by a web interface, making a locally installed software unnecessary. Support from large computer clusters and pipeline processing
make GenDB and SAMS powerful tools, and a well defined and open programming interface allows developers to add functionality with ease. Both systems
support complete automated annotation of sequence data and may also guide manual
annotation by a group of users.
Most large sequencing and bioinformatics centres have developed their own processing pipelines and are offering other institutes access to their systems. Much of
this development has been carried out in the context of large eukaryotic genome
projects. The Ensembl system is a good example of this kind of system (Flicek
