346
V. Mittard-Runte et al.
KEGG and KOBAS: The Kyoto Encyclopedia of Genes and Genomes (Kanehisa
and Goto 2000) combines databases that represent molecular interaction networks
in the context of biochemical pathways. These feature enzymes, compounds, and
connection points to related pathways. The recently published KOBAS (Mao et al.
2005) system uses the KEGG Ontology (KO) to produce automated annotations of
large sets of genes. KO is a controlled vocabulary similar to the Gene Ontology
(Ashburner et al. 2000) that offers the connection of genes to metabolic pathways present in the KEGG database. KOBAS is an automated annotation system
written in Python, which is able to identify and annotate complete pathways in
sets of sequences. KO terms are assigned to query genes if they show a high
sequence similarity to genes that are already annotated and connected to KEGG
maps.
STRING: The STRING (von Mering et al. 2005) database stores information
about genomic associations between genes. The most relevant associations between
two genes are conserved chromosomal proximity of genes in phylogenetically
distant organisms and gene fusion events. In addition, information about similar
expression patterns in microarray experiments and the co-occurrence of gene names
in the literature are included in the database. The combined information can be used
to predict unknown protein–protein interactions. The database is a pre-computed
global resource for the exploration of functional interactions between genes and is
available online.
COG: The COG database (Clusters of Orthologous Groups of proteins) (Tatusov
et al. 2003) was created by Tatusov and co-workers in 1997. It was designed to classify genes from completely sequenced organisms based on their common origin.
Initially a set of 21 complete genomes was used and an all-against-all sequence comparison was performed. Clusters of Orthologous Groups were created by applying
the criterion of genome specific best hits to a comparison of all coding sequences.
The COG categories were derived from the clusters. They act as a controlled hierarchical vocabulary that can be used to describe the function of proteins. For the
initial creation of the database, 2,091 COGs could be created that included up to
83% of the gene products of a single organism. The clusters have continuously been
extended with the genes of novel sequenced organisms.
GO: The Gene Ontology (GO) project was started as a collaborative effort to
address the need for consistent descriptions of gene products in different databases
(Ashburner et al. 2000). Driven by the fact that functional conservation of genes
can be found across all three domains of life, a common language for the annotation of gene products has been established. Thereby, the interoperability of genomic
databases is simplified. Three structured controlled vocabularies (ontologies) have
been defined that describe gene products in terms of their associated biological
processes, cellular components, and molecular functions in a species-independent
manner. The Gene Ontology can be found online at http://www.geneontology.org,
existing terms and their relations can be browsed via the AmiGO ontology
browser.
Précédent

- 357/410

Suivant