9 Genomic Techniques and How to Apply Them to Marine Questions
351
(e) Third Party Annotation
The INSDC Third Party Annotation (TPA) section (Cochrane et al. 2006) was created in 2002 in order to accept high-quality annotations of nucleotide sequences
from submitters who have not themselves generated those nucleotide sequences.
The quality of the annotation is supported by experimental or inferential analyses.
TPA data are divided into two tiers:
• with experimental evidence (e.g., BK000016)
• with inferential evidence – where the source molecule or its product(s) have not
been the subject of direct experimentation (e.g., BK000554).
More information about TPA data can be found at the DDBJ, EMBL, or GenBank
websites.
9.3.5.2 Major Public Protein Sequences Database: UniProt
Dr. Margaret Oakley Dayhoff (1925–1983) was a pioneer in the bioinformatics field
of computers in chemistry and biology, and produced the Atlas of Protein Sequence
and Structure, published from 1965 to 1978 by the National Biomedical Research
Foundation (NBRF). The Protein Information Resource (PIR) was then established
in 1984 by the NBRF as a first resource of protein sequence information. It produced
the protein sequence database called PIR-PSD until the end of 2004.
Swiss-Prot, an annotated sequence database, was created in 1986 by Amos
Bairoch in Geneva, Switzerland and the first entry was human cytochrome
c (http://www.expasy.ch/uniprot/P99999). The Swiss Institute of Bioinformatics
(SIB) has hosted the Swiss-Prot group since 1998 and also currently
maintains the Expert Protein Analysis System (ExPASy) proteomics server
(http://www.expasy.org). This server is dedicated for more than 15 years to the
analysis of protein sequences and structures as well as two-dimensional gel
electrophoresis (2D Page electrophoresis).
To cope with the growing amount of sequence data, TrEMBL was created at the
EBI (European Bioinformatics Institute) in 1996. It is a computer-annotated protein
sequence database containing translations of all coding sequences (CDS) present in
the DDBJ/EMBL/GenBank Nucleotide Sequence Databases, which are not yet in
Swiss-Prot. SIB and EBI combined efforts from the beginning to jointly produce
Swiss-Prot and TrEMBL.
In 2002, the three institutes (EBI, PIR, and SIB) pooled their resources and expertise. The Universal Protein Resource (UniProt) Consortium (Consortium 2008)
was born to provide a single database of high-quality and comprehensive protein
sequence and functional information (http://www.uniprot.org).
(a) Organisation of the UniProt Databases (Table 9.11)
The UniProt Knowledgebase (UniProtKB) is composed of two databases:
UniProtKB/Swiss-Prot, a high-quality manually annotated and non-redundant
Précédent

- 362/410

Suivant