9 Genomic Techniques and How to Apply Them to Marine Questions
347
9.3.5 Major Public Sequences Databases and Other Resources
This section presents the major public sequence databases and other resources that
aim to provide a comprehensive coverage of sequences and annotations available to
the scientific community.
In the first part, we will initially introduce the three leading nucleotide sequence
centres and explain how to access the collected data. Following this, the different
procedures of submission required by these centres, depending on the type of data
being submitted, are described. The second part focuses on UniProt, which is a
database for protein sequences. Finally, an introduction is provided to some other
resources that may be useful to marine biologists.
9.3.5.1 Major Public Nucleotide Sequences Databases
The three major public database centres are located in Europe, Japan, and the
USA (see Table 9.6). These databases form the International Nucleotide Sequence
Database Collaboration (INSDC, www.insdc.org) and have a long-established collaboration of over 18 years involving the daily exchange of new and updated
nucleotide sequence and annotation data (Brunak et al. 2002). Collectively, the
databases aim to provide a comprehensive coverage of sequences and annotations
available within the public domain.
Table 9.6 The major public nucleotide sequences databases and their database centres
Major public database centres
Europe
Japan
USA
EBI (European Bioinformatics
Institute)
CIB (Center for Information
Biology) at the NIG
(National Institute of
Genetics)
NCBI (National Center for
Biotechnology Information)
Established in 1993
Established in 1995
Established in 1988
www.ebi.ac.uk
www.ddbj.nig.ac.jp
www.ncbi.nlm.nih.gov
EMBL-Bank created in 1981
(Cochrane et al. 2008)
DDBJ created in 1986
(Sugawara et al. 2008)
GenBank created in 1982
(Benson et al. 2008)
Three activities are central to the ongoing INSDC effort. Firstly, the databases
provide submission services for data generators, so that information can be submitted as easily as possible, while retaining important contextual biological information
(such as the biological source of the sequence) and functional interpretations of the
sequence data in the form of an annotation. Secondly, the databases develop structures and formats through which the sequence and annotation data can be accurately
and concisely represented. The focus is on the usability for the users; instruments
include the INSDC Feature Table Definition document and associated vocabularies,
Précédent

- 358/410

Suivant