348
V. Mittard-Runte et al.
such as those in the /country and /db_xref qualifiers. Finally, the databases strive to
promote public availability of sequence and annotation data through a collaboration
with publishers.
Each centre strives to impose a quality control upon submitted sequences and
annotation, but the data generator retains editorial responsibility for the biological
content of his/her entries. Increasingly, in the face of ever growing data volumes,
these quality control measures rely on automated validation procedures. Unique
and permanent database entry accession numbers (e.g., AF123456) are provided by
the receiving database upon submission to allow future identification of the entry.
Database accession numbers are required prior to publication by the publishers of
the major molecular biology journals.
Accession number prefixes depend on the issuing database and the type of data
submitted (e.g. direct submission from DDBJ, genome project data from EMBL,
EST from GenBank, patent from JPO, EPO, or USPTO). A complete list of prefix
codes used so far can be found here: http://www.ddbj.nig.ac.jp/sub/prefix.html.
Note that the GI prefix identifiers (e.g. GI:26117688) are internal to Genbank and
are not primary INSDC accession numbers. To resolve a GenInfo identifier or GI,
the Entrez website from NCBI can be used to retrieve the primary INSDC accession
number (in this case it is U00089).
(a) DDBJ at CIB
The DNA Data Bank of Japan (DDBJ) (Sugawara et al. 2008) was created in 1986 at
the National Institute of Genetics (NIG) which was then reorganized as the Center
for Information Biology and DNA Data Bank of Japan (CIB-DDBJ) in 2001. The
centre mainly collects submissions from Japanese laboratories.
A sample entry in DDBJ flat file format can be found in the regularly updated
DDBJ/EMBL/GenBank Feature Table file on the DDBJ website.
Several useful DDBJ annotation examples (e.g. ribosomal RNA, EST,
microsatellite, transposon) can be found on their website.
The DDBJ nucleotide sequence database can be accessed in different ways (see
Table 9.7).
Table 9.7 The different ways to retrieve data at DDBJ
Sequence retrieval at DDBJ
Getentry
Data retrieval of nucleotide sequences mainly by accession numbers
ARSA
All-round Retrieval of Sequence and Annotation. Search of sequence
libraries such as DDBJ and UniProt, sequence-related libraries such
as PROSITE and Pfam, protein 3D structures, and metabolic
pathways
SRS
Sequence Retrieval System offering term search
GIB
Genome information broker or data retrieval and comparative analysis
system for completed genomes
V. Mittard-Runte et al.
such as those in the /country and /db_xref qualifiers. Finally, the databases strive to
promote public availability of sequence and annotation data through a collaboration
with publishers.
Each centre strives to impose a quality control upon submitted sequences and
annotation, but the data generator retains editorial responsibility for the biological
content of his/her entries. Increasingly, in the face of ever growing data volumes,
these quality control measures rely on automated validation procedures. Unique
and permanent database entry accession numbers (e.g., AF123456) are provided by
the receiving database upon submission to allow future identification of the entry.
Database accession numbers are required prior to publication by the publishers of
the major molecular biology journals.
Accession number prefixes depend on the issuing database and the type of data
submitted (e.g. direct submission from DDBJ, genome project data from EMBL,
EST from GenBank, patent from JPO, EPO, or USPTO). A complete list of prefix
codes used so far can be found here: http://www.ddbj.nig.ac.jp/sub/prefix.html.
Note that the GI prefix identifiers (e.g. GI:26117688) are internal to Genbank and
are not primary INSDC accession numbers. To resolve a GenInfo identifier or GI,
the Entrez website from NCBI can be used to retrieve the primary INSDC accession
number (in this case it is U00089).
(a) DDBJ at CIB
The DNA Data Bank of Japan (DDBJ) (Sugawara et al. 2008) was created in 1986 at
the National Institute of Genetics (NIG) which was then reorganized as the Center
for Information Biology and DNA Data Bank of Japan (CIB-DDBJ) in 2001. The
centre mainly collects submissions from Japanese laboratories.
A sample entry in DDBJ flat file format can be found in the regularly updated
DDBJ/EMBL/GenBank Feature Table file on the DDBJ website.
Several useful DDBJ annotation examples (e.g. ribosomal RNA, EST,
microsatellite, transposon) can be found on their website.
The DDBJ nucleotide sequence database can be accessed in different ways (see
Table 9.7).
Table 9.7 The different ways to retrieve data at DDBJ
Sequence retrieval at DDBJ
Getentry
Data retrieval of nucleotide sequences mainly by accession numbers
ARSA
All-round Retrieval of Sequence and Annotation. Search of sequence
libraries such as DDBJ and UniProt, sequence-related libraries such
as PROSITE and Pfam, protein 3D structures, and metabolic
pathways
SRS
Sequence Retrieval System offering term search
GIB
Genome information broker or data retrieval and comparative analysis
system for completed genomes
