9 Genomic Techniques and How to Apply Them to Marine Questions
353
(INSDC). UniMES is available on the FTP site in FASTA format with a “UniMES
matches to InterPro methods” file.
The UniProt Archive (UniParc) is a non-redundant database that aims to capture all publicly available protein sequences. UniParc stores each unique sequence
only once and gives it a stable and unique identifier (UPI), making it possible to
identify the same protein from different source databases. A UPI is never removed,
changed, or reassigned. The basic information stored within each UniParc entry is
the identifier, the sequence, cyclic redundancy check number, source database(s)
with accession and version numbers, and a time stamp. If a UniParc entry does not
have a cross-reference to a UniProtKB entry, the reason for the exclusion of that
sequence from UniProtKB is provided (e.g. patent).
More information about UniParc can be found on the UniProt website.
(b) Release, Submission, and Access
UniProt is updated every 3 weeks and a major release is also produced three times
per year. The first major release was in December 2003. The latest release (release
13, February 26, 2008) includes 5,751,608 entries as follows:
• 356,194 UniProtKB/Swiss-Prot entries (release 55) and 11,290 different species,
• 5,395,414 UniProtKB/TrEMBL entries (release 38) and 1,552,882 different
species.
If you have a new protein sequence you can submit it directly to UniProtKB
using the web-based tool SPIN at the EBI.
The UniProt website can be accessed at http://www.uniprot.org where examples
with the protein accession number, the UniRef entries, and UniParc unique identifier
can be found.
9.3.5.3 RefSeq
The Reference Sequence (RefSeq) database is a non-redundant collection of
transcripts, proteins, and genomic regions (http://www.ncbi.nlm.nih.gov/RefSeq/)
produced by NCBI (Pruitt et al. 2007). RefSeq is limited to major organisms for
which sufficient data is available (almost 5,000 distinct organisms as of January
2008, release 27), while GenBank includes sequences for any organism submitted
(more than 250,000 different named organisms). Each RefSeq entry represents a
single sequence from one organism. The annotation status within the entries varies
and includes not annotated (inferred, model, predicted, provisional or WGS) or
annotated (validated and reviewed) entries.
The Reference Sequence (RefSeq) database can be accessed in different ways,
either directly by querying or indirectly through links provided from several NCBI
resources including Gene, Entrez, PubMed, and Map Viewer.
RefSeq uses the following prefixes with two characters followed by an underscore character (“_”) such as NP_010000. More information on the RefSeq
accession number format can be found on the RefSeq website.
Précédent

- 364/410

Suivant