352
V. Mittard-Runte et al.
Table 9.11 Organisation of the UniProt databases
The universal protein resource
UniProtKB: protein
knowledgebase
UniRef: sequence
clusters
UniMES
UniParc
UniProtKB/SwissProt (manually
curated
annotation)
UniProtKB/TrEMBL
(automatic
annotation)
UniRef100
UniRef90 (at least
90% sequence
identity)
UniRef50 (at least
50% sequence
identity)
Metagenomic and
environmental
sample sequences
UniProt archive
protein sequence database, which brings together experimental results and computed features and UniProtKB/TrEMBL, a high quality automatically annotated
database. TrEMBL entries are manually annotated and integrated into Swiss-Prot,
keeping their unique accession number.
UniProtKB contains all the protein sequences available except for the following
ones:
• Most non-germline immunoglobulins and T-cell receptors
• Synthetic sequences
• Most patent application sequences
• Small fragments encoded from nucleotide sequence (<8 amino acids)
• Pseudogenes
• Fusion/truncated proteins
• Not a real protein (when enough evidence that the existence of a protein looks
dubious)
The first five types of sequences are identified automatically during the creation
of UniProtKB/TrEMBL. The last two types are manually identified by curators (e.g.,
sequences derived from gene predictions from genomic sequences, which were
wrongly predicted to code for proteins) and then removed.
All these seven types of excluded sequences are available in UniParc with a
reason for their exclusion from UniProtKB (see paragraph below).
More information about UniProtKB can be found on the UniProt website.
The UniProt Reference Clusters (UniRef) provide clustered sets of closely
related sequences from the UniProt Knowledgebase to allow fast searches.
UniRef90 and UniRef50 are composed of sequences that have at least 90% or 50%
sequence identity, respectively. More information about UniRef can be found on the
UniProt website.
The UniProt Metagenomic and Environmental Sequences database (UniMES)
currently contains data from the Global Ocean Sampling Expedition (GOS), which
were originally submitted to the International Nucleotide Sequence Databases
V. Mittard-Runte et al.
Table 9.11 Organisation of the UniProt databases
The universal protein resource
UniProtKB: protein
knowledgebase
UniRef: sequence
clusters
UniMES
UniParc
UniProtKB/SwissProt (manually
curated
annotation)
UniProtKB/TrEMBL
(automatic
annotation)
UniRef100
UniRef90 (at least
90% sequence
identity)
UniRef50 (at least
50% sequence
identity)
Metagenomic and
environmental
sample sequences
UniProt archive
protein sequence database, which brings together experimental results and computed features and UniProtKB/TrEMBL, a high quality automatically annotated
database. TrEMBL entries are manually annotated and integrated into Swiss-Prot,
keeping their unique accession number.
UniProtKB contains all the protein sequences available except for the following
ones:
• Most non-germline immunoglobulins and T-cell receptors
• Synthetic sequences
• Most patent application sequences
• Small fragments encoded from nucleotide sequence (<8 amino acids)
• Pseudogenes
• Fusion/truncated proteins
• Not a real protein (when enough evidence that the existence of a protein looks
dubious)
The first five types of sequences are identified automatically during the creation
of UniProtKB/TrEMBL. The last two types are manually identified by curators (e.g.,
sequences derived from gene predictions from genomic sequences, which were
wrongly predicted to code for proteins) and then removed.
All these seven types of excluded sequences are available in UniParc with a
reason for their exclusion from UniProtKB (see paragraph below).
More information about UniProtKB can be found on the UniProt website.
The UniProt Reference Clusters (UniRef) provide clustered sets of closely
related sequences from the UniProt Knowledgebase to allow fast searches.
UniRef90 and UniRef50 are composed of sequences that have at least 90% or 50%
sequence identity, respectively. More information about UniRef can be found on the
UniProt website.
The UniProt Metagenomic and Environmental Sequences database (UniMES)
currently contains data from the Global Ocean Sampling Expedition (GOS), which
were originally submitted to the International Nucleotide Sequence Databases
