199
High-Resolution Proteomic Analyses
is also detrimental to the search outcomes. Currently,
there are multiple disjoint sources of protein sequences in
Xenopus, which we list in Table 13.1. These databases are
largely but not entirely mutually redundant. One of the outstanding challenges for the feld is to arrive at a single gold
standard set of sequences to be used across studies. Selecting
such a set has not been accomplished even in more common
model species so far, and for a good reason. The genomebased sequences are likely to provide “phantom” sequences
due to imperfect gene models and missing sequences due to
unidentifed splice variants. The federated collections such
as UniProt will have curated a high-quality set of sequences,
but it is bound to be incomplete. Presence of allo-alleles in
X. laevis presents a separate challenge since allele-specif c
expression might be of interest. As an alternative to all these,
we took a genome-free approach in which we compiled a
set of all transcripts detected in X. laevis and used coding
frame-aware translation to predict expressed proteins (Wühr
et al., 2014).
There are several groups of sequences that warrant special consideration when selecting and/or compiling a set of
reference sequences, provided in the following in a list that
is by no means complete:
• Genes which are simply missing in the genomic
sequence because the DNA sequencing pipeline is
imperfect. Until genome assembly improves, this
issue will remain problematic.
• Genes which are in the current genome assembly
scaffolds but are missed by gene modeling and thus
are missing from the set of predicted mRNAs and
proteins. This issue might be rectifed in the next
release of the genome assembly. If you are aware
of such a gene/protein, just add its sequence to the
database before you search against it.
• Proteins that undergo atypical post-translational
sequence editing (not covalent addition of a certain functional group like phosphorylation), such as
amino acid sequence cleavage on the way to a mature
protein, for example, signaling peptides or proteasomally processed immuno-peptides. Translating
cDNA in these cases does not predict the correct
protein sequence. Currently, in these cases, the correct protein or peptide sequence needs to be added
to the database.
• Selenoproteins—a few dozen proteins (e.g. selenoproteins S, K, N, O, and I) use a non-canonical
amino acid—“selenocysteine”—coded for by one
of three stop codons, and this peculiarity of context-dependent genetic code is seldom taken into
account when gene models are built. Usually, this
imperfection results in a premature termination of
a reference sequence, which is thus missing a portion of the actual sequence.
• There are 13 abundant and functionally important
proteins encoded by mitochondrial DNA which are
not part of nuclear genome-derived gene models
and should always be appended to a database.
13.3. NOMENCLATURE AND GENE SYMBOLS
Accurate detection of proteins is the necessary frst step, but
biological interpretation of the observations requires biological knowledge of the proteins. For high-dimensional data,
often the frst analysis is to look for enrichment of different
pathways, cell locations, or protein functions. In order to do
this, gene symbol annotation is necessary. Systems biology
rationalizes the organization of biological function at the
level of gene sets such as molecular pathways or protein complexes. Such sets are mainly defned and studied using human
cells and human gene nomenclature. Sequences in other
species are traditionally named using homology to already
named sequences with the assumption that sequence homology often implies functional homology. When gene models
are released with a genome, the assigned gene symbols take
into account species- and f eld-specifc context. For example, a gene identifed as Xelaev18018806m at the release of
X.laevis genome v1.8 is assigned the symbol nodal5.3.L ,which
refects the Xenopus -specif c nodal gene family expansion.
However, what matters for gene set analysis is that this gene
is best matched by the human sp|Q96S42|NODAL_HUMAN
TABLE 13.1
A List of Available Resources to Obtain a Reference Set of Protein Sequences for Peptide-Spectra Matching of
X. laevis Data
Source
Description
Source URL
Size
Remarks and Criticism
Genome-derived
Based on gene models
http://ftp.xenbase.org/pub/
45K seqs
Only as complete and accurate as gene
Genomics/JGI/
21 Mb
models are
“Phrog”
RNA-seq derived
https://scholar.princeton.edu/
80K seqs
Redundant, wildcards in sequences
( Wühr et al., 2014 )
wuehr/sample_prep
30.5 Mb
UniProt
Mostly uncurated, assembled
www.uniprot.org/uniprot/?query= 61K seqs
Severely incomplete and redundant, only
from misc.
organism%3A%22XENLA
23 Mb
5K reviewed
Marcotte Lab
Compiled from miscellaneous
https://github.com/marcottelab/
25 K seqs
Contains allo-alleles fused into a single
sources, then curated
pivo
18 Mb
sequence from non-redundant peptides
High-Resolution Proteomic Analyses
is also detrimental to the search outcomes. Currently,
there are multiple disjoint sources of protein sequences in
Xenopus, which we list in Table 13.1. These databases are
largely but not entirely mutually redundant. One of the outstanding challenges for the feld is to arrive at a single gold
standard set of sequences to be used across studies. Selecting
such a set has not been accomplished even in more common
model species so far, and for a good reason. The genomebased sequences are likely to provide “phantom” sequences
due to imperfect gene models and missing sequences due to
unidentifed splice variants. The federated collections such
as UniProt will have curated a high-quality set of sequences,
but it is bound to be incomplete. Presence of allo-alleles in
X. laevis presents a separate challenge since allele-specif c
expression might be of interest. As an alternative to all these,
we took a genome-free approach in which we compiled a
set of all transcripts detected in X. laevis and used coding
frame-aware translation to predict expressed proteins (Wühr
et al., 2014).
There are several groups of sequences that warrant special consideration when selecting and/or compiling a set of
reference sequences, provided in the following in a list that
is by no means complete:
• Genes which are simply missing in the genomic
sequence because the DNA sequencing pipeline is
imperfect. Until genome assembly improves, this
issue will remain problematic.
• Genes which are in the current genome assembly
scaffolds but are missed by gene modeling and thus
are missing from the set of predicted mRNAs and
proteins. This issue might be rectifed in the next
release of the genome assembly. If you are aware
of such a gene/protein, just add its sequence to the
database before you search against it.
• Proteins that undergo atypical post-translational
sequence editing (not covalent addition of a certain functional group like phosphorylation), such as
amino acid sequence cleavage on the way to a mature
protein, for example, signaling peptides or proteasomally processed immuno-peptides. Translating
cDNA in these cases does not predict the correct
protein sequence. Currently, in these cases, the correct protein or peptide sequence needs to be added
to the database.
• Selenoproteins—a few dozen proteins (e.g. selenoproteins S, K, N, O, and I) use a non-canonical
amino acid—“selenocysteine”—coded for by one
of three stop codons, and this peculiarity of context-dependent genetic code is seldom taken into
account when gene models are built. Usually, this
imperfection results in a premature termination of
a reference sequence, which is thus missing a portion of the actual sequence.
• There are 13 abundant and functionally important
proteins encoded by mitochondrial DNA which are
not part of nuclear genome-derived gene models
and should always be appended to a database.
13.3. NOMENCLATURE AND GENE SYMBOLS
Accurate detection of proteins is the necessary frst step, but
biological interpretation of the observations requires biological knowledge of the proteins. For high-dimensional data,
often the frst analysis is to look for enrichment of different
pathways, cell locations, or protein functions. In order to do
this, gene symbol annotation is necessary. Systems biology
rationalizes the organization of biological function at the
level of gene sets such as molecular pathways or protein complexes. Such sets are mainly defned and studied using human
cells and human gene nomenclature. Sequences in other
species are traditionally named using homology to already
named sequences with the assumption that sequence homology often implies functional homology. When gene models
are released with a genome, the assigned gene symbols take
into account species- and f eld-specifc context. For example, a gene identifed as Xelaev18018806m at the release of
X.laevis genome v1.8 is assigned the symbol nodal5.3.L ,which
refects the Xenopus -specif c nodal gene family expansion.
However, what matters for gene set analysis is that this gene
is best matched by the human sp|Q96S42|NODAL_HUMAN
TABLE 13.1
A List of Available Resources to Obtain a Reference Set of Protein Sequences for Peptide-Spectra Matching of
X. laevis Data
Source
Description
Source URL
Size
Remarks and Criticism
Genome-derived
Based on gene models
http://ftp.xenbase.org/pub/
45K seqs
Only as complete and accurate as gene
Genomics/JGI/
21 Mb
models are
“Phrog”
RNA-seq derived
https://scholar.princeton.edu/
80K seqs
Redundant, wildcards in sequences
( Wühr et al., 2014 )
wuehr/sample_prep
30.5 Mb
UniProt
Mostly uncurated, assembled
www.uniprot.org/uniprot/?query= 61K seqs
Severely incomplete and redundant, only
from misc.
organism%3A%22XENLA
23 Mb
5K reviewed
Marcotte Lab
Compiled from miscellaneous
https://github.com/marcottelab/
25 K seqs
Contains allo-alleles fused into a single
sources, then curated
pivo
18 Mb
sequence from non-redundant peptides
