198
Xenopus
What is so challenging about this technology? The f rst
issue is that large amounts of protein material are required.
This is because there is currently no way to amplify proteins (analogous to reverse transcription PCR followed by
deep sequencing for RNA quantif cation, i.e., RNAseq) and
because the sample preparation, and especially ionization,
results in material loss. In this case, the Xenopus model system, and in particular X. laevis, is an ideal model since a
single egg or embryo contains ~25 μg of non-yolk protein
(Gurdon et al., 1986). A single egg is suffcient for shallow
protein measurement, and hundreds more sibling embryos
can easily be collected from a single clutch, which is more
than suffcient for deeper measurement. Starting material
amount is even more of a concern when measuring posttranslational modifcations. This is because an enrichment
step has to be done to study posttranslational modif cations
since even the most abundant modif cation, phosphorylation, is present only in about 1% of peptides (Peuchen et al.,
2016; Peuchen et al., 2017; Presler et al., 2017). The second
challenge is depth of measurement. Unlike RNAseq, tandem mass spectrometry is not a parallelized measurement;
the measurement of each peptide comes at the expense of
an unmeasured one. Usually, the most abundant peptides
are measured. Fortunately, egg/embryo lysis conditions
that effectively remove abundant yolk proteins have been
developed (Gupta et al., 2018; Wühr et al., 2014), which
has made mass spectrometry-based proteomics possible
in Xenopus. Measuring more deeply requires more instrument time, which means higher cost per experiment. A third
challenge is the bioinformatics. In nucleotide sequencing,
a priori knowledge of the genome or transcriptome is not
needed to obtain nucleotide reads. In contrast, protein mass
spectrometry measures the mass and charge of peptides and
peptide fragments. The measurements that result from intact
and fragmented peptides are matched to a set of theoretical
digested peptides from a database of protein references. This
means that having an accurate reference database is critical
to identify and then measure a specifc peptide. To make
matters worse, mere completeness of the reference database
is not suffcient; inclusion of artifact sequences (that could
never be observed in vivo) actually hampers identif cation
of other peptides.
There are other non-technical reasons successful application of mass spectrometry to perform quantitative proteomics
in Xenopus has materialized only recently. Proteomics
by mass spectrometry is lagging behind transcriptomics,
which has been made routine mostly by the availability of
Illumina sequencing machines and standardized protocols
and standard operating procedures throughout industrial
and academic sequencing facilities. The best MS instruments produced by Thermo-Fisher are expensive and are
impossible for a non-specialist to use. When modern MS
instruments are available via a “fee-for-service” model at
facility cores, both the pre- and post-instrument processing
pipelines are tailored to human and mouse samples. Perhaps
most importantly, the bioinformatics support for proteomics
in non-model organisms is lacking. Given this situation, it is
not surprising that successful applications of mass spectrometry proteomics in Xenopus so far have resulted from close
collaborations between Xenopus experts and MS specialists.
A few notable examples of such collaborations are the labs
of Moody with Nemes (Lombard-Banek et al., 2016; Onjiko
et al., 2015), Huber with Dovichi (Peuchen et al., 2017;
Sun et al., 2016), Kirschner with Gygi (Wühr et al., 2014;
Peshkin et al., 2015), Veenstra with Vermeulen (Lindeboom
et al., 2019; Smits et al., 2014), and Klein with Garcia (SahaShah et al., 2019). Very few laboratories have the expertise
in both advanced mass spectrometry and Xenopus methods,
with the one notable exception being Martin Wühr’s lab at
Princeton (Gupta et al., 2018; Sonnett et al., 2018a, Sonnett
et al., 2018b).
One key milestone on the way towards Xenopus mass
spectrometry proteomics was sequencing the genomes of
both Xenopus species actively used in research—X. laevis and X. tropicalis (Hellsten et al., 2010; Session et al.,
2016)—which naturally benefts systems-level analysis at
the genome, RNA, and protein levels. Being able to attribute
spectra to peptides relies on the availability of a complete
and accurate list of protein sequences, for which a complete
genome with a high-quality set of gene models is the key.
At the point of being released, the genome quality is judged
by various metrics of the DNA sequence itself, not so much
the quality of gene models. Notably, proteomics studies feed
back to the gene model quality by providing an accurate catalogue of detected peptides and respective proteins that can
be used to validate and adjust the intron-exon-UTR annotation of gene structures. Both the X. laevis and X. tropicalis
genomes have undergone many iterations of genome assembly and even more of refning the gene models. X. laevis
is an allotetraploid species, whereas X. tropicalis is a true
diploid. Unsurprisingly, given their respective complexity,
the annotation of the X. tropicalis genome (~24K gene models) is currently in a much better state than that of X. laevis
(~44K gene models), which presents a challenge when comparing the number of distinct proteins characterized across
different studies.
13.2. THE DATABASE
For the methods described in this perspective, one can only
detect a protein of a pre-defned sequence, and thus the
spectra must be matched against a set of potential peptide
sequences. Thus, having a sound and complete reference
set of sequences (sometimes referred to as “the database”
in the feld) is particularly important. The reason for soundness deserves a special explanation, since one might expect
that mixing in arbitrary sequences (e.g. sea urchin, since sea
urchin peptide material is never observed in Xenopus samples) with Xenopus ones would be benign. However, such
nonsense sequences affect the peptide-to-spectra matching
algorithms by producing spuriously matched peptides, and
at a given false discovery rate (FDR), fewer real proteins will
remain. Having identical or similar sequences (from alloalleles or highly conserved protein families such as histones)
Xenopus
What is so challenging about this technology? The f rst
issue is that large amounts of protein material are required.
This is because there is currently no way to amplify proteins (analogous to reverse transcription PCR followed by
deep sequencing for RNA quantif cation, i.e., RNAseq) and
because the sample preparation, and especially ionization,
results in material loss. In this case, the Xenopus model system, and in particular X. laevis, is an ideal model since a
single egg or embryo contains ~25 μg of non-yolk protein
(Gurdon et al., 1986). A single egg is suffcient for shallow
protein measurement, and hundreds more sibling embryos
can easily be collected from a single clutch, which is more
than suffcient for deeper measurement. Starting material
amount is even more of a concern when measuring posttranslational modifcations. This is because an enrichment
step has to be done to study posttranslational modif cations
since even the most abundant modif cation, phosphorylation, is present only in about 1% of peptides (Peuchen et al.,
2016; Peuchen et al., 2017; Presler et al., 2017). The second
challenge is depth of measurement. Unlike RNAseq, tandem mass spectrometry is not a parallelized measurement;
the measurement of each peptide comes at the expense of
an unmeasured one. Usually, the most abundant peptides
are measured. Fortunately, egg/embryo lysis conditions
that effectively remove abundant yolk proteins have been
developed (Gupta et al., 2018; Wühr et al., 2014), which
has made mass spectrometry-based proteomics possible
in Xenopus. Measuring more deeply requires more instrument time, which means higher cost per experiment. A third
challenge is the bioinformatics. In nucleotide sequencing,
a priori knowledge of the genome or transcriptome is not
needed to obtain nucleotide reads. In contrast, protein mass
spectrometry measures the mass and charge of peptides and
peptide fragments. The measurements that result from intact
and fragmented peptides are matched to a set of theoretical
digested peptides from a database of protein references. This
means that having an accurate reference database is critical
to identify and then measure a specifc peptide. To make
matters worse, mere completeness of the reference database
is not suffcient; inclusion of artifact sequences (that could
never be observed in vivo) actually hampers identif cation
of other peptides.
There are other non-technical reasons successful application of mass spectrometry to perform quantitative proteomics
in Xenopus has materialized only recently. Proteomics
by mass spectrometry is lagging behind transcriptomics,
which has been made routine mostly by the availability of
Illumina sequencing machines and standardized protocols
and standard operating procedures throughout industrial
and academic sequencing facilities. The best MS instruments produced by Thermo-Fisher are expensive and are
impossible for a non-specialist to use. When modern MS
instruments are available via a “fee-for-service” model at
facility cores, both the pre- and post-instrument processing
pipelines are tailored to human and mouse samples. Perhaps
most importantly, the bioinformatics support for proteomics
in non-model organisms is lacking. Given this situation, it is
not surprising that successful applications of mass spectrometry proteomics in Xenopus so far have resulted from close
collaborations between Xenopus experts and MS specialists.
A few notable examples of such collaborations are the labs
of Moody with Nemes (Lombard-Banek et al., 2016; Onjiko
et al., 2015), Huber with Dovichi (Peuchen et al., 2017;
Sun et al., 2016), Kirschner with Gygi (Wühr et al., 2014;
Peshkin et al., 2015), Veenstra with Vermeulen (Lindeboom
et al., 2019; Smits et al., 2014), and Klein with Garcia (SahaShah et al., 2019). Very few laboratories have the expertise
in both advanced mass spectrometry and Xenopus methods,
with the one notable exception being Martin Wühr’s lab at
Princeton (Gupta et al., 2018; Sonnett et al., 2018a, Sonnett
et al., 2018b).
One key milestone on the way towards Xenopus mass
spectrometry proteomics was sequencing the genomes of
both Xenopus species actively used in research—X. laevis and X. tropicalis (Hellsten et al., 2010; Session et al.,
2016)—which naturally benefts systems-level analysis at
the genome, RNA, and protein levels. Being able to attribute
spectra to peptides relies on the availability of a complete
and accurate list of protein sequences, for which a complete
genome with a high-quality set of gene models is the key.
At the point of being released, the genome quality is judged
by various metrics of the DNA sequence itself, not so much
the quality of gene models. Notably, proteomics studies feed
back to the gene model quality by providing an accurate catalogue of detected peptides and respective proteins that can
be used to validate and adjust the intron-exon-UTR annotation of gene structures. Both the X. laevis and X. tropicalis
genomes have undergone many iterations of genome assembly and even more of refning the gene models. X. laevis
is an allotetraploid species, whereas X. tropicalis is a true
diploid. Unsurprisingly, given their respective complexity,
the annotation of the X. tropicalis genome (~24K gene models) is currently in a much better state than that of X. laevis
(~44K gene models), which presents a challenge when comparing the number of distinct proteins characterized across
different studies.
13.2. THE DATABASE
For the methods described in this perspective, one can only
detect a protein of a pre-defned sequence, and thus the
spectra must be matched against a set of potential peptide
sequences. Thus, having a sound and complete reference
set of sequences (sometimes referred to as “the database”
in the feld) is particularly important. The reason for soundness deserves a special explanation, since one might expect
that mixing in arbitrary sequences (e.g. sea urchin, since sea
urchin peptide material is never observed in Xenopus samples) with Xenopus ones would be benign. However, such
nonsense sequences affect the peptide-to-spectra matching
algorithms by producing spuriously matched peptides, and
at a given false discovery rate (FDR), fewer real proteins will
remain. Having identical or similar sequences (from alloalleles or highly conserved protein families such as histones)
