During the last decade, important methodological and technological improvements in protein sample preparation (extraction,
separation, and fractionation) and mass spectrometry field together
with such exponential increase of genome sequencing projects for
different organisms, and nucleotide and protein sequences in public
databases, have made the production of high-throughput and highquality proteomic data a less challenging task than it used to
be. There has also been important advance in proteomic data
analysis, generally after bringing and adjusting to proteomic field
models and methods that were implemented and successfully used
in transcriptomics. Not surprisingly, these advances have been
mainly exploited in research projects involving model organisms
with abundant genomic and proteomic information available in
public databases, a necessary requirement to get a high number of
proteins accurately identified and quantified through mass spectrometry data analysis. The situation has been more challenging
when working with non-model organisms due to their poor representation in nucleotide and protein sequence databases. However
different alternatives exist to overcome this limitation. It is actually
becoming easier and easier to prepare customized protein databases. For example, the translation of nucleotide sequences
obtained from species and tissue-specific transcriptomic projects
has proven to be a very useful approximation to significantly
increase the number and quality of protein identifications. Thanks
to important advances in massive parallel sequencing technology
together with a continuously dropping of sequencing prices, it is
now affordable for most laboratories to run genomic and transcriptomic analyses, hence to produce customized protein databases for
any organism. Most remarkably, the complementarity and synergies
generated by the joint interpretation of the results obtained from
analyses performed at different omic levels (multiomics) are allowing to address a broad range of complex biological questions that
cannot be properly answered using a single level approach and have
actually led to the emergence of a novel field known as proteogenomics. Initially, the term was coined to encompass all studies
where proteomic data are used to improve and expand genomic
annotations, but now it has been expanded to other studies describing much more applications. In fact, there are important examples
in the literature that indicate the great utility and advantages of
following a proteogenomic approach for both model and
non-model organisms [7–9].
Quantitative proteomic analysis has traditionally relied on two
different workflows. A first one based on two-dimensional electrophoresis (2-DE) coupled to mass spectrometry (MS) analysis for
protein separation and identification, respectively, and a second one
based on a gel-free LC–MS approach where peptides, resulting
from the enzymatic in-solution digestion (usually with trypsin) of
total protein content, are separated and usually fractionated
Shotgun Proteomics in Non-model Organisms
79
separation, and fractionation) and mass spectrometry field together
with such exponential increase of genome sequencing projects for
different organisms, and nucleotide and protein sequences in public
databases, have made the production of high-throughput and highquality proteomic data a less challenging task than it used to
be. There has also been important advance in proteomic data
analysis, generally after bringing and adjusting to proteomic field
models and methods that were implemented and successfully used
in transcriptomics. Not surprisingly, these advances have been
mainly exploited in research projects involving model organisms
with abundant genomic and proteomic information available in
public databases, a necessary requirement to get a high number of
proteins accurately identified and quantified through mass spectrometry data analysis. The situation has been more challenging
when working with non-model organisms due to their poor representation in nucleotide and protein sequence databases. However
different alternatives exist to overcome this limitation. It is actually
becoming easier and easier to prepare customized protein databases. For example, the translation of nucleotide sequences
obtained from species and tissue-specific transcriptomic projects
has proven to be a very useful approximation to significantly
increase the number and quality of protein identifications. Thanks
to important advances in massive parallel sequencing technology
together with a continuously dropping of sequencing prices, it is
now affordable for most laboratories to run genomic and transcriptomic analyses, hence to produce customized protein databases for
any organism. Most remarkably, the complementarity and synergies
generated by the joint interpretation of the results obtained from
analyses performed at different omic levels (multiomics) are allowing to address a broad range of complex biological questions that
cannot be properly answered using a single level approach and have
actually led to the emergence of a novel field known as proteogenomics. Initially, the term was coined to encompass all studies
where proteomic data are used to improve and expand genomic
annotations, but now it has been expanded to other studies describing much more applications. In fact, there are important examples
in the literature that indicate the great utility and advantages of
following a proteogenomic approach for both model and
non-model organisms [7–9].
Quantitative proteomic analysis has traditionally relied on two
different workflows. A first one based on two-dimensional electrophoresis (2-DE) coupled to mass spectrometry (MS) analysis for
protein separation and identification, respectively, and a second one
based on a gel-free LC–MS approach where peptides, resulting
from the enzymatic in-solution digestion (usually with trypsin) of
total protein content, are separated and usually fractionated
Shotgun Proteomics in Non-model Organisms
79
