3.8 Databases
Selection, Including
Proteogenomic
Approaches
In order to maximize the number and quality of protein identifications, a proper protein database needs to be selected when running
MS data analysis (it is advisable to repeat MS data analysis using
different databases when working with non-model species). When
working with non-model organisms, which are poorly represented
in public protein databases, different and complementary strategies
are needed in order to overcome important limitations during
protein identification and quantification analysis.
1. Automatically annotated—non-reviewed (TrEMBL)—and
manually annotated—reviewed (SwissProt), protein sequence
databases can be downloaded from UniProtKB repositories
(https://www.uniprot.org), while a nonredundant protein
sequence (nr) preformatted BLAST database from NCBI that
includes all non-redundant GenBank CDS (coding sequences)
translations, PDB (Protein Data Bank), SwissProt, PIR (Protein Information Resource), and PRF (Protein Research Foundation) sequences can be downloaded from ftp://ftp.ncbi.nlm.
nih.gov/blast/db. Speed and accuracy of protein identification
analyses are greatly improved if databases are reduced to a
specific taxonomic level.
2. Search for Expression Sequence Tags (ESTs; short, usually less
than 1 kbp, single-pass sequence reads from mRNA converted
to cDNA) available in NCBI for your organisms/species (see
Note 18). Go to https://www.ncbi.nlm.nih.gov/nuccore,
type organism or species name in the search box (use Boolean
operators to refine/customize your search if necessary), then
filter results by selecting “EST” option within “sequence type”
field (additionally you can check also for “mRNA” option
within “molecular types” field). Finally, use the option “send
to” in order to download the full set of selected sequences in
FASTA format. In order to create your customized protein
database from mRNA sequences, either find and extract Open
Reading Frames from all downloaded sequences using
EMBOSS getorf tool, or simply proceed to translate nucleic
acid sequences by using all possible reading frames with
EMBOSS transeq tool. Getorf and transeq bioinformatic
tools are available for free at https://usegalaxy.org. Please
bear in mind that high redundancy in downloaded sequences
is expected, so this can be reduced by using different bioinformatic programs such as CD-HIT (http://weizhongli-lab.org/
cd-hit).
3. Search for RNA-seq projects related to your organism (or very
closely related species) with raw sequencing data deposited in
the Sequence Read Archive (SRA): https://www.ncbi.nlm.nih.
gov/sra (filter by “RNA” option within “Source” field). It is
preferable these data sets were obtained from similar tissues,
90
Angel P. Diz and Paula Sa ´ nchez-Marı ´n
Selection, Including
Proteogenomic
Approaches
In order to maximize the number and quality of protein identifications, a proper protein database needs to be selected when running
MS data analysis (it is advisable to repeat MS data analysis using
different databases when working with non-model species). When
working with non-model organisms, which are poorly represented
in public protein databases, different and complementary strategies
are needed in order to overcome important limitations during
protein identification and quantification analysis.
1. Automatically annotated—non-reviewed (TrEMBL)—and
manually annotated—reviewed (SwissProt), protein sequence
databases can be downloaded from UniProtKB repositories
(https://www.uniprot.org), while a nonredundant protein
sequence (nr) preformatted BLAST database from NCBI that
includes all non-redundant GenBank CDS (coding sequences)
translations, PDB (Protein Data Bank), SwissProt, PIR (Protein Information Resource), and PRF (Protein Research Foundation) sequences can be downloaded from ftp://ftp.ncbi.nlm.
nih.gov/blast/db. Speed and accuracy of protein identification
analyses are greatly improved if databases are reduced to a
specific taxonomic level.
2. Search for Expression Sequence Tags (ESTs; short, usually less
than 1 kbp, single-pass sequence reads from mRNA converted
to cDNA) available in NCBI for your organisms/species (see
Note 18). Go to https://www.ncbi.nlm.nih.gov/nuccore,
type organism or species name in the search box (use Boolean
operators to refine/customize your search if necessary), then
filter results by selecting “EST” option within “sequence type”
field (additionally you can check also for “mRNA” option
within “molecular types” field). Finally, use the option “send
to” in order to download the full set of selected sequences in
FASTA format. In order to create your customized protein
database from mRNA sequences, either find and extract Open
Reading Frames from all downloaded sequences using
EMBOSS getorf tool, or simply proceed to translate nucleic
acid sequences by using all possible reading frames with
EMBOSS transeq tool. Getorf and transeq bioinformatic
tools are available for free at https://usegalaxy.org. Please
bear in mind that high redundancy in downloaded sequences
is expected, so this can be reduced by using different bioinformatic programs such as CD-HIT (http://weizhongli-lab.org/
cd-hit).
3. Search for RNA-seq projects related to your organism (or very
closely related species) with raw sequencing data deposited in
the Sequence Read Archive (SRA): https://www.ncbi.nlm.nih.
gov/sra (filter by “RNA” option within “Source” field). It is
preferable these data sets were obtained from similar tissues,
90
Angel P. Diz and Paula Sa ´ nchez-Marı ´n
