3.6 Identification
of Mass
Spectrometry Data
1. Convert raw LC–MS/MS data files to MGF peak lists with
msConvertGUI (ProteoWizard) with the following filter
parameter: msLevel peakPicking true.
2. For thorough protein identification, run the generated peak
lists against the appropriate protein database using MASCOT
(version 2.5.1, Matrix Science Ltd., UK) and X! Tandem
(ALANINE, 2017.02.01, The Global Proteome Machine
Organization) search engines. Set the precursor and fragment
mass tolerance at 20 ppm. Database-searching parameters
include tryptic digestion and modifications: fixed cysteine carbamidomethylation and variable methionine oxidation. For X!
Tandem also select parameters that allowed quick check for
protein N-terminal residue acetylation, peptide N-terminal
glutamine ammonia loss or peptide N-terminal glutamic acid
water loss.
3. Submit the resulting files to Scaffold 4 software (version 4.2.1)
for validation and further analysis. Use the local false discovery
rate scoring algorithm with standard experiment-wide protein
grouping. For the evaluation of peptide hits, select a false
discovery rate less than 1% for peptides and proteins.
4. Export filtered results at the level of the spectrum (PSM),
peptide, or protein identification.
3.7 Proteogenomic
Analysis
1. Download an appropriategenome with its annotation and add
required modifications as separate sequences: pseudogenes,
known genes with alternative start sites, known genes with
possible mutations. Sometimes it is easier to use Biostrings
library of R programming language.
2. Translate the resulting nucleic acid sequences into amino acid
sequences with Biostrings library of R programming language.
Select genetic code #9 for Bacterial, Archaeal and Plant Plastid.
3. Use translated protein database for peptide and protein identification via MASCOT, X! Tandem and Scaffold 4 software as
described in Subheading 3.6. Export results as PSM (peptide to
spectrum match).
4. Analyze the confirmation of pseudogene possible peptide
identifications.
5. Compare the results of known protein identifications and confirmations for alternative start sites and possible mutations.
6. For more reliable validation of the results, synthesize several of
the most questionable of the identified peptides and measure
their MS/MS spectra using the same LC–MS/MS device and
parameters. Confirm each peptide identification comparing
MS/MS spectra of synthetic peptide and the experimental
one using MSnbase library in the R programming language.
Proteogenomic Analysis of Mycobacteria
199
of Mass
Spectrometry Data
1. Convert raw LC–MS/MS data files to MGF peak lists with
msConvertGUI (ProteoWizard) with the following filter
parameter: msLevel peakPicking true.
2. For thorough protein identification, run the generated peak
lists against the appropriate protein database using MASCOT
(version 2.5.1, Matrix Science Ltd., UK) and X! Tandem
(ALANINE, 2017.02.01, The Global Proteome Machine
Organization) search engines. Set the precursor and fragment
mass tolerance at 20 ppm. Database-searching parameters
include tryptic digestion and modifications: fixed cysteine carbamidomethylation and variable methionine oxidation. For X!
Tandem also select parameters that allowed quick check for
protein N-terminal residue acetylation, peptide N-terminal
glutamine ammonia loss or peptide N-terminal glutamic acid
water loss.
3. Submit the resulting files to Scaffold 4 software (version 4.2.1)
for validation and further analysis. Use the local false discovery
rate scoring algorithm with standard experiment-wide protein
grouping. For the evaluation of peptide hits, select a false
discovery rate less than 1% for peptides and proteins.
4. Export filtered results at the level of the spectrum (PSM),
peptide, or protein identification.
3.7 Proteogenomic
Analysis
1. Download an appropriategenome with its annotation and add
required modifications as separate sequences: pseudogenes,
known genes with alternative start sites, known genes with
possible mutations. Sometimes it is easier to use Biostrings
library of R programming language.
2. Translate the resulting nucleic acid sequences into amino acid
sequences with Biostrings library of R programming language.
Select genetic code #9 for Bacterial, Archaeal and Plant Plastid.
3. Use translated protein database for peptide and protein identification via MASCOT, X! Tandem and Scaffold 4 software as
described in Subheading 3.6. Export results as PSM (peptide to
spectrum match).
4. Analyze the confirmation of pseudogene possible peptide
identifications.
5. Compare the results of known protein identifications and confirmations for alternative start sites and possible mutations.
6. For more reliable validation of the results, synthesize several of
the most questionable of the identified peptides and measure
their MS/MS spectra using the same LC–MS/MS device and
parameters. Confirm each peptide identification comparing
MS/MS spectra of synthetic peptide and the experimental
one using MSnbase library in the R programming language.
Proteogenomic Analysis of Mycobacteria
199
