Chapter 12
Proteogenomic Approach for Mycobacterium tuberculosis
Investigation
Julia Bespyatykh, Georgij Arapidi, and Egor Shitikov
Abstract
Recent advances in MS/MS technology have made it possible to use proteomic data to predict proteincoding sequences. This approach is called proteogenomics, and it allows to correctly translate start and stop
sites and to reveal new open reading frames. Here, we focus on using proteogenomics to improve the
annotation of Mycobacteriumtuberculosis strains. We also describe detail procedures of the extraction of
proteins and their further preparation for LC–MS/MS analysis and outline the main steps of data analysis.
Key words Proteome, Beijing B0/W148, Proteomic, Label-free, Genome
1 Introduction
Nowadays, the starting point in the study of any living organism is
sequencing and annotation of its genome. The correct deciphering
of genomic information may be further facilitated by proteome and
transcriptome experimental data sets.
Most mass spectrometric techniques rely on databases that
contain annotated amino acid sequences of the proteins. The correct genome annotation is necessary for omics studies of clinically
relevant pathogens, e.g., Mycobacterium tuberculosis, as well as for
the progress in drug design and in silico biology. Proteogenomic
studies of M. tuberculosis mainly use the well-characterized reference strain H37Rv [1–3]. The genome of this strain was fully
sequenced at the Sanger Institute in 1998 [4] and shown to contain
3924 open-reading frames (ORF). However, a few years later, the
authors reported an increase in the number of ORFs to 3995
[5]. Several major proteogenomic studies of this species have
been performed using MS approaches in recent years, describing
20–40 novel proteins per study and correcting the current genome
annotations [2, 6]. At the same time, the studies of other
M. tuberculosis strains remain scarce [7].
Mo ´ nica Carrera and Jesu ´ s Mateos (eds.), Shotgun Proteomics: Methods and Protocols, Methods in Molecular Biology, vol. 2259,
https://doi.org/10.1007/978-1-0716-1178-4_12, © Springer Science+Business Media, LLC, part of Springer Nature 2021
191
Proteogenomic Approach for Mycobacterium tuberculosis
Investigation
Julia Bespyatykh, Georgij Arapidi, and Egor Shitikov
Abstract
Recent advances in MS/MS technology have made it possible to use proteomic data to predict proteincoding sequences. This approach is called proteogenomics, and it allows to correctly translate start and stop
sites and to reveal new open reading frames. Here, we focus on using proteogenomics to improve the
annotation of Mycobacteriumtuberculosis strains. We also describe detail procedures of the extraction of
proteins and their further preparation for LC–MS/MS analysis and outline the main steps of data analysis.
Key words Proteome, Beijing B0/W148, Proteomic, Label-free, Genome
1 Introduction
Nowadays, the starting point in the study of any living organism is
sequencing and annotation of its genome. The correct deciphering
of genomic information may be further facilitated by proteome and
transcriptome experimental data sets.
Most mass spectrometric techniques rely on databases that
contain annotated amino acid sequences of the proteins. The correct genome annotation is necessary for omics studies of clinically
relevant pathogens, e.g., Mycobacterium tuberculosis, as well as for
the progress in drug design and in silico biology. Proteogenomic
studies of M. tuberculosis mainly use the well-characterized reference strain H37Rv [1–3]. The genome of this strain was fully
sequenced at the Sanger Institute in 1998 [4] and shown to contain
3924 open-reading frames (ORF). However, a few years later, the
authors reported an increase in the number of ORFs to 3995
[5]. Several major proteogenomic studies of this species have
been performed using MS approaches in recent years, describing
20–40 novel proteins per study and correcting the current genome
annotations [2, 6]. At the same time, the studies of other
M. tuberculosis strains remain scarce [7].
Mo ´ nica Carrera and Jesu ´ s Mateos (eds.), Shotgun Proteomics: Methods and Protocols, Methods in Molecular Biology, vol. 2259,
https://doi.org/10.1007/978-1-0716-1178-4_12, © Springer Science+Business Media, LLC, part of Springer Nature 2021
191
