A comparative analysis has shown that annotations based on
genomic data are often incomplete and contain some errors
[8]. For example, proteomic studies have greatly improved the
genomic annotation of M. tuberculosis H37Rv strain by presenting
experimental evidence for several genes that had not been previously annotated or genes for which the transcription initiation sites
had been incorrectly identified, as well as by simply confirming the
existing ORFs [9]. Of note, the H37Rv strain is traditionally used
as a reference in omics studies of any strains and lineages of
M. tuberculosis. This is not always correct because the genome of
most closely related strains should be used for the proper identification of proteins in proteomic data [10].
In our work, we focused on the Beijing B0/W148 cluster, also
known as a “successful” clone that possesses distinctive pathogenic
properties and has a unique genomic organization. Only one complete genome of a representative of this cluster, W-148, was available in the NCBI database (GenBank accession number
CP012090.1) by the year 2018. Its annotation contained a considerable number of mistakes and discrepancies, exemplified by a
considerable number of pseudogenes. In our study, we have conducted a label-free LC–MS/MS shotgun proteome analysis of
56 Beijing B0/W148 strains using a six-frame translation of the
W148 genome [11]. Noteworthy, a low genetic diversity within the
cluster in the absence of horizontal gene transfer allowed to use
combine proteomic data for several strains. The obtained data
reemphasized the need for proteomic data to complement genome
annotations. In the same way, we also used a proteogenomic
approach to improve the annotation of the RUS_B0 strain of
M. tuberculosisBeijing B0/W148 cluster [12, 13]. These results
allowed us to get the most complete annotation of this strain
(GenBank accession number CP030093.1), which is widespread
in the Russian Federation.
2 Materials
2.1 Purification
of Mycobacterial Cells
1. Microcentrifuge/vortex.
2. Water bath.
3. Microbiological loop 1 μL, sterile.
4. Microbiological loop 10 μL, sterile.
5. Tubes (e.g. Falcon) 15 mL.
6. Centrifuge for 14,000 rpm (20,800 Â g), refrigerated.
7. Variable volume pipette 1,000–5,000 μL.
8. Variable volume pipette 200–1,000 μL.
9. Tips 200–1,000 μL.
192
Julia Bespyatykh et al.
genomic data are often incomplete and contain some errors
[8]. For example, proteomic studies have greatly improved the
genomic annotation of M. tuberculosis H37Rv strain by presenting
experimental evidence for several genes that had not been previously annotated or genes for which the transcription initiation sites
had been incorrectly identified, as well as by simply confirming the
existing ORFs [9]. Of note, the H37Rv strain is traditionally used
as a reference in omics studies of any strains and lineages of
M. tuberculosis. This is not always correct because the genome of
most closely related strains should be used for the proper identification of proteins in proteomic data [10].
In our work, we focused on the Beijing B0/W148 cluster, also
known as a “successful” clone that possesses distinctive pathogenic
properties and has a unique genomic organization. Only one complete genome of a representative of this cluster, W-148, was available in the NCBI database (GenBank accession number
CP012090.1) by the year 2018. Its annotation contained a considerable number of mistakes and discrepancies, exemplified by a
considerable number of pseudogenes. In our study, we have conducted a label-free LC–MS/MS shotgun proteome analysis of
56 Beijing B0/W148 strains using a six-frame translation of the
W148 genome [11]. Noteworthy, a low genetic diversity within the
cluster in the absence of horizontal gene transfer allowed to use
combine proteomic data for several strains. The obtained data
reemphasized the need for proteomic data to complement genome
annotations. In the same way, we also used a proteogenomic
approach to improve the annotation of the RUS_B0 strain of
M. tuberculosisBeijing B0/W148 cluster [12, 13]. These results
allowed us to get the most complete annotation of this strain
(GenBank accession number CP030093.1), which is widespread
in the Russian Federation.
2 Materials
2.1 Purification
of Mycobacterial Cells
1. Microcentrifuge/vortex.
2. Water bath.
3. Microbiological loop 1 μL, sterile.
4. Microbiological loop 10 μL, sterile.
5. Tubes (e.g. Falcon) 15 mL.
6. Centrifuge for 14,000 rpm (20,800 Â g), refrigerated.
7. Variable volume pipette 1,000–5,000 μL.
8. Variable volume pipette 200–1,000 μL.
9. Tips 200–1,000 μL.
192
Julia Bespyatykh et al.
