164
capable of sequencing a peptide even from small amounts of unpurified raw sample
(Molinski 2010). MS-based approaches were used to investigate the secondary
metabolomes of several microbes including fungi such as Penicillium and Furcatum
(Krug and Müller 2014; Vansteelandt et al. 2012). The tandem mass spectroscopy
(MS/MS) uses two mass spectrometers each to separate and identify fragments of
compounds subject to investigation and is found very powerful to obtain characteristic spectra which can be compared to comprehensive libraries or databases to
facilitate novel compound identification (Tautenhahn et al. 2012; Wang et al. 2016).
Improved detection techniques are also being developed that includes spatial imaging to improve the sensitivity (Fang and Dorrestein 2014). MS-based proteomics
can also be used to validate predicted compounds and identify their interactions
(Smits and Vermeulen 2016) and modifications (Ribeiro et al. 2017).
With advances in computational technology, computational tools are available to
connect gene clusters to its product based on MS/MS data analysis. Proteomic
Investigation of Secondary Metabolism (PrISM) is a proteomics approach developed to detect NRPSs whose primary sequences were used to generate primers to
amplify corresponding gene clusters (Bumpus et al. 2009). For unsequenced
genomes, a strategy was developed that works based on grouping BGCs and MS/
MS spectrums into corresponding families by similarities (Nguyen et al. 2013).
Gene cluster families and molecular families are accordingly identified and correlated based on MS/MS techniques and networking by connecting them to gene cluster families of sequenced organisms. The original study was used in 60 unsequenced
but related strains of bacteria but is easily adaptable to study other different
organisms.
Addressing the issues present in the sequencing of non-linear peptides, the first
NRP identification algorithm known as NRPquest was developed in 2014 (Mohimani
et al. 2014). With a sequenced genome and MS dataset as inputs, the algorithm is
implemented in four steps: (a) genome mining using NRPSpredictor2 and construction of putative genomic-NRP database, (b) matching of spectral data to database,
(c) computation of statistical significance and ranking of peptide-spectrum matches
and (d) construction of spectral network to enhance identified natural product sets
and uncover related NRP families through dereplication. The same research group
later implemented an extended tool known as Dereplicator that aims at the analysis
of peptidic natural products including NRPs with modified algorithms using dereplication techniques (Mohimani et al. 2017). For a provided chemical database, the
algorithm works by generating in silico mass spectra of compounds via prediction
of their fragmentation during MS. The resultant spectra are then compared to experimental LC-MS/MS to identify similarities, based on which significant matches are
reported.
With an enormous amount of research in MS-based analysis of NPs, it is vital to
implement a platform for the sharing and management of data. Global Natural
Products Social (GNPS) Molecular Networking was thus introduced as an open
access knowledge base to allow researchers to share raw, processed or identified MS
data. Crowd-sourced analyses and curation are supposed to assist improved annotation of the datasets (Wang et al. 2016).
D. Subramanian et al.
capable of sequencing a peptide even from small amounts of unpurified raw sample
(Molinski 2010). MS-based approaches were used to investigate the secondary
metabolomes of several microbes including fungi such as Penicillium and Furcatum
(Krug and Müller 2014; Vansteelandt et al. 2012). The tandem mass spectroscopy
(MS/MS) uses two mass spectrometers each to separate and identify fragments of
compounds subject to investigation and is found very powerful to obtain characteristic spectra which can be compared to comprehensive libraries or databases to
facilitate novel compound identification (Tautenhahn et al. 2012; Wang et al. 2016).
Improved detection techniques are also being developed that includes spatial imaging to improve the sensitivity (Fang and Dorrestein 2014). MS-based proteomics
can also be used to validate predicted compounds and identify their interactions
(Smits and Vermeulen 2016) and modifications (Ribeiro et al. 2017).
With advances in computational technology, computational tools are available to
connect gene clusters to its product based on MS/MS data analysis. Proteomic
Investigation of Secondary Metabolism (PrISM) is a proteomics approach developed to detect NRPSs whose primary sequences were used to generate primers to
amplify corresponding gene clusters (Bumpus et al. 2009). For unsequenced
genomes, a strategy was developed that works based on grouping BGCs and MS/
MS spectrums into corresponding families by similarities (Nguyen et al. 2013).
Gene cluster families and molecular families are accordingly identified and correlated based on MS/MS techniques and networking by connecting them to gene cluster families of sequenced organisms. The original study was used in 60 unsequenced
but related strains of bacteria but is easily adaptable to study other different
organisms.
Addressing the issues present in the sequencing of non-linear peptides, the first
NRP identification algorithm known as NRPquest was developed in 2014 (Mohimani
et al. 2014). With a sequenced genome and MS dataset as inputs, the algorithm is
implemented in four steps: (a) genome mining using NRPSpredictor2 and construction of putative genomic-NRP database, (b) matching of spectral data to database,
(c) computation of statistical significance and ranking of peptide-spectrum matches
and (d) construction of spectral network to enhance identified natural product sets
and uncover related NRP families through dereplication. The same research group
later implemented an extended tool known as Dereplicator that aims at the analysis
of peptidic natural products including NRPs with modified algorithms using dereplication techniques (Mohimani et al. 2017). For a provided chemical database, the
algorithm works by generating in silico mass spectra of compounds via prediction
of their fragmentation during MS. The resultant spectra are then compared to experimental LC-MS/MS to identify similarities, based on which significant matches are
reported.
With an enormous amount of research in MS-based analysis of NPs, it is vital to
implement a platform for the sharing and management of data. Global Natural
Products Social (GNPS) Molecular Networking was thus introduced as an open
access knowledge base to allow researchers to share raw, processed or identified MS
data. Crowd-sourced analyses and curation are supposed to assist improved annotation of the datasets (Wang et al. 2016).
D. Subramanian et al.
