162
information from siderophores whose substrates were experimentally determined or
predicted from the domain architecture. Novel NRPSs are aligned against the
database to predict their substrates. The analysis focuses mainly on the Stachelhaus
code or the 10 residues in the active site of A domain that were identified to be specific for each substrate (Stachelhaus et al. 1999). For bacterial domain specificity
prediction, this approach was found considerably efficient and was used by NRPSPKS (Ansari et al. 2004). Other tools such as NRPSpredictor2 (Röttig et al. 2011)
and NRPSsp (Prieto et al. 2011) were also developed based on similar principles but
using machine learning techniques like support vector machines (SVM) and profile
HMMs. Another prediction tool SEQL-NRPS (Knudsen et al. 2015) also claims to
be able to decipher substrate specificity using machine learning but does not necessarily rely on active sites.
All the above concepts were found to perform well on bacterial NRPS prediction
and have been used to predict novel siderophores from bacterial genomes
(Etchegaray et al. 2004; Komaki et al. 2018; Aleti et al. 2015) however, their efficiency to predict substrate specificities in the case of fungal NRPSs is still debatable
(Sørensen et al. 2014). The lack of characterized data of fungal origin for training is
a major disadvantage to develop machine learning–based tools for the said purpose.
It is important to obtain enough data primarily of fungal NRPS and more specifically of characterized fungal siderophores to train prediction models for fungal siderophore genes. To overcome the issue of lack of prior specificity data in fungi, a
virtual screening approach was tested (Verne Lee et al. 2015) and found to be promising. Nevertheless for the approach to be efficiently exploited, the homology model
build and assessment need to be enhanced.
Unlike the previously discussed tools that were developed based on domain
homology, a tool called CASSIS predicts secondary metabolite gene clusters by
taking into account co-regulation of cluster genes. However, it is not a genomic
approach by itself but instead takes as input a given anchor/backbone gene under
consideration. To facilitate the identification of such anchor genes from the genome,
the SMIPS interface was developed along with CASSIS. Whole genome sequence
is processed by SMIPS and the results could be loaded to CASSIS for gene cluster
prediction. These servers are dedicated to mining eukaryotic genomes and hence are
found useful for fungal genome mining (Wolf et al. 2015).
Generally, fungal NRPS domains put forward certain difficulties in substrate prediction due to its non-linear and iterative architecture. In such cases, the ‘A’ domain
could not be mapped directly to a specific amino acid of the product as found in
bacteria (Mootz et  al. 2002). Additionally, post-translational modifications in the
cyclic or branched peptides also make it difficult to determine the sequence of
incorporation (Caboche et al. 2007). Ultimately, substrate specificity prediction of
fungal NRPS and particularly of siderophore-producing genes is still largely unexplored leaving scope for collection of sufficient experimental data, appropriate curation, analysis and training.
D. Subramanian et al.
Précédent

- 168/220

Suivant