drug-like molecules that could be synthesized is estimated to be around 10
60 , the
current experimental techniques cannot possibly screen all of these within reasonable time and expense. Computational methods such as docking calculations address
this to some extent; however, the accuracy of the scoring functions behind these
algorithms is still not good enough to efficiently narrow the search space that can be
explored by experiments. Recently, it has been shown that machine learning (ML)based scoring functions can predict binding affinities better than the classical scoring
functions that are primarily used in computer-aided drug design [64]. Wojcikowski
et al. recently reported a systematic study on the performance of ML-based scoring
functions and compared it to well-known established methods [65]. They proposed a
scoring function based on the random forest method (RF-Score-VS) that was
trained on about 15,000 active and 900,000 inactive molecules against about 100
different drug targets . However, the authors do indicate that use of better molecular
representations and descriptors will further increase the success of machine learning
scoring functions. In addition, Kinnings et al. also showed that the support vector
machines (SVMs) can be used to improve the performance of scoring functions.
They constructed two prediction models; one is a regression model to predict the
IC 50 values, and the second is a classification model that was shown to perform very
well across the entire data set [66]. Ragoza et al. have proposed a CNN-based model
for scoring functions that can be used in structure-based drug design [67]. The model
used the existing three-dimensional structures of protein–ligand complexes to train a
model that predicts the binding affinity corresponding to any protein and ligand. The
model was systematically trained by including a series of structural and binding
variations such as high affinity binders, low affinity binders, correct binding pose and
incorrect poses. They found that the scoring function obtained based on the CNN
algorithm performs significantly better than AutoDock Vina in terms of predicting
both binding poses and affinities. Recently, Dror and co-workers proposed a method
named Siamese Atomic Surfacelet Network (SASNet) that applies CNN to predict
protein–protein binding interfaces with high accuracy compared to the previously
available knowledge-based and ML-based methods [58]. Interestingly, the training
was done on a biased data where binding-driven conformational changes are not part
of the data set; however, the model was shown to perform very well suggesting that
the model has possibly learned the inherent structural and dynamic properties of
proteins in general. In addition to the traditional neural network-based algorithms,
there are other deep learning methods such as reinforcement learning that have been
found to be very effective in drug design. Reinforcement learning is based on two
neural networks, namely the generative and predictive neural networks. Popova et al.
have recently proposed, Reinforcement Learning for Structural Evolution
(ReLeaSE), a de novo method based on reinforcement learning [68]. Initially, the
generative and predictive networks are trained individually using one of the
supervised learning methods followed by training of both models together. This
allows for predicting new chemical structures with desired biological activities. They
have shown test cases by generating libraries of molecules with desired melting
points, hydrophobicity and biological activities. They propose that it is possible to
use a similar approach for optimization of multiple properties such as biological
Recent Advancements in Computing Reliable Binding Free Energies …
241
Précédent

- 251/413

Suivant