7 Machine Learning (ML)-Based Approaches for Drug
Discovery
Application of machine learning (ML) methods to problems in chemistry, biology,
materials, etc., has taken a huge leap during the last few years. Specifically, a number
of problems related to accurate intermolecular potentials, [53] drug design, [54]
protein–protein interaction, [55] viable retrosynthetic pathways, [56] stability of
solids, [57] potential energy surfaces [58], etc., are being addressed [59]. Advances
that are being made in this space in terms of tackling problems in a way that was not
thought about even few years are rapid, and the number of papers that are being
published in this area is increasing exponentially. Unlike in the most research areas
of science and technology, traditional ML methods such as single-layer neural
networks or random forest have been applied in the area of computer-aided drug
design long time ago. However, modern deep learning methods within ML are
expected to make significant contributions to the area of drug design in the coming
days [54, 60, 61]. Given that the last fifteen years have witnessed prolific generation
of experimental data in terms of synthesizable compounds, their pharmacodynamic
and pharmacokinetic properties, application of data-driven methods is likely to
advance the field significantly. The following sections give a brief account of ML
and some of the recent successes in application of ML methods in areas relevant to
various drug design projects, including off-targets [62].
There are two fundamentally different methods in ML: supervised and unsupervised learning. Given a large data of inputs and outputs, supervised learning
methods try to learn a function so that given a set of inputs, output may be predicted. Supervised learning methods such as the artificial neural networks (ANNs)
are pertinent in quite a few drug discovery applications. On the other hand,
unsupervised learning methods learn structure within the data when only inputs are
available, which are typically applied dimensional reduction, pattern recognition,
etc. Most of the ML methods are based on ANNs that connect the input and the
output layers via an interconnected neural network (hidden layer(s)). The ANNs
consist of a number of layers with each containing a number of neurons. Output of
one of the layers is taken as input of the next layer, and the output values are
calculated using an activation function. Fully connected deep neural network
(DNN), recurrent neural network (RNN), convolutional neural network (CNN) and
autoencoders are some of the variants of ANNs that are very successful as efficient
methods for statistical modelling in a variety of fields. For a detailed account of
different machine learning methods relevant to drug discovery, the readers may
refer to the review by Lavecchia [63].
7.1 Structure-Based ML Approaches in Drug Design
As explained in previous sections, molecular recognition is a fundamental phenomenon behind all biological processes and in drug binding. While the number of
240
N. A. Murugan et al.
Précédent

- 250/413

Suivant