carry out the same reaction. Hence, we can assume that such mutations are seldom
seen, and if at all they occur, a compensatory mutation(s) will be seen to counter the
detrimental effects of those mutations [11, 12]. Another strategy adopted by
organisms is to increase the production of drug-metabolizing enzymes that modify
the drugs to their inactive form eventually leading to their elimination. A classic
example of this is the inactivation of penicillin by the enzyme b-lactamase.
1.2 Overview of Computational Methods to Study Drug
Resistance
Broadly, computer-assisted methods used to study drug resistance can be classified
into two categories based on the information they require and the output they return.
The first category of methods requires only 1D sequence data as input and the
output is generally a classification type, i.e. the test sequence is classified as a
resistant or a non-resistant sequence. Thus, the methods grouped under this class are
collectively called as “sequence-based” methods [13]. The workflow of these
methods is akin to machine learning or QSAR type classification methods. In a
nutshell, sequence-based methods require sequences with the corresponding biological activity data (K i or IC 50 or any other suitable numerical value) for the drug
under study. Such data can be curated from databases like HIVDB (for HIV
resistance, curated and maintained by Stanford University; [14, 15]) CancerDR (for
cancer resistance, curated by CSIR Institute of Microbial Technology and OSDD,
India; [16]), tuberculosis resistance mutation database (curated and maintained by
various departments and schools with Harvard University; [17], and many other
such databases. The data is then split into training and test sets to develop and
validate the predictive models. The advantage of such methods is that it is not
necessary to know the tertiary structure of the protein or the drug-receptor interactions. Therefore, sequence-based methods are computationally inexpensive and
large amount of data can be trained to obtain decent quality predictive models in a
short time. However, they suffer from two major drawbacks; (1) a lot of a priori
information on drug-resistant mutations is needed to train/develop predictive
models and (2) no mechanistic insights or atomistic details can be obtained.
The drawbacks seen in the sequence-based methods are efficiently overcome by
structure-based methods [13, 18, 19]. Further, structure-based methods are the
methods of choice when atomistic details are desired. However, these additional
details come at an added computational cost and require high-resolution protein
structures to be able to make accurate and reliable predictions. However, unlike the
sequence-based methods, they do not require large a priori information on mutations; on the contrary, they can be applied to systems where no data on mutation is
available. To assess the binding stability which is the basis for predictions, these
methods employ either empirical scoring functions that implicitly try to reflect the
free energy of binding or use techniques that compute the free energy of binding
Free Energy-Based Methods to Understand Drug Resistance Mutations
3
seen, and if at all they occur, a compensatory mutation(s) will be seen to counter the
detrimental effects of those mutations [11, 12]. Another strategy adopted by
organisms is to increase the production of drug-metabolizing enzymes that modify
the drugs to their inactive form eventually leading to their elimination. A classic
example of this is the inactivation of penicillin by the enzyme b-lactamase.
1.2 Overview of Computational Methods to Study Drug
Resistance
Broadly, computer-assisted methods used to study drug resistance can be classified
into two categories based on the information they require and the output they return.
The first category of methods requires only 1D sequence data as input and the
output is generally a classification type, i.e. the test sequence is classified as a
resistant or a non-resistant sequence. Thus, the methods grouped under this class are
collectively called as “sequence-based” methods [13]. The workflow of these
methods is akin to machine learning or QSAR type classification methods. In a
nutshell, sequence-based methods require sequences with the corresponding biological activity data (K i or IC 50 or any other suitable numerical value) for the drug
under study. Such data can be curated from databases like HIVDB (for HIV
resistance, curated and maintained by Stanford University; [14, 15]) CancerDR (for
cancer resistance, curated by CSIR Institute of Microbial Technology and OSDD,
India; [16]), tuberculosis resistance mutation database (curated and maintained by
various departments and schools with Harvard University; [17], and many other
such databases. The data is then split into training and test sets to develop and
validate the predictive models. The advantage of such methods is that it is not
necessary to know the tertiary structure of the protein or the drug-receptor interactions. Therefore, sequence-based methods are computationally inexpensive and
large amount of data can be trained to obtain decent quality predictive models in a
short time. However, they suffer from two major drawbacks; (1) a lot of a priori
information on drug-resistant mutations is needed to train/develop predictive
models and (2) no mechanistic insights or atomistic details can be obtained.
The drawbacks seen in the sequence-based methods are efficiently overcome by
structure-based methods [13, 18, 19]. Further, structure-based methods are the
methods of choice when atomistic details are desired. However, these additional
details come at an added computational cost and require high-resolution protein
structures to be able to make accurate and reliable predictions. However, unlike the
sequence-based methods, they do not require large a priori information on mutations; on the contrary, they can be applied to systems where no data on mutation is
available. To assess the binding stability which is the basis for predictions, these
methods employ either empirical scoring functions that implicitly try to reflect the
free energy of binding or use techniques that compute the free energy of binding
Free Energy-Based Methods to Understand Drug Resistance Mutations
3
