Strategy 4. Generate 3D conformers
Drugs act through physical interaction with specific biological targets. Such
interactions are determined primarily by complementarity of shape and properties
between the interacting molecules. On that account, the biological activity of a drug
depends on its three-dimensional structure. In solution, most drug molecules are
flexible and exist as an ensemble of low-energy conformations (shapes) in equilibrium with one another. The biologically active conformation (target-bound) can
either be similar to the conformations present in solution or can be induced by target
binding [80]. Also, different target proteins or cellular environments can induce
different conformations of the same ligand. Thus, conformational adaptation is an
important aspect in pharmacophore modeling, rigid docking, shape-based screening, 3D-QSAR, and virtual screening and must be considered carefully during lead
optimization. There are various small-molecule conformer generation tools available such as BALLOON, CONFAB, FROG2, RDKIT, and OMEGA. These tools
use diverse algorithms which may be based on approaches including
knowledge-based rule sets [81, 82], random coordinate changes [83], random torsional angle changes [84, 85], and/or distance geometry [86, 87]. Numerous studies
have reported the use of conformer generating tools for predicting bioactive conformations [88, 89]. In view of this, we generated a multi-conformer database of
5447 compounds using Universal Force Field (UFF) to facilitate epigenetic ligands
with maximum coverage of conformational space [46]. Conformers were generated
using RDKit, an open-source toolkit that comes under the permissive Berkeley
Software Distribution (BSD) license and employs the distance geometry approach
[90]. For each compound, we generated 50 conformers with RMSD cutoff of 0.5.
For some compounds, the number of conformers generated was less than 50 as a
result of RMSD criteria. In total, 269,052 conformers were generated which provide
the structural information about the conformational states of all epigenetic ligands
in the database and can be exploited further for in silico drug designing.
Strategy 5. Perform clustering analysis
Clustering is a powerful tool to identify homogeneous subsets within a heterogeneous compound dataset using structural characteristics [76, 91, 92]. For
example, it can be utilized to explore a large dataset to correlate compounds based
on their biological activity/property and scaffold hopping [93]. Clustering tools are
applied widely in drug discovery for chemical diversity, compound selection, and
data reduction in libraries. In times, it is advantageous to cluster small subsets of
compounds together to perform assays where it is not feasible to perform the
high-throughput screening. An ideal clustering process creates a series of clusters
from a larger library of compounds. Each cluster consists of compounds with
similar datapoints clubbed together as per the chosen criteria for similarity [94].
JKlustor, ChemMine Tools, and PKOM are some examples of the available clustering tools. In EpiDBase, ChemMine Tools were utilized to perform clustering
analysis. ChemMine has an online workbench that provides three important clustering methods including hierarchical clustering, multidimensional scaling (MDS),
Integrated Chemoinformatics Approaches …
259
Drugs act through physical interaction with specific biological targets. Such
interactions are determined primarily by complementarity of shape and properties
between the interacting molecules. On that account, the biological activity of a drug
depends on its three-dimensional structure. In solution, most drug molecules are
flexible and exist as an ensemble of low-energy conformations (shapes) in equilibrium with one another. The biologically active conformation (target-bound) can
either be similar to the conformations present in solution or can be induced by target
binding [80]. Also, different target proteins or cellular environments can induce
different conformations of the same ligand. Thus, conformational adaptation is an
important aspect in pharmacophore modeling, rigid docking, shape-based screening, 3D-QSAR, and virtual screening and must be considered carefully during lead
optimization. There are various small-molecule conformer generation tools available such as BALLOON, CONFAB, FROG2, RDKIT, and OMEGA. These tools
use diverse algorithms which may be based on approaches including
knowledge-based rule sets [81, 82], random coordinate changes [83], random torsional angle changes [84, 85], and/or distance geometry [86, 87]. Numerous studies
have reported the use of conformer generating tools for predicting bioactive conformations [88, 89]. In view of this, we generated a multi-conformer database of
5447 compounds using Universal Force Field (UFF) to facilitate epigenetic ligands
with maximum coverage of conformational space [46]. Conformers were generated
using RDKit, an open-source toolkit that comes under the permissive Berkeley
Software Distribution (BSD) license and employs the distance geometry approach
[90]. For each compound, we generated 50 conformers with RMSD cutoff of 0.5.
For some compounds, the number of conformers generated was less than 50 as a
result of RMSD criteria. In total, 269,052 conformers were generated which provide
the structural information about the conformational states of all epigenetic ligands
in the database and can be exploited further for in silico drug designing.
Strategy 5. Perform clustering analysis
Clustering is a powerful tool to identify homogeneous subsets within a heterogeneous compound dataset using structural characteristics [76, 91, 92]. For
example, it can be utilized to explore a large dataset to correlate compounds based
on their biological activity/property and scaffold hopping [93]. Clustering tools are
applied widely in drug discovery for chemical diversity, compound selection, and
data reduction in libraries. In times, it is advantageous to cluster small subsets of
compounds together to perform assays where it is not feasible to perform the
high-throughput screening. An ideal clustering process creates a series of clusters
from a larger library of compounds. Each cluster consists of compounds with
similar datapoints clubbed together as per the chosen criteria for similarity [94].
JKlustor, ChemMine Tools, and PKOM are some examples of the available clustering tools. In EpiDBase, ChemMine Tools were utilized to perform clustering
analysis. ChemMine has an online workbench that provides three important clustering methods including hierarchical clustering, multidimensional scaling (MDS),
Integrated Chemoinformatics Approaches …
259
