and binning clustering. ChemMine tool first calculates the similarity matrix based
on atom pair descriptors for each compound using the Tanimoto coefficient. The
Tanimoto coefficient may range between 0 and 1, indicating ‘1’ as the highest
similarity and ‘0’ as no similarity. The similarity matrix is then converted into a
distance matrix by deducting the similarity values from 1. The hierarchical clustering arranges the similar compounds in a tree with branch representation, where
branch lengths are proportional to the similarity between compounds. However,
MDS represents compounds as a scatter plot. The binning cluster displays the
results as a table where similar compounds are grouped together for a user-definable
cutoff. In EpiDBase, binning clustering was performed using a similarity cutoff of
0.6 (Tanimoto coefficient) to define the chemical diversity and space coverage of
epigenetic ligands. For example, 622 ligands of SIRT2 were clustered using
ChemMine to classify similar compounds and scaffolds into groups. Out of 108
clusters formed, 65 were single molecule clusters represented by unique ligands.
Such clusters provide a rare opportunity for medicinal chemists to populate them
using rational design and structure-activity relationship (SAR) studies to identify
potential therapeutics.
Strategy 6. Optimization of fragment-based library
Fragment-based drug design has emerged as a powerful technique in lead discovery
paradigm which is used as an alternative or often complementary to traditional
high-throughput screening (HTS). Fragments are ‘atom-efficient’ binders that can
be further expanded into high-affinity lead compounds [95]. In comparison with
compounds, fragments exhibit weak affinity toward the target and need high-end
instruments for their detection. Furthermore, fragment screening needs a high
concentration of protein and fragments. These challenges can be overcome by using
a wide variety of computational approaches to identify potential fragments and
binding sites. In EpiDBase, Retrosynthetic Combinatorial Analysis Procedure
(RECAP) [96] was used to generate the fragment database for SIRT2 modulators.
RECAP cleaves along the bonds using chemical knowledge and generates a collection of fragments suitable for combinatorial library synthesis. Such fragments
can be further clustered to generate effective libraries for epigenetic drug discovery.
Strategy 7. Optimization of commercial library for HTS
High-throughput screening (HTS) is a robust approach to drug discovery that
allows the assaying of a large number of small molecules against a validated target.
The aim of HTS is to accelerate drug discovery by identifying active chemical
series. It is essential to enrich small-molecule libraries with high chemical diversity
to increase the hit rate. The compound enrichment can be done using a multitude of
techniques including scaffold tree classification, virtual docking or general structure
and property filters (Fig. 4) [97, 98]. In our earlier study [99], various commercially
available chemical libraries were analyzed for their exclusiveness, drug-likeness,
and scaffold similarities (using asymmetrical metrics). The study demonstrates that
260
S. Loharch et al.
on atom pair descriptors for each compound using the Tanimoto coefficient. The
Tanimoto coefficient may range between 0 and 1, indicating ‘1’ as the highest
similarity and ‘0’ as no similarity. The similarity matrix is then converted into a
distance matrix by deducting the similarity values from 1. The hierarchical clustering arranges the similar compounds in a tree with branch representation, where
branch lengths are proportional to the similarity between compounds. However,
MDS represents compounds as a scatter plot. The binning cluster displays the
results as a table where similar compounds are grouped together for a user-definable
cutoff. In EpiDBase, binning clustering was performed using a similarity cutoff of
0.6 (Tanimoto coefficient) to define the chemical diversity and space coverage of
epigenetic ligands. For example, 622 ligands of SIRT2 were clustered using
ChemMine to classify similar compounds and scaffolds into groups. Out of 108
clusters formed, 65 were single molecule clusters represented by unique ligands.
Such clusters provide a rare opportunity for medicinal chemists to populate them
using rational design and structure-activity relationship (SAR) studies to identify
potential therapeutics.
Strategy 6. Optimization of fragment-based library
Fragment-based drug design has emerged as a powerful technique in lead discovery
paradigm which is used as an alternative or often complementary to traditional
high-throughput screening (HTS). Fragments are ‘atom-efficient’ binders that can
be further expanded into high-affinity lead compounds [95]. In comparison with
compounds, fragments exhibit weak affinity toward the target and need high-end
instruments for their detection. Furthermore, fragment screening needs a high
concentration of protein and fragments. These challenges can be overcome by using
a wide variety of computational approaches to identify potential fragments and
binding sites. In EpiDBase, Retrosynthetic Combinatorial Analysis Procedure
(RECAP) [96] was used to generate the fragment database for SIRT2 modulators.
RECAP cleaves along the bonds using chemical knowledge and generates a collection of fragments suitable for combinatorial library synthesis. Such fragments
can be further clustered to generate effective libraries for epigenetic drug discovery.
Strategy 7. Optimization of commercial library for HTS
High-throughput screening (HTS) is a robust approach to drug discovery that
allows the assaying of a large number of small molecules against a validated target.
The aim of HTS is to accelerate drug discovery by identifying active chemical
series. It is essential to enrich small-molecule libraries with high chemical diversity
to increase the hit rate. The compound enrichment can be done using a multitude of
techniques including scaffold tree classification, virtual docking or general structure
and property filters (Fig. 4) [97, 98]. In our earlier study [99], various commercially
available chemical libraries were analyzed for their exclusiveness, drug-likeness,
and scaffold similarities (using asymmetrical metrics). The study demonstrates that
260
S. Loharch et al.
