fundamental set from {C 1 , C 2 } to {C 1 , C 3 } while the set of all cycles will still
remain {C 1 , C 2 , C 3 }. This, arguably, is just an instance of change of basis in the
cycle vector space.
It will now suffice to check the sizes of the cycles in the fundamental set against
the required sizes and keep or discard the generated structure accordingly. This
decision made, considering the fundamental set only, is in accordance with the
IUPAC convention of the number of rings in polycyclic systems [26] where the
number of rings is equal to the minimum number of scissions required to convert
the system into an open chain compound or structure. Following this convention of
ring count, the example corresponding to Fig. 4 will be a valid structure against the
cycle size restriction either being 5 or 6.
(c) Removal of duplicate cyclic structures using graph canonicalization:
Although the trees generated by the algorithm given by Beyer et al. [16] are
non-isomorphic (hence distinct structures), it is easy to comprehend that introduction of edges may lead to generating more than one chemical structure of same
topology. As the entire process starts with tree structure, consider the case of the
rightmost tree representation shown in Fig. 2, and two different edge introductions
for a given cycle size constraint of 6 and cycle count constraint of 1 as shown in
Fig. 5.
Although the presented example is basic in nature, the problem aggravates when
the number of nodes is fairly large and such node pairs lie in different branches,
sometimes far apart. For example, the molecules with 30 or more non-hydrogen
atoms are fairly common in organic compounds developed as pharmaceutical
entities. Moreover, even when the graph topology is uniquely fixed, the combinatorial imposition of node colours for imparting heterogeneity by introducing
different atoms and the imposition of multiplicity of bonds can again lead to
duplicate structures. Hence, any duplicate elimination strategy should consider the
complete graph along with heterogeneity and bond multiplicity.
In the above context, molecular graph canonicalization algorithms can be used to
identify the duplicate structures and eliminate them during generation. As we intend
to store the molecules in SMILES notation format, it has been decided to use the
algorithm proposed for generation of unique SMILES by Weininger et al. [27],
which tackles the molecular graph canonicalization by extended connectivity
through an unambiguous function using product of primes.
Fig. 5 Duplicate cyclic structures
Combinatorial Drug Discovery from Activity-Related Substructure …
85
Précédent

- 96/413

Suivant