383
12 Consensus Drug Design Using IT Microcosm
A study of information redundancy of QL showed that a set of simple descriptors
of ranks 1–4 with an indication of their number is almost identical to the molecular
graph; therefore, only these 11 types of QL descriptors are used in IT Microcosm for
the calculation of decision rules and the prediction of activity.
The meaning of a QL-based description is easily derived from the symbols of
descriptors. For instance:
{–CH3 2 ○ ○}
This is a methyl group in a two-bond chain with arbitrary type bonds, with conjugation present or absent (“○”
stands for any elementary descriptor);
{–N < 5 –CH3 ○}
This is a tertiary amino group bonded to a methyl group
by a five-bond chain with arbitrary type bonds, with conjugation present or absent;
{–N = –1 CycAr05 …1} A secondary imino group included into a five-membered
aromatic ring.
The structure of an organic compound of medium complexity is usually described
by 50-1,000 types of QL descriptors.
For example, the QL description of the structure of Thiomedan, an antiepileptic
drug, includes 165 types of QL descriptors.
For a desired activity type, models of generalized patterns of active/inactive
compounds are constructed from a training set in the form of a substructural descriptor matrix, which is a table where the lines feature the symbols for unique QL
descriptors of 11 types of the first four ranks and the columns indicate the numbers
of compounds. Each table cell shows the number of QL descriptors of this type in
the structure of a certain compound. In the matrix, the descriptors are placed in the
order of increasing rank in lexicographic order. The activity of the compounds in the
training set is annotated in a separate file.
12.2.3 Prediction Methods
To obtain a spectrum of intermediate prediction estimates, IT Microcosm utilizes
four original classification methods that show a consistent performance in megadimensional spaces. The prediction estimates are binary variables; they can only
assume two values: A or N (in numerical expression, 1 or 0, correspondingly). Each
method yields 11 prediction estimates (according to the number of QL descriptor
12 Consensus Drug Design Using IT Microcosm
A study of information redundancy of QL showed that a set of simple descriptors
of ranks 1–4 with an indication of their number is almost identical to the molecular
graph; therefore, only these 11 types of QL descriptors are used in IT Microcosm for
the calculation of decision rules and the prediction of activity.
The meaning of a QL-based description is easily derived from the symbols of
descriptors. For instance:
{–CH3 2 ○ ○}
This is a methyl group in a two-bond chain with arbitrary type bonds, with conjugation present or absent (“○”
stands for any elementary descriptor);
{–N < 5 –CH3 ○}
This is a tertiary amino group bonded to a methyl group
by a five-bond chain with arbitrary type bonds, with conjugation present or absent;
{–N = –1 CycAr05 …1} A secondary imino group included into a five-membered
aromatic ring.
The structure of an organic compound of medium complexity is usually described
by 50-1,000 types of QL descriptors.
For example, the QL description of the structure of Thiomedan, an antiepileptic
drug, includes 165 types of QL descriptors.
For a desired activity type, models of generalized patterns of active/inactive
compounds are constructed from a training set in the form of a substructural descriptor matrix, which is a table where the lines feature the symbols for unique QL
descriptors of 11 types of the first four ranks and the columns indicate the numbers
of compounds. Each table cell shows the number of QL descriptors of this type in
the structure of a certain compound. In the matrix, the descriptors are placed in the
order of increasing rank in lexicographic order. The activity of the compounds in the
training set is annotated in a separate file.
12.2.3 Prediction Methods
To obtain a spectrum of intermediate prediction estimates, IT Microcosm utilizes
four original classification methods that show a consistent performance in megadimensional spaces. The prediction estimates are binary variables; they can only
assume two values: A or N (in numerical expression, 1 or 0, correspondingly). Each
method yields 11 prediction estimates (according to the number of QL descriptor
