376
P. M. Vassiliev et al.
Therefore, a representation of a chemical structure should be multilevel, and the
variables should reflect both the local and integral properties of the compound.
The generalized pattern of a class of compounds with a desired property is a set
of all of the compounds that showing this property described by the set of parameters that characterize this compound [94, 101, 110]. The cardinality of this generalized pattern goes to infinity because its elements include both synthesized (tested)
and nonsynthesized (untested) active compounds when the number of parameters
is not limited. The more compounds in the training set and the greater number of
contrast variables of varying degrees of complexity describing their structure, the
more adequate the model of the generalized pattern.
If we unite these concepts, three important consequences ensue:
1. the biological activity shown by a chemical compound is not necessarily related
to its interaction with a specific biological target;
2. the chemical compound is regarded as a whole. There are no “significant” or
“insignificant” fragments in its structure; likewise, there are no “significant” or
“insignificant” variables describing this structure; and
3. a parametric description of the model of a generalized pattern is context-independent from the method of data analysis because it is redundant; the model also
does not presuppose the involvement of any procedures for detection of “informative” variables.
Thus, the parameter space of the models of generalized pattern of a class of compounds with a desired property is of extremely great dimension; it is not divided
into “informative” and “noninformative” subspaces. With this method of representation, no information that determines the individual specifics of the chemical structures to be recognized is lost, and this retention of information allows an effective
extrapolation of the obtained QSAR regularities to the area of new or poorly studied
compounds with nontrivial specifics of action.
A multidescriptor, hierarchical, multilevel representation of the structure of
a chemical compound consists of description methods that vary in their complexity and physicochemical meaning; in other words, groups of parameters that are
simultaneously divided into several levels of chemical structure representation with
increasing complexity, where each subsequent level of greater complexity is generated by the preceding one. The redundancy of description expansion in the parameters provides for an increased cardinality and improves the resulting model of a
generalized pattern of a compound class with the desired properties. This concept
was implemented when developing QL [110], a specialized language that describes
the structure of chemical compounds using 11 groups of QL descriptors at four
levels of complexity, with each of the descriptors varying in their physicochemical
meaning. The higher-ranking descriptor is formed as a combination of the preceding rank and the elementary QL descriptors.
The mega-dimensional space is a nonlinear space with a variable curve of extra-large dimensionality [105]. As a consequence, it is strongly correlated while
it is neither orthogonal nor normed. Generally speaking, such a space is mixed
and discrete-continuous. Mega-dimensional spaces are, in fact, an “impression”
P. M. Vassiliev et al.
Therefore, a representation of a chemical structure should be multilevel, and the
variables should reflect both the local and integral properties of the compound.
The generalized pattern of a class of compounds with a desired property is a set
of all of the compounds that showing this property described by the set of parameters that characterize this compound [94, 101, 110]. The cardinality of this generalized pattern goes to infinity because its elements include both synthesized (tested)
and nonsynthesized (untested) active compounds when the number of parameters
is not limited. The more compounds in the training set and the greater number of
contrast variables of varying degrees of complexity describing their structure, the
more adequate the model of the generalized pattern.
If we unite these concepts, three important consequences ensue:
1. the biological activity shown by a chemical compound is not necessarily related
to its interaction with a specific biological target;
2. the chemical compound is regarded as a whole. There are no “significant” or
“insignificant” fragments in its structure; likewise, there are no “significant” or
“insignificant” variables describing this structure; and
3. a parametric description of the model of a generalized pattern is context-independent from the method of data analysis because it is redundant; the model also
does not presuppose the involvement of any procedures for detection of “informative” variables.
Thus, the parameter space of the models of generalized pattern of a class of compounds with a desired property is of extremely great dimension; it is not divided
into “informative” and “noninformative” subspaces. With this method of representation, no information that determines the individual specifics of the chemical structures to be recognized is lost, and this retention of information allows an effective
extrapolation of the obtained QSAR regularities to the area of new or poorly studied
compounds with nontrivial specifics of action.
A multidescriptor, hierarchical, multilevel representation of the structure of
a chemical compound consists of description methods that vary in their complexity and physicochemical meaning; in other words, groups of parameters that are
simultaneously divided into several levels of chemical structure representation with
increasing complexity, where each subsequent level of greater complexity is generated by the preceding one. The redundancy of description expansion in the parameters provides for an increased cardinality and improves the resulting model of a
generalized pattern of a compound class with the desired properties. This concept
was implemented when developing QL [110], a specialized language that describes
the structure of chemical compounds using 11 groups of QL descriptors at four
levels of complexity, with each of the descriptors varying in their physicochemical
meaning. The higher-ranking descriptor is formed as a combination of the preceding rank and the elementary QL descriptors.
The mega-dimensional space is a nonlinear space with a variable curve of extra-large dimensionality [105]. As a consequence, it is strongly correlated while
it is neither orthogonal nor normed. Generally speaking, such a space is mixed
and discrete-continuous. Mega-dimensional spaces are, in fact, an “impression”
