As an initial foray into extracting useful information from the LDMs using the
PCA transformation we looked at a series of carboxylic acids that extended the set
used by Matta et al. [20, 21] to evaluate the potential of LDMs to predict physical
properties (e.g. pK a ) of the acids. One of our first observations is that as long as the
pair-wise values for the LI (e.g. C1 to C1, O7:O7, etc.) and the DI (C1to O7,O7:H4,
etc.) are retained, then the ordering of the LDM does not affect the eigenvalues that
are produced from the LDM (Table 3.3).
Table 3.4 shows the largest six eigenvalues extracted using the PCA method. For
most of these acids the six largest eigenvalues account for >95 % of the variance in
the LDM. Almost all the unaccounted-for variance (especially in the larger molecules) is due to the hydrogen atoms.
As pointed out by Matta [20], a threshold of dissimilarity is assumed between
the members of a molecular set for the construction of a QSAR model. Similarity is
commonly quantified on the basis of amino acid sequence matching [76], mismatching of 2-dimentional chemical graphs [8], on 3-D molecular skeleton
superpositions [77], and on point-by-point comparison of electron density—pioneered by Carbó [78–81] or of the molecular electrostatic potential [82–94].
Through relationships (5) and (6) in Ref. [20], a similarity of ρ(r) necessarily leads
to the similarity of all other ground-state properties, and hence the most fundamental molecular comparisons are those effected at the electron density level.
One way to appreciate the similarity of a group of molecules would be to map
the molecules in n-dimensional abstract mathematical space and determine if the
eigenvalues resulting from the PCA of the LDM coincide with chemical intuition.
The visualization of such multi-dimensional data is not feasible beyond three
dimensions. Thus even for the 6-dimensional descriptor space that can be constructed from the principal components listed in Table 3.4, the reduction of
dimensionality is a must to visualize distance similarity relationships between
molecules.
Such dimensional reduction may be achieved by “Multidimensional Scaling
(MDS)” techniques. MDS projects the n-dimensional distance in a
lower-dimensional space (2- or 3-dimensions) under the constraint of maximizing
the retention of the structure of the inter-molecular distance matrix. The representation of the n-dimensional space is optimized in the lower-dimensional projection by minimizing what is known as “stress”. The smaller the stress the better
the projection up to the (generally unattainable) limit of zero which is a projection
that preserves the distance matrix completely. The quality of the projection can be
gleaned from what is known as a “Shepard diagram”.
The Shepard diagram is a scatter plot in which the dissimilarities between the
molecules, measured as the distances in the full n-dimensional space, are compared
with the corresponding distances in the projected 2-dimensional space. A large
spread is an indicator of a poor MDS projection and vice versa up to the
unreachable extreme where all points fall on one line which indicates a perfect
MDS projection.
Using the data in Table 3.4 and the MDS algorithm in the software package
XLSTAT™, the Shepard diagram displayed in Fig. 3.6 is obtained. One can glean
76
C.F. Matta et al.
Précédent

- 84/582

Suivant