7 A Review of Feature Reduction Methods …
121
are developed on the assumption of a similarity principle. That is, molecules with
similar structures (and descriptors, consequently) will have similar biological activity
[4, 5]. A SAR model to predict toxicity (T ) is given in Eq. (1)
T = g
D f
(1)
where
D f
represents the feature space of molecular descriptors as chemical properties and g is a function that relates T to
D f
[2]. The accuracy of the model or
function g has been shown to depend on the most representative set of molecular
descriptors that will encode the useful properties of the molecules for prediction.
Molecular descriptors, being numerical features extracted from molecular structures, are the most common variables used for SAR-based toxicity prediction modeling [6]. The information encoded by descriptors depends on the molecular representation or “dimensionality” of the compound as well as the algorithm used to
calculate the descriptors [7]. One-dimensional (1D) descriptors are scalars encoding
physiochemical properties (molecular weight, logP) and constitutional parameters,
such as number of atoms, bond count, atom type, ring count, and fragment counts.
1D descriptors are insensitive to the topology of the molecule and tend to be similar
for distinct compounds. As a result, they are often used in combination with other
descriptors. Two-dimensional (2D) descriptors are more frequently used for chemical space description. 2D descriptors, including topological indices and structural
fragments, are calculated from the connection table (chemical graph) representation
of a molecule. They are not only independent of the conformation of the molecule
but also graph invariant (not sensitive to altering the number of graph nodes). Threedimensional (3D) descriptors provide a more complete characterization of molecular structures. 3D descriptors require conformational searching and can discriminate between isomers; this comes at the price of being computationally expensive.
The ability to discriminate between isomers can translate to less redundant features.
Examples of 3D descriptors include geometric, electrostatic, quantum chemical, and
WHIM & GETAWAY. Four-dimensional (4D) descriptors are much like 3D descriptors that evaluate multiple structural conformations simultaneously. Fingerprints are
another form of molecular descriptors [7–9]. Commonly used fingerprints include
the Molecular ACCess System (MACCS) [10] substructure fingerprints, PubChem
[11], and extended-connectivity fingerprints (ECFP) [12]. These fingerprints and 2D
descriptors were widely used in the Tox21 data challenge [13] where the winning
submissions used over 2500 predefined features covering a wide range of data from
topological and physical properties to fingerprints [14].
As shown above, the chemical structures used in SAR modeling are characterized
by many molecular descriptors. It is common to generate thousands of descriptors
for a single molecule [14]. It is well known that the accuracy of predictive models
is not positively correlated to the dimensionality of the data, as overfitting tends to
become an issue [15–17]. High-dimensional spaces are prone to include irrelevant
and noisy features [18]. SARs developed using such features tend to focus on the
peculiarities of molecules and fail to be generalizable [19]. In the chemical space
121
are developed on the assumption of a similarity principle. That is, molecules with
similar structures (and descriptors, consequently) will have similar biological activity
[4, 5]. A SAR model to predict toxicity (T ) is given in Eq. (1)
T = g
D f
(1)
where
D f
represents the feature space of molecular descriptors as chemical properties and g is a function that relates T to
D f
[2]. The accuracy of the model or
function g has been shown to depend on the most representative set of molecular
descriptors that will encode the useful properties of the molecules for prediction.
Molecular descriptors, being numerical features extracted from molecular structures, are the most common variables used for SAR-based toxicity prediction modeling [6]. The information encoded by descriptors depends on the molecular representation or “dimensionality” of the compound as well as the algorithm used to
calculate the descriptors [7]. One-dimensional (1D) descriptors are scalars encoding
physiochemical properties (molecular weight, logP) and constitutional parameters,
such as number of atoms, bond count, atom type, ring count, and fragment counts.
1D descriptors are insensitive to the topology of the molecule and tend to be similar
for distinct compounds. As a result, they are often used in combination with other
descriptors. Two-dimensional (2D) descriptors are more frequently used for chemical space description. 2D descriptors, including topological indices and structural
fragments, are calculated from the connection table (chemical graph) representation
of a molecule. They are not only independent of the conformation of the molecule
but also graph invariant (not sensitive to altering the number of graph nodes). Threedimensional (3D) descriptors provide a more complete characterization of molecular structures. 3D descriptors require conformational searching and can discriminate between isomers; this comes at the price of being computationally expensive.
The ability to discriminate between isomers can translate to less redundant features.
Examples of 3D descriptors include geometric, electrostatic, quantum chemical, and
WHIM & GETAWAY. Four-dimensional (4D) descriptors are much like 3D descriptors that evaluate multiple structural conformations simultaneously. Fingerprints are
another form of molecular descriptors [7–9]. Commonly used fingerprints include
the Molecular ACCess System (MACCS) [10] substructure fingerprints, PubChem
[11], and extended-connectivity fingerprints (ECFP) [12]. These fingerprints and 2D
descriptors were widely used in the Tox21 data challenge [13] where the winning
submissions used over 2500 predefined features covering a wide range of data from
topological and physical properties to fingerprints [14].
As shown above, the chemical structures used in SAR modeling are characterized
by many molecular descriptors. It is common to generate thousands of descriptors
for a single molecule [14]. It is well known that the accuracy of predictive models
is not positively correlated to the dimensionality of the data, as overfitting tends to
become an issue [15–17]. High-dimensional spaces are prone to include irrelevant
and noisy features [18]. SARs developed using such features tend to focus on the
peculiarities of molecules and fail to be generalizable [19]. In the chemical space
