2 Background, Tasks, Modeling Methods …
27
of the chemicals. It deserves mentioning that it is correlation rather than causality
between the features and targeted properties that QSARs eventually obtain. To learn
more about the history and/or future perspectives of QSAR, some literature reviews
can be consulted [25, 73, 74].
From a mathematical point of view, a QSAR model has three basic elements: the
features (X), the properties to be predicted (Y), and the algorithm adopted to map
the features onto the properties (f ). In brief, a QSAR model can always be expressed
as Y = f (X). As for X, if features are extracted only from chemical molecules
themselves, then these features are typically called molecular (structural) descriptors. Details about molecular descriptors can be found in Todeschini and Consonni’s
Molecular Descriptor for Chemoinformatics [75]. Besides, the application domain
beyond which QSAR models could not be reliable also depends strongly on the feature space of chemicals in the training set. As for Y, if the properties have discrete
categories/levels, then it refers to a classification task, resulting in a predictive model
termed as classifier. If the properties have continuous values it refers to a regression
task, resulting in a regressor. The quality of QSAR models heavily relies on the quality of the input data sets. Therefore, collection and curation of data need particular
patience and caution. As for f, nowadays, a large number of algorithms for classification or regression are available, e.g., multivariate linear regression, partial least
regression, support vector machine, decision tree, naïve Bayes, artificial neural network, etc., as well as ensemble algorithms such as random forest [76], among which
computational toxicologists may find some suitable for their specialized cases.
It is also notable that molecular models could inspire novel descriptors that can
better characterize the underlying pattern for certain toxicological phenomena. For
example, simulation of the interaction between halogenated compounds and human
transthyretin protein with QM/MM methods indicated specific descriptors for QSAR
modeling [77]. Similarly, QSARs can translate the time-consuming computational
chemistry models into empirical rules and mathematics, thus resulting in more efficient predictive models. For example, Rydberg and Olsen et al. performed a series
of QM/DFT simulations on the active sites of P450 enzymes with various types of
small molecules [78–80], based on which a web server named SMARTCyp has been
developed that is capable of rapidly predicting sites of metabolism and associated
reaction energy barriers [81].
Currently, the rapid development of modern machine learning algorithms could
most likely promote a renaissance of the field of QSAR modeling. Previously unnoticed details might also be recaptured by innovative feature extraction and wellestablished machine learning algorithms, making scientists’ conclusions or predictions less arbitrary and more robust.
2.4 Challenges for Computational Toxicology
Although modeling frameworks seem to have been nicely established, challenges in
computational toxicology still remain for many aspects.
27
of the chemicals. It deserves mentioning that it is correlation rather than causality
between the features and targeted properties that QSARs eventually obtain. To learn
more about the history and/or future perspectives of QSAR, some literature reviews
can be consulted [25, 73, 74].
From a mathematical point of view, a QSAR model has three basic elements: the
features (X), the properties to be predicted (Y), and the algorithm adopted to map
the features onto the properties (f ). In brief, a QSAR model can always be expressed
as Y = f (X). As for X, if features are extracted only from chemical molecules
themselves, then these features are typically called molecular (structural) descriptors. Details about molecular descriptors can be found in Todeschini and Consonni’s
Molecular Descriptor for Chemoinformatics [75]. Besides, the application domain
beyond which QSAR models could not be reliable also depends strongly on the feature space of chemicals in the training set. As for Y, if the properties have discrete
categories/levels, then it refers to a classification task, resulting in a predictive model
termed as classifier. If the properties have continuous values it refers to a regression
task, resulting in a regressor. The quality of QSAR models heavily relies on the quality of the input data sets. Therefore, collection and curation of data need particular
patience and caution. As for f, nowadays, a large number of algorithms for classification or regression are available, e.g., multivariate linear regression, partial least
regression, support vector machine, decision tree, naïve Bayes, artificial neural network, etc., as well as ensemble algorithms such as random forest [76], among which
computational toxicologists may find some suitable for their specialized cases.
It is also notable that molecular models could inspire novel descriptors that can
better characterize the underlying pattern for certain toxicological phenomena. For
example, simulation of the interaction between halogenated compounds and human
transthyretin protein with QM/MM methods indicated specific descriptors for QSAR
modeling [77]. Similarly, QSARs can translate the time-consuming computational
chemistry models into empirical rules and mathematics, thus resulting in more efficient predictive models. For example, Rydberg and Olsen et al. performed a series
of QM/DFT simulations on the active sites of P450 enzymes with various types of
small molecules [78–80], based on which a web server named SMARTCyp has been
developed that is capable of rapidly predicting sites of metabolism and associated
reaction energy barriers [81].
Currently, the rapid development of modern machine learning algorithms could
most likely promote a renaissance of the field of QSAR modeling. Previously unnoticed details might also be recaptured by innovative feature extraction and wellestablished machine learning algorithms, making scientists’ conclusions or predictions less arbitrary and more robust.
2.4 Challenges for Computational Toxicology
Although modeling frameworks seem to have been nicely established, challenges in
computational toxicology still remain for many aspects.
