26
Z. Wang and J. Chen
Molecular models are different from the previously introduced macroscopic models in many aspects. First of all, when a chemical substance is queried in a macroscopic model, it always refers to a population of a huge number of same-type
molecules. In chemistry or physics, Avogadro’s constant is employed to translate
the microscopically huge number of molecules into macroscopic amount of substances with mole as the unit. Noticeably, even a nanomole of substance corresponds
to 6.02 × 10
14 molecules, which is definitely an astronomical figure far beyond the
capacity of molecular simulations. In a molecular simulation, a queried chemical
is typically represented by a single (in some cases a few) molecule(s) of its type.
For example, when a MD simulation is carried out to calculate binding free energy
of a ligand to a receptor, there is typically just one ligand molecule in the simulated system. A macroscopic binding affinity bioassay usually gives an inhibitory
concentration–response curve, from which a half-inhibition concentration (IC 50 ) or
inhibition constant (K i ) can be calculated as quantitative features for the binding
affinity. However, molecular simulation, with only one ligand molecule, would only
give the free energy difference between or the probability distribution of the binding
state and the free state of the ligand, which could be later translated into equilibrium
constants such as K i . In conclusion, molecular simulation checks the relative potential energies of important microstates, and then uses the obtained potential energy
surface to explain the macroscopic phenomena with principles of statistical mechanics or statistical thermodynamics. Currently, molecular simulation mainly aims at
explaining experimental phenomena by revealing molecular mechanisms, and qualitatively foreseeing tendencies of some properties for a series of congeners. But, as
computational power is increasing and novel efficient algorithms are being implemented, larger systems with thousands of atoms would ultimately be simulated with
more accurate and sophisticated empirical FF or even ab initio QM/DFT methods,
which should give much better predictions to convincingly fill the data gap required
by chemicals risk assessment.
2.3.5 QSAR Models
With the so-called toxicological big data [13], machine learning algorithms have
also proven their usefulness in predicting parameters required by the macroscale
models or data required by chemicals risk assessment, which typically refer to QSAR
modeling or supervised learning in the field of computational toxicology or machine
learning, respectively. QSARs are based on the linear free energy relationship theory
suggested by Hammett [71], Hansch et al. [72] or even on the chemistry intuition
that “chemicals with similar structures have similar properties.”
Distinct from all previously introduced models, QSARs do not simulate a particular chemical object in a physical process/event. In a QSAR study, features are firstly
extracted from a series of chemical objects (so-called training and testing sets) or a
series of molecular scenarios where each queried chemical interacts with certain situational objects, and then these features are employed to predict certain properties
Z. Wang and J. Chen
Molecular models are different from the previously introduced macroscopic models in many aspects. First of all, when a chemical substance is queried in a macroscopic model, it always refers to a population of a huge number of same-type
molecules. In chemistry or physics, Avogadro’s constant is employed to translate
the microscopically huge number of molecules into macroscopic amount of substances with mole as the unit. Noticeably, even a nanomole of substance corresponds
to 6.02 × 10
14 molecules, which is definitely an astronomical figure far beyond the
capacity of molecular simulations. In a molecular simulation, a queried chemical
is typically represented by a single (in some cases a few) molecule(s) of its type.
For example, when a MD simulation is carried out to calculate binding free energy
of a ligand to a receptor, there is typically just one ligand molecule in the simulated system. A macroscopic binding affinity bioassay usually gives an inhibitory
concentration–response curve, from which a half-inhibition concentration (IC 50 ) or
inhibition constant (K i ) can be calculated as quantitative features for the binding
affinity. However, molecular simulation, with only one ligand molecule, would only
give the free energy difference between or the probability distribution of the binding
state and the free state of the ligand, which could be later translated into equilibrium
constants such as K i . In conclusion, molecular simulation checks the relative potential energies of important microstates, and then uses the obtained potential energy
surface to explain the macroscopic phenomena with principles of statistical mechanics or statistical thermodynamics. Currently, molecular simulation mainly aims at
explaining experimental phenomena by revealing molecular mechanisms, and qualitatively foreseeing tendencies of some properties for a series of congeners. But, as
computational power is increasing and novel efficient algorithms are being implemented, larger systems with thousands of atoms would ultimately be simulated with
more accurate and sophisticated empirical FF or even ab initio QM/DFT methods,
which should give much better predictions to convincingly fill the data gap required
by chemicals risk assessment.
2.3.5 QSAR Models
With the so-called toxicological big data [13], machine learning algorithms have
also proven their usefulness in predicting parameters required by the macroscale
models or data required by chemicals risk assessment, which typically refer to QSAR
modeling or supervised learning in the field of computational toxicology or machine
learning, respectively. QSARs are based on the linear free energy relationship theory
suggested by Hammett [71], Hansch et al. [72] or even on the chemistry intuition
that “chemicals with similar structures have similar properties.”
Distinct from all previously introduced models, QSARs do not simulate a particular chemical object in a physical process/event. In a QSAR study, features are firstly
extracted from a series of chemical objects (so-called training and testing sets) or a
series of molecular scenarios where each queried chemical interacts with certain situational objects, and then these features are employed to predict certain properties
