while the rest is not obvious to interpret. QSAR, however, exploits all data collected
(both positive and negative results) with the aim of predicting the behaviour of
compounds of interest. Perhaps the simplest and most easily interpretable form is the
linear free energy relationship, which has been used by organic chemists since the
1930s following the pioneering work from Hammett [53]. In the context of asymmetric catalysis and enantioselectivity, the quantitative structure-selectivity relationship (QSSR) asserts that there is a fundamental relationship between selectivity and a
quantitative description of the reactants, catalyst and other reaction conditions.
Although the physical basis underlying such a relationship typically results from
common mechanistic features shared across the different reactions, model creation in
itself does not require a mechanistic hypothesis to be formulated. QSSR and QSAR
share a common approach, with the exception that the endpoints differ (i.e. activity
and selectivity).
5.1 QSAR Best Practices
The experimental design of a QSAR study requires careful consideration at each
stage: data collection, parameter selection, model construction, validation and interpretation. A substantial body of literature exists to steer chemists clear of common
pitfalls in this process. Tropsha has summarized a rigorous protocol for QSAR
practices. Dearden, Cronin and Kaiser published a critical review on QSAR model
construction and helpfully include a checklist of errors to be aware of supported by
case studies [54]. Sigman has recently published a review on multivariate linear
regression (MLR) for reaction development [55]. Denmark and co-workers have
recently published a comprehensive review featuring multiple types of QSSR
models used in enantioselective catalysis, particularly the molecular interaction
field 3D-QSSR [56]. Detailed terminologies, definitions and practices regarding
QSAR models and case studies can be found in Organization for Economic
Co-operation and Development (OECD) guidelines, underlining the importance of
the approach far outside of chemistry [57]. Distilled from the literature, we now
summarize the most widely accepted QSAR practice, as applied to the development
of catalytic reactions, in a flow chart displayed below (Fig. 12).
This process can be divided into four stages starting with data collection and
processing, model building, model validation (which is intertwined with model
building in a feedback loop) and application. In each step on the flow chart below,
requirements and generally accepted criteria are summarized in a form of checklist
boxes. For example, in the first step of data processing, it is crucial that all the data
are homogeneous, i.e. with identical units and collected in identical manner to avoid
any random artefacts incurred during data collection. In addition, repeats are important to show that the data is reproducible and to minimize random error. Moving
forwards to model construction, data (dependent parameter, y i ) is split into training
set (used in model construction) and testing set (to validate the model). Chemical
descriptors for the training data are collected, and most relevant are selected during
Ligand Design for Asymmetric Catalysis: Combining Mechanistic and. . .
171
(both positive and negative results) with the aim of predicting the behaviour of
compounds of interest. Perhaps the simplest and most easily interpretable form is the
linear free energy relationship, which has been used by organic chemists since the
1930s following the pioneering work from Hammett [53]. In the context of asymmetric catalysis and enantioselectivity, the quantitative structure-selectivity relationship (QSSR) asserts that there is a fundamental relationship between selectivity and a
quantitative description of the reactants, catalyst and other reaction conditions.
Although the physical basis underlying such a relationship typically results from
common mechanistic features shared across the different reactions, model creation in
itself does not require a mechanistic hypothesis to be formulated. QSSR and QSAR
share a common approach, with the exception that the endpoints differ (i.e. activity
and selectivity).
5.1 QSAR Best Practices
The experimental design of a QSAR study requires careful consideration at each
stage: data collection, parameter selection, model construction, validation and interpretation. A substantial body of literature exists to steer chemists clear of common
pitfalls in this process. Tropsha has summarized a rigorous protocol for QSAR
practices. Dearden, Cronin and Kaiser published a critical review on QSAR model
construction and helpfully include a checklist of errors to be aware of supported by
case studies [54]. Sigman has recently published a review on multivariate linear
regression (MLR) for reaction development [55]. Denmark and co-workers have
recently published a comprehensive review featuring multiple types of QSSR
models used in enantioselective catalysis, particularly the molecular interaction
field 3D-QSSR [56]. Detailed terminologies, definitions and practices regarding
QSAR models and case studies can be found in Organization for Economic
Co-operation and Development (OECD) guidelines, underlining the importance of
the approach far outside of chemistry [57]. Distilled from the literature, we now
summarize the most widely accepted QSAR practice, as applied to the development
of catalytic reactions, in a flow chart displayed below (Fig. 12).
This process can be divided into four stages starting with data collection and
processing, model building, model validation (which is intertwined with model
building in a feedback loop) and application. In each step on the flow chart below,
requirements and generally accepted criteria are summarized in a form of checklist
boxes. For example, in the first step of data processing, it is crucial that all the data
are homogeneous, i.e. with identical units and collected in identical manner to avoid
any random artefacts incurred during data collection. In addition, repeats are important to show that the data is reproducible and to minimize random error. Moving
forwards to model construction, data (dependent parameter, y i ) is split into training
set (used in model construction) and testing set (to validate the model). Chemical
descriptors for the training data are collected, and most relevant are selected during
Ligand Design for Asymmetric Catalysis: Combining Mechanistic and. . .
171
