present. It is possible that higher R-factors observed for more challenging structures
will be offset by the fact that there will be an increasing number of very good values
resulting from advances in instrumentation, technique and software and therefore the
net R-factor remains approximately the same. However, the consistent value at
around 5% may also be defined by a general requirement from some journals for
structures to be of a certain, relatively high, standard and that the community is not
embracing the ‘data-fit-for-purpose’ concept. It is acknowledged that the R-factor is
just one indicator of data quality and others should be taken into account when the
data is considered for re-use. Of course, the chemical sense of the final model should
not be overlooked, via a consideration of atom assignment, analysis of molecular
geometries and an appraisal of anisotropic displacement parameters (ADPs). Using a
range of metrics would enable an overall assessment of the final structure quality
over the whole process from collection to refinement [99].
The increase in the size of molecules, together with improvements in technique
and software, has resulted in an increase in structures with modelled disorder. If the
current trend continues, it is estimated that 50% of new structures added to the CSD
will have modelled disorder by 2060. Modelled disorder appears to be more common in metal-organic structures, as 70% of disordered structures contain a metal.
Highly symmetric components and those without strong intermolecular interactions
to neighbouring molecules may have greater freedom to occupy multiple orientations throughout the crystal and therefore exhibit disorder which is difficult to model.
Perhaps due to the rise in popularity of MOFs, there has been a significant increase in
the number of structures that contain a co-crystallised solvent, or solvents, in the
structure. These too are invariably very disordered, often adopting many orientations
in large void spaces in this type of structure. It is now possible to omit such entities
from the calculation of structure factors and hence from the model itself. Since
A.L. Spek’s initial paper [100] on the BYPASS routine and the subsequent implementation in PLATON [62] and Olex2 [84], there has been a dramatic uptake in the
use of these approaches. This is illustrated by the increase in the percentage of entries
using the SQUEEZE or MASK algorithm in refinement to remove solvent and
produce a better model fit for the part of the structure of primary interest (Fig. 8).
This has allowed complex structures to be determined that previously would not
have been possible, such as those with severe disorder and those where a satisfactory
model for some of the electron density cannot be found.
One effect that the use of improved apparatus as described in previous sections of
this article can have is in the observation of improved precision for bond lengths.
This could be illustrated by the increased percentage of structures with low average
estimated standard deviation on C-C bonds. Increased precision is valuable for a
variety of reasons; however, it has a significant effect on the re-use and repurposing
of data in building more accurate models and knowledgebases, which are amongst
concepts that will be discussed later in this review. The inclusion of hydrogen atoms
in the refinement model is also valuable for the re-use of data, especially when
examining hydrogen-bonding motifs. Improvements in data quality and refinement
procedures have resulted in a dramatic increase in the inclusion of hydrogen atoms
110
S. J. Coles et al.
will be offset by the fact that there will be an increasing number of very good values
resulting from advances in instrumentation, technique and software and therefore the
net R-factor remains approximately the same. However, the consistent value at
around 5% may also be defined by a general requirement from some journals for
structures to be of a certain, relatively high, standard and that the community is not
embracing the ‘data-fit-for-purpose’ concept. It is acknowledged that the R-factor is
just one indicator of data quality and others should be taken into account when the
data is considered for re-use. Of course, the chemical sense of the final model should
not be overlooked, via a consideration of atom assignment, analysis of molecular
geometries and an appraisal of anisotropic displacement parameters (ADPs). Using a
range of metrics would enable an overall assessment of the final structure quality
over the whole process from collection to refinement [99].
The increase in the size of molecules, together with improvements in technique
and software, has resulted in an increase in structures with modelled disorder. If the
current trend continues, it is estimated that 50% of new structures added to the CSD
will have modelled disorder by 2060. Modelled disorder appears to be more common in metal-organic structures, as 70% of disordered structures contain a metal.
Highly symmetric components and those without strong intermolecular interactions
to neighbouring molecules may have greater freedom to occupy multiple orientations throughout the crystal and therefore exhibit disorder which is difficult to model.
Perhaps due to the rise in popularity of MOFs, there has been a significant increase in
the number of structures that contain a co-crystallised solvent, or solvents, in the
structure. These too are invariably very disordered, often adopting many orientations
in large void spaces in this type of structure. It is now possible to omit such entities
from the calculation of structure factors and hence from the model itself. Since
A.L. Spek’s initial paper [100] on the BYPASS routine and the subsequent implementation in PLATON [62] and Olex2 [84], there has been a dramatic uptake in the
use of these approaches. This is illustrated by the increase in the percentage of entries
using the SQUEEZE or MASK algorithm in refinement to remove solvent and
produce a better model fit for the part of the structure of primary interest (Fig. 8).
This has allowed complex structures to be determined that previously would not
have been possible, such as those with severe disorder and those where a satisfactory
model for some of the electron density cannot be found.
One effect that the use of improved apparatus as described in previous sections of
this article can have is in the observation of improved precision for bond lengths.
This could be illustrated by the increased percentage of structures with low average
estimated standard deviation on C-C bonds. Increased precision is valuable for a
variety of reasons; however, it has a significant effect on the re-use and repurposing
of data in building more accurate models and knowledgebases, which are amongst
concepts that will be discussed later in this review. The inclusion of hydrogen atoms
in the refinement model is also valuable for the re-use of data, especially when
examining hydrogen-bonding motifs. Improvements in data quality and refinement
procedures have resulted in a dramatic increase in the inclusion of hydrogen atoms
110
S. J. Coles et al.
