• Greater scientific insight could be gained – e.g. secondary study of diffuse
scattering purposefully not accounted for in the original model.
• It is necessary to support the review of claims made in publication – this would in
turn support the data-fit-for-purpose concept (see above).
• Software developers could have a range of examples to develop their algorithms
and code.
• Valuable examples arise that can be used for training in advanced data processing
techniques.
With a more developed attitude to post processing, new approaches will also
emerge. One development that would be particularly useful to improve quality and
support dynamic crystallography is the practice of merging data collections to get a
better composite dataset, e.g. in an approach to that used in macromolecular and
serial crystallography communities.
The Data Landscape
‘Chemical space’ is a concept used by data scientists and cheminformaticians that
refers to the property space spanned by all possible molecules, or compounds, for a
given property. Chemical space is very large, e.g. pharmacological chemical space is
estimated to cover 10
63 molecules [226], and this even has many restrictions,
e.g. does not include molecular weight >500 and only includes simple atoms
(C, H, O, N, S) – many of the compounds are yet to exist. The Chemical Abstracts
Service [227], which extracts compounds from the scientific literature, contains
158 million entries.
In comparison there are somewhat less than two million structures in the space
covered by crystallographic databases. There are a number of reasons why these
levels are so different – primarily that only a subset of materials are crystalline.
However, a huge contributing factor is that most small molecule structures are
determined as a service for synthesis chemists and subject to academic publishing
rules in order to be available. There are further factors related to this phenomenon,
e.g. that parameters/variables for crystallisation are not comprehensively tested; not
all structures of a homologous/related family of compounds are deemed necessary
(only one representative compound is suitable for publication – others may/may not
have been determined); a structure may not be of suitable quality for publication;
some compounds are ‘not academically interesting enough’ (from the perspective of
a funder or synthesis chemist).
For these, and other, reasons, the crystallographic databases could be considered
to have ‘gaps’ in them – particularly from the perspective of a data scientist. For
some data science research to be valid or achievable, these gaps would have to be
filled. There is a need for a change of mindset so that some of these gaps can begin to
be filled – if the incentives were different, there is no reason why these gaps could
not be identified and this used to drive synthesis programs and change culture. If a
compound has been synthesised and there is ready access to crystallography, then
130
S. J. Coles et al.
Précédent

- 138/285

Suivant