The main message of this review has the implication that there should be much
more integration between the experiment and the database. With the correct infrastructure, a far greater number of structures could be included in the databases. For
example, a lot of data mining and machine learning studies merely need to know that
a structure exists, and its quality is of secondary importance – with a relatively small
development and a change of publishing mindset, a vast number of new types of
structures could be incorporated into databases. This ‘data-fit-for-purpose’ concept
has the potential to hugely transform the amount of data available for follow-on
studies; however, it requires further development of validation procedures and more
accurate classification in databases, so that the integrity of a good-quality (subset)
collection is not compromised for other areas of study. These developments could
enable a clear, simple, fast, automated route from diffractometer to database which
could in turn empower data mining, statistical and machine learning methods. These
approaches are powerful not only in terms of performing structural chemistry
analyses but also in making connections and correlations with data from other
disciplines. There is a different data infrastructure that is required for this kind of
work, and it is necessary to develop this for our subject. Currently data science
involves a significant amount of data cleaning and transformation before the techniques can be applied – and this is often a very manual process. In order to make data
science involving crystal structures seamless, it will be necessary to understand and
implement new ways to interface between data collections – from both a metadata
and descriptor perspective as well as via scripting and automation, e.g. via Application Programming Interfaces (API). Furthermore, changing the nature of the interactions between laboratory and database and data re-user will drive other
developments, for example, the use of the database more interactively for structure
refinement. It will also enhance integration between collaborators, complementary
techniques and facilities.
Finally, it is necessary to extend the data infrastructure to enable greater inclusion
of raw data. Currently there is no culture of sharing raw data in chemical
crystallography – and in fact there is even a significant diversity in which individual
facilities manage their own raw data. This lack of comprehensive management leads
to difficulties in accessing the data (locally or globally) over time. Many other
disciplines now routinely make raw data available, and there is an increasing
pressure from funders and other stakeholders for this to be routine practice. There
are many cultural, political and financial barriers to overcome for this to happen, but
also some technical matters around description, validation, quality and storage
would have to be addressed. Furthermore, it is worth considering whether it is
necessary to make ALL raw data available, e.g. would it be necessary in the case
where diffraction was very clean and all Bragg data had been accounted for? The
IUCr CommDat [60] are considering these matters, and as a first step, the development of a CheckCIF-type utility for raw data is being investigated. Nevertheless, it
would be very beneficial for a data infrastructure to support raw data where:
• Others, or future developments, could do a better job – in theory models could
automatically be updated if better processing software were available.
Leading Edge Chemical Crystallography Service Provision and Its Impact on. . .
129
more integration between the experiment and the database. With the correct infrastructure, a far greater number of structures could be included in the databases. For
example, a lot of data mining and machine learning studies merely need to know that
a structure exists, and its quality is of secondary importance – with a relatively small
development and a change of publishing mindset, a vast number of new types of
structures could be incorporated into databases. This ‘data-fit-for-purpose’ concept
has the potential to hugely transform the amount of data available for follow-on
studies; however, it requires further development of validation procedures and more
accurate classification in databases, so that the integrity of a good-quality (subset)
collection is not compromised for other areas of study. These developments could
enable a clear, simple, fast, automated route from diffractometer to database which
could in turn empower data mining, statistical and machine learning methods. These
approaches are powerful not only in terms of performing structural chemistry
analyses but also in making connections and correlations with data from other
disciplines. There is a different data infrastructure that is required for this kind of
work, and it is necessary to develop this for our subject. Currently data science
involves a significant amount of data cleaning and transformation before the techniques can be applied – and this is often a very manual process. In order to make data
science involving crystal structures seamless, it will be necessary to understand and
implement new ways to interface between data collections – from both a metadata
and descriptor perspective as well as via scripting and automation, e.g. via Application Programming Interfaces (API). Furthermore, changing the nature of the interactions between laboratory and database and data re-user will drive other
developments, for example, the use of the database more interactively for structure
refinement. It will also enhance integration between collaborators, complementary
techniques and facilities.
Finally, it is necessary to extend the data infrastructure to enable greater inclusion
of raw data. Currently there is no culture of sharing raw data in chemical
crystallography – and in fact there is even a significant diversity in which individual
facilities manage their own raw data. This lack of comprehensive management leads
to difficulties in accessing the data (locally or globally) over time. Many other
disciplines now routinely make raw data available, and there is an increasing
pressure from funders and other stakeholders for this to be routine practice. There
are many cultural, political and financial barriers to overcome for this to happen, but
also some technical matters around description, validation, quality and storage
would have to be addressed. Furthermore, it is worth considering whether it is
necessary to make ALL raw data available, e.g. would it be necessary in the case
where diffraction was very clean and all Bragg data had been accounted for? The
IUCr CommDat [60] are considering these matters, and as a first step, the development of a CheckCIF-type utility for raw data is being investigated. Nevertheless, it
would be very beneficial for a data infrastructure to support raw data where:
• Others, or future developments, could do a better job – in theory models could
automatically be updated if better processing software were available.
Leading Edge Chemical Crystallography Service Provision and Its Impact on. . .
129
