392
B. Hardy et al.
method and protocol description, optimally also in a computer-readable version,
data quality assurance measures and data management are at the heart of creditable
scientific practice. This fact is now generally acknowledged resulting in the FAIR
data principles (findable, accessible, interoperable, and reusable) endorsed by almost
all major funding agencies and many high-impact journals as well as the Guidance
on Good Data and Record Management Practices by the WHO [11]. Well-prepared
data are valuable resources that can be used and reused with high confidence. In this
way, data sharing facilitates new scientific inquiry, avoids duplication of experiments
and data collection, provides rich real-life resources for method validation as well
as education and training and, most importantly for the topic here, allows for the
complete comparison of all steps from raw data to results when trying to reproduce
the study. To allow for data usage without the need to refer to external sources
like publications, final reports, working papers, or laboratory books, all important
information should be included in the data submission. Good data documentation
includes information on:
• author and affiliation contact, relevant dates, reference to test method and protocols
as well as publications and reports;
• the context of data collection: project history, aim, objectives, and hypotheses;
• data collection methods, dataset structure of data files, study cases, relationships
between files;
• data validation, checking, cleaning, and quality assurance procedures carried out;
• changes made to data over time since their original creation and identification of
different versions of data files;
• information on access and use conditions or data confidentiality;
• names, labels, and descriptions for variables, records, and their values;
• explanation or definition of codes and classification schemes used;
• codes of, and reasons for, missing values;
• derived data created after collection, with code, algorithm, or command file.
It is clear that data completeness, quality and reproducibility is mainly influenced
by procedures adopted during data collection and documentation of how data are
collected provides evidence of such quality. The digitization and entering of data, the
documentation of data manipulation, the processing, and most importantly complete
in silico approaches like read-across and QSAR need to follow high-quality data
standards. Errors during input and unintentional modification or disruption can be
avoided or at least be minimized by standardized and consistent procedures with
clear instructions. The procedures described below will provide guidance to follow
these standards, provide tools for on-the-fly verification, as well as minimize human
intervention by:
• setting up validation rules for data entry software as well as providing validation
tools to be executed after each data modification;
• using automatic protocoling of data manipulations and software usage including
hardware and software setup, needed data transformation procedures and programspecific run time parameters;
Précédent

- 397/416

Suivant