19 OpenTox Principles and Best Practices for Trusted Reproducible …
393
• supporting the usage of controlled vocabularies taken from and referencing standard ontologies, code lists, and choice lists to minimize manual data entry;
• eliminating manual file format transformations by data harmonization and enhancing the interoperability of the integrated analysis software.
19.6 Proposed Technology Solutions for Data Processing
In the domain of software engineering, the intersecting problems of versioning and
collaboration were solved by the introduction of Distributed Version Control Systems
(DVCS), most notably among them git [12]. These systems work by recording all
changes to all source code files of a project over the entire history of the project,
starting with an empty folder and arriving at the latest development version. When a
changeset is created by the user, the DVCS compares the status of all files with the
last recorded changeset and creates a list of differences across all files. This list of
differences is then combined with metadata about the changeset (timestamp, author
information, the hash of the previous changeset, etc.). The resulting data is hashed,
meaning that a unique numeric identifier of fixed length is calculated and added to
the inventory of changesets.
We envision a similar system for data and data transformations, to enable verification and reproduction of data and workflows operating on them. This would consist
of a standardized description of metadata about the data and data transformation
steps, linked together as a merkle tree [13] similar to the way git or bitcoin operates.
In a simple case, a downloaded raw data file would be hashed, metadata added
(e.g., the URL the dataset was retrieved from, a timestamp of the retrieval), and the
result recorded as a first step alongside the downloaded data. An analysis that would
then be run on top of this as an R script would similarly be recorded into the system,
but in this case, the metadata would include the path to the R script in a git repository
and the hash of the version used when running the analysis. Crucially, the metadata
of this second step would include the hash of the first step to link these two steps
together.
We acknowledge that risk assessment is an environment where the use of such
systems is not yet widespread, data manipulation is often done offline, and thus,
no full chain of linked processing description steps exist from the final result to
the raw experimental data (though such a full chain would certainly be possible
and desirable). As such, we anticipate a verification system concentrating on the
“tail end” of the analysis initially (e.g., starting with a downloaded dataset about
which nothing more is known than the URL where it was downloaded from and then
doing one or two analysis steps to reach a conclusion). As the usefulness of such an
approach is recognized in the wider community, we hope that data managers either
responsible for public or in-house data would themselves, in collaboration with the
experimental scientists, start to record such information based on the approaches
Précédent

- 398/416

Suivant