394
B. Hardy et al.
developed for the verification system. Summary datasets prepared like this would
contain recordings of all data transformations that were used to derive this data from
the raw experimental data allowing for an in-depth validation and reproduction of
the procedure during the review process of a regulatory submission. The technology
primitives for such systems are well established and are in use in git, bitcoin and other
blockchains, and IPFS (the InterPlanetary File System) and can be used to express
nonlinear relationships between workflow steps. Another often volatile component
of any data-based workflow is the system utilities and libraries installed as part of
the operating system (OS). With the recent development of NixOS Linux [14], even
these parts could be recorded into the merkle tree, because NixOS provides capability
to hash all the tools available in an OS at a given point in time and to restore this
exact state at a later point in time.
19.7 Locating the Source of Irreproducibility: Sub-tasks
and Intermediate Datasets
The system we envision would contain the necessary information not just for a
human to audit the data steps and be able to reference the precise files and software
versions to reproduce a data analysis or transformation—it would indeed go further
and describe the necessary execution steps in enough machine readable detail to
reproduce the workflow automatically (probably by extending an existing workflow
tool like NextFlow or similar). In such a scenario, the analysis could be replayed
by the workflow tool one to one. A successful demonstration of such an approach
is the literate programming environment R Markdown, which can regenerate an
entire publication including all data downloads from original URLs, through all
intermediate data cleaning and processing steps until the final generation of the
report pdf. The default way of comparing any intermediate step would be to compare
the hashes—this would allow a user reproducing the workflow to see if there is
at any point a divergence between the original publication or analysis and their
reproduction but would not give them more information. To make the system more
powerful, we want to follow the example of git and make the actual diffing algorithm
interchangeable. In this way, a data format aware diffing could be used to compare
files with a deeper understanding of its content—for example, when two csv files
(the first from the original study submitted with the report and the second from the
validation study) differ, the divergence could be highlighted in the file instead of
giving just a yes/no answer to the question “Are these files binary identical?”. Such
a flexible diffing would allow users that reproduce the data to define a threshold of
similarity, up until which they would still regard the result as being reproduced (e.g.,
if there is a newer original dataset that is used to run the same analysis).
B. Hardy et al.
developed for the verification system. Summary datasets prepared like this would
contain recordings of all data transformations that were used to derive this data from
the raw experimental data allowing for an in-depth validation and reproduction of
the procedure during the review process of a regulatory submission. The technology
primitives for such systems are well established and are in use in git, bitcoin and other
blockchains, and IPFS (the InterPlanetary File System) and can be used to express
nonlinear relationships between workflow steps. Another often volatile component
of any data-based workflow is the system utilities and libraries installed as part of
the operating system (OS). With the recent development of NixOS Linux [14], even
these parts could be recorded into the merkle tree, because NixOS provides capability
to hash all the tools available in an OS at a given point in time and to restore this
exact state at a later point in time.
19.7 Locating the Source of Irreproducibility: Sub-tasks
and Intermediate Datasets
The system we envision would contain the necessary information not just for a
human to audit the data steps and be able to reference the precise files and software
versions to reproduce a data analysis or transformation—it would indeed go further
and describe the necessary execution steps in enough machine readable detail to
reproduce the workflow automatically (probably by extending an existing workflow
tool like NextFlow or similar). In such a scenario, the analysis could be replayed
by the workflow tool one to one. A successful demonstration of such an approach
is the literate programming environment R Markdown, which can regenerate an
entire publication including all data downloads from original URLs, through all
intermediate data cleaning and processing steps until the final generation of the
report pdf. The default way of comparing any intermediate step would be to compare
the hashes—this would allow a user reproducing the workflow to see if there is
at any point a divergence between the original publication or analysis and their
reproduction but would not give them more information. To make the system more
powerful, we want to follow the example of git and make the actual diffing algorithm
interchangeable. In this way, a data format aware diffing could be used to compare
files with a deeper understanding of its content—for example, when two csv files
(the first from the original study submitted with the report and the second from the
validation study) differ, the divergence could be highlighted in the file instead of
giving just a yes/no answer to the question “Are these files binary identical?”. Such
a flexible diffing would allow users that reproduce the data to define a threshold of
similarity, up until which they would still regard the result as being reproduced (e.g.,
if there is a newer original dataset that is used to run the same analysis).
