386
B. Hardy et al.
specify the metadata needed, implement interoperability layers and workflows, and
obtain community contributions and acceptance. Achievement of such goals will
increase the reliability of predictions, increase the ability to reproduce a prediction
and determine reasons for deviations, and support the independent verification of
resources used.
19.2 Proposed Principle of Reproducibility for In Silico
Modeling and Workflows and Implementation
in Practice
We state the principle of reproducibility here as:
The Principle of Reproducibility states that close agreement in scientific results can be
obtained when a sufficiently well-described protocol is competently executed. The protocol
can be experimental (in vitro or in vivo) or computational (in silico) or a combination of
such protocols.
The main concern of the current paper addresses the achievement in practice of
reproducible computational (in silico) protocols. Bioinformatics and data science
in general are fraught with challenge that hamper the achievement of the principle
of reproducibility in practice. These include correctly identifying the dataset one
is working with, identifying the transformations and manipulations that may have
been done to it, validating these, as well as similarly recording new transformations that one applies. The ease of making, modifying, and distributing copies of
digital data leads to the proliferation of multiple versions of datasets, which may
have ambiguous origin and meaning. At the same time, technologies such as formal
ontologies and blockchain provide opportunities to address this problem. At their
core, the blockchain methods build on calculating hashes (checksums) of the data
and the software used for data processing (Sect. 19.6 gives a more in-depth outline
of this approach).
We suggest a solution structured around verification and reproducibility annotations, implemented according to the following practices (see Fig. 19.1):
• Formally identifying datasets, and versions of these, with a suitable hash function;
• Similarly identifying tools (computational steps that produce new or derived versions of data) and versions of these;
• Generating an “audit trail” that describes all transformations that have been applied
to data that is retrieved, for example, through APIs, by referring to hashes of
specific versions of necessary tools, inputs, and intermediate data. This would at
a minimum be in a way that allows such transformations to be verified and ideally
in a way that allows them to be reproduced. When verification fails, the supporting
tools might suggest how to update or repair the workflow;
• Making it trivially easy to integrate or introduce new transformations (such as
simple R or Python scripts) into this kind of auditing system, such as to not interfere
with existing work habits;
B. Hardy et al.
specify the metadata needed, implement interoperability layers and workflows, and
obtain community contributions and acceptance. Achievement of such goals will
increase the reliability of predictions, increase the ability to reproduce a prediction
and determine reasons for deviations, and support the independent verification of
resources used.
19.2 Proposed Principle of Reproducibility for In Silico
Modeling and Workflows and Implementation
in Practice
We state the principle of reproducibility here as:
The Principle of Reproducibility states that close agreement in scientific results can be
obtained when a sufficiently well-described protocol is competently executed. The protocol
can be experimental (in vitro or in vivo) or computational (in silico) or a combination of
such protocols.
The main concern of the current paper addresses the achievement in practice of
reproducible computational (in silico) protocols. Bioinformatics and data science
in general are fraught with challenge that hamper the achievement of the principle
of reproducibility in practice. These include correctly identifying the dataset one
is working with, identifying the transformations and manipulations that may have
been done to it, validating these, as well as similarly recording new transformations that one applies. The ease of making, modifying, and distributing copies of
digital data leads to the proliferation of multiple versions of datasets, which may
have ambiguous origin and meaning. At the same time, technologies such as formal
ontologies and blockchain provide opportunities to address this problem. At their
core, the blockchain methods build on calculating hashes (checksums) of the data
and the software used for data processing (Sect. 19.6 gives a more in-depth outline
of this approach).
We suggest a solution structured around verification and reproducibility annotations, implemented according to the following practices (see Fig. 19.1):
• Formally identifying datasets, and versions of these, with a suitable hash function;
• Similarly identifying tools (computational steps that produce new or derived versions of data) and versions of these;
• Generating an “audit trail” that describes all transformations that have been applied
to data that is retrieved, for example, through APIs, by referring to hashes of
specific versions of necessary tools, inputs, and intermediate data. This would at
a minimum be in a way that allows such transformations to be verified and ideally
in a way that allows them to be reproduced. When verification fails, the supporting
tools might suggest how to update or repair the workflow;
• Making it trivially easy to integrate or introduce new transformations (such as
simple R or Python scripts) into this kind of auditing system, such as to not interfere
with existing work habits;
