19 OpenTox Principles and Best Practices for Trusted Reproducible …
385
In data science and computational modeling, this problem has reached new magnitudes of concern, due to the diversity of approaches, software tools, and hardware
architectures involved. Many methods are emerging in the research field of alternative
testing methods supporting safety assessment, but industry and regulators are finding
the establishment of reproducible application, evaluation, and guidance challenging,
and current regulatory acceptance is as a result moving slowly.
Our aim in this work and initiative is to establish a practice and guidance for
tracking and reporting modern in silico data analyses in a reproducible manner.
Reproducible research supports the concept that data analyses, and more generally,
scientific claims, are published with their raw data and software code so that others
may verify the findings and build upon them. Reproducibility allows people to focus
on the actual content of a data analysis, rather than on superficial details reported
in a written methods summary. In addition, reproducibility makes an analysis more
useful to others because the data and code used to conduct the analysis are available.
This paper focuses on literate statistical analysis tools which allow one to publish
data analyses in a single repository that allows others to easily execute the same
analysis to obtain the same results. Although many of the principles described here
apply broadly to scientific research and associated data science and principles, we
will focus on the domains of toxicology and risk assessment where data is generated
and used and models are developed primarily for the purpose of ensuring the safe
level of use of a chemical or drug in a human or the environment.
The work proposed here aims to exploit background developments related to the
OpenTox community (http://www.opentox.net/), and its associated current infrastructure project OpenRiskNet (https://openrisknet.org/), so that information generated and processed by in silico methods are more suitable for purpose for industrial
use and regulatory acceptance including:
(a) Establishment of application programming interfaces (APIs) for scientific data
within the context of an open infrastructure providing reliable quality-controlled
access to harmonized scientific data;
(b) Development of best practice in silico workflows, processing the data with the
principle of reproducibility as established by the context of use;
(c) Engagement of the scientific community to develop and contribute best practice
in silico workflows;
(d) Establishment within the OpenTox community of guidance to best practices for
in silico workflows in predictive toxicology.
Our objective is to develop guidance, methods, and best practices supporting
reproducible in silico computational toxicology and safety assessment. Implementations within OpenTox and OpenRiskNet will be established against use cases involving model building, validation and integrated testing. The approach will be extended
to include additional contributions from the scientific and regulatory communities for
elaboration and consensus building. This approach should support the independent
verification of resources used in producing results as toxicological evidence.
Currently, basic research and outputs exist, but the approaches for regulatory use
and acceptance are missing from practice. We need to develop reliable access to data,
385
In data science and computational modeling, this problem has reached new magnitudes of concern, due to the diversity of approaches, software tools, and hardware
architectures involved. Many methods are emerging in the research field of alternative
testing methods supporting safety assessment, but industry and regulators are finding
the establishment of reproducible application, evaluation, and guidance challenging,
and current regulatory acceptance is as a result moving slowly.
Our aim in this work and initiative is to establish a practice and guidance for
tracking and reporting modern in silico data analyses in a reproducible manner.
Reproducible research supports the concept that data analyses, and more generally,
scientific claims, are published with their raw data and software code so that others
may verify the findings and build upon them. Reproducibility allows people to focus
on the actual content of a data analysis, rather than on superficial details reported
in a written methods summary. In addition, reproducibility makes an analysis more
useful to others because the data and code used to conduct the analysis are available.
This paper focuses on literate statistical analysis tools which allow one to publish
data analyses in a single repository that allows others to easily execute the same
analysis to obtain the same results. Although many of the principles described here
apply broadly to scientific research and associated data science and principles, we
will focus on the domains of toxicology and risk assessment where data is generated
and used and models are developed primarily for the purpose of ensuring the safe
level of use of a chemical or drug in a human or the environment.
The work proposed here aims to exploit background developments related to the
OpenTox community (http://www.opentox.net/), and its associated current infrastructure project OpenRiskNet (https://openrisknet.org/), so that information generated and processed by in silico methods are more suitable for purpose for industrial
use and regulatory acceptance including:
(a) Establishment of application programming interfaces (APIs) for scientific data
within the context of an open infrastructure providing reliable quality-controlled
access to harmonized scientific data;
(b) Development of best practice in silico workflows, processing the data with the
principle of reproducibility as established by the context of use;
(c) Engagement of the scientific community to develop and contribute best practice
in silico workflows;
(d) Establishment within the OpenTox community of guidance to best practices for
in silico workflows in predictive toxicology.
Our objective is to develop guidance, methods, and best practices supporting
reproducible in silico computational toxicology and safety assessment. Implementations within OpenTox and OpenRiskNet will be established against use cases involving model building, validation and integrated testing. The approach will be extended
to include additional contributions from the scientific and regulatory communities for
elaboration and consensus building. This approach should support the independent
verification of resources used in producing results as toxicological evidence.
Currently, basic research and outputs exist, but the approaches for regulatory use
and acceptance are missing from practice. We need to develop reliable access to data,
