19 OpenTox Principles and Best Practices for Trusted Reproducible …
387
Fig. 19.1 Measures to
improve reproducibility of in
silico workflows
• At the same time developing and providing advanced methods for the full integration and documentation of the workflows with repositories like GitHub and
finally by containerization of data including raw and intermediate data as well as
the software tools for producing a complete snapshot of the knowledge generation
process.
If these practices can be successfully realized (and the success of systems like
the blockchain, GitHub, distributed file systems like IPFS, and containerization
approaches like docker suggests that they can), then these reproducibility annotations may also form the basis for future applications that may, for example, actively
generate and regenerate the latest versions of data, results, and conclusions using
the latest tools and input data sources, according to well-defined recipes and in this
way protocoling the scientific advances. Reproducibility annotations could also be a
cornerstone of significant efficiency improvements in organizations that manipulate
biological data.
19.3 Example 1: QSAR Model Building and Validation
The increasing complexity of molecular descriptors and machine learning algorithms
presents a challenge for the reproducibility of QSAR models. The use of nonlinear
algorithms is increasing, a practice that is assisted by the advancements in fields
like neural networks and support vector machines among others. Not only does
this require the explicit documentation of optimization methods, but also algorithm
parameters, including random number seeds, must be noted down. Furthermore, the
construction of metamodels, such as utilizing a bootstrap aggregation (bagging) and
boosting protocols, adds yet another layer of complexity to the model building and
Précédent

- 392/416

Suivant