388
B. Hardy et al.
documentation making it harder to reproduce. Unless the authors contribute a saved
copy of their model, it is almost impossible to reproduce the QSAR model based
on a text description alone. A standard format (the QSAR model reporting format
(QMRF)) has been suggested to report on the structure of a QSAR model. It takes
the form of a harmonized template for summarizing and reporting key information
on QSAR models, including the results of any validation studies. The information is
structured according to the OECD QSAR validation principles. At the time of writing,
the QMRF database published by the European Commission’s Joint Research Centre
(http://qsardb.jrc.it/qmrf/) includes over one hundred such reports.
Thousands of chemical descriptors have been documented [1]. While many
descriptors are chemically intuitive, for instance “number of hydrogen-bond acceptors,” their algorithmic implementation is often not. Many software packages have
different interpretations of the same SMILES representation of a chemical structure
resulting in a different count of what is to be considered “hydrogen-bond acceptor.”
This problem may often be aggravated by the poor description of the structure cleaning steps (such as whether the modelers renormalized aromaticity, re-optimized 3D
structures, or neutralized charges). As a minimum requirement for reproducibility,
the name of the software tool and its version have to be documented. A more practical
approach is to internally store (and possibly publish) an exact copy of the software
program and the preprocessing workflows that were used to calculate descriptors.
Such a step is more easily attained for open software than proprietary programs due
to the availability of the source code and the inviting license model.
The importance of validating QSAR models for regulatory acceptance cannot
be overstated. A validated model is one that is able to produce consistent results
when tested on external validation set(s) not involved in its training. For external
researchers and regulators to validate a QSAR model, they must be able to reproduce
it with sufficient accuracy. Many online platforms for QSAR model building and validation have been developed: OpenTox, OCHEM [2], Chembench [3], and AMBIT
[4], among others. Many of these platforms can be managed through application
programming interfaces (APIs) allowing the scripting control of complicated QSAR
model building workflows by passing the parameters needed for machine learning
algorithms, descriptor packages, variables selection, and preprocessing steps. Modelers can therefore document an entire workflow using such a script, while the online
platforms store the underlying binaries and versions. OCHEM also offers its XML
format for documenting such settings in a reproducible manner.
OpenTox proposed a best practice for building validated in silico QSAR models accompanied by specifications for APIs and supported by open standards and
ontology (see Fig. 19.2) for harmonized knowledge descriptions and communications between components for data, algorithms, models, and validation. We propose
to update and upgrade these specifications and best practices through making a proposal with accompanying case study examples to the community to include the goals
of reproducibility, trust, and provenance discussed in this paper. Such a QSAR may
also be deployed in an ITS as described in the example in the next section.
Précédent

- 393/416

Suivant