level of formalization to be automatically leveraged by a digital simulation. And on
the models for how the system operates, we are often faced with the competing
approaches of trying a bottom-up approach, component by component, that fails to
capture the system dynamics that actually defines what happens; and an end-to-end
heuristic model delivering insights and correlations devoid of any physical causality,
amplifying noise in measurements more than modeling the real world.
Success in this area requires a convergence of practices in different disciplines
which is still very emerging.
3.7 Data
Many digital twins fail due to a lack of data from previous runs of the target process
and equipment. Without enough volumes of data many of the modern machine
learning algorithms become unusable. While we have seen success from pure
mechanistic modeling in many discrete manufacturing operations (Automotive,
Aeronautics,. . .), in Biopharma we seem to still require real-world experiments to
provide the context for simulations.
Capturing data is not an issue by itself but there are two challenges:
• Providing experimental data requires several actual experimentations. For new
operations, this can only come from development activities. To be noted that
some industries have successfully created the ability to do virtual experimentation
(a case in point is autonomous driving where most major players have successfully created virtual playgrounds to increase the learning of the driving machine
learning model). In Biopharma, based on the complexity of the system we shall
most probably have to go in the direction of a mixed model.
• Beyond the measurement, you need to understand what it means and be able to
reuse it in different context. This is an overall question of appropriately “contextualizing” data. In many cases data has been captured but without tracking the
different changes that were done to the system while doing experiments. Thus,
you need to “realign” those datasets to make them comparable. The different
capabilities discussed above in point IV. and V. are what helps you achieve this.
Generating data will in many cases be one of the longest parts of a digital twin
project or barriers to doing it.
3.8 Data Modeling and Ontologies
Good data modeling is always important for any kind of data analysis project. In
cases where we are manipulating complex data with high variability, it is critical to
raise the bar even more.
174
M. Canzoneri et al.
Précédent

- 180/260

Suivant