Material Agnostic Data-Driven Framework to Develop Structure-Property Linkages
255
other features of interest in simulated data is trivial in most cases, experimental data
often requires segmentation of images to properly identify a given feature of interest.
As needed, one might set criterion to eliminate spurious or questionable data (e.g.,
the data that does not conform to known physics). In this step, the inputs (process
parameters) are also clearly associated with the outputs (microstructure data).
It is acknowledged that the robustness of this preprocessing step is a direct
function of the uncertainties introduced during data collection, experimentation, and
data processing techniques. For example, measurement uncertainty is immediately
introduced with data collection due to instrument accuracy and human error.
Additional uncertainty is introduced as these data are processed into human and
computer interpretable forms using various data processing algorithms. For example, the uncertainty associated with imaging and image processing of microscopic
data has been studied and published for a variety of techniques. Uncertainty is
introduced at each stage of the processes depending on the methods and algorithms
employed. While understanding these uncertainties is important to the success of
the framework proposed here, there is much ongoing research in this area, and this
topic is not addressed in detail here but in another chapter in this volume.
In the second step, microstructures are quantified to obtain salient statistical
measures of microstructures. In a data science approach, it is desirable to capture
a very large set of measures at this stage. Consequently, it is preferable to adopt a
microstructure quantification framework that allows one to increase systematically
the numbers of potential features included in the analyses. In this regard, the
framework of n-point spatial correlations offers tremendous promise because of
its scalability (ability to define an infinite number of microstructural features) and
organization (value of n can start with one and increase).
The third step in the workflow focuses on reducing the dimensionality of
microstructure representation using data science approaches. Some of the established dimensionality reduction techniques include principal component analysis,
factor analysis, projection pursuit, and independent component analysis, among
others. These methods are designed to reduce dataset dimensions, while losing
only the smallest amounts of information. The use of dimensionality reduction
leads to savings in both computational time and storage and identification of salient
features that can be used to establish models. For example, in prior work, PCA has
proven to be remarkably efficient in producing high-value, low-order representation
of microstructures that are ideally suited to establishing P-S-P linkages in a broad
variety of material systems.
The last step of the workflow focuses on establishing and validating a reliable
and robust process-structure (P-S) or structure-property (S-P) linkage. This step
typically involves an iterative process of model selection. The first part of this
step requires establishing a model using a variety of machine learning techniques
ranging from simple regression to sophisticated M5 model trees and support
vector machines. It is important to recognize that the models developed are indeed
dependent on the available data. Therefore, the model itself can change as one adds
more data. Validation of the model established in this step is typically performed
using accuracy estimation methods. Cross-validation has been found to be quite
255
other features of interest in simulated data is trivial in most cases, experimental data
often requires segmentation of images to properly identify a given feature of interest.
As needed, one might set criterion to eliminate spurious or questionable data (e.g.,
the data that does not conform to known physics). In this step, the inputs (process
parameters) are also clearly associated with the outputs (microstructure data).
It is acknowledged that the robustness of this preprocessing step is a direct
function of the uncertainties introduced during data collection, experimentation, and
data processing techniques. For example, measurement uncertainty is immediately
introduced with data collection due to instrument accuracy and human error.
Additional uncertainty is introduced as these data are processed into human and
computer interpretable forms using various data processing algorithms. For example, the uncertainty associated with imaging and image processing of microscopic
data has been studied and published for a variety of techniques. Uncertainty is
introduced at each stage of the processes depending on the methods and algorithms
employed. While understanding these uncertainties is important to the success of
the framework proposed here, there is much ongoing research in this area, and this
topic is not addressed in detail here but in another chapter in this volume.
In the second step, microstructures are quantified to obtain salient statistical
measures of microstructures. In a data science approach, it is desirable to capture
a very large set of measures at this stage. Consequently, it is preferable to adopt a
microstructure quantification framework that allows one to increase systematically
the numbers of potential features included in the analyses. In this regard, the
framework of n-point spatial correlations offers tremendous promise because of
its scalability (ability to define an infinite number of microstructural features) and
organization (value of n can start with one and increase).
The third step in the workflow focuses on reducing the dimensionality of
microstructure representation using data science approaches. Some of the established dimensionality reduction techniques include principal component analysis,
factor analysis, projection pursuit, and independent component analysis, among
others. These methods are designed to reduce dataset dimensions, while losing
only the smallest amounts of information. The use of dimensionality reduction
leads to savings in both computational time and storage and identification of salient
features that can be used to establish models. For example, in prior work, PCA has
proven to be remarkably efficient in producing high-value, low-order representation
of microstructures that are ideally suited to establishing P-S-P linkages in a broad
variety of material systems.
The last step of the workflow focuses on establishing and validating a reliable
and robust process-structure (P-S) or structure-property (S-P) linkage. This step
typically involves an iterative process of model selection. The first part of this
step requires establishing a model using a variety of machine learning techniques
ranging from simple regression to sophisticated M5 model trees and support
vector machines. It is important to recognize that the models developed are indeed
dependent on the available data. Therefore, the model itself can change as one adds
more data. Validation of the model established in this step is typically performed
using accuracy estimation methods. Cross-validation has been found to be quite
