Material Agnostic Data-Driven Framework to Develop Structure-Property Linkages
253
Fig. 2 A schematic representation of PCA analysis summarizing the procedural logic of dimensionality reduction based on variance of data
data reduction involve hand-selecting features based on scientific insight. As an
example, one might quantify the polycrystalline microstructure using grain size or
shape distributions while studying yield strength in metals due to Hall-Petch effect
[14]. However, such approaches do not have a common set of low-dimensional
representations that can be applied universally across all material systems.
The recent explosion of dataset size, in terms of both number of records and
number of features/attributes, has triggered the advancement of dimensionality
reduction algorithms [2]. Dimensionality reduction techniques, such as Random
Forest/Ensemble Trees, Principal Component Analysis, and Backward/Forward
Feature Elimination, are extensively employed for image analysis in computer
vision technology as well as other scientific fields. One such data dimensionality reduction technique heavily employed in material science field is principal
component analysis (PCA), extensively used in formulating P-S-P linkages. PCA
is a statistical analysis that transforms the original k coordinates of datasets
into a new set of n coordinates called principal components. As a result of the
transformation, the first principal component has the largest possible variance; each
succeeding component has the highest possible variance under the constraint that
it is orthogonal (i.e., uncorrelated) to the preceding components. In other words,
PCA is a distance-preserving linear map, which involves the derivation of a new
set of orthonormal feature vectors (basis) that are linear combinations of existing
feature vectors, whose optimization is determined by maximizing the variance of
the data along the principal component vectors (see Fig. 2). PCA, as a statistical
modeling tool, is capable of efficient representation of complex, nonlinear data
without the need for a separate identification of the parameter typically required
for conventional modeling. For example, Cord et al. employed PCA to reduce pixel
texture descriptions of metallic surfaces and then employ a supervised learning
approach to distinguish between nominal background and anomalous or defect
structures on the imaged surfaces [15].
Précédent

- 265/416

Suivant