Data Structures and Workflows for ICME
43
Fig. 14 The strain components from the final time step of the DEFORM ® simulation. The
underlying mesh has been emphasized to better show the quadrilateral geometry
medoid is a representative datum for that cluster. This approach is reminiscent of
the classic k-means algorithm, in which data are placed in clusters with the closets
cluster mean [66]. For a set of d-dimensional data points X = {x i } and k clusters
C = {c k }, k-medoids attempts to solve the following minimization:
min
C
k
i=1
x i ∈c k
d (x i , m k )
where m k is the medoid of cluster c k and d(a, b) is some distance metric between
points a and b. A benefit of k-medoids is the ability to customize the choice
of metric, which is useful for problems where the standard l 2 norm may be
inappropriate. This optimization is computationally intensive; however, several
heuristic algorithms exist that perform well in practice. We utilize the partition
around medoids algorithm, which iteratively minimizes the total distances within
each cluster by recursively checking medoid candidates for a given partition,
reassigning points to new clusters as medoids are moved [65]. k-medoids has the
advantage of being unsupervised: the clustering model does not require training
data other than the input. However, the choice of k is problem dependent, and
imprudent choices of k may lead to spurious results. Various quality metrics exist
43
Fig. 14 The strain components from the final time step of the DEFORM ® simulation. The
underlying mesh has been emphasized to better show the quadrilateral geometry
medoid is a representative datum for that cluster. This approach is reminiscent of
the classic k-means algorithm, in which data are placed in clusters with the closets
cluster mean [66]. For a set of d-dimensional data points X = {x i } and k clusters
C = {c k }, k-medoids attempts to solve the following minimization:
min
C
k
i=1
x i ∈c k
d (x i , m k )
where m k is the medoid of cluster c k and d(a, b) is some distance metric between
points a and b. A benefit of k-medoids is the ability to customize the choice
of metric, which is useful for problems where the standard l 2 norm may be
inappropriate. This optimization is computationally intensive; however, several
heuristic algorithms exist that perform well in practice. We utilize the partition
around medoids algorithm, which iteratively minimizes the total distances within
each cluster by recursively checking medoid candidates for a given partition,
reassigning points to new clusters as medoids are moved [65]. k-medoids has the
advantage of being unsupervised: the clustering model does not require training
data other than the input. However, the choice of k is problem dependent, and
imprudent choices of k may lead to spurious results. Various quality metrics exist
