5 Machine Learning for IoT
311
Fig. 5.66 An example of K-means clustering algorithm
kilograms and height in centimeters. Then, there is a weighting bias on height versus
weight, which can be solved by a data preparation method called normalization.
Training the K-means algorithm is based on iterative refinement and centroid
calculation that consists of the following steps (see Fig. 5.66):
1. Specify the number of clusters (K) and then randomly select a centroid for each
cluster.
2. Assign/update each data point to the closest cluster (based on the distance of the
data point and centroids).
3. Calculate/update centroid of the cluster. The centroid of each cluster is the
average of all data points in the cluster.
4. Iteratively execute steps 2 and 3. Eventually, the algorithm is terminated when
the change of objective function is below a threshold value.
Note that the accuracy of clustering in the presence of variations depends on the
selection of an appropriate number of clusters k. Researchers usually use the elbow
method for estimating k to make a tradeoff between the accuracy of the clustering
method and the number of clusters.
5.7.2 Hierarchical Clustering
The method of hierarchical clustering produces a specific number of overlapping
clusters of various sizes across a tree that creates a hierarchical system of classification. This clustering technique can be accomplished using a variety of methods, with
the most widely used methods being divisive approach and agglomerative approach.
An agglomerative method works from the bottom-up and it consists of the following
steps:
Précédent

- 317/647

Suivant