Clustering and Classification
127
As can be seen, without any calculation and merely from observing the distribution of the points, one can see that we are dealing with four clusters (types) of
persons, which can be defined as follows:
Cluster 1 (Red): People with small height (<170 cm) and small weight (<70 kg)
Cluster 2 (Yellow): People with small height (<170 cm) and high weight (>70 kg)
Cluster 3 (Green): People with high height (>170 m) and average weight
(>70 and <90 kg)
Cluster 4 (Blue): People with high height (>170 m) and high weight (>90 kg)
Looking at health conditions of each cluster, one can see that people belonging
to clusters 1 and 4 are more healthy than people in clusters 2 and 3. This means
that while people in clusters 2 and 3 need to make some changes in their diet to
become healthier, no particular recommendations might be made to the people
in clusters 2 and 3. The clustering results as well as the dietary recommendations
resulting from these clustering processes are rather intuitive but useful. This means
that now the model can be used to analyze the examples that have not been seen
by the clustering method. In other words, now that the clusters are formed, one
can use the resulting rules to define clusters for assigning new examples to one of
the groups. For example, in a classification process based on the resulting groups,
a person whose height and weight are 180 cm and 82 kg, respectively, is mapped
to cluster 4. Based on our intuitive analysis of cluster 4, people assigned to this
cluster are rather healthy and no particular dietary recommendations are needed.
It is important to note that this model was created without using any labeled
examples, i.e., the training samples used to develop the model were not labeled
by a supervisor.
As another biomedical example, consider the discovery of the genes that are
involved in the biological pathways involved in a biological process. Essentially,
this problem is often simplified to the identification and grouping of the genes
that are activated and suppressed at the same time throughout the course of a
biological process. In such studies, having no previous example, one simply finds
the genes having similar patterns and groups them together as a cluster. When
the clusters are formed, the clustering technique can extract some rules or mathematical methods to separate the clusters from each other. This type of learning
in which no labeled examples are given and the method is designed to find the
clusters without any supervision is called “unsupervised learning.”
This chapter first introduces different methods of extracting useful features from
signals and images and then presents some popular techniques for classification and
clustering.
7.3 FEATURE EXTRACTION
The first step in both clustering and classification is finding the best features to represent an example (sample). For instance, in order to distinguish healthy people from
the ones suffering from flu, one needs to collect relevant features such as temperature and blood pressure. The choice of the right features can dramatically affect the
outcome of diagnosis. In general, there are two types of features that are often used
in biomedical sciences that are discussed in the following text.
Précédent

- 154/412

Suivant