7
Clustering and
Classification
7.1 INTRODUCTION AND OVERVIEW
This section is dedicated to clustering and classification techniques. The first issue to
address is the exact definitions of “clustering” and “classification” and the comparison of the two processes. We start this section with the description of these concepts
and their importance in biomedical signal and image processing. Then, we discuss
several popular clustering and classification techniques including Bayesian methods,
K-means, and neural networks.
7.2 CLUSTERING VERSUS CLASSIFICATION
In classification, one is provided with some examples from two or more groups of
objects. For example, assume that in a study of cardiovascular diseases, features such
as heart rate and cardiac output for a number of healthy persons as well as patients
with some known diseases are available. This means that each example (i.e., a set of
features taken from a person) is labeled either as healthy or as a particular disease.
A classifier is then trained with the “labeled examples” to create a set of rules or
mathematical model that can then look at the features captured from a new person
and label the case as healthy or a particular disease. Since the set of examples provided to the classifier is used to tune and train the classifier, this set is also referred
to as the “training set.” Any other case (i.e., set of features captured from the patient)
that has not been seen by the classifier (i.e., was not included in the training set) will
then be used to “test” the quality of the classifier. In testing the classifier, the features from a person are provided to the trained classifier and the classifier is asked to
predict a label for the case (i.e., predict if the case is healthy or a particular type of
disease). Then, this prediction is compared with the true label of the case and if the
two labels match, the classifier is said to have learned the concept. Often, in order
to have a more reliable testing, a number of testing examples are used. The set of
examples used to test the trained classifier is often referred to as testing set.
Since the labels for the examples in the training set are known, the training process followed to train a classifier is often called “supervised learning.” This means
that a “supervisor” has discovered the labels beforehand and provides these labels
during the training process. The supervisor can also provide some labeled examples
to be treated as the testing set.
Supervised training and classification are heavily used in biomedical sciences. In many applications, physicians can provide biomedical engineers with
125
Clustering and
Classification
7.1 INTRODUCTION AND OVERVIEW
This section is dedicated to clustering and classification techniques. The first issue to
address is the exact definitions of “clustering” and “classification” and the comparison of the two processes. We start this section with the description of these concepts
and their importance in biomedical signal and image processing. Then, we discuss
several popular clustering and classification techniques including Bayesian methods,
K-means, and neural networks.
7.2 CLUSTERING VERSUS CLASSIFICATION
In classification, one is provided with some examples from two or more groups of
objects. For example, assume that in a study of cardiovascular diseases, features such
as heart rate and cardiac output for a number of healthy persons as well as patients
with some known diseases are available. This means that each example (i.e., a set of
features taken from a person) is labeled either as healthy or as a particular disease.
A classifier is then trained with the “labeled examples” to create a set of rules or
mathematical model that can then look at the features captured from a new person
and label the case as healthy or a particular disease. Since the set of examples provided to the classifier is used to tune and train the classifier, this set is also referred
to as the “training set.” Any other case (i.e., set of features captured from the patient)
that has not been seen by the classifier (i.e., was not included in the training set) will
then be used to “test” the quality of the classifier. In testing the classifier, the features from a person are provided to the trained classifier and the classifier is asked to
predict a label for the case (i.e., predict if the case is healthy or a particular type of
disease). Then, this prediction is compared with the true label of the case and if the
two labels match, the classifier is said to have learned the concept. Often, in order
to have a more reliable testing, a number of testing examples are used. The set of
examples used to test the trained classifier is often referred to as testing set.
Since the labels for the examples in the training set are known, the training process followed to train a classifier is often called “supervised learning.” This means
that a “supervisor” has discovered the labels beforehand and provides these labels
during the training process. The supervisor can also provide some labeled examples
to be treated as the testing set.
Supervised training and classification are heavily used in biomedical sciences. In many applications, physicians can provide biomedical engineers with
125
