10 Layer Segmentation and Analysis for Retina with Diseases
269
method searches for other grey levels in the 13 directions (mentioned above) across
multiple planes and constructs 13 GLCMs. Here, the GLCMs in the 13 directions of
every 5 × 5 × 5 block are constructed. Then, the following four features are calculated: (1) the contrast, which measures the local contrast of the volumetric image and
is expected to be higher when a large grey-level difference occurs more frequently;
(2) the correlation, which provides a correlation between the two voxels in a voxel
pair and is expected to be higher when the grey levels of a voxel pair are more correlated; (3) the energy, which measures the number of repeated voxel pairs and is
expected to be higher if the occurrence of repeated voxel pairs is higher; and (4)
the homogeneity, which measures the local homogeneity of a voxel pair and will be
larger when the grey levels of each voxel pair are more similar.
Based on above definitions, we have a total of 57 features extracted for each voxel
in the VOIs. To reduce the dimensionality of the feature vector and describe the
inter-correlated quantitative dependence of the features, a feature selection procedure
based on the PCA is performed. In our experiments, the first 10 principle components
are selected as the new features; they represent more than 90% of the information in
the original features.
10.4.2.4 Adaboost Algorithm and the Under-Sampling Based
Integrated Classifier
In this study, the number of non-disrupted samples in the EZ region is far greater
than the number of disrupted ones. The disrupted EZ samples and non-disrupted
EZ samples belong to the minority and majority classes, respectively. This is a typical imbalanced classification problem, which means the class distribution is highly
skewed. Most traditional single classifiers, such as the support vector machine, the
k-nearest neighbour classifier, quadratic discriminate analysis, and the decision tree
classifier, tend to show a strong bias towards the majority class and do not work
well for this type of problem because they aim to maximize the overall accuracy.
The Adaboost algorithm based integrated classifier [70–72] is one solution to overcome this problem at the algorithm level; it integrates multiple weak classifiers into a
strong classifier and is therefore more sensitive to the minority. Hence, the Adaboost
algorithm is adopted in this study.
To further improve the classification performance at the data level, the training
datasets are balanced by under-sampling majority samples. In the training step, the
Adaboost algorithm-based classifier model is calculated according to leave-oneout cross-validation, using all the disrupted samples and an equivalent number of
randomly selected non-disrupted samples. In the testing stage, each voxel in the
VOIs is classified as disrupted or non-disrupted using the trained Adaboost model.
269
method searches for other grey levels in the 13 directions (mentioned above) across
multiple planes and constructs 13 GLCMs. Here, the GLCMs in the 13 directions of
every 5 × 5 × 5 block are constructed. Then, the following four features are calculated: (1) the contrast, which measures the local contrast of the volumetric image and
is expected to be higher when a large grey-level difference occurs more frequently;
(2) the correlation, which provides a correlation between the two voxels in a voxel
pair and is expected to be higher when the grey levels of a voxel pair are more correlated; (3) the energy, which measures the number of repeated voxel pairs and is
expected to be higher if the occurrence of repeated voxel pairs is higher; and (4)
the homogeneity, which measures the local homogeneity of a voxel pair and will be
larger when the grey levels of each voxel pair are more similar.
Based on above definitions, we have a total of 57 features extracted for each voxel
in the VOIs. To reduce the dimensionality of the feature vector and describe the
inter-correlated quantitative dependence of the features, a feature selection procedure
based on the PCA is performed. In our experiments, the first 10 principle components
are selected as the new features; they represent more than 90% of the information in
the original features.
10.4.2.4 Adaboost Algorithm and the Under-Sampling Based
Integrated Classifier
In this study, the number of non-disrupted samples in the EZ region is far greater
than the number of disrupted ones. The disrupted EZ samples and non-disrupted
EZ samples belong to the minority and majority classes, respectively. This is a typical imbalanced classification problem, which means the class distribution is highly
skewed. Most traditional single classifiers, such as the support vector machine, the
k-nearest neighbour classifier, quadratic discriminate analysis, and the decision tree
classifier, tend to show a strong bias towards the majority class and do not work
well for this type of problem because they aim to maximize the overall accuracy.
The Adaboost algorithm based integrated classifier [70–72] is one solution to overcome this problem at the algorithm level; it integrates multiple weak classifiers into a
strong classifier and is therefore more sensitive to the minority. Hence, the Adaboost
algorithm is adopted in this study.
To further improve the classification performance at the data level, the training
datasets are balanced by under-sampling majority samples. In the training step, the
Adaboost algorithm-based classifier model is calculated according to leave-oneout cross-validation, using all the disrupted samples and an equivalent number of
randomly selected non-disrupted samples. In the testing stage, each voxel in the
VOIs is classified as disrupted or non-disrupted using the trained Adaboost model.
