114
F. Rathke et al.
An important parameter during the inference is the variance of p(b k, j |b \ j ), which
balances the influence of appearance and shape. Artificially increasing this parameter
results in broader normal distributions (that is wider stripes in Fig. 5.12b), which
makes q c less dependent of ¯
μ during inference, and q b less sensitive to q c when
estimating ¯
μ.
5.3.1.3 Error Measures and Test Framework
For each boundary as well as the entire scan we computed the unsigned distance
in µm (1 px= 3.87 µm) between estimates ˆ
c k, j = E q c [c k, j ] and manual segmentations s k, j ∈ R, that is
E
k
unsgn =
1
M
M
j=1
| ˆ
c k, j − s k, j |,
E unsgn =
1
K
K
k=1
E
k
unsgn .
For volumes we additionally averaged over all B-Scans in the volume.
Results were obtained via cross-validation: After splitting each data set into a
number of subsets, each subset in turn is used as a test set, while the remaining subsets
are used for training: We used 10-fold cross-validation for the healthy circular scans
and leave-one-out cross-validation for our 3-D dataset, to maximize the number of
training examples in each split. For the glaucoma dataset we trained a model on all
healthy circular scans and similar for the Tian et al. datasets we used all our volumes
for training, without any cross-validation.
5.3.1.4 Performance Evaluation
Results for all datasets are summarized in Table 5.3. Datasets are reported in the same
order as in Table 5.1.
2-D. In general, boundaries 1 and 6–9 turned out to be easier to segment than boundaries 2–5. While boundary 1 has an easily detectable texture, boundaries 6–9 with
their regular shape particularly profit from the shape regularization. Boundaries 2-5
on the other hand pose a harder challenge with their high variability of texture and
shape. Figure 5.5a depicts an example segmentation.
For the pathological scans segmentation performance decreased with the progression of the disease. However, the average error for the first three classes was
still smaller than or equal to one pixel. The decline in performance had several reasons: Since glaucoma is known to cause a thinning of the nerve fiber layer (NFL), the
shape prior trained on healthy scans encounters difficulties adapting to very abnormal
shapes. Furthermore, we observed a reduced scan quality for glaucomatous scans,
also reported by others (e.g. [6, 7]). In the most advanced stage of the disease the
NFL can vanish at some locations. Since this anomaly is not part of the training data,
the model failed in these regions. We discuss possible modifications to overcome this
problem in Sect. 5.4. The right panel in Fig. 5.5 shows an example of a PGA-type
scan and its segmentation.
F. Rathke et al.
An important parameter during the inference is the variance of p(b k, j |b \ j ), which
balances the influence of appearance and shape. Artificially increasing this parameter
results in broader normal distributions (that is wider stripes in Fig. 5.12b), which
makes q c less dependent of ¯
μ during inference, and q b less sensitive to q c when
estimating ¯
μ.
5.3.1.3 Error Measures and Test Framework
For each boundary as well as the entire scan we computed the unsigned distance
in µm (1 px= 3.87 µm) between estimates ˆ
c k, j = E q c [c k, j ] and manual segmentations s k, j ∈ R, that is
E
k
unsgn =
1
M
M
j=1
| ˆ
c k, j − s k, j |,
E unsgn =
1
K
K
k=1
E
k
unsgn .
For volumes we additionally averaged over all B-Scans in the volume.
Results were obtained via cross-validation: After splitting each data set into a
number of subsets, each subset in turn is used as a test set, while the remaining subsets
are used for training: We used 10-fold cross-validation for the healthy circular scans
and leave-one-out cross-validation for our 3-D dataset, to maximize the number of
training examples in each split. For the glaucoma dataset we trained a model on all
healthy circular scans and similar for the Tian et al. datasets we used all our volumes
for training, without any cross-validation.
5.3.1.4 Performance Evaluation
Results for all datasets are summarized in Table 5.3. Datasets are reported in the same
order as in Table 5.1.
2-D. In general, boundaries 1 and 6–9 turned out to be easier to segment than boundaries 2–5. While boundary 1 has an easily detectable texture, boundaries 6–9 with
their regular shape particularly profit from the shape regularization. Boundaries 2-5
on the other hand pose a harder challenge with their high variability of texture and
shape. Figure 5.5a depicts an example segmentation.
For the pathological scans segmentation performance decreased with the progression of the disease. However, the average error for the first three classes was
still smaller than or equal to one pixel. The decline in performance had several reasons: Since glaucoma is known to cause a thinning of the nerve fiber layer (NFL), the
shape prior trained on healthy scans encounters difficulties adapting to very abnormal
shapes. Furthermore, we observed a reduced scan quality for glaucomatous scans,
also reported by others (e.g. [6, 7]). In the most advanced stage of the disease the
NFL can vanish at some locations. Since this anomaly is not part of the training data,
the model failed in these regions. We discuss possible modifications to overcome this
problem in Sect. 5.4. The right panel in Fig. 5.5 shows an example of a PGA-type
scan and its segmentation.
