116
F. Rathke et al.
3-D. In contrast to 2-D scans, the labeling of OCT volumes is very time consuming,
hence our dataset only consisted of 35 samples. Thus we were left with less data
points to train a shape model of much higher dimension. Consequently, we observed
a reduced ability of p(b) respectively q b (b) to generalize well to unseen scans. We
tackled this problem by suppressing the connectivity between different B-scans inside
the volume, which corresponds to a block-diagonal covariance matrix , where each
block is obtained separately using PPCA. This significantly reduced the amount of
parameters that had to be determined, and improved accuracy significantly.
We tested our approach on our in-house dataset as well as the two datasets published by Tian et al. For the latter two we had to deal with the problem, that no position
inside the volume was given for the B-Scans. Since this information in necessary to
pick the correct shape prior, we used the following approach: We run the model for
each region 1–17 (see Fig. 5.4b for the distribution of regions) and then picked the
region (a) with the lowest error (supervised) and (b) with the highest model likelihood (unsupervised). This gave an upper and lower bound on the true error, and we
averaged over these two to obtain the final result.
In Table 5.3 we also report results from Tian et al., which we could outperform
in both datasets, as well as those of Duan et al. [13] for the first Tian dataset, which
are very good as well. Unfortunately, they only provided specific results for two
boundaries besides the average error, so its difficult to fully grasp their strong and
weak points. For all 3-D datasets results are worst for the retinopathy dataset due
to its pathological nature. Still, the segmentation of our approach is fairly accurate,
with an average error of around 1 pixel.
5.3.2 Pathology Detection
5.3.2.1 Model Likelihoods
A key property of our model is the inference of full probability distributions over
segmentations q c and q b , instead of only modes thereof. Figure 5.6 shows boxblots
of four terms (b–e) of the objective function J (q b , q c ) and compares them to the
unsigned error (a), broken down for the healthy scans as well as the different stages
of glaucoma. Singleton entropy (b) and mutual information (c) are the two summands
of the negative entropy of q c , see (5.33). The data (d) and shape (e) terms represent
the first two summands of J (q b , q c ) (5.12), presented in details in the Sections “First
Summand log P(c|y) of J (q b , q c )” and “Second Summand log P(c|b) of J (q b , q c )”
in Appendix.
The shape term, which measures how much the data-driven distribution q c differs
from the shape-driven expectation E q b [log p(c|b)], is highly discriminative between
healthy and pathological scans. The data term on the other hand measures how well
the appearance terms fit the actual segmentation and correlates well with the unsigned
error.
F. Rathke et al.
3-D. In contrast to 2-D scans, the labeling of OCT volumes is very time consuming,
hence our dataset only consisted of 35 samples. Thus we were left with less data
points to train a shape model of much higher dimension. Consequently, we observed
a reduced ability of p(b) respectively q b (b) to generalize well to unseen scans. We
tackled this problem by suppressing the connectivity between different B-scans inside
the volume, which corresponds to a block-diagonal covariance matrix , where each
block is obtained separately using PPCA. This significantly reduced the amount of
parameters that had to be determined, and improved accuracy significantly.
We tested our approach on our in-house dataset as well as the two datasets published by Tian et al. For the latter two we had to deal with the problem, that no position
inside the volume was given for the B-Scans. Since this information in necessary to
pick the correct shape prior, we used the following approach: We run the model for
each region 1–17 (see Fig. 5.4b for the distribution of regions) and then picked the
region (a) with the lowest error (supervised) and (b) with the highest model likelihood (unsupervised). This gave an upper and lower bound on the true error, and we
averaged over these two to obtain the final result.
In Table 5.3 we also report results from Tian et al., which we could outperform
in both datasets, as well as those of Duan et al. [13] for the first Tian dataset, which
are very good as well. Unfortunately, they only provided specific results for two
boundaries besides the average error, so its difficult to fully grasp their strong and
weak points. For all 3-D datasets results are worst for the retinopathy dataset due
to its pathological nature. Still, the segmentation of our approach is fairly accurate,
with an average error of around 1 pixel.
5.3.2 Pathology Detection
5.3.2.1 Model Likelihoods
A key property of our model is the inference of full probability distributions over
segmentations q c and q b , instead of only modes thereof. Figure 5.6 shows boxblots
of four terms (b–e) of the objective function J (q b , q c ) and compares them to the
unsigned error (a), broken down for the healthy scans as well as the different stages
of glaucoma. Singleton entropy (b) and mutual information (c) are the two summands
of the negative entropy of q c , see (5.33). The data (d) and shape (e) terms represent
the first two summands of J (q b , q c ) (5.12), presented in details in the Sections “First
Summand log P(c|y) of J (q b , q c )” and “Second Summand log P(c|b) of J (q b , q c )”
in Appendix.
The shape term, which measures how much the data-driven distribution q c differs
from the shape-driven expectation E q b [log p(c|b)], is highly discriminative between
healthy and pathological scans. The data term on the other hand measures how well
the appearance terms fit the actual segmentation and correlates well with the unsigned
error.
