9 Choroidal OCT Analytics
225
performance comparison among various algorithms. For fair comparison, ideally,
there should be standard datasets, fairly representing possible OCT images, as
well as standardized performance measures. Indeed, desired standardization has
been achieved in certain fields, such as stereo vision [39] and electrocardiogram
(ECG) signal analysis [40]. However, such standardization requires enormous
resources. Pending similar standardization in the study of choroid segmentation,
to improve contextual comprehension, SSIM-based method compared results
generated by the proposed automated algorithm against observer repeatability
figures on the same datasets as those figures implicitly reflect image quality.
Various aspects of the reported literature including SSIM-based method is presented in Fig. 9.9.
To highlight the issues, consider comparing the SSIM-based algorithm [26]
against that reported by dual-gradient-based method by Alonso-Caneiro et
al. [19], which is considered as state-of-the-art prior to SSIM-based algorithm. Specifically, we have rival mean difference (MD) values of −16.63 versus 2.35 µm, and rival standard deviation on difference (SDD) 19.79 versus
15.48 µm (see Fig. 9.7). If one considers the above numerical figures alone,
one would infer the superiority of the latter algorithm. However, such simplistic comparison inherently ignores the fact that the respective datasets on which
the competing algorithms are applied do not necessarily pose similar level of
difficulty in choroidal delineation. Indeed, when presented to the human expert,
those datasets exhibit respective standard deviation (SDD) values of 14.53 and
6.08 µm on observer repeatability, indicating the relatively higher difficulty level
posed by the former (SSIM-based method’s) dataset. In other words, SDD for
SSIM-based algorithm worsens by 36% compared to manual methods, while
such worsening factor is 155% for Alonso-Caneiro et al. Thus, with reference
to manual methods, SSIM-based approach exhibits less relative dispersion.
Now similar comparison is made based on other performance criteria. In terms of
correlation coefficient (CC), SSIM-based method achieves an overall mean CC
(MCC) value of 99.54%, which compares well with the corresponding observer
repeatability value of 99.77%. Further, our algorithmic MCC value of 99.54%
is much higher than the value 93% reported by Hu et al. [22] (refer to Fig. 9.9).
However, as seen in the previous paragraph, the above comparison is not necessarily fair in the absence of the observer repeatability value for the later method.
Turning to Dice coefficient (DC), the results reported by Alonso-Caneiro et al.
[19], appears to improve on the earlier methods (refer to Fig. 9.9), albeit in an
absolute sense, because observer repeatability values are not available. Referring
to Fig. 9.7, SSIM-based method obtained an overall mean DC (MDC) of 94.65%,
which compares well with the corresponding observer repeatability of 96.73%,
but is lower than the algorithmic MDC of 96.7% reported by Alonso-Caneiro
et al. Unfortunately, their observer repeatability value, expected to be higher, is
unavailable, ruling out fair comparison.
B. Performance quotients: Proceeding further, quotient measures that incorporate
observer repeatability are proposed, which not only facilitate performance comparison with reported algorithms, but also to set benchmarks for future research.
Précédent

- 231/387

Suivant