128
F. Rathke et al.
Since ˜
P is independent of q c and depends only on the sufficient statistics of p(b),
we do not have to update it while optimizing J (q b , q c ). Furthermore, since it is
composed of linear combinations of submatrices of K , it can be expressed solely in
terms of W and σ
2 I .
Optimization With Respect to ¯
μ. Optimizing (5.15) with respect to μ yields
min
¯
μ
1
2
K + ˜
P, ¯
μ( ¯
μ − 2μ)
T
+ ˜
p
T
¯
μ,
(5.35)
with the solution ¯
μ = μ − (K + ˜
P)
−1
˜
p. Again ˜
p captures the dependencies of ω 1, j
and k, j , this time on ¯
μ, and is derived below. ˜
P is the same as above. To minimize
(5.35), we use conjugate gradient descent which enables us to calculate ¯
μ using
(K + ˜
P) instead of (K + ˜
P)
−1 .
Derivation of ˜
p. Only considering terms in ω 1, j depending on ¯
μ, we obtain
(ω 1, j ) n ( ¯
μ) = −
1
2(E j|\ j ) 1,1
2(n − μ 1, j )λ
j
1 ¯
μ \ j + λ
j
1 ( ¯
μ \ j ¯
μ
T
\ j − 2μ \ j ¯
μ
T
\ j )(λ
j
1 )
T
and accordingly for (( k, j ) m,n ( ¯
μ). The first term is dependent on n and thus on q c ,
whereas the remaining terms are again independent and q c marginalizes out as above.
Using again ( ˜
λ
j
1 )
T as the extended version of λ
j
1 (see (5.34)) we obtain
−
M
j=1
(q c;1, j ) T ω 1, j ( ¯
μ) +
K
k=2
(q c;k∧k−1, j ) T , , k, j ( ¯
μ)
=
1
2
M
j=1
K
k=1
1
(E j|\ j ) k,k
2
E q c [c k, j ] − μ k, j
( ˜
λ
j
k ) T ¯
μ +
˜
λ
j
k ( ˜
λ
j
k ) T , ¯
μ( ¯
μ − 2μ) T
=
1
2
M
j=1
K
k=1
2 ˜
p T
k, j ¯
μ +
˜
P k, j , ¯
μ( ¯
μ − 2μ) T
= ˜
p T ¯
μ +
1
2
˜
P, ¯
μ( ¯
μ − 2μ) T
.
Since ˜
p is dependent on q c , it is updated at every iteration.
References
1. D. Huang, E.A. Swanson, C.P. Lin, J.S. Schuman, W.G. Stinson, W. Chang, M.R. Hee, T.
Flotte, K. Gregory, C.A. Puliafito et al., Optical coherence tomography. Science 254(5035),
1178–1181 (1991)
2. J. Schuman, C. Puliafito, J. Fujimoto, Optical Coherence Tomography of Ocular Diseases
(Slack Incorporated, Thorofare, 2004)
F. Rathke et al.
Since ˜
P is independent of q c and depends only on the sufficient statistics of p(b),
we do not have to update it while optimizing J (q b , q c ). Furthermore, since it is
composed of linear combinations of submatrices of K , it can be expressed solely in
terms of W and σ
2 I .
Optimization With Respect to ¯
μ. Optimizing (5.15) with respect to μ yields
min
¯
μ
1
2
K + ˜
P, ¯
μ( ¯
μ − 2μ)
T
+ ˜
p
T
¯
μ,
(5.35)
with the solution ¯
μ = μ − (K + ˜
P)
−1
˜
p. Again ˜
p captures the dependencies of ω 1, j
and k, j , this time on ¯
μ, and is derived below. ˜
P is the same as above. To minimize
(5.35), we use conjugate gradient descent which enables us to calculate ¯
μ using
(K + ˜
P) instead of (K + ˜
P)
−1 .
Derivation of ˜
p. Only considering terms in ω 1, j depending on ¯
μ, we obtain
(ω 1, j ) n ( ¯
μ) = −
1
2(E j|\ j ) 1,1
2(n − μ 1, j )λ
j
1 ¯
μ \ j + λ
j
1 ( ¯
μ \ j ¯
μ
T
\ j − 2μ \ j ¯
μ
T
\ j )(λ
j
1 )
T
and accordingly for (( k, j ) m,n ( ¯
μ). The first term is dependent on n and thus on q c ,
whereas the remaining terms are again independent and q c marginalizes out as above.
Using again ( ˜
λ
j
1 )
T as the extended version of λ
j
1 (see (5.34)) we obtain
−
M
j=1
(q c;1, j ) T ω 1, j ( ¯
μ) +
K
k=2
(q c;k∧k−1, j ) T , , k, j ( ¯
μ)
=
1
2
M
j=1
K
k=1
1
(E j|\ j ) k,k
2
E q c [c k, j ] − μ k, j
( ˜
λ
j
k ) T ¯
μ +
˜
λ
j
k ( ˜
λ
j
k ) T , ¯
μ( ¯
μ − 2μ) T
=
1
2
M
j=1
K
k=1
2 ˜
p T
k, j ¯
μ +
˜
P k, j , ¯
μ( ¯
μ − 2μ) T
= ˜
p T ¯
μ +
1
2
˜
P, ¯
μ( ¯
μ − 2μ) T
.
Since ˜
p is dependent on q c , it is updated at every iteration.
References
1. D. Huang, E.A. Swanson, C.P. Lin, J.S. Schuman, W.G. Stinson, W. Chang, M.R. Hee, T.
Flotte, K. Gregory, C.A. Puliafito et al., Optical coherence tomography. Science 254(5035),
1178–1181 (1991)
2. J. Schuman, C. Puliafito, J. Fujimoto, Optical Coherence Tomography of Ocular Diseases
(Slack Incorporated, Thorofare, 2004)
