240
Y. Zhu et al.
on the combination of the pre-processing network and the generating network, which
holds a dominate position. In Fig. 1, Network 1 is the CRFasRNN network. Network 2
is a generated confrontation network designed for this paper. The Fully connected CRF
used in CRF as RNN calculates the relationship of any two pixels, as shown in Formula
(1).
Fig. 1. Network structure
E(x) =
i
ϕ q (x i ) +
i =i
ϕ p
x i , x j
(1)
The conditional random field probability function P(X=x|I) of the input image X can
be constructed by E(x) as the Formula (2), and the minimized E(x) corresponds to the
maximum posterior probability P(X=x|I), thereby obtaining the optimal segmentation
result, and finally outputting the segmentation map Z.
P(X = x|I) = 1/2 exp (−E(x|I ))
(2)
2.2 Generator Design
Network 2 is Generative Adversarial Network. The four networks designed in this paper
contain a generator G and an discriminator D.
In Fig. 2, GA and GB adopt the same U-NET architecture; the architecture of GC is
similar to that of GA, and the initial input channel is 3; GD first uses two convolutional
layers to extract the early features of visible image and target segmentation map, and
uses NIN [4] after cascading. The network fuses the features and finally derives the
infrared image through the U-NETD network.
Fig. 2. The Design of Generator
The identification models of the four algorithms are consistent, all adopt Patch Priminator, and the input is the whole image. In the process, it will be divided into image
Y. Zhu et al.
on the combination of the pre-processing network and the generating network, which
holds a dominate position. In Fig. 1, Network 1 is the CRFasRNN network. Network 2
is a generated confrontation network designed for this paper. The Fully connected CRF
used in CRF as RNN calculates the relationship of any two pixels, as shown in Formula
(1).
Fig. 1. Network structure
E(x) =
i
ϕ q (x i ) +
i =i
ϕ p
x i , x j
(1)
The conditional random field probability function P(X=x|I) of the input image X can
be constructed by E(x) as the Formula (2), and the minimized E(x) corresponds to the
maximum posterior probability P(X=x|I), thereby obtaining the optimal segmentation
result, and finally outputting the segmentation map Z.
P(X = x|I) = 1/2 exp (−E(x|I ))
(2)
2.2 Generator Design
Network 2 is Generative Adversarial Network. The four networks designed in this paper
contain a generator G and an discriminator D.
In Fig. 2, GA and GB adopt the same U-NET architecture; the architecture of GC is
similar to that of GA, and the initial input channel is 3; GD first uses two convolutional
layers to extract the early features of visible image and target segmentation map, and
uses NIN [4] after cascading. The network fuses the features and finally derives the
infrared image through the U-NETD network.
Fig. 2. The Design of Generator
The identification models of the four algorithms are consistent, all adopt Patch Priminator, and the input is the whole image. In the process, it will be divided into image
