234
X. Qi et al.
2 Related Works
In order to improve the effect and efficiency of algorithm of image translation and style
transfer, researchers have proposed CGAN [2], DCGAN [3], Pix2Pix [4], etc. Among
them, cycle-consistent adversarial network (CycleGAN) [5] has excellent performance
in image-to-image translation. While it has unsatisfactory performance in one-to-many
translation, because it is designed for one-to-one mode.
This paper proposes a one-to-many transfer model, MCGAN, which takes advantage
of both CGAN and CycleGAN. It is discussed in detail in Sect. 3.
3 Multi-conditional Cycle Generative Adversarial Networks
3.1 Algorithm Framework
As shown in Fig. 1, our MCGAN network consists of a pair of generative adversarial
models with a cycle framework, which include five condition branches so as to generate
infrared images of five different periods at once. According to the initial generation
model and the source of the input images, the training of the network can be divided into
two parts. The first part is A → B process. The second part is B → A process. B → A
process converts a fixed-period infrared image to multi-period infrared images.
Fig. 1. The framework of Multi-conditional Cycle Generative Adversarial Networks. X: the
infrared images of a fixed-period. Yi: the real infrared images in different periods of time. Zi:
a vector representing a certain period of time. KZ: a vector representing the period of time of X.
3.2 The Loss Functions
This paper aims to increase both the size and diversity of infrared image datasets. The W
distance [6, 7] represents the minimum cost of moving from one distribution to another,
which can still reflect the distance between two distributions even if the support sets of
them overlap little or none. Hence, W distance is chose to construct the loss function of
MCGAN.
In order to improve the details of the derived images, this paper uses a discriminator
to process the image into blocks [8, 9], which would reduce the overall similarity between
the generated image and the real image. Therefore, we introduce the L1 norm [10] as an
additional penalty term to supervise the global consistency of the image, consequently
X. Qi et al.
2 Related Works
In order to improve the effect and efficiency of algorithm of image translation and style
transfer, researchers have proposed CGAN [2], DCGAN [3], Pix2Pix [4], etc. Among
them, cycle-consistent adversarial network (CycleGAN) [5] has excellent performance
in image-to-image translation. While it has unsatisfactory performance in one-to-many
translation, because it is designed for one-to-one mode.
This paper proposes a one-to-many transfer model, MCGAN, which takes advantage
of both CGAN and CycleGAN. It is discussed in detail in Sect. 3.
3 Multi-conditional Cycle Generative Adversarial Networks
3.1 Algorithm Framework
As shown in Fig. 1, our MCGAN network consists of a pair of generative adversarial
models with a cycle framework, which include five condition branches so as to generate
infrared images of five different periods at once. According to the initial generation
model and the source of the input images, the training of the network can be divided into
two parts. The first part is A → B process. The second part is B → A process. B → A
process converts a fixed-period infrared image to multi-period infrared images.
Fig. 1. The framework of Multi-conditional Cycle Generative Adversarial Networks. X: the
infrared images of a fixed-period. Yi: the real infrared images in different periods of time. Zi:
a vector representing a certain period of time. KZ: a vector representing the period of time of X.
3.2 The Loss Functions
This paper aims to increase both the size and diversity of infrared image datasets. The W
distance [6, 7] represents the minimum cost of moving from one distribution to another,
which can still reflect the distance between two distributions even if the support sets of
them overlap little or none. Hence, W distance is chose to construct the loss function of
MCGAN.
In order to improve the details of the derived images, this paper uses a discriminator
to process the image into blocks [8, 9], which would reduce the overall similarity between
the generated image and the real image. Therefore, we introduce the L1 norm [10] as an
additional penalty term to supervise the global consistency of the image, consequently
