212
A. Kubota et al.
Among the six nested loops, the outer z loop and the inner x loop are parallelized
by XcalableMP and OpenMP. In other words, the nested loops are parallelized
among inter-nodes by the z loop level and also parallelized within each node by
the x loop level. The parallelized kernel loop of three-dimensional atomic image
reconstruction is shown in Fig. 5. Other combinations of parallelizing x, y, and z
loops are also performed and discussed in Sect 4.
In this study, z loop is parallelized by XcalableMP directives and the output
arrays are distributed contiguously, namely BLOCK distribution, among nodes
in z dimension as shown in Fig. 6. All of the input data are read on every
node simultaneously and stored replicated. The data of atomic images, which are
distributed among nodes in z dimension, are calculated on each node, aggregated to
one root node as shown in Fig. 7, and finally written to the output file on the root
node.
4 Performance Evaluation
In this section, we show the performance results of parallel runs of reconstruction of
two-dimensional atomic images parallelized by OpenMP executed on a single node
and reconstruction of three-dimensional atomic images parallelized by XcalableMP
and OpenMP on a multi-node PC cluster. Comparison of XcalableMP and MPI
for multi-node parallelization with respect to the performance and productivity of
programming are also demonstrated.
The PC cluster consist of eight nodes and each node has two sockets of sixcore Intel Xeon X5660 2.8 GHz and 24 GB main memory. The nodes are connected
with InfiniBand DDR (4 Gbps) and Gigabit Ethernet. The size of smart cache on
Xeon X5660 is 12 MB. The program is compiled with XcalableMP 1.2.2 and Intel
Compiler 18.0.1 with -O3 -XHOST optimization option.
4.1 Performance Results of Reconstruction
of Two-Dimensional Atomic Images
The program of reconstruction of two-dimensional atomic images is executed on
a single node of the PC cluster. The input data are obtained in an experiment in
which lead zirconate titanate (PZT) is used as a sample in the experiment and 21
types of incident X-rays are entered to the sample while varying angles ranging
from θ = 1 to 179 ◦ and from φ = 0 to 359 ◦ . The output is an two-dimensional array
[x][y] = [192][192] for a certain z.
In Table 1, the execution time of the original, array reference of trigonometric
function calls, and loop interchange on one node. Execution time and speed-up ratio
A. Kubota et al.
Among the six nested loops, the outer z loop and the inner x loop are parallelized
by XcalableMP and OpenMP. In other words, the nested loops are parallelized
among inter-nodes by the z loop level and also parallelized within each node by
the x loop level. The parallelized kernel loop of three-dimensional atomic image
reconstruction is shown in Fig. 5. Other combinations of parallelizing x, y, and z
loops are also performed and discussed in Sect 4.
In this study, z loop is parallelized by XcalableMP directives and the output
arrays are distributed contiguously, namely BLOCK distribution, among nodes
in z dimension as shown in Fig. 6. All of the input data are read on every
node simultaneously and stored replicated. The data of atomic images, which are
distributed among nodes in z dimension, are calculated on each node, aggregated to
one root node as shown in Fig. 7, and finally written to the output file on the root
node.
4 Performance Evaluation
In this section, we show the performance results of parallel runs of reconstruction of
two-dimensional atomic images parallelized by OpenMP executed on a single node
and reconstruction of three-dimensional atomic images parallelized by XcalableMP
and OpenMP on a multi-node PC cluster. Comparison of XcalableMP and MPI
for multi-node parallelization with respect to the performance and productivity of
programming are also demonstrated.
The PC cluster consist of eight nodes and each node has two sockets of sixcore Intel Xeon X5660 2.8 GHz and 24 GB main memory. The nodes are connected
with InfiniBand DDR (4 Gbps) and Gigabit Ethernet. The size of smart cache on
Xeon X5660 is 12 MB. The program is compiled with XcalableMP 1.2.2 and Intel
Compiler 18.0.1 with -O3 -XHOST optimization option.
4.1 Performance Results of Reconstruction
of Two-Dimensional Atomic Images
The program of reconstruction of two-dimensional atomic images is executed on
a single node of the PC cluster. The input data are obtained in an experiment in
which lead zirconate titanate (PZT) is used as a sample in the experiment and 21
types of incident X-rays are entered to the sample while varying angles ranging
from θ = 1 to 179 ◦ and from φ = 0 to 359 ◦ . The output is an two-dimensional array
[x][y] = [192][192] for a certain z.
In Table 1, the execution time of the original, array reference of trigonometric
function calls, and loop interchange on one node. Execution time and speed-up ratio
