Coarrays in the Context of XcalableMP
117
4.3 Application Program
The Himeno benchmark is a part of the 2D Poisson equation solver using the Jacobi
iteration method [7]. The MPI version of the Himeno benchmark is a strong scaling
program distributing up to three-dimensional nodes. The K computer shown in
Table 2 was used in this evaluation.
4.3.1 Coarray Version of the Himeno Benchmark
For comparison, we prepared the following three versions of Himeno programs.
MPI/original The original MPI version of Himeno benchmark was used as a
two-dimensionally distributed in the y and z axes. The x axis was automatically
SIMD-vectorized by the Fortran compiler. The program executes the computation and communication parts repetitively. The communication part consists of
two steps: z-axis direction communication and y-axis direction communication,
as shown in Fig. 8a. Each communication is written with non-blocking MPI
message passing and completion wait at the end of each step.
MPI/non-blocking The two-step communication was replaced by non-blocking
scrambled communication, as shown in Fig. 8b. With this replacement, the
Fig. 8 Two algorithms of stencil communication in the Himeno benchmark. (a) Original MPI
version. (b) Non-blocking MPI and coarray versions
117
4.3 Application Program
The Himeno benchmark is a part of the 2D Poisson equation solver using the Jacobi
iteration method [7]. The MPI version of the Himeno benchmark is a strong scaling
program distributing up to three-dimensional nodes. The K computer shown in
Table 2 was used in this evaluation.
4.3.1 Coarray Version of the Himeno Benchmark
For comparison, we prepared the following three versions of Himeno programs.
MPI/original The original MPI version of Himeno benchmark was used as a
two-dimensionally distributed in the y and z axes. The x axis was automatically
SIMD-vectorized by the Fortran compiler. The program executes the computation and communication parts repetitively. The communication part consists of
two steps: z-axis direction communication and y-axis direction communication,
as shown in Fig. 8a. Each communication is written with non-blocking MPI
message passing and completion wait at the end of each step.
MPI/non-blocking The two-step communication was replaced by non-blocking
scrambled communication, as shown in Fig. 8b. With this replacement, the
Fig. 8 Two algorithms of stencil communication in the Himeno benchmark. (a) Original MPI
version. (b) Non-blocking MPI and coarray versions
