Hybrid-View Programming of Nuclear Fusion Simulation Code in XcalableMP
201
5 Related Research
GTC or GTC-P have been executed and optimized on some platforms. X. Liao
et al.[7] optimized GTC to use offload-programming model for the Intel Xeon
Phi accelerator, and evaluated the performance on MilkyWay-2 supercomputer. K.
Madduri et al.[8] described the optimization for multi- and many-core systems,
and evaluated on some systems including Graphic Processing Unit (GPU) based
on NVIDIA Fermi architectures. In our study, we focus on the evaluation of not
only the performance but also the productivity for GTC-P.
PIC method is often implemented some PGAS parallel programming languages.
H. Sakagami and T. Mizuno[12] implemented 2D particle code, ESPAC2: 2D
electrostatic plasma, based on PIC method using High Performance Fortran (HPF)
[5] which is directive-based language similar to OpenMP and supports the globalview model. The particle data is distributed into the block, while the electrostatic
field is replicated onto each process. After the distributed particle data is calculated
in each process, the reduction operation is executed to update the particle data of
electrostatic field on each time step. The data distribution is an easy expression
which is annotated by directives in HPF. R. Preissl et al.[11] introduced hybrid
PGAS+OpenMP approach for 3D PIC code, Gyrokinetic Tokamak Simulation
(GTS) which is implemented in MPI+OpenMP. As PGAS parallel programming
language, they used Coarray Fortran. The one-sided communication in Coarray
Fortran is simple and more intuitive notation compared with MPI programming
because it is expressed in the form of array assignment statement. However, the
description of data distribution is same as MPI. To use simple coarray communication and easy data distribution by directives, we consider a hybrid-view approach,
which combines the global-view and local-view models in XMP.
6 Conclusion
In this study, we implemented two versions of GTC-P, a large-scale nuclear fusion
simulation code, using the global-view and local-view programming models in
XMP for parallel programming languages, and we evaluated their performance and
productivity. The first version, XMP-localview, only uses coarray communication
in the local-view programming model, which simply replaces MPI point-to-point
communication, except for collective communication such as MPI_Allreduce. The
second version, XMP-hybridview, uses the distribution of the calculation domain
and the reflect directive in the global-view programming model, as well as
coarray communication for particle motion in the local-view programming model.
Experimental evaluations showed that the XMP-localview implementation obtained
approximately the same performance as MPI, whereas the XMP-hybridview implementation degraded the performance by 20%. In addition, we obtained high
201
5 Related Research
GTC or GTC-P have been executed and optimized on some platforms. X. Liao
et al.[7] optimized GTC to use offload-programming model for the Intel Xeon
Phi accelerator, and evaluated the performance on MilkyWay-2 supercomputer. K.
Madduri et al.[8] described the optimization for multi- and many-core systems,
and evaluated on some systems including Graphic Processing Unit (GPU) based
on NVIDIA Fermi architectures. In our study, we focus on the evaluation of not
only the performance but also the productivity for GTC-P.
PIC method is often implemented some PGAS parallel programming languages.
H. Sakagami and T. Mizuno[12] implemented 2D particle code, ESPAC2: 2D
electrostatic plasma, based on PIC method using High Performance Fortran (HPF)
[5] which is directive-based language similar to OpenMP and supports the globalview model. The particle data is distributed into the block, while the electrostatic
field is replicated onto each process. After the distributed particle data is calculated
in each process, the reduction operation is executed to update the particle data of
electrostatic field on each time step. The data distribution is an easy expression
which is annotated by directives in HPF. R. Preissl et al.[11] introduced hybrid
PGAS+OpenMP approach for 3D PIC code, Gyrokinetic Tokamak Simulation
(GTS) which is implemented in MPI+OpenMP. As PGAS parallel programming
language, they used Coarray Fortran. The one-sided communication in Coarray
Fortran is simple and more intuitive notation compared with MPI programming
because it is expressed in the form of array assignment statement. However, the
description of data distribution is same as MPI. To use simple coarray communication and easy data distribution by directives, we consider a hybrid-view approach,
which combines the global-view and local-view models in XMP.
6 Conclusion
In this study, we implemented two versions of GTC-P, a large-scale nuclear fusion
simulation code, using the global-view and local-view programming models in
XMP for parallel programming languages, and we evaluated their performance and
productivity. The first version, XMP-localview, only uses coarray communication
in the local-view programming model, which simply replaces MPI point-to-point
communication, except for collective communication such as MPI_Allreduce. The
second version, XMP-hybridview, uses the distribution of the calculation domain
and the reflect directive in the global-view programming model, as well as
coarray communication for particle motion in the local-view programming model.
Experimental evaluations showed that the XMP-localview implementation obtained
approximately the same performance as MPI, whereas the XMP-hybridview implementation degraded the performance by 20%. In addition, we obtained high
