182
K. Tsugane et al.
1 Introduction
In XMP, the global-view model allows programmers to define global arrays, which
are distributed to processors by adding the directives. Some typical communication
patterns are supported by directives, such as data exchange between neighbor
processors in stencil computations. In contrast to the global-view model, the localview model describes remote memory access using the node (processor) index. This
operation is implemented as one-sided communication. XMP employs the coarray
concept from Coarray Fortran as a local-view programming model. A coarray is
a distributed data object, which is indexed by the coarray dimension that maps
indices to processors. In XMP, the coarray is defined in C as well as Fortran. In the
local-view model, a thread on each processor executes its own local computations
independently with remote memory access to data located in different processors
by coarray access. The local-view model requires that programmers define their
algorithms by explicitly decomposing the data structures and controlling the flow in
each processor. The data view is similar to that in MPI, but coarray remote access
provides a more intuitive view of accessing the data in different processors, thereby
increasing productivity.
In this chapter, 1 we consider a hybrid-view programming approach, which
combines the global-view and local-view models in XMP according to the characteristics of the distributed data structure of the target application. The global-view
model allows programmers to express regular parallel computations such as domain
decomposition with stencil computation in a highly intuitive manner simply by
adding directives to a serial version of code. However, it is difficult to describe
parallel programs in the global-view model when more irregular communication
patterns and complex load balancing are required on the processing. Thus, localview programming is necessary in these situations.
We apply this hybrid-view programming for Gyrokinetic Toroidal Code -
Princeton (GTC-P)[4], which is a large-scale plasma turbulence code that can
be applied at the International Thermonuclear Experimental Reactor (ITER [13])
scale and beyond for next-generation nuclear fusion simulation. The GTC-P is an
improved version of the original GTC[2] and it is a type of gyrokinetic Particlein-Cell (PIC) code with two basic data arrays: global grid data that corresponds to
the physical problem space and particle data that corresponds to particles moving
around the grid space. The original GTC-P was written in C as a form of hybrid
programming with OpenMP and MPI. In this code, the grid data and particle data
are mapped onto MPI processes and exchanged. As found with most codes of this
type, it is difficult to manage complex data distributions and communication for both
grid data and particle data during code development. Furthermore, to simulate the
1 The original version of this chapter was published in: Keisuke Tsugane, Taisuke Boku, Hitoshi
Murai, Mitsuhisa Sato, William M. Tang, Bei Wang, “Hybrid-view programming of nuclear fusion
simulation code in the PGAS parallel programming language XcalableMP.” Parallel Computing
57: 37–51 (2016).
K. Tsugane et al.
1 Introduction
In XMP, the global-view model allows programmers to define global arrays, which
are distributed to processors by adding the directives. Some typical communication
patterns are supported by directives, such as data exchange between neighbor
processors in stencil computations. In contrast to the global-view model, the localview model describes remote memory access using the node (processor) index. This
operation is implemented as one-sided communication. XMP employs the coarray
concept from Coarray Fortran as a local-view programming model. A coarray is
a distributed data object, which is indexed by the coarray dimension that maps
indices to processors. In XMP, the coarray is defined in C as well as Fortran. In the
local-view model, a thread on each processor executes its own local computations
independently with remote memory access to data located in different processors
by coarray access. The local-view model requires that programmers define their
algorithms by explicitly decomposing the data structures and controlling the flow in
each processor. The data view is similar to that in MPI, but coarray remote access
provides a more intuitive view of accessing the data in different processors, thereby
increasing productivity.
In this chapter, 1 we consider a hybrid-view programming approach, which
combines the global-view and local-view models in XMP according to the characteristics of the distributed data structure of the target application. The global-view
model allows programmers to express regular parallel computations such as domain
decomposition with stencil computation in a highly intuitive manner simply by
adding directives to a serial version of code. However, it is difficult to describe
parallel programs in the global-view model when more irregular communication
patterns and complex load balancing are required on the processing. Thus, localview programming is necessary in these situations.
We apply this hybrid-view programming for Gyrokinetic Toroidal Code -
Princeton (GTC-P)[4], which is a large-scale plasma turbulence code that can
be applied at the International Thermonuclear Experimental Reactor (ITER [13])
scale and beyond for next-generation nuclear fusion simulation. The GTC-P is an
improved version of the original GTC[2] and it is a type of gyrokinetic Particlein-Cell (PIC) code with two basic data arrays: global grid data that corresponds to
the physical problem space and particle data that corresponds to particles moving
around the grid space. The original GTC-P was written in C as a form of hybrid
programming with OpenMP and MPI. In this code, the grid data and particle data
are mapped onto MPI processes and exchanged. As found with most codes of this
type, it is difficult to manage complex data distributions and communication for both
grid data and particle data during code development. Furthermore, to simulate the
1 The original version of this chapter was published in: Keisuke Tsugane, Taisuke Boku, Hitoshi
Murai, Mitsuhisa Sato, William M. Tang, Bei Wang, “Hybrid-view programming of nuclear fusion
simulation code in the PGAS parallel programming language XcalableMP.” Parallel Computing
57: 37–51 (2016).
