Hybrid-View Programming of Nuclear Fusion Simulation Code in XcalableMP
185
Fig. 2 Conceptual image of the three-dimensional torus space in GTC-P[10]
as the poloidal plane. GTC-P is a modified version of GTC, where there are several
differences in the parallelization scheme. Moreover, GTC-P is implemented as two
versions in the C and Fortran languages, whereas the original GTC is coded in
Fortran. In this study, we focus on the C implementation of GTC-P.
GTC parallelizes the problem according to three levels. Processing on the
space grid domain in the toroidal direction and the processing of particles in each
domain are mapped onto MPI processes. Also, the grid-related calculation and
particles in each distributed domain are further subdivided using OpenMP for each
process. GTC-P has four levels of parallelism with additional parallelism in the
radial direction. The total number of MPI processes that need to be executed is
N t × N r × N rp , where N t is the number of domains decomposed in the toroidal
direction, N r is the number of domains decomposed in the radial direction, and N rp
is the number of particles decomposed in each of the distributed domains.
There is a difference in the number of grid points on the poloidal plane, as
demonstrated in Fig. 3 (left). The toroidal domain can be distributed with equally
sized intervals, but the radial domain cannot be distributed with equally sized
intervals due to the large difference in the domain size depending on its position
in space. Therefore, in order to align as much as possible the number of grid points
to be mapped on each process, the outer area of the radial domain is distributed as
short radial interval and its inner area is distributed as long radial interval, such as
Fig. 3 (right).
GTC-P has mainly six computational kernels. The charge kernel deposits the
charge from particles onto the grid using the four-point approximation of nearby
185
Fig. 2 Conceptual image of the three-dimensional torus space in GTC-P[10]
as the poloidal plane. GTC-P is a modified version of GTC, where there are several
differences in the parallelization scheme. Moreover, GTC-P is implemented as two
versions in the C and Fortran languages, whereas the original GTC is coded in
Fortran. In this study, we focus on the C implementation of GTC-P.
GTC parallelizes the problem according to three levels. Processing on the
space grid domain in the toroidal direction and the processing of particles in each
domain are mapped onto MPI processes. Also, the grid-related calculation and
particles in each distributed domain are further subdivided using OpenMP for each
process. GTC-P has four levels of parallelism with additional parallelism in the
radial direction. The total number of MPI processes that need to be executed is
N t × N r × N rp , where N t is the number of domains decomposed in the toroidal
direction, N r is the number of domains decomposed in the radial direction, and N rp
is the number of particles decomposed in each of the distributed domains.
There is a difference in the number of grid points on the poloidal plane, as
demonstrated in Fig. 3 (left). The toroidal domain can be distributed with equally
sized intervals, but the radial domain cannot be distributed with equally sized
intervals due to the large difference in the domain size depending on its position
in space. Therefore, in order to align as much as possible the number of grid points
to be mapped on each process, the outer area of the radial domain is distributed as
short radial interval and its inner area is distributed as long radial interval, such as
Fig. 3 (right).
GTC-P has mainly six computational kernels. The charge kernel deposits the
charge from particles onto the grid using the four-point approximation of nearby
