Hybrid-View Programming of Nuclear Fusion Simulation Code in XcalableMP
187
1 double f[X][Y]; /* Electric field data */
2 double p[N]; /* Particle data */
3 double send[N], recv[N]:[*];
4 #pragma xmp align f[i][j] with T(i, j)
5 #pragma xmp shadow f[1:1][1:1]
6
7 for(t=0; t 8
/* Calculate the grid−related work */
9
#pragma xmp reflect(f)
10
/* Calculate the particle−related work */
11
/* Pack the communication elements from array "p" to array "send" */
12
/* Calculate the destination process "pe" and communication size "icount" */
13
recv[0:icount]:[pe] = send[0:icount];
14
xmp_sync_all(NULL); /* Synchronization */
15 }
Fig. 4 Example showing the implementation of a gyrokinetic PIC simulation with XMP
distribution and each block has a sleeve area, which is used to calculate the field with
the nearby grid points based on the shadow directive. The particle movement is
represented by the coarray notation where the communication elements are packed
in the array send. Based on the above, we describe the two implementations of
GTC-P using XMP. First, we implement the XMP-localview version using coarray
communication, which is equivalent to using MPI point-to-point communication
with the exception of MPI collective communication (as shown below). Next, the
XMP-hybridview version is implemented by describing the fields using a distributed
array with the reflect directive for overlapped sleeve area communication and
the distributed data in the global-view programming model, as well as using the
coarray notation to move the particle data. In addition, we use the bcast and
reduction directives instead of MPI collective communication (MPI_Bcast
and MPI_Allreduce) in both versions.
3.2 Implementation Based on the XMP-Localview Model:
XMP-localview
In GTC-P, the communication processes required to move particles between
grids and to exchange grid points are represented by MPI_Sendrecv or
MPI_Isend/Irecv, where most of the communication is performed between
adjacent processes in one dimension. GTC-P has the steady state exchange
of particles between neighboring subdomains. Because the number of particles
changes dynamically, this implementation uses the coarray notation in the localview programming model.
187
1 double f[X][Y]; /* Electric field data */
2 double p[N]; /* Particle data */
3 double send[N], recv[N]:[*];
4 #pragma xmp align f[i][j] with T(i, j)
5 #pragma xmp shadow f[1:1][1:1]
6
7 for(t=0; t 8
/* Calculate the grid−related work */
9
#pragma xmp reflect(f)
10
/* Calculate the particle−related work */
11
/* Pack the communication elements from array "p" to array "send" */
12
/* Calculate the destination process "pe" and communication size "icount" */
13
recv[0:icount]:[pe] = send[0:icount];
14
xmp_sync_all(NULL); /* Synchronization */
15 }
Fig. 4 Example showing the implementation of a gyrokinetic PIC simulation with XMP
distribution and each block has a sleeve area, which is used to calculate the field with
the nearby grid points based on the shadow directive. The particle movement is
represented by the coarray notation where the communication elements are packed
in the array send. Based on the above, we describe the two implementations of
GTC-P using XMP. First, we implement the XMP-localview version using coarray
communication, which is equivalent to using MPI point-to-point communication
with the exception of MPI collective communication (as shown below). Next, the
XMP-hybridview version is implemented by describing the fields using a distributed
array with the reflect directive for overlapped sleeve area communication and
the distributed data in the global-view programming model, as well as using the
coarray notation to move the particle data. In addition, we use the bcast and
reduction directives instead of MPI collective communication (MPI_Bcast
and MPI_Allreduce) in both versions.
3.2 Implementation Based on the XMP-Localview Model:
XMP-localview
In GTC-P, the communication processes required to move particles between
grids and to exchange grid points are represented by MPI_Sendrecv or
MPI_Isend/Irecv, where most of the communication is performed between
adjacent processes in one dimension. GTC-P has the steady state exchange
of particles between neighboring subdomains. Because the number of particles
changes dynamically, this implementation uses the coarray notation in the localview programming model.
