Hybrid-View Programming of Nuclear Fusion Simulation Code in XcalableMP
189
1 double Xsendr[nloc_over],Xrecvl[nloc_over]:[*];
2
3 for(i=0;i 4 Xsendr[i]=phitmp[i*(mzeta+1)+mzeta];
5
6 Xrecvl[0:nloc_over]:[right_pe]=Xsendr[0:nloc_over];
7 xmp_sync_image(right_pe, NULL);
Fig. 8 Exchange of grid points using the coarray notation in GTC-P
3.3 Implementation Based on the XMP-Hybridview Model:
XMP-Hybridview
In the XMP-hybridview implementation, all of the space grid data are denoted by
a global-view model with compile-time mapping and the sleeve data are exchanged
by XMP directives, whereas the particle data movements are denoted by a localview model with the coarray notation, as shown Fig. 6. It is necessary to represent
an unequal block size for domain decomposition in the radial dimension. Because
this dimension’s space grid is denoted in the global-view model, we apply the
gblock notation to represent it correctly in the same manner as the original MPI
implementation. The gblock notation can control the variable block size of each
domain on the mapped space position. This feature is especially important for
porting GTC-P onto XMP with a global-view model. Figure 10 shows an example of
the GTC-P implementation with the XMP global-view programming model using
gblock. The 11th line of this example denotes the block size distribution in the
radial dimension. Because of describing the data distribution by global-view model,
we can describe the loop distribution only to insert loop directive onto the serial
code that is from the 28th to 30th lines of this example. In addition, OpenMP
directives can be combined with XMP such as the 27th line.
The calculation of the grid-related works, such as the deposit of the charge
from particles onto the grid using a four-point approximation of grid points, the
computation of an electronic field, and the interpolation of the electronic field onto
particles, are similar to four-point stencil calculation on the poloidal plain. In these
codes, we can describe the loop parallelization by inserting loop directive onto the
serial version. Appropriate directives are used for each dimension of the distributed
array in XMP, and we further synchronize the sleeve data that overlap at each end of
the distributed domain, which we can describe simply using the reflect directive.
Figure 9 shows an example of the reflect directive, which is the same as the
communication described in Figs. 7 and 8. Thus, we can describe it using a directive
on one line, which is much simpler compared with the MPI notation in Figs. 7 and 8.
When the width clause is specified, it can be designated as part of the sleeve
elements and the periodic is used to update the sleeve area of the global lower
(upper) bound based on the global upper (lower) bound (Fig. 10).
Précédent

- 194/265

Suivant