XcalableACC: An Integration of XcalableMP and OpenACC
145
1 #pragma acc parallel loop collapse(4) present(v_out, u, v)
2 for(int t=0;t 3
for(int z=0;z 4
for(int y=0;y 5
for(int x=0;x 6
int tt = (t + 1) % NT;
7
v_out[tt][z][y][x].v[0][0][0] = ... ;
Fig. 25 Code modification of WD() in OpenACC
1 #pragma xmp loop (t,z) on t[t][z]
2 #pragma acc parallel loop collapse(4) present(v_out, u, v)
3 for(int t=0;t 4
for(int z=0;z 5
for(int y=0;y 6
for(int x=0;x 7
int tt = t + 1;
8
v_out[tt][z][y][x].v[0][0][0] = ... ;
Fig. 26 Code modification of WD() in XcalableACC
Table 3 SLOC of Lattice QCD implementations
MPI+CUDA
MPI+OpenACC
XcalableACC
SLOC
1091
1002
996
#XcalableMP
–
–
122
#OpenACC
–
26
16
#XcalableACC
–
–
3
#MPI function
39
39
–
parallelization. In addition, there are two lines for modification shown in “b” of
Table 2. It is a very fine modification for OpenACC constraints, which keeps the
semantics of the base code.
As basic information, we count the source lines of codes (SLOC) of each of
the Lattice QCD implementations. Table 3 shows the SLOC excluding comments
and blank lines, as well as the numbers of each directive and MPI functions
included in their SLOC. For reader information, SLOC of the serial version Lattice
QCD code is 842. Table 3 shows that the 122 XMP directives are used in the
XACC implementation, many of which are declarations for function arguments. To
reduce the XMP directives, we are planning to develop a new syntax that combines
declarations with the same attribute into one directive. Figure 27 shows an example
of the new syntax applied to the declarations in Fig. 17. Since the arrays v_out and v
have the same attribute, they can be declared into a single XMP directive. Moreover,
the shadow directive attribute is added to the align directive as its clause. When
applying the new directive to XACC implementation, the number of XMP directives
Précédent

- 152/265

Suivant