138
A. Tabuchi et al.
1 #pragma xmp reflect(v) width(/periodic/1:1,/periodic/1:1,0,0) orthogonal acc
2 #pragma xmp reflect(u) width(0,/periodic/1:0,/periodic/1:0,0,0) orthogonal acc
3 WD(tmp_v, u, v);
4 #pragma xmp reflect(tmp_v) width(/periodic/1:1,/periodic/1:1,0,0) orthogonal acc
5 WD(v, u, tmp_v);
Fig. 16 Calling Wilson-Dirac operator
1 void WD(Quark_t v_out[NT][NZ][NY][NX], const Gluon_t u[4][NT][NZ][NY][NX],
const Quark_t v[NT][NZ][NY][NX])
2 {
3 #pragma xmp align v_out[i][j][*][*] with t[i][j]
4 #pragma xmp align u[*][i][j][*][*] with t[i][j]
5 #pragma xmp align v[i][j][*][*] with t[i][j]
6 #pragma xmp shadow v_out[1:1][1:1][0][0]
7 #pragma xmp shadow u[0][1:1][1:1][0][0]
8 #pragma xmp shadow v[1:1][1:1][0][0]
9 ...
10 #pragma xmp loop (t,z) on t[t][z]
11 #pragma acc parallel loop collapse(4) present(v_out, u, v)
12 for(int t=0;t
13
for(int z=0;z
14
for(int y=0;y
15
for(int x=0;x
Fig. 17 A portion of Wilson-Dirac operator
in WD(). Moreover, the orthogonal clause is added because diagonal updates of the
arrays are not required in WD().
Figure 17 shows a part of the Wilson-Dirac operator code. All arguments in WD()
are distributed arrays. In XMP and XACC, distributed arrays which are used as
arguments must be redeclared in function to pass their information to a compiler.
Thus, the align and shadow directives are used in WD(). In line 10, the loop
directive parallelizes the outer two loop statements. In line 11, the parallel loop
directive parallelizes all loop statements. In the loop statements, a calculation needs
neighboring and orthogonal elements. Note that while the WD() updates only the
v_out, it only refers the u and v.
Figure 18 shows L2 norm calculation code in the CG method. In line 8,
the reduction clause performs a reduction operation for the variable a in each
accelerator when finishing the next loop statement. The calculated variable a is
located in both host memory and accelerator memory. However, at this point, all
nodes have individual values of a. To obtain the total value of the variable a, the
XMP reduction directive in line 18 also performs a reduction operation among
nodes. Since the total value is used on only host after this function, the XMP
reduction directive does not have acc clause.
A. Tabuchi et al.
1 #pragma xmp reflect(v) width(/periodic/1:1,/periodic/1:1,0,0) orthogonal acc
2 #pragma xmp reflect(u) width(0,/periodic/1:0,/periodic/1:0,0,0) orthogonal acc
3 WD(tmp_v, u, v);
4 #pragma xmp reflect(tmp_v) width(/periodic/1:1,/periodic/1:1,0,0) orthogonal acc
5 WD(v, u, tmp_v);
Fig. 16 Calling Wilson-Dirac operator
1 void WD(Quark_t v_out[NT][NZ][NY][NX], const Gluon_t u[4][NT][NZ][NY][NX],
const Quark_t v[NT][NZ][NY][NX])
2 {
3 #pragma xmp align v_out[i][j][*][*] with t[i][j]
4 #pragma xmp align u[*][i][j][*][*] with t[i][j]
5 #pragma xmp align v[i][j][*][*] with t[i][j]
6 #pragma xmp shadow v_out[1:1][1:1][0][0]
7 #pragma xmp shadow u[0][1:1][1:1][0][0]
8 #pragma xmp shadow v[1:1][1:1][0][0]
9 ...
10 #pragma xmp loop (t,z) on t[t][z]
11 #pragma acc parallel loop collapse(4) present(v_out, u, v)
12 for(int t=0;t
for(int z=0;z
for(int y=0;y
for(int x=0;x
in WD(). Moreover, the orthogonal clause is added because diagonal updates of the
arrays are not required in WD().
Figure 17 shows a part of the Wilson-Dirac operator code. All arguments in WD()
are distributed arrays. In XMP and XACC, distributed arrays which are used as
arguments must be redeclared in function to pass their information to a compiler.
Thus, the align and shadow directives are used in WD(). In line 10, the loop
directive parallelizes the outer two loop statements. In line 11, the parallel loop
directive parallelizes all loop statements. In the loop statements, a calculation needs
neighboring and orthogonal elements. Note that while the WD() updates only the
v_out, it only refers the u and v.
Figure 18 shows L2 norm calculation code in the CG method. In line 8,
the reduction clause performs a reduction operation for the variable a in each
accelerator when finishing the next loop statement. The calculated variable a is
located in both host memory and accelerator memory. However, at this point, all
nodes have individual values of a. To obtain the total value of the variable a, the
XMP reduction directive in line 18 also performs a reduction operation among
nodes. Since the total value is used on only host after this function, the XMP
reduction directive does not have acc clause.
