142
A. Tabuchi et al.
1.50
1.25
1.00
0.75
0.50
0.25
0.00
Performance ratio for XcalableACC
60
50
40
30
20
10
0
Time per CG solve (msec.)
MPI+CUDA
MPI+OpenACC
XcalableACC
MPI+CUDA ratio
MPI+OpenACC ratio
2x1 2x2 4x2 4x4 8x4 8x8 16x8 16x16
Number of nodes (NODE_T x NODE_Z)
Fig. 22 Time excluding halo updating time
Figure 22 shows the overall time excluding the halo updating time, where
performance levels of MPI+CUDA are the best, and those of XACC are almost
the same as those of MPI+OpenACC. The reason for the difference is due to how
to use GPU threads. In the CUDA implementation, we assign loop iterations to
GPU threads in a cyclic-manner manually. In contrast, in the OpenACC and XACC
implementations, how to assign GPU threads is an implementation dependent on
an OpenACC compiler. In the Omni OpenACC compiler, initially loop iterations
are assigned to GPU threads by a gang (threadblock) in a block manner, and then
are also assigned to them by a vector (thread) in a cyclic manner. With the gang
clause with the static argument proposed in the OpenACC specification version 2.0,
programmers can determine how to use GPU threads to some extent, but the Omni
OpenACC compiler does not yet support it.
Figure 23 shows the ratio of the halo updating time to overall time. As can be
seen, as the number of nodes increases, the ratio increases as well. Therefore,
when a large number of nodes are used, there is little difference in performance
level of Fig. 19 among the three implementations. The reason why the ratio of
MPI+CUDA is slightly larger than those of the others is that the time excluding
the halo communication of MPI+CUDA in Fig. 20 is relatively small.
Précédent

- 149/265

Suivant