140
A. Tabuchi et al.
1x1 2x1 2x2 4x2 4x4 8x4 8x8 16x8 16x16
120
100
80
60
40
20
0
Time per CG solve (msec.)
1.50
1.25
1.00
0.75
0.50
0.25
0.00
Performance ratio for XcalableACC
MPI+CUDA ratio
MPI+OpenACC ratio
Number of nodes (NODE_T x NODE_Z)
MPI+CUDA
MPI+OpenACC
XcalableACC
Fig. 19 Performance results
use the MPI+CUDA implementation type because it provides a balance of versatility
and performance.
Figure 19 shows the performance results that indicate the time required to
solve one CG iteration as well as the performance ratio values that indicate the
comparative performance of XACC and other languages. When the performance
ratio value of a language is greater than 1.00, the performance result of the language
is better than that of XACC. Figure 19 shows that the performance ratio values
of MPI+CUDA are between 1.04 and 1.18, and that those of MPI+OpenACC are
between 0.99 and 1.04. Moreover, Fig. 19 also shows that the performance results
of both MPI+CUDA and MPI+OpenACC become closer to those of XACC as the
number of nodes increases.
5.2 Discussion
To examine the performance levels in detail, we measure the time required for the
halo updating operation for two nodes and more. The halo updating operation
consists of the communication and pack/unpack processes for non-contiguous
regions in the XACC runtime.
While Fig. 20 describes communication time of the halo updating time of Fig 19,
Fig. 21 describes pack/unpack time of it. Figure 20 shows that the communication
performance levels of all implementations are almost the same. However, Fig. 21
shows that the pack/unpack performance levels of MPI+CUDA are better than those
of XACC, and that those of MPI+OpenACC are worse than those of XACC. The
reason for the pack/unpack operation performance level difference is that the XACC
A. Tabuchi et al.
1x1 2x1 2x2 4x2 4x4 8x4 8x8 16x8 16x16
120
100
80
60
40
20
0
Time per CG solve (msec.)
1.50
1.25
1.00
0.75
0.50
0.25
0.00
Performance ratio for XcalableACC
MPI+CUDA ratio
MPI+OpenACC ratio
Number of nodes (NODE_T x NODE_Z)
MPI+CUDA
MPI+OpenACC
XcalableACC
Fig. 19 Performance results
use the MPI+CUDA implementation type because it provides a balance of versatility
and performance.
Figure 19 shows the performance results that indicate the time required to
solve one CG iteration as well as the performance ratio values that indicate the
comparative performance of XACC and other languages. When the performance
ratio value of a language is greater than 1.00, the performance result of the language
is better than that of XACC. Figure 19 shows that the performance ratio values
of MPI+CUDA are between 1.04 and 1.18, and that those of MPI+OpenACC are
between 0.99 and 1.04. Moreover, Fig. 19 also shows that the performance results
of both MPI+CUDA and MPI+OpenACC become closer to those of XACC as the
number of nodes increases.
5.2 Discussion
To examine the performance levels in detail, we measure the time required for the
halo updating operation for two nodes and more. The halo updating operation
consists of the communication and pack/unpack processes for non-contiguous
regions in the XACC runtime.
While Fig. 20 describes communication time of the halo updating time of Fig 19,
Fig. 21 describes pack/unpack time of it. Figure 20 shows that the communication
performance levels of all implementations are almost the same. However, Fig. 21
shows that the pack/unpack performance levels of MPI+CUDA are better than those
of XACC, and that those of MPI+OpenACC are worse than those of XACC. The
reason for the pack/unpack operation performance level difference is that the XACC
