XcalableACC: An Integration of XcalableMP and OpenACC
143
100
80
60
40
20
0
Updating halo ratio (%)
2x1 2x2 4x2 4x4 8x4 8x8 16x8 16x16
Number of nodes (NODE_T x NODE_Z)
MPI+CUDA
MPI+OpenACC
XcalableACC
Fig. 23 Updating halo ratio
6 Productivity Improvement
6.1 Requirement for Productive Parallel Language
In Sect. 5, we developed three Lattice QCD codes using MPI+CUDA,
MPI+OpenACC, and XACC. Figure 24 shows our procedure for developing each
code where we first develop the code for an accelerator from the serial code, and
then extend it to handle an accelerated cluster.
To parallelize the serial code for an accelerator using CUDA requires large
code changes (“a” in Fig. 24), most of which are necessary to create new kernel
functions and to make 1D arrays out of multi-dimensional arrays. By contrast,
OpenACC accomplishes the same parallelization with just small code changes (“b”),
because OpenACC’s directive-based approach encourages reuse of an existing code.
Besides, to parallelize the code for a distributed memory system, MPI also requires
large changes (“c” and “d”), primarily to convert global indices into local indices.
Serial
CUDA
OpenACC
MPI+CUDA
MPI+OpenACC
XcalableACC
a
b
c
d
e
Fig. 24 Application development order on accelerated cluster
Précédent

- 150/265

Suivant