XcalableACC: An Integration of XcalableMP and OpenACC
135
3 Omni XcalableACC Compiler
We have developed the Omni XACC compiler as the reference implementation of
XACC compilers.
Figure 13 shows the compile flow of Omni XACC. First, Omni XACC accepts
XACC source codes and translates them into those in the base languages with
runtime calls. Next, the translated code is compiled by a native compiler, which
supports OpenACC, to generate an object file. Finally, the object files and the
runtime library are linked by the native compiler to generate an execution file.
In particular, for the data transfer between NVIDIA GPUs across nodes, we have
implemented the following three methods in Omni XACC:
(a) TCA/IB hybrid communication
(b) GPUDirect RDMA with CUDA-Aware MPI
(c) MPI and CUDA
Item (a) performs communication with the smallest latency, but it requires a
computing environment equipped with the Tightly Coupled Accelerator (TCA)
feature[1, 7]. Item (b) is superior in performance to Item (c), but also requires
specific software and hardware (e.g., MVAPICH2-GDR and Mellanox InfiniBand).
Whereas (a) and (b) can realize direct communication between GPUs without the
intervention of CPU, Item (c) cannot. It copies the data from accelerator memory to
host memory using CUDA and then transfers the data to other compute nodes using
MPI. Therefore, although its performacnce is the lowest, it requires neither specific
software nor hardware.
Frontend
Omni XACC Compiler
Base language (C or Fortran)
+ OpenACC directive
+ XcalableMP directive
+ XcalableACC directive
Runtime library
Translator
Backend
User code
Translated code
Execution binary
Modified base language
+ Modified OpenACC directive
+ Runtime call
Fig. 13 Compile flow of Omni XcalableACC compiler
Précédent

- 142/265

Suivant