XcalableACC: An Integration of XcalableMP and OpenACC
139
1 double norm(const Quark_t v[NT][NZ][NY][NX])
2 {
3 #pragma xmp align v[i][j][*][*] with t[i][j]
4 #pragma xmp shadow v[1:1][1:1][0][0]
5 double a = 0.0;
6
7 #pragma xmp loop (t,z) on t[t][z]
8 #pragma acc parallel loop collapse(7) present(v) reduction(+:a)
9 for(int t=0;t 10
for(int z=0;z 11
for(int y=0;y 12
for(int x=0;x 13
for(int i=0;i<4;i++)
14
for(int j=0;j<3;j++)
15
for(int k=0;k<2;k++)
16
a += v[t][z][y][x].v[i][j][k]*v[t][z][y][x].v[i][j][k];
17
18 #pragma xmp reduction (+:a)
19 return a;
20 }
Fig. 18 L2 norm calculation code
5 Performance Evaluation
5.1 Result
This section evaluates the performance level of XACC on the Lattice QCD code. For
comparison purposes, those of MPI+CUDA and MPI+OpenACC are also evaluated.
For performance evaluation, we use the HA-PACS/TCA system[7], the hardware
specifications and software environments of which are shown in Table 1. Since each
compute node has four GPUs, we assign four nodes per compute node and direct
each node to deal with a single GPU. We use the Omni OpenACC compiler[8] as a
backend compiler in the Omni XACC compiler. We execute the Lattice QCD codes
with strong scaling in regions (32,32,32,32) as (NT,NZ,NY,NX). The Omni XACC
compiler provides various types of data communication among accelerators[9]. We
Table 1 Evaluation environment
CPU
Intel Xeon-E5 2680v2 2.8 GHz × 2 Sockets
Memory
DDR3 1866 MHz 59.7 GB/s 128 GB
GPU
NVIDIA Tesla K20X (GDDR5 250 GB/s 6 GB) × 4 GPUs
Network
InfiniBand Mellanox Connect-X3 Dual-port QDR 8 GB/s
Software
Intel 16.0.2, CUDA 7.5.18, Omni OpenACC compiler 1.1
MVAPICH2 2.1
Précédent

- 146/265

Suivant