Three-Dimensional Fluid Code with XcalableMP
171
45 !$OMP PARALLEL DO REDUCTION(max:ram) PRIVATE(iy,ix)
46 do iz = 1, lz
47 do iy = 1, ly
48 do ix = 1, lx
49
ram = max( ram, ... )
2.2 Performance on the K Computer
As one node consists of 8 cores in the K computer, one MPI process is dispatched
onto each node and each process performs computations with 8 threads. We run
both XMP and MPI codes with three different decomposition methods and evaluate
the weak scaling on the K computer using Omni XcalableMP 0.7.0 and Fujitsu
Fortran K-1.2.0.15. A number of cores for execution and corresponding simulation
parameters are summarized in Table 1.
Performance are measured by a hardware monitor installed on the K computer,
and MFLOPS/PEAK, Memory throughput/PEAK and SIMD execution usage are
obtained.
2.2.1 Comparison with Hand-Coded MPI Program
MFLOPS/PEAK values for all six cases, namely (MPI, XMP) × (only Z, both Y
and Z, all of X, Y, and Z) are shown in Fig. 2.
Performances of XMP codes are the same as those of MPI codes, and small
differences among three decomposition methods are found. But we can get only
8~9% of peak performance of the K computer. From the hardware monitor, we
found that SIMD execution usage was less than 5% in all cases, and this could
degrade the performance. Most cost intensive DO loops in IMPACT-3D include
IF statements, which are needed to correctly treat extremely slow fluid velocity
regardless of XMP and MPI codes, and the IF statement interrupts the native Fortran
compiler to generate SIMD instructions inside the DO loop. Thus relatively low
performance is obtained.
Table 1 Simulation parameters for global-view programming model
Only Z
Both Y and Z
All of X, Y and Z
#Core
lx=ly=lz
nz
ny
nz
nx
ny
nz
256
1024
32
8
4
4
4
2
2048
2048
256
16
16
8
8
4
16,384
4096
64
32
16
16
8
131,072
8192
128
128
32
32
16
Précédent

- 177/265

Suivant