Three-Dimensional Fluid Code with XcalableMP
177
9 ...
10 SYNC ALL
11 physval(:,:,:,lsz+1)= physval1(:,:,:,1)[linzp]
12 physval(:,:,:,0)= physval1(:,:,:,lsz)[linzm]
13 ...
3.2 Performance on the K Computer
We run three XMP codes using the global-view programming model, and the
local-view programming model with put and get communications, and localview communications are implemented on Fujitsu RDMA. Each process performs
computations with 8 threads just like before and we evaluate the weak scaling on
the K computer using Omni XcalableMP 0.9.1 and Fujitsu Fortran K-1.2.0.18.
A number of cores for execution and corresponding simulation parameters are
summarized in Table 2 and performance are also measured by the hardware monitor
installed on the K computer. As versions of both Omni XcalableMP and Fujitsu
Fortran compilers are different from those of previous section, we also rerun the
global-view programming model code. MFLOPS/PEAK values for all cases are
shown in Fig. 5.
Performances using the global-view programming model are almost same as
those in previous section, but the local-view programming model shows very low
performances, namely 3% of peak performance of the K computer. From the
hardware monitor, we found that SIMD execution usage was less than 0.2% in
local-view programming model cases, this means that cost intensive DO loops
in IMPACT-3D are not SIMDized at all even with simd=2 and prefetch_stride
native Fortran compiler options. All Fortran allocatable coarrays in the local-view
programming model codes are converted to pointer arrays by the XMP/F compiler.
The pointer array prevents the native Fortran compiler from SIMDizing the DO
loop even it is forced to SIMDize the loop by simd=2 compiler option because
the compiler thinks that variables may be overlapped and SIMD execution causes
incorrect calculations. To tell the compiler that variables are not overlapped, we can
specify noalias option and SIMD execution usage is improved to 15%. But prefetch
instructions are still suppressed and the pointer array may prevent other compiler
optimizations, performances are not improved at all.
Table 2 Simulation
parameters for local-view
programming model
All of X, Y and Z
#Core lx=ly=lz nx ny nz
256
1024
4
4
2
2048
2048
8
8
4
16,384 4096
16 16
8
177
9 ...
10 SYNC ALL
11 physval(:,:,:,lsz+1)= physval1(:,:,:,1)[linzp]
12 physval(:,:,:,0)= physval1(:,:,:,lsz)[linzm]
13 ...
3.2 Performance on the K Computer
We run three XMP codes using the global-view programming model, and the
local-view programming model with put and get communications, and localview communications are implemented on Fujitsu RDMA. Each process performs
computations with 8 threads just like before and we evaluate the weak scaling on
the K computer using Omni XcalableMP 0.9.1 and Fujitsu Fortran K-1.2.0.18.
A number of cores for execution and corresponding simulation parameters are
summarized in Table 2 and performance are also measured by the hardware monitor
installed on the K computer. As versions of both Omni XcalableMP and Fujitsu
Fortran compilers are different from those of previous section, we also rerun the
global-view programming model code. MFLOPS/PEAK values for all cases are
shown in Fig. 5.
Performances using the global-view programming model are almost same as
those in previous section, but the local-view programming model shows very low
performances, namely 3% of peak performance of the K computer. From the
hardware monitor, we found that SIMD execution usage was less than 0.2% in
local-view programming model cases, this means that cost intensive DO loops
in IMPACT-3D are not SIMDized at all even with simd=2 and prefetch_stride
native Fortran compiler options. All Fortran allocatable coarrays in the local-view
programming model codes are converted to pointer arrays by the XMP/F compiler.
The pointer array prevents the native Fortran compiler from SIMDizing the DO
loop even it is forced to SIMDize the loop by simd=2 compiler option because
the compiler thinks that variables may be overlapped and SIMD execution causes
incorrect calculations. To tell the compiler that variables are not overlapped, we can
specify noalias option and SIMD execution usage is improved to 15%. But prefetch
instructions are still suppressed and the pointer array may prevent other compiler
optimizations, performances are not improved at all.
Table 2 Simulation
parameters for local-view
programming model
All of X, Y and Z
#Core lx=ly=lz nx ny nz
256
1024
4
4
2
2048
2048
8
8
4
16,384 4096
16 16
8
