Three-Dimensional Fluid Code with XcalableMP
173
Fig. 3 MFLOPS/PEAK measured by the hardware monitor installed on the K computer with
SIMD optimization by simd=2 compiler option. Solid and dash lines indicate performances of
XMP and MPI codes, respectively. Colors of light gray, gray, and black indicate only Z domain
decomposition, both Y and Z domain decomposition, and all of X, Y, and Z domain decomposition
methods, respectively
of MPI codes are 55%, those of XMP codes are only 37% and this low memory
throughput is one of candidates for low sustained performance.
2.2.3 Optimization for Allocatable Arrays
In the converted code by the XMP/F compiler, all Fortran arrays are treated as
allocatable arrays even the original code uses static arrays. The allocatable array
prevents the native Fortran compiler from optimizing the DO loop with prefetch
instructions because the array size cannot be determined at compilation time, and
it could cause low memory throughput. All Fortran arrays in the hand-coded MPI
code for XYZ decomposition are just replaced by allocatable arrays and we check a
performance difference. Performance of the MPI code are shown in Fig. 4 for static
arrays (light gray dash) and allocatable arrays (gray dash).
MFLOPS/PEAK values are dropped from 20% to 15%, and this performance
degradation without the prefetch instructions is confirmed. To force the native
Fortran compiler to perform the prefetch optimization, we can use prefetch_stride
compiler option. All codes are recompiled with prefetch_stride compiler option and
rerun. Performance improvements by this compiler option are shown in Fig. 4 for
both MPI (gray dash to black dash) and XMP (gray solid to black solid) codes.
MFLOPS/PEAK values are improved by 2~3% with the prefetch optimization.
Finally we can get almost the same performance with XMP as that of MPI when
allocatable arrays are used, but efforts to shrink the performance gap between static
and allocatable arrays are still needed.
Précédent

- 179/265

Suivant