178
H. Sakagami
Fig. 5 MFLOPS/PEAK measured by the hardware monitor installed on the K computer. Gray
dash line indicates performance of the global-view programming model. Gray and black solid lines
indicate performances of the local-view programming model with put and get communications,
respectively
4 Summary
We have parallelized a three-dimensional fluid code with XMP Fortran using the
global-view programming model and compared XMP performances with those of
the hand-coded MPI program on the K computer. We found that performances of
XMP programs are the same as those of MPI programs but these are only 8~9% of
peak performance of the K computer. It was found that this relative low performance
is due to lack of SIMD execution according to SIMD execution usage by the
hardware monitor. We forced the native Fortran compiler to SIMDize loops with
the specific compiler option, and found that performance of MPI programs reach
to 20% of peak performance even those of XMP programs remain around 15%.
It was found that this relative low performance is due to low memory throughput
according to Memory throughput/PEAK by the hardware monitor. Finally we could
get almost the same performance of XMP codes as those of MPI codes by using
additional specific compiler option of the native Fortran compiler.
Next we parallelized the code using the local-view programming model, and also
measured its performance on the K computer. We found that translated programs
prevent SIMDization by the native Fortran compiler and show only 3% of peak
performance of the K computer, much lower performance than that of the globalview programming model programs. This degradation cannot be solved by simply
specifying native compiler options at this moment, and improvements of the XMP/F
compiler are expected.
These kinds of advanced performance optimization techniques of the native
Fortran compiler are not clear and may be somewhat difficult for computational
Précédent

- 184/265

Suivant