172
H. Sakagami
Fig. 2 MFLOPS/PEAK measured by the hardware monitor installed on the K computer. Solid and
dash lines indicate performances of XMP and MPI codes, respectively. Colors of light gray, gray,
and black indicate only Z domain decomposition, both Y and Z domain decomposition, and all of
X, Y, and Z domain decomposition methods, respectively
2.2.2 Optimization for SIMD
As the true rate of the IF statement is nearly 100% in IMPACT-3D, speculative
execution of SIMD instruction causes almost no overhead. So forcing the compiler
to generate the SIMD instructions could be useful to enhance the performance, and
it can be done with simd=2 compiler option. All codes are recompiled with that
option and rerun. SIMD execution usage increases up to around 50% in all cases,
and we can expect performance improvement. MFLOPS/PEAK values for all cases
are shown in Fig. 3.
Small differences among three decomposition methods are also found with this
compiler option. MPI code performance is improved and we can get up to 20%
of the peak performance. XMP code performance is also improved, but these
are below 15% even MPI and XMP code performance is almost same without
simd=2 option. According to compiler diagnostic of the native Fortran compiler,
the software pipelining is adopted for cost intensive DO loops in the MPI code, but
it is not applied for the source code converted by the XMP/F compiler from the
XMP code. As the XMP/F compiler converts a simple DO statement of “do i = is,
ie” to more general form “do i1 = xmp_s1, xmp_e1, xmp_d1” and the native Fortran
compiler cannot optimize the DO loop because do increment is given by a variable
and it is unknown at compilation time. So we improved the XMP/F compiler to
generate “do i1 = xmp_s1, xmp_e1, 1” form when the do increment is not given
and supposed to be one in the XMP code. As a result, the software pipelining is
also adopted for cost intensive DO loops converted by the XMP/F compiler, but no
performance improvement is obtained. Although Memory throughput/PEAK values
H. Sakagami
Fig. 2 MFLOPS/PEAK measured by the hardware monitor installed on the K computer. Solid and
dash lines indicate performances of XMP and MPI codes, respectively. Colors of light gray, gray,
and black indicate only Z domain decomposition, both Y and Z domain decomposition, and all of
X, Y, and Z domain decomposition methods, respectively
2.2.2 Optimization for SIMD
As the true rate of the IF statement is nearly 100% in IMPACT-3D, speculative
execution of SIMD instruction causes almost no overhead. So forcing the compiler
to generate the SIMD instructions could be useful to enhance the performance, and
it can be done with simd=2 compiler option. All codes are recompiled with that
option and rerun. SIMD execution usage increases up to around 50% in all cases,
and we can expect performance improvement. MFLOPS/PEAK values for all cases
are shown in Fig. 3.
Small differences among three decomposition methods are also found with this
compiler option. MPI code performance is improved and we can get up to 20%
of the peak performance. XMP code performance is also improved, but these
are below 15% even MPI and XMP code performance is almost same without
simd=2 option. According to compiler diagnostic of the native Fortran compiler,
the software pipelining is adopted for cost intensive DO loops in the MPI code, but
it is not applied for the source code converted by the XMP/F compiler from the
XMP code. As the XMP/F compiler converts a simple DO statement of “do i = is,
ie” to more general form “do i1 = xmp_s1, xmp_e1, xmp_d1” and the native Fortran
compiler cannot optimize the DO loop because do increment is given by a variable
and it is unknown at compilation time. So we improved the XMP/F compiler to
generate “do i1 = xmp_s1, xmp_e1, 1” form when the do increment is not given
and supposed to be one in the XMP code. As a result, the software pipelining is
also adopted for cost intensive DO loops converted by the XMP/F compiler, but no
performance improvement is obtained. Although Memory throughput/PEAK values
