Implementation and Performance Evaluation of Omni Compiler
89
4096 compute nodes on the K computer, and 47.32 TFlops (70.02% of the peak
performance) for 128 compute nodes on the COMA system. The values of the
performance ratio are between 0.95 and 1.09 on the K computer, and between 0.99
and 1.06 on the COMA system.
5.4 Global Fast Fourier Transform
5.4.1 Design
FFT evaluates the performance for a double-precision complex one-dimensional
discrete Fourier transform. We implement a six-step FFT algorithm [7, 8] using
FFTE library [9]. The six-step FFT algorithm is also used in the MPI implementation. In the six-step FFT algorithm, both the computing performance and
the all-to-all communication performance for a matrix transpose are important.
The six-step FFT algorithm reduces the cache-miss ratio by expression of a twodimensional array. In order to develop the XMP implementation, we use XMP in
Fortran because FFTE library is written in Fortran and therefore it is easy to call it.
In addition, we use the XMP intrinsic subroutine xmp_transpose() to transpose a
distributed array in the global-view memory model. Figure 14 shows an example of
xmp_transpose(). The first argument is an output array, and the second argument
is the input array. The third argument is an option to save memory, and is “0”
or “1.” If it is “0,” an input array must not be changed. If it is “1,” an input
1 complex*16 a(4,12), b(12,4)
2 !$xmp template ty(12)
3 !$xmp template tx(4)
4 !$xmp nodes p(4)
5 !$xmp distribute ty(block) onto p
6 !$xmp distribute tx(block) onto p
7 !$xmp align a(*,i) with ty(i)
8 !$xmp align b(*,i) with tx(i)
9 call xmp_transpose(a, b, 1)
p(1)
p(3)
p(2)
p(4)
a(4,12)
b(12,4)
Fig. 14 Action of subroutine xmp_transpose() [5]
Précédent

- 96/265

Suivant