90
M. Nakao and H. Murai
array may be changed but less memory may be used. Thus, we use “1” in the
XMP implementation. In Fig. 14, the second dimensions of arrays a() and b() are
distributed in the block manner, and array b() is transposed to array a(). For example,
elements b(1:3,2) on p(2) are transferred to elements a(2,1:3) on p(1).
5.4.2 Implementation
Figure 15 shows a part of the XMP implementation. In lines 1–9, arrays a(), b(),
and w() are distributed in a block manner. The a() is aligned with template ty, and
the b() and w() arrays are aligned with template tx. In lines 16–20, each thread on
all nodes calls the FFTE subroutine zfft1d(), which applies the distributed array b().
Note that the subroutine zfft1d() executes with a single thread locally. In lines 22–
23, the XMP loop directive and the OpenMP parallel directive parallelize the loop
statement. In line 30, xmp_transpose() is used to transpose the distributed twodimensional array.
1 complex*16 a(nx,ny), b(ny,nx), w(ny,nx)
2 !$xmp template ty(ny)
3 !$xmp template tx(nx)
4 !$xmp nodes p(*)
5 !$xmp distribute ty(block) onto p
6 !$xmp distribute tx(block) onto p
7 !$xmp align a(*,i) with ty(i)
8 !$xmp align b(*,i) with tx(i)
9 !$xmp align w(*,i) with tx(i)
10
11 integer, save :: ithread
12 !$omp threadprivate (ithread)
13 !$omp parallel
14
ithread = omp_get_thread_num()
15 !$omp end parallel
16 !$xmp loop on tx(i)
17 !$omp parallel do
18 do i=1,nx
19
call zfft1d(b(1,i),ny,−1,cy(1,ithread))
20 end do
21
22 !$xmp loop on tx(i)
23 !$omp parallel do
24 do i=1,nx
25
do j=1,ny
26
b(j,i)=b(j,i)*w(j,i)
27
end do
28 end do
29
30 call xmp_transpose(a,b,1)
Fig. 15 Part of the FFT code [5]
M. Nakao and H. Murai
array may be changed but less memory may be used. Thus, we use “1” in the
XMP implementation. In Fig. 14, the second dimensions of arrays a() and b() are
distributed in the block manner, and array b() is transposed to array a(). For example,
elements b(1:3,2) on p(2) are transferred to elements a(2,1:3) on p(1).
5.4.2 Implementation
Figure 15 shows a part of the XMP implementation. In lines 1–9, arrays a(), b(),
and w() are distributed in a block manner. The a() is aligned with template ty, and
the b() and w() arrays are aligned with template tx. In lines 16–20, each thread on
all nodes calls the FFTE subroutine zfft1d(), which applies the distributed array b().
Note that the subroutine zfft1d() executes with a single thread locally. In lines 22–
23, the XMP loop directive and the OpenMP parallel directive parallelize the loop
statement. In line 30, xmp_transpose() is used to transpose the distributed twodimensional array.
1 complex*16 a(nx,ny), b(ny,nx), w(ny,nx)
2 !$xmp template ty(ny)
3 !$xmp template tx(nx)
4 !$xmp nodes p(*)
5 !$xmp distribute ty(block) onto p
6 !$xmp distribute tx(block) onto p
7 !$xmp align a(*,i) with ty(i)
8 !$xmp align b(*,i) with tx(i)
9 !$xmp align w(*,i) with tx(i)
10
11 integer, save :: ithread
12 !$omp threadprivate (ithread)
13 !$omp parallel
14
ithread = omp_get_thread_num()
15 !$omp end parallel
16 !$xmp loop on tx(i)
17 !$omp parallel do
18 do i=1,nx
19
call zfft1d(b(1,i),ny,−1,cy(1,ithread))
20 end do
21
22 !$xmp loop on tx(i)
23 !$omp parallel do
24 do i=1,nx
25
do j=1,ny
26
b(j,i)=b(j,i)*w(j,i)
27
end do
28 end do
29
30 call xmp_transpose(a,b,1)
Fig. 15 Part of the FFT code [5]
